System

A system using a camera and microphone to analyze pet behavior and sounds with AI provides real-time emotional and health assessments, addressing the inaccuracy of current tools and enabling timely care for pets.

JP2026025719APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128531
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Current communication tools for pets, such as birds, rabbits, and reptiles, are inaccurate and fail to detect subtle changes in behavior or vocalizations, making it difficult to identify signs of illness or injury, and owners struggle to provide timely care.

Method used

A system comprising a camera and microphone to record pet behavior and sounds, preprocessing the data, transmitting it to a cloud server for analysis using object and voice recognition algorithms, and employing a generative AI model to determine emotions and health conditions, with feedback to users.

Benefits of technology

Enables real-time monitoring and accurate assessment of pet emotions and health, allowing for early detection of abnormalities and appropriate care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025719000001_ABST
    Figure 2026025719000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system for determining an emotion or health condition of a pet, comprising: a camera module configured to record a behavior of the pet; a microphone module configured to record a cry of the pet; a transmission module configured to transmit the pre-processed information to a cloud server; an analysis module configured to analyze the information received by the cloud server; a generative AI model configured to determine the emotion or health condition of the pet based on an analysis result; and a feedback module configured to notify a user of a determination result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] As the pet market expands, owners are increasingly required to accurately understand their pets' health conditions and emotions. However, current communication tools are inaccurate and primarily limited to dogs and cats, making them inapplicable to other pets, such as birds, rabbits, and reptiles. Furthermore, it is difficult to detect subtle changes in a pet's behavior or vocalizations early on and discover signs of illness or injury. The purpose of the present invention is to solve these problems and provide a comfortable life for both pets and their owners. [Means for solving the problem]

[0005] The present invention provides a system including a camera means for recording a pet's behavior, a microphone means for recording the pet's cries, a terminal means for preprocessing the recorded data, a transmission means for transmitting the preprocessed data to a cloud server, an analysis means for analyzing the data received by the cloud server, a generative AI model for determining the pet's emotions and health condition based on the analysis results, and a feedback means for notifying the user of the determination results. This makes it possible to monitor subtle changes in a pet's behavior and cries in real time and accurately grasp the pet's health condition and emotions based on the analysis results. Furthermore, by promptly notifying the user if an abnormality is detected, early action can be taken and appropriate care of the pet can be achieved.

[0006] The "camera means" is a device for recording the behavior of a pet in real time.

[0007] The "microphone means" is a device that records the sounds of pets and saves them as audio data.

[0008] "Terminal means" refers to a device for preprocessing data collected by the camera means and microphone means.

[0009] The "transmitting means" is a communication device that has a function for transmitting the preprocessed data to the cloud server.

[0010] A "cloud server" is a remote server for receiving, analyzing, and storing data.

[0011] The "analysis means" is an algorithm that analyzes the data received by the cloud server and determines the behavior and cries of the pet.

[0012] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on data obtained through analytical means.

[0013] "Feedback means" is a function for notifying users of the results of the generative AI model's judgment.

[0014] An "object recognition algorithm" is an analytical method for identifying specific objects and their movements from video data.

[0015] A "voice recognition algorithm" is an analytical method for identifying specific sounds and their characteristics from audio data.

[0016] A "user terminal" is a device that receives judgment results and notifications and displays them to the user. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[0039] Specific operations of program processing

[0040] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[0041] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[0042] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[0043] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[0044] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[0045] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[0046] Specific examples

[0047] 1. Data collection

[0048] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[0049] 2. Data transmission and preprocessing

[0050] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[0051] 3. Data Analysis

[0052] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0053] 4. Emotional and health status assessment

[0054] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0055] 5. Feedback and Notifications

[0056] The cloud server sends the results to the user's device, and a dedicated application notifies them, saying, "Your cat is stressed. Attention is needed."

[0057] The introduction of this system will enable owners to monitor their pets' health and emotions in real time and provide appropriate care at an early stage, marking a major step towards providing a comfortable and healthy life for both pets and their owners.

[0058] The processing flow will be explained below.

[0059] Step 1:

[0060] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[0061] Step 2:

[0062] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[0063] Step 3:

[0064] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[0065] Step 4:

[0066] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[0067] Step 5:

[0068] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[0069] Step 6:

[0070] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[0071] Step 7:

[0072] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[0073] Step 8:

[0074] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[0075] Step 9:

[0076] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's mood and health status and take appropriate action if necessary.

[0077] Example 1

[0078] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0079] A pet's behavior and vocalizations are important indicators for understanding its emotions and health status, but it is difficult for owners to constantly monitor them and respond appropriately. Furthermore, if owners are unable to notice abnormalities in their pets early, appropriate care may be delayed, increasing the risk of overlooking pet health problems. Given this situation, there is a demand for a system that can automatically monitor and analyze pet behavior and vocalizations in real time.

[0080] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0081] In this invention, the server includes a recording means for recording the behavior of the pet, an audio acquisition means for recording the sounds of the pet, a processing means for preprocessing the recorded data, a transmission means for transmitting the preprocessed data via a network, an analysis means for analyzing the data received via the network, a generation algorithm model for determining the emotions and health condition of the pet based on the analysis results, and a notification means for notifying the user of the determination results. This makes it possible to automatically monitor and analyze the behavior and sounds of the pet in real time, accurately grasp the emotions and health condition of the pet, and provide necessary care early.

[0082] "Recording means" refers to a device or function for recording the behavior of a pet.

[0083] "Audio acquisition means" refers to a device or function for recording the sounds of pets.

[0084] "Processing means" refers to devices or functions for pre-processing the recorded data.

[0085] "Transmitting means" refers to a device or function for transmitting preprocessed data over a network.

[0086] "Analysis means" refers to a device or function for analyzing data received via a network.

[0087] A "generative algorithm model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[0088] "Notification means" refers to a device or function for notifying the user of the judgment result.

[0089] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[0090] First, the user installs a pet-specific device in the pet's living space. This device has a built-in camera and microphone, which allows it to record the pet's behavior and sounds in real time. For example, the user can install the device in the living room and adjust the camera's orientation to capture the area where the pet is often found.

[0091] Next, the device captures footage of the pet's behavior with a camera and simultaneously records its meows with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. For example, the device may capture a video of a cat playing by a window and record its meows. At this time, unnecessary parts of the video are deleted and background noise is removed from the audio.

[0092] The preprocessed data is then sent by the device to a cloud server. This transmission is performed using wireless communication such as Wi-Fi or 4G / 5G. To send the preprocessed data to the cloud server, the device connects to the home Wi-Fi network, encrypts the data, and transmits it. After a few seconds, the cloud server receives the data.

[0093] The cloud server applies an object recognition algorithm to the video data it receives to identify the pet's body parts and analyze its behavioral patterns. It also applies a voice recognition algorithm to the audio data to analyze the characteristics of its meows. For example, the cloud server analyzes the video data to identify the behavior of a cat licking its hind paws, and analyzes the audio data to detect that short, sharp meows are a sign of stress.

[0094] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, the generative AI model compares a cat's past data with its current meows and behavioral patterns and determines that it is likely to be stressed.

[0095] The results of the assessment are sent to the user's device via a feedback mechanism. The user's device may be, for example, a smartphone or tablet, and the results are displayed through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, the cloud server may send a notification to the user's device stating, "Your cat is feeling stressed," and a pop-up notification will appear on the user's smartphone. When the user opens the app, detailed analysis results will be displayed.

[0096] For example, when a device records a cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone picks up short, sharp meows. This data is preprocessed and sent to a cloud server where it is analyzed using object recognition and voice recognition algorithms. The generative AI model then determines that the cat is stressed and sends a notification to the user's device saying, "Your cat is stressed. Attention is needed." In this way, users can monitor their pet's health in real time and provide appropriate care early.

[0097] Example prompts for generative AI models

[0098] Below is an example of a prompt sentence to input to the generative AI model.

[0099] "Your cat is frequently licking its hind paws and making short, sharp meowing noises. Use these data to assess your cat's emotional state and health."

[0100] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0101] Step 1: The user places the pet-specific device in the pet's living space. This device has a built-in camera and microphone. The input is the device's installation location and the pet's range of movement. The output is the device's readiness to record the pet's movements and sounds. The user places the device in a location in the living room where the pet is often found and turns it on. The user adjusts the camera's orientation to capture the pet at the optimal angle.

[0102] Step 2: The device captures the pet's behavior with a camera and simultaneously records its cries with a microphone. The input is the pet's behavior and cries. The output is the recorded video and audio data. For example, the device can record a video of the pet playing and simultaneously record the cries it makes.

[0103] Step 3: The device preprocesses the recorded data. This preprocessing includes trimming unnecessary parts of the video data and removing noise. The input is the recorded video data and audio data. The output is the preprocessed video data and audio data. The device removes unnecessary parts of the video data and removes background noise from the audio data.

[0104] Step 4: The device sends the preprocessed data to the cloud server. This is done using wireless communication such as Wi-Fi or 4G / 5G. The input is the preprocessed video and audio data. The output is the data sent to the cloud server. For example, the device sends the data via a home Wi-Fi network, and the cloud server receives it.

[0105] Step 5: The cloud server analyzes the received video data. It applies an object recognition algorithm to identify the pet's body parts and analyzes its behavioral patterns. The input is the video data sent to the cloud server. The output is the analysis results (behavioral patterns). The cloud server analyzes the video data and identifies, for example, the pet licking its hind paws.

[0106] Step 6: The cloud server analyzes the audio data. It applies a speech recognition algorithm to analyze the characteristics of the bird's calls. The input is the audio data sent to the cloud server. The output is the analysis result (the characteristics of the bird's calls). The cloud server analyzes the audio data and detects, for example, that the bird's calls are a sign of stress.

[0107] Step 7: The cloud server inputs the analysis results into the generative AI model. This generative AI model compares it with past data to determine the pet's emotions and health condition. The input is the analysis results of the video and audio data. The output is the determination of the pet's emotions and health condition. The generative AI model compares it with past data and determines that the pet is likely to be feeling stressed.

[0108] Step 8: The cloud server sends the judgment result to the user device via a feedback means. The user device displays the result through a dedicated application. The input is the judgment result of emotions and health status. The output is a notification displayed on the user device. The cloud server sends a notification to the user device saying "The cat is feeling stressed," and the user checks the notification on their smartphone.

[0109] This allows users to understand their pet's condition in real time and provide appropriate care early on.

[0110] (Application example 1)

[0111] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0112] Conventional pet management systems simply record pet behavior and sounds, and are unable to assess a pet's emotions or health status in real time. Additionally, there is a lack of information that allows owners to select appropriate care products and food based on their pet's health status. Furthermore, specific care suggestions based on analysis results are rarely provided, making it difficult for owners to provide detailed care for their pets.

[0113] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0114] In this invention, the server includes a video recording means for recording the behavior of the pet, an audio recording means for recording the sounds of the pet, an information processing means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the information server, an analysis means for analyzing the data received by the information server, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a suggestion means for suggesting appropriate care products and food based on the judgment results, and a notification means for notifying the user of the suggested content. This makes it possible to analyze the emotions and health condition of the pet and notify the user of specific care suggestions and products based on the pet's health condition.

[0115] "Video recording means" is a function for recording the behavior of pets using a camera.

[0116] The "audio recording means" is a function for recording the sounds of pets using a microphone.

[0117] The "information processing means" is a computer system for preprocessing the recorded video data and audio data.

[0118] The "communication means" is a wireless communication function for transmitting preprocessed data to an information server such as a cloud server.

[0119] The "analysis means" refers to a group of algorithms and programs for analyzing data received by the information server.

[0120] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[0121] The "suggestion means" is a function for suggesting care products and food suitable for pets based on the judgment results.

[0122] "Notification means" refers to a notification system for informing users of the proposed content.

[0123] An "object recognition algorithm" is an algorithm that analyzes pet behavior data and recognizes specific actions and movements.

[0124] The "voice recognition algorithm" is an algorithm for analyzing pet cry data and recognizing specific voice patterns.

[0125] The present invention relates to a system that records a pet's behavior and cries in real time, analyzes this data to determine the pet's emotions and health condition, and then suggests appropriate care products and food based on the results of this determination.

[0126] The system consists of the following main tools:

[0127] 1. Video recording means: Recording pet behavior using a camera. For example, a camera device specifically for pets is included.

[0128] 2. Audio recording means: Record the sounds your pet makes using a microphone. This also includes a dedicated audio recording device for pets.

[0129] 3. Information processing means: A computer system for preprocessing recorded video and audio data, such as noise removal and data trimming.

[0130] 4. Communication means: Equipped with wireless communication functions for transmitting preprocessed data to the information server. Specific examples include Wi-Fi and 4G / 5G communication modules.

[0131] 5. Analysis means: A group of algorithms and programs for analyzing data received by the information server. This includes object recognition algorithms and voice recognition algorithms.

[0132] 6. Generative AI model: An artificial intelligence model that determines a pet's emotions and health status based on the analysis results.

[0133] 7. Suggestion means: Based on the judgment results, it has a function to suggest care products and food suitable for pets.

[0134] 8. Notification method: A notification system to inform users of the proposed content. For example, a notification is sent to a smartphone.

[0135] The server performs the following main tasks:

[0136] 1. Analysis of video and audio data. Characterize pet behavior and sounds using object and audio recognition algorithms.

[0137] 2. The analysis results are input into a generative AI model to determine the pet's emotions and health status.

[0138] 3. Based on the judgment results, the appropriate care products and food for the pet are determined using a suggestion method.

[0139] 4. The proposal will be communicated to the user via notification means.

[0140] For example, if a pet frequently scratches its ears, the server can analyze the behavior and determine that this is a sign of an ear infection. Based on this determination, appropriate care products (e.g., ear cleaning wipes or anti-infection shampoo) can be suggested and notified to the user.

[0141] An example prompt is:

[0142] Prompt for the generative AI model:

[0143] You are an AI model that analyzes your pet's behavior and vocalizations to determine its emotions and health status. Analyze the following data to determine your pet's emotions and health status:

[0144] Action: {action}

[0145] Sound: {sound}

[0146] Based on the results, we will recommend the appropriate care products and food for your pet.

[0147] This system allows users to understand their pet's health condition and emotions in real time, enabling them to provide prompt and accurate physical and psychological care.

[0148] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0149] Step 1:

[0150] The terminal records the pet's behavior using a video recording means and records the pet's cries using an audio recording means. The input includes video data from the camera and audio data from the microphone. These data are temporarily stored for subsequent processing.

[0151] Step 2:

[0152] The device preprocesses the recorded video and audio data. Specific preprocessing operations include trimming the video data, removing noise, and adjusting the resolution. For audio data, noise removal and extraction of necessary frequency components are also performed. Clean video and audio data are generated as the output of the preprocessing.

[0153] Step 3:

[0154] The terminal transmits the preprocessed video and audio data to the information server using a communication means, with the input comprising the preprocessed data and the output being data packets received by the information server.

[0155] Step 4:

[0156] The server analyzes the received data using an analysis means. Specifically, an object recognition algorithm is applied to the video data to detect the specific behavior of the pet. A voice recognition algorithm is applied to the audio data to extract the characteristics and patterns of the pet's cries. The input includes the received data, and the output generates data on behavioral patterns and voice patterns, which are the analysis results.

[0157] Step 5:

[0158] The server inputs the analysis results into a generative AI model and performs calculations to determine the pet's emotions and health condition. The input includes the analysis result data, and the output generates a judgment result on the pet's emotions and health condition. For example, based on the analysis results, it may be determined that the pet is "stressed."

[0159] Step 6:

[0160] The server uses a suggestion tool to determine the appropriate pet care products and food based on the assessment results. The input includes the assessment results of the pet's emotions and health condition, and the output generates a list of suggested products. For example, foods and toys with a relaxing effect are suggested for a stressed pet.

[0161] Step 7:

[0162] The server notifies the user of the proposed content using a notification method. Specifically, it sends a notification to the user's smartphone or tablet. The input includes a list of proposed products, and the output is generated as notification information to be displayed on the user's device. The user can then receive this notification and check the details of the proposed products.

[0163] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0164] This invention combines a system that records pet behavior and sounds in real time and analyzes this data to determine the pet's emotions and health condition with an emotion engine that recognizes the user's emotions. This system allows for accurate understanding of the condition of both the pet and the owner, enabling more appropriate care and communication.

[0165] Specific operations of program processing

[0166] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[0167] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[0168] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[0169] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[0170] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[0171] Furthermore, the system incorporates an emotion engine that recognizes the user's emotions. The user inputs text and voice data into the emotion engine, which then analyzes the user's emotional state. It also uses a facial recognition algorithm to analyze the user's facial expressions and evaluate their emotional state.

[0172] The cloud server correlates the user's emotional state with the pet's state and determines the overall communication state. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress state to determine how the two are affecting each other.

[0173] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[0174] Specific examples

[0175] 1. Data collection

[0176] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[0177] 2. Data transmission and preprocessing

[0178] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[0179] 3. Data Analysis

[0180] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0181] 4. Emotional and health status assessment

[0182] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0183] 5. Recognizing user emotions

[0184] Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[0185] 6. Overall Judgment

[0186] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0187] 7. Feedback and Notifications

[0188] The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is feeling stressed. Attention is needed." Users can also check their own emotional state and take necessary action.

[0189] This system allows for a comprehensive understanding of the emotions and health status of both pets and users, enabling appropriate care and communication.

[0190] The processing flow will be explained below.

[0191] Step 1:

[0192] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[0193] Step 2:

[0194] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[0195] Step 3:

[0196] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[0197] Step 4:

[0198] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[0199] Step 5:

[0200] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[0201] Step 6:

[0202] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[0203] Step 7:

[0204] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[0205] Step 8:

[0206] The user inputs text or voice data into the emotion engine, which then analyzes the user's emotional state.

[0207] Step 9:

[0208] The user's facial expressions are captured by a camera, and the emotion engine uses a facial recognition algorithm to analyze the user's facial expression data, thereby assessing the user's emotional state.

[0209] Step 10:

[0210] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0211] Step 11:

[0212] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[0213] Step 12:

[0214] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's emotional and health status, as well as their own emotional state, and take necessary actions.

[0215] Example 2

[0216] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0217] Pet owners face the challenge of understanding their pets' emotions and health status from their behavior and vocalizations. There is a need for a system that can properly recognize the emotional state of not only pets but also their owners, and grasp the overall state of communication with their pets. This will enable them to effectively manage the health and happiness of both pets and their owners, and provide appropriate care and advice.

[0218] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0219] In this invention, the server includes a camera for recording the pet's behavior, an audio collection device for recording the pet's cries, a processing device for preprocessing the recorded data, a communication device for transmitting the preprocessed data to a cloud server, an analysis device for analyzing the data received by the cloud server, a generative AI model for determining the pet's emotions and health status based on the analysis results, a transmission device for notifying the user of the determination results, an emotion engine for recognizing the user's emotions, and a comprehensive determination engine for correlating the user's emotional state with the pet's emotional state and making a comprehensive determination. This makes it possible to accurately grasp the pet's emotions and health status in real time and also evaluate the owner's emotional state. This allows for a comprehensive understanding of the pet's communication status and provides appropriate care and advice.

[0220] "Photographing means" refers to a device used to record the behavior of a pet.

[0221] An "audio collection means" is a device used to record the sounds of a pet.

[0222] "Processing equipment" is equipment used to pre-process recorded data.

[0223] A "communication device" is a device used to transmit pre-processed data to a cloud server.

[0224] An "analysis device" is a device used to analyze data received by a cloud server.

[0225] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on analysis results.

[0226] A "communication device" is a device used to notify the user of the determination result.

[0227] An "emotion engine" is a device or software used to recognize a user's emotions.

[0228] The "comprehensive judgment engine" is a device or software used to correlate and comprehensively judge the emotional state of the user and the emotional state of the pet.

[0229] This invention combines a system that determines a pet's emotions and health condition by recording and analyzing the pet's behavior and cries in real time with an emotion engine that recognizes the user's emotions. The following describes an embodiment of this system.

[0230] The first thing users need to do is to install a dedicated pet device in their pet's living space. This device is equipped with a high-resolution camera (photography) and a highly sensitive microphone (audio collection), allowing it to record the pet's behavior and sounds in real time.

[0231] The device then captures images of the pet's behavior with a camera and records its cries with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. This process produces data that is easy to analyze.

[0232] The preprocessed data is sent from the device to the cloud server using Wi-Fi or 4G / 5G wireless communication (communication equipment), allowing the data to reach the cloud server safely and quickly.

[0233] To analyze the received data, the cloud server first applies an object recognition algorithm to the video data. This algorithm identifies the pet's body parts and analyzes its behavioral patterns. It also applies a voice recognition algorithm to the audio data, analyzing the characteristics of the pet's cries. This allows for a detailed analysis of your pet's behavior.

[0234] The cloud server then uses a generative AI model to determine the pet's emotions and health based on the analyzed data. The generative AI model compares the analysis results with past data to determine, for example, whether the pet is stressed or in pain.

[0235] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotions. The user can input text or voice data into the emotion engine, which then analyzes the user's emotional state. The system can also record the user's facial expressions using a camera and evaluate their emotional state using a facial recognition algorithm.

[0236] The cloud server correlates the pet's emotional and health status with the user's emotional status and uses a comprehensive assessment engine to determine the overall communication status. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress status to determine how the two are affecting each other.

[0237] The results are sent from the cloud server to the user's device and displayed via a dedicated application. Notifications are displayed, allowing users to check their pet's current condition and whether any abnormalities have been detected.

[0238] Specific examples

[0239] As a concrete example, consider the following scenario.

[0240] 1. The device records the cat's behavior: the camera captures the cat frequently licking its hind paws, and the microphone picks up the cat's short, sharp meows.

[0241] 2. The collected data is pre-processed on the device, trimmed, and denoised before being sent to the cloud server.

[0242] 3. The cloud server analyzes the video data using an object recognition algorithm to confirm the cat's licking of its hind paw, and analyzes the audio data using a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0243] 4. The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0244] 5. Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[0245] 6. The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0246] 7. The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is stressed. Attention is needed." The user can also check their own emotional state and take necessary action.

[0247] In this way, this system allows for a comprehensive understanding of the emotions and health status of both the pet and the user, enabling appropriate care and communication.

[0248] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0249] Step 1:

[0250] The user installs a pet-specific device in the pet's living space.

[0251] How it works: The user places a device containing a high-resolution camera and a sensitive microphone in a location where the pet is frequently present or resting.

[0252] Step 2:

[0253] The device captures your pet's behavior with a camera and records its cries with a microphone.

[0254] How it works: The camera captures pet movements in high resolution, and the microphone records pet sounds with high sensitivity, providing real-time data on pet behavior and sounds.

[0255] Input: Pet actions and sounds

[0256] Output: Raw video and audio data

[0257] Step 3:

[0258] Preprocessing is performed on the data collected by the terminal.

[0259] How it works: The device trims unnecessary parts of the video data and removes noise from the audio data. This preprocessing produces data that is easy to analyze.

[0260] Input: Raw video and audio data

[0261] Output: Pre-processed video and audio data

[0262] Step 4:

[0263] The terminal transmits the preprocessed data to the cloud server.

[0264] How it works: The device uses Wi-Fi or 4G / 5G to send pre-processed data to the cloud server, where it is encrypted to ensure security.

[0265] Input: Preprocessed video and audio data

[0266] Output: Data sent to the cloud server

[0267] Step 5:

[0268] The cloud server analyzes the received data.

[0269] How it works: The cloud server first applies an object recognition algorithm to the video data to identify the pet's body parts and analyze its behavioral patterns, and then applies a voice recognition algorithm to the audio data to analyze the characteristics of the pet's cries.

[0270] Input: Video and audio data sent to the cloud server

[0271] Output: Analyzed behavioral and audio data

[0272] Step 6:

[0273] The cloud server uses the generated AI model to determine the pet's emotions and health condition based on the analysis results.

[0274] How it works: The generative AI model compares your pet's behavior and vocalizations with past data to determine their emotions and health status. For example, it analyzes the frequency of vocalizations and behavioral patterns to determine whether your pet is stressed or in pain.

[0275] Input: Analyzed behavioral and audio data

[0276] Output: Pet's emotions and health status

[0277] Step 7:

[0278] Users input text or voice into the emotion engine, which analyzes their emotional state.

[0279] How it works: Users input text or voice to the emotion engine to recognize their emotional state, and a facial recognition algorithm is used to assess the user's emotional state from facial expression data captured by the camera.

[0280] Input: User text, voice, and facial expression data

[0281] Output: Analysis of the user's emotional state

[0282] Step 8:

[0283] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0284] Specific operation: The cloud server uses a comprehensive judgment engine to correlate and analyze the emotional state of the pet and the user. For example, when the user is feeling stressed, the cloud server will also analyze the pet's stress state and evaluate the overall communication state.

[0285] Input: Pet's emotional and health status assessment results, analysis results of user's emotional state

[0286] Output: Overall communication status assessment

[0287] Step 9:

[0288] The cloud server sends the results to the user's device, and a dedicated application displays a notification.

[0289] Specific operation: The cloud server uses a transmission device to send the judgment result to the user's device. A dedicated application on the user's device displays a notification, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, a notification such as "Your cat is stressed. Attention is required" may be displayed.

[0290] Input: Overall communication status assessment result

[0291] Output: Notification displayed on the user's device

[0292] This series of steps creates a system that can comprehensively understand the emotions and health status of both the pet and the user, and provide appropriate care and communication.

[0293] (Application example 2)

[0294] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0295] In today's world, it is extremely important to properly understand pet health and emotional states and facilitate smooth communication with owners. However, many current systems simply record pet behavior and vocalizations, lacking the ability to comprehensively analyze pets' emotions and health. Furthermore, no systems provide appropriate feedback that takes into account the user's emotional state. Therefore, there is a need for a system that can analyze the state of both pets and users in real time to achieve optimal care and communication.

[0296] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a photographing means for recording the behavior of the pet, a sound collecting means for recording the sounds of the pet, a processing device means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the data processing device, an analysis device means for analyzing the data received by the data processing device, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a notification device means for notifying the user of the determination results, an emotion analysis means for analyzing the user's emotional state, and an association means for associating the pet's state with the user's emotional state. This makes it possible to analyze detailed data including the behavior and sounds of the pet and provide comprehensive advice that also takes the user's emotional state into consideration.

[0297] "Photographing means" refers to a camera device used to record the behavior of a pet.

[0298] The "audio collection means" is a microphone device for recording the sounds of pets.

[0299] The "processing device means" is a device for pre-processing data recorded by the imaging means and sound collection means.

[0300] "Communication means" refers to a wireless or wired communication device for transmitting pre-processed data to a data processing device.

[0301] "Analysis device means" is a general term for hardware and software for analyzing received data in a data processing device.

[0302] A "generative AI model" is a machine learning model that determines a pet's emotions and health status based on analysis results.

[0303] "Notification device means" is a device for notifying the user of the results determined by the generative AI model.

[0304] "Emotion analysis means" is a system for analyzing the emotional state of a user.

[0305] The "association means" is a mechanism for associating and analyzing the pet's emotions and health status with the user's emotional state.

[0306] This invention relates to a system that records pet behavior and sounds in real time and provides optimal feedback by taking into account the user's emotional state. The system includes three main components: a pet-specific device, a user terminal, and a cloud server.

[0307] Hardware and software used

[0308] Hardware:

[0309] Recording method: Camera equipment to record pet behavior

[0310] Sound collection means: microphone device for recording pet sounds

[0311] Processing Unit: A local device (e.g., Raspberry Pi) for preprocessing data.

[0312] Communication: Wi-Fi or 4G / 5G communication module for sending data to a cloud server

[0313] software:

[0314] TensorFlow: Deep learning object and speech recognition algorithms

[0315] OpenCV: Image processing library

[0316] Pydub: A library for preprocessing audio data

[0317] NLTK: A natural language processing library for analyzing user sentiment

[0318] System operation explanation

[0319] 1. Data Collection:

[0320] The pet device uses a camera to capture images of your pet's behavior and a microphone to record your pet's sounds, and this data is collected in real time.

[0321] 2. Data preprocessing:

[0322] Processing means within the terminal trims the collected video data and removes noise from the audio data.

[0323] 3. Data transmission:

[0324] The preprocessed data is sent to a cloud server via a communication method such as Wi-Fi or 4G / 5G.

[0325] 4. Data Analysis:

[0326] The cloud server analyzes the received data using an analysis device means. An object recognition algorithm using TensorFlow is applied to the video data, and a voice recognition algorithm also using TensorFlow is applied to the voice data.

[0327] 5. Emotional and health status assessment:

[0328] Based on the analysis results, the generative AI model determines the pet's emotions and health status. For example, if the pet repeatedly behaves in a certain way or makes a certain sound, stress or health problems may be suspected.

[0329] 6. User Emotion Recognition:

[0330] By inputting text or voice from the user's device, the user's emotional state is analyzed by the emotion analysis means. The emotion analysis is performed using the NLTK library.

[0331] 7. Overall Judgment and Association:

[0332] The system analyzes the pet's emotions and health status in relation to the user's emotional state, and generates optimal feedback based on the results.

[0333] 8. Feedback and Notifications:

[0334] The cloud server sends the results to the user's device, where they are displayed via a dedicated application, allowing the user to receive feedback and take appropriate action.

[0335] Specific examples

[0336] For example, the device records the cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone records the cat's sharp meows. This data is sent to a cloud server and analyzed using object recognition and voice recognition algorithms. The generative AI model determines that the cat is likely feeling stressed. The user enters text into the emotion engine, which analyzes the cat's stress. The cloud server correlates this data and makes a comprehensive judgment, and a notification such as "Your cat is feeling stressed. Attention is required" is displayed on the user's device.

[0337] Example prompts to input to the generative AI model:

[0338] Pet Behavior: Cat repeatedly licks the same spot

[0339] Call: High-pitched, sharp call

[0340] User Emotion: Stressed

[0341] Determine your pet's emotional and health status.

[0342] This system makes it possible to perform detailed analysis of a pet's behavior and cries, and provide optimal actions that take into account the user's emotional state.

[0343] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0344] Step 1:

[0345] The device uses a camera and microphone to record your pet's movements and sounds in real time. This input data is used to generate video and audio files, which are then temporarily stored on the device.

[0346] Step 2:

[0347] The processing unit of the terminal performs preprocessing on the recorded video data and audio data, such as trimming the video data and adjusting the resolution, and removing noise from the audio data. As a result of the preprocessing, preprocessed video and audio files are generated.

[0348] Step 3:

[0349] The communication means transmits the preprocessed data to the cloud server. Specifically, the data is uploaded using Wi-Fi or 4G / 5G. After transmission, the data is stored in the cloud server.

[0350] Step 4:

[0351] The cloud server analyzes the received video data and audio data using an analysis device. An object recognition algorithm using TensorFlow is applied to the video data to extract the pet's behavioral characteristics. A voice recognition algorithm also using TensorFlow is applied to the audio data to extract the characteristics of the pet's cries. As a result of the analysis, data on the pet's behavioral characteristics and audio characteristics is generated.

[0352] Step 5:

[0353] The cloud server uses a generative AI model based on the analysis results to determine the pet's emotions and health condition. The feature data from the analysis results is used as input, and the generative AI model makes inferences based on it. The output is an evaluation of the pet's emotions and health condition.

[0354] Step 6:

[0355] Users input text or audio data into the emotion analysis tool to analyze their emotional state. The input data includes the user's text messages and audio files, and emotion analysis is performed using the NLTK library. The analysis results in an evaluation of the user's emotional state.

[0356] Step 7:

[0357] The cloud server performs a comprehensive assessment to correlate the pet's emotions and health status with the user's emotional state. The pet and user's emotional assessment results are used as input, and comprehensive feedback is generated based on the correlation. The output data includes a comprehensive diagnosis and recommended actions.

[0358] Step 8:

[0359] The cloud server sends the comprehensive feedback to the user's device and notifies the user through a dedicated application, allowing the user to check the pet's current status and recommended actions.

[0360] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0361] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0362] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0363] [Second embodiment]

[0364] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0365] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0366] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0367] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0368] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0369] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0370] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0371] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0372] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0373] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0374] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0375] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0376] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[0377] Specific operations of program processing

[0378] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[0379] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[0380] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[0381] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[0382] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[0383] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[0384] Specific examples

[0385] 1. Data collection

[0386] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[0387] 2. Data transmission and preprocessing

[0388] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[0389] 3. Data Analysis

[0390] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0391] 4. Emotional and health status assessment

[0392] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0393] 5. Feedback and Notifications

[0394] The cloud server sends the results to the user's device, and a dedicated application notifies them, saying, "Your cat is stressed. Attention is needed."

[0395] The introduction of this system will enable owners to monitor their pets' health and emotions in real time and provide appropriate care at an early stage, marking a major step towards providing a comfortable and healthy life for both pets and their owners.

[0396] The processing flow will be explained below.

[0397] Step 1:

[0398] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[0399] Step 2:

[0400] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[0401] Step 3:

[0402] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[0403] Step 4:

[0404] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[0405] Step 5:

[0406] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[0407] Step 6:

[0408] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[0409] Step 7:

[0410] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[0411] Step 8:

[0412] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[0413] Step 9:

[0414] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's mood and health status and take appropriate action if necessary.

[0415] Example 1

[0416] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0417] A pet's behavior and vocalizations are important indicators for understanding its emotions and health status, but it is difficult for owners to constantly monitor them and respond appropriately. Furthermore, if owners are unable to notice abnormalities in their pets early, appropriate care may be delayed, increasing the risk of overlooking pet health problems. Given this situation, there is a demand for a system that can automatically monitor and analyze pet behavior and vocalizations in real time.

[0418] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0419] In this invention, the server includes a recording means for recording the behavior of the pet, an audio acquisition means for recording the sounds the pet makes, a processing means for preprocessing the recorded data, a transmission means for transmitting the preprocessed data via a network, an analysis means for analyzing the data received via the network, a generation algorithm model for determining the emotions and health condition of the pet based on the analysis results, and a notification means for notifying the user of the determination results. This makes it possible to automatically monitor and analyze the behavior and sounds of the pet in real time, accurately grasp the emotions and health condition of the pet, and provide necessary care early.

[0420] "Recording means" refers to a device or function for recording the behavior of a pet.

[0421] "Audio acquisition means" refers to a device or function for recording the sounds of pets.

[0422] "Processing means" refers to devices or functions for pre-processing the recorded data.

[0423] "Transmitting means" refers to a device or function for transmitting preprocessed data over a network.

[0424] "Analysis means" refers to a device or function for analyzing data received via a network.

[0425] A "generative algorithm model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[0426] "Notification means" refers to a device or function for notifying the user of the judgment result.

[0427] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[0428] First, the user installs a pet-specific device in the pet's living space. This device has a built-in camera and microphone, which allows it to record the pet's behavior and sounds in real time. For example, the user can install the device in the living room and adjust the camera's orientation to capture the area where the pet is often found.

[0429] Next, the device captures footage of the pet's behavior with a camera and simultaneously records its meows with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. For example, the device may capture a video of a cat playing by a window and record its meows. At this time, unnecessary parts of the video are deleted and background noise is removed from the audio.

[0430] The preprocessed data is then sent by the device to a cloud server. This transmission is performed using wireless communication such as Wi-Fi or 4G / 5G. To send the preprocessed data to the cloud server, the device connects to the home Wi-Fi network, encrypts the data, and transmits it. After a few seconds, the cloud server receives the data.

[0431] The cloud server applies an object recognition algorithm to the video data it receives to identify the pet's body parts and analyze its behavioral patterns. It also applies a voice recognition algorithm to the audio data to analyze the characteristics of its meows. For example, the cloud server analyzes the video data to identify the behavior of a cat licking its hind paws, and analyzes the audio data to detect that short, sharp meows are a sign of stress.

[0432] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, the generative AI model compares a cat's past data with its current meows and behavioral patterns and determines that it is likely to be stressed.

[0433] The results of the assessment are sent to the user's device via a feedback mechanism. The user's device may be, for example, a smartphone or tablet, and the results are displayed through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, the cloud server may send a notification to the user's device stating, "Your cat is feeling stressed," and a pop-up notification will appear on the user's smartphone. When the user opens the app, detailed analysis results will be displayed.

[0434] For example, when a device records a cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone picks up short, sharp meows. This data is preprocessed and sent to a cloud server where it is analyzed using object recognition and voice recognition algorithms. The generative AI model then determines that the cat is stressed and sends a notification to the user's device saying, "Your cat is stressed. Attention is needed." In this way, users can monitor their pet's health in real time and provide appropriate care early.

[0435] Example prompts for generative AI models

[0436] Below is an example of a prompt sentence to input to the generative AI model.

[0437] "Your cat is frequently licking its hind paws and making short, sharp meowing noises. Use these data to assess your cat's emotional state and health."

[0438] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0439] Step 1: The user places the pet-specific device in the pet's living space. This device has a built-in camera and microphone. The input is the device's installation location and the pet's range of movement. The output is the device's readiness to record the pet's movements and sounds. The user places the device in a location in the living room where the pet is often found and turns it on. The user adjusts the camera's orientation to capture the pet at the optimal angle.

[0440] Step 2: The device captures the pet's behavior with a camera and simultaneously records its cries with a microphone. The input is the pet's behavior and cries. The output is the recorded video and audio data. For example, the device can record a video of the pet playing and simultaneously record the cries it makes.

[0441] Step 3: The device preprocesses the recorded data. This preprocessing includes trimming unnecessary parts of the video data and removing noise. The input is the recorded video data and audio data. The output is the preprocessed video data and audio data. The device removes unnecessary parts of the video data and removes background noise from the audio data.

[0442] Step 4: The device sends the preprocessed data to the cloud server. This is done using wireless communication such as Wi-Fi or 4G / 5G. The input is the preprocessed video and audio data. The output is the data sent to the cloud server. For example, the device sends the data via a home Wi-Fi network, and the cloud server receives it.

[0443] Step 5: The cloud server analyzes the received video data. It applies an object recognition algorithm to identify the pet's body parts and analyzes its behavioral patterns. The input is the video data sent to the cloud server. The output is the analysis results (behavioral patterns). The cloud server analyzes the video data and identifies, for example, the pet licking its hind paws.

[0444] Step 6: The cloud server analyzes the audio data. It applies a speech recognition algorithm to analyze the characteristics of the bird's calls. The input is the audio data sent to the cloud server. The output is the analysis result (the characteristics of the bird's calls). The cloud server analyzes the audio data and detects, for example, that the bird's calls are a sign of stress.

[0445] Step 7: The cloud server inputs the analysis results into the generative AI model. This generative AI model compares it with past data to determine the pet's emotions and health condition. The input is the analysis results of the video and audio data. The output is the determination of the pet's emotions and health condition. The generative AI model compares it with past data and determines that the pet is likely to be feeling stressed.

[0446] Step 8: The cloud server sends the judgment result to the user device via a feedback means. The user device displays the result through a dedicated application. The input is the judgment result of emotions and health status. The output is a notification displayed on the user device. The cloud server sends a notification to the user device saying "The cat is feeling stressed," and the user checks the notification on their smartphone.

[0447] This allows users to understand their pet's condition in real time and provide appropriate care early on.

[0448] (Application example 1)

[0449] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0450] Conventional pet management systems simply record pet behavior and sounds, and are unable to assess a pet's emotions or health status in real time. Additionally, there is a lack of information that allows owners to select appropriate care products and food based on their pet's health status. Furthermore, specific care suggestions based on analysis results are rarely provided, making it difficult for owners to provide detailed care for their pets.

[0451] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0452] In this invention, the server includes a video recording means for recording the behavior of the pet, an audio recording means for recording the sounds of the pet, an information processing means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the information server, an analysis means for analyzing the data received by the information server, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a suggestion means for suggesting appropriate care products and food based on the judgment results, and a notification means for notifying the user of the suggested content. This makes it possible to analyze the emotions and health condition of the pet and notify the user of specific care suggestions and products based on the pet's health condition.

[0453] "Video recording means" is a function for recording the behavior of pets using a camera.

[0454] The "audio recording means" is a function for recording the sounds of pets using a microphone.

[0455] The "information processing means" is a computer system for preprocessing the recorded video data and audio data.

[0456] The "communication means" is a wireless communication function for transmitting preprocessed data to an information server such as a cloud server.

[0457] The "analysis means" refers to a group of algorithms and programs for analyzing data received by the information server.

[0458] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[0459] The "suggestion means" is a function for suggesting care products and food suitable for pets based on the judgment results.

[0460] "Notification means" refers to a notification system for informing users of the proposed content.

[0461] An "object recognition algorithm" is an algorithm that analyzes pet behavior data and recognizes specific actions and movements.

[0462] The "voice recognition algorithm" is an algorithm for analyzing pet cry data and recognizing specific voice patterns.

[0463] The present invention relates to a system that records a pet's behavior and cries in real time, analyzes this data to determine the pet's emotions and health condition, and then suggests appropriate care products and food based on the results of this determination.

[0464] The system consists of the following main tools:

[0465] 1. Video recording means: Recording pet behavior using a camera. For example, a camera device specifically for pets is included.

[0466] 2. Audio recording means: Record the sounds your pet makes using a microphone. This also includes a dedicated audio recording device for pets.

[0467] 3. Information processing means: A computer system for preprocessing recorded video and audio data, such as noise removal and data trimming.

[0468] 4. Communication means: Equipped with wireless communication functions for transmitting preprocessed data to the information server. Specific examples include Wi-Fi and 4G / 5G communication modules.

[0469] 5. Analysis means: A group of algorithms and programs for analyzing data received by the information server. This includes object recognition algorithms and voice recognition algorithms.

[0470] 6. Generative AI model: An artificial intelligence model that determines a pet's emotions and health status based on analysis results.

[0471] 7. Suggestion means: Based on the judgment results, it has a function to suggest care products and food suitable for pets.

[0472] 8. Notification method: A notification system to inform users of the proposed content. For example, a notification is sent to a smartphone.

[0473] The server performs the following main tasks:

[0474] 1. Analysis of video and audio data. Characterize pet behavior and sounds using object and audio recognition algorithms.

[0475] 2. The analysis results are input into a generative AI model to determine the pet's emotions and health status.

[0476] 3. Based on the judgment results, the appropriate care products and food for the pet are determined using a suggestion method.

[0477] 4. The proposal will be communicated to the user via notification means.

[0478] For example, if a pet frequently scratches its ears, the server can analyze the behavior and determine that this is a sign of an ear infection. Based on this determination, appropriate care products (e.g., ear cleaning wipes or anti-infection shampoo) can be suggested and notified to the user.

[0479] An example prompt is:

[0480] Prompt for the generative AI model:

[0481] You are an AI model that analyzes your pet's behavior and vocalizations to determine its emotions and health status. Analyze the following data to determine your pet's emotions and health status:

[0482] Action: {action}

[0483] Sound: {sound}

[0484] Based on the results, we will recommend the appropriate care products and food for your pet.

[0485] This system allows users to understand their pet's health condition and emotions in real time, enabling them to provide prompt and accurate physical and psychological care.

[0486] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0487] Step 1:

[0488] The terminal records the pet's behavior using a video recording means and records the pet's cries using an audio recording means. The input includes video data from the camera and audio data from the microphone. These data are temporarily stored for subsequent processing.

[0489] Step 2:

[0490] The device preprocesses the recorded video and audio data. Specific preprocessing operations include trimming the video data, removing noise, and adjusting the resolution. For audio data, noise removal and extraction of necessary frequency components are also performed. Clean video and audio data are generated as the output of the preprocessing.

[0491] Step 3:

[0492] The terminal transmits the preprocessed video and audio data to the information server using a communication means, with the input comprising the preprocessed data and the output being data packets received by the information server.

[0493] Step 4:

[0494] The server analyzes the received data using an analysis means. Specifically, an object recognition algorithm is applied to the video data to detect the specific behavior of the pet. A voice recognition algorithm is applied to the audio data to extract the characteristics and patterns of the pet's cries. The input includes the received data, and the output generates data on behavioral patterns and voice patterns, which are the analysis results.

[0495] Step 5:

[0496] The server inputs the analysis results into a generative AI model and performs calculations to determine the pet's emotions and health condition. The input includes the analysis result data, and the output generates a judgment result on the pet's emotions and health condition. For example, based on the analysis results, it may be determined that the pet is "stressed."

[0497] Step 6:

[0498] The server uses a suggestion tool to determine the appropriate pet care products and food based on the assessment results. The input includes the assessment results of the pet's emotions and health condition, and the output generates a list of suggested products. For example, foods and toys with a relaxing effect are suggested for a stressed pet.

[0499] Step 7:

[0500] The server notifies the user of the proposed content using a notification method. Specifically, it sends a notification to the user's smartphone or tablet. The input includes a list of proposed products, and the output is generated as notification information to be displayed on the user's device. The user can then receive this notification and check the details of the proposed products.

[0501] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0502] This invention combines a system that records pet behavior and sounds in real time and analyzes this data to determine the pet's emotions and health condition with an emotion engine that recognizes the user's emotions. This system allows for accurate understanding of the condition of both the pet and the owner, enabling more appropriate care and communication.

[0503] Specific operations of program processing

[0504] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[0505] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[0506] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[0507] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[0508] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[0509] Furthermore, the system incorporates an emotion engine that recognizes the user's emotions. The user inputs text and voice data into the emotion engine, which then analyzes the user's emotional state. It also uses a facial recognition algorithm to analyze the user's facial expressions and evaluate their emotional state.

[0510] The cloud server correlates the user's emotional state with the pet's state and determines the overall communication state. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress state to determine how the two are affecting each other.

[0511] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[0512] Specific examples

[0513] 1. Data collection

[0514] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[0515] 2. Data transmission and preprocessing

[0516] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[0517] 3. Data Analysis

[0518] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0519] 4. Emotional and health status assessment

[0520] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0521] 5. Recognizing user emotions

[0522] Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[0523] 6. Overall Judgment

[0524] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0525] 7. Feedback and Notifications

[0526] The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is feeling stressed. Attention is needed." Users can also check their own emotional state and take necessary action.

[0527] This system allows for a comprehensive understanding of the emotions and health status of both pets and users, enabling appropriate care and communication.

[0528] The processing flow will be explained below.

[0529] Step 1:

[0530] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[0531] Step 2:

[0532] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[0533] Step 3:

[0534] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[0535] Step 4:

[0536] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[0537] Step 5:

[0538] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[0539] Step 6:

[0540] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[0541] Step 7:

[0542] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[0543] Step 8:

[0544] The user inputs text or voice data into the emotion engine, which then analyzes the user's emotional state.

[0545] Step 9:

[0546] The user's facial expressions are captured by a camera, and the emotion engine uses a facial recognition algorithm to analyze the user's facial expression data, thereby assessing the user's emotional state.

[0547] Step 10:

[0548] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0549] Step 11:

[0550] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[0551] Step 12:

[0552] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's emotional and health status, as well as their own emotional state, and take necessary actions.

[0553] Example 2

[0554] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0555] Pet owners face the challenge of understanding their pets' emotions and health status from their behavior and vocalizations. There is a need for a system that can properly recognize the emotional state of not only pets but also their owners, and grasp the overall state of communication with their pets. This will enable them to effectively manage the health and happiness of both pets and their owners, and provide appropriate care and advice.

[0556] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0557] In this invention, the server includes a camera for recording the pet's behavior, an audio collection device for recording the pet's cries, a processing device for preprocessing the recorded data, a communication device for transmitting the preprocessed data to a cloud server, an analysis device for analyzing the data received by the cloud server, a generative AI model for determining the pet's emotions and health status based on the analysis results, a transmission device for notifying the user of the determination results, an emotion engine for recognizing the user's emotions, and a comprehensive determination engine for correlating the user's emotional state with the pet's emotional state and making a comprehensive determination. This makes it possible to accurately grasp the pet's emotions and health status in real time and also evaluate the owner's emotional state. This allows for a comprehensive understanding of the pet's communication status and provides appropriate care and advice.

[0558] "Photographing means" refers to a device used to record the behavior of a pet.

[0559] An "audio collection means" is a device used to record the sounds of a pet.

[0560] "Processing equipment" is equipment used to pre-process recorded data.

[0561] A "communication device" is a device used to transmit pre-processed data to a cloud server.

[0562] An "analysis device" is a device used to analyze data received by a cloud server.

[0563] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on analysis results.

[0564] A "communication device" is a device used to notify the user of the determination result.

[0565] An "emotion engine" is a device or software used to recognize a user's emotions.

[0566] The "comprehensive judgment engine" is a device or software used to correlate and comprehensively judge the emotional state of the user and the emotional state of the pet.

[0567] This invention combines a system that determines a pet's emotions and health condition by recording and analyzing the pet's behavior and cries in real time with an emotion engine that recognizes the user's emotions. The following describes an embodiment of this system.

[0568] The first thing users need to do is to install a dedicated pet device in their pet's living space. This device is equipped with a high-resolution camera (photography) and a highly sensitive microphone (audio collection), allowing it to record the pet's behavior and sounds in real time.

[0569] The device then captures images of the pet's behavior with a camera and records its cries with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. This process produces data that is easy to analyze.

[0570] The preprocessed data is sent from the device to the cloud server using Wi-Fi or 4G / 5G wireless communication (communication equipment), allowing the data to reach the cloud server safely and quickly.

[0571] To analyze the received data, the cloud server first applies an object recognition algorithm to the video data. This algorithm identifies the pet's body parts and analyzes its behavioral patterns. It also applies a voice recognition algorithm to the audio data, analyzing the characteristics of the pet's cries. This allows for a detailed analysis of your pet's behavior.

[0572] The cloud server then uses a generative AI model to determine the pet's emotions and health based on the analyzed data. The generative AI model compares the analysis results with past data to determine, for example, whether the pet is stressed or in pain.

[0573] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotions. The user can input text or voice data into the emotion engine, which then analyzes the user's emotional state. The system can also record the user's facial expressions using a camera and evaluate their emotional state using a facial recognition algorithm.

[0574] The cloud server correlates the pet's emotional and health status with the user's emotional status and uses a comprehensive assessment engine to determine the overall communication status. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress status to determine how the two are affecting each other.

[0575] The results are sent from the cloud server to the user's device and displayed via a dedicated application. Notifications are displayed, allowing users to check their pet's current condition and whether any abnormalities have been detected.

[0576] Specific examples

[0577] As a concrete example, consider the following scenario.

[0578] 1. The device records the cat's behavior: the camera captures the cat frequently licking its hind paws, and the microphone picks up the cat's short, sharp meows.

[0579] 2. The collected data is pre-processed on the device, trimmed, and denoised before being sent to the cloud server.

[0580] 3. The cloud server analyzes the video data using an object recognition algorithm to confirm the cat's licking of its hind paw, and analyzes the audio data using a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0581] 4. The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0582] 5. Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[0583] 6. The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0584] 7. The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is stressed. Attention is needed." The user can also check their own emotional state and take necessary action.

[0585] In this way, this system allows for a comprehensive understanding of the emotions and health status of both the pet and the user, enabling appropriate care and communication.

[0586] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0587] Step 1:

[0588] The user installs a pet-specific device in the pet's living space.

[0589] How it works: The user places a device containing a high-resolution camera and a sensitive microphone in a location where the pet is frequently present or resting.

[0590] Step 2:

[0591] The device captures your pet's behavior with a camera and records its cries with a microphone.

[0592] How it works: The camera captures pet movements in high resolution, and the microphone records pet sounds with high sensitivity, providing real-time data on pet behavior and sounds.

[0593] Input: Pet actions and sounds

[0594] Output: Raw video and audio data

[0595] Step 3:

[0596] Preprocessing is performed on the data collected by the terminal.

[0597] How it works: The device trims unnecessary parts of the video data and removes noise from the audio data. This preprocessing produces data that is easy to analyze.

[0598] Input: Raw video and audio data

[0599] Output: Pre-processed video and audio data

[0600] Step 4:

[0601] The terminal transmits the preprocessed data to the cloud server.

[0602] How it works: The device uses Wi-Fi or 4G / 5G to send pre-processed data to the cloud server, where it is encrypted to ensure security.

[0603] Input: Preprocessed video and audio data

[0604] Output: Data sent to the cloud server

[0605] Step 5:

[0606] The cloud server analyzes the received data.

[0607] How it works: The cloud server first applies an object recognition algorithm to the video data to identify the pet's body parts and analyze its behavioral patterns, and then applies a voice recognition algorithm to the audio data to analyze the characteristics of the pet's cries.

[0608] Input: Video and audio data sent to the cloud server

[0609] Output: Analyzed behavioral and audio data

[0610] Step 6:

[0611] The cloud server uses the generated AI model to determine the pet's emotions and health condition based on the analysis results.

[0612] How it works: The generative AI model compares your pet's behavior and vocalizations with past data to determine their emotions and health status. For example, it analyzes the frequency of vocalizations and behavioral patterns to determine whether your pet is stressed or in pain.

[0613] Input: Analyzed behavioral and audio data

[0614] Output: Pet's emotions and health status

[0615] Step 7:

[0616] Users input text or voice into the emotion engine, which analyzes their emotional state.

[0617] How it works: Users input text or voice to the emotion engine to recognize their emotional state, and a facial recognition algorithm is used to assess the user's emotional state from facial expression data captured by the camera.

[0618] Input: User text, voice, and facial expression data

[0619] Output: Analysis of the user's emotional state

[0620] Step 8:

[0621] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0622] Specific operation: The cloud server uses a comprehensive judgment engine to correlate and analyze the emotional state of the pet and the user. For example, when the user is feeling stressed, the cloud server will also analyze the pet's stress state and evaluate the overall communication state.

[0623] Input: Pet's emotional and health status assessment results, analysis results of user's emotional state

[0624] Output: Overall communication status assessment

[0625] Step 9:

[0626] The cloud server sends the results to the user's device, and a dedicated application displays a notification.

[0627] Specific operation: The cloud server uses a transmission device to send the judgment result to the user's device. A dedicated application on the user's device displays a notification, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, a notification such as "Your cat is stressed. Attention is required" may be displayed.

[0628] Input: Overall communication status assessment result

[0629] Output: Notification displayed on the user's device

[0630] This series of steps creates a system that can comprehensively understand the emotions and health status of both the pet and the user, and provide appropriate care and communication.

[0631] (Application example 2)

[0632] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0633] In today's world, it is extremely important to properly understand pet health and emotional states and facilitate smooth communication with owners. However, many current systems simply record pet behavior and vocalizations, lacking the ability to comprehensively analyze pets' emotions and health. Furthermore, no systems provide appropriate feedback that takes into account the user's emotional state. Therefore, there is a need for a system that can analyze the state of both pets and users in real time to achieve optimal care and communication.

[0634] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a photographing means for recording the behavior of the pet, a sound collecting means for recording the sounds of the pet, a processing device means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the data processing device, an analysis device means for analyzing the data received by the data processing device, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a notification device means for notifying the user of the determination results, an emotion analysis means for analyzing the user's emotional state, and an association means for associating the pet's state with the user's emotional state. This makes it possible to analyze detailed data including the behavior and sounds of the pet and provide comprehensive advice that also takes the user's emotional state into consideration.

[0635] "Photographing means" refers to a camera device used to record the behavior of a pet.

[0636] The "audio collection means" is a microphone device for recording the sounds of pets.

[0637] The "processing device means" is a device for pre-processing data recorded by the imaging means and sound collection means.

[0638] "Communication means" refers to a wireless or wired communication device for transmitting pre-processed data to a data processing device.

[0639] "Analysis device means" is a general term for hardware and software for analyzing received data in a data processing device.

[0640] A "generative AI model" is a machine learning model that determines a pet's emotions and health status based on analysis results.

[0641] "Notification device means" is a device for notifying the user of the results determined by the generative AI model.

[0642] "Emotion analysis means" is a system for analyzing the emotional state of a user.

[0643] The "association means" is a mechanism for associating and analyzing the pet's emotions and health status with the user's emotional state.

[0644] This invention relates to a system that records pet behavior and sounds in real time and provides optimal feedback by taking into account the user's emotional state. The system includes three main components: a pet-specific device, a user terminal, and a cloud server.

[0645] Hardware and software used

[0646] Hardware:

[0647] Recording method: Camera equipment to record pet behavior

[0648] Sound collection means: microphone device for recording pet sounds

[0649] Processing Unit: A local device (e.g., Raspberry Pi) for preprocessing data.

[0650] Communication: Wi-Fi or 4G / 5G communication module for sending data to a cloud server

[0651] software:

[0652] TensorFlow: Deep learning object and speech recognition algorithms

[0653] OpenCV: Image processing library

[0654] Pydub: A library for preprocessing audio data

[0655] NLTK: A natural language processing library for analyzing user sentiment

[0656] System operation explanation

[0657] 1. Data Collection:

[0658] The pet device uses a camera to capture images of your pet's behavior and a microphone to record your pet's sounds, and this data is collected in real time.

[0659] 2. Data preprocessing:

[0660] Processing means within the terminal trims the collected video data and removes noise from the audio data.

[0661] 3. Data transmission:

[0662] The preprocessed data is sent to a cloud server via a communication method such as Wi-Fi or 4G / 5G.

[0663] 4. Data Analysis:

[0664] The cloud server analyzes the received data using an analysis device means. An object recognition algorithm using TensorFlow is applied to the video data, and a voice recognition algorithm also using TensorFlow is applied to the voice data.

[0665] 5. Emotional and health status assessment:

[0666] Based on the analysis results, the generative AI model determines the pet's emotions and health status. For example, if the pet repeatedly behaves in a certain way or makes a certain sound, stress or health problems may be suspected.

[0667] 6. User Emotion Recognition:

[0668] By inputting text or voice from the user's device, the user's emotional state is analyzed by the emotion analysis means. The emotion analysis is performed using the NLTK library.

[0669] 7. Overall Judgment and Association:

[0670] The system analyzes the pet's emotions and health status in relation to the user's emotional state, and generates optimal feedback based on the results.

[0671] 8. Feedback and Notifications:

[0672] The cloud server sends the results to the user's device, where they are displayed via a dedicated application, allowing the user to receive feedback and take appropriate action.

[0673] Specific examples

[0674] For example, the device records the cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone records the cat's sharp meows. This data is sent to a cloud server and analyzed using object recognition and voice recognition algorithms. The generative AI model determines that the cat is likely feeling stressed. The user enters text into the emotion engine, which analyzes the cat's stress. The cloud server correlates this data and makes a comprehensive judgment, and a notification such as "Your cat is feeling stressed. Attention is required" is displayed on the user's device.

[0675] Example prompts to input to the generative AI model:

[0676] Pet Behavior: Cat repeatedly licks the same spot

[0677] Call: High-pitched, sharp call

[0678] User Emotion: Stressed

[0679] Determine your pet's emotional and health status.

[0680] This system makes it possible to perform detailed analysis of a pet's behavior and cries, and provide optimal actions that take into account the user's emotional state.

[0681] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0682] Step 1:

[0683] The device uses a camera and microphone to record your pet's movements and sounds in real time. This input data is used to generate video and audio files, which are then temporarily stored on the device.

[0684] Step 2:

[0685] The processing unit of the terminal performs preprocessing on the recorded video data and audio data, such as trimming the video data and adjusting the resolution, and removing noise from the audio data. As a result of the preprocessing, preprocessed video and audio files are generated.

[0686] Step 3:

[0687] The communication means transmits the preprocessed data to the cloud server. Specifically, the data is uploaded using Wi-Fi or 4G / 5G. After transmission, the data is stored in the cloud server.

[0688] Step 4:

[0689] The cloud server analyzes the received video data and audio data using an analysis device. An object recognition algorithm using TensorFlow is applied to the video data to extract the pet's behavioral characteristics. A voice recognition algorithm also using TensorFlow is applied to the audio data to extract the characteristics of the pet's cries. As a result of the analysis, data on the pet's behavioral characteristics and audio characteristics is generated.

[0690] Step 5:

[0691] The cloud server uses a generative AI model based on the analysis results to determine the pet's emotions and health condition. The feature data from the analysis results is used as input, and the generative AI model makes inferences based on it. The output is an evaluation of the pet's emotions and health condition.

[0692] Step 6:

[0693] Users input text or audio data into the emotion analysis tool to analyze their emotional state. The input data includes the user's text messages and audio files, and emotion analysis is performed using the NLTK library. The analysis results in an evaluation of the user's emotional state.

[0694] Step 7:

[0695] The cloud server performs a comprehensive assessment to correlate the pet's emotions and health status with the user's emotional state. The pet and user's emotional assessment results are used as input, and comprehensive feedback is generated based on the correlation. The output data includes a comprehensive diagnosis and recommended actions.

[0696] Step 8:

[0697] The cloud server sends the comprehensive feedback to the user's device and notifies the user through a dedicated application, allowing the user to check the pet's current status and recommended actions.

[0698] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0699] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0700] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0701] [Third embodiment]

[0702] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0703] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0704] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0705] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0706] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0707] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0708] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0709] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0710] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0711] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0712] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0713] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0714] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[0715] Specific operations of program processing

[0716] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[0717] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[0718] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[0719] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[0720] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[0721] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[0722] Specific examples

[0723] 1. Data collection

[0724] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[0725] 2. Data transmission and preprocessing

[0726] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[0727] 3. Data Analysis

[0728] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0729] 4. Emotional and health status assessment

[0730] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0731] 5. Feedback and Notifications

[0732] The cloud server sends the results to the user's device, and a dedicated application notifies them, saying, "Your cat is stressed. Attention is needed."

[0733] The introduction of this system will enable owners to monitor their pets' health and emotions in real time and provide appropriate care at an early stage, marking a major step towards providing a comfortable and healthy life for both pets and their owners.

[0734] The processing flow will be explained below.

[0735] Step 1:

[0736] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[0737] Step 2:

[0738] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[0739] Step 3:

[0740] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[0741] Step 4:

[0742] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[0743] Step 5:

[0744] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[0745] Step 6:

[0746] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[0747] Step 7:

[0748] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[0749] Step 8:

[0750] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[0751] Step 9:

[0752] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's mood and health status and take appropriate action if necessary.

[0753] Example 1

[0754] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0755] A pet's behavior and vocalizations are important indicators for understanding its emotions and health status, but it is difficult for owners to constantly monitor them and respond appropriately. Furthermore, if owners are unable to notice abnormalities in their pets early, appropriate care may be delayed, increasing the risk of overlooking pet health problems. Given this situation, there is a demand for a system that can automatically monitor and analyze pet behavior and vocalizations in real time.

[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0757] In this invention, the server includes a recording means for recording the behavior of the pet, an audio acquisition means for recording the sounds the pet makes, a processing means for preprocessing the recorded data, a transmission means for transmitting the preprocessed data via a network, an analysis means for analyzing the data received via the network, a generation algorithm model for determining the emotions and health condition of the pet based on the analysis results, and a notification means for notifying the user of the determination results. This makes it possible to automatically monitor and analyze the behavior and sounds of the pet in real time, accurately grasp the emotions and health condition of the pet, and provide necessary care early.

[0758] "Recording means" refers to a device or function for recording the behavior of a pet.

[0759] "Audio acquisition means" refers to a device or function for recording the sounds of pets.

[0760] "Processing means" refers to devices or functions for pre-processing the recorded data.

[0761] "Transmitting means" refers to a device or function for transmitting preprocessed data over a network.

[0762] "Analysis means" refers to a device or function for analyzing data received via a network.

[0763] A "generative algorithm model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[0764] "Notification means" refers to a device or function for notifying the user of the judgment result.

[0765] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[0766] First, the user installs a pet-specific device in the pet's living space. This device has a built-in camera and microphone, which allows it to record the pet's behavior and sounds in real time. For example, the user can install the device in the living room and adjust the camera's orientation to capture the area where the pet is often found.

[0767] Next, the device captures footage of the pet's behavior with a camera and simultaneously records its meows with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. For example, the device may capture a video of a cat playing by a window and record its meows. At this time, unnecessary parts of the video are deleted and background noise is removed from the audio.

[0768] The preprocessed data is then sent by the device to a cloud server. This transmission is performed using wireless communication such as Wi-Fi or 4G / 5G. To send the preprocessed data to the cloud server, the device connects to the home Wi-Fi network, encrypts the data, and transmits it. After a few seconds, the cloud server receives the data.

[0769] The cloud server applies an object recognition algorithm to the video data it receives to identify the pet's body parts and analyze its behavioral patterns. It also applies a voice recognition algorithm to the audio data to analyze the characteristics of its meows. For example, the cloud server analyzes the video data to identify the behavior of a cat licking its hind paws, and analyzes the audio data to detect that short, sharp meows are a sign of stress.

[0770] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, the generative AI model compares a cat's past data with its current meows and behavioral patterns and determines that it is likely to be stressed.

[0771] The results of the assessment are sent to the user's device via a feedback mechanism. The user's device may be, for example, a smartphone or tablet, and the results are displayed through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, the cloud server may send a notification to the user's device stating, "Your cat is feeling stressed," and a pop-up notification will appear on the user's smartphone. When the user opens the app, detailed analysis results will be displayed.

[0772] For example, when a device records a cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone picks up short, sharp meows. This data is preprocessed and sent to a cloud server where it is analyzed using object recognition and voice recognition algorithms. The generative AI model then determines that the cat is stressed and sends a notification to the user's device saying, "Your cat is stressed. Attention is needed." In this way, users can monitor their pet's health in real time and provide appropriate care early.

[0773] Example prompts for generative AI models

[0774] Below is an example of a prompt sentence to input to the generative AI model.

[0775] "Your cat is frequently licking its hind paws and making short, sharp meowing noises. Use these data to assess your cat's emotional state and health."

[0776] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0777] Step 1: The user places the pet-specific device in the pet's living space. This device has a built-in camera and microphone. The input is the device's installation location and the pet's range of movement. The output is the device's readiness to record the pet's movements and sounds. The user places the device in a location in the living room where the pet is often found and turns it on. The user adjusts the camera's orientation to capture the pet at the optimal angle.

[0778] Step 2: The device captures the pet's behavior with a camera and simultaneously records its cries with a microphone. The input is the pet's behavior and cries. The output is the recorded video and audio data. For example, the device can record a video of the pet playing and simultaneously record the cries it makes.

[0779] Step 3: The device preprocesses the recorded data. This preprocessing includes trimming unnecessary parts of the video data and removing noise. The input is the recorded video data and audio data. The output is the preprocessed video data and audio data. The device removes unnecessary parts of the video data and removes background noise from the audio data.

[0780] Step 4: The device sends the preprocessed data to the cloud server. This is done using wireless communication such as Wi-Fi or 4G / 5G. The input is the preprocessed video and audio data. The output is the data sent to the cloud server. For example, the device sends the data via a home Wi-Fi network, and the cloud server receives it.

[0781] Step 5: The cloud server analyzes the received video data. It applies an object recognition algorithm to identify the pet's body parts and analyzes its behavioral patterns. The input is the video data sent to the cloud server. The output is the analysis results (behavioral patterns). The cloud server analyzes the video data and identifies, for example, the pet licking its hind paws.

[0782] Step 6: The cloud server analyzes the audio data. It applies a speech recognition algorithm to analyze the characteristics of the bird's calls. The input is the audio data sent to the cloud server. The output is the analysis result (the characteristics of the bird's calls). The cloud server analyzes the audio data and detects, for example, that the bird's calls are a sign of stress.

[0783] Step 7: The cloud server inputs the analysis results into the generative AI model. This generative AI model compares it with past data to determine the pet's emotions and health condition. The input is the analysis results of the video and audio data. The output is the determination of the pet's emotions and health condition. The generative AI model compares it with past data and determines that the pet is likely to be feeling stressed.

[0784] Step 8: The cloud server sends the judgment result to the user device via a feedback means. The user device displays the result through a dedicated application. The input is the judgment result of emotions and health status. The output is a notification displayed on the user device. The cloud server sends a notification to the user device saying "The cat is feeling stressed," and the user checks the notification on their smartphone.

[0785] This allows users to understand their pet's condition in real time and provide appropriate care early on.

[0786] (Application example 1)

[0787] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0788] Conventional pet management systems simply record pet behavior and sounds, and are unable to assess a pet's emotions or health status in real time. Additionally, there is a lack of information that allows owners to select appropriate care products and food based on their pet's health status. Furthermore, specific care suggestions based on analysis results are rarely provided, making it difficult for owners to provide detailed care for their pets.

[0789] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0790] In this invention, the server includes a video recording means for recording the behavior of the pet, an audio recording means for recording the sounds of the pet, an information processing means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the information server, an analysis means for analyzing the data received by the information server, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a suggestion means for suggesting appropriate care products and food based on the judgment results, and a notification means for notifying the user of the suggested content. This makes it possible to analyze the emotions and health condition of the pet and notify the user of specific care suggestions and products based on the pet's health condition.

[0791] "Video recording means" is a function for recording the behavior of pets using a camera.

[0792] The "audio recording means" is a function for recording the sounds of pets using a microphone.

[0793] The "information processing means" is a computer system for preprocessing the recorded video data and audio data.

[0794] The "communication means" is a wireless communication function for transmitting preprocessed data to an information server such as a cloud server.

[0795] The "analysis means" refers to a group of algorithms and programs for analyzing data received by the information server.

[0796] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[0797] The "suggestion means" is a function for suggesting care products and food suitable for pets based on the judgment results.

[0798] "Notification means" refers to a notification system for informing users of the proposed content.

[0799] An "object recognition algorithm" is an algorithm that analyzes pet behavior data and recognizes specific actions and movements.

[0800] The "voice recognition algorithm" is an algorithm for analyzing pet cry data and recognizing specific voice patterns.

[0801] The present invention relates to a system that records a pet's behavior and cries in real time, analyzes this data to determine the pet's emotions and health condition, and then suggests appropriate care products and food based on the results of this determination.

[0802] The system consists of the following main tools:

[0803] 1. Video recording means: Recording pet behavior using a camera. For example, a camera device specifically for pets is included.

[0804] 2. Audio recording means: Record the sounds your pet makes using a microphone. This also includes a dedicated audio recording device for pets.

[0805] 3. Information processing means: A computer system for preprocessing recorded video and audio data, such as noise removal and data trimming.

[0806] 4. Communication means: Equipped with wireless communication functions for transmitting preprocessed data to the information server. Specific examples include Wi-Fi and 4G / 5G communication modules.

[0807] 5. Analysis means: A group of algorithms and programs for analyzing data received by the information server. This includes object recognition algorithms and voice recognition algorithms.

[0808] 6. Generative AI model: An artificial intelligence model that determines a pet's emotions and health status based on the analysis results.

[0809] 7. Suggestion means: Based on the judgment results, it has a function to suggest care products and food suitable for pets.

[0810] 8. Notification method: A notification system to inform users of the proposed content. For example, a notification is sent to a smartphone.

[0811] The server performs the following main tasks:

[0812] 1. Analysis of video and audio data. Characterize pet behavior and sounds using object and audio recognition algorithms.

[0813] 2. The analysis results are input into a generative AI model to determine the pet's emotions and health status.

[0814] 3. Based on the judgment results, the appropriate care products and food for the pet are determined using a suggestion method.

[0815] 4. The proposal will be communicated to the user via notification means.

[0816] For example, if a pet frequently scratches its ears, the server can analyze the behavior and determine that this is a sign of an ear infection. Based on this determination, appropriate care products (e.g., ear cleaning wipes or anti-infection shampoo) can be suggested and notified to the user.

[0817] An example prompt is:

[0818] Prompt for the generative AI model:

[0819] You are an AI model that analyzes your pet's behavior and vocalizations to determine its emotions and health status. Analyze the following data to determine your pet's emotions and health status:

[0820] Action: {action}

[0821] Sound: {sound}

[0822] Based on the results, we will recommend the appropriate care products and food for your pet.

[0823] This system allows users to understand their pet's health and emotions in real time, enabling them to provide prompt and appropriate physical and psychological care.

[0824] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0825] Step 1:

[0826] The terminal records the pet's behavior using a video recording means and records the pet's cries using an audio recording means. The input includes video data from the camera and audio data from the microphone. These data are temporarily stored for subsequent processing.

[0827] Step 2:

[0828] The device preprocesses the recorded video and audio data. Specific preprocessing operations include trimming the video data, removing noise, and adjusting the resolution. For audio data, noise removal and extraction of necessary frequency components are also performed. Clean video and audio data are generated as the output of the preprocessing.

[0829] Step 3:

[0830] The terminal transmits the preprocessed video and audio data to the information server using a communication means, with the input comprising the preprocessed data and the output being data packets received by the information server.

[0831] Step 4:

[0832] The server analyzes the received data using an analysis means. Specifically, an object recognition algorithm is applied to the video data to detect the specific behavior of the pet. A voice recognition algorithm is applied to the audio data to extract the characteristics and patterns of the pet's cries. The input includes the received data, and the output generates data on behavioral patterns and voice patterns, which are the analysis results.

[0833] Step 5:

[0834] The server inputs the analysis results into a generative AI model and performs calculations to determine the pet's emotions and health condition. The input includes the analysis result data, and the output generates a judgment result on the pet's emotions and health condition. For example, based on the analysis results, it may be determined that the pet is "stressed."

[0835] Step 6:

[0836] The server uses a suggestion tool to determine the appropriate pet care products and food based on the assessment results. The input includes the assessment results of the pet's emotions and health condition, and the output generates a list of suggested products. For example, foods and toys with a relaxing effect are suggested for a stressed pet.

[0837] Step 7:

[0838] The server notifies the user of the proposed content using a notification method. Specifically, it sends a notification to the user's smartphone or tablet. The input includes a list of proposed products, and the output is generated as notification information to be displayed on the user's device. The user can then receive this notification and check the details of the proposed products.

[0839] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0840] This invention combines a system that records pet behavior and sounds in real time and analyzes this data to determine the pet's emotions and health condition with an emotion engine that recognizes the user's emotions. This system allows for accurate understanding of the condition of both the pet and the owner, enabling more appropriate care and communication.

[0841] Specific operations of program processing

[0842] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[0843] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[0844] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[0845] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[0846] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[0847] Furthermore, the system incorporates an emotion engine that recognizes the user's emotions. The user inputs text and voice data into the emotion engine, which then analyzes the user's emotional state. It also uses a facial recognition algorithm to analyze the user's facial expressions and evaluate their emotional state.

[0848] The cloud server correlates the user's emotional state with the pet's state and determines the overall communication state. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress state to determine how the two are affecting each other.

[0849] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[0850] Specific examples

[0851] 1. Data collection

[0852] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[0853] 2. Data transmission and preprocessing

[0854] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[0855] 3. Data Analysis

[0856] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0857] 4. Emotional and health status assessment

[0858] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0859] 5. Recognizing user emotions

[0860] Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[0861] 6. Overall Judgment

[0862] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0863] 7. Feedback and Notifications

[0864] The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is feeling stressed. Attention is needed." Users can also check their own emotional state and take necessary action.

[0865] This system allows for a comprehensive understanding of the emotions and health status of both pets and users, enabling appropriate care and communication.

[0866] The processing flow will be explained below.

[0867] Step 1:

[0868] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[0869] Step 2:

[0870] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[0871] Step 3:

[0872] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[0873] Step 4:

[0874] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[0875] Step 5:

[0876] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[0877] Step 6:

[0878] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[0879] Step 7:

[0880] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[0881] Step 8:

[0882] The user inputs text or voice data into the emotion engine, which then analyzes the user's emotional state.

[0883] Step 9:

[0884] The user's facial expressions are captured by a camera, and the emotion engine uses a facial recognition algorithm to analyze the user's facial expression data, thereby assessing the user's emotional state.

[0885] Step 10:

[0886] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0887] Step 11:

[0888] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[0889] Step 12:

[0890] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's emotional and health status, as well as their own emotional state, and take necessary actions.

[0891] Example 2

[0892] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0893] Pet owners face the challenge of understanding their pets' emotions and health status from their behavior and vocalizations. There is a need for a system that can properly recognize the emotional state of not only pets but also their owners, and grasp the overall state of communication with their pets. This will enable them to effectively manage the health and happiness of both pets and their owners, and provide appropriate care and advice.

[0894] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0895] In this invention, the server includes a camera for recording the pet's behavior, an audio collection device for recording the pet's cries, a processing device for preprocessing the recorded data, a communication device for transmitting the preprocessed data to a cloud server, an analysis device for analyzing the data received by the cloud server, a generative AI model for determining the pet's emotions and health status based on the analysis results, a transmission device for notifying the user of the determination results, an emotion engine for recognizing the user's emotions, and a comprehensive determination engine for correlating the user's emotional state with the pet's emotional state and making a comprehensive determination. This makes it possible to accurately grasp the pet's emotions and health status in real time and also evaluate the owner's emotional state. This allows for a comprehensive understanding of the pet's communication status and provides appropriate care and advice.

[0896] "Photographing means" refers to a device used to record the behavior of a pet.

[0897] An "audio collection means" is a device used to record the sounds of a pet.

[0898] "Processing equipment" is equipment used to pre-process recorded data.

[0899] A "communication device" is a device used to transmit pre-processed data to a cloud server.

[0900] An "analysis device" is a device used to analyze data received by a cloud server.

[0901] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on analysis results.

[0902] A "communication device" is a device used to notify the user of the determination result.

[0903] An "emotion engine" is a device or software used to recognize a user's emotions.

[0904] The "comprehensive judgment engine" is a device or software used to correlate and comprehensively judge the emotional state of the user and the emotional state of the pet.

[0905] This invention combines a system that determines a pet's emotions and health condition by recording and analyzing the pet's behavior and cries in real time with an emotion engine that recognizes the user's emotions. The following describes an embodiment of this system.

[0906] The first thing users need to do is to install a dedicated pet device in their pet's living space. This device is equipped with a high-resolution camera (photography) and a highly sensitive microphone (audio collection), allowing it to record the pet's behavior and sounds in real time.

[0907] The device then captures images of the pet's behavior with a camera and records its cries with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. This process produces data that is easy to analyze.

[0908] The preprocessed data is sent from the device to the cloud server using Wi-Fi or 4G / 5G wireless communication (communication equipment), allowing the data to reach the cloud server safely and quickly.

[0909] To analyze the received data, the cloud server first applies an object recognition algorithm to the video data. This algorithm identifies the pet's body parts and analyzes its behavioral patterns. It also applies a voice recognition algorithm to the audio data, analyzing the characteristics of the pet's cries. This allows for a detailed analysis of your pet's behavior.

[0910] The cloud server then uses a generative AI model to determine the pet's emotions and health based on the analyzed data. The generative AI model compares the analysis results with past data to determine, for example, whether the pet is stressed or in pain.

[0911] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotions. The user can input text or voice data into the emotion engine, which then analyzes the user's emotional state. The system can also record the user's facial expressions using a camera and evaluate their emotional state using a facial recognition algorithm.

[0912] The cloud server correlates the pet's emotional and health status with the user's emotional status and uses a comprehensive assessment engine to determine the overall communication status. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress status to determine how the two are affecting each other.

[0913] The results are sent from the cloud server to the user's device and displayed via a dedicated application. Notifications are displayed, allowing users to check their pet's current condition and whether any abnormalities have been detected.

[0914] Specific examples

[0915] As a concrete example, consider the following scenario.

[0916] 1. The device records the cat's behavior: the camera captures the cat frequently licking its hind paws, and the microphone picks up the cat's short, sharp meows.

[0917] 2. The collected data is pre-processed on the device, trimmed, and denoised before being sent to the cloud server.

[0918] 3. The cloud server analyzes the video data using an object recognition algorithm to confirm the cat's licking of its hind paw, and analyzes the audio data using a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[0919] 4. The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[0920] 5. Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[0921] 6. The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0922] 7. The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is stressed. Attention is needed." The user can also check their own emotional state and take necessary action.

[0923] In this way, this system allows for a comprehensive understanding of the emotions and health status of both the pet and the user, enabling appropriate care and communication.

[0924] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0925] Step 1:

[0926] The user installs a pet-specific device in the pet's living space.

[0927] How it works: The user places a device containing a high-resolution camera and a sensitive microphone in a location where the pet is frequently present or resting.

[0928] Step 2:

[0929] The device captures your pet's behavior with a camera and records its cries with a microphone.

[0930] How it works: The camera captures pet movements in high resolution, and the microphone records pet sounds with high sensitivity, providing real-time data on pet behavior and sounds.

[0931] Input: Pet actions and sounds

[0932] Output: Raw video and audio data

[0933] Step 3:

[0934] Preprocessing is performed on the data collected by the terminal.

[0935] How it works: The device trims unnecessary parts of the video data and removes noise from the audio data. This preprocessing produces data that is easy to analyze.

[0936] Input: Raw video and audio data

[0937] Output: Pre-processed video and audio data

[0938] Step 4:

[0939] The terminal transmits the preprocessed data to the cloud server.

[0940] How it works: The device uses Wi-Fi or 4G / 5G to send pre-processed data to the cloud server, where it is encrypted to ensure security.

[0941] Input: Preprocessed video and audio data

[0942] Output: Data sent to the cloud server

[0943] Step 5:

[0944] The cloud server analyzes the received data.

[0945] How it works: The cloud server first applies an object recognition algorithm to the video data to identify the pet's body parts and analyze its behavioral patterns, and then applies a voice recognition algorithm to the audio data to analyze the characteristics of the pet's cries.

[0946] Input: Video and audio data sent to the cloud server

[0947] Output: Analyzed behavioral and audio data

[0948] Step 6:

[0949] The cloud server uses the generated AI model to determine the pet's emotions and health condition based on the analysis results.

[0950] How it works: The generative AI model compares your pet's behavior and vocalizations with past data to determine their emotions and health status. For example, it analyzes the frequency of vocalizations and behavioral patterns to determine whether your pet is stressed or in pain.

[0951] Input: Analyzed behavioral and audio data

[0952] Output: Pet's emotions and health status

[0953] Step 7:

[0954] Users input text or voice into the emotion engine, which analyzes their emotional state.

[0955] How it works: Users input text or voice to the emotion engine to recognize their emotional state, and a facial recognition algorithm is used to assess the user's emotional state from facial expression data captured by the camera.

[0956] Input: User text, voice, and facial expression data

[0957] Output: Analysis of the user's emotional state

[0958] Step 8:

[0959] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[0960] Specific operation: The cloud server uses a comprehensive judgment engine to correlate and analyze the emotional state of the pet and the user. For example, when the user is feeling stressed, the cloud server will also analyze the pet's stress state and evaluate the overall communication state.

[0961] Input: Pet's emotional and health status assessment results, analysis results of user's emotional state

[0962] Output: Overall communication status assessment

[0963] Step 9:

[0964] The cloud server sends the results to the user's device, and a dedicated application displays a notification.

[0965] Specific operation: The cloud server uses a transmission device to send the judgment result to the user's device. A dedicated application on the user's device displays a notification, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, a notification such as "Your cat is stressed. Attention is required" may be displayed.

[0966] Input: Overall communication status assessment result

[0967] Output: Notification displayed on the user's device

[0968] This series of steps creates a system that can comprehensively understand the emotions and health status of both the pet and the user, and provide appropriate care and communication.

[0969] (Application example 2)

[0970] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0971] In today's world, it is extremely important to properly understand pet health and emotional states and facilitate smooth communication with owners. However, many current systems simply record pet behavior and vocalizations, lacking the ability to comprehensively analyze pets' emotions and health. Furthermore, no systems provide appropriate feedback that takes into account the user's emotional state. Therefore, there is a need for a system that can analyze the state of both pets and users in real time to achieve optimal care and communication.

[0972] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a photographing means for recording the behavior of the pet, a sound collecting means for recording the sounds of the pet, a processing device means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the data processing device, an analysis device means for analyzing the data received by the data processing device, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a notification device means for notifying the user of the determination results, an emotion analysis means for analyzing the user's emotional state, and an association means for associating the pet's state with the user's emotional state. This makes it possible to analyze detailed data including the behavior and sounds of the pet and provide comprehensive advice that also takes the user's emotional state into consideration.

[0973] "Photographing means" refers to a camera device used to record the behavior of a pet.

[0974] The "audio collection means" is a microphone device for recording the sounds of pets.

[0975] The "processing device means" is a device for pre-processing data recorded by the imaging means and sound collection means.

[0976] "Communication means" refers to a wireless or wired communication device for transmitting pre-processed data to a data processing device.

[0977] "Analysis device means" is a general term for hardware and software for analyzing received data in a data processing device.

[0978] A "generative AI model" is a machine learning model that determines a pet's emotions and health status based on analysis results.

[0979] "Notification device means" is a device for notifying the user of the results determined by the generative AI model.

[0980] "Emotion analysis means" is a system for analyzing the emotional state of a user.

[0981] The "association means" is a mechanism for associating and analyzing the pet's emotions and health status with the user's emotional state.

[0982] This invention relates to a system that records pet behavior and sounds in real time and provides optimal feedback by taking into account the user's emotional state. The system includes three main components: a pet-specific device, a user terminal, and a cloud server.

[0983] Hardware and software used

[0984] Hardware:

[0985] Recording method: Camera equipment to record pet behavior

[0986] Sound collection means: microphone device for recording pet sounds

[0987] Processing Unit: A local device (e.g., Raspberry Pi) for preprocessing data.

[0988] Communication: Wi-Fi or 4G / 5G communication module for sending data to a cloud server

[0989] software:

[0990] TensorFlow: Deep learning object and speech recognition algorithms

[0991] OpenCV: Image processing library

[0992] Pydub: A library for preprocessing audio data

[0993] NLTK: A natural language processing library for analyzing user sentiment

[0994] System operation explanation

[0995] 1. Data Collection:

[0996] The pet device uses a camera to capture images of your pet's behavior and a microphone to record your pet's sounds, and this data is collected in real time.

[0997] 2. Data preprocessing:

[0998] Processing means within the terminal trims the collected video data and removes noise from the audio data.

[0999] 3. Data transmission:

[1000] The preprocessed data is sent to a cloud server via a communication method such as Wi-Fi or 4G / 5G.

[1001] 4. Data Analysis:

[1002] The cloud server analyzes the received data using an analysis device means, applying an object recognition algorithm using TensorFlow to the video data and a voice recognition algorithm also using TensorFlow to the voice data.

[1003] 5. Emotional and health status assessment:

[1004] Based on the analysis results, the generative AI model determines the pet's emotions and health status. For example, if the pet repeatedly behaves in a certain way or makes a certain sound, stress or health problems may be suspected.

[1005] 6. User Emotion Recognition:

[1006] By inputting text or voice from the user's device, the user's emotional state is analyzed by the emotion analysis means. The emotion analysis is performed using the NLTK library.

[1007] 7. Overall Judgment and Association:

[1008] The system analyzes the pet's emotions and health status in relation to the user's emotional state, and generates optimal feedback based on the results.

[1009] 8. Feedback and Notifications:

[1010] The cloud server sends the results to the user's device, where they are displayed via a dedicated application, allowing the user to receive feedback and take appropriate action.

[1011] Specific examples

[1012] For example, the device records the cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone records the cat's sharp meows. This data is sent to a cloud server and analyzed using object recognition and voice recognition algorithms. The generative AI model determines that the cat is likely feeling stressed. The user enters text into the emotion engine, which analyzes the cat's stress. The cloud server correlates this data and makes a comprehensive judgment, and a notification such as "Your cat is feeling stressed. Attention is required" is displayed on the user's device.

[1013] Example prompts to input to the generative AI model:

[1014] Pet Behavior: Cat repeatedly licks the same spot

[1015] Call: High-pitched, sharp call

[1016] User Emotion: Stressed

[1017] Determine your pet's emotional and health status.

[1018] This system makes it possible to perform detailed analysis of a pet's behavior and cries, and provide optimal actions that take into account the user's emotional state.

[1019] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1020] Step 1:

[1021] The device uses a camera and microphone to record your pet's movements and sounds in real time. This input data is used to generate video and audio files, which are then temporarily stored on the device.

[1022] Step 2:

[1023] The processing unit of the terminal performs preprocessing on the recorded video data and audio data, such as trimming the video data and adjusting the resolution, and removing noise from the audio data. As a result of the preprocessing, preprocessed video and audio files are generated.

[1024] Step 3:

[1025] The communication means transmits the preprocessed data to the cloud server. Specifically, the data is uploaded using Wi-Fi or 4G / 5G. After transmission, the data is stored in the cloud server.

[1026] Step 4:

[1027] The cloud server analyzes the received video data and audio data using an analysis device. An object recognition algorithm using TensorFlow is applied to the video data to extract the pet's behavioral characteristics. A voice recognition algorithm also using TensorFlow is applied to the audio data to extract the characteristics of the pet's cries. As a result of the analysis, data on the pet's behavioral characteristics and audio characteristics is generated.

[1028] Step 5:

[1029] The cloud server uses a generative AI model based on the analysis results to determine the pet's emotions and health condition. The feature data from the analysis results is used as input, and the generative AI model makes inferences based on it. The output is an evaluation of the pet's emotions and health condition.

[1030] Step 6:

[1031] Users input text or audio data into the emotion analysis tool to analyze their emotional state. The input data includes the user's text messages and audio files, and emotion analysis is performed using the NLTK library. The analysis results in an evaluation of the user's emotional state.

[1032] Step 7:

[1033] The cloud server performs a comprehensive assessment to correlate the pet's emotions and health status with the user's emotional state. The pet and user's emotional assessment results are used as input, and comprehensive feedback is generated based on the correlation. The output data includes a comprehensive diagnosis and recommended actions.

[1034] Step 8:

[1035] The cloud server sends the comprehensive feedback to the user's device and notifies the user through a dedicated application, allowing the user to check the pet's current status and recommended actions.

[1036] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1037] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1038] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1039] [Fourth embodiment]

[1040] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1041] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1042] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1043] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1044] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1045] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1046] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1047] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1048] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1049] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1050] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1051] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1052] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1053] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[1054] Specific operations of program processing

[1055] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[1056] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[1057] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[1058] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[1059] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[1060] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[1061] Specific examples

[1062] 1. Data collection

[1063] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[1064] 2. Data transmission and preprocessing

[1065] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[1066] 3. Data Analysis

[1067] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[1068] 4. Emotional and health status assessment

[1069] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[1070] 5. Feedback and Notifications

[1071] The cloud server sends the results to the user's device, and a dedicated application notifies them, saying, "Your cat is stressed. Attention is needed."

[1072] The introduction of this system will enable owners to monitor their pets' health and emotions in real time and provide appropriate care at an early stage, marking a major step towards providing a comfortable and healthy life for both pets and their owners.

[1073] The processing flow will be explained below.

[1074] Step 1:

[1075] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[1076] Step 2:

[1077] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[1078] Step 3:

[1079] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[1080] Step 4:

[1081] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[1082] Step 5:

[1083] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[1084] Step 6:

[1085] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[1086] Step 7:

[1087] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[1088] Step 8:

[1089] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[1090] Step 9:

[1091] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's mood and health status and take appropriate action if necessary.

[1092] Example 1

[1093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1094] A pet's behavior and vocalizations are important indicators for understanding its emotions and health status, but it is difficult for owners to constantly monitor them and respond appropriately. Furthermore, if owners are unable to notice abnormalities in their pets early, appropriate care may be delayed, increasing the risk of overlooking pet health problems. Given this situation, there is a demand for a system that can automatically monitor and analyze pet behavior and vocalizations in real time.

[1095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1096] In this invention, the server includes a recording means for recording the behavior of the pet, an audio acquisition means for recording the sounds of the pet, a processing means for preprocessing the recorded data, a transmission means for transmitting the preprocessed data via a network, an analysis means for analyzing the data received via the network, a generation algorithm model for determining the emotions and health condition of the pet based on the analysis results, and a notification means for notifying the user of the determination results. This makes it possible to automatically monitor and analyze the behavior and sounds of the pet in real time, accurately grasp the emotions and health condition of the pet, and provide necessary care early.

[1097] "Recording means" refers to a device or function for recording the behavior of a pet.

[1098] "Audio acquisition means" refers to a device or function for recording the sounds of pets.

[1099] "Processing means" refers to devices or functions for pre-processing the recorded data.

[1100] "Transmitting means" refers to a device or function for transmitting preprocessed data over a network.

[1101] "Analysis means" refers to a device or function for analyzing data received via a network.

[1102] A "generative algorithm model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[1103] "Notification means" refers to a device or function for notifying the user of the judgment result.

[1104] The present invention relates to a system that records pet behavior and sounds in real time and analyzes the data to determine the pet's emotions and health condition. This system allows pet owners to accurately understand the condition of their pets and provide appropriate care.

[1105] First, the user installs a pet-specific device in the pet's living space. This device has a built-in camera and microphone, which allows it to record the pet's behavior and sounds in real time. For example, the user can install the device in the living room and adjust the camera's orientation to capture the area where the pet is often found.

[1106] Next, the device captures footage of the pet's behavior with a camera and simultaneously records its meows with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. For example, the device may capture a video of a cat playing by a window and record its meows. At this time, unnecessary parts of the video are deleted and background noise is removed from the audio.

[1107] The preprocessed data is then sent by the device to a cloud server. This transmission is performed using wireless communication such as Wi-Fi or 4G / 5G. To send the preprocessed data to the cloud server, the device connects to the home Wi-Fi network, encrypts the data, and transmits it. After a few seconds, the cloud server receives the data.

[1108] The cloud server applies an object recognition algorithm to the video data it receives to identify the pet's body parts and analyze its behavioral patterns. It also applies a voice recognition algorithm to the audio data to analyze the characteristics of its meows. For example, the cloud server analyzes the video data to identify the behavior of a cat licking its hind paws, and analyzes the audio data to detect that short, sharp meows are a sign of stress.

[1109] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, the generative AI model compares a cat's past data with its current meows and behavioral patterns and determines that it is likely to be stressed.

[1110] The results of the assessment are sent to the user's device via a feedback mechanism. The user's device may be, for example, a smartphone or tablet, and the results are displayed through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, the cloud server may send a notification to the user's device stating, "Your cat is feeling stressed," and a pop-up notification will appear on the user's smartphone. When the user opens the app, detailed analysis results will be displayed.

[1111] For example, when a device records a cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone picks up short, sharp meows. This data is preprocessed and sent to a cloud server where it is analyzed using object recognition and voice recognition algorithms. The generative AI model then determines that the cat is stressed and sends a notification to the user's device saying, "Your cat is stressed. Attention is needed." In this way, users can monitor their pet's health in real time and provide appropriate care early.

[1112] Example prompts for generative AI models

[1113] Below is an example of a prompt sentence to input to the generative AI model.

[1114] "Your cat is frequently licking its hind paws and making short, sharp meowing noises. Use these data to assess your cat's emotional state and health."

[1115] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1116] Step 1: The user places the pet-specific device in the pet's living space. This device has a built-in camera and microphone. The input is the device's installation location and the pet's range of movement. The output is the device's readiness to record the pet's movements and sounds. The user places the device in a location in the living room where the pet is often found and turns it on. The user adjusts the camera's orientation to capture the pet at the optimal angle.

[1117] Step 2: The device captures the pet's behavior with a camera and simultaneously records its cries with a microphone. The input is the pet's behavior and cries. The output is the recorded video and audio data. For example, the device can record a video of the pet playing and simultaneously record the cries it makes.

[1118] Step 3: The device preprocesses the recorded data. This preprocessing includes trimming unnecessary parts of the video data and removing noise. The input is the recorded video data and audio data. The output is the preprocessed video data and audio data. The device removes unnecessary parts of the video data and removes background noise from the audio data.

[1119] Step 4: The device sends the preprocessed data to the cloud server. This is done using wireless communication such as Wi-Fi or 4G / 5G. The input is the preprocessed video and audio data. The output is the data sent to the cloud server. For example, the device sends the data via a home Wi-Fi network, and the cloud server receives it.

[1120] Step 5: The cloud server analyzes the received video data. It applies an object recognition algorithm to identify the pet's body parts and analyzes its behavioral patterns. The input is the video data sent to the cloud server. The output is the analysis results (behavioral patterns). The cloud server analyzes the video data and identifies, for example, the pet licking its hind paws.

[1121] Step 6: The cloud server analyzes the audio data. It applies a speech recognition algorithm to analyze the characteristics of the bird's calls. The input is the audio data sent to the cloud server. The output is the analysis result (the characteristics of the bird's calls). The cloud server analyzes the audio data and detects, for example, that the bird's calls are a sign of stress.

[1122] Step 7: The cloud server inputs the analysis results into the generative AI model. This generative AI model compares it with past data to determine the pet's emotions and health condition. The input is the analysis results of the video and audio data. The output is the determination of the pet's emotions and health condition. The generative AI model compares it with past data and determines that the pet is likely to be feeling stressed.

[1123] Step 8: The cloud server sends the judgment result to the user device via a feedback means. The user device displays the result through a dedicated application. The input is the judgment result of emotions and health status. The output is a notification displayed on the user device. The cloud server sends a notification to the user device saying "The cat is feeling stressed," and the user checks the notification on their smartphone.

[1124] This allows users to understand their pet's condition in real time and provide appropriate care early on.

[1125] (Application example 1)

[1126] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1127] Conventional pet management systems simply record pet behavior and sounds, and are unable to assess a pet's emotions or health status in real time. Additionally, there is a lack of information that allows owners to select appropriate care products and food based on their pet's health status. Furthermore, specific care suggestions based on analysis results are rarely provided, making it difficult for owners to provide detailed care for their pets.

[1128] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1129] In this invention, the server includes a video recording means for recording the behavior of the pet, an audio recording means for recording the sounds of the pet, an information processing means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the information server, an analysis means for analyzing the data received by the information server, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a suggestion means for suggesting appropriate care products and food based on the judgment results, and a notification means for notifying the user of the suggested content. This makes it possible to analyze the emotions and health condition of the pet and notify the user of specific care suggestions and products based on the pet's health condition.

[1130] "Video recording means" is a function for recording the behavior of pets using a camera.

[1131] The "audio recording means" is a function for recording the sounds of pets using a microphone.

[1132] The "information processing means" is a computer system for preprocessing the recorded video data and audio data.

[1133] The "communication means" is a wireless communication function for transmitting preprocessed data to an information server such as a cloud server.

[1134] The "analysis means" refers to a group of algorithms and programs for analyzing data received by the information server.

[1135] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on the analysis results.

[1136] The "suggestion means" is a function for suggesting care products and food suitable for pets based on the judgment results.

[1137] "Notification means" refers to a notification system for informing users of the proposed content.

[1138] An "object recognition algorithm" is an algorithm that analyzes pet behavior data and recognizes specific actions and movements.

[1139] The "voice recognition algorithm" is an algorithm for analyzing pet cry data and recognizing specific voice patterns.

[1140] The present invention relates to a system that records a pet's behavior and cries in real time, analyzes this data to determine the pet's emotions and health condition, and then suggests appropriate care products and food based on the results of this determination.

[1141] The system consists of the following main tools:

[1142] 1. Video recording means: Recording pet behavior using a camera. For example, a camera device specifically for pets is included.

[1143] 2. Audio recording means: Record the sounds your pet makes using a microphone. This also includes a dedicated audio recording device for pets.

[1144] 3. Information processing means: A computer system for preprocessing recorded video and audio data, such as noise removal and data trimming.

[1145] 4. Communication means: Equipped with wireless communication functions for transmitting preprocessed data to the information server. Specific examples include Wi-Fi and 4G / 5G communication modules.

[1146] 5. Analysis means: A group of algorithms and programs for analyzing data received by the information server. This includes object recognition algorithms and voice recognition algorithms.

[1147] 6. Generative AI model: An artificial intelligence model that determines a pet's emotions and health status based on the analysis results.

[1148] 7. Suggestion means: Based on the judgment results, it has a function to suggest care products and food suitable for pets.

[1149] 8. Notification method: A notification system to inform users of the proposed content. For example, a notification is sent to a smartphone.

[1150] The server performs the following main tasks:

[1151] 1. Analysis of video and audio data. Characterize pet behavior and sounds using object and audio recognition algorithms.

[1152] 2. The analysis results are input into a generative AI model to determine the pet's emotions and health status.

[1153] 3. Based on the judgment results, the appropriate care products and food for the pet are determined using a suggestion method.

[1154] 4. The proposal will be communicated to the user via notification means.

[1155] For example, if a pet frequently scratches its ears, the server can analyze the behavior and determine that this is a sign of an ear infection. Based on this determination, appropriate care products (e.g., ear cleaning wipes or anti-infection shampoo) can be suggested and notified to the user.

[1156] An example prompt is:

[1157] Prompt for the generative AI model:

[1158] You are an AI model that analyzes your pet's behavior and vocalizations to determine its emotions and health status. Analyze the following data to determine your pet's emotions and health status:

[1159] Action: {action}

[1160] Sound: {sound}

[1161] Based on the results, we will recommend the appropriate care products and food for your pet.

[1162] This system allows users to understand their pet's health condition and emotions in real time, enabling them to provide prompt and accurate physical and psychological care.

[1163] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1164] Step 1:

[1165] The terminal records the pet's behavior using a video recording means and records the pet's cries using an audio recording means. The input includes video data from the camera and audio data from the microphone. These data are temporarily stored for subsequent processing.

[1166] Step 2:

[1167] The device preprocesses the recorded video and audio data. Specific preprocessing operations include trimming the video data, removing noise, and adjusting the resolution. For audio data, noise removal and extraction of necessary frequency components are also performed. Clean video and audio data are generated as the output of the preprocessing.

[1168] Step 3:

[1169] The terminal transmits the preprocessed video and audio data to the information server using a communication means, with the input comprising the preprocessed data and the output being data packets received by the information server.

[1170] Step 4:

[1171] The server analyzes the received data using an analysis means. Specifically, an object recognition algorithm is applied to the video data to detect the specific behavior of the pet. A voice recognition algorithm is applied to the audio data to extract the characteristics and patterns of the pet's cries. The input includes the received data, and the output generates data on behavioral patterns and voice patterns, which are the analysis results.

[1172] Step 5:

[1173] The server inputs the analysis results into a generative AI model and performs calculations to determine the pet's emotions and health condition. The input includes the analysis result data, and the output generates a judgment result on the pet's emotions and health condition. For example, based on the analysis results, it may be determined that the pet is "stressed."

[1174] Step 6:

[1175] The server uses a suggestion tool to determine the appropriate pet care products and food based on the assessment results. The input includes the assessment results of the pet's emotions and health condition, and the output generates a list of suggested products. For example, foods and toys with a relaxing effect are suggested for a stressed pet.

[1176] Step 7:

[1177] The server notifies the user of the proposed content using a notification method. Specifically, it sends a notification to the user's smartphone or tablet. The input includes a list of proposed products, and the output is generated as notification information to be displayed on the user's device. The user can then receive this notification and check the details of the proposed products.

[1178] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1179] This invention combines a system that records pet behavior and sounds in real time and analyzes this data to determine the pet's emotions and health condition with an emotion engine that recognizes the user's emotions. This system allows for accurate understanding of the condition of both the pet and the owner, enabling more appropriate care and communication.

[1180] Specific operations of program processing

[1181] The first thing users do is to install a pet-specific device (hereafter referred to as the device) in their pet's living space. This device has a built-in camera and microphone, and can record their pet's behavior and sounds in real time.

[1182] The device captures images of your pet's behavior with a camera and simultaneously records its cries with a microphone. The resulting video and audio data is then pre-processed within the device, which includes trimming unnecessary parts of the video data and removing noise.

[1183] The preprocessed data is then sent by the device to the cloud server using wireless communication such as Wi-Fi or 4G / 5G. This allows the data to reach the cloud server from the device.

[1184] To analyze the data received by the cloud server, an object recognition algorithm is applied to the video data, which identifies the pet's body parts and analyzes its behavioral patterns. A voice recognition algorithm is also applied to the audio data, which analyzes the characteristics of the pet's cries.

[1185] Based on the analysis results, the cloud server inputs this data into a generative AI model, which compares it with past data to determine the pet's emotions and health status. For example, it analyzes the frequency of its cries and behavioral patterns to determine whether the pet is stressed or in pain.

[1186] Furthermore, the system incorporates an emotion engine that recognizes the user's emotions. The user inputs text and voice data into the emotion engine, which then analyzes the user's emotional state. It also uses a facial recognition algorithm to analyze the user's facial expressions and evaluate their emotional state.

[1187] The cloud server correlates the user's emotional state with the pet's state and determines the overall communication state. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress state to determine how the two are affecting each other.

[1188] The results of the assessment are sent to the user's device via a feedback mechanism. The user device, such as a smartphone or tablet, displays the results through a dedicated application. A notification is displayed, allowing the user to check the pet's current condition and whether any abnormalities have been detected.

[1189] Specific examples

[1190] 1. Data collection

[1191] The device records the cat's behavior: a camera captures the cat's frequent licking of its hind paws, and a microphone captures the cat's short, sharp meows.

[1192] 2. Data transmission and preprocessing

[1193] The collected data is pre-processed on the device, including trimming and noise removal, before being sent to a cloud server.

[1194] 3. Data Analysis

[1195] The cloud server analyzes the video data with an object recognition algorithm to identify the cat's licking behavior, and the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[1196] 4. Emotional and health status assessment

[1197] The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[1198] 5. Recognizing user emotions

[1199] Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[1200] 6. Overall Judgment

[1201] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[1202] 7. Feedback and Notifications

[1203] The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is feeling stressed. Attention is needed." Users can also check their own emotional state and take necessary action.

[1204] This system allows for a comprehensive understanding of the emotions and health status of both pets and users, enabling appropriate care and communication.

[1205] The processing flow will be explained below.

[1206] Step 1:

[1207] The device records your pet's behavior with a camera, which captures the behavior in high quality and saves it as video data.

[1208] Step 2:

[1209] The device records your pet's cries using a microphone, which also picks up surrounding sounds and stores them as audio data.

[1210] Step 3:

[1211] The device preprocesses the video and audio data, trimming unnecessary parts from the video data and removing noise from the audio data.

[1212] Step 4:

[1213] The device sends the preprocessed data to a cloud server via Wi-Fi or 4G / 5G.

[1214] Step 5:

[1215] The cloud server analyzes the received data using an object recognition algorithm, identifying the pet's body parts from the video data and analyzing its behavioral patterns.

[1216] Step 6:

[1217] The cloud server analyzes the audio data using a speech recognition algorithm, analyzing the frequency and characteristics of the calls to assess the animal's emotional state.

[1218] Step 7:

[1219] The cloud server inputs the analysis results into a generative AI model, which compares them with past data to determine the pet's emotions and health status.

[1220] Step 8:

[1221] The user inputs text or voice data into the emotion engine, which then analyzes the user's emotional state.

[1222] Step 9:

[1223] The user's facial expressions are captured by a camera, and the emotion engine uses a facial recognition algorithm to analyze the user's facial expression data, thereby assessing the user's emotional state.

[1224] Step 10:

[1225] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[1226] Step 11:

[1227] The cloud server sends the judgment result to the user's device, which is then displayed to the user in the form of a notification.

[1228] Step 12:

[1229] The user's device will receive a notification and display detailed results in a dedicated application, allowing the user to check their pet's emotional and health status, as well as their own emotional state, and take necessary actions.

[1230] Example 2

[1231] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1232] Pet owners face the challenge of understanding their pets' emotions and health status from their behavior and vocalizations. There is a need for a system that can properly recognize the emotional state of not only pets but also their owners, and grasp the overall state of communication with their pets. This will enable them to effectively manage the health and happiness of both pets and their owners, and provide appropriate care and advice.

[1233] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1234] In this invention, the server includes a camera for recording the pet's behavior, an audio collection device for recording the pet's cries, a processing device for preprocessing the recorded data, a communication device for transmitting the preprocessed data to a cloud server, an analysis device for analyzing the data received by the cloud server, a generative AI model for determining the pet's emotions and health status based on the analysis results, a transmission device for notifying the user of the determination results, an emotion engine for recognizing the user's emotions, and a comprehensive determination engine for correlating the user's emotional state with the pet's emotional state and making a comprehensive determination. This makes it possible to accurately grasp the pet's emotions and health status in real time and also evaluate the owner's emotional state. This allows for a comprehensive understanding of the pet's communication status and provides appropriate care and advice.

[1235] "Photographing means" refers to a device used to record the behavior of a pet.

[1236] An "audio collection means" is a device used to record the sounds of a pet.

[1237] "Processing equipment" is equipment used to pre-process recorded data.

[1238] A "communication device" is a device used to transmit pre-processed data to a cloud server.

[1239] An "analysis device" is a device used to analyze data received by a cloud server.

[1240] A "generative AI model" is an artificial intelligence model that determines a pet's emotions and health condition based on analysis results.

[1241] A "communication device" is a device used to notify the user of the determination result.

[1242] An "emotion engine" is a device or software used to recognize a user's emotions.

[1243] The "comprehensive judgment engine" is a device or software used to correlate and comprehensively judge the emotional state of the user and the emotional state of the pet.

[1244] This invention combines a system that determines a pet's emotions and health condition by recording and analyzing the pet's behavior and cries in real time with an emotion engine that recognizes the user's emotions. The following describes an embodiment of this system.

[1245] The first thing users need to do is to install a dedicated pet device in their pet's living space. This device is equipped with a high-resolution camera (photography) and a highly sensitive microphone (audio collection), allowing it to record the pet's behavior and sounds in real time.

[1246] The device then captures images of the pet's behavior with a camera and records its cries with a microphone. The resulting video and audio data is pre-processed within the device. This pre-processing includes trimming unnecessary parts of the video data and removing noise. This process produces data that is easy to analyze.

[1247] The preprocessed data is sent from the device to the cloud server using Wi-Fi or 4G / 5G wireless communication (communication equipment), allowing the data to reach the cloud server safely and quickly.

[1248] To analyze the received data, the cloud server first applies an object recognition algorithm to the video data. This algorithm identifies the pet's body parts and analyzes its behavioral patterns. It also applies a voice recognition algorithm to the audio data, analyzing the characteristics of the pet's cries. This allows for a detailed analysis of your pet's behavior.

[1249] The cloud server then uses a generative AI model to determine the pet's emotions and health based on the analyzed data. The generative AI model compares the analysis results with past data to determine, for example, whether the pet is stressed or in pain.

[1250] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotions. The user can input text or voice data into the emotion engine, which then analyzes the user's emotional state. The system can also record the user's facial expressions using a camera and evaluate their emotional state using a facial recognition algorithm.

[1251] The cloud server correlates the pet's emotional and health status with the user's emotional status and uses a comprehensive assessment engine to determine the overall communication status. For example, if the user is feeling stressed, the cloud server also analyzes the pet's stress status to determine how the two are affecting each other.

[1252] The results are sent from the cloud server to the user's device and displayed via a dedicated application. Notifications are displayed, allowing users to check their pet's current condition and whether any abnormalities have been detected.

[1253] Specific examples

[1254] As a concrete example, consider the following scenario.

[1255] 1. The device records the cat's behavior: the camera captures the cat's frequent licking of its hind paws, and the microphone captures the cat's short, sharp meows.

[1256] 2. The collected data is pre-processed on the device, trimmed, and denoised before being sent to the cloud server.

[1257] 3. The cloud server analyzes the video data with an object recognition algorithm to confirm the cat's licking behavior, and analyzes the audio data with a voice recognition algorithm to detect whether the cat's cries are a sign of stress.

[1258] 4. The generative AI model determines that the cat is likely stressed and experiencing discomfort in its hind legs.

[1259] 5. Users input text or voice into the emotion engine to analyze their emotional state, and facial recognition algorithms evaluate emotions based on facial expressions.

[1260] 6. The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[1261] 7. The cloud server sends the results to the user's device, and a dedicated application displays a notification such as, "Your cat is stressed. Attention is needed." The user can also check their own emotional state and take necessary action.

[1262] In this way, this system allows for a comprehensive understanding of the emotions and health status of both the pet and the user, enabling appropriate care and communication.

[1263] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1264] Step 1:

[1265] The user installs a pet-specific device in the pet's living space.

[1266] How it works: The user places a device containing a high-resolution camera and a sensitive microphone in a location where the pet is frequently present or resting.

[1267] Step 2:

[1268] The device captures your pet's behavior with a camera and records its cries with a microphone.

[1269] How it works: The camera captures pet movements in high resolution, and the microphone records pet sounds with high sensitivity, providing real-time data on pet behavior and sounds.

[1270] Input: Pet actions and sounds

[1271] Output: Raw video and audio data

[1272] Step 3:

[1273] Preprocessing is performed on the data collected by the terminal.

[1274] How it works: The device trims unnecessary parts of the video data and removes noise from the audio data. This preprocessing produces data that is easy to analyze.

[1275] Input: Raw video and audio data

[1276] Output: Pre-processed video and audio data

[1277] Step 4:

[1278] The terminal transmits the preprocessed data to the cloud server.

[1279] How it works: The device uses Wi-Fi or 4G / 5G to send pre-processed data to the cloud server, where it is encrypted to ensure security.

[1280] Input: Preprocessed video and audio data

[1281] Output: Data sent to the cloud server

[1282] Step 5:

[1283] The cloud server analyzes the received data.

[1284] How it works: The cloud server first applies an object recognition algorithm to the video data to identify the pet's body parts and analyze its behavioral patterns, and then applies a voice recognition algorithm to the audio data to analyze the characteristics of the pet's cries.

[1285] Input: Video and audio data sent to the cloud server

[1286] Output: Analyzed behavioral and audio data

[1287] Step 6:

[1288] The cloud server uses the generated AI model to determine the pet's emotions and health condition based on the analysis results.

[1289] How it works: The generative AI model compares your pet's behavior and vocalizations with past data to determine their emotions and health status. For example, it analyzes the frequency of vocalizations and behavioral patterns to determine whether your pet is stressed or in pain.

[1290] Input: Analyzed behavioral and audio data

[1291] Output: Pet's emotions and health status

[1292] Step 7:

[1293] Users input text or voice into the emotion engine, which analyzes their emotional state.

[1294] How it works: Users input text or voice to the emotion engine to recognize their emotional state, and a facial recognition algorithm is used to assess the user's emotional state from facial expression data captured by the camera.

[1295] Input: User text, voice, and facial expression data

[1296] Output: Analysis of the user's emotional state

[1297] Step 8:

[1298] The cloud server correlates the pet's emotional and health status with the user's emotional state and makes a comprehensive judgment.

[1299] Specific operation: The cloud server uses a comprehensive judgment engine to correlate and analyze the emotional state of the pet and the user. For example, when the user is feeling stressed, the cloud server will also analyze the pet's stress state and evaluate the overall communication state.

[1300] Input: Pet's emotional and health status assessment results, analysis results of user's emotional state

[1301] Output: Overall communication status assessment

[1302] Step 9:

[1303] The cloud server sends the results to the user's device, and a dedicated application displays a notification.

[1304] Specific operation: The cloud server uses a transmission device to send the judgment result to the user's device. A dedicated application on the user's device displays a notification, allowing the user to check the pet's current condition and whether there are any abnormalities. For example, a notification such as "Your cat is stressed. Attention is required" may be displayed.

[1305] Input: Overall communication status assessment result

[1306] Output: Notification displayed on the user's device

[1307] This series of steps creates a system that can comprehensively understand the emotions and health status of both the pet and the user, and provide appropriate care and communication.

[1308] (Application example 2)

[1309] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1310] In today's world, it is extremely important to properly understand pet health and emotional states and facilitate smooth communication with owners. However, many current systems simply record pet behavior and vocalizations, lacking the ability to comprehensively analyze pets' emotions and health. Furthermore, no systems provide appropriate feedback that takes into account the user's emotional state. Therefore, there is a need for a system that can analyze the state of both pets and users in real time to achieve optimal care and communication.

[1311] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a photographing means for recording the behavior of the pet, a sound collecting means for recording the sounds of the pet, a processing device means for preprocessing the recorded data, a communication means for transmitting the preprocessed data to the data processing device, an analysis device means for analyzing the data received by the data processing device, a generative AI model for determining the emotions and health condition of the pet based on the analysis results, a notification device means for notifying the user of the determination results, an emotion analysis means for analyzing the user's emotional state, and an association means for associating the pet's state with the user's emotional state. This makes it possible to analyze detailed data including the behavior and sounds of the pet and provide comprehensive advice that also takes the user's emotional state into consideration.

[1312] "Photographing means" refers to a camera device used to record the behavior of a pet.

[1313] The "audio collection means" is a microphone device for recording the sounds of pets.

[1314] The "processing device means" is a device for pre-processing data recorded by the imaging means and sound collection means.

[1315] "Communication means" refers to a wireless or wired communication device for transmitting pre-processed data to a data processing device.

[1316] "Analysis device means" is a general term for hardware and software for analyzing received data in a data processing device.

[1317] A "generative AI model" is a machine learning model that determines a pet's emotions and health status based on analysis results.

[1318] "Notification device means" is a device for notifying the user of the results determined by the generative AI model.

[1319] "Emotion analysis means" is a system for analyzing the emotional state of a user.

[1320] The "association means" is a mechanism for associating and analyzing the pet's emotions and health status with the user's emotional state.

[1321] This invention relates to a system that records pet behavior and sounds in real time and provides optimal feedback by taking into account the user's emotional state. The system includes three main components: a pet-specific device, a user terminal, and a cloud server.

[1322] Hardware and software used

[1323] Hardware:

[1324] Recording method: Camera equipment to record pet behavior

[1325] Sound collection means: microphone device for recording pet sounds

[1326] Processing Unit: A local device (e.g., Raspberry Pi) for preprocessing data.

[1327] Communication: Wi-Fi or 4G / 5G communication module for sending data to a cloud server

[1328] software:

[1329] TensorFlow: Deep learning object and speech recognition algorithms

[1330] OpenCV: Image processing library

[1331] Pydub: A library for preprocessing audio data

[1332] NLTK: A natural language processing library for analyzing user sentiment

[1333] System operation explanation

[1334] 1. Data Collection:

[1335] The pet device uses a camera to capture images of your pet's behavior and a microphone to record your pet's sounds, and this data is collected in real time.

[1336] 2. Data preprocessing:

[1337] Processing means within the terminal trims the collected video data and removes noise from the audio data.

[1338] 3. Data transmission:

[1339] The preprocessed data is sent to a cloud server via a communication method such as Wi-Fi or 4G / 5G.

[1340] 4. Data Analysis:

[1341] The cloud server analyzes the received data using an analysis device means. An object recognition algorithm using TensorFlow is applied to the video data, and a voice recognition algorithm also using TensorFlow is applied to the voice data.

[1342] 5. Emotional and health status assessment:

[1343] Based on the analysis results, the generative AI model determines the pet's emotions and health status. For example, if the pet repeatedly behaves in a certain way or makes a certain sound, stress or health problems may be suspected.

[1344] 6. User Emotion Recognition:

[1345] By inputting text or voice from the user's device, the user's emotional state is analyzed by the emotion analysis means. The emotion analysis is performed using the NLTK library.

[1346] 7. Overall Judgment and Association:

[1347] The system analyzes the pet's emotions and health status in relation to the user's emotional state, and generates optimal feedback based on the results.

[1348] 8. Feedback and Notifications:

[1349] The cloud server sends the results to the user's device, where they are displayed via a dedicated application, allowing the user to receive feedback and take appropriate action.

[1350] Specific examples

[1351] For example, the device records the cat's behavior, the camera captures the cat frequently licking its hind paws, and the microphone records the cat's sharp meows. This data is sent to a cloud server and analyzed using object recognition and voice recognition algorithms. The generative AI model determines that the cat is likely feeling stressed. The user enters text into the emotion engine, which analyzes the cat's stress. The cloud server correlates this data and makes a comprehensive judgment, and a notification such as "Your cat is feeling stressed. Attention is required" is displayed on the user's device.

[1352] Example prompts to input to the generative AI model:

[1353] Pet Behavior: Cat repeatedly licks the same spot

[1354] Call: High-pitched, sharp call

[1355] User Emotion: Stressed

[1356] Determine your pet's emotional and health status.

[1357] This system makes it possible to perform detailed analysis of a pet's behavior and cries, and provide optimal actions that take into account the user's emotional state.

[1358] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1359] Step 1:

[1360] The device uses a camera and microphone to record your pet's movements and sounds in real time. This input data is used to generate video and audio files, which are then temporarily stored on the device.

[1361] Step 2:

[1362] The processing unit of the terminal performs preprocessing on the recorded video data and audio data, such as trimming the video data and adjusting the resolution, and removing noise from the audio data. As a result of the preprocessing, preprocessed video and audio files are generated.

[1363] Step 3:

[1364] The communication means transmits the preprocessed data to the cloud server. Specifically, the data is uploaded using Wi-Fi or 4G / 5G. After transmission, the data is stored in the cloud server.

[1365] Step 4:

[1366] The cloud server analyzes the received video data and audio data using an analysis device. An object recognition algorithm using TensorFlow is applied to the video data to extract the pet's behavioral characteristics. A voice recognition algorithm also using TensorFlow is applied to the audio data to extract the characteristics of the pet's cries. As a result of the analysis, data on the pet's behavioral characteristics and audio characteristics is generated.

[1367] Step 5:

[1368] The cloud server uses a generative AI model based on the analysis results to determine the pet's emotions and health condition. The feature data from the analysis results is used as input, and the generative AI model makes inferences based on it. The output is an evaluation of the pet's emotions and health condition.

[1369] Step 6:

[1370] Users input text or audio data into the emotion analysis tool to analyze their emotional state. The input data includes the user's text messages and audio files, and emotion analysis is performed using the NLTK library. The analysis results in an evaluation of the user's emotional state.

[1371] Step 7:

[1372] The cloud server performs a comprehensive assessment to correlate the pet's emotions and health status with the user's emotional state. The pet and user's emotional assessment results are used as input, and comprehensive feedback is generated based on the correlation. The output data includes a comprehensive diagnosis and recommended actions.

[1373] Step 8:

[1374] The cloud server sends the comprehensive feedback to the user's device and notifies the user through a dedicated application, allowing the user to check the pet's current status and recommended actions.

[1375] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1376] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1377] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1378] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1379] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1380] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1381] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1382] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1383] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1384] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1385] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1386] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1387] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1388] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1389] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1390] The hardware resource for executing a specific process can be any of the following processors: A CPU is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A dedicated electrical circuit, such as a field-programmable gate array (FPGA), a programmable logic device (PLD), or an application-specific integrated circuit (ASIC), is a processor with a circuit configuration specifically designed to execute a specific process. Each processor has built-in or connected memory, and uses the memory to execute the specific process.

[1391] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1392] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1393] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1394] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1395] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1396] The following is further disclosed regarding the above embodiment.

[1397] (Claim 1)

[1398] camera means for recording the behavior of the pet;

[1399] a microphone means for recording the sounds of pets;

[1400] terminal means for pre-processing the recorded data;

[1401] a transmitting means for transmitting the preprocessed data to a cloud server;

[1402] an analysis means for analyzing the received data in the cloud server;

[1403] A generative AI model that determines pets' emotions and health status based on the analysis results, and

[1404] A feedback means for notifying the user of the judgment result;

[1405] A system including:

[1406] (Claim 2)

[1407] 2. The system according to claim 1, wherein the analyzing means includes means for analyzing the behavioral data of the pet using an object recognition algorithm, and means for analyzing the cries of the pet using a voice recognition algorithm.

[1408] (Claim 3)

[1409] 2. The system according to claim 1, wherein the camera means and the microphone means are means for recording the behavior and sounds of a pet in real time.

[1410] (Claim 4)

[1411] 2. The system according to claim 1, wherein the feedback means includes means for notifying a user terminal of the analysis result.

[1412] (Claim 5)

[1413] 2. The system of claim 1, wherein the generative AI model includes means for determining the current state of the pet in relation to past data.

[1414] "Example 1"

[1415] (Claim 1)

[1416] a recording means for recording the behavior of the pet;

[1417] an audio capture means for recording the sounds of pets;

[1418] processing means for pre-processing the recorded data;

[1419] transmitting means for transmitting the preprocessed data over a network;

[1420] an analysis means for analyzing data received via a network;

[1421] A generative algorithm model that determines the pet's emotions and health status based on the analysis results, and

[1422] a notification means for notifying the user of the determination result;

[1423] A system including:

[1424] (Claim 2)

[1425] 2. The system according to claim 1, wherein the analyzing means includes a function of analyzing the behavior data of the pet using an object recognition algorithm and a function of analyzing the cries of the pet using a voice recognition algorithm.

[1426] (Claim 3)

[1427] 2. The system according to claim 1, wherein the recording means and the sound acquisition means have a function of recording the behavior and sounds of a pet in real time.

[1428] "Application Example 1"

[1429] (Claim 1)

[1430] a video recording means for recording the behavior of the pet;

[1431] an audio recording means for recording the sounds of the pet;

[1432] information processing means for pre-processing the recorded data;

[1433] a communication means for transmitting the preprocessed data to an information server;

[1434] an analysis means for analyzing the received data in the information server;

[1435] A generative AI model that determines pets' emotions and health status based on the analysis results, and

[1436] A suggestion means for suggesting appropriate care products and foods based on the judgment results;

[1437] a notification means for notifying the user of the proposed content;

[1438] A system including:

[1439] (Claim 2)

[1440] 2. The system according to claim 1, wherein the analyzing means includes means for analyzing the behavioral data of the pet using an object recognition algorithm, and means for analyzing the cries of the pet using a voice recognition algorithm.

[1441] (Claim 3)

[1442] 2. The system according to claim 1, wherein the video recording means and the audio recording means are means for recording the behavior and sounds of a pet in real time.

[1443] "Example 2: Combining Emotion Engines"

[1444] (Claim 1)

[1445] A photographing means for recording the behavior of a pet;

[1446] an audio collection means for recording the sounds of pets;

[1447] a processing device for pre-processing the recorded data;

[1448] a communication device that transmits the preprocessed data to a cloud server;

[1449] an analysis device that analyzes the received data in the cloud server;

[1450] A generative AI model that determines pets' emotions and health status based on the analysis results, and

[1451] A communication device that notifies the user of the determination result;

[1452] an emotion engine for recognizing user emotions;

[1453] a comprehensive judgment engine that correlates the emotional state of the user with the emotional state of the pet and makes a comprehensive judgment;

[1454] A system including:

[1455] (Claim 2)

[1456] 2. The system according to claim 1, wherein the analysis device includes means for analyzing the behavioral data of the pet using an object recognition algorithm, and means for analyzing the cries of the pet using a voice recognition algorithm.

[1457] (Claim 3)

[1458] 2. The system according to claim 1, wherein the photographing means and the sound collecting means are means for recording the behavior and sounds of a pet in real time.

[1459] "Application example 2 when combining emotion engines"

[1460] (Claim 1)

[1461] A photographing means for recording the behavior of a pet;

[1462] an audio collection means for recording the sounds of pets;

[1463] processor means for pre-processing the recorded data;

[1464] communication means for transmitting the preprocessed data to a data processing device;

[1465] analyzer means for analyzing the received data in the data processing device;

[1466] A generative AI model that determines pets' emotions and health status based on the analysis results, and

[1467] a notification device means for notifying a user of the determination result;

[1468] emotion analysis means for analyzing the emotional state of a user;

[1469] an association means for associating the state of the pet with the emotional state of the user;

[1470] A system including:

[1471] (Claim 2)

[1472] 2. The system according to claim 1, wherein the analysis device means includes means for analyzing the behavioral data of the pet using an object recognition algorithm, and means for analyzing the cries of the pet using a voice recognition algorithm.

[1473] (Claim 3)

[1474] 2. The system according to claim 1, wherein the photographing means and the sound collecting means are means for recording the behavior and sounds of a pet in real time. [Explanation of symbols]

[1475] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. camera means for recording the behavior of the pet; a microphone means for recording the sounds of pets; terminal means for pre-processing the recorded data; a transmitting means for transmitting the preprocessed data to a cloud server; an analysis means for analyzing the received data in the cloud server; A generative AI model that determines pets' emotions and health status based on the analysis results, and A feedback means for notifying the user of the judgment result; A system including:

2. 2. The system according to claim 1, wherein the analyzing means includes means for analyzing the behavioral data of the pet using an object recognition algorithm, and means for analyzing the cries of the pet using a voice recognition algorithm.

3. 2. The system according to claim 1, wherein said camera means and said microphone means are means for recording the behavior and sounds of a pet in real time.

4. 2. The system according to claim 1, wherein the feedback means includes means for notifying a user terminal of the analysis result.

5. The system of claim 1 , wherein the generative AI model includes means for determining the current state of the pet in relation to past data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A