system

The system uses AI to analyze pet cries and characteristics for comprehensive understanding of pet conditions, facilitating effective communication and timely responses to health issues.

JP7785887B2Active Publication Date: 2025-12-15SOFTBANK GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024161845
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-09-19
Filing Date
2024-09-19
Publication Date
2025-12-15
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

Existing methods for understanding pet conditions from their cries are limited and one-sided, making it difficult for humans to accurately comprehend their health and emotional states.

Method used

A system that collects audio data and animal characteristics using AI to facilitate communication between humans and pets, enabling accurate understanding of pet conditions and emotional states through data analysis and conversion of cries into understandable messages.

Benefits of technology

Enhances the ability to comprehend pet conditions and emotional states, allowing for timely responses to health issues and improved communication between humans and pets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785887000001
    Figure 0007785887000001
  • Figure 0007785887000002
    Figure 0007785887000002
  • Figure 0007785887000003
    Figure 0007785887000003
Patent Text Reader

Abstract

To provide a system for promoting communication between a pet and an owner.SOLUTION: A system includes: means for receiving voice data; means for collecting characteristics of a pet; means for identifying an animal type; means for collecting characteristics data on general animals; means for generating a prompt sentence for instructing generation of information on at least one of a state and a request of the pet on the basis of the voice data, the characteristics of the pet, the identified animal type, and characteristics data; means for generating information about at least one of the state and the request of the pet on the basis of the generated prompt sentence and a generative AI model; and means for generating a prompt sentence for instructing to translate an instruction to the pet into a form understandable to the pet on the basis of a feeling state of a user identified using a feeling engine, and means for translating and outputting the instruction spoken to the pet into the form understandable to the pet using the generated prompt sentence and the generative AI model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Traditionally, pet owners have tried to understand their pets' condition from their cries, but cries alone are limited and it is difficult to accurately understand their condition. In addition, communication between humans and pets is one-sided, making it difficult to understand the pet's intentions. [Means for solving the problem]

[0005] This invention collects audio data (animal cries), characteristics of pet animals (personality, cries, etc.), animal species, and general animal characteristics, and uses this data to understand the condition of pets from their cries. Furthermore, it utilizes AI to enable conversations between people and their pets, and understands the pet's health and emotional state. This allows for a more accurate understanding of the pet's condition and improves communication between people and pets. [Brief explanation of the drawings]

[0006] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 2 is a sequence diagram showing a flow of processing in the data processing system according to the first embodiment of the first form example. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Embodiment 1. [Figure 13] FIG. 10 is a sequence diagram showing a processing flow of a data processing system in a second embodiment of the second form example. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Embodiment Example 2. [Figure 15]FIG. 10 is a sequence diagram showing the flow of processing in a data processing system according to a third embodiment of the third embodiment. [Figure 16] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Embodiment 3. [Figure 17] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the first embodiment of the first form example when an emotion engine is combined. [Figure 18] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. [Figure 19] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the second embodiment of the second form example when an emotion engine is combined. [Figure 20] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Form Example 2 when an emotion engine is combined. [Figure 21] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the third embodiment of the third form example when an emotion engine is combined. [Figure 22] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Form Example 3 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0007] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0008] First, the terms used in the following description will be explained.

[0009] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (TENSOR PROCESSING UNIT (registered trademark)).

[0010] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0011] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0012] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0014] [First embodiment]

[0015] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0016] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0017] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0018] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0019] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0021] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0022] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0023] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0024] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0025] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0026] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[0027] "Example 1"

[0028] In one embodiment of the present invention, pet cries are collected using an audio data collection device such as a microphone. The collected audio data is sent to an audio analysis device and identified as the cry of a specific animal. In addition, the characteristics of the pet (personality, cries, etc.) are collected using input from the owner or sensors that observe the pet's behavior. The type of animal is identified using input from the owner or image recognition technology. General animal characteristic data is obtained from a database or the like. This data is analyzed by AI to understand the pet's condition.

[0029] "Example 2"

[0030] Another embodiment of the present invention provides a system that utilizes AI to enable conversations between people and pets. Specifically, AI analyzes a pet's cries and converts them into words that humans can understand. Conversely, AI analyzes human language and converts it into cries and actions that pets can understand. For example, if a pet expresses its hunger state through cries, the AI ​​analyzes it and conveys the message "your pet is hungry" to the owner. Also, if the owner says "let's play," the AI ​​analyzes it and conveys the message in a form that the pet can understand.

[0031] "Example 3"

[0032] As a further embodiment of the present invention, a system for understanding the health and emotional state of a pet is provided. Specifically, AI analyzes the health and emotional state of a pet from its sounds and behavior. For example, if a pet is suffering from an illness, its sounds and behavior may be different from normal. The AI ​​detects these changes and issues a warning to the owner. Also, if a pet is happy, its sounds and behavior may become more active. The AI ​​detects these changes and conveys the owner's joy.

[0033] The processing flow of each embodiment will be described below.

[0034] "Example 1"

[0035] Step 1: Collect your pet's cries using a sound data collection device such as a microphone.

[0036] Step 2: The collected audio data is sent to an audio analyzer to identify it as the sound of a specific animal.

[0037] Step 3: Collect information about pet characteristics (personality, vocalizations, etc.) using input from owners and sensors that observe pet behavior.

[0038] Step 4: Identify the animal's species using input from the owner and image recognition technology.

[0039] Step 5: Obtain general animal characteristic data from a database, etc.

[0040] Step 6: This data is analyzed using AI to understand the pet's condition.

[0041] "Example 2"

[0042] Step 1: AI analyzes your pet's cries and converts them into words that humans can understand. Step 2: AI analyzes human words and converts them into cries and actions that your pet can understand.

[0043] Step 3: For example, if a pet expresses its hunger state through a cry, the AI ​​will analyze it and convey the message to the owner that "your pet is hungry."

[0044] Step 4: Also, if the owner says "Let's play," the AI ​​will analyze it and convey the message in a way that the pet can understand.

[0045] "Example 3"

[0046] Step 1: AI analyzes your pet's health and emotional state based on its sounds and behavior.

[0047] Step 2: For example, if a pet is suffering from an illness, its sounds and behavior may be different from normal. AI can detect these changes and alert the owner.

[0048] Step 3: Also, if your pet is happy, its vocalizations and behavior may become more active. AI will detect these changes and communicate its happiness to its owner.

[0049] Example 1

[0050] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0051] Conventional pet condition monitoring systems mainly collect and analyze voice data and animal characteristic data separately, making it difficult to integrate this data to comprehensively understand the pet's condition.In addition, there was a lack of means to accurately understand the pet's health and emotional state, making it difficult for owners to appropriately understand and respond to their pet's condition.

[0052] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0053] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data, means for analyzing the voice data, means for collecting characteristics of pet animals, means for identifying the type of animal, means for collecting general animal characteristic data, and means for understanding the condition of pet animals from their cries based on the collected data. This makes it possible to comprehensively understand the condition of pets by integrating the voice data and animal characteristic data.

[0054] "Audio data" refers to data in which audio signals such as animal cries are recorded in digital format.

[0055] "Collecting means" refers to a device or method for acquiring audio data and animal characteristic data.

[0056] "Transmitting means" refers to a device or method for transferring collected data to another device or server.

[0057] An "analyzing means" is a device or method for processing collected data and extracting specific information.

[0058] "Animal characteristics" refers to individual attributes of pet animals, such as their personalities and sounds.

[0059] "Means for identification" refers to devices or methods for determining the type or characteristics of an animal.

[0060] "General animal characteristic data" refers to data such as characteristics and behavioral patterns common to a particular animal species.

[0061] "Pet condition" refers to the overall condition of the pet, including its health and emotional state.

[0062] "AI" is a technology that uses artificial intelligence to analyze data and make specific judgments and predictions.

[0063] "Means for realizing conversation" refers to devices or methods that allow people and pets to communicate.

[0064] "Health Status" refers to information about the physical health of your pet.

[0065] "Emotional state" refers to a pet's psychological state or mood.

[0066] The present invention is a system for collecting sounds of pets and analyzing the data to understand the state of the pets. A specific embodiment of this system will be described below.

[0067] First, a user uses the smartphone's built-in microphone or a dedicated external microphone to collect the sound of their pet. For example, the user starts a recording app on their smartphone and records the sound of their pet. This audio data is then sent to the server by the device. Specifically, the device sends the audio data to the server's API endpoint using an HTTP POST request.

[0068] The server analyzes the received audio data using a voice analyzer (for example, Google® Cloud Speech-to-Text API). The server sends the audio data to the API, analyzes the returned text data, and identifies it as the sound of a specific animal.

[0069] Next, the user enters the characteristics of their pet (personality, sounds, etc.) into the app. For example, the user might enter information such as "Personality: Active, Sound: High-pitched" into the app's form. The app also uses sensors connected to the device (such as an accelerometer or camera) to observe the pet's behavior and collect data. The device periodically sends the data from the sensors to a server.

[0070] To identify the type of animal, the user can input the type of pet into the app, or the device can use image recognition technology (such as TENSORFLOW (registered trademark) or OpenCV) to identify the type of animal. For example, the user can input "dog" or the device can send an image taken with its camera to the server, which then uses image recognition technology to identify it as a "dog."

[0071] The server retrieves general animal characteristics data from a database (e.g., AWS® RDS or Google BigQuery). The server executes SQL queries to retrieve the required data from the database.

[0072] Finally, the server analyzes the collected voice data, pet characteristic data, animal species data, and general animal characteristic data using AI (for example, TensorFlow or PyTorch). The server integrates this data and inputs it into an AI model to understand the pet's condition. For example, the server analyzes data such as "the dog is barking loudly and is behaving actively" and determines that "the dog is excited."

[0073] Examples and prompts

[0074] As a concrete example, the bark of a user's dog is collected using the microphone on their smartphone and sent to a server. The server analyzes the bark using the Google Cloud Speech-to-Text API and identifies it as a dog's bark. The user enters the dog's personality and characteristics of its bark into the app and observes the dog's behavior using a camera connected to the device. The server retrieves data on general dog characteristics from AWS RDS, analyzes the data using TensorFlow, and understands the dog's condition.

[0075] An example of a prompt sentence is as follows:

[0076] "Please analyze my dog's barks to understand his condition. My dog ​​has an active personality and barks a lot. Please analyze the following audio data."

[0077] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0078] Step 1:

[0079] A user collects the sounds of their pets.

[0080] Specifically, the user launches a recording app on their smartphone and records the sound of their pet's cries.

[0081] Input: Pet noises

[0082] Output: Audio data file (e.g. .wav format)

[0083] Step 2:

[0084] The device sends the collected voice data to the server.

[0085] Specifically, the device sends audio data to the server's API endpoint using an HTTP POST request.

[0086] Input: Audio data file

[0087] Output: Audio data sent to the server

[0088] Step 3:

[0089] The server analyzes the received audio data.

[0090] Specifically, the server uses a voice analysis device (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text data and identify it as the sound of a specific animal.

[0091] Input: Audio data sent to the server

[0092] Output: Parsed text data (e.g. "dog barking")

[0093] Step 4:

[0094] Collect characteristics of the animals that the user keeps.

[0095] Specifically, the user enters information such as "personality: active, bark: high-pitched" into a form on the app. The app also uses sensors connected to the device (e.g., accelerometer and camera) to observe the pet's behavior and collect data.

[0096] Input: User-entered animal feature data, behavior data from sensors

[0097] Output: Collected animal trait data

[0098] Step 5:

[0099] Identify types of animals.

[0100] Specifically, the user enters the type of animal they own into the app, or the device uses image recognition technology (e.g., TensorFlow or OpenCV) to identify the type of animal.

[0101] Input: Animal species data or image data entered by the user

[0102] Output: Identified animal type data (e.g. "dog")

[0103] Step 6:

[0104] The server obtains general animal characteristic data.

[0105] Specifically, the server executes an SQL query from a database (e.g., AWS RDS or Google BigQuery) to retrieve the required data.

[0106] Input: Animal species data

[0107] Output: General animal traits data

[0108] Step 7:

[0109] The server analyzes the collected data and determines the pet's condition.

[0110] Specifically, the server analyzes collected voice data, animal characteristics data, animal species data, and general animal characteristics data using AI (e.g., TensorFlow or PyTorch). The server integrates this data and inputs it into an AI model to determine the pet's condition.

[0111] Input: Audio data, animal characteristics data, animal species data, general animal characteristics data

[0112] Output: Analysis result (e.g. "The dog is excited")

[0113] (Application example 1)

[0114] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0115] It is difficult for pet owners to monitor their pets' abnormal behavior and health status in real time while they are out or not watching over them. In addition, there is a lack of means to detect abnormalities in pets early and respond quickly, making it difficult to ensure the health and safety of pets.

[0116] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0117] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying the type of animal, means for collecting general animal characteristic data, means for understanding the pet's condition from its cries based on the collected data, and means for sending a notification when an abnormality is detected. This makes it possible to understand the pet's abnormal behavior and health condition in real time, and to quickly notify the owner when an abnormality is detected.

[0118] "Audio data" refers to information collected in digital format from audio signals such as animal cries.

[0119] "Collecting means" refers to a device or method for acquiring audio data and animal characteristic data.

[0120] "Animal characteristics" refers to information that refers to the animal's unique characteristics, such as its personality and cries.

[0121] "Means for identification" refers to devices or methods for determining the type or characteristics of an animal.

[0122] "General animal characteristic data" refers to information such as personality and behavioral patterns common to a specific animal species.

[0123] "Means of understanding" refers to devices and methods used to analyze collected data and understand the condition of animals.

[0124] "Means for detecting abnormalities" refers to devices or methods for identifying abnormalities in animal behavior or vocalizations and detecting the occurrence of an abnormality.

[0125] The "means for sending a notification" refers to a device or method for conveying information to the owner or other person when an abnormality is detected.

[0126] The system for carrying out this invention collects and analyzes the sounds of pets to understand their condition, and if an abnormality is detected, sends a notification to the pet owner. A specific embodiment of the system will be described below.

[0127] Hardware Configuration

[0128] The system includes the following hardware:

[0129] Microphone: Used to collect pet sounds.

[0130] Smartphone or smart glasses: Used to analyze collected voice data and detect anomalies.

[0131] Server: Analyzes voice data and sends notifications.

[0132] Software Configuration

[0133] The system includes the following software:

[0134] TensorFlow: Runs AI models to analyze audio data.

[0135] sounddevice: Used to collect audio data.

[0136] smtplib: Used to send notifications when an anomaly is detected.

[0137] Data processing and calculation

[0138] 1. Audio data collection: The device (smartphone or smart glasses) uses a microphone to collect the pet's cries. The collected audio data is stored digitally.

[0139] 2. Audio data analysis: The server inputs the collected audio data into the TensorFlow model for analysis. Based on the analysis results, the pet's condition is determined.

[0140] 3. Anomaly detection: The server determines whether an anomaly has been detected based on the analysis results. If an anomaly is detected, it sends a notification to the owner.

[0141] 4. Sending notifications: The server uses smtplib to send notifications to the pet owner, including information about the pet's abnormal behavior and health status.

[0142] Specific examples

[0143] For example, if a pet makes an unusual noise while the owner is out, the system will collect and analyze the sound. If an abnormality is detected based on the analysis results, a notification will be sent to the owner's smartphone, allowing the owner to take prompt action.

[0144] Prompt Sentence Examples

[0145] "Collect pet sounds and analyze them using an AI model. If an abnormality is detected, create a program that notifies the owner."

[0146] In this way, the present invention makes it possible to grasp abnormal behavior and health conditions of pets in real time, and to promptly notify the owner if an abnormality is detected.

[0147] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0148] Step 1:

[0149] The device (smartphone or smart glasses) uses a microphone to collect the pet's barks. The input is the pet's barks, and the output is digital audio data. Specifically, the device uses software for audio data collection (sounddevice library) to record audio for a certain period of time and save it as digital data.

[0150] Step 2:

[0151] The device sends the collected voice data to a server. The input is digital voice data, and the output is the voice data sent to the server. Specifically, the device uses an internet connection to upload the voice data to the server.

[0152] Step 3:

[0153] The server inputs the received audio data into the TensorFlow model for analysis. The input is digital audio data, and the output is the analysis result (the pet's state). Specifically, the server uses the TensorFlow library to input the audio data into the AI ​​model and analyze the characteristics of the pet's cries.

[0154] Step 4:

[0155] The server determines whether an anomaly has been detected based on the analysis results. The input is the analysis results, and the output is information about whether an anomaly has been detected. Specifically, the server compares the analysis results with a threshold and sets a flag if an anomaly is detected.

[0156] Step 5:

[0157] If an abnormality is detected, the server sends a notification to the pet owner. The input is information about whether an abnormality exists, and the output is a notification to the pet owner. Specifically, the server uses the smtplib library to send an abnormality notification to the pet owner's email address. The notification contains detailed information about the pet's abnormal behavior and health status.

[0158] Step 6:

[0159] The user (owner) checks the received notification and takes action to check the pet's status if necessary. The input is the notification from the server, and the output is the user's response. Specifically, the user checks the notification on their smartphone and takes action such as returning home to check the pet's status or contacting the pet sitter.

[0160] Example 2

[0161] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0162] Conventional methods of communication between pets and humans have been limited, making it difficult to accurately understand a pet's intentions and state from its cries and behavior. Furthermore, there has been a lack of ways to communicate human language to pets in a way that they can understand. This has made it difficult to properly manage a pet's health and emotional state, resulting in insufficient communication between pets and their owners.

[0163] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the status of the pet from its cries based on the collected data, means for converting the voice data into text data using voice recognition software, means for analyzing the text data using a generative AI model to understand the pet's intentions, means for generating a message based on the analysis result and notifying the user, means for collecting the user's words and converting it into text data using voice recognition software, means for analyzing the text data using a generative AI model to generate a message to be conveyed to the pet, and means for converting the message into cries and actions that the pet can understand using voice synthesis software and transmitting it. This makes it possible to accurately understand the intentions and status of the pet from its cries and actions, and to convey human language in a form that the pet can understand.

[0164] "Audio data" refers to data that has been recorded in digital format from sounds made by animals or humans.

[0165] "Animal characteristics" refers to characteristics unique to individual animals, such as their personalities, sounds, and behavioral patterns.

[0166] "Type of animal" indicates the classification of animals, such as dog, cat, bird, etc.

[0167] "General animal characteristic data" refers to data such as personality and behavioral patterns common to a particular animal species.

[0168] "Speech recognition software" is software that analyzes voice data and converts it into text data.

[0169] A "generative AI model" is a model that uses artificial intelligence to analyze text data and understand its meaning.

[0170] "Text data" is text information converted from voice data by voice recognition software.

[0171] "Message" refers to the content of notifications or instructions generated based on the analysis results.

[0172] "Speech synthesis software" is software for converting text data into speech.

[0173] "User" refers to a person who uses this system.

[0174] "Pet" refers to an animal kept in a household.

[0175] This invention is a system that uses AI to enable conversation between humans and pets. This system has the functions of analyzing pet cries and converting them into words that humans can understand, and analyzing human words and converting them into cries and actions that pets can understand.

[0176] Hardware and software used

[0177] Hardware

[0178] Microphone: Used to capture pet sounds and user speech.

[0179] Speaker: Used to transmit generated sounds to your pet.

[0180] Camera: Used to monitor pet behavior if necessary.

[0181] Server: Used to analyze and process data.

[0182] software

[0183] Speech recognition software: used to convert voice data into text data (e.g., Google Speech-to-Text).

[0184] Generative AI models: used to analyze text data and understand your pet's intent (e.g., GPT-4®).

[0185] Text-to-speech software: Used to convert text data into speech (e.g., Google Text-to-Speech).

[0186] Specific operation of the system

[0187] Pet cry analysis

[0188] The device's microphone captures the sound of your pet barking. For example, if your dog barks, "woof woof," the audio data is collected through the microphone and sent to the server in real time.

[0189] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice "woof woof" is converted into the text "woof woof." The converted text data is sent to the next analysis step.

[0190] The server uses a generative AI model to analyze the text data and understand the pet's intentions. For example, the text "woof woof" is interpreted as meaning "I'm hungry." The analysis results are sent to the message generation step.

[0191] Based on the analysis results, the server generates a message saying "your pet is hungry" and notifies the user. For example, a notification is sent to the user via a smartphone app. The user can check the notification on the app and understand the status of their pet.

[0192] Human language analysis

[0193] The microphone on the device captures the user's words. For example, when the user says "Let's play," the voice data is collected through the microphone. The collected voice data is sent to the server in real time.

[0194] The server analyzes the captured voice data using speech recognition software and converts it into text data. For example, the voice saying "Let's play" is converted into the text "Let's play." The converted text data is sent to the next analysis step.

[0195] The server uses a generative AI model to analyze the text data and generate a message to convey to the pet. For example, the text "Let's play" is analyzed to mean "It's play time." The analysis result is sent to the message generation step.

[0196] The server uses speech synthesis software to convert the generated message into sounds and actions that the pet can understand, and communicates them to the pet through a speaker. For example, the message "It's playtime" is converted into sounds and actions that a dog can understand, and played back from the speaker. The pet understands the sounds and actions, and prepares to play with the user.

[0197] Examples of specific examples and prompts

[0198] Specific examples

[0199] Pet Sound Analysis:

[0200] If the pet barks "woof woof," the microphone captures the sound and the server notifies the user that "your pet is hungry."

[0201] Human language analysis:

[0202] When the user says "let's play," a microphone captures the voice and the server communicates through the speaker with sounds and actions that the pet understands as "play time."

[0203] Prompt Sentence Examples

[0204] Pet Sound Analysis:

[0205] If your pet barks "woof woof," explain the process of analyzing that sound and informing the user that "your pet is hungry."

[0206] Human language analysis:

[0207] When a user says "Let's play," describe the process for analyzing that speech and communicating "It's playtime" to your pet in a way that makes sense.

[0208] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0209] Step 1: Capture your pet's sounds

[0210] The device's microphone captures the sound of your pet barking. For example, if your dog barks "woof woof," the audio data is collected through the microphone. The input is your pet's bark, and the output is audio data. The collected audio data is sent to the server in real time.

[0211] Step 2: Analyzing bird calls and converting them into text

[0212] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice "woof woof" is converted into the text "woof woof." The input is voice data and the output is text data. The converted text data is sent to the next analysis step.

[0213] Step 3: Semantic analysis of the text

[0214] The server uses a generative AI model to analyze the text data and understand the pet's intention. For example, the text "woof woof" is analyzed to mean "I'm hungry." The input is the text data, and the output is the analysis result indicating the pet's intention. The analysis result is sent to the message generation step.

[0215] Step 4: Message generation and notification

[0216] The server generates a message based on the analysis results, saying "Your pet is hungry," and notifies the user. For example, the notification is sent to the user via a smartphone app. The input is the analysis results, and the output is a notification message to the user. The user can check the notification on the app to understand the status of their pet.

[0217] Step 5: Capturing human language

[0218] The microphone on the device captures the user's words. For example, when a user says "Let's play," the voice data is collected through the microphone. The input is the user's words, and the output is voice data. The collected voice data is sent to the server in real time.

[0219] Step 6: Word analysis and text conversion

[0220] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice saying "Let's play" is converted into the text "Let's play." The input is voice data and the output is text data. The converted text data is sent to the next analysis step.

[0221] Step 7: Semantic analysis of the text

[0222] The server uses a generative AI model to analyze the text data and generate a message to be conveyed to the pet. For example, the text "Let's play" is analyzed to mean "It's play time." The input is text data, and the output is the message to be conveyed to the pet. The analysis result is sent to the message generation step.

[0223] Step 8: Generate a message and send it to your pet

[0224] The server uses speech synthesis software to convert the generated message into sounds and actions that the pet can understand, and communicates them to the pet through a speaker. For example, the message "It's playtime" is converted into sounds and actions that the dog can understand, and played back from the speaker. The input is the message to be communicated to the pet, and the output is sounds and actions that the pet can understand. The pet understands the sounds and actions, and gets ready to play with the user.

[0225] (Application example 2)

[0226] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0227] Conventional pet care systems have had problems such as difficulty in accurately understanding pet sounds and behavior and taking appropriate action. They also lacked a means to notify owners when their pets were hungry and quickly order pet food. Furthermore, they lacked an effective means to facilitate smooth communication between owners and their pets.

[0228] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from the cries based on the collected data, means for notifying the owner based on the pet's condition, and means for ordering pet food. This makes it possible to accurately determine the pet's condition and respond quickly. It also facilitates communication between the owner and the pet and automates pet food ordering.

[0229] "Audio data" refers to sound information, including animal sounds.

[0230] "Characteristics of pets" refers to individual characteristics such as the personality and sounds of pets.

[0231] "Means for identifying the type of animal" refers to a method or device for identifying the type of pet.

[0232] "General animal characteristic data" refers to information about widely known animal characteristics.

[0233] "Means for understanding the condition of pets from their cries" refers to methods and devices for analyzing pet cries to understand their condition.

[0234] "Means for notifying owners based on the status of their pets" refers to methods or devices for informing owners of the status of their pets.

[0235] "Pet food ordering means" refers to a method or device for purchasing pet food online.

[0236] "Means for enabling conversation between humans and pets using AI" refers to methods and devices for exchanging information between humans and pets using artificial intelligence.

[0237] "Means for understanding the health and emotional state of a pet" refers to a method or device for understanding the physical condition and emotions of a pet.

[0238] The following system configuration will be described as an embodiment of the present invention.

[0239] System Configuration

[0240] This system includes a means for collecting audio data (animal cries), a means for collecting characteristics of pet animals (personality, cries, etc.), a means for identifying the type of animal, a means for collecting general data on animal characteristics, a means for determining the condition of pets from their cries based on the collected data, a means for notifying owners based on the condition of their pets, and a means for ordering pet food.

[0241] Hardware and software used

[0242] Hardware: Smartphone, microphone

[0243] Software: Python, SpeechRecognition library, Requests library

[0244] Data processing and calculation

[0245] Audio data collection and analysis

[0246] The server records the pet's cries using the smartphone's microphone. The recorded audio data is converted to text using the SpeechRecognition library. It is then analyzed using a generative AI model to understand the pet's condition.

[0247] Collecting and analyzing owner speech

[0248] The server uses the smartphone's microphone to record the owner's words, converts the recorded voice data into text using the SpeechRecognition library, and then analyzes it using a generative AI model to generate a message to be conveyed to the pet.

[0249] Pet food orders

[0250] The server notifies the owner based on the pet's condition. For example, if the pet cries "I'm hungry," the server generates a prompt to order pet food and notifies the owner. When the owner orders pet food through the app, recommended products are displayed taking into account the pet's preferences and allergies.

[0251] Specific examples

[0252] Analyzing pet cries: When a pet cries "I'm hungry," the server analyzes the cry and notifies the owner with a message that "Your pet is hungry."

[0253] Analyzing owner's words: When the owner says "Let's play," the server analyzes the words and conveys the message in a way that the pet can understand.

[0254] Prompt Sentence Examples

[0255] "Analyze your pet's cries and let us know their condition."

[0256] "Analyze the owner's words and generate a message to convey to the pet."

[0257] In this way, the present invention allows for accurate understanding of the pet's condition and prompt response, facilitates communication between pet owners and their pets, and automates pet food ordering.

[0258] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0259] Step 1:

[0260] The server records the pet's cries using the smartphone's microphone. The input is the pet's cries, and the output is an audio data file. This audio data file is used for subsequent analysis.

[0261] Step 2:

[0262] The server converts the recorded audio data into text data using the SpeechRecognition library. The input is an audio data file, and the output is text data. This text data is input into the generative AI model.

[0263] Step 3:

[0264] The server analyzes the text data using a generative AI model to understand the pet's state. The input is text data, and the output is information indicating the pet's state. For example, the analyzed state is "hungry."

[0265] Step 4:

[0266] The server notifies the owner based on the pet's status. The input is information indicating the pet's status, and the output is a notification message to the owner. For example, a message saying "your pet is hungry" is sent to the owner.

[0267] Step 5:

[0268] The server records the owner's words using the smartphone's microphone. The input is the owner's words, and the output is an audio data file. This audio data file is used for subsequent analysis.

[0269] Step 6:

[0270] The server converts the recorded voice data of the owner into text data using the SpeechRecognition library. The input is a voice data file, and the output is text data. This text data is input into the generative AI model.

[0271] Step 7:

[0272] The server uses a generative AI model to analyze the text data and generate a message to be conveyed to the pet. The input is text data, and the output is a message to be conveyed to the pet. For example, the command "Let's play" is analyzed.

[0273] Step 8:

[0274] The server provides a means for ordering pet food based on the pet's condition. The input is information indicating the pet's condition, and the output is pet food ordering information. For example, if the pet cries "I'm hungry," an order for pet food is automatically placed.

[0275] Step 9:

[0276] The server displays recommended products to the pet owner, taking into account the pet's preferences and allergy information. The input is the pet's preferences and allergy information, and the output is a list of recommended products. This allows the pet owner to select the appropriate pet food.

[0277] Step 10:

[0278] The server notifies the pet owner that the pet food order has been completed. The input is the order completion information, and the output is a notification message to the pet owner. For example, a message saying "Pet food order completed" is sent to the pet owner.

[0279] Example 3

[0280] Next, a description will be given of a third embodiment of the third embodiment. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0281] Conventional systems for understanding pet health and emotional states only analyze voice data and do not take behavioral data into account, making it difficult to accurately understand a pet's condition. Additionally, there is a lack of a way to quickly notify the owner of the analysis results, making it difficult to detect abnormalities in a pet early on.

[0282] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[0283] In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for determining the status of pet animals from their cries based on the collected data, means for collecting animal behavior data, means for analyzing the collected behavior data, and means for notifying the results of the analysis. By analyzing both the voice data and the behavior data, it is possible to more accurately determine the health and emotional state of pets and promptly notify their owners.

[0284] "Audio Data" refers to animal sounds and other audio information.

[0285] "Animal characteristics" refers to individual characteristics such as an animal's personality, vocalizations, and behavioral patterns.

[0286] "Type of animal" refers to a classification of animals such as dogs, cats, birds, etc.

[0287] "General animal characteristic data" refers to data such as personality and behavioral patterns common to a particular type of animal.

[0288] "Animal condition" refers to the animal's physical and emotional state.

[0289] "Behavioral data" refers to information about an animal's movements and behavior patterns.

[0290] "Analysis results" refers to the diagnosis and evaluation output by the AI ​​model based on the collected voice data and behavioral data.

[0291] "Notification means" refers to a method or device for notifying the owner of the analysis results.

[0292] The present invention is a system for understanding the health and emotional state of a pet. A specific embodiment of this system will be described below.

[0293] First, the user uses a device such as a smartphone or tablet to collect the sounds and behavior of their pet. The device is equipped with a microphone and camera, and these devices are used to capture audio data and behavior data in real time. For example, when a user launches the smartphone app, points it at their pet, and presses the "Start Recording" button, the microphone begins to collect the pet's sounds. At the same time, the camera captures the pet's movements.

[0294] The device then transmits the collected data to a server in real time using Wi-Fi or mobile data, and the data is encrypted to ensure security.

[0295] The server inputs the received audio and behavior data into an AI model, which is built using machine learning frameworks such as TensorFlow and PyTorch. The AI ​​model analyzes the frequency and volume of the pet's cries, as well as the speed and direction of its movements, to detect abnormal patterns and patterns of delight.

[0296] The analysis results are sent from the server to the device, and the user can check the results through a smartphone app. For example, the server may send a warning message such as "Your pet may be sick." If the pet is happy, it may send a message such as "Your pet is happy."

[0297] As a concrete example, the following prompt sentence could be input to a generative AI model:

[0298] "Analyze your pet's vocalizations and behavioral data to determine their health and emotional state. Alert you if they're sick and let you know if they're happy."

[0299] Using this prompt, the AI ​​model can accurately analyze the pet's condition and provide appropriate feedback to the user.

[0300] As described above, by analyzing both the voice data and the behavioral data, the present invention makes it possible to grasp the health condition and emotional state of the pet more accurately and to notify the owner promptly. The flow of the identification process in the third embodiment will be described with reference to FIG.

[0301] Step 1:

[0302] The user launches the smartphone app and collects the sounds and behavior of their pet. Specifically, when the user presses the "Start Recording" button, the device's microphone begins collecting the pet's sounds. At the same time, the camera captures the pet's movements. The input is the pet's sounds and behavior, and the output is the collected audio data and behavior data.

[0303] Step 2:

[0304] The device transmits the collected voice and behavioral data to a server. Specifically, the device transmits data in real time using Wi-Fi or mobile data. The input is the collected voice and behavioral data, and the output is the data transmitted to the server.

[0305] Step 3:

[0306] The server inputs the received voice data and behavioral data into an AI model. Specifically, the server inputs the data into an AI model built using machine learning frameworks such as TensorFlow and PyTorch. The input is the voice data and behavioral data sent to the server, and the output is the analysis results by the AI ​​model.

[0307] Step 4:

[0308] The server determines the pet's health and emotional state based on the analysis results of the AI ​​model. Specifically, the AI ​​model analyzes the frequency and volume of the pet's cries, as well as the speed and direction of its movements, to detect abnormal patterns and patterns of joy. The input is the analysis results of the AI ​​model, and the output is a judgment of the pet's health and emotional state.

[0309] Step 5:

[0310] The server sends the analysis results to the device. Specifically, the server sends a message generated based on the analysis results to the device. The input is the judgment result of the pet's health condition and emotional state, and the output is the message sent to the device.

[0311] Step 6:

[0312] The user checks the analysis results through the device. Specifically, the user opens the smartphone app and checks the notified message. For example, a warning message such as "Your pet may be sick" or a message such as "Your pet is happy" may be displayed. The input is the message sent to the device, and the output is the analysis result checked by the user.

[0313] (Application example 3)

[0314] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0315] Understanding a pet's health and emotional state is important for pet owners, but existing methods make it difficult to detect abnormalities early on. There are also limited ways to accurately determine whether a pet is happy. This can lead to inadequate health management and understanding of a pet's emotions, potentially reducing the pet's quality of life. Furthermore, the lack of a way to quickly notify owners of abnormalities can delay emergency response.

[0316] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[0317] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from its cries based on the collected data, means for analyzing the pet's condition and sending a notification if an abnormality is detected, means for determining the pet's health and emotional state in real time, means for issuing an alert to the owner if an abnormality is detected, and means for informing the owner if the pet is happy. This makes it possible to quickly and accurately determine the pet's health and emotional state, and to immediately notify the owner if an abnormality occurs.

[0318] "Audio data" refers to data that has been recorded in digital format, such as animal cries.

[0319] "Animal characteristics" refers to the animal's unique characteristics, such as its personality or sounds.

[0320] "Means for identifying animal species" refers to methods or techniques for identifying the species of an animal.

[0321] "General animal characteristic data" refers to data about widely known animal characteristics.

[0322] "Means for understanding the condition of pets" refers to methods for analyzing the health and emotional state of pets based on collected data.

[0323] "Means for analyzing a pet's condition" refers to a method for analyzing a pet's cries and behavior to assess its condition.

[0324] "Means for sending a notification when an abnormality is detected" refers to a method for alerting the owner when an abnormality is detected in the pet.

[0325] "Means for understanding a pet's health and emotional state in real time" refers to a method for instantly analyzing a pet's condition and providing the results in real time.

[0326] "Means for issuing a warning to the owner when an abnormality is detected" refers to a method for issuing a warning to the owner when an abnormality is detected in the pet.

[0327] The "means for conveying information about a happy pet to its owner" refers to a method for detecting a happy state of a pet and conveying that information to its owner.

[0328] As an embodiment of the present invention, a pet security alert system will be described as an example. This system analyzes the sounds and behavior of pets in real time and notifies the owner if an abnormality is detected.

[0329] System Configuration

[0330] Hardware:

[0331] Smartphone

[0332] microphone

[0333] software:

[0334] TensorFlow

[0335] librosa

[0336] smtplib

[0337] Data processing and calculation

[0338] Audio data collection:

[0339] The device (smartphone) uses a microphone to collect the sounds your pet makes, and the audio data is recorded in digital format.

[0340] Analysis of audio data:

[0341] The server converts the collected audio data into MFCC (Mel-Frequency Cepstrum Coefficients) using librosa, which allows the extraction of audio data features.

[0342] Analysis by AI model:

[0343] The server inputs the audio data into a generative AI model pre-trained with TensorFlow to analyze the pet's health and emotional state, which is then used to detect abnormalities in the pet.

[0344] Send notifications:

[0345] If an abnormality is detected, the server uses smtplib to send an email notification to the owner, containing detailed information about the pet's abnormal condition.

[0346] Specific examples

[0347] For example, you can record your pet's cries and input the audio data into the application. The AI ​​model analyzes the audio data, and if it detects an abnormality, it will notify the owner by email. This notification will include a message such as, "An abnormality has been detected with your pet. Please check immediately."

[0348] Prompt Sentence Examples

[0349] Examples of prompts to be input to a generative AI model include:

[0350] Write a Python program that analyzes pet noise data and notifies owners if an abnormality is detected. Use TensorFlow and librosa to analyze the audio data, and smtplib to send email notifications.

[0351] In this way, a system can be provided that can grasp the health and emotional state of a pet in real time and respond quickly if there is an abnormality.

[0352] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[0353] Step 1:

[0354] Audio data collection

[0355] The device (smartphone) collects pet sounds using a microphone. The input is the pet's sound, and the output is digital audio data. This audio data is used in subsequent analysis steps.

[0356] Step 2:

[0357] Audio data preprocessing

[0358] The server converts the collected audio data into MFCC (Mel-Frequency Cepstrum Coefficients) using librosa. The input is digital audio data, and the output is MFCC features. This conversion extracts the features of the audio data.

[0359] Step 3:

[0360] Analysis using AI models

[0361] The server inputs MFCC features into a generative AI model pre-trained using TensorFlow to analyze the pet's health and emotional state. The input is the MFCC features, and the output is a prediction result about the pet's condition. This analysis evaluates the pet's abnormalities and emotional state.

[0362] Step 4:

[0363] Anomaly detection

[0364] The server determines whether an abnormality has been detected in the pet based on the prediction results of the AI ​​model. The input is the prediction result of the AI ​​model, and the output is information about whether an abnormality has been detected. If an abnormality is detected, the server proceeds to the next step.

[0365] Step 5:

[0366] Sending notifications

[0367] The server uses smtplib to send notifications to the pet owner via email. The input is the pet owner's email address and information about the presence or absence of an abnormality. The output is the notification email sent. This notification contains detailed information about the pet's abnormal condition.

[0368] Step 6:

[0369] Real-time status monitoring

[0370] The server monitors the pet's health and emotional state in real time and provides the information to the owner as needed. The input is continuously collected voice data, and the output is real-time status information, allowing the owner to always be aware of the pet's condition.

[0371] In this way, it is possible to quickly and accurately grasp the health and emotional state of a pet, and to immediately notify the owner if any abnormalities occur.

[0372] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0373] "Example 1"

[0374] The system described in claim 4 includes an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's emotions from the user's tone of voice, facial expressions, behavior, etc. For example, when the user is angry at a pet, the emotion engine recognizes the anger and adjusts the instructions to the pet. When the user is happy, the emotion engine recognizes the joy and adjusts the instructions to the pet. This allows for smoother communication between the pet and the user.

[0375] "Example 2"

[0376] The system described in claim 5 includes a means for utilizing AI to realize conversation between a person and a pet, and an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's emotions from the user's tone of voice, facial expressions, behavior, etc., and feeds the results back to the AI. The AI ​​adjusts its instructions to the pet based on this feedback. For example, when the user is angry at the pet, the AI ​​recognizes the anger and adjusts its instructions to the pet. Also, when the user is happy, the AI ​​recognizes the joy and adjusts its instructions to the pet. This allows for smoother communication between the pet and the user.

[0377] "Example 3"

[0378] The system described in claim 6 includes a means for grasping the pet's state and an emotion engine for recognizing the user's emotions. The emotion engine analyzes the user's emotions from the user's tone of voice, facial expressions, behavior, etc., and feeds the results back to the means for grasping the pet's state. The means for grasping the pet's state adjusts the pet's state based on this feedback. For example, when the user is angry with the pet, the means for grasping the pet's state recognizes the anger and adjusts the pet's behavior. Also, when the user is happy, the means for grasping the pet's state recognizes the joy and adjusts the pet's behavior. This allows for smoother communication between the pet and the user.

[0379] The processing flow of each embodiment will be described below.

[0380] "Example 1"

[0381] Step 1: The emotion engine analyzes the user's emotions based on their tone of voice, facial expressions, behavior, etc.

[0382] Step 2: Based on the analysis results of the emotion engine, adjust the instructions given to the pet.

[0383] Step 3: The system operates to facilitate smooth communication between the pet and the user.

[0384] "Example 2"

[0385] Step 1: The emotion engine analyzes the user's emotions based on their tone of voice, facial expressions, behavior, etc.

[0386] Step 2: The emotion engine feeds back the results of its analysis to the AI.

[0387] Step 3: The AI ​​uses this feedback to adjust its instructions to the pet.

[0388] Step 4: The system operates to facilitate smooth communication between the pet and the user.

[0389] "Example 3"

[0390] Step 1: The emotion engine analyzes the user's emotions based on their tone of voice, facial expressions, behavior, etc.

[0391] Step 2: The results of the emotion engine's analysis are fed back into a means of understanding the pet's condition.

[0392] Step 3: The means to understand the pet's condition will use this feedback to adjust the pet's condition.

[0393] Step 4: The system operates to facilitate smooth communication between the pet and the user.

[0394] Example 1

[0395] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0396] Conventional pet management systems simply record the sounds and behavior of pets, making it difficult to respond appropriately based on the pet's condition or the user's emotions. Furthermore, they lacked a means for smooth communication between the user and the pet. This made it difficult to properly understand the pet's health and emotional state and give instructions to the pet based on the user's emotions.

[0397] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0398] In this invention, the server includes means for collecting voice data, means for collecting characteristics of pet animals, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the state of the pet from its cries based on the collected data, means for recognizing the user's emotions, and means for adjusting instructions to the pet based on the user's emotions. This makes it possible to understand the state of the pet and the user's emotions in real time and to give appropriate instructions to the pet.

[0399] "Audio data" refers to data in which sounds, such as animal cries or the user's voice, are recorded in digital format.

[0400] "Characteristics of pets" refers to specific information about pets, such as their personalities, sounds, and behavioral patterns.

[0401] "Type of animal" is information indicating the type of pet, such as dog, cat, bird, etc.

[0402] "General animal characteristic data" refers to data such as personality and behavior patterns common to a particular animal species.

[0403] "Pet status" is information that indicates the current state of the pet, such as the pet's health and emotional state.

[0404] "User emotion" refers to the emotional state of the user as analyzed from the user's tone of voice, facial expression, behavior, etc.

[0405] "Adjusting instructions" means changing the instructions given to the pet based on the user's emotions and the state of the pet.

[0406] "Collecting means" refers to devices and methods for acquiring voice data, animal characteristic data, etc.

[0407] "Means for identification" refers to devices or methods for identifying the type and characteristics of animals based on collected data.

[0408] The "analysis means" refers to a device or method for analyzing the collected data and understanding the pet's condition and the user's emotions.

[0409] This invention is a system that collects and analyzes data on the cries and behavior of pets to understand their state and adjust instructions to the pet based on the user's emotions. Specific embodiments of this system will be described below.

[0410] Audio data collection

[0411] Users use the microphone on their smartphone to collect their pet's cries. The collected audio data is sent from the device to the server in real time. Specifically, the user starts the smartphone app and presses the "Start Recording" button to begin collecting audio data.

[0412] Analysis of audio data

[0413] The server uses voice analysis software to analyze the audio data. Specifically, it converts the collected audio data into text data using a common voice analysis API. For example, it can use the Google Cloud Speech-to-Text API. Based on the analysis results, it identifies whether the audio is the sound of a specific animal.

[0414] Collection of animal characteristic data

[0415] Users input their pet's personality and vocalizations into a smartphone app. The device also uses sensors to observe the pet's behavior. Specifically, it uses behavioral observation sensors such as Nest Cam to collect data on the pet's behavior. The collected data is then sent to a server.

[0416] Identifying Animal Species

[0417] The user inputs the pet's type into the app, and the device uses image recognition technology to identify the pet's type. Specifically, a general image recognition API can be used. For example, Google Cloud Vision API can be used to analyze the pet's image and identify the animal type.

[0418] Analyzing data and understanding your pet's condition

[0419] The server uses AI models to analyze the collected data, specifically machine learning frameworks such as TensorFlow, to predict the pet's condition and understand its health and emotional state.

[0420] User Emotion Recognition

[0421] The device uses emotion analysis software to recognize the user's emotions. Specifically, it uses the Emotion API to analyze the user's tone of voice and facial expressions. Based on the analysis results, the server recognizes whether the user is angry or happy with their pet.

[0422] Adjusting pet instructions

[0423] The server adjusts the instructions to the pet based on the user's emotions: if the user is angry, it generates a gentle instruction, and if the user is happy, it generates a reinforced instruction. The generated instructions are transmitted to the pet via the terminal.

[0424] Specific examples

[0425] Audio data collection: The user launches the app on their smartphone and presses the "Start Recording" button to record the sound of their dog barking.

[0426] Voice data analysis: The server uses the Google Cloud Speech-to-Text API to analyze the recorded voice data and obtain the text data "Woof woof."

[0427] Animal characteristic data collection: The user enters "The dog's personality is gentle" into the app's input form, and the device uses Nest Cam to collect the dog's behavior data.

[0428] Animal species identification: A user uploads a photo of a dog to the app, and the device identifies it as a "Golden Retriever" using the Google Cloud Vision API.

[0429] Data analysis and pet condition assessment: The server uses TensorFlow to analyze the collected data and determine whether the dog is experiencing stress.

[0430] User emotion recognition: The device uses the Emotion API to analyze the user's tone of voice and facial expressions and recognize that the user is angry.

[0431] Pet command adjustment: The server analyzes the user's emotional data and generates gentle commands such as "sit," which the device then relays to the dog via voice.

[0432] Prompt Sentence Examples

[0433] "Collect dog barks and analyze them with a common audio analysis API. Based on the analysis results, use a machine learning framework to understand the dog's state, and use emotion analysis software to recognize the user's emotions and adjust instructions for the pet."

[0434] In this way, the system specifically executes each processing step to facilitate smooth communication between the pet and the user.

[0435] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0436] Step 1:

[0437] Audio data collection

[0438] The user collects the sounds of their pet using the microphone on their smartphone. Specifically, the user starts the smartphone app and presses the "Start Recording" button to begin collecting audio data.

[0439] Input: Pet noises

[0440] Output: Recorded audio data

[0441] Step 2:

[0442] Sending audio data

[0443] The device transmits the collected voice data to a server in real time. Specifically, the smartphone app uploads the recorded voice data to the server via the Internet.

[0444] Input: Recorded audio data

[0445] Output: Audio data sent to the server

[0446] Step 3:

[0447] Analysis of audio data

[0448] The server uses speech analysis software to analyze the voice data. Specifically, it converts the collected voice data into text data using a common speech analysis API. For example, Google Cloud Speech-to-Text API can be used.

[0449] Input: Audio data sent to the server

[0450] Output: Parsed text data

[0451] Step 4:

[0452] Collection of animal characteristic data

[0453] Users input their pet's personality and vocalizations into a smartphone app. The device also uses sensors to observe the pet's behavior. Specifically, it uses behavioral observation sensors such as Nest Cam to collect data on the pet's behavior. The collected data is then sent to a server.

[0454] Input: Pet's personality, vocalizations, and behavioral data

[0455] Output: Feature and behavioral data sent to the server

[0456] Step 5:

[0457] Identifying Animal Species

[0458] The user inputs the pet's type into the app, and the device uses image recognition technology to identify the pet's type. Specifically, a general image recognition API can be used. For example, Google Cloud Vision API can be used to analyze the pet's image and identify the animal type.

[0459] Input: Pet image

[0460] Output: Analyzed animal species

[0461] Step 6:

[0462] Analyzing data and understanding your pet's condition

[0463] The server uses AI models to analyze the collected data, specifically machine learning frameworks such as TensorFlow, to predict the pet's condition and understand its health and emotional state.

[0464] Input: feature data, behavior data, animal type

[0465] Output: Parsed pet status

[0466] Step 7:

[0467] User Emotion Recognition

[0468] The device uses emotion analysis software to recognize the user's emotions. Specifically, it uses the Emotion API to analyze the user's tone of voice and facial expressions. Based on the analysis results, the server recognizes whether the user is angry or happy with their pet.

[0469] Input: User's tone of voice and facial expressions

[0470] Output: Parsed user sentiment

[0471] Step 8:

[0472] Adjusting pet instructions

[0473] The server adjusts the instructions to the pet based on the user's emotions: if the user is angry, it generates a gentle instruction, and if the user is happy, it generates a reinforced instruction. The generated instructions are transmitted to the pet via the terminal.

[0474] Input: Analyzed pet state, user emotion

[0475] Output: Adjusted instructions

[0476] In this way, the system specifically executes each processing step to facilitate smooth communication between the pet and the user.

[0477] (Application example 1)

[0478] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0479] Conventional pet care systems only analyze pets' sounds and behaviors, and are unable to provide advice or services that take the user's emotions into account. This makes it difficult to communicate smoothly between pets and users, making it difficult to properly manage pets' health and emotional state. Furthermore, in brick-and-mortar stores, it is difficult to provide appropriate services tailored to the pet's condition.

[0480] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for understanding the pet's condition from its cries based on the collected data, means for recognizing the user's emotions, means for generating advice based on the pet's condition and the user's emotions, and means for providing the generated advice to the user. This makes it possible to comprehensively analyze the pet's condition and the user's emotions and provide appropriate advice and services.

[0481] "Audio data" refers to audio information such as animal cries.

[0482] "Characteristics of pet animals" refers to information about the characteristics of the animals kept by the owner, such as their personalities and sounds.

[0483] "Means for identifying the type of animal" refers to a method for identifying the type of animal using image recognition technology or input from the owner.

[0484] "General animal characteristic data" refers to characteristic information about general animals obtained from a database or the like.

[0485] "Means for understanding a pet's condition" refers to a method for analyzing a pet's health and emotional state based on collected voice data and animal characteristic data.

[0486] "Means for recognizing user emotions" refers to a method for identifying a user's emotions by analyzing the user's tone of voice, facial expressions, behavior, etc.

[0487] The "means for generating advice" refers to a method for generating appropriate advice or services based on the pet's condition and the user's emotions.

[0488] The "means for providing the generated advice to the user" refers to a method for communicating the generated advice or service to the user.

[0489] A system for implementing this invention includes means for collecting audio data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying the type of animal, means for collecting general animal characteristic data, means for understanding the condition of the pet from its cries based on the collected data, means for recognizing the user's emotions, means for generating advice based on the pet's condition and the user's emotions, and means for providing the generated advice to the user.

[0490] Hardware and software used

[0491] Hardware:

[0492] Microphone: Used to collect audio data.

[0493] Camera: Used to collect video data of the user.

[0494] software:

[0495] speech_recognition library: Used to perform speech recognition.

[0496] cv2 (OpenCV): Used to collect and analyze video data.

[0497] emotion_recognition: An emotion recognition engine to recognize user emotions.

[0498] pet_behavior_analysis: A pet behavior analysis engine for analyzing pet sounds and behavior.

[0499] advice_generator: An advice generation engine for generating advice based on the pet's state and the user's emotions.

[0500] System processing overview

[0501] 1. Audio data collection:

[0502] The server collects pet sounds using a microphone, and the collected audio data is analyzed using the speech_recognition library.

[0503] 2. Video data collection:

[0504] The server collects video data of the user using a camera, and the collected video data is analyzed using cv2 (OpenCV).

[0505] 3. Pet status analysis:

[0506] The server analyzes the collected voice data using the pet_behavior_analysis engine to understand the pet's condition (e.g., hunger, stress, joy, etc.).

[0507] 4. User Emotion Recognition:

[0508] The server analyzes the collected video data using the emotion_recognition engine to recognize the user's emotions (e.g., anger, joy, sadness, etc.).

[0509] 5. Generating Advice:

[0510] The server generates appropriate advice using the advice_generator engine based on the pet's condition and the user's emotions.

[0511] 6. Providing advice:

[0512] The server provides the generated advice to the user, for example, if the pet is stressed, it suggests items to help the pet relax.

[0513] Specific examples

[0514] If your pet is stressed:

[0515] The server generates advice such as "Your pet is stressed. Try using a toy that will help it relax." and provides it to the user.

[0516] If the user is angry:

[0517] The server generates advice such as "The user is angry. Please be kind to your pet." and provides it to the user.

[0518] Prompt Sentence Examples

[0519] Analyze the pet's meow data and the user's video data, and generate appropriate advice based on the pet's condition and the user's emotions. If the pet is stressed, suggest items that will help the pet relax, and if the user is angry, advise the user to be gentle with the pet.

[0520] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0521] Step 1:

[0522] The server collects pet sounds using a microphone. The collected audio data is analyzed using the speech_recognition library. The input is the pet's sound, and the output is the analyzed audio data. Through this analysis, the characteristics of the pet's sound are extracted.

[0523] Step 2:

[0524] The server collects the user's video data using a camera. The collected video data is analyzed using cv2 (OpenCV). The input is the user's video data, and the output is the analyzed video data. This analysis extracts the user's facial expressions and behavioral characteristics.

[0525] Step 3:

[0526] The server analyzes the collected voice data using the pet_behavior_analysis engine to understand the pet's state. The input is the analyzed voice data, and the output is the pet's state (e.g., hunger, stress, joy, etc.). This analysis identifies the pet's health and emotional state.

[0527] Step 4:

[0528] The server analyzes the collected video data using the emotion_recognition engine to recognize the user's emotions. The input is the analyzed video data, and the output is the user's emotion (e.g., anger, joy, sadness, etc.). This analysis identifies the user's emotional state.

[0529] Step 5:

[0530] The server generates appropriate advice using the advice_generator engine based on the pet's state and the user's emotion. The input is the pet's state and the user's emotion, and the output is the generated advice. This generation creates specific advice that is tailored to the pet and user's situation.

[0531] Step 6:

[0532] The server provides the generated advice to the user. The input is the generated advice, and the output is the advice provided to the user. This allows the user to take appropriate action according to the state of their pet.

[0533] Example 2

[0534] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0535] Conventional methods of communicating with pets have had the problem that it is difficult to accurately understand the sounds and behavior of pets, and it takes time to grasp the pet's condition and requests. Also, when a user gives instructions to a pet, the pet often does not understand the instructions, which makes communication difficult.

[0536] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the status of the pet from its cries based on the collected data, means for analyzing the collected voice data, means for notifying the user of the analysis results, means for collecting the user's voice instructions, means for analyzing the voice instruction data, and means for communicating the analysis results to the pet. This allows the pet's cries and the user's voice instructions to be accurately analyzed, enabling smooth communication between the pet and the user.

[0537] "Audio data" is data in which sounds made by a pet or a user are recorded in digital format.

[0538] "Animal characteristics" refers to information that refers to the unique characteristics of each individual animal, such as the pet's personality or the sounds it makes.

[0539] "Animal type" refers to the biological classification to which a pet belongs, such as a dog or a cat.

[0540] "General animal characteristic data" is data that includes information such as behaviors and sounds common to a particular type of animal.

[0541] "Means for understanding the state of a pet from its cries" refers to a method or device for analyzing a pet's cries and identifying the state or needs of the pet indicated by the cries.

[0542] The "means for analyzing collected voice data" refers to a method or device for converting collected voice data into text data and analyzing the content of the text data.

[0543] The "means for notifying the user of the analysis results" refers to a method or device for conveying the analyzed information to the user.

[0544] The "means for collecting user's voice instructions" refers to a method or device for recording the voice instructions given by the user and collecting the data.

[0545] The "means for analyzing voice instruction data" refers to a method or device for converting a user's voice instruction into text data and analyzing the content of that text data.

[0546] The "means for communicating the analysis results to the pet" refers to a method or device for communicating the analyzed user's instructions in a form that the pet can understand.

[0547] This invention is a system for realizing smooth communication between a pet and a user. This system includes a series of processes for collecting and analyzing voice data and communicating the results to the user and the pet.

[0548] Hardware and software used

[0549] Subject: Server

[0550] The server uses a speech recognition API and a natural language processing (NLP) model to analyze the pet's cries and the user's voice commands. Specifically, it uses Google Cloud's Speech-to-Text API to convert the voice data into text data, and then analyzes the text data using the NLP model. As a result of the analysis, it identifies the pet's status, requests, and user commands, and conveys them to the user and the pet in an appropriate manner.

[0551] Subject: Terminal

[0552] The device is equipped with a microphone to collect the pet's cries and the user's voice commands. The collected voice data is sent to a server via the Internet. It also has a display and speaker to notify the user of the analysis results received from the server. For example, a smartphone or tablet can play this role.

[0553] Subject: User

[0554] The user inputs the sound of their pet barking into the terminal and also issues their own voice commands to the terminal. For example, if their pet barks "woof woof," the user inputs the sound into the terminal, and the terminal sends the voice data to the server. Similarly, if the user says "let's play," the voice command is sent to the server via the terminal.

[0555] Specific examples

[0556] Example 1: Pet cry analysis

[0557] 1. Your pet barks "woof woof."

[0558] 2. The device's microphone captures the bark.

[0559] 3. The device sends the bird call data to the server.

[0560] 4. The server converts the audio data into text using Google Cloud's Speech-to-Text API.

[0561] 5. The server uses an NLP model to analyze the text data and generate a message saying "Your pet is hungry."

[0562] 6. The server generates a message and sends it to the device.

[0563] 7. The device displays a notification to the user that "Your pet is hungry."

[0564] Example 2: User voice command analysis

[0565] 1. The user says "Let's play."

[0566] 2. The device's microphone captures the audio.

[0567] 3. The device sends the audio data to the server.

[0568] 4. The server converts the audio data into text using Google Cloud's Speech-to-Text API.

[0569] 5. The server uses an NLP model to analyze the text data and convert the command "Let's play" into a form that the pet can understand.

[0570] 6. The server sends the converted instructions to the terminal.

[0571] 7. The device will play a specific sound to tell your pet to "play."

[0572] Prompt Sentence Examples

[0573] "If your pet barks 'woof woof,' we can analyze the sound and tell you what condition your pet is in."

[0574] "When the user says 'Let's play,' analyze that speech and generate commands for the pet."

[0575] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0576] Step 1: Collect your pet sounds

[0577] Subject: Terminal

[0578] The device collects the pet's barks through a microphone. For example, if the pet barks "woof woof," the device's microphone captures the sound. The input is the pet's bark, and the output is digital audio data.

[0579] Step 2: Send the bird sound data to the server

[0580] Subject: Terminal

[0581] The terminal sends the collected sound data to a server via the Internet. The data is sent in audio file format (e.g., WAV format). The input is digital audio data, and the output is the audio data sent to the server.

[0582] Step 3: Analyze the bird call data on the server

[0583] Subject: Server

[0584] The server converts the received bark data into text data using a speech recognition API. Specifically, it uses Google Cloud's Speech-to-Text API. The input is audio data and the output is text data. The NLP model then analyzes the text data to identify the pet's status and requests. For example, it analyzes that a bark of "woof woof" means "I'm hungry." The input is text data and the output is the analysis result.

[0585] Step 4: Notify the user of the analysis results

[0586] Subject: Server

[0587] The server sends the analysis results to the user's device. For example, it generates a message saying "your pet is hungry" and notifies the user's smartphone. The input is the analysis results, and the output is the notification message sent to the user's device.

[0588] Step 5: Collect user voice commands

[0589] Subject: Terminal

[0590] The device collects the user's voice commands through a microphone. For example, when the user says "Let's play," the device captures the voice. The input is the user's voice command, and the output is digital voice data.

[0591] Step 6: Send voice instructions to the server

[0592] Subject: Terminal

[0593] The terminal transmits the collected voice instruction data to a server via the Internet. The data is transmitted in audio file format (e.g., WAV format). The input is digital audio data, and the output is the audio data transmitted to the server.

[0594] Step 7: Parse the voice command data on the server

[0595] Subject: Server

[0596] The server converts the received voice instruction data into text data using a speech recognition API. Specifically, it uses Google Cloud's Speech-to-Text API. The input is voice data and the output is text data. The text data is then analyzed using an NLP model and converted into a form that the pet can understand. For example, the instruction "Let's play" is converted into sounds and action instructions that the pet can understand. The input is text data and the output is the analysis results.

[0597] Step 8: Communicate the results to your pet

[0598] Subject: Terminal

[0599] Based on the analysis results received from the server, the device generates sounds and behavioral instructions that the pet can understand and communicates them to the pet. For example, the device can play a specific sound to tell the pet to "play." The input is the analysis results, and the output is the instruction to be communicated to the pet.

[0600] (Application example 2)

[0601] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0602] Conventional methods of communication between pets and their owners rely on owners' intuitive understanding of their pets' sounds and behavior, making accurate communication difficult. It's also difficult to accurately grasp a pet's health and emotional state, which can delay appropriate responses. Furthermore, in brick-and-mortar stores like pet shops, staff are required to quickly understand a pet's condition and take appropriate action, but currently there is a lack of effective methods for doing so.

[0603] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0604] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from its cries based on the collected data, means including an emotion engine for recognizing the user's emotions, means for adjusting instructions to the pet based on the user's emotions, means for analyzing the pet's cries and determining the pet's condition and requests, means for converting human speech into a form the pet can understand, and means including an application to be installed on a smartphone. This enables accurate communication between pets and their owners, allowing the pet's health and emotional state to be quickly understood and appropriate measures to be taken. Furthermore, in physical stores, store staff can quickly determine the pet's condition and take appropriate measures.

[0605] "Audio data" refers to digitally recorded sound information, including animal sounds.

[0606] "Animal characteristics" refers to the individual characteristics of pet animals, such as their personalities and sounds.

[0607] "Means for identifying animal species" refers to methods or techniques for distinguishing between specific animal species.

[0608] "General animal characteristic data" refers to information about widely known animal characteristics.

[0609] "Means for understanding the condition of pets from their cries" refers to methods and techniques for analyzing pet cries to understand their health and emotional state.

[0610] An "emotion engine that recognizes the user's emotions" refers to technology that analyzes emotions from the user's tone of voice, facial expressions, behavior, etc.

[0611] The term "means for adjusting instructions to a pet based on the user's emotions" refers to a method or technology for changing instructions to a pet depending on the user's emotional state.

[0612] "Means for analyzing pet cries and understanding the pet's condition and requests" refers to methods and techniques for analyzing pet cries and understanding the pet's condition and requests.

[0613] "Means for converting human language into a form that pets can understand" refers to methods and techniques for converting human language into sounds and actions that pets can easily understand.

[0614] "Applications installed on a smartphone" refers to software programs that run on a smartphone.

[0615] A system for implementing this invention includes means for collecting audio data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for understanding the state of a pet from its cries based on the collected data, means including an emotion engine for recognizing the user's emotions, means for adjusting instructions to the pet based on the user's emotions, means for analyzing the pet's cries and understanding the pet's state and requests, means for converting human words into a form that the pet can understand, and means including an application to be installed on a smartphone.

[0616] The server collects voice data and has a database for identifying the characteristics and species of animals. This allows it to analyze the pet's cries and understand its status and requests. Furthermore, using an emotion engine that recognizes the user's emotions, it analyzes the user's tone of voice, facial expressions, and behavior, and adjusts instructions to the pet based on the results.

[0617] The device (smartphone) sends the collected voice data to the server and receives the analysis results. The analysis results are displayed to the user in a form that notifies them of the pet's status and requests. When the user gives instructions to the pet, the instructions are converted into a form that the pet can understand and communicated to the pet.

[0618] A specific example is a scenario in which a pet shop clerk uses a smartphone to analyze the cries of a pet and understand that the pet is saying, "I'm hungry." Another scenario is when a clerk says, "Let's play," and the command is communicated to the pet in a way that is easy for the pet to understand.

[0619] Examples of prompts to input to a generative AI model include:

[0620] "Analyze your pet's cries and let us know their condition."

[0621] "Translate human language into a form your pet can understand."

[0622] This system enables accurate communication between pets and their owners, allowing them to quickly understand the pet's health and emotional state and take appropriate action.It also enables store staff in physical stores to quickly understand the pet's condition and take appropriate action.

[0623] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0624] Step 1:

[0625] The server receives audio data (animal sounds) sent from the device (smartphone). The input is audio data, and the output is data ready for analysis. Specifically, the server stores the audio data in digital format and passes it to the analysis engine.

[0626] Step 2:

[0627] The server analyzes the audio data and identifies the characteristics and type of animal. The input is the audio data received in step 1, and the output is information about the animal's type and characteristics. Specifically, the server uses a voice recognition algorithm to convert the sound into text, and then searches a database for the animal's type and characteristics based on that text.

[0628] Step 3:

[0629] The server understands the pet's condition and requests based on the collected animal characteristic data. The input is the animal characteristic data obtained in step 2, and the output is information about the pet's condition and requests. Specifically, the server runs an algorithm to estimate the pet's health and emotional state based on the analysis results.

[0630] Step 4:

[0631] The server uses an emotion engine that recognizes the user's emotions to analyze the user's emotions from their tone of voice, facial expressions, and behavior. The input is the user's voice data and video data, and the output is information about the user's emotional state. Specifically, the server runs an emotion recognition algorithm to analyze the user's emotions.

[0632] Step 5:

[0633] The server adjusts the instructions to the pet based on the user's emotional state. The input is the information about the user's emotional state obtained in step 4, and the output is the adjusted instructions. Specifically, the server executes an algorithm that changes the instructions depending on the user's emotions.

[0634] Step 6:

[0635] The server analyzes the pet's cries and understands the pet's status and requests. The input is the information about the pet's status and requests obtained in step 3, and the output is a message to notify the user. Specifically, the server generates a message to notify the user based on the analysis results.

[0636] Step 7:

[0637] The server converts human speech into a form that the pet can understand. The input is the user's words, and the output is instructions converted into sounds and actions that the pet can understand. Specifically, the server uses a generative AI model to translate the user's words into a form that the pet can understand.

[0638] Step 8:

[0639] The device (smartphone) displays the analysis results and instructions received from the server to the user. The input is the message sent from the server, and the output is the information displayed to the user. Specifically, the device displays the received message on the screen and notifies the user.

[0640] Step 9:

[0641] The user checks the pet's status and requests through the terminal and takes necessary action. The input is the information displayed on the terminal, and the output is the user's actions. In concrete terms, the user takes appropriate action toward the pet based on the information displayed on the terminal.

[0642] Example 3

[0643] Next, a description will be given of a third embodiment of the third embodiment. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0644] Conventional systems for understanding a pet's health and emotional state mainly analyze only the pet's cries and behavior, and have the problem of being unable to adjust the pet's behavior to take the user's emotions into account. Furthermore, there are insufficient means for facilitating communication between the pet and the user. This has resulted in insufficient health management and emotional understanding of the pet, leading to problems with smooth communication with the user.

[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[0646] In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the state of the pet from its cries based on the collected data, means for collecting the user's emotions, means for analyzing the user's emotions, and means for adjusting the pet's behavior based on the results of the user's emotion analysis, thereby making it possible to more accurately understand the health and emotional state of the pet and adjust the pet's behavior in accordance with the user's emotions.

[0647] "Means for collecting audio data" refers to a function for collecting pet sounds using an audio input device such as a microphone.

[0648] "Means for collecting animal characteristics" refers to the function of collecting individual characteristics such as a pet's personality and cries using sensors, cameras, etc.

[0649] "Means for identifying animal species" refers to algorithms or identification technologies that identify the species of pet based on collected data.

[0650] "Means for collecting general animal characteristics data" refers to the ability to collect data on general animal characteristics from databases or external sources.

[0651] "Means of understanding the condition of your pet from its cries based on the collected data" refers to an algorithm that analyzes the collected audio data and characteristic data to determine the health and emotional state of your pet.

[0652] "Means for collecting user emotions" refers to the function of collecting the user's voice and facial expressions using a microphone or camera.

[0653] "Means for analyzing user emotions" refers to algorithms and analytical technologies that analyze collected user voice and facial expression data to determine the user's emotional state.

[0654] "Means for adjusting the behavior of a pet based on the results of analyzing the user's emotions" refers to a control algorithm or instruction function for appropriately adjusting the behavior of a pet based on the results of analyzing the user's emotions.

[0655] A "generative AI model" refers to an artificial intelligence model that analyzes the pet's condition and the user's emotions based on collected data.

[0656] The present invention is a system for understanding the health and emotional state of a pet and adjusting the behavior of the pet in accordance with the emotions of a user. Specific embodiments of this system will be described below.

[0657] Hardware and software used

[0658] Hardware: microphone, camera, sensors (to detect pet movement)

[0659] Software: Generative AI models (e.g., general artificial intelligence models), sentiment engines (e.g., general sentiment analysis APIs)

[0660] System configuration

[0661] 1. The device is equipped with a microphone to capture the sounds of your pet and a camera to capture images of your pet's behavior. These devices collect your pet's sounds and behavior data in real time.

[0662] 2. The device temporarily stores the collected data and sends it to the server at regular intervals.

[0663] 3. The server inputs the received data into the generative AI model to analyze the pet's health and emotional state. The generative AI model determines the pet's condition based on the collected voice and behavioral data.

[0664] 4. The server sends the analysis results to the device, which then notifies the user, for example by displaying a message on a smartphone app saying, "Your pet may be unwell."

[0665] 5. The device is equipped with a microphone to collect the user's voice and a camera to capture the user's facial expressions. These devices collect the user's voice and facial expression data.

[0666] 6. The device sends the collected user data to the server, and the server uses an emotion engine to analyze the user's emotions.

[0667] 7. The server sends the results of the user's emotion analysis to the device, which then adjusts the pet's behavior. For example, if the user is angry, the device will instruct the pet to "calm down."

[0668] Specific examples

[0669] Pet Health Analysis:

[0670] For example, if a pet is making an unusual sound, a generative AI model can detect the change and notify the owner that their pet may be unwell.

[0671] User emotion recognition and pet behavior regulation:

[0672] Example: If a user is angry with their pet, the emotion engine will recognize that anger and the pet state awareness will adjust the pet's behavior to calm it down.

[0673] Prompt Sentence Examples

[0674] Pet Health Analysis:

[0675] "Analyze your pet's sounds and behavior to determine their health status."

[0676] User emotion recognition and pet behavior regulation:

[0677] "Analyze the user's voice and facial expressions to recognize their emotions. Adjust your pet's behavior based on the results."

[0678] This system makes it possible to grasp the status of the pet and the user in real time and take appropriate measures. The flow of the identification process in the third embodiment will be described with reference to FIG.

[0679] Step 1:

[0680] The device collects the sounds your pet makes.

[0681] Input: Pet noises

[0682] How it works: A microphone on the device collects your pet's sounds in real time.

[0683] Output: Collected audio data

[0684] Step 2:

[0685] The device captures your pet's behavior.

[0686] Input: pet behavior

[0687] How it works: The camera on the device captures your pet's movements in real time.

[0688] Output: Collected video data

[0689] Step 3:

[0690] The data collected by the device is temporarily stored and sent to the server at regular intervals.

[0691] Input: Audio data, video data

[0692] Operation: The data collected by the device is temporarily stored in the internal memory and sent to the server in batch processing at regular intervals.

[0693] Output: Data sent to the server

[0694] Step 4:

[0695] The server inputs the received data into a generative AI model to analyze the pet's health and emotional state.

[0696] Input: Audio data, video data

[0697] How it works: The server inputs data into a generative AI model that analyzes your pet's health and emotional state.

[0698] Output: Analysis results (pet's health and emotional state)

[0699] Step 5:

[0700] The server sends the analysis results to the terminal, which then notifies the user.

[0701] Input: Analysis results

[0702] How it works: The server sends the analysis results to the device, which then sends a push notification to the user's smartphone.

[0703] Output: Analysis results reported to the user

[0704] Step 6:

[0705] The device collects user feedback.

[0706] Input: User's voice

[0707] How it works: A microphone on the device collects the user's voice.

[0708] Output: Collected audio data

[0709] Step 7:

[0710] The device captures the user's facial expression.

[0711] Input: User's facial expression

[0712] How it works: The device's camera captures the user's facial expression.

[0713] Output: Collected video data

[0714] Step 8:

[0715] The terminal transmits the collected user data to the server.

[0716] Input: Audio data, video data

[0717] Operation: The device sends the collected data to the server.

[0718] Output: Data sent to the server

[0719] Step 9:

[0720] The server analyzes the user's emotions using an emotion engine.

[0721] Input: Audio data, video data

[0722] How it works: The server inputs data into the emotion engine to analyze the user's emotional state.

[0723] Output: Analysis result (user's emotional state)

[0724] Step 10:

[0725] The server sends the results of the user's emotion analysis to the terminal, which then adjusts the pet's behavior.

[0726] Input: Analysis results

[0727] Operation: The server sends the analysis results to the device, and the device gives the appropriate instructions to the pet.

[0728] Output: Adjusted pet behavior

[0729] In this way, the system can grasp the status of the pet and the user in real time and take appropriate action.

[0730] (Application example 3)

[0731] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0732] Conventional pet care systems have difficulty accurately understanding the health and emotional state of pets, making it difficult for owners and store staff to properly manage their pets. Furthermore, they lack the functionality to adjust the pet's behavior based on the user's emotions, which hinders smooth communication between the pet and the user.

[0733] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for collecting audio data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for understanding the pet's condition from its cries based on the collected data, means for collecting and analyzing image data, means for recognizing the user's emotions, and means for adjusting the pet's behavior based on the user's emotions. This makes it possible to accurately understand the pet's health condition and emotional state and adjust the pet's behavior in accordance with the user's emotions.

[0734] "Audio data" refers to audio information such as animal cries, and is data collected to analyze the condition of pets.

[0735] "Animal characteristics" refers to the unique characteristics of each animal, such as personality and vocalizations, and is information collected to understand the condition of your pet.

[0736] "Means for identifying animal species" refers to technology that identifies the type of animal a pet is based on collected data.

[0737] "General animal characteristic data" is data on widely known animal characteristics, and is information used as a basis for analyzing the condition of pets.

[0738] "Means for understanding the condition of pets from their cries" is a technology that analyzes collected cry data to determine the health and emotional state of pets.

[0739] The "means for collecting and analyzing image data" refers to a technology for collecting images of a pet and analyzing the images to understand the pet's condition.

[0740] "Means for recognizing user emotions" refers to technology that analyzes the user's tone of voice, facial expressions, etc. to determine the user's emotional state.

[0741] The "means for adjusting the behavior of a pet based on the user's emotions" is a technique for appropriately changing the behavior of a pet in accordance with the recognized emotions of the user.

[0742] As an embodiment of the present invention, a pet care assistant system will be described. This system is intended for use in physical stores such as pet shops and pet cafes. The system is installed on a smartphone, smart glasses, or a head-mounted display, and analyzes the health and emotional state of pets in real time, notifying store staff and owners.

[0743] Hardware and software used

[0744] Hardware

[0745] Smartphone

[0746] Smart Glasses

[0747] head-mounted display

[0748] software

[0749] OpenCV: Image processing library

[0750] Keras: A deep learning library

[0751] librosa: Audio processing library

[0752] soundfile: Sound file reading library

[0753] Data processing and calculation

[0754] Analysis of audio data

[0755] The server collects audio data (animal sounds) and extracts MFCC features using librosa. This converts the audio data into numerical data and feeds it into a model trained using Keras. The model predicts the pet's health and emotional state.

[0756] Image data analysis

[0757] The server collects image data, resizes and preprocesses it using OpenCV, and then feeds the preprocessed image data into a model trained using Keras to analyze the pet's condition.

[0758] User Emotion Recognition

[0759] The server collects the user's video data and extracts frames using OpenCV, which are then fed into an emotion recognition model trained with Keras to predict the user's emotional state.

[0760] Pet behavior adjustment

[0761] The server adjusts the pet's behavior based on the user's emotional state, for example, adjusting the pet's behavior to be calmer if the user is angry, and increasing the pet's behavior if the user is happy.

[0762] Specific examples

[0763] Usage example at a pet shop

[0764] If a pet in a pet shop is feeling stressed, the system will analyze audio and image data to detect the pet's stress state, alerting store staff and urging them to take better care of the pet.

[0765] Example of use at a pet cafe

[0766] At the pet cafe, the system analyzes the user's emotional state while playing with their pet. If the user is happy, the system will activate the pet's behavior and promote communication with the user.

[0767] Prompt Sentence Examples

[0768] Develop a system that analyzes your pet's sounds and behavior to notify you of its health and emotional state. Include the ability to recognize your emotions and adjust your pet's behavior accordingly.

[0769] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[0770] Step 1:

[0771] The server collects pet voice data. Specifically, it records pet cries using microphones installed in smartphones, smart glasses, and head-mounted displays. The input is voice data, and the output is a recorded voice file.

[0772] Step 2:

[0773] The server analyzes the collected audio data. It uses librosa to extract MFCC features from the audio data and convert them into numerical data. The input is the recorded audio file, and the output is numerical data including MFCC features.

[0774] Step 3:

[0775] The server inputs MFCC features into a voice analysis model trained using Keras to predict the pet's health and emotional state. The input is numerical data including MFCC features, and the output is a prediction result indicating the pet's health and emotional state.

[0776] Step 4:

[0777] The server collects image data of pets. Images of pets are taken using a camera mounted on a smartphone, smart glasses, or head-mounted display. The input is image data, and the output is the captured image file.

[0778] Step 5:

[0779] The server analyzes the collected image data. It resizes and preprocesses the images using OpenCV and inputs them into an image analysis model trained using Keras. The input is the captured image file, and the output is the preprocessed image data.

[0780] Step 6:

[0781] The server inputs the preprocessed image data into a model trained with Keras to analyze the pet's condition. The input is the preprocessed image data, and the output is a prediction result indicating the pet's condition.

[0782] Step 7:

[0783] The server collects the user's video data. It captures the user's facial expressions using a camera mounted on a smartphone, smart glasses, or head-mounted display. The input is video data, and the output is the captured video file.

[0784] Step 8:

[0785] The server analyzes the collected video data, extracts frames from the video using OpenCV, and inputs them into an emotion recognition model trained with Keras. The input is the captured video file, and the output is the extracted frame data.

[0786] Step 9:

[0787] The server inputs the extracted frame data into an emotion recognition model trained with Keras to predict the user's emotional state. The input is the extracted frame data, and the output is a prediction result indicating the user's emotional state.

[0788] Step 10:

[0789] The server adjusts the pet's behavior based on the user's emotional state. For example, if the user is angry, the server adjusts the pet's behavior to be calmer, and if the user is happy, the server adjusts the pet's behavior to be more active. The input is a prediction result indicating the user's emotional state, and the output is the adjusted pet's behavior.

[0790] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0791] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0792] Another example of generative AI is Gemini (registered trademark) (Internet search engine). <url: https: gemini.google.com ?hl="ja">) are listed.

[0793] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0794] [Second embodiment]

[0795] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0796] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0797] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0798] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0799] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0800] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0801] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0802] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0803] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0804] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0805] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0806] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[0807] "Example 1"

[0808] In one embodiment of the present invention, pet cries are collected using an audio data collection device such as a microphone. The collected audio data is sent to an audio analysis device and identified as the cry of a specific animal. In addition, the characteristics of the pet (personality, cries, etc.) are collected using input from the owner or sensors that observe the pet's behavior. The type of animal is identified using input from the owner or image recognition technology. General animal characteristic data is obtained from a database or the like. This data is analyzed by AI to understand the pet's condition.

[0809] "Example 2"

[0810] Another embodiment of the present invention provides a system that utilizes AI to enable conversations between people and pets. Specifically, AI analyzes a pet's cries and converts them into words that humans can understand. Conversely, AI analyzes human language and converts it into cries and actions that pets can understand. For example, if a pet expresses its hunger state through cries, the AI ​​analyzes it and conveys the message "your pet is hungry" to the owner. Also, if the owner says "let's play," the AI ​​analyzes it and conveys the message in a form that the pet can understand.

[0811] "Example 3"

[0812] As a further embodiment of the present invention, a system for understanding the health and emotional state of a pet is provided. Specifically, AI analyzes the health and emotional state of a pet from its sounds and behavior. For example, if a pet is suffering from an illness, its sounds and behavior may be different from normal. The AI ​​detects these changes and issues a warning to the owner. Also, if a pet is happy, its sounds and behavior may become more active. The AI ​​detects these changes and conveys the owner's joy.

[0813] The processing flow of each embodiment will be described below.

[0814] "Example 1"

[0815] Step 1: Collect your pet's cries using a sound data collection device such as a microphone.

[0816] Step 2: The collected audio data is sent to an audio analyzer to identify it as the sound of a specific animal.

[0817] Step 3: Collect information about pet characteristics (personality, vocalizations, etc.) using input from owners and sensors that observe pet behavior.

[0818] Step 4: Identify the animal's species using input from the owner and image recognition technology.

[0819] Step 5: Obtain general animal characteristic data from a database, etc.

[0820] Step 6: This data is analyzed using AI to understand the pet's condition.

[0821] "Example 2"

[0822] Step 1: AI analyzes your pet's cries and converts them into words that humans can understand. Step 2: AI analyzes human words and converts them into cries and actions that your pet can understand.

[0823] Step 3: For example, if a pet expresses its hunger state through a cry, the AI ​​will analyze it and convey the message to the owner that "your pet is hungry."

[0824] Step 4: Also, if the owner says "Let's play," the AI ​​will analyze it and convey the message in a way that the pet can understand.

[0825] "Example 3"

[0826] Step 1: AI analyzes your pet's health and emotional state based on its sounds and behavior.

[0827] Step 2: For example, if a pet is suffering from an illness, its sounds and behavior may be different from normal. AI can detect these changes and alert the owner.

[0828] Step 3: Also, if your pet is happy, its vocalizations and behavior may become more active. AI will detect these changes and communicate its happiness to its owner.

[0829] Example 1

[0830] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0831] Conventional pet condition monitoring systems mainly collect and analyze voice data and animal characteristic data separately, making it difficult to integrate this data to comprehensively understand the pet's condition.In addition, there was a lack of means to accurately understand the pet's health and emotional state, making it difficult for owners to appropriately understand and respond to their pet's condition.

[0832] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0833] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data, means for analyzing the voice data, means for collecting characteristics of pet animals, means for identifying the type of animal, means for collecting general animal characteristic data, and means for understanding the condition of pet animals from their cries based on the collected data. This makes it possible to comprehensively understand the condition of pets by integrating the voice data and animal characteristic data.

[0834] "Audio data" refers to data in which audio signals such as animal cries are recorded in digital format.

[0835] "Collecting means" refers to a device or method for acquiring audio data and animal characteristic data.

[0836] "Transmitting means" refers to a device or method for transferring collected data to another device or server.

[0837] An "analyzing means" is a device or method for processing collected data and extracting specific information.

[0838] "Animal characteristics" refers to individual attributes of pet animals, such as their personalities and sounds.

[0839] "Means for identification" refers to devices or methods for determining the type or characteristics of an animal.

[0840] "General animal characteristic data" refers to data such as characteristics and behavioral patterns common to a particular animal species.

[0841] "Pet condition" refers to the overall condition of the pet, including its health and emotional state.

[0842] "AI" is a technology that uses artificial intelligence to analyze data and make specific judgments and predictions.

[0843] "Means for realizing conversation" refers to devices or methods that allow people and pets to communicate.

[0844] "Health Status" refers to information about the physical health of your pet.

[0845] "Emotional state" refers to a pet's psychological state or mood.

[0846] The present invention is a system for collecting sounds of pets and analyzing the data to understand the state of the pets. A specific embodiment of this system will be described below.

[0847] First, a user uses the smartphone's built-in microphone or a dedicated external microphone to collect the sound of their pet. For example, the user starts a recording app on their smartphone and records the sound of their pet. This audio data is then sent to the server by the device. Specifically, the device sends the audio data to the server's API endpoint using an HTTP POST request.

[0848] The server analyzes the received audio data using a voice analyzer (for example, Google Cloud Speech-to-Text API). The server sends the audio data to the API, analyzes the returned text data, and identifies it as the sound of a specific animal.

[0849] Next, the user enters the characteristics of their pet (personality, sounds, etc.) into the app. For example, the user might enter information such as "Personality: Active, Sound: High-pitched" into the app's form. The app also uses sensors connected to the device (such as an accelerometer or camera) to observe the pet's behavior and collect data. The device periodically sends the data from the sensors to a server.

[0850] To identify the type of animal, the user can either input the type of pet into the app, or the device can use image recognition technology (such as TensorFlow or OpenCV) to identify the type of animal. For example, the user can input "dog," or the device can send an image taken with its camera to the server, which then uses image recognition technology to identify it as a "dog."

[0851] The server retrieves general animal trait data from a database (e.g., AWS RDS or Google BigQuery). The server runs SQL queries to retrieve the required data from the database.

[0852] Finally, the server analyzes the collected voice data, pet characteristic data, animal species data, and general animal characteristic data using AI (for example, TensorFlow or PyTorch). The server integrates this data and inputs it into an AI model to understand the pet's condition. For example, the server analyzes data such as "the dog is barking loudly and is behaving actively" and determines that "the dog is excited."

[0853] Examples and prompts

[0854] As a concrete example, the bark of a user's dog is collected using the microphone on their smartphone and sent to a server. The server analyzes the bark using the Google Cloud Speech-to-Text API and identifies it as a dog's bark. The user enters the dog's personality and characteristics of its bark into the app and observes the dog's behavior using a camera connected to the device. The server retrieves data on general dog characteristics from AWS RDS, analyzes the data using TensorFlow, and understands the dog's condition.

[0855] An example of a prompt sentence is as follows:

[0856] "Please analyze my dog's barks to understand his condition. My dog ​​has an active personality and barks a lot. Please analyze the following audio data."

[0857] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0858] Step 1:

[0859] A user collects the sounds of their pets.

[0860] Specifically, the user launches a recording app on their smartphone and records the sound of their pet's cries.

[0861] Input: Pet noises

[0862] Output: Audio data file (e.g. .wav format)

[0863] Step 2:

[0864] The device sends the collected voice data to the server.

[0865] Specifically, the device sends audio data to the server's API endpoint using an HTTP POST request.

[0866] Input: Audio data file

[0867] Output: Audio data sent to the server

[0868] Step 3:

[0869] The server analyzes the received audio data.

[0870] Specifically, the server uses a voice analysis device (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text data and identify it as the sound of a specific animal.

[0871] Input: Audio data sent to the server

[0872] Output: Parsed text data (e.g. "dog barking")

[0873] Step 4:

[0874] Collect characteristics of the animals that the user keeps.

[0875] Specifically, the user enters information such as "personality: active, bark: high-pitched" into a form on the app. The app also uses sensors connected to the device (e.g., accelerometer and camera) to observe the pet's behavior and collect data.

[0876] Input: User-entered animal feature data, behavior data from sensors

[0877] Output: Collected animal trait data

[0878] Step 5:

[0879] Identify types of animals.

[0880] Specifically, the user enters the type of animal they own into the app, or the device uses image recognition technology (e.g., TensorFlow or OpenCV) to identify the type of animal.

[0881] Input: Animal species data or image data entered by the user

[0882] Output: Identified animal type data (e.g. "dog")

[0883] Step 6:

[0884] The server obtains general animal characteristic data.

[0885] Specifically, the server executes an SQL query from a database (e.g., AWS RDS or Google BigQuery) to retrieve the required data.

[0886] Input: Animal species data

[0887] Output: General animal traits data

[0888] Step 7:

[0889] The server analyzes the collected data and determines the pet's condition.

[0890] Specifically, the server analyzes collected voice data, animal characteristics data, animal species data, and general animal characteristics data using AI (e.g., TensorFlow or PyTorch). The server integrates this data and inputs it into an AI model to determine the pet's condition.

[0891] Input: Audio data, animal characteristics data, animal species data, general animal characteristics data

[0892] Output: Analysis result (e.g. "The dog is excited")

[0893] (Application example 1)

[0894] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0895] It is difficult for pet owners to monitor their pets' abnormal behavior and health status in real time while they are out or not watching over them. In addition, there is a lack of means to detect abnormalities in pets early and respond quickly, making it difficult to ensure the health and safety of pets.

[0896] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0897] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying the type of animal, means for collecting general animal characteristic data, means for understanding the pet's condition from its cries based on the collected data, and means for sending a notification when an abnormality is detected. This makes it possible to understand the pet's abnormal behavior and health condition in real time, and to quickly notify the owner when an abnormality is detected.

[0898] "Audio data" refers to information collected in digital format from audio signals such as animal cries.

[0899] "Collecting means" refers to a device or method for acquiring audio data and animal characteristic data.

[0900] "Animal characteristics" refers to information that refers to the animal's unique characteristics, such as its personality and cries.

[0901] "Means for identification" refers to devices or methods for determining the type or characteristics of an animal.

[0902] "General animal characteristic data" refers to information such as personality and behavioral patterns common to a specific animal species.

[0903] "Means of understanding" refers to devices and methods used to analyze collected data and understand the condition of animals.

[0904] "Means for detecting abnormalities" refers to devices or methods for identifying abnormalities in animal behavior or vocalizations and detecting the occurrence of an abnormality.

[0905] The "means for sending a notification" refers to a device or method for conveying information to the owner or other person when an abnormality is detected.

[0906] The system for carrying out this invention collects and analyzes the sounds of pets to understand their condition, and if an abnormality is detected, sends a notification to the pet owner. A specific embodiment of the system will be described below.

[0907] Hardware Configuration

[0908] The system includes the following hardware:

[0909] Microphone: Used to collect pet sounds.

[0910] Smartphone or smart glasses: Used to analyze collected voice data and detect anomalies.

[0911] Server: Analyzes voice data and sends notifications.

[0912] Software Configuration

[0913] The system includes the following software:

[0914] TensorFlow: Runs AI models to analyze audio data.

[0915] sounddevice: Used to collect audio data.

[0916] smtplib: Used to send notifications when an anomaly is detected.

[0917] Data processing and calculation

[0918] 1. Audio data collection: The device (smartphone or smart glasses) uses a microphone to collect the pet's cries. The collected audio data is stored digitally.

[0919] 2. Audio data analysis: The server inputs the collected audio data into the TensorFlow model for analysis. Based on the analysis results, the pet's condition is determined.

[0920] 3. Anomaly detection: The server determines whether an anomaly has been detected based on the analysis results. If an anomaly is detected, it sends a notification to the owner.

[0921] 4. Sending notifications: The server uses smtplib to send notifications to the pet owner, including information about the pet's abnormal behavior and health status.

[0922] Specific examples

[0923] For example, if a pet makes an unusual noise while the owner is out, the system will collect and analyze the sound. If an abnormality is detected based on the analysis results, a notification will be sent to the owner's smartphone, allowing the owner to take prompt action.

[0924] Prompt Sentence Examples

[0925] "Collect pet sounds and analyze them using an AI model. If an abnormality is detected, create a program that notifies the owner."

[0926] In this way, the present invention makes it possible to grasp abnormal behavior and health conditions of pets in real time, and to promptly notify the owner if an abnormality is detected.

[0927] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0928] Step 1:

[0929] The device (smartphone or smart glasses) uses a microphone to collect the pet's barks. The input is the pet's barks, and the output is digital audio data. Specifically, the device uses software for audio data collection (sounddevice library) to record audio for a certain period of time and save it as digital data.

[0930] Step 2:

[0931] The device sends the collected voice data to a server. The input is digital voice data, and the output is the voice data sent to the server. Specifically, the device uses an internet connection to upload the voice data to the server.

[0932] Step 3:

[0933] The server inputs the received audio data into the TensorFlow model for analysis. The input is digital audio data, and the output is the analysis result (the pet's state). Specifically, the server uses the TensorFlow library to input the audio data into the AI ​​model and analyze the characteristics of the pet's cries.

[0934] Step 4:

[0935] The server determines whether an anomaly has been detected based on the analysis results. The input is the analysis results, and the output is information about whether an anomaly has been detected. Specifically, the server compares the analysis results with a threshold and sets a flag if an anomaly is detected.

[0936] Step 5:

[0937] If an abnormality is detected, the server sends a notification to the pet owner. The input is information about whether an abnormality exists, and the output is a notification to the pet owner. Specifically, the server uses the smtplib library to send an abnormality notification to the pet owner's email address. The notification contains detailed information about the pet's abnormal behavior and health status.

[0938] Step 6:

[0939] The user (owner) checks the received notification and takes action to check the pet's status if necessary. The input is the notification from the server, and the output is the user's response. Specifically, the user checks the notification on their smartphone and takes action such as returning home to check the pet's status or contacting the pet sitter.

[0940] Example 2

[0941] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0942] Conventional methods of communication between pets and humans have been limited, making it difficult to accurately understand a pet's intentions and state from its cries and behavior. Furthermore, there has been a lack of ways to communicate human language to pets in a way that they can understand. This has made it difficult to properly manage a pet's health and emotional state, resulting in insufficient communication between pets and their owners.

[0943] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the status of the pet from its cries based on the collected data, means for converting the voice data into text data using voice recognition software, means for analyzing the text data using a generative AI model to understand the pet's intentions, means for generating a message based on the analysis result and notifying the user, means for collecting the user's words and converting it into text data using voice recognition software, means for analyzing the text data using a generative AI model to generate a message to be conveyed to the pet, and means for converting the message into cries and actions that the pet can understand using voice synthesis software and transmitting it. This makes it possible to accurately understand the intentions and status of the pet from its cries and actions, and to convey human language in a form that the pet can understand.

[0944] "Audio data" refers to data that has been recorded in digital format from sounds made by animals or humans.

[0945] "Animal characteristics" refers to characteristics unique to individual animals, such as their personalities, sounds, and behavioral patterns.

[0946] "Type of animal" indicates the classification of animals, such as dog, cat, bird, etc.

[0947] "General animal characteristic data" refers to data such as personality and behavioral patterns common to a particular animal species.

[0948] "Speech recognition software" is software that analyzes voice data and converts it into text data.

[0949] A "generative AI model" is a model that uses artificial intelligence to analyze text data and understand its meaning.

[0950] "Text data" is text information converted from voice data by voice recognition software.

[0951] "Message" refers to the content of notifications or instructions generated based on the analysis results.

[0952] "Speech synthesis software" is software for converting text data into speech.

[0953] "User" refers to a person who uses this system.

[0954] "Pet" refers to an animal kept in a household.

[0955] This invention is a system that uses AI to enable conversation between humans and pets. This system has the functions of analyzing pet cries and converting them into words that humans can understand, and analyzing human words and converting them into cries and actions that pets can understand.

[0956] Hardware and software used

[0957] Hardware

[0958] Microphone: Used to capture pet sounds and user speech.

[0959] Speaker: Used to transmit generated sounds to your pet.

[0960] Camera: Used to monitor pet behavior if necessary.

[0961] Server: Used to analyze and process data.

[0962] software

[0963] Speech recognition software: used to convert voice data into text data (e.g., Google Speech-to-Text).

[0964] Generative AI models: Analyze text data and use it to understand your pet's intentions (e.g., GPT-4).

[0965] Text-to-speech software: Used to convert text data into speech (e.g., Google Text-to-Speech).

[0966] Specific operation of the system

[0967] Pet cry analysis

[0968] The device's microphone captures the sound of your pet barking. For example, if your dog barks, "woof woof," the audio data is collected through the microphone and sent to the server in real time.

[0969] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice "woof woof" is converted into the text "woof woof." The converted text data is sent to the next analysis step.

[0970] The server uses a generative AI model to analyze the text data and understand the pet's intentions. For example, the text "woof woof" is interpreted as meaning "I'm hungry." The analysis results are sent to the message generation step.

[0971] Based on the analysis results, the server generates a message saying "your pet is hungry" and notifies the user. For example, a notification is sent to the user via a smartphone app. The user can check the notification on the app and understand the status of their pet.

[0972] Human language analysis

[0973] The microphone on the device captures the user's words. For example, when the user says "Let's play," the voice data is collected through the microphone. The collected voice data is sent to the server in real time.

[0974] The server analyzes the captured voice data using speech recognition software and converts it into text data. For example, the voice saying "Let's play" is converted into the text "Let's play." The converted text data is sent to the next analysis step.

[0975] The server uses a generative AI model to analyze the text data and generate a message to convey to the pet. For example, the text "Let's play" is analyzed to mean "It's play time." The analysis result is sent to the message generation step.

[0976] The server uses speech synthesis software to convert the generated message into sounds and actions that the pet can understand, and communicates them to the pet through a speaker. For example, the message "It's playtime" is converted into sounds and actions that a dog can understand, and played back from the speaker. The pet understands the sounds and actions, and prepares to play with the user.

[0977] Examples of specific examples and prompts

[0978] Specific examples

[0979] Pet Sound Analysis:

[0980] If the pet barks "woof woof," the microphone captures the sound and the server notifies the user that "your pet is hungry."

[0981] Human language analysis:

[0982] When the user says "let's play," a microphone captures the voice and the server communicates through the speaker with sounds and actions that the pet understands as "play time."

[0983] Prompt Sentence Examples

[0984] Pet Sound Analysis:

[0985] If your pet barks "woof woof," explain the process of analyzing that sound and informing the user that "your pet is hungry."

[0986] Human language analysis:

[0987] When a user says "Let's play," describe the process for analyzing that speech and communicating "It's playtime" to your pet in a way that makes sense.

[0988] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0989] Step 1: Capture your pet's sounds

[0990] The device's microphone captures the sound of your pet barking. For example, if your dog barks "woof woof," the audio data is collected through the microphone. The input is your pet's bark, and the output is audio data. The collected audio data is sent to the server in real time.

[0991] Step 2: Analyzing bird calls and converting them into text

[0992] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice "woof woof" is converted into the text "woof woof." The input is voice data and the output is text data. The converted text data is sent to the next analysis step.

[0993] Step 3: Semantic analysis of the text

[0994] The server uses a generative AI model to analyze the text data and understand the pet's intention. For example, the text "woof woof" is analyzed to mean "I'm hungry." The input is the text data, and the output is the analysis result indicating the pet's intention. The analysis result is sent to the message generation step.

[0995] Step 4: Message generation and notification

[0996] The server generates a message based on the analysis results, saying "Your pet is hungry," and notifies the user. For example, the notification is sent to the user via a smartphone app. The input is the analysis results, and the output is a notification message to the user. The user can check the notification on the app to understand the status of their pet.

[0997] Step 5: Capturing human language

[0998] The microphone on the device captures the user's words. For example, when a user says "Let's play," the voice data is collected through the microphone. The input is the user's words, and the output is voice data. The collected voice data is sent to the server in real time.

[0999] Step 6: Word analysis and text conversion

[1000] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice saying "Let's play" is converted into the text "Let's play." The input is voice data and the output is text data. The converted text data is sent to the next analysis step.

[1001] Step 7: Semantic analysis of the text

[1002] The server uses a generative AI model to analyze the text data and generate a message to be conveyed to the pet. For example, the text "Let's play" is analyzed to mean "It's play time." The input is text data, and the output is the message to be conveyed to the pet. The analysis result is sent to the message generation step.

[1003] Step 8: Generate a message and send it to your pet

[1004] The server uses speech synthesis software to convert the generated message into sounds and actions that the pet can understand, and communicates them to the pet through a speaker. For example, the message "It's playtime" is converted into sounds and actions that the dog can understand, and played back from the speaker. The input is the message to be communicated to the pet, and the output is sounds and actions that the pet can understand. The pet understands the sounds and actions, and gets ready to play with the user.

[1005] (Application example 2)

[1006] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1007] Conventional pet care systems have had problems such as difficulty in accurately understanding pet sounds and behavior and taking appropriate action. They also lacked a means to notify owners when their pets were hungry and quickly order pet food. Furthermore, they lacked an effective means to facilitate smooth communication between owners and their pets.

[1008] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from the cries based on the collected data, means for notifying the owner based on the pet's condition, and means for ordering pet food. This makes it possible to accurately determine the pet's condition and respond quickly. It also facilitates communication between the owner and the pet and automates pet food ordering.

[1009] "Audio data" refers to sound information, including animal sounds.

[1010] "Characteristics of pets" refers to individual characteristics such as the personality and sounds of pets.

[1011] "Means for identifying the type of animal" refers to a method or device for identifying the type of pet.

[1012] "General animal characteristic data" refers to information about widely known animal characteristics.

[1013] "Means for understanding the condition of pets from their cries" refers to methods and devices for analyzing pet cries to understand their condition.

[1014] "Means for notifying owners based on the status of their pets" refers to methods or devices for informing owners of the status of their pets.

[1015] "Pet food ordering means" refers to a method or device for purchasing pet food online.

[1016] "Means for enabling conversation between humans and pets using AI" refers to methods and devices for exchanging information between humans and pets using artificial intelligence.

[1017] "Means for understanding the health and emotional state of a pet" refers to a method or device for understanding the physical condition and emotions of a pet.

[1018] The following system configuration will be described as an embodiment of the present invention.

[1019] System Configuration

[1020] This system includes a means for collecting audio data (animal cries), a means for collecting characteristics of pet animals (personality, cries, etc.), a means for identifying the type of animal, a means for collecting general data on animal characteristics, a means for determining the condition of pets from their cries based on the collected data, a means for notifying owners based on the condition of their pets, and a means for ordering pet food.

[1021] Hardware and software used

[1022] Hardware: Smartphone, microphone

[1023] Software: Python, SpeechRecognition library, Requests library

[1024] Data processing and calculation

[1025] Audio data collection and analysis

[1026] The server records the pet's cries using the smartphone's microphone. The recorded audio data is converted to text using the SpeechRecognition library. It is then analyzed using a generative AI model to understand the pet's condition.

[1027] Collecting and analyzing owner speech

[1028] The server uses the smartphone's microphone to record the owner's words, converts the recorded voice data into text using the SpeechRecognition library, and then analyzes it using a generative AI model to generate a message to be conveyed to the pet.

[1029] Pet food orders

[1030] The server notifies the owner based on the pet's condition. For example, if the pet cries "I'm hungry," the server generates a prompt to order pet food and notifies the owner. When the owner orders pet food through the app, recommended products are displayed taking into account the pet's preferences and allergies.

[1031] Specific examples

[1032] Analyzing pet cries: When a pet cries "I'm hungry," the server analyzes the cry and notifies the owner with a message that "Your pet is hungry."

[1033] Analyzing owner's words: When the owner says "Let's play," the server analyzes the words and conveys the message in a way that the pet can understand.

[1034] Prompt Sentence Examples

[1035] "Analyze your pet's cries and let us know their condition."

[1036] "Analyze the owner's words and generate a message to convey to the pet."

[1037] In this way, the present invention allows for accurate understanding of the pet's condition and prompt response, facilitates communication between pet owners and their pets, and automates pet food ordering.

[1038] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1039] Step 1:

[1040] The server records the pet's cries using the smartphone's microphone. The input is the pet's cries, and the output is an audio data file. This audio data file is used for subsequent analysis.

[1041] Step 2:

[1042] The server converts the recorded audio data into text data using the SpeechRecognition library. The input is an audio data file, and the output is text data. This text data is input into the generative AI model.

[1043] Step 3:

[1044] The server analyzes the text data using a generative AI model to understand the pet's state. The input is text data, and the output is information indicating the pet's state. For example, the analyzed state is "hungry."

[1045] Step 4:

[1046] The server notifies the owner based on the pet's status. The input is information indicating the pet's status, and the output is a notification message to the owner. For example, a message saying "your pet is hungry" is sent to the owner.

[1047] Step 5:

[1048] The server records the owner's words using the smartphone's microphone. The input is the owner's words, and the output is an audio data file. This audio data file is used for subsequent analysis.

[1049] Step 6:

[1050] The server converts the recorded voice data of the owner into text data using the SpeechRecognition library. The input is a voice data file, and the output is text data. This text data is input into the generative AI model.

[1051] Step 7:

[1052] The server uses a generative AI model to analyze the text data and generate a message to be conveyed to the pet. The input is text data, and the output is a message to be conveyed to the pet. For example, the command "Let's play" is analyzed.

[1053] Step 8:

[1054] The server provides a means for ordering pet food based on the pet's condition. The input is information indicating the pet's condition, and the output is pet food ordering information. For example, if the pet cries "I'm hungry," an order for pet food is automatically placed.

[1055] Step 9:

[1056] The server displays recommended products to the pet owner, taking into account the pet's preferences and allergy information. The input is the pet's preferences and allergy information, and the output is a list of recommended products. This allows the pet owner to select the appropriate pet food.

[1057] Step 10:

[1058] The server notifies the pet owner that the pet food order has been completed. The input is the order completion information, and the output is a notification message to the pet owner. For example, a message saying "Pet food order completed" is sent to the pet owner.

[1059] Example 3

[1060] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1061] Conventional systems for understanding pet health and emotional states only analyze voice data and do not take behavioral data into account, making it difficult to accurately understand a pet's condition. Additionally, there is a lack of a way to quickly notify the owner of the analysis results, making it difficult to detect abnormalities in a pet early on.

[1062] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[1063] In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for determining the status of pet animals from their cries based on the collected data, means for collecting animal behavior data, means for analyzing the collected behavior data, and means for notifying the results of the analysis. By analyzing both the voice data and the behavior data, it is possible to more accurately determine the health and emotional state of pets and promptly notify their owners.

[1064] "Audio Data" refers to animal sounds and other audio information.

[1065] "Animal characteristics" refers to individual characteristics such as an animal's personality, vocalizations, and behavioral patterns.

[1066] "Type of animal" refers to a classification of animals such as dogs, cats, birds, etc.

[1067] "General animal characteristic data" refers to data such as personality and behavioral patterns common to a particular type of animal.

[1068] "Animal condition" refers to the animal's physical and emotional state.

[1069] "Behavioral data" refers to information about an animal's movements and behavior patterns.

[1070] "Analysis results" refers to the diagnosis and evaluation output by the AI ​​model based on the collected voice data and behavioral data.

[1071] "Notification means" refers to a method or device for notifying the owner of the analysis results.

[1072] The present invention is a system for understanding the health and emotional state of a pet. A specific embodiment of this system will be described below.

[1073] First, the user uses a device such as a smartphone or tablet to collect the sounds and behavior of their pet. The device is equipped with a microphone and camera, and these devices are used to capture audio data and behavior data in real time. For example, when a user launches the smartphone app, points it at their pet, and presses the "Start Recording" button, the microphone begins to collect the pet's sounds. At the same time, the camera captures the pet's movements.

[1074] The device then transmits the collected data to a server in real time using Wi-Fi or mobile data, and the data is encrypted to ensure security.

[1075] The server inputs the received audio and behavior data into an AI model, which is built using machine learning frameworks such as TensorFlow and PyTorch. The AI ​​model analyzes the frequency and volume of the pet's cries, as well as the speed and direction of its movements, to detect abnormal patterns and patterns of delight.

[1076] The analysis results are sent from the server to the device, and the user can check the results through a smartphone app. For example, the server may send a warning message such as "Your pet may be sick." If the pet is happy, it may send a message such as "Your pet is happy."

[1077] As a concrete example, the following prompt sentence could be input to a generative AI model:

[1078] "Analyze your pet's vocalizations and behavioral data to determine their health and emotional state. Alert you if they're sick and let you know if they're happy."

[1079] Using this prompt, the AI ​​model can accurately analyze the pet's condition and provide appropriate feedback to the user.

[1080] As described above, by analyzing both the voice data and the behavioral data, the present invention makes it possible to grasp the health condition and emotional state of the pet more accurately and to notify the owner promptly. The flow of the identification process in the third embodiment will be described with reference to FIG.

[1081] Step 1:

[1082] The user launches the smartphone app and collects the sounds and behavior of their pet. Specifically, when the user presses the "Start Recording" button, the device's microphone begins collecting the pet's sounds. At the same time, the camera captures the pet's movements. The input is the pet's sounds and behavior, and the output is the collected audio data and behavior data.

[1083] Step 2:

[1084] The device transmits the collected voice and behavioral data to a server. Specifically, the device transmits data in real time using Wi-Fi or mobile data. The input is the collected voice and behavioral data, and the output is the data transmitted to the server.

[1085] Step 3:

[1086] The server inputs the received voice data and behavioral data into an AI model. Specifically, the server inputs the data into an AI model built using machine learning frameworks such as TensorFlow and PyTorch. The input is the voice data and behavioral data sent to the server, and the output is the analysis results by the AI ​​model.

[1087] Step 4:

[1088] The server determines the pet's health and emotional state based on the analysis results of the AI ​​model. Specifically, the AI ​​model analyzes the frequency and volume of the pet's cries, as well as the speed and direction of its movements, to detect abnormal patterns and patterns of joy. The input is the analysis results of the AI ​​model, and the output is a judgment of the pet's health and emotional state.

[1089] Step 5:

[1090] The server sends the analysis results to the device. Specifically, the server sends a message generated based on the analysis results to the device. The input is the judgment result of the pet's health condition and emotional state, and the output is the message sent to the device.

[1091] Step 6:

[1092] The user checks the analysis results through the device. Specifically, the user opens the smartphone app and checks the notified message. For example, a warning message such as "Your pet may be sick" or a message such as "Your pet is happy" may be displayed. The input is the message sent to the device, and the output is the analysis result checked by the user.

[1093] (Application example 3)

[1094] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1095] Understanding a pet's health and emotional state is important for pet owners, but existing methods make it difficult to detect abnormalities early on. There are also limited ways to accurately determine whether a pet is happy. This can lead to inadequate health management and understanding of a pet's emotions, potentially reducing the pet's quality of life. Furthermore, the lack of a way to quickly notify owners of abnormalities can delay emergency response.

[1096] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[1097] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from its cries based on the collected data, means for analyzing the pet's condition and sending a notification if an abnormality is detected, means for determining the pet's health and emotional state in real time, means for issuing an alert to the owner if an abnormality is detected, and means for informing the owner if the pet is happy. This makes it possible to quickly and accurately determine the pet's health and emotional state, and to immediately notify the owner if an abnormality occurs.

[1098] "Audio data" refers to data that has been recorded in digital format, such as animal cries.

[1099] "Animal characteristics" refers to the animal's unique characteristics, such as its personality or sounds.

[1100] "Means for identifying animal species" refers to methods or techniques for identifying the species of an animal.

[1101] "General animal characteristic data" refers to data about widely known animal characteristics.

[1102] "Means for understanding the condition of pets" refers to methods for analyzing the health and emotional state of pets based on collected data.

[1103] "Means for analyzing a pet's condition" refers to a method for analyzing a pet's cries and behavior to assess its condition.

[1104] "Means for sending a notification when an abnormality is detected" refers to a method for alerting the owner when an abnormality is detected in the pet.

[1105] "Means for understanding a pet's health and emotional state in real time" refers to a method for instantly analyzing a pet's condition and providing the results in real time.

[1106] "Means for issuing a warning to the owner when an abnormality is detected" refers to a method for issuing a warning to the owner when an abnormality is detected in the pet.

[1107] The "means for conveying information about a happy pet to its owner" refers to a method for detecting a happy state of a pet and conveying that information to its owner.

[1108] As an embodiment of the present invention, a pet security alert system will be described as an example. This system analyzes the sounds and behavior of pets in real time and notifies the owner if an abnormality is detected.

[1109] System Configuration

[1110] Hardware:

[1111] Smartphone

[1112] microphone

[1113] software:

[1114] TensorFlow

[1115] librosa

[1116] smtplib

[1117] Data processing and calculation

[1118] Audio data collection:

[1119] The device (smartphone) uses a microphone to collect the sounds your pet makes, and the audio data is recorded in digital format.

[1120] Analysis of audio data:

[1121] The server converts the collected audio data into MFCC (Mel-Frequency Cepstrum Coefficients) using librosa, which allows the extraction of audio data features.

[1122] Analysis by AI model:

[1123] The server inputs the audio data into a generative AI model pre-trained with TensorFlow to analyze the pet's health and emotional state, which is then used to detect abnormalities in the pet.

[1124] Send notifications:

[1125] If an abnormality is detected, the server uses smtplib to send an email notification to the owner, containing detailed information about the pet's abnormal condition.

[1126] Specific examples

[1127] For example, you can record your pet's cries and input the audio data into the application. The AI ​​model analyzes the audio data, and if it detects an abnormality, it will notify the owner by email. This notification will include a message such as, "An abnormality has been detected with your pet. Please check immediately."

[1128] Prompt Sentence Examples

[1129] Examples of prompts to be input to a generative AI model include:

[1130] Write a Python program that analyzes pet noise data and notifies owners if an abnormality is detected. Use TensorFlow and librosa to analyze the audio data, and smtplib to send email notifications.

[1131] In this way, a system can be provided that can grasp the health and emotional state of a pet in real time and respond quickly if there is an abnormality.

[1132] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[1133] Step 1:

[1134] Audio data collection

[1135] The device (smartphone) collects pet sounds using a microphone. The input is the pet's sound, and the output is digital audio data. This audio data is used in subsequent analysis steps.

[1136] Step 2:

[1137] Audio data preprocessing

[1138] The server converts the collected audio data into MFCC (Mel-Frequency Cepstrum Coefficients) using librosa. The input is digital audio data, and the output is MFCC features. This conversion extracts the features of the audio data.

[1139] Step 3:

[1140] Analysis using AI models

[1141] The server inputs MFCC features into a generative AI model pre-trained using TensorFlow to analyze the pet's health and emotional state. The input is the MFCC features, and the output is a prediction result about the pet's condition. This analysis evaluates the pet's abnormalities and emotional state.

[1142] Step 4:

[1143] Anomaly detection

[1144] The server determines whether an abnormality has been detected in the pet based on the prediction results of the AI ​​model. The input is the prediction result of the AI ​​model, and the output is information about whether an abnormality has been detected. If an abnormality is detected, the server proceeds to the next step.

[1145] Step 5:

[1146] Sending notifications

[1147] The server uses smtplib to send notifications to the pet owner via email. The input is the pet owner's email address and information about the presence or absence of an abnormality. The output is the notification email sent. This notification contains detailed information about the pet's abnormal condition.

[1148] Step 6:

[1149] Real-time status monitoring

[1150] The server monitors the pet's health and emotional state in real time and provides the information to the owner as needed. The input is continuously collected voice data, and the output is real-time status information, allowing the owner to always be aware of the pet's condition.

[1151] In this way, it is possible to quickly and accurately grasp the health and emotional state of a pet, and to immediately notify the owner if any abnormalities occur.

[1152] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1153] "Example 1"

[1154] The system described in claim 4 includes an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's emotions from the user's tone of voice, facial expressions, behavior, etc. For example, when the user is angry at a pet, the emotion engine recognizes the anger and adjusts the instructions to the pet. When the user is happy, the emotion engine recognizes the joy and adjusts the instructions to the pet. This allows for smoother communication between the pet and the user.

[1155] "Example 2"

[1156] The system described in claim 5 includes a means for utilizing AI to realize conversation between a person and a pet, and an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's emotions from the user's tone of voice, facial expressions, behavior, etc., and feeds the results back to the AI. The AI ​​adjusts its instructions to the pet based on this feedback. For example, when the user is angry at the pet, the AI ​​recognizes the anger and adjusts its instructions to the pet. Also, when the user is happy, the AI ​​recognizes the joy and adjusts its instructions to the pet. This allows for smoother communication between the pet and the user.

[1157] "Example 3"

[1158] The system described in claim 6 includes a means for grasping the pet's state and an emotion engine for recognizing the user's emotions. The emotion engine analyzes the user's emotions from the user's tone of voice, facial expressions, behavior, etc., and feeds the results back to the means for grasping the pet's state. The means for grasping the pet's state adjusts the pet's state based on this feedback. For example, when the user is angry with the pet, the means for grasping the pet's state recognizes the anger and adjusts the pet's behavior. Also, when the user is happy, the means for grasping the pet's state recognizes the joy and adjusts the pet's behavior. This allows for smoother communication between the pet and the user.

[1159] The processing flow of each embodiment will be described below.

[1160] "Example 1"

[1161] Step 1: The emotion engine analyzes the user's emotions based on their tone of voice, facial expressions, behavior, etc.

[1162] Step 2: Based on the analysis results of the emotion engine, adjust the instructions given to the pet.

[1163] Step 3: The system operates to facilitate smooth communication between the pet and the user.

[1164] "Example 2"

[1165] Step 1: The emotion engine analyzes the user's emotions based on their tone of voice, facial expressions, behavior, etc.

[1166] Step 2: The emotion engine feeds back the results of its analysis to the AI.

[1167] Step 3: The AI ​​uses this feedback to adjust its instructions to the pet.

[1168] Step 4: The system operates to facilitate smooth communication between the pet and the user.

[1169] "Example 3"

[1170] Step 1: The emotion engine analyzes the user's emotions based on their tone of voice, facial expressions, behavior, etc.

[1171] Step 2: The results of the emotion engine's analysis are fed back into a means of understanding the pet's condition.

[1172] Step 3: The means to understand the pet's condition will use this feedback to adjust the pet's condition.

[1173] Step 4: The system operates to facilitate smooth communication between the pet and the user.

[1174] Example 1

[1175] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1176] Conventional pet management systems simply record the sounds and behavior of pets, making it difficult to respond appropriately based on the pet's condition or the user's emotions. Furthermore, they lacked a means for smooth communication between the user and the pet. This made it difficult to properly understand the pet's health and emotional state and give instructions to the pet based on the user's emotions.

[1177] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1178] In this invention, the server includes means for collecting voice data, means for collecting characteristics of pet animals, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the state of the pet from its cries based on the collected data, means for recognizing the user's emotions, and means for adjusting instructions to the pet based on the user's emotions. This makes it possible to understand the state of the pet and the user's emotions in real time and to give appropriate instructions to the pet.

[1179] "Audio data" refers to data in which sounds, such as animal cries or the user's voice, are recorded in digital format.

[1180] "Characteristics of pets" refers to specific information about pets, such as their personalities, sounds, and behavioral patterns.

[1181] "Type of animal" is information indicating the type of pet, such as dog, cat, bird, etc.

[1182] "General animal characteristic data" refers to data such as personality and behavior patterns common to a particular animal species.

[1183] "Pet status" is information that indicates the current state of the pet, such as the pet's health and emotional state.

[1184] "User emotion" refers to the emotional state of the user as analyzed from the user's tone of voice, facial expression, behavior, etc.

[1185] "Adjusting instructions" means changing the instructions given to the pet based on the user's emotions and the state of the pet.

[1186] "Collecting means" refers to devices and methods for acquiring voice data, animal characteristic data, etc.

[1187] "Means for identification" refers to devices or methods for identifying the type and characteristics of animals based on collected data.

[1188] The "analysis means" refers to a device or method for analyzing the collected data and understanding the pet's condition and the user's emotions.

[1189] This invention is a system that collects and analyzes data on the cries and behavior of pets to understand their state and adjust instructions to the pet based on the user's emotions. Specific embodiments of this system will be described below.

[1190] Audio data collection

[1191] Users use the microphone on their smartphone to collect their pet's cries. The collected audio data is sent from the device to the server in real time. Specifically, the user starts the smartphone app and presses the "Start Recording" button to begin collecting audio data.

[1192] Analysis of audio data

[1193] The server uses voice analysis software to analyze the audio data. Specifically, it converts the collected audio data into text data using a common voice analysis API. For example, it can use the Google Cloud Speech-to-Text API. Based on the analysis results, it identifies whether the audio is the sound of a specific animal.

[1194] Collection of animal characteristic data

[1195] Users input their pet's personality and vocalizations into a smartphone app. The device also uses sensors to observe the pet's behavior. Specifically, it uses behavioral observation sensors such as Nest Cam to collect data on the pet's behavior. The collected data is then sent to a server.

[1196] Identifying Animal Species

[1197] The user inputs the pet's type into the app, and the device uses image recognition technology to identify the pet's type. Specifically, a general image recognition API can be used. For example, Google Cloud Vision API can be used to analyze the pet's image and identify the animal type.

[1198] Analyzing data and understanding your pet's condition

[1199] The server uses AI models to analyze the collected data, specifically machine learning frameworks such as TensorFlow, to predict the pet's condition and understand its health and emotional state.

[1200] User Emotion Recognition

[1201] The device uses emotion analysis software to recognize the user's emotions. Specifically, it uses the Emotion API to analyze the user's tone of voice and facial expressions. Based on the analysis results, the server recognizes whether the user is angry or happy with their pet.

[1202] Adjusting pet instructions

[1203] The server adjusts the instructions to the pet based on the user's emotions: if the user is angry, it generates a gentle instruction, and if the user is happy, it generates a reinforced instruction. The generated instructions are transmitted to the pet via the terminal.

[1204] Specific examples

[1205] Audio data collection: The user launches the app on their smartphone and presses the "Start Recording" button to record the sound of their dog barking.

[1206] Voice data analysis: The server uses the Google Cloud Speech-to-Text API to analyze the recorded voice data and obtain the text data "Woof woof."

[1207] Animal characteristic data collection: The user enters "The dog's personality is gentle" into the app's input form, and the device uses Nest Cam to collect the dog's behavior data.

[1208] Animal species identification: A user uploads a photo of a dog to the app, and the device identifies it as a "Golden Retriever" using the Google Cloud Vision API.

[1209] Data analysis and pet condition assessment: The server uses TensorFlow to analyze the collected data and determine whether the dog is experiencing stress.

[1210] User emotion recognition: The device uses the Emotion API to analyze the user's tone of voice and facial expressions and recognize that the user is angry.

[1211] Pet command adjustment: The server analyzes the user's emotional data and generates gentle commands such as "sit," which the device then relays to the dog via voice.

[1212] Prompt Sentence Examples

[1213] "Collect dog barks and analyze them with a common audio analysis API. Based on the analysis results, use a machine learning framework to understand the dog's state, and use emotion analysis software to recognize the user's emotions and adjust instructions for the pet."

[1214] In this way, the system specifically executes each processing step to facilitate smooth communication between the pet and the user.

[1215] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1216] Step 1:

[1217] Audio data collection

[1218] The user collects the sounds of their pet using the microphone on their smartphone. Specifically, the user starts the smartphone app and presses the "Start Recording" button to begin collecting audio data.

[1219] Input: Pet noises

[1220] Output: Recorded audio data

[1221] Step 2:

[1222] Sending audio data

[1223] The device transmits the collected voice data to a server in real time. Specifically, the smartphone app uploads the recorded voice data to the server via the Internet.

[1224] Input: Recorded audio data

[1225] Output: Audio data sent to the server

[1226] Step 3:

[1227] Analysis of audio data

[1228] The server uses speech analysis software to analyze the voice data. Specifically, it converts the collected voice data into text data using a common speech analysis API. For example, Google Cloud Speech-to-Text API can be used.

[1229] Input: Audio data sent to the server

[1230] Output: Parsed text data

[1231] Step 4:

[1232] Collection of animal characteristic data

[1233] Users input their pet's personality and vocalizations into a smartphone app. The device also uses sensors to observe the pet's behavior. Specifically, it uses behavioral observation sensors such as Nest Cam to collect data on the pet's behavior. The collected data is then sent to a server.

[1234] Input: Pet's personality, vocalizations, and behavioral data

[1235] Output: Feature and behavioral data sent to the server

[1236] Step 5:

[1237] Identifying Animal Species

[1238] The user inputs the pet's type into the app, and the device uses image recognition technology to identify the pet's type. Specifically, a general image recognition API can be used. For example, Google Cloud Vision API can be used to analyze the pet's image and identify the animal type.

[1239] Input: Pet image

[1240] Output: Analyzed animal species

[1241] Step 6:

[1242] Analyzing data and understanding your pet's condition

[1243] The server uses AI models to analyze the collected data, specifically machine learning frameworks such as TensorFlow, to predict the pet's condition and understand its health and emotional state.

[1244] Input: feature data, behavior data, animal type

[1245] Output: Parsed pet status

[1246] Step 7:

[1247] User Emotion Recognition

[1248] The device uses emotion analysis software to recognize the user's emotions. Specifically, it uses the Emotion API to analyze the user's tone of voice and facial expressions. Based on the analysis results, the server recognizes whether the user is angry or happy with their pet.

[1249] Input: User's tone of voice and facial expressions

[1250] Output: Parsed user sentiment

[1251] Step 8:

[1252] Adjusting pet instructions

[1253] The server adjusts the instructions to the pet based on the user's emotions: if the user is angry, it generates a gentle instruction, and if the user is happy, it generates a reinforced instruction. The generated instructions are transmitted to the pet via the terminal.

[1254] Input: Analyzed pet state, user emotion

[1255] Output: Adjusted instructions

[1256] In this way, the system specifically executes each processing step to facilitate smooth communication between the pet and the user.

[1257] (Application example 1)

[1258] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1259] Conventional pet care systems only analyze pets' sounds and behaviors, and are unable to provide advice or services that take the user's emotions into account. This makes it difficult to communicate smoothly between pets and users, making it difficult to properly manage pets' health and emotional state. Furthermore, in brick-and-mortar stores, it is difficult to provide appropriate services tailored to the pet's condition.

[1260] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for understanding the pet's condition from its cries based on the collected data, means for recognizing the user's emotions, means for generating advice based on the pet's condition and the user's emotions, and means for providing the generated advice to the user. This makes it possible to comprehensively analyze the pet's condition and the user's emotions and provide appropriate advice and services.

[1261] "Audio data" refers to audio information such as animal cries.

[1262] "Characteristics of pet animals" refers to information about the characteristics of the animals kept by the owner, such as their personalities and sounds.

[1263] "Means for identifying the type of animal" refers to a method for identifying the type of animal using image recognition technology or input from the owner.

[1264] "General animal characteristic data" refers to characteristic information about general animals obtained from a database or the like.

[1265] "Means for understanding a pet's condition" refers to a method for analyzing a pet's health and emotional state based on collected voice data and animal characteristic data.

[1266] "Means for recognizing user emotions" refers to a method for identifying a user's emotions by analyzing the user's tone of voice, facial expressions, behavior, etc.

[1267] The "means for generating advice" refers to a method for generating appropriate advice or services based on the pet's condition and the user's emotions.

[1268] The "means for providing the generated advice to the user" refers to a method for communicating the generated advice or service to the user.

[1269] A system for implementing this invention includes means for collecting audio data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying the type of animal, means for collecting general animal characteristic data, means for understanding the condition of the pet from its cries based on the collected data, means for recognizing the user's emotions, means for generating advice based on the pet's condition and the user's emotions, and means for providing the generated advice to the user.

[1270] Hardware and software used

[1271] Hardware:

[1272] Microphone: Used to collect audio data.

[1273] Camera: Used to collect video data of the user.

[1274] software:

[1275] speech_recognition library: Used to perform speech recognition.

[1276] cv2 (OpenCV): Used to collect and analyze video data.

[1277] emotion_recognition: An emotion recognition engine to recognize user emotions.

[1278] pet_behavior_analysis: A pet behavior analysis engine for analyzing pet sounds and behavior.

[1279] advice_generator: An advice generation engine for generating advice based on the pet's state and the user's emotions.

[1280] System processing overview

[1281] 1. Audio data collection:

[1282] The server collects pet sounds using a microphone, and the collected audio data is analyzed using the speech_recognition library.

[1283] 2. Video data collection:

[1284] The server collects video data of the user using a camera, and the collected video data is analyzed using cv2 (OpenCV).

[1285] 3. Pet status analysis:

[1286] The server analyzes the collected voice data using the pet_behavior_analysis engine to understand the pet's condition (e.g., hunger, stress, joy, etc.).

[1287] 4. User Emotion Recognition:

[1288] The server analyzes the collected video data using the emotion_recognition engine to recognize the user's emotions (e.g., anger, joy, sadness, etc.).

[1289] 5. Generating Advice:

[1290] The server generates appropriate advice using the advice_generator engine based on the pet's condition and the user's emotions.

[1291] 6. Providing advice:

[1292] The server provides the generated advice to the user, for example, if the pet is stressed, it suggests items to help the pet relax.

[1293] Specific examples

[1294] If your pet is stressed:

[1295] The server generates advice such as "Your pet is stressed. Try using a toy that will help it relax." and provides it to the user.

[1296] If the user is angry:

[1297] The server generates advice such as "The user is angry. Please be kind to your pet." and provides it to the user.

[1298] Prompt Sentence Examples

[1299] Analyze the pet's meow data and the user's video data, and generate appropriate advice based on the pet's condition and the user's emotions. If the pet is stressed, suggest items that will help the pet relax, and if the user is angry, advise the user to be gentle with the pet.

[1300] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1301] Step 1:

[1302] The server collects pet sounds using a microphone. The collected audio data is analyzed using the speech_recognition library. The input is the pet's sound, and the output is the analyzed audio data. Through this analysis, the characteristics of the pet's sound are extracted.

[1303] Step 2:

[1304] The server collects the user's video data using a camera. The collected video data is analyzed using cv2 (OpenCV). The input is the user's video data, and the output is the analyzed video data. This analysis extracts the user's facial expressions and behavioral characteristics.

[1305] Step 3:

[1306] The server analyzes the collected voice data using the pet_behavior_analysis engine to understand the pet's state. The input is the analyzed voice data, and the output is the pet's state (e.g., hunger, stress, joy, etc.). This analysis identifies the pet's health and emotional state.

[1307] Step 4:

[1308] The server analyzes the collected video data using the emotion_recognition engine to recognize the user's emotions. The input is the analyzed video data, and the output is the user's emotion (e.g., anger, joy, sadness, etc.). This analysis identifies the user's emotional state.

[1309] Step 5:

[1310] The server generates appropriate advice using the advice_generator engine based on the pet's state and the user's emotion. The input is the pet's state and the user's emotion, and the output is the generated advice. This generation creates specific advice that is tailored to the pet and user's situation.

[1311] Step 6:

[1312] The server provides the generated advice to the user. The input is the generated advice, and the output is the advice provided to the user. This allows the user to take appropriate action according to the state of their pet.

[1313] Example 2

[1314] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1315] Conventional methods of communicating with pets have had the problem that it is difficult to accurately understand the sounds and behavior of pets, and it takes time to grasp the pet's condition and requests. Also, when a user gives instructions to a pet, the pet often does not understand the instructions, which makes communication difficult.

[1316] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the status of the pet from its cries based on the collected data, means for analyzing the collected voice data, means for notifying the user of the analysis results, means for collecting the user's voice instructions, means for analyzing the voice instruction data, and means for communicating the analysis results to the pet. This allows the pet's cries and the user's voice instructions to be accurately analyzed, enabling smooth communication between the pet and the user.

[1317] "Audio data" is data in which sounds made by a pet or a user are recorded in digital format.

[1318] "Animal characteristics" refers to information that refers to the unique characteristics of each individual animal, such as the pet's personality or the sounds it makes.

[1319] "Animal type" refers to the biological classification to which a pet belongs, such as a dog or a cat.

[1320] "General animal characteristic data" is data that includes information such as behaviors and sounds common to a particular type of animal.

[1321] "Means for understanding the state of a pet from its cries" refers to a method or device for analyzing a pet's cries and identifying the state or needs of the pet indicated by the cries.

[1322] The "means for analyzing collected voice data" refers to a method or device for converting collected voice data into text data and analyzing the content of the text data.

[1323] The "means for notifying the user of the analysis results" refers to a method or device for conveying the analyzed information to the user.

[1324] The "means for collecting user's voice instructions" refers to a method or device for recording the voice instructions given by the user and collecting the data.

[1325] The "means for analyzing voice instruction data" refers to a method or device for converting a user's voice instruction into text data and analyzing the content of that text data.

[1326] The "means for communicating the analysis results to the pet" refers to a method or device for communicating the analyzed user's instructions in a form that the pet can understand.

[1327] This invention is a system for realizing smooth communication between a pet and a user. This system includes a series of processes for collecting and analyzing voice data and communicating the results to the user and the pet.

[1328] Hardware and software used

[1329] Subject: Server

[1330] The server uses a speech recognition API and a natural language processing (NLP) model to analyze the pet's cries and the user's voice commands. Specifically, it uses Google Cloud's Speech-to-Text API to convert the voice data into text data, and then analyzes the text data using the NLP model. As a result of the analysis, it identifies the pet's status, requests, and user commands, and conveys them to the user and the pet in an appropriate manner.

[1331] Subject: Terminal

[1332] The device is equipped with a microphone to collect the pet's cries and the user's voice commands. The collected voice data is sent to a server via the Internet. It also has a display and speaker to notify the user of the analysis results received from the server. For example, a smartphone or tablet can play this role.

[1333] Subject: User

[1334] The user inputs the sound of their pet barking into the terminal and also issues their own voice commands to the terminal. For example, if their pet barks "woof woof," the user inputs the sound into the terminal, and the terminal sends the voice data to the server. Similarly, if the user says "let's play," the voice command is sent to the server via the terminal.

[1335] Specific examples

[1336] Example 1: Pet cry analysis

[1337] 1. Your pet barks "woof woof."

[1338] 2. The device's microphone captures the bark.

[1339] 3. The device sends the bird call data to the server.

[1340] 4. The server converts the audio data into text using Google Cloud's Speech-to-Text API.

[1341] 5. The server uses an NLP model to analyze the text data and generate a message saying "Your pet is hungry."

[1342] 6. The server generates a message and sends it to the device.

[1343] 7. The device displays a notification to the user that "Your pet is hungry."

[1344] Example 2: User voice command analysis

[1345] 1. The user says "Let's play."

[1346] 2. The device's microphone captures the audio.

[1347] 3. The device sends the audio data to the server.

[1348] 4. The server converts the audio data into text using Google Cloud's Speech-to-Text API.

[1349] 5. The server uses an NLP model to analyze the text data and convert the command "Let's play" into a form that the pet can understand.

[1350] 6. The server sends the converted instructions to the terminal.

[1351] 7. The device will play a specific sound to tell your pet to "play."

[1352] Prompt Sentence Examples

[1353] "If your pet barks 'woof woof,' we can analyze the sound and tell you what condition your pet is in."

[1354] "When the user says 'Let's play,' analyze that speech and generate commands for the pet."

[1355] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1356] Step 1: Collect your pet sounds

[1357] Subject: Terminal

[1358] The device collects the pet's barks through a microphone. For example, if the pet barks "woof woof," the device's microphone captures the sound. The input is the pet's bark, and the output is digital audio data.

[1359] Step 2: Send the bird sound data to the server

[1360] Subject: Terminal

[1361] The terminal sends the collected sound data to a server via the Internet. The data is sent in audio file format (e.g., WAV format). The input is digital audio data, and the output is the audio data sent to the server.

[1362] Step 3: Analyze the bird call data on the server

[1363] Subject: Server

[1364] The server converts the received bark data into text data using a speech recognition API. Specifically, it uses Google Cloud's Speech-to-Text API. The input is audio data and the output is text data. The NLP model then analyzes the text data to identify the pet's status and requests. For example, it analyzes that a bark of "woof woof" means "I'm hungry." The input is text data and the output is the analysis result.

[1365] Step 4: Notify the user of the analysis results

[1366] Subject: Server

[1367] The server sends the analysis results to the user's device. For example, it generates a message saying "your pet is hungry" and notifies the user's smartphone. The input is the analysis results, and the output is the notification message sent to the user's device.

[1368] Step 5: Collect user voice commands

[1369] Subject: Terminal

[1370] The device collects the user's voice commands through a microphone. For example, when the user says "Let's play," the device captures the voice. The input is the user's voice command, and the output is digital voice data.

[1371] Step 6: Send voice instructions to the server

[1372] Subject: Terminal

[1373] The terminal transmits the collected voice instruction data to a server via the Internet. The data is transmitted in audio file format (e.g., WAV format). The input is digital audio data, and the output is the audio data transmitted to the server.

[1374] Step 7: Parse the voice command data on the server

[1375] Subject: Server

[1376] The server converts the received voice instruction data into text data using a speech recognition API. Specifically, it uses Google Cloud's Speech-to-Text API. The input is voice data and the output is text data. The text data is then analyzed using an NLP model and converted into a form that the pet can understand. For example, the instruction "Let's play" is converted into sounds and action instructions that the pet can understand. The input is text data and the output is the analysis results.

[1377] Step 8: Communicate the results to your pet

[1378] Subject: Terminal

[1379] Based on the analysis results received from the server, the device generates sounds and behavioral instructions that the pet can understand and communicates them to the pet. For example, the device can play a specific sound to tell the pet to "play." The input is the analysis results, and the output is the instruction to be communicated to the pet.

[1380] (Application example 2)

[1381] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1382] Conventional methods of communication between pets and their owners rely on owners' intuitive understanding of their pets' cries and behavior, making accurate communication difficult. It is also difficult to accurately grasp a pet's health and emotional state, which can delay appropriate responses. Furthermore, in brick-and-mortar stores such as pet shops, staff are required to quickly understand the pet's condition and take appropriate action, but currently there is a lack of effective methods for doing so.

[1383] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1384] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from its cries based on the collected data, means including an emotion engine for recognizing the user's emotions, means for adjusting instructions to the pet based on the user's emotions, means for analyzing the pet's cries and determining the pet's condition and requests, means for converting human speech into a form the pet can understand, and means including an application to be installed on a smartphone. This enables accurate communication between pets and their owners, allowing the pet's health and emotional state to be quickly understood and appropriate measures to be taken. Furthermore, in physical stores, store staff can quickly determine the pet's condition and take appropriate measures.

[1385] "Audio data" refers to digitally recorded sound information, including animal sounds.

[1386] "Animal characteristics" refers to the individual characteristics of pet animals, such as their personalities and sounds.

[1387] "Means for identifying animal species" refers to methods or techniques for distinguishing between specific animal species.

[1388] "General animal characteristic data" refers to information about widely known animal characteristics.

[1389] "Means for understanding the condition of pets from their cries" refers to methods and techniques for analyzing pet cries to understand their health and emotional state.

[1390] An "emotion engine that recognizes the user's emotions" refers to technology that analyzes emotions from the user's tone of voice, facial expressions, behavior, etc.

[1391] The term "means for adjusting instructions to a pet based on the user's emotions" refers to a method or technology for changing instructions to a pet depending on the user's emotional state.

[1392] "Means for analyzing pet cries and understanding the pet's condition and requests" refers to methods and techniques for analyzing pet cries and understanding the pet's condition and requests.

[1393] "Means for converting human language into a form that pets can understand" refers to methods and techniques for converting human language into sounds and actions that pets can easily understand.

[1394] "Applications installed on a smartphone" refers to software programs that run on a smartphone.

[1395] A system for implementing this invention includes means for collecting audio data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for understanding the state of a pet from its cries based on the collected data, means including an emotion engine for recognizing the user's emotions, means for adjusting instructions to the pet based on the user's emotions, means for analyzing the pet's cries and understanding the pet's state and requests, means for converting human words into a form that the pet can understand, and means including an application to be installed on a smartphone.

[1396] The server collects voice data and has a database for identifying the characteristics and species of animals. This allows it to analyze the pet's cries and understand its status and requests. Furthermore, using an emotion engine that recognizes the user's emotions, it analyzes the user's tone of voice, facial expressions, and behavior, and adjusts instructions to the pet based on the results.

[1397] The device (smartphone) sends the collected voice data to the server and receives the analysis results. The analysis results are displayed to the user in a form that notifies them of the pet's status and requests. When the user gives instructions to the pet, the instructions are converted into a form that the pet can understand and communicated to the pet.

[1398] A specific example is a scenario in which a pet shop clerk uses a smartphone to analyze the cries of a pet and understand that the pet is saying, "I'm hungry." Another scenario is when a clerk says, "Let's play," and the command is communicated to the pet in a way that is easy for the pet to understand.

[1399] Examples of prompts to input to a generative AI model include:

[1400] "Analyze your pet's cries and let us know their condition."

[1401] "Translate human language into a form your pet can understand."

[1402] This system enables accurate communication between pets and their owners, allowing them to quickly understand the pet's health and emotional state and take appropriate action.It also enables store staff in physical stores to quickly understand the pet's condition and take appropriate action.

[1403] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1404] Step 1:

[1405] The server receives audio data (animal sounds) sent from the device (smartphone). The input is audio data, and the output is data ready for analysis. Specifically, the server stores the audio data in digital format and passes it to the analysis engine.

[1406] Step 2:

[1407] The server analyzes the audio data and identifies the characteristics and type of animal. The input is the audio data received in step 1, and the output is information about the animal's type and characteristics. Specifically, the server uses a voice recognition algorithm to convert the sound into text, and then searches a database for the animal's type and characteristics based on that text.

[1408] Step 3:

[1409] The server understands the pet's condition and requests based on the collected animal characteristic data. The input is the animal characteristic data obtained in step 2, and the output is information about the pet's condition and requests. Specifically, the server runs an algorithm to estimate the pet's health and emotional state based on the analysis results.

[1410] Step 4:

[1411] The server uses an emotion engine that recognizes the user's emotions to analyze the user's emotions from their tone of voice, facial expressions, and behavior. The input is the user's voice data and video data, and the output is information about the user's emotional state. Specifically, the server runs an emotion recognition algorithm to analyze the user's emotions.

[1412] Step 5:

[1413] The server adjusts the instructions to the pet based on the user's emotional state. The input is the information about the user's emotional state obtained in step 4, and the output is the adjusted instructions. Specifically, the server executes an algorithm that changes the instructions depending on the user's emotions.

[1414] Step 6:

[1415] The server analyzes the pet's cries and understands the pet's status and requests. The input is the information about the pet's status and requests obtained in step 3, and the output is a message to notify the user. Specifically, the server generates a message to notify the user based on the analysis results.

[1416] Step 7:

[1417] The server converts human speech into a form that the pet can understand. The input is the user's words, and the output is instructions converted into sounds and actions that the pet can understand. Specifically, the server uses a generative AI model to translate the user's words into a form that the pet can understand.

[1418] Step 8:

[1419] The device (smartphone) displays the analysis results and instructions received from the server to the user. The input is the message sent from the server, and the output is the information displayed to the user. Specifically, the device displays the received message on the screen and notifies the user.

[1420] Step 9:

[1421] The user checks the pet's status and requests through the terminal and takes necessary action. The input is the information displayed on the terminal, and the output is the user's actions. In concrete terms, the user takes appropriate action toward the pet based on the information displayed on the terminal.

[1422] Example 3

[1423] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1424] Conventional systems for understanding a pet's health and emotional state mainly analyze only the pet's cries and behavior, and have the problem of being unable to adjust the pet's behavior to take the user's emotions into account. Furthermore, there are insufficient means for facilitating communication between the pet and the user. This has resulted in insufficient health management and emotional understanding of the pet, leading to problems with smooth communication with the user.

[1425] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[1426] In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the state of the pet from its cries based on the collected data, means for collecting the user's emotions, means for analyzing the user's emotions, and means for adjusting the pet's behavior based on the results of the user's emotion analysis, thereby making it possible to more accurately understand the health and emotional state of the pet and adjust the pet's behavior in accordance with the user's emotions.

[1427] "Means for collecting audio data" refers to a function for collecting pet sounds using an audio input device such as a microphone.

[1428] "Means for collecting animal characteristics" refers to the function of collecting individual characteristics such as a pet's personality and cries using sensors, cameras, etc.

[1429] "Means for identifying animal species" refers to algorithms or identification technologies that identify the species of pet based on collected data.

[1430] "Means for collecting general animal characteristics data" refers to the ability to collect data on general animal characteristics from databases or external sources.

[1431] "Means of understanding the condition of your pet from its cries based on the collected data" refers to an algorithm that analyzes the collected audio data and characteristic data to determine the health and emotional state of your pet.

[1432] "Means for collecting user emotions" refers to the function of collecting the user's voice and facial expressions using a microphone or camera.

[1433] "Means for analyzing user emotions" refers to algorithms and analytical technologies that analyze collected user voice and facial expression data to determine the user's emotional state.

[1434] "Means for adjusting the behavior of a pet based on the results of analyzing the user's emotions" refers to a control algorithm or instruction function for appropriately adjusting the behavior of a pet based on the results of analyzing the user's emotions.

[1435] A "generative AI model" refers to an artificial intelligence model that analyzes the pet's condition and the user's emotions based on collected data.

[1436] The present invention is a system for understanding the health and emotional state of a pet and adjusting the behavior of the pet in accordance with the emotions of a user. Specific embodiments of this system will be described below.

[1437] Hardware and software used

[1438] Hardware: microphone, camera, sensors (to detect pet movement)

[1439] Software: Generative AI models (e.g., general artificial intelligence models), sentiment engines (e.g., general sentiment analysis APIs)

[1440] System configuration

[1441] 1. The device is equipped with a microphone to capture the sounds of your pet and a camera to capture images of your pet's behavior. These devices collect your pet's sounds and behavior data in real time.

[1442] 2. The device temporarily stores the collected data and sends it to the server at regular intervals.

[1443] 3. The server inputs the received data into the generative AI model to analyze the pet's health and emotional state. The generative AI model determines the pet's condition based on the collected voice and behavioral data.

[1444] 4. The server sends the analysis results to the device, which then notifies the user, for example by displaying a message on a smartphone app saying, "Your pet may be unwell."

[1445] 5. The device is equipped with a microphone to collect the user's voice and a camera to capture the user's facial expressions. These devices collect the user's voice and facial expression data.

[1446] 6. The device sends the collected user data to the server, and the server uses an emotion engine to analyze the user's emotions.

[1447] 7. The server sends the results of the user's emotion analysis to the device, which then adjusts the pet's behavior. For example, if the user is angry, the device will instruct the pet to "calm down."

[1448] Specific examples

[1449] Pet Health Analysis:

[1450] For example, if a pet is making an unusual sound, a generative AI model can detect the change and notify the owner that their pet may be unwell.

[1451] User emotion recognition and pet behavior regulation:

[1452] Example: If a user is angry with their pet, the emotion engine will recognize that anger and the pet state awareness will adjust the pet's behavior to calm it down.

[1453] Prompt Sentence Examples

[1454] Pet Health Analysis:

[1455] "Analyze your pet's sounds and behavior to determine their health status."

[1456] User emotion recognition and pet behavior regulation:

[1457] "Analyze the user's voice and facial expressions to recognize their emotions. Adjust your pet's behavior based on the results."

[1458] This system makes it possible to grasp the status of the pet and the user in real time and take appropriate measures. The flow of the identification process in the third embodiment will be described with reference to FIG.

[1459] Step 1:

[1460] The device collects the sounds your pet makes.

[1461] Input: Pet noises

[1462] How it works: A microphone on the device collects your pet's sounds in real time.

[1463] Output: Collected audio data

[1464] Step 2:

[1465] The device captures your pet's behavior.

[1466] Input: pet behavior

[1467] How it works: The camera on the device captures your pet's movements in real time.

[1468] Output: Collected video data

[1469] Step 3:

[1470] The data collected by the device is temporarily stored and sent to the server at regular intervals.

[1471] Input: Audio data, video data

[1472] Operation: The data collected by the device is temporarily stored in the internal memory and sent to the server in batch processing at regular intervals.

[1473] Output: Data sent to the server

[1474] Step 4:

[1475] The server inputs the received data into a generative AI model to analyze the pet's health and emotional state.

[1476] Input: Audio data, video data

[1477] How it works: The server inputs data into a generative AI model that analyzes your pet's health and emotional state.

[1478] Output: Analysis results (pet's health and emotional state)

[1479] Step 5:

[1480] The server sends the analysis results to the terminal, which then notifies the user.

[1481] Input: Analysis results

[1482] How it works: The server sends the analysis results to the device, which then sends a push notification to the user's smartphone.

[1483] Output: Analysis results reported to the user

[1484] Step 6:

[1485] The device collects user feedback.

[1486] Input: User's voice

[1487] How it works: A microphone on the device collects the user's voice.

[1488] Output: Collected audio data

[1489] Step 7:

[1490] The device captures the user's facial expression.

[1491] Input: User's facial expression

[1492] How it works: The device's camera captures the user's facial expression.

[1493] Output: Collected video data

[1494] Step 8:

[1495] The terminal transmits the collected user data to the server.

[1496] Input: Audio data, video data

[1497] Operation: The device sends the collected data to the server.

[1498] Output: Data sent to the server

[1499] Step 9:

[1500] The server analyzes the user's emotions using an emotion engine.

[1501] Input: Audio data, video data

[1502] How it works: The server inputs data into the emotion engine to analyze the user's emotional state.

[1503] Output: Analysis result (user's emotional state)

[1504] Step 10:

[1505] The server sends the results of the user's emotion analysis to the terminal, which then adjusts the pet's behavior.

[1506] Input: Analysis results

[1507] Operation: The server sends the analysis results to the device, and the device gives the appropriate instructions to the pet.

[1508] Output: Adjusted pet behavior

[1509] In this way, the system can grasp the status of the pet and the user in real time and take appropriate action.

[1510] (Application example 3)

[1511] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1512] Conventional pet care systems have difficulty accurately understanding the health and emotional state of pets, making it difficult for owners and store staff to properly manage their pets. Furthermore, they lack the functionality to adjust the pet's behavior based on the user's emotions, which hinders smooth communication between the pet and the user.

[1513] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for collecting audio data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for understanding the pet's condition from its cries based on the collected data, means for collecting and analyzing image data, means for recognizing the user's emotions, and means for adjusting the pet's behavior based on the user's emotions. This makes it possible to accurately understand the pet's health condition and emotional state and adjust the pet's behavior in accordance with the user's emotions.

[1514] "Audio data" refers to audio information such as animal cries, and is data collected to analyze the condition of pets.

[1515] "Animal characteristics" refers to the unique characteristics of each animal, such as personality and vocalizations, and is information collected to understand the condition of your pet.

[1516] "Means for identifying animal species" refers to technology that identifies the type of animal a pet is based on collected data.

[1517] "General animal characteristic data" is data on widely known animal characteristics, and is information used as a basis for analyzing the condition of pets.

[1518] "Means for understanding the condition of pets from their cries" is a technology that analyzes collected cry data to determine the health and emotional state of pets.

[1519] The "means for collecting and analyzing image data" refers to a technology for collecting images of a pet and analyzing the images to understand the pet's condition.

[1520] "Means for recognizing user emotions" refers to technology that analyzes the user's tone of voice, facial expressions, etc. to determine the user's emotional state.

[1521] The "means for adjusting the behavior of a pet based on the user's emotions" is a technique for appropriately changing the behavior of a pet in accordance with the recognized emotions of the user.

[1522] As an embodiment of the present invention, a pet care assistant system will be described. This system is intended for use in physical stores such as pet shops and pet cafes. The system is installed on a smartphone, smart glasses, or a head-mounted display, and analyzes the health and emotional state of pets in real time, notifying store staff and owners.

[1523] Hardware and software used

[1524] Hardware

[1525] Smartphone

[1526] Smart Glasses

[1527] head-mounted display

[1528] software

[1529] OpenCV: Image processing library

[1530] Keras: A deep learning library

[1531] librosa: Audio processing library

[1532] soundfile: Sound file reading library

[1533] Data processing and calculation

[1534] Analysis of audio data

[1535] The server collects audio data (animal sounds) and extracts MFCC features using librosa. This converts the audio data into numerical data and feeds it into a model trained using Keras. The model predicts the pet's health and emotional state.

[1536] Image data analysis

[1537] The server collects image data, resizes and preprocesses it using OpenCV, and then feeds the preprocessed image data into a model trained using Keras to analyze the pet's condition.

[1538] User Emotion Recognition

[1539] The server collects the user's video data and extracts frames using OpenCV, which are then fed into an emotion recognition model trained with Keras to predict the user's emotional state.

[1540] Pet behavior adjustment

[1541] The server adjusts the pet's behavior based on the user's emotional state, for example, adjusting the pet's behavior to be calmer if the user is angry, and increasing the pet's behavior if the user is happy.

[1542] Specific examples

[1543] Usage example at a pet shop

[1544] If a pet in a pet shop is feeling stressed, the system will analyze audio and image data to detect the pet's stress state, alerting store staff and urging them to take better care of the pet.

[1545] Example of use at a pet cafe

[1546] At the pet cafe, the system analyzes the user's emotional state while playing with their pet. If the user is happy, the system will activate the pet's behavior and promote communication with the user.

[1547] Prompt Sentence Examples

[1548] Develop a system that analyzes your pet's sounds and behavior to notify you of its health and emotional state. Include the ability to recognize your emotions and adjust your pet's behavior accordingly.

[1549] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[1550] Step 1:

[1551] The server collects pet voice data. Specifically, it records pet cries using microphones installed in smartphones, smart glasses, and head-mounted displays. The input is voice data, and the output is a recorded voice file.

[1552] Step 2:

[1553] The server analyzes the collected audio data. It uses librosa to extract MFCC features from the audio data and convert them into numerical data. The input is the recorded audio file, and the output is numerical data including MFCC features.

[1554] Step 3:

[1555] The server inputs MFCC features into a voice analysis model trained using Keras to predict the pet's health and emotional state. The input is numerical data including MFCC features, and the output is a prediction result indicating the pet's health and emotional state.

[1556] Step 4:

[1557] The server collects image data of pets. Images of pets are taken using a camera mounted on a smartphone, smart glasses, or head-mounted display. The input is image data, and the output is the captured image file.

[1558] Step 5:

[1559] The server analyzes the collected image data. It resizes and preprocesses the images using OpenCV and inputs them into an image analysis model trained using Keras. The input is the captured image file, and the output is the preprocessed image data.

[1560] Step 6:

[1561] The server inputs the preprocessed image data into a model trained with Keras to analyze the pet's condition. The input is the preprocessed image data, and the output is a prediction result indicating the pet's condition.

[1562] Step 7:

[1563] The server collects the user's video data. It captures the user's facial expressions using a camera mounted on a smartphone, smart glasses, or head-mounted display. The input is video data, and the output is the captured video file.

[1564] Step 8:

[1565] The server analyzes the collected video data, extracts frames from the video using OpenCV, and inputs them into an emotion recognition model trained with Keras. The input is the captured video file, and the output is the extracted frame data.

[1566] Step 9:

[1567] The server inputs the extracted frame data into an emotion recognition model trained with Keras to predict the user's emotional state. The input is the extracted frame data, and the output is a prediction result indicating the user's emotional state.

[1568] Step 10:

[1569] The server adjusts the pet's behavior based on the user's emotional state. For example, if the user is angry, the server adjusts the pet's behavior to be calmer, and if the user is happy, the server adjusts the pet's behavior to be more active. The input is a prediction result indicating the user's emotional state, and the output is the adjusted pet's behavior.

[1570] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1571] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1572] Another example of generative AI is Gemini (internet search engine). <url: https: gemini.google.com ?hl="ja">) are listed.

[1573] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1574] [Third embodiment]

[1575] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1576] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1577] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1578] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1579] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1580] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1581] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1582] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1583] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1584] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1585] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1586] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[1587] "Example 1"

[1588] In one embodiment of the present invention, pet cries are collected using an audio data collection device such as a microphone. The collected audio data is sent to an audio analysis device and identified as the cry of a specific animal. In addition, the characteristics of the pet (personality, cries, etc.) are collected using input from the owner or sensors that observe the pet's behavior. The type of animal is identified using input from the owner or image recognition technology. General animal characteristic data is obtained from a database or the like. This data is analyzed by AI to understand the pet's condition.

[1589] "Example 2"

[1590] Another embodiment of the present invention provides a system that utilizes AI to enable conversations between people and pets. Specifically, AI analyzes a pet's cries and converts them into words that humans can understand. Conversely, AI analyzes human language and converts it into cries and actions that pets can understand. For example, if a pet expresses its hunger state through cries, the AI ​​analyzes it and conveys the message "your pet is hungry" to the owner. Also, if the owner says "let's play," the AI ​​analyzes it and conveys the message in a form that the pet can understand.

[1591] "Example 3"

[1592] As a further embodiment of the present invention, a system for understanding the health and emotional state of a pet is provided. Specifically, AI analyzes the health and emotional state of a pet from its sounds and behavior. For example, if a pet is suffering from an illness, its sounds and behavior may be different from normal. The AI ​​detects these changes and issues a warning to the owner. Also, if a pet is happy, its sounds and behavior may become more active. The AI ​​detects these changes and conveys the owner's joy.

[1593] The processing flow of each embodiment will be described below.

[1594] "Example 1"

[1595] Step 1: Collect your pet's cries using a sound data collection device such as a microphone.

[1596] Step 2: The collected audio data is sent to an audio analyzer to identify it as the sound of a specific animal.

[1597] Step 3: Collect information about pet characteristics (personality, vocalizations, etc.) using input from owners and sensors that observe pet behavior.

[1598] Step 4: Identify the animal's species using input from the owner and image recognition technology.

[1599] Step 5: Obtain general animal characteristic data from a database, etc.

[1600] Step 6: This data is analyzed using AI to understand the pet's condition.

[1601] "Example 2"

[1602] Step 1: AI analyzes your pet's cries and converts them into words that humans can understand. Step 2: AI analyzes human words and converts them into cries and actions that your pet can understand.

[1603] Step 3: For example, if a pet expresses its hunger state through a cry, the AI ​​will analyze it and convey the message to the owner that "your pet is hungry."

[1604] Step 4: Also, if the owner says "Let's play," the AI ​​will analyze it and convey the message in a way that the pet can understand.

[1605] "Example 3"

[1606] Step 1: AI analyzes your pet's health and emotional state based on its sounds and behavior.

[1607] Step 2: For example, if a pet is suffering from an illness, its sounds and behavior may be different from normal. AI can detect these changes and alert the owner.

[1608] Step 3: Also, if your pet is happy, its vocalizations and behavior may become more active. AI will detect these changes and communicate its happiness to its owner.

[1609] Example 1

[1610] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1611] Conventional pet condition monitoring systems mainly collect and analyze voice data and animal characteristic data separately, making it difficult to integrate this data to comprehensively understand the pet's condition.In addition, there was a lack of means to accurately understand the pet's health and emotional state, making it difficult for owners to appropriately understand and respond to their pet's condition.

[1612] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1613] In this invention, the server includes means for collecting voice data, means for transmitting the collected voice data, means for analyzing the voice data, means for collecting characteristics of pet animals, means for identifying the type of animal, means for collecting general animal characteristic data, and means for understanding the condition of pet animals from their cries based on the collected data. This makes it possible to comprehensively understand the condition of pets by integrating the voice data and animal characteristic data.

[1614] "Audio data" refers to data in which audio signals such as animal cries are recorded in digital format.

[1615] "Collecting means" refers to a device or method for acquiring audio data and animal characteristic data.

[1616] "Transmitting means" refers to a device or method for transferring collected data to another device or server.

[1617] An "analyzing means" is a device or method for processing collected data and extracting specific information.

[1618] "Animal characteristics" refers to individual attributes of pet animals, such as their personalities and sounds.

[1619] "Means for identification" refers to devices or methods for determining the type or characteristics of an animal.

[1620] "General animal characteristic data" refers to data such as characteristics and behavioral patterns common to a particular animal species.

[1621] "Pet condition" refers to the overall condition of the pet, including its health and emotional state.

[1622] "AI" is a technology that uses artificial intelligence to analyze data and make specific judgments and predictions.

[1623] "Means for realizing conversation" refers to devices or methods that allow people and pets to communicate.

[1624] "Health Status" refers to information about the physical health of your pet.

[1625] "Emotional state" refers to a pet's psychological state or mood.

[1626] The present invention is a system for collecting sounds of pets and analyzing the data to understand the state of the pets. A specific embodiment of this system will be described below.

[1627] First, a user uses the smartphone's built-in microphone or a dedicated external microphone to collect the sound of their pet. For example, the user starts a recording app on their smartphone and records the sound of their pet. This audio data is then sent to the server by the device. Specifically, the device sends the audio data to the server's API endpoint using an HTTP POST request.

[1628] The server analyzes the received audio data using a voice analyzer (for example, Google Cloud Speech-to-Text API). The server sends the audio data to the API, analyzes the returned text data, and identifies it as the sound of a specific animal.

[1629] Next, the user enters the characteristics of their pet (personality, sounds, etc.) into the app. For example, the user might enter information such as "Personality: Active, Sound: High-pitched" into the app's form. The app also uses sensors connected to the device (such as an accelerometer or camera) to observe the pet's behavior and collect data. The device periodically sends the data from the sensors to a server.

[1630] To identify the type of animal, the user can either input the type of pet into the app, or the device can use image recognition technology (such as TensorFlow or OpenCV) to identify the type of animal. For example, the user can input "dog," or the device can send an image taken with its camera to the server, which then uses image recognition technology to identify it as a "dog."

[1631] The server retrieves general animal trait data from a database (e.g., AWS RDS or Google BigQuery). The server runs SQL queries to retrieve the required data from the database.

[1632] Finally, the server analyzes the collected voice data, pet characteristic data, animal species data, and general animal characteristic data using AI (for example, TensorFlow or PyTorch). The server integrates this data and inputs it into an AI model to understand the pet's condition. For example, the server analyzes data such as "the dog is barking loudly and is behaving actively" and determines that "the dog is excited."

[1633] Examples and prompts

[1634] As a concrete example, the bark of a user's dog is collected using the microphone on their smartphone and sent to a server. The server analyzes the bark using the Google Cloud Speech-to-Text API and identifies it as a dog's bark. The user enters the dog's personality and characteristics of its bark into the app and observes the dog's behavior using a camera connected to the device. The server retrieves data on general dog characteristics from AWS RDS, analyzes the data using TensorFlow, and understands the dog's condition.

[1635] An example of a prompt sentence is as follows:

[1636] "Please analyze my dog's barks to understand his condition. My dog ​​has an active personality and barks a lot. Please analyze the following audio data."

[1637] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1638] Step 1:

[1639] A user collects the sounds of their pets.

[1640] Specifically, the user launches a recording app on their smartphone and records the sound of their pet's cries.

[1641] Input: Pet noises

[1642] Output: Audio data file (e.g. .wav format)

[1643] Step 2:

[1644] The device sends the collected voice data to the server.

[1645] Specifically, the device sends audio data to the server's API endpoint using an HTTP POST request.

[1646] Input: Audio data file

[1647] Output: Audio data sent to the server

[1648] Step 3:

[1649] The server analyzes the received audio data.

[1650] Specifically, the server uses a voice analysis device (e.g., Google Cloud Speech-to-Text API) to convert the voice data into text data and identify it as the sound of a specific animal.

[1651] Input: Audio data sent to the server

[1652] Output: Parsed text data (e.g. "dog barking")

[1653] Step 4:

[1654] Collect characteristics of the animals that the user keeps.

[1655] Specifically, the user enters information such as "personality: active, bark: high-pitched" into a form on the app. The app also uses sensors connected to the device (e.g., accelerometer and camera) to observe the pet's behavior and collect data.

[1656] Input: User-entered animal feature data, behavior data from sensors

[1657] Output: Collected animal trait data

[1658] Step 5:

[1659] Identify types of animals.

[1660] Specifically, the user enters the type of animal they own into the app, or the device uses image recognition technology (e.g., TensorFlow or OpenCV) to identify the type of animal.

[1661] Input: Animal species data or image data entered by the user

[1662] Output: Identified animal type data (e.g. "dog")

[1663] Step 6:

[1664] The server obtains general animal characteristic data.

[1665] Specifically, the server executes an SQL query from a database (e.g., AWS RDS or Google BigQuery) to retrieve the required data.

[1666] Input: Animal species data

[1667] Output: General animal traits data

[1668] Step 7:

[1669] The server analyzes the collected data and determines the pet's condition.

[1670] Specifically, the server analyzes collected voice data, animal characteristics data, animal species data, and general animal characteristics data using AI (e.g., TensorFlow or PyTorch). The server integrates this data and inputs it into an AI model to determine the pet's condition.

[1671] Input: Audio data, animal characteristics data, animal species data, general animal characteristics data

[1672] Output: Analysis result (e.g. "The dog is excited")

[1673] (Application example 1)

[1674] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1675] It is difficult for pet owners to monitor their pets' abnormal behavior and health status in real time while they are out or not watching over them. In addition, there is a lack of means to detect abnormalities in pets early and respond quickly, making it difficult to ensure the health and safety of pets.

[1676] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1677] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying the type of animal, means for collecting general animal characteristic data, means for understanding the pet's condition from its cries based on the collected data, and means for sending a notification when an abnormality is detected. This makes it possible to understand the pet's abnormal behavior and health condition in real time, and to quickly notify the owner when an abnormality is detected.

[1678] "Audio data" refers to information collected in digital format from audio signals such as animal cries.

[1679] "Collecting means" refers to a device or method for acquiring audio data and animal characteristic data.

[1680] "Animal characteristics" refers to information that refers to the animal's unique characteristics, such as its personality and cries.

[1681] "Means for identification" refers to devices or methods for determining the type or characteristics of an animal.

[1682] "General animal characteristic data" refers to information such as personality and behavioral patterns common to a specific animal species.

[1683] "Means of understanding" refers to devices and methods used to analyze collected data and understand the condition of animals.

[1684] "Means for detecting abnormalities" refers to devices or methods for identifying abnormalities in animal behavior or vocalizations and detecting the occurrence of an abnormality.

[1685] The "means for sending a notification" refers to a device or method for conveying information to the owner or other person when an abnormality is detected.

[1686] The system for carrying out this invention collects and analyzes the sounds of pets to understand their condition, and if an abnormality is detected, sends a notification to the pet owner. A specific embodiment of the system will be described below.

[1687] Hardware Configuration

[1688] The system includes the following hardware:

[1689] Microphone: Used to collect pet sounds.

[1690] Smartphone or smart glasses: Used to analyze collected voice data and detect anomalies.

[1691] Server: Analyzes voice data and sends notifications.

[1692] Software Configuration

[1693] The system includes the following software:

[1694] TensorFlow: Runs AI models to analyze audio data.

[1695] sounddevice: Used to collect audio data.

[1696] smtplib: Used to send notifications when an anomaly is detected.

[1697] Data processing and calculation

[1698] 1. Audio data collection: The device (smartphone or smart glasses) uses a microphone to collect the pet's cries. The collected audio data is stored digitally.

[1699] 2. Audio data analysis: The server inputs the collected audio data into the TensorFlow model for analysis. Based on the analysis results, the pet's condition is determined.

[1700] 3. Anomaly detection: The server determines whether an anomaly has been detected based on the analysis results. If an anomaly is detected, it sends a notification to the owner.

[1701] 4. Sending notifications: The server uses smtplib to send notifications to the pet owner, including information about the pet's abnormal behavior and health status.

[1702] Specific examples

[1703] For example, if a pet makes an unusual noise while the owner is out, the system will collect and analyze the sound. If an abnormality is detected based on the analysis results, a notification will be sent to the owner's smartphone, allowing the owner to take prompt action.

[1704] Prompt Sentence Examples

[1705] "Collect pet sounds and analyze them using an AI model. If an abnormality is detected, create a program that notifies the owner."

[1706] In this way, the present invention makes it possible to grasp abnormal behavior and health conditions of pets in real time, and to promptly notify the owner if an abnormality is detected.

[1707] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1708] Step 1:

[1709] The device (smartphone or smart glasses) uses a microphone to collect the pet's barks. The input is the pet's barks, and the output is digital audio data. Specifically, the device uses software for audio data collection (sounddevice library) to record audio for a certain period of time and save it as digital data.

[1710] Step 2:

[1711] The device sends the collected voice data to a server. The input is digital voice data, and the output is the voice data sent to the server. Specifically, the device uses an internet connection to upload the voice data to the server.

[1712] Step 3:

[1713] The server inputs the received audio data into the TensorFlow model for analysis. The input is digital audio data, and the output is the analysis result (the pet's state). Specifically, the server uses the TensorFlow library to input the audio data into the AI ​​model and analyze the characteristics of the pet's cries.

[1714] Step 4:

[1715] The server determines whether an anomaly has been detected based on the analysis results. The input is the analysis results, and the output is information about whether an anomaly has been detected. Specifically, the server compares the analysis results with a threshold and sets a flag if an anomaly is detected.

[1716] Step 5:

[1717] If an abnormality is detected, the server sends a notification to the pet owner. The input is information about whether an abnormality exists, and the output is a notification to the pet owner. Specifically, the server uses the smtplib library to send an abnormality notification to the pet owner's email address. The notification contains detailed information about the pet's abnormal behavior and health status.

[1718] Step 6:

[1719] The user (owner) checks the received notification and takes action to check the pet's status if necessary. The input is the notification from the server, and the output is the user's response. Specifically, the user checks the notification on their smartphone and takes action such as returning home to check the pet's status or contacting the pet sitter.

[1720] Example 2

[1721] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1722] Conventional methods of communication between pets and humans have been limited, making it difficult to accurately understand a pet's intentions and state from its cries and behavior. Furthermore, there has been a lack of ways to communicate human language to pets in a way that they can understand. This has made it difficult to properly manage a pet's health and emotional state, resulting in insufficient communication between pets and their owners.

[1723] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for understanding the status of the pet from its cries based on the collected data, means for converting the voice data into text data using voice recognition software, means for analyzing the text data using a generative AI model to understand the pet's intentions, means for generating a message based on the analysis result and notifying the user, means for collecting the user's words and converting it into text data using voice recognition software, means for analyzing the text data using a generative AI model to generate a message to be conveyed to the pet, and means for converting the message into cries and actions that the pet can understand using voice synthesis software and transmitting it. This makes it possible to accurately understand the intentions and status of the pet from its cries and actions, and to convey human language in a form that the pet can understand.

[1724] "Audio data" refers to data that has been recorded in digital format from sounds made by animals or humans.

[1725] "Animal characteristics" refers to characteristics unique to individual animals, such as their personalities, sounds, and behavioral patterns.

[1726] "Type of animal" indicates the classification of animals, such as dog, cat, bird, etc.

[1727] "General animal characteristic data" refers to data such as personality and behavioral patterns common to a particular animal species.

[1728] "Speech recognition software" is software that analyzes voice data and converts it into text data.

[1729] A "generative AI model" is a model that uses artificial intelligence to analyze text data and understand its meaning.

[1730] "Text data" is text information converted from voice data by voice recognition software.

[1731] "Message" refers to the content of notifications or instructions generated based on the analysis results.

[1732] "Speech synthesis software" is software for converting text data into speech.

[1733] "User" refers to a person who uses this system.

[1734] "Pet" refers to an animal kept in a household.

[1735] This invention is a system that uses AI to enable conversation between humans and pets. This system has the functions of analyzing pet cries and converting them into words that humans can understand, and analyzing human words and converting them into cries and actions that pets can understand.

[1736] Hardware and software used

[1737] Hardware

[1738] Microphone: Used to capture pet sounds and user speech.

[1739] Speaker: Used to transmit generated sounds to your pet.

[1740] Camera: Used to monitor pet behavior if necessary.

[1741] Server: Used to analyze and process data.

[1742] software

[1743] Speech recognition software: used to convert voice data into text data (e.g., Google Speech-to-Text).

[1744] Generative AI models: Analyze text data and use it to understand your pet's intentions (e.g., GPT-4).

[1745] Text-to-speech software: Used to convert text data into speech (e.g., Google Text-to-Speech).

[1746] Specific operation of the system

[1747] Pet cry analysis

[1748] The device's microphone captures the sound of your pet barking. For example, if your dog barks, "woof woof," the audio data is collected through the microphone and sent to the server in real time.

[1749] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice "woof woof" is converted into the text "woof woof." The converted text data is sent to the next analysis step.

[1750] The server uses a generative AI model to analyze the text data and understand the pet's intentions. For example, the text "woof woof" is interpreted as meaning "I'm hungry." The analysis results are sent to the message generation step.

[1751] Based on the analysis results, the server generates a message saying "your pet is hungry" and notifies the user. For example, a notification is sent to the user via a smartphone app. The user can check the notification on the app and understand the status of their pet.

[1752] Human language analysis

[1753] The microphone on the device captures the user's words. For example, when the user says "Let's play," the voice data is collected through the microphone. The collected voice data is sent to the server in real time.

[1754] The server analyzes the captured voice data using speech recognition software and converts it into text data. For example, the voice saying "Let's play" is converted into the text "Let's play." The converted text data is sent to the next analysis step.

[1755] The server uses a generative AI model to analyze the text data and generate a message to convey to the pet. For example, the text "Let's play" is analyzed to mean "It's play time." The analysis result is sent to the message generation step.

[1756] The server uses speech synthesis software to convert the generated message into sounds and actions that the pet can understand, and communicates them to the pet through a speaker. For example, the message "It's playtime" is converted into sounds and actions that a dog can understand, and played back from the speaker. The pet understands the sounds and actions, and prepares to play with the user.

[1757] Examples of specific examples and prompts

[1758] Specific examples

[1759] Pet Sound Analysis:

[1760] If the pet barks "woof woof," the microphone captures the sound and the server notifies the user that "your pet is hungry."

[1761] Human language analysis:

[1762] When the user says "let's play," a microphone captures the voice and the server communicates through the speaker with sounds and actions that the pet understands as "play time."

[1763] Prompt Sentence Examples

[1764] Pet Sound Analysis:

[1765] If your pet barks "woof woof," explain the process of analyzing that sound and informing the user that "your pet is hungry."

[1766] Human language analysis:

[1767] When a user says "Let's play," describe the process for analyzing that speech and communicating "It's playtime" to your pet in a way that makes sense.

[1768] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1769] Step 1: Capture your pet's sounds

[1770] The device's microphone captures the sound of your pet barking. For example, if your dog barks "woof woof," the audio data is collected through the microphone. The input is your pet's bark, and the output is audio data. The collected audio data is sent to the server in real time.

[1771] Step 2: Analyzing bird calls and converting them into text

[1772] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice "woof woof" is converted into the text "woof woof." The input is voice data and the output is text data. The converted text data is sent to the next analysis step.

[1773] Step 3: Semantic analysis of the text

[1774] The server uses a generative AI model to analyze the text data and understand the pet's intention. For example, the text "woof woof" is analyzed to mean "I'm hungry." The input is the text data, and the output is the analysis result indicating the pet's intention. The analysis result is sent to the message generation step.

[1775] Step 4: Message generation and notification

[1776] The server generates a message based on the analysis results, saying "Your pet is hungry," and notifies the user. For example, the notification is sent to the user via a smartphone app. The input is the analysis results, and the output is a notification message to the user. The user can check the notification on the app to understand the status of their pet.

[1777] Step 5: Capturing human language

[1778] The microphone on the device captures the user's words. For example, when a user says "Let's play," the voice data is collected through the microphone. The input is the user's words, and the output is voice data. The collected voice data is sent to the server in real time.

[1779] Step 6: Word analysis and text conversion

[1780] The server analyzes the captured voice data using voice recognition software and converts it into text data. For example, the voice saying "Let's play" is converted into the text "Let's play." The input is voice data and the output is text data. The converted text data is sent to the next analysis step.

[1781] Step 7: Semantic analysis of the text

[1782] The server uses a generative AI model to analyze the text data and generate a message to be conveyed to the pet. For example, the text "Let's play" is analyzed to mean "It's play time." The input is text data, and the output is the message to be conveyed to the pet. The analysis result is sent to the message generation step.

[1783] Step 8: Generate a message and send it to your pet

[1784] The server uses speech synthesis software to convert the generated message into sounds and actions that the pet can understand, and communicates them to the pet through a speaker. For example, the message "It's playtime" is converted into sounds and actions that the dog can understand, and played back from the speaker. The input is the message to be communicated to the pet, and the output is sounds and actions that the pet can understand. The pet understands the sounds and actions, and gets ready to play with the user.

[1785] (Application example 2)

[1786] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1787] Conventional pet care systems have had problems such as difficulty in accurately understanding pet sounds and behavior and taking appropriate action. They also lacked a means to notify owners when their pets were hungry and quickly order pet food. Furthermore, they lacked an effective means to facilitate smooth communication between owners and their pets.

[1788] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from the cries based on the collected data, means for notifying the owner based on the pet's condition, and means for ordering pet food. This makes it possible to accurately determine the pet's condition and respond quickly. It also facilitates communication between the owner and the pet and automates pet food ordering.

[1789] "Audio data" refers to sound information, including animal sounds.

[1790] "Characteristics of pets" refers to individual characteristics such as the personality and sounds of pets.

[1791] "Means for identifying the type of animal" refers to a method or device for identifying the type of pet.

[1792] "General animal characteristic data" refers to information about widely known animal characteristics.

[1793] "Means for understanding the condition of pets from their cries" refers to methods and devices for analyzing pet cries to understand their condition.

[1794] "Means for notifying owners based on the status of their pets" refers to methods or devices for informing owners of the status of their pets.

[1795] "Pet food ordering means" refers to a method or device for purchasing pet food online.

[1796] "Means for enabling conversation between humans and pets using AI" refers to methods and devices for exchanging information between humans and pets using artificial intelligence.

[1797] "Means for understanding the health and emotional state of a pet" refers to a method or device for understanding the physical condition and emotions of a pet.

[1798] The following system configuration will be described as an embodiment of the present invention.

[1799] System Configuration

[1800] This system includes a means for collecting audio data (animal cries), a means for collecting characteristics of pet animals (personality, cries, etc.), a means for identifying the type of animal, a means for collecting general data on animal characteristics, a means for determining the condition of pets from their cries based on the collected data, a means for notifying owners based on the condition of their pets, and a means for ordering pet food.

[1801] Hardware and software used

[1802] Hardware: Smartphone, microphone

[1803] Software: Python, SpeechRecognition library, Requests library

[1804] Data processing and calculation

[1805] Audio data collection and analysis

[1806] The server records the pet's cries using the smartphone's microphone. The recorded audio data is converted to text using the SpeechRecognition library. It is then analyzed using a generative AI model to understand the pet's condition.

[1807] Collecting and analyzing owner speech

[1808] The server uses the smartphone's microphone to record the owner's words, converts the recorded voice data into text using the SpeechRecognition library, and then analyzes it using a generative AI model to generate a message to be conveyed to the pet.

[1809] Pet food orders

[1810] The server notifies the owner based on the pet's condition. For example, if the pet cries "I'm hungry," the server generates a prompt to order pet food and notifies the owner. When the owner orders pet food through the app, recommended products are displayed taking into account the pet's preferences and allergies.

[1811] Specific examples

[1812] Analyzing pet cries: When a pet cries "I'm hungry," the server analyzes the cry and notifies the owner with a message that "Your pet is hungry."

[1813] Analyzing owner's words: When the owner says "Let's play," the server analyzes the words and conveys the message in a way that the pet can understand.

[1814] Prompt Sentence Examples

[1815] "Analyze your pet's cries and let us know their condition."

[1816] "Analyze the owner's words and generate a message to convey to the pet."

[1817] In this way, the present invention allows for accurate understanding of the pet's condition and prompt response, facilitates communication between pet owners and their pets, and automates pet food ordering.

[1818] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1819] Step 1:

[1820] The server records the pet's cries using the smartphone's microphone. The input is the pet's cries, and the output is an audio data file. This audio data file is used for subsequent analysis.

[1821] Step 2:

[1822] The server converts the recorded audio data into text data using the SpeechRecognition library. The input is an audio data file, and the output is text data. This text data is input into the generative AI model.

[1823] Step 3:

[1824] The server analyzes the text data using a generative AI model to understand the pet's state. The input is text data, and the output is information indicating the pet's state. For example, the analyzed state is "hungry."

[1825] Step 4:

[1826] The server notifies the owner based on the pet's status. The input is information indicating the pet's status, and the output is a notification message to the owner. For example, a message saying "your pet is hungry" is sent to the owner.

[1827] Step 5:

[1828] The server records the owner's words using the smartphone's microphone. The input is the owner's words, and the output is an audio data file. This audio data file is used for subsequent analysis.

[1829] Step 6:

[1830] The server converts the recorded voice data of the owner into text data using the SpeechRecognition library. The input is a voice data file, and the output is text data. This text data is input into the generative AI model.

[1831] Step 7:

[1832] The server uses a generative AI model to analyze the text data and generate a message to be conveyed to the pet. The input is text data, and the output is a message to be conveyed to the pet. For example, the command "Let's play" is analyzed.

[1833] Step 8:

[1834] The server provides a means for ordering pet food based on the pet's condition. The input is information indicating the pet's condition, and the output is pet food ordering information. For example, if the pet cries "I'm hungry," an order for pet food is automatically placed.

[1835] Step 9:

[1836] The server displays recommended products to the pet owner, taking into account the pet's preferences and allergy information. The input is the pet's preferences and allergy information, and the output is a list of recommended products. This allows the pet owner to select the appropriate pet food.

[1837] Step 10:

[1838] The server notifies the pet owner that the pet food order has been completed. The input is the order completion information, and the output is a notification message to the pet owner. For example, a message saying "Pet food order completed" is sent to the pet owner.

[1839] Example 3

[1840] Next, a third embodiment of the third embodiment will be described. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1841] Conventional systems for understanding pet health and emotional states only analyze voice data and do not take behavioral data into account, making it difficult to accurately understand a pet's condition. Additionally, there is a lack of a way to quickly notify the owner of the analysis results, making it difficult to detect abnormalities in a pet early on.

[1842] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[1843] In this invention, the server includes means for collecting voice data, means for collecting animal characteristics, means for identifying animal species, means for collecting general animal characteristic data, means for determining the status of pet animals from their cries based on the collected data, means for collecting animal behavior data, means for analyzing the collected behavior data, and means for notifying the results of the analysis. By analyzing both the voice data and the behavior data, it is possible to more accurately determine the health and emotional state of pets and promptly notify their owners.

[1844] "Audio Data" refers to animal sounds and other audio information.

[1845] "Animal characteristics" refers to individual characteristics such as an animal's personality, vocalizations, and behavioral patterns.

[1846] "Type of animal" refers to a classification of animals such as dogs, cats, birds, etc.

[1847] "General animal characteristic data" refers to data such as personality and behavioral patterns common to a particular type of animal.

[1848] "Animal condition" refers to the animal's physical and emotional state.

[1849] "Behavioral data" refers to information about an animal's movements and behavior patterns.

[1850] "Analysis results" refers to the diagnosis and evaluation output by the AI ​​model based on the collected voice data and behavioral data.

[1851] "Notification means" refers to a method or device for notifying the owner of the analysis results.

[1852] The present invention is a system for understanding the health and emotional state of a pet. A specific embodiment of this system will be described below.

[1853] First, the user uses a device such as a smartphone or tablet to collect the sounds and behavior of their pet. The device is equipped with a microphone and camera, and these devices are used to capture audio data and behavior data in real time. For example, when a user launches the smartphone app, points it at their pet, and presses the "Start Recording" button, the microphone begins to collect the pet's sounds. At the same time, the camera captures the pet's movements.

[1854] The device then transmits the collected data to a server in real time using Wi-Fi or mobile data, and the data is encrypted to ensure security.

[1855] The server inputs the received audio and behavior data into an AI model, which is built using machine learning frameworks such as TensorFlow and PyTorch. The AI ​​model analyzes the frequency and volume of the pet's cries, as well as the speed and direction of its movements, to detect abnormal patterns and patterns of delight.

[1856] The analysis results are sent from the server to the device, and the user can check the results through a smartphone app. For example, the server may send a warning message such as "Your pet may be sick." If the pet is happy, it may send a message such as "Your pet is happy."

[1857] As a concrete example, the following prompt sentence could be input to a generative AI model:

[1858] "Analyze your pet's vocalizations and behavioral data to determine their health and emotional state. Alert you if they're sick and let you know if they're happy."

[1859] Using this prompt, the AI ​​model can accurately analyze the pet's condition and provide appropriate feedback to the user.

[1860] As described above, by analyzing both the voice data and the behavioral data, the present invention makes it possible to grasp the health condition and emotional state of the pet more accurately and to notify the owner promptly. The flow of the identification process in the third embodiment will be described with reference to FIG.

[1861] Step 1:

[1862] The user launches the smartphone app and collects the sounds and behavior of their pet. Specifically, when the user presses the "Start Recording" button, the device's microphone begins collecting the pet's sounds. At the same time, the camera captures the pet's movements. The input is the pet's sounds and behavior, and the output is the collected audio data and behavior data.

[1863] Step 2:

[1864] The device transmits the collected voice and behavioral data to a server. Specifically, the device transmits data in real time using Wi-Fi or mobile data. The input is the collected voice and behavioral data, and the output is the data transmitted to the server.

[1865] Step 3:

[1866] The server inputs the received voice data and behavioral data into an AI model. Specifically, the server inputs the data into an AI model built using machine learning frameworks such as TensorFlow and PyTorch. The input is the voice data and behavioral data sent to the server, and the output is the analysis results by the AI ​​model.

[1867] Step 4:

[1868] The server determines the pet's health and emotional state based on the analysis results of the AI ​​model. Specifically, the AI ​​model analyzes the frequency and volume of the pet's cries, as well as the speed and direction of its movements, to detect abnormal patterns and patterns of joy. The input is the analysis results of the AI ​​model, and the output is a judgment of the pet's health and emotional state.

[1869] Step 5:

[1870] The server sends the analysis results to the device. Specifically, the server sends a message generated based on the analysis results to the device. The input is the judgment result of the pet's health condition and emotional state, and the output is the message sent to the device.

[1871] Step 6:

[1872] The user checks the analysis results through the device. Specifically, the user opens the smartphone app and checks the notified message. For example, a warning message such as "Your pet may be sick" or a message such as "Your pet is happy" may be displayed. The input is the message sent to the device, and the output is the analysis result checked by the user.

[1873] (Application example 3)

[1874] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1875] Understanding a pet's health and emotional state is important for pet owners, but existing methods make it difficult to detect abnormalities early on. There are also limited ways to accurately determine whether a pet is happy. This can lead to inadequate health management and understanding of a pet's emotions, potentially reducing the pet's quality of life. Furthermore, the lack of a way to quickly notify owners of abnormalities can delay emergency response.

[1876] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[1877] In this invention, the server includes means for collecting voice data (animal cries), means for collecting characteristics of pet animals (personality, cries, etc.), means for identifying animal species, means for collecting general animal characteristic data, means for determining the pet's condition from its cries based on the collected data, means for analyzing the pet's condition and sending a notification if an abnormality is detected, means for determining the pet's health and emotional state in real time, means for issuing an alert to the owner if an abnormality is detected, and means for informing the owner if the pet is happy. This makes it possible to quickly and accurately determine the pet's health and emotional state, and to immediately notify the owner if an abnormality occurs.

[1878] "Audio data" refers to data that has been recorded in digital format, such as animal cries.

[1879] "Animal characteristics" refers to the animal's unique characteristics, such as its personality or sounds.

[1880] "Means for identifying animal species" refers to methods or techniques for identifying the species of an animal.

[1881] "General animal characteristic data" refers to data about widely known animal characteristics.

[1882] "Means for understanding the condition of pets" refers to methods for analyzing the health and emotional state of pets based on collected data.

[1883] "Means for analyzing a pet's condition" refers to a method for analyzing a pet's cries and behavior to assess its condition.

[1884] "Means for sending a notification when an abnormality is detected" refers to a method for alerting the owner when an abnormality is detected in the pet.

[1885] "Means for understanding a pet's health and emotional state in real time" refers to a method for instantly analyzing a pet's condition and providing the results in real time.

[1886] "Means for issuing a warning to the owner when an abnormality is detected" refers to a method for issuing a warning to the owner when an abnormality is detected in the pet.

[1887] The "means for conveying information about a happy pet to its owner" refers to a method for detecting a happy state of a pet and conveying that information to its owner.

[1888] As an embodiment of the present invention, a pet security alert system will be described as an example. This system analyzes the sounds and behavior of pets in real time and notifies the owner if an abnormality is detected.

[1889] System Configuration

[1890] Hardware:

[1891] Smartphone

[1892] microphone

[1893] software:

[1894] TensorFlow

[1895] librosa

[1896] smtplib

[1897] Data processing and calculation

[1898] Audio data collection:

[1899] The device (smartphone) uses a microphone to collect the sounds your pet makes, and the audio data is recorded in digital format.

[1900] Analysis of audio data:

[1901] The server converts the collected audio data into MFCC (Mel-Frequency Cepstrum Coefficients) using librosa, which allows the extraction of audio data features.

[1902] Analysis by AI model:

[1903] The server inputs the audio data into a generative AI model pre-trained with TensorFlow to analyze the pet's health and emotional state, which is then used to detect abnormalities in the pet.

[1904] Send notifications:

[1905] If an abnormality is detected, the server uses smtplib to send an email notification to the owner, containing detailed information about the pet's abnormal condition.

[1906] Specific examples

[1907] For example, you can record your pet's cries and input the audio data into the application. The AI ​​model analyzes the audio data, and if it detects an abnormality, it will notify the owner by email. This notification will include a message such as, "An abnormality has been detected with your pet. Please check immediately."

[1908] Prompt Sentence Examples

[1909] Examples of prompts to be input to a generative AI model include:

[1910] Write a Python program that analyzes pet noise data and notifies owners if an abnormality is detected. Use TensorFlow and librosa to analyze the audio data, and smtplib to send email notifications.

[1911] In this way, a system can be provided that can grasp the health and emotional state of a pet in real time and respond quickly if there is an abnormality.

[1912] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[1913] Step 1:

[1914] Audio data collection

[1915] The device (smartphone) collects pet sounds usi...

Claims

1. means for receiving audio data including animal sounds and user speech; means for collecting characteristics, including personalities, of animals kept by users; a means for identifying the type of animal; means for collecting trait data about general animal characteristics; means for generating a prompt sentence instructing the generation of information regarding at least one of the status and requirements of the animal being raised by the user, based on the voice data, the characteristics of the personality of the animal being raised by the user, the identified animal type, and the collected characteristic data; means for generating information regarding at least one of the status and requirements of the animals kept by the user using the generated prompt sentence and a generative AI model; means for notifying a user of the generated information; means for identifying an emotional state of a user using an emotion engine; a means for generating a prompt sentence based on the identified emotional state of the user, the prompt sentence instructing the animal to translate an instruction spoken by the user into a form that the animal can understand; a means for translating, using the generated prompt sentence and a generative AI model, an instruction spoken by the user to an animal kept by the user into a form that the animal can understand; means for outputting the translated voice instructions; A system including:

2. 2. The system according to claim 1, wherein the means for translating instructions uttered by the user to the animal into a form understandable to the animal outputs the instructions as an image understandable to the animal.

3. The system of claim 1 , wherein the means for generating information about at least one of the condition and needs of the animal kept by the user identifies at least one of the health and emotional state of the animal.

Citation Information

Patent Citations

  • Translation module and speech translation device using the same

    JP2004212685A

  • Speech translation method and system using a multilingual text-to-speech synthesis model

    JP2021511534A

  • Persona chatbot control method and system

    JP2022180282A

  • Device and method for judging dog's feeling from cry vocal character analysis

    WO2003015076A1