System

A system using sensors and machine learning to translate pets' behavior and sounds into understandable feedback for users addresses the challenge of pet communication, enabling prompt and appropriate responses.

JP2026037921APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024141255
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Pet owners face difficulties in understanding their pets' intentions and emotions due to the lack of precise and real-time analysis of animal behavior and sounds, making it challenging to communicate effectively with them.

Method used

A system that collects animal data in real-time using sensors and microphones, analyzes it using machine learning algorithms, and translates the meaning of the animal's behavior and voice through a dedicated language model, providing feedback to the user.

Benefits of technology

Enables quick and appropriate responses to pets' requests and emotions by accurately analyzing their movements, behavior, and sounds, facilitating effective communication between users and pets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026037921000001_ABST
    Figure 2026037921000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: a terminal including a sensor and a microphone for sensing movement, behavior, and voice of an animal; a server for receiving and analyzing data of movement, behavior, and voice of the animal transmitted from the terminal; a means for updating a language model for translating a meaning of behavior and voice of the animal based on the analyzed data; and a means for generating feedback information to a user based on the language model and transmitting the feedback information to the terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Pet owners often find it difficult to understand their pets' intentions and emotions, especially when communicating with animals that do not have the ability to speak. Therefore, there is a need for effective methods to properly understand and quickly respond to pets' requests. Furthermore, existing technologies lack the precision and real-time capabilities to analyze animal behavior and sounds, making them impractical for communicating with users. [Means for solving the problem]

[0005] The system of the present invention collects animal data in real time using a terminal equipped with sensors and a microphone that detect the animal's movements, behavior, and voice. The collected data is sent from the terminal to a server, where it is analyzed. The analyzed data is used to translate the meaning of the animal's behavior and voice using a machine learning algorithm, and a language model dedicated to the animal is generated and updated based on this. Furthermore, feedback information for the user is generated based on the updated language model and notified to the user via the terminal, making it easier for the user to understand the pet's intentions. This system improves communication between the user and the pet, enabling quick and appropriate responses to the pet's requests and emotions.

[0006] Understood. Below are definitions of important terms contained in the claims.

[0007] "Animals" refers to living creatures such as pets kept by users, including dogs and cats.

[0008] "Terminal" refers to a device equipped with sensors and a microphone that detects animal movements, behavior, and sounds.

[0009] "Sensor" refers to a device used to detect and collect data on animal movements.

[0010] "Microphone" refers to a device used to record animal sounds and collect the data.

[0011] "Data" refers to information about animal movements, behavior, and sounds collected by the device.

[0012] The term "server" refers to a computer system that receives and analyzes data sent from a terminal.

[0013] "Analysis" refers to the process of determining the meaning of animal behavior and vocalizations based on the data received.

[0014] "Machine learning algorithms" refer to mathematical models that learn patterns in data and use them to analyze animal behavior and vocalizations.

[0015] A "language model" refers to a model that translates an animal's actions and voices into human language and uses that to express the animal's intentions.

[0016] "Feedback information" refers to information generated based on the analysis results and used to provide the user with information about the animal's intentions and emotions.

[0017] "User" refers to an individual who keeps a pet or a user of the pet. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds, and a server that performs the analysis and translation.

[0040] About program processing

[0041] 1. Data collection

[0042] The device uses sensors and microphones to collect real-time information on the animal's movements, behavior, and voices. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphones record the frequency and volume of its calls. In addition, cameras and other devices are used to capture information on the animal's facial expressions and posture.

[0043] 2. Data transmission

[0044] The device converts the collected data into a certain format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS), which provides the data to the server in real time.

[0045] 3. Data Analysis

[0046] The server pre-processes the data it receives, then analyzes it using machine learning algorithms that detect specific patterns in animal behavior and vocalizations and infer what they mean.

[0047] 4. Creating and updating a language model

[0048] The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[0049] 5. Feedback Generation

[0050] The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[0051] 6. User Notices

[0052] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, it can also use a voice notification function to notify the user more clearly.

[0053] Specific examples

[0054] Below is a concrete example of a dog standing on its hind legs and barking "woof woof."

[0055] 1. The device uses a motion sensor to detect the dog standing up on its hind legs, and simultaneously records the dog's bark "woof woof" with a microphone.

[0056] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol.

[0057] 3. The server cleans and normalizes the data it receives and uses machine learning algorithms to determine that the "reaching sounds" mean "a request for something."

[0058] 4. The server updates the language model based on the results of the classification and interprets this behavior as meaning "wanting a snack."

[0059] 5. The server generates feedback information such as "The dog wants something. For example, it might be a good idea to give it a treat," and sends this information to the terminal.

[0060] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[0061] In this way, the system of the present invention effectively supports communication between the user and the pet, and enables quick and appropriate responses to the pet's requests and emotions.

[0062] The processing flow will be explained below.

[0063] Step 1:

[0064] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time: the motion sensors capture the animal's location, speed, and direction, the camera captures facial expressions and posture, and the microphone records the frequency and volume of the animal's calls.

[0065] Step 2:

[0066] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[0067] Step 3:

[0068] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS), in real time and optimized to minimize latency.

[0069] Step 4:

[0070] The server pre-processes the received data before analyzing it, which involves cleaning the data (removing noise and missing data) and normalizing it (aligning the numerical data to the same scale).

[0071] Step 5:

[0072] The server then feeds the pre-processed data into a machine learning algorithm that identifies patterns in the animal's behavior and vocalizations and uses these patterns to parse the animal's intentions and emotions using pre-trained models.

[0073] Step 6:

[0074] The server generates and updates a language model specifically for animals based on the analysis results, and insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation.

[0075] Step 7:

[0076] The server generates feedback information for the user based on the updated language model, written in natural language, to help the user understand the meaning of the animal's actions and sounds.

[0077] Step 8:

[0078] The server sends the generated feedback information to the terminal, allowing the terminal to obtain the latest information in real time.

[0079] Step 9:

[0080] The device then notifies the user of the received feedback information, which can be displayed on the smartphone screen or provided as a voice message using the voice notification function, such as "Your dog wants to play."

[0081] Step 10:

[0082] Based on the feedback information provided by the device, the user can respond appropriately to the pet's requests, such as giving treats or taking the pet out to play, thereby increasing the pet's satisfaction.

[0083] In this way, the system of the present invention goes through detailed processing steps, allowing the user to understand the intentions and emotions of their pet and take prompt and appropriate action.

[0084] Example 1

[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0086] In modern times, communication between animals and humans remains difficult, particularly due to limited means for properly understanding an animal's intentions and emotions. This makes it difficult for pet owners to understand their pet's requests and emotions and respond quickly and appropriately. Furthermore, conventional systems have struggled to accurately analyze an animal's behavior and voice and provide real-time feedback. Therefore, there is a need for a system that can accurately analyze an animal's movements, behavior, and voice, and then translate the animal's intentions and emotions and notify the user.

[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0088] In this invention, the server includes means for generating and updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and sending it to the terminal, and means for notifying the user of the feedback information to their smartphone. This enables the animal's movements, actions, and voices to be analyzed with high accuracy, allowing the user to quickly and appropriately understand the animal's intentions and emotions.

[0089] A "terminal" is a device equipped with a sensor and microphone that detects the movements, behavior, and sounds of animals, and transmits the data collected from these to a server.

[0090] A "sensor" is a device that detects physical changes and converts them into electrical signals, and in this invention it is used to sense the movement and location of animals.

[0091] A "microphone" is a device that converts sound into an electrical signal, and in this invention it is used to record the frequency and volume of animal cries.

[0092] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[0093] "Data" refers to information about animal movements, behaviors, and sounds collected by sensors and microphones.

[0094] "JSON format" is an abbreviation for JavaScript (registered trademark) Object Notation, and is a lightweight data exchange format for structurally representing data.

[0095] "HTTPS" stands for Hypertext Transfer Protocol Secure, a protocol for encrypting data communications.

[0096] A "machine learning algorithm" is a computer algorithm used to analyze data and recognize patterns.

[0097] A "language model" is a collection of databases and algorithms that translate the meaning of animal behaviors and vocalizations based on analyzed data.

[0098] "Feedback information" is a message in natural language that explains the animal's intentions and emotions and is communicated to the user.

[0099] A "smartphone" is a portable information terminal equipped with mobile communication and computer functions.

[0100] The present invention provides a system that facilitates communication with animals by collecting, analyzing, and translating data on animal movements, behaviors, and voices, and notifying the user of the translation. The following describes in detail the embodiments of the present invention.

[0101] System Overview

[0102] The system consists of a device equipped with sensors and microphones that detect animal movements, behavior, and sounds, and a server that performs analysis and translation. Specifically, the device includes a motion sensor, microphone, and camera. These devices acquire information on the animal's location, speed, direction, frequency, and volume of its calls, as well as its facial expressions and posture. The collected data is converted into JSON format and securely sent to the server using the HTTPS protocol.

[0103] Hardware and Software

[0104] Device: Equipped with motion sensors, a microphone, and a camera to detect animal movements, behavior, and sounds. Includes software to convert this data into JSON format.

[0105] Server: Contains the processing unit and storage for parsing the received data, as well as software for data cleaning, normalization, and analyzing the data using machine learning algorithms, including algorithms for generating and updating language models.

[0106] Analyzing data and generating feedback

[0107] The server preprocesses the received data, for example by removing outliers and normalizing the data. The preprocessed data is then analyzed by machine learning algorithms. These algorithms use neural networks and deep learning models to recognize patterns in animal behavior and vocalizations and infer their meaning. For example, if a dog stands on its hind legs and barks "woof woof," it may determine that this behavior indicates "it wants a treat."

[0108] Based on the analysis results, the server generates and updates a language model specific to the animal. This language model is a collection of databases and algorithms for translating the meaning of the animal's actions and vocalizations. The server then generates feedback information for the user, which explains the animal's intentions and emotions to the user in natural language. For example, a message such as "The dog is asking for a treat" may be generated.

[0109] User Notifications

[0110] The device receives the feedback information sent from the server and notifies the user's smartphone. Notifications are displayed as push notifications, and audio notifications are also possible if necessary. For example, a notification such as "Your dog wants something. Please give it a treat" may appear on the smartphone.

[0111] Specific examples

[0112] A specific example of the system's operation is shown below.

[0113] When a dog stands on its hind legs and barks "woof woof"

[0114] 1. The device's motion sensor detects the dog standing up on its hind legs, and at the same time, the microphone records the dog's bark.

[0115] 2. The device converts the collected data into JSON format and sends it to the server via HTTPS protocol.

[0116] 3. The server cleans and normalizes the data, compares it with historical data, and analyzes it. Based on this analysis, it recognizes that the "standing up on hind legs" and "woof woof" indicate a request for a treat.

[0117] 4. The server updates the language model and generates feedback and sends it to the device.

[0118] 5. The device notifies the user of the feedback information on their smartphone.

[0119] Examples of prompts for generative AI models

[0120] "What does it mean when a dog stands on its hind legs and whines?"

[0121] "When a cat wags its tail and meows, what emotion does it express?"

[0122] "What does the behavior of a bird flapping its wings and chirping loudly indicate?"

[0123] Such prompts can be used to query the generative AI model about the meaning of an animal's behavior, allowing the system to provide appropriate feedback.

[0124] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0125] Step 1: Collect data

[0126] The device uses sensors, microphones, and cameras to collect animal movements, behaviors, and sounds in real time.

[0127] Input: Animal movements, behaviors, and sounds.

[0128] Specific operations: The device's built-in sensors detect the animal's position, speed, and direction, the microphone captures the frequency and volume of the animal's calls, and the camera captures images of its facial expressions and posture.

[0129] Output: Raw data from sensors, microphones, and cameras.

[0130] Step 2: Sending data

[0131] The terminal converts the collected data into JSON format and sends it to the server using the HTTPS protocol.

[0132] Input: Raw data from sensors, microphones, and cameras.

[0133] Specific operation: The terminal software converts the raw data into JSON format and transmits it securely via HTTPS.

[0134] Output: Data converted to JSON format.

[0135] Step 3: Preprocessing the data

[0136] The server cleans and normalizes the received data.

[0137] Input: JSON formatted data.

[0138] Specific operation: The server performs outlier removal, noise filtering, and data scaling.

[0139] Output: Cleaned and normalized data.

[0140] Step 4: Analyze the data

[0141] The server analyzes the data using machine learning algorithms.

[0142] Input: Cleaned and normalized data.

[0143] How it works: The server inputs data into a machine learning model (e.g., a neural network) to analyze the animal's behavior and vocal patterns.

[0144] Output: Analysis results of behavioral and vocal patterns.

[0145] Step 5: Generate and update a language model

[0146] Based on the analysis results, the server translates the meaning of the animal's behavior and voice, and generates and updates a language model specifically for the animal.

[0147] Input: Analysis results.

[0148] Specific operation: The server compares the analysis results with the existing language model and adds new behavioral patterns to the model.

[0149] Output: An updated language model.

[0150] Step 6: Generate feedback information

[0151] The server generates feedback information for the user.

[0152] Input: The updated language model.

[0153] Specific operation: The server converts the analysis results into natural language and generates a message that is easy for the user to understand.

[0154] Output: The generated feedback information.

[0155] Step 7: User Notification

[0156] The terminal notifies the user's smartphone of the feedback information from the server.

[0157] Input: The generated feedback information.

[0158] Specific operation: The device displays a push notification on the smartphone and, if necessary, plays a voice notification.

[0159] Output: The notification message that is displayed to the user.

[0160] This embodied processing allows the system to accurately analyze animal movements, behaviors, and sounds and provide information to the user.

[0161] (Application example 1)

[0162] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0163] While conventional pet communication systems can recognize pets' movements and voices, they are unable to recommend appropriate products and services to users based on that information. This makes it difficult for users to find products and services that meet their pets' needs and to establish effective communication with their pets. Virtual stores, in particular, lack product and service recommendations tailored to pets' specific needs, limiting the user experience.

[0164] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0165] In this invention, the server includes a terminal equipped with a sensor and a microphone that detects the movements, behavior, and voice of the animal, means for receiving and analyzing data on the movements, behavior, and voice of the animal transmitted from the terminal, means for updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for automatically suggesting related products and services in a virtual store based on the feedback information, thereby enabling the user to understand the meaning of their pet's actions and voices and easily find products and services in the virtual store that meet their pet's needs.

[0166] A "terminal" is a device equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[0167] The "server" is a system that receives and analyzes data on animal movements, behavior, and voices sent from the terminal.

[0168] A "language model" refers to the algorithms and datasets used to translate the meaning of animal behavior and vocalizations.

[0169] A "virtual store" is a virtual store operated on the Internet, which is a platform where real products can be purchased and services can be provided.

[0170] "Feedback information" refers to information provided to the user that is generated based on the meaning of the analyzed animal's behavior and voice.

[0171] "JSON format" is a data serialization format, an abbreviation for JavaScript Object Notation, and is a standard format for expressing structured data.

[0172] A "machine learning algorithm" is an algorithm that learns patterns and features from data and makes predictions and classifications for new data.

[0173] A "sensor" is a device that senses physical phenomena or conditions and collects that information as data.

[0174] A "microphone" is a device that converts sound waves into electrical signals and is used to collect animal sounds.

[0175] The "HTTPS protocol" is an abbreviation for Hypertext Transfer Protocol Secure, and is a communication protocol for securely sending and receiving data over the Internet.

[0176] "Animal-specific language models" are specialized algorithms and datasets for analyzing animal behavior and vocal patterns and translating their meaning.

[0177] This invention is a system that collects, analyzes, and translates animal movements, behaviors, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behaviors, and sounds, and a server that performs the analysis and translation.

[0178] System configuration

[0179] Hardware

[0180] Terminal: Equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[0181] Server: Carries out analysis and translation processing. Implements machine learning algorithms.

[0182] software

[0183] On the device:

[0184] A program that collects movements and voices in real time and converts the data into JSON format.

[0185] A communications program that sends data to a server using the HTTPS protocol.

[0186] Server side:

[0187] A program that receives and preprocesses data.

[0188] A program that uses machine learning algorithms to analyze data and translate its meaning.

[0189] A program that generates feedback information based on translation results and sends it to the terminal.

[0190] A program that automatically suggests related products and services in a virtual store based on feedback information.

[0191] Operation flow

[0192] 1. The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphone records the frequency and volume of its calls. Cameras and other devices also capture information on the animal's facial expressions and posture.

[0193] 2. The device converts the collected data into JSON format and sends it to the server using a secure protocol (HTTPS), which provides the data to the server in real time.

[0194] 3. The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning and normalizing the data. The machine learning algorithms detect specific patterns in the animals' behavior and vocalizations and infer what they mean.

[0195] 4. The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[0196] 5. The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[0197] 6. The device receives the feedback information sent from the server and notifies the user. For example, the device can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, a voice notification function can be used to notify the user more clearly.

[0198] 7. The server automatically suggests related products and services in the virtual store based on the feedback information. For example, based on the translation result that the dog "wants to play," the server suggests a new toy in the virtual store.

[0199] Specific examples

[0200] 1. The user uses a smartphone to observe the dog's behavior in real time.

[0201] 2. Your pet wags its tail and barks "woof woof."

[0202] 3. The app will notify you, "Your dog seems happy and might want to play with his new toy."

[0203] 4. Virtual stores suggest toys for pets.

[0204] Prompt Sentence Examples

[0205] Input: The smartphone detects the dog's tail wagging and the sound of it barking "woof woof."

[0206] Processing: Data is sent to the server and analyzed by the translation system. Feedback information is generated based on the translation results and sent to the smartphone.

[0207] Output: Your dog seems happy and may be eager to play with his new toy.

[0208] This configuration makes it possible to construct a system within the scope of the invention and realize effective communication between pets and users.

[0209] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0210] Step 1:

[0211] The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. Inputs include location information, speed, and direction of movement from the sensors, as well as the frequency and volume of sounds from the microphones. This data is output as information on the animal's behavior and sounds detected through the sensors and microphones.

[0212] Step 2:

[0213] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. The input is the behavioral and voice data obtained in step 1. These data are packaged in JSON format and sent to the server via a secure communication channel for output.

[0214] Step 3:

[0215] The server preprocesses the data it receives. It takes as input the JSON data sent in step 2. It cleans and normalizes the data, resulting in clean data that can be meaningfully analyzed.

[0216] Step 4:

[0217] The server analyzes the data using machine learning algorithms. The input is the clean, pre-processed data. The machine learning algorithms detect specific patterns in the animal's behavior and vocalizations and infer their meaning. The output is the meaning of the animal's behavior and vocalizations.

[0218] Step 5:

[0219] The server translates the meaning of the animal's actions and voices based on the analysis results and updates the animal-specific language model. The input is the analysis results from step 4. Based on this data, the server generates translations corresponding to the animal's actions and voices and updates the language model. The output is the latest language model.

[0220] Step 6:

[0221] The server generates feedback information for the user based on the updated language model. The inputs are the latest language model and data on the animal's behavior and voice. The feedback information is formed as a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand. The output is the feedback information.

[0222] Step 7:

[0223] The device receives feedback information sent from the server and notifies the user. The input is the feedback information sent from the server. This is displayed on the screen of the user's smartphone or smart glasses, providing feedback to the user. The output is a notification sent to the user.

[0224] Step 8:

[0225] The server automatically suggests related products and services in the virtual store based on the feedback information. The input is the feedback information. For example, based on the feedback information "Your dog wants to play," the virtual store suggests a new toy. The output is information about the suggested products and services.

[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0227] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and voices to facilitate communication with users, and provides more efficient feedback by also analyzing the user's emotions. This system is composed of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and voices, a server that performs analysis and translation, and an emotion engine that recognizes the user's emotions.

[0228] About program processing

[0229] 1. Data collection

[0230] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The sensors capture the animal's location, speed, and direction, the camera captures its facial expressions and posture, and the microphone records the frequency and volume of its calls. At the same time, the user's voice and facial expressions are also collected using an emotion engine.

[0231] 2. Data transmission

[0232] The device converts the collected data into a specific format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS). This allows the data to be provided to the server in real time, along with the user's emotional data.

[0233] 3. Data Analysis

[0234] The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning the data (removing noise and missing data) and normalizing it (aligning numerical data to the same scale). Not only animal behavior and voices, but also user emotion data are analyzed.

[0235] 4. Creating and updating a language model

[0236] The server uses the analysis results to translate the meaning of the animal's behavior and voice, and generates and updates a language model specifically for animals. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. Additionally, the user's emotional data is also reflected in the language model.

[0237] 5. Feedback Generation

[0238] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[0239] 6. User Notices

[0240] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play. You seem busy, but let's take a moment to play with him." If necessary, it can also be provided as a voice message using the voice notification function.

[0241] Specific examples

[0242] Below is a specific example of a dog standing on its hind legs and barking "woof woof."

[0243] 1. The device uses a motion sensor to detect when the dog stands up on its hind legs, and simultaneously records the dog's bark with a microphone. The camera also captures the user's facial expressions, and the audio is analyzed by an emotion engine.

[0244] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol, along with the user's emotional data.

[0245] 3. The server cleans and normalizes the received data and uses machine learning algorithms to determine whether the "reaching on its hind legs" is a "request for something." User sentiment data is also analyzed.

[0246] 4. The server updates the language model based on the results of the classification and translates this behavior as meaning "wanting a snack." The user's emotional data is also reflected in the language model.

[0247] 5. The server generates feedback information such as "The dog wants something. For example, it would be good to give it a treat. It seems you are a little busy, but please make time to respond," and sends this information to the terminal.

[0248] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[0249] In this way, the system of the present invention effectively supports communication between the user and the pet, enabling the system to respond quickly and appropriately to the pet's needs and emotions. Furthermore, by taking the user's emotions into consideration, the system provides more natural and efficient feedback.

[0250] The processing flow will be explained below.

[0251] Step 1:

[0252] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The motion sensor captures the animal's location, speed, and direction of movement. The camera captures the animal's facial expressions and posture, and the microphone records the frequency and volume of its calls. The device also collects the user's voice and facial expressions using an emotion engine.

[0253] Step 2:

[0254] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[0255] Step 3:

[0256] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS) in real time and optimized to minimize latency, along with the user's emotional data.

[0257] Step 4:

[0258] The server preprocesses the received data, which includes removing noise and missing data and normalizing the data (aligning the numerical data to the same scale), thereby improving the accuracy of the analysis.

[0259] Step 5:

[0260] The server then inputs the pre-processed data into a machine learning algorithm, which identifies patterns in the animal's behavior and vocalizations and uses these patterns to analyze the animal's intentions and emotions, while simultaneously analyzing the user's emotional data.

[0261] Step 6:

[0262] The server generates and updates a language model specifically for animals based on the analysis results. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. User emotion data is also reflected in the language model.

[0263] Step 7:

[0264] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[0265] Step 8:

[0266] The server sends the generated feedback information to the device, allowing the device to obtain the latest information in real time. For example, it generates a message such as, "Your dog wants to play. It seems you're a little busy, so please take a moment to play with him."

[0267] Step 9:

[0268] The device notifies the user of the received feedback information. This notification is displayed on the smartphone screen and also provided as a voice message using the voice notification function, allowing the user to receive feedback both visually and audibly.

[0269] Step 10:

[0270] The user can respond appropriately to their pet's needs based on the feedback information provided by the device, for example, by giving them treats or taking them outside to increase their pet's satisfaction. Further data is collected based on the user's actions, improving the accuracy of the system.

[0271] Example 2

[0272] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0273] Conventional animal-human communication systems simply analyze the animal's behavior and voice, and are unable to provide appropriate feedback that takes into account the user's emotions. This makes it difficult to accurately understand the animal's needs and emotions and respond quickly and appropriately, making it difficult to build a trusting relationship between the user and the animal.

[0274] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing data on animal movements, behaviors, and voices and user emotion data, means for analyzing the preprocessed data using a machine learning algorithm, means for generating and updating a language model based on the analysis results, means for generating feedback information for the user and transmitting it to the terminal, and means for notifying the user of the feedback information. This makes it possible to accurately understand the requests and emotions of the animal and provide quick and accurate feedback taking the user's emotions into consideration.

[0275] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[0276] The "server" is a central processing unit that receives and analyzes data sent from the terminal, generates feedback information, and sends it to the terminal.

[0277] "Data preprocessing" refers to the process of removing noise and missing data from the received data and aligning the numerical data to the same scale.

[0278] A "machine learning algorithm" is a mathematical method that analyzes received and preprocessed data to infer the meaning of animal behavior and vocalizations.

[0279] A "language model" is a model for translating the meaning of animal behavior and voices based on analysis results, and is an updatable data structure.

[0280] "Feedback information" is information generated based on a language model that conveys to the user the meaning of the animal's behavior and voice.

[0281] "User emotion data" is data relating to the user's emotional state obtained by analyzing the user's voice and facial expressions.

[0282] "JSON format" is an abbreviation for JavaScript Object Notation, and is a format that represents data in a lightweight text-based format.

[0283] The "HTTPS protocol" is an abbreviation for HyperText Transfer Protocol Secure, and is a protocol for securely communicating data.

[0284] The present invention is a system that collects and analyzes information on animal movements, behavior, and voices, and provides feedback to the user based on that information. This system has the advantage of being able to more accurately understand animal behavior and also analyze the user's emotions in order to communicate with the user more efficiently. Specific embodiments for implementing this system are described below.

[0285] Hardware and Software Configuration

[0286] The device includes the following hardware:

[0287] 1. Motion sensor: Detects animal location, speed, and direction.

[0288] 2. Camera: Captures the animal's facial expressions and posture.

[0289] 3. Microphone: Records the frequency and volume of animal sounds.

[0290] 4. Emotion engine: Analyzes the user's voice and facial expressions to generate emotion data.

[0291] The server includes the following processing capabilities:

[0292] 1. Data preprocessing: Remove noise and missing data, and normalize numerical data.

[0293] 2. Machine learning algorithms: Used to analyze animal movements, behaviors, and sounds, as well as user emotional data.

[0294] 3. Generating and updating a language model: Based on the analysis results, the meaning of the animal's behavior and voice is translated and the model is updated.

[0295] Explanation of program processing

[0296] 1. Data collection

[0297] The device detects the animal's movements, behavior, and voice in real time, and also collects the user's facial expressions and voice. For example, a motion sensor detects the animal's location, a camera captures its posture, and a microphone records its calls. The device also analyzes the user's voice and facial expressions using the user's emotion engine to generate emotional data.

[0298] 2. Data transmission

[0299] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. Data encoded in JSON format represents each data item (e.g., animal location information, call frequency, user facial expression data, etc.) as a key-value pair.

[0300] 3. Data Preprocessing and Analysis

[0301] The server preprocesses the received data, removing noise and missing data and normalizing the numerical data. The preprocessed data is then fed into a machine learning algorithm to analyze the meaning of the animal's behavior and vocalizations. The analysis also includes the user's emotional data, allowing for a more accurate understanding of the animal's needs and emotions.

[0302] 4. Creating and updating language models and providing feedback

[0303] The server generates and updates a language model based on the analysis results, interprets the meaning of the animal's behavior and vocalizations, and generates feedback information for the user based on this. For example, it generates information such as, "If a dog stands on its hind legs and barks, it means it is asking for a treat."

[0304] 5. User Notices

[0305] The device receives the feedback information sent from the server and notifies the user's smartphone, allowing the user to take appropriate action. For example, the smartphone screen might display, "Your dog wants a treat."

[0306] Specific examples

[0307] For example, consider a case where a dog stands on its hind legs and barks "woof woof." The specific processing flow in this case is as follows:

[0308] Example prompt sentence:

[0309] "If a dog stands on its hind legs and barks 'woof woof', explain the processing flow to analyze the meaning of this behavior and provide appropriate feedback to the user."

[0310] In this way, it is possible to comprehensively analyze the animal's movements and the user's emotions and provide prompt and appropriate feedback, which will lead to smoother and deeper communication between the user and the animal.

[0311] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0312] Step 1: Collect data

[0313] The device uses motion sensors, cameras, microphones, and an emotion engine to collect animal movements, behaviors, and voices, as well as the user's facial expressions and voice, in real time. The inputs are the animal's location, posture, and calls, and the user's voice and facial expression data. These data are captured by the device and used for subsequent processing. The output is structured raw data.

[0314] Specifically, the motion sensor captures the animal's position and movement vector, the camera captures its posture as a still image, and the microphone records its cries as an audio file. At the same time, the emotion engine analyzes the user's voice and facial expressions to generate emotion data.

[0315] Step 2: Convert and send data

[0316] The terminal converts the collected data into JSON format. The input is structured raw data, and the data is encoded as key-value pairs based on this. The output is a JSON object. Specifically, each data item, such as collected location information, call frequency, and facial expression data, is converted into a JSON-formatted string.

[0317] The converted JSON data is securely sent to the server using the HTTPS protocol. Here, the input is JSON format data, and the output is safe and secure data transmission to the server.

[0318] Step 3: Preprocessing the data

[0319] The server preprocesses the received JSON data. The input is JSON-formatted data received from the device. Preprocessing removes noise and missing data, and normalizes numerical data to the same scale. For example, background noise is filtered from audio data, and motion data is scaled to the range 0 to 1. The output is a clean, normalized dataset.

[0320] Specifically, the data cleansing function removes noise and runs algorithms to fill in duplicate and missing data.

[0321] Step 4: Analyze the data

[0322] The server inputs the preprocessed data into a machine learning algorithm for analysis. The input is a clean, normalized dataset. This is used to analyze what the animal's movements and sounds mean and how they are affected by the user's emotional state. The output is the analysis results, such as determining that "standing up on one's hind legs" indicates "a request for something."

[0323] Specifically, the machine learning model analyzes the data and infers behavioral patterns and their meaning.

[0324] Step 5: Generate and update a language model

[0325] The server generates and updates a language model for translating the meaning of animal behavior and vocalizations based on the analysis results. The input is the analysis result data, and the output is an updated language model. Specifically, the server uses newly collected data to provide feedback to the existing language model, improving the model's accuracy.

[0326] For example, "standing up on hind legs" is translated as "wanting a treat" and incorporated into the model.

[0327] Step 6: Generate feedback information

[0328] The server generates feedback information for the user based on the updated language model. The input is the updated language model and the analysis results. The output is feedback information. Specifically, it is written in natural language in the form of "The dog is asking for something. For example, it would be good to give it a treat."

[0329] The user's emotional data is also taken into consideration here, and feedback such as "I know you seem busy, but please take a moment to respond" is also provided.

[0330] Step 7: User Notification

[0331] The device receives the feedback information sent from the server and notifies the user. The input is the feedback information, and the output is the feedback displayed or notified on the user's smartphone or audio device. Specifically, the message "The dog wants a treat" is displayed on the smartphone screen, or a voice message is notified.

[0332] This allows users to understand the animal's needs and emotions in real time and respond appropriately.

[0333] (Application example 2)

[0334] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0335] Conventional animal behavior analysis systems can analyze and translate animal behavior and voices, but they cannot take into account the user's emotions. Furthermore, they lack the information necessary for store staff to respond appropriately to customers with pets. This results in a decline in the quality of service for customers with pets, making it difficult to improve customer satisfaction.

[0336] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0337] In this invention, the server includes means for receiving and analyzing data on the animal's movements, behavior, and voice and data on the user's emotion transmitted from the terminal, means for updating a language model that translates the meaning of the animal's behavior and voice based on the analyzed data and the user's emotion data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for displaying the behavior and emotions of the pet and its owner to store staff in real time at physical stores where pets are allowed. This enables translation based on the pet's behavior and voice and feedback that takes into account the user's emotion, making it possible to provide more appropriate and higher quality service to customers with pets.

[0338] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[0339] The "server" is a device that receives and analyzes data on animal movements, actions, and voices and user emotion data transmitted from the terminal, updates the language model, and generates feedback information.

[0340] A "language model" is a data model used to translate animal behavior and vocalizations, and is updated based on analyzed data and user emotion data.

[0341] "User emotion data" is emotion information obtained by analyzing the user's facial expressions and voice.

[0342] "Feedback information" is information that conveys the meaning of an animal's behavior and voice to the user based on the analyzed data and updated language model.

[0343] A "pet-friendly brick-and-mortar store" is a physical store that caters to customers who bring their pets with them.

[0344] "Store associates" are employees who work in physical stores and provide services to customers in general and customers with pets.

[0345] "Real-time display" refers to providing and displaying the results of an analysis of the behavior and emotions of pets and their owners to store staff immediately.

[0346] This invention is a system that collects and analyzes animal movement, behavior, and voice data, as well as user emotion data, to generate feedback information in real time and improve interaction with customers who bring their pets to physical stores.

[0347] System Program

[0348] This system uses the following hardware and software:

[0349] Hardware used

[0350] 1. Smart glasses: A device equipped with a camera, microphone, and motion sensors that collects pet and user data in real time.

[0351] 2. Server: A device that analyzes the received data and generates feedback information.

[0352] Software used

[0353] 1. Analysis engine: Software containing machine learning algorithms for animal behavior analysis and user sentiment analysis.

[0354] 2. Emotion Engine: An AI module for analyzing user emotion data.

[0355] 3. Communication protocol: JSON format and HTTPS protocol are used to send and receive data.

[0356] Data collection and analysis

[0357] The smart glasses, which function as the device, detect animal movements and behaviors using a camera and motion sensors, record voices using a microphone, and collect the user's facial expressions and voice, then analyze the user's emotional data using an analysis engine.

[0358] The collected data is converted into JSON format and sent to the server using the HTTPS protocol. The server receives the received data, first preprocesses it (cleaning and normalizing it), and then applies machine learning algorithms to analyze animal behavior, vocalizations, and user emotions.

[0359] Generating feedback information

[0360] The server updates the animal-specific language model based on the analyzed data and the user's emotional data, which can more accurately translate the meaning of the animal's actions and voices based on the newly acquired data.

[0361] Based on the updated language model, feedback information is generated, including suggestions that take into account the meaning of the animal's actions and sounds, as well as the user's emotions.

[0362] The generated feedback information is displayed in real time on the smart glasses, for example, if the pet is calm, the glasses can provide suggestions to the store clerk such as "The pet is calm, please help the owner to take their time choosing products."

[0363] Specific examples

[0364] For example, if a pet is sitting, wagging its tail, and barking, the smart glasses will detect this and send the data to the server. The server will analyze the data and generate feedback such as, "The pet is excited, but the owner seems relaxed. Please suggest a pet toy." This will enable the store clerk to respond to the customer's needs quickly and appropriately.

[0365] Prompt Sentence Examples

[0366] Pet behavior data: "Sitting, wagging tail, barking"

[0367] User emotion data: "Relaxed"

[0368] In this way, the present invention provides an effective system that analyzes the behavior and emotions of pets and their owners in real time and supports customer service for pets in physical stores.

[0369] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0370] Step 1:

[0371] The device detects animal movements, behaviors, and sounds.

[0372] Input: Real-time animal movements, actions, and sounds, user facial expressions, and voice.

[0373] Processing: The camera built into the smart glasses captures the animal's posture, movements, and behavior, and the microphone records its voice. The analysis engine also captures the user's facial expressions and analyzes the voice with the emotion engine.

[0374] Output: Collected animal and user data (image data, audio data, emotion data).

[0375] Step 2:

[0376] The device sends the data to the server.

[0377] Input: Collected animal movement, behavior, and vocalization data, and user emotion data.

[0378] Processing: The device converts the collected data into JSON format and sends it securely to the server using the HTTPS protocol.

[0379] Output: Data sent in JSON format.

[0380] Step 3:

[0381] The server preprocesses the received data.

[0382] Input: Animal movement, behavior, and vocalization data, and user emotion data, sent in JSON format.

[0383] Processing: The server deserializes the received data, cleans it (removes noise and missing data) and normalizes it (aligns the numerical data to the same scale).

[0384] Output: Preprocessed and clean data.

[0385] Step 4:

[0386] The server analyzes the pre-processed data.

[0387] Input: Preprocessed and clean data.

[0388] Processing: The server uses an analysis engine and an emotion engine to analyze the animal's behavior and voice, as well as the user's emotion data. Specifically, it uses machine learning models to identify patterns in the animal's behavior and voice and infer the user's emotion.

[0389] Output: Analysis results (meaning of animal behavior and vocalizations, emotional state of the user).

[0390] Step 5:

[0391] The server updates the language model.

[0392] Input: Analysis results (meaning of animal behavior and vocalizations, user emotional state).

[0393] Processing: The server updates the animal-specific language model based on the analysis results. Newly collected data is reflected in the model, enabling more accurate translations.

[0394] Output: An updated language model.

[0395] Step 6:

[0396] The server generates the feedback information.

[0397] Input: The updated language model.

[0398] Processing: Based on the updated language model, the server generates feedback information that takes into account the meaning of the animal's actions and sounds, as well as the user's emotions. For example, it might say, "Your pet is making noise, so I suggest a toy."

[0399] Output: Feedback information.

[0400] Step 7:

[0401] The terminal notifies the user of the feedback information.

[0402] Input: Feedback information sent by the server.

[0403] Processing: The device displays feedback information on the smart glasses, such as "Your pet wants something. Please suggest a toy to its owner."

[0404] Output: The displayed feedback information to the user.

[0405] In this way, the system includes a flow that starts with data collection at the terminal, goes through data analysis and feedback information generation at the server, and finally provides real-time feedback to the user.

[0406] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0407] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0408] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0409] [Second embodiment]

[0410] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0411] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0412] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0413] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0414] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0415] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0416] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0417] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0418] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0419] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0420] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0421] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0422] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds, and a server that performs the analysis and translation.

[0423] About program processing

[0424] 1. Data collection

[0425] The device uses sensors and microphones to collect real-time information on the animal's movements, behavior, and voices. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphones record the frequency and volume of its calls. In addition, cameras and other devices are used to capture information on the animal's facial expressions and posture.

[0426] 2. Data transmission

[0427] The device converts the collected data into a certain format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS), which provides the data to the server in real time.

[0428] 3. Data Analysis

[0429] The server pre-processes the data it receives, then analyzes it using machine learning algorithms that detect specific patterns in animal behavior and vocalizations and infer what they mean.

[0430] 4. Creating and updating a language model

[0431] The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[0432] 5. Feedback Generation

[0433] The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[0434] 6. User Notices

[0435] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, it can also use a voice notification function to notify the user more clearly.

[0436] Specific examples

[0437] Below is a concrete example of a dog standing on its hind legs and barking "woof woof."

[0438] 1. The device uses a motion sensor to detect the dog standing up on its hind legs, and simultaneously records the dog's bark "woof woof" with a microphone.

[0439] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol.

[0440] 3. The server cleans and normalizes the data it receives and uses machine learning algorithms to determine that the "reaching sounds" mean "a request for something."

[0441] 4. The server updates the language model based on the results of the classification and interprets this behavior as meaning "wanting a snack."

[0442] 5. The server generates feedback information such as "The dog wants something. For example, it might be a good idea to give it a treat," and sends this information to the terminal.

[0443] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[0444] In this way, the system of the present invention effectively supports communication between the user and the pet, and enables quick and appropriate responses to the pet's requests and emotions.

[0445] The processing flow will be explained below.

[0446] Step 1:

[0447] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time: the motion sensors capture the animal's location, speed, and direction, the camera captures facial expressions and posture, and the microphone records the frequency and volume of the animal's calls.

[0448] Step 2:

[0449] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[0450] Step 3:

[0451] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS), in real time and optimized to minimize latency.

[0452] Step 4:

[0453] The server pre-processes the received data before analyzing it, which involves cleaning the data (removing noise and missing data) and normalizing it (aligning the numerical data to the same scale).

[0454] Step 5:

[0455] The server then feeds the pre-processed data into a machine learning algorithm that identifies patterns in the animal's behavior and vocalizations and uses these patterns to parse the animal's intentions and emotions using pre-trained models.

[0456] Step 6:

[0457] The server generates and updates a language model specifically for animals based on the analysis results, and insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation.

[0458] Step 7:

[0459] The server generates feedback information for the user based on the updated language model, written in natural language, to help the user understand the meaning of the animal's actions and sounds.

[0460] Step 8:

[0461] The server sends the generated feedback information to the terminal, allowing the terminal to obtain the latest information in real time.

[0462] Step 9:

[0463] The device then notifies the user of the received feedback information, which can be displayed on the smartphone screen or provided as a voice message using the voice notification function, such as "Your dog wants to play."

[0464] Step 10:

[0465] Based on the feedback information provided by the device, the user can respond appropriately to the pet's requests, such as giving treats or taking the pet out to play, thereby increasing the pet's satisfaction.

[0466] In this way, the system of the present invention goes through detailed processing steps, allowing the user to understand the intentions and emotions of their pet and take prompt and appropriate action.

[0467] Example 1

[0468] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0469] In modern times, communication between animals and humans remains difficult, particularly due to limited means for properly understanding an animal's intentions and emotions. This makes it difficult for pet owners to understand their pet's requests and emotions and respond quickly and appropriately. Furthermore, conventional systems have struggled to accurately analyze an animal's behavior and voice and provide real-time feedback. Therefore, there is a need for a system that can accurately analyze an animal's movements, behavior, and voice, and then translate the animal's intentions and emotions and notify the user.

[0470] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0471] In this invention, the server includes means for generating and updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and sending it to the terminal, and means for notifying the user of the feedback information to their smartphone. This enables the animal's movements, actions, and voices to be analyzed with high accuracy, allowing the user to quickly and appropriately understand the animal's intentions and emotions.

[0472] A "terminal" is a device equipped with a sensor and microphone that detects the movements, behavior, and sounds of animals, and transmits the data collected from these to a server.

[0473] A "sensor" is a device that detects physical changes and converts them into electrical signals, and in this invention it is used to sense the movement and location of animals.

[0474] A "microphone" is a device that converts sound into an electrical signal, and in this invention it is used to record the frequency and volume of animal cries.

[0475] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[0476] "Data" refers to information about animal movements, behaviors, and sounds collected by sensors and microphones.

[0477] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format for representing data in a structured manner.

[0478] "HTTPS" stands for Hypertext Transfer Protocol Secure, a protocol for encrypting data communications.

[0479] A "machine learning algorithm" is a computer algorithm used to analyze data and recognize patterns.

[0480] A "language model" is a collection of databases and algorithms that translate the meaning of animal behaviors and vocalizations based on analyzed data.

[0481] "Feedback information" is a message in natural language that explains the animal's intentions and emotions and is communicated to the user.

[0482] A "smartphone" is a portable information terminal equipped with mobile communication and computer functions.

[0483] The present invention provides a system that facilitates communication with animals by collecting, analyzing, and translating data on animal movements, behaviors, and voices, and notifying the user of the translation. The following describes in detail the embodiments of the present invention.

[0484] System Overview

[0485] The system consists of a device equipped with sensors and microphones that detect animal movements, behavior, and sounds, and a server that performs analysis and translation. Specifically, the device includes a motion sensor, microphone, and camera. These devices acquire information on the animal's location, speed, direction, frequency, and volume of its calls, as well as its facial expressions and posture. The collected data is converted into JSON format and securely sent to the server using the HTTPS protocol.

[0486] Hardware and Software

[0487] Device: Equipped with motion sensors, a microphone, and a camera to detect animal movements, behavior, and sounds. Includes software to convert this data into JSON format.

[0488] Server: Contains the processing unit and storage for parsing the received data, as well as software for data cleaning, normalization, and analyzing the data using machine learning algorithms, including algorithms for generating and updating language models.

[0489] Analyzing data and generating feedback

[0490] The server preprocesses the received data, for example by removing outliers and normalizing the data. The preprocessed data is then analyzed by machine learning algorithms. These algorithms use neural networks and deep learning models to recognize patterns in animal behavior and vocalizations and infer their meaning. For example, if a dog stands on its hind legs and barks "woof woof," it may determine that this behavior indicates "it wants a treat."

[0491] Based on the analysis results, the server generates and updates a language model specific to the animal. This language model is a collection of databases and algorithms for translating the meaning of the animal's actions and vocalizations. The server then generates feedback information for the user, which explains the animal's intentions and emotions to the user in natural language. For example, a message such as "The dog is asking for a treat" may be generated.

[0492] User Notifications

[0493] The device receives the feedback information sent from the server and notifies the user's smartphone. Notifications are displayed as push notifications, and audio notifications are also possible if necessary. For example, a notification such as "Your dog wants something. Please give it a treat" may appear on the smartphone.

[0494] Specific examples

[0495] A specific example of the system's operation is shown below.

[0496] When a dog stands on its hind legs and barks "woof woof"

[0497] 1. The device's motion sensor detects the dog standing up on its hind legs, and at the same time, the microphone records the dog's bark.

[0498] 2. The device converts the collected data into JSON format and sends it to the server via HTTPS protocol.

[0499] 3. The server cleans and normalizes the data, compares it with historical data, and analyzes it. Based on this analysis, it recognizes that the "standing up on hind legs" and "woof woof" indicate a request for a treat.

[0500] 4. The server updates the language model and generates feedback and sends it to the device.

[0501] 5. The device notifies the user of the feedback information on their smartphone.

[0502] Examples of prompts for generative AI models

[0503] "What does it mean when a dog stands on its hind legs and whines?"

[0504] "When a cat wags its tail and meows, what emotion does it express?"

[0505] "What does the behavior of a bird flapping its wings and chirping loudly indicate?"

[0506] Such prompts can be used to query the generative AI model about the meaning of an animal's behavior, allowing the system to provide appropriate feedback.

[0507] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0508] Step 1: Collect data

[0509] The device uses sensors, microphones, and cameras to collect animal movements, behaviors, and sounds in real time.

[0510] Input: Animal movements, behaviors, and sounds.

[0511] Specific operations: The device's built-in sensors detect the animal's position, speed, and direction, the microphone captures the frequency and volume of the animal's calls, and the camera captures images of its facial expressions and posture.

[0512] Output: Raw data from sensors, microphones, and cameras.

[0513] Step 2: Sending data

[0514] The terminal converts the collected data into JSON format and sends it to the server using the HTTPS protocol.

[0515] Input: Raw data from sensors, microphones, and cameras.

[0516] Specific operation: The terminal software converts the raw data into JSON format and transmits it securely via HTTPS.

[0517] Output: Data converted to JSON format.

[0518] Step 3: Preprocessing the data

[0519] The server cleans and normalizes the received data.

[0520] Input: JSON formatted data.

[0521] Specific operation: The server performs outlier removal, noise filtering, and data scaling.

[0522] Output: Cleaned and normalized data.

[0523] Step 4: Analyze the data

[0524] The server analyzes the data using machine learning algorithms.

[0525] Input: Cleaned and normalized data.

[0526] How it works: The server inputs data into a machine learning model (e.g., a neural network) to analyze the animal's behavior and vocal patterns.

[0527] Output: Analysis results of behavioral and vocal patterns.

[0528] Step 5: Generate and update a language model

[0529] Based on the analysis results, the server translates the meaning of the animal's behavior and voice, and generates and updates a language model specifically for the animal.

[0530] Input: Analysis results.

[0531] Specific operation: The server compares the analysis results with the existing language model and adds new behavioral patterns to the model.

[0532] Output: An updated language model.

[0533] Step 6: Generate feedback information

[0534] The server generates feedback information for the user.

[0535] Input: The updated language model.

[0536] Specific operation: The server converts the analysis results into natural language and generates a message that is easy for the user to understand.

[0537] Output: The generated feedback information.

[0538] Step 7: User Notification

[0539] The terminal notifies the user's smartphone of the feedback information from the server.

[0540] Input: The generated feedback information.

[0541] Specific operation: The device displays a push notification on the smartphone and, if necessary, plays a voice notification.

[0542] Output: The notification message that is displayed to the user.

[0543] This embodied processing allows the system to accurately analyze animal movements, behaviors, and sounds and provide information to the user.

[0544] (Application example 1)

[0545] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0546] While conventional pet communication systems can recognize pets' movements and voices, they are unable to recommend appropriate products and services to users based on that information. This makes it difficult for users to find products and services that meet their pets' needs and to establish effective communication with their pets. Virtual stores, in particular, lack product and service recommendations tailored to pets' specific needs, limiting the user experience.

[0547] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0548] In this invention, the server includes a terminal equipped with a sensor and a microphone that detects the movements, behavior, and voice of the animal, means for receiving and analyzing data on the movements, behavior, and voice of the animal transmitted from the terminal, means for updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for automatically suggesting related products and services in a virtual store based on the feedback information, thereby enabling the user to understand the meaning of their pet's actions and voices and easily find products and services in the virtual store that meet their pet's needs.

[0549] A "terminal" is a device equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[0550] The "server" is a system that receives and analyzes data on animal movements, behavior, and voices sent from the terminal.

[0551] A "language model" refers to the algorithms and datasets used to translate the meaning of animal behavior and vocalizations.

[0552] A "virtual store" is a virtual store operated on the Internet, which is a platform where real products can be purchased and services can be provided.

[0553] "Feedback information" refers to information provided to the user that is generated based on the meaning of the analyzed animal's behavior and voice.

[0554] "JSON format" is a data serialization format, an abbreviation for JavaScript Object Notation, and is a standard format for expressing structured data.

[0555] A "machine learning algorithm" is an algorithm that learns patterns and features from data and makes predictions and classifications for new data.

[0556] A "sensor" is a device that senses physical phenomena or conditions and collects that information as data.

[0557] A "microphone" is a device that converts sound waves into electrical signals and is used to collect animal sounds.

[0558] The "HTTPS protocol" is an abbreviation for Hypertext Transfer Protocol Secure, and is a communication protocol for securely sending and receiving data over the Internet.

[0559] "Animal-specific language models" are specialized algorithms and datasets for analyzing animal behavior and vocal patterns and translating their meaning.

[0560] This invention is a system that collects, analyzes, and translates animal movements, behaviors, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behaviors, and sounds, and a server that performs the analysis and translation.

[0561] System configuration

[0562] Hardware

[0563] Terminal: Equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[0564] Server: Carries out analysis and translation processing. Implements machine learning algorithms.

[0565] software

[0566] On the device:

[0567] A program that collects movements and voices in real time and converts the data into JSON format.

[0568] A communications program that sends data to a server using the HTTPS protocol.

[0569] Server side:

[0570] A program that receives and preprocesses data.

[0571] A program that uses machine learning algorithms to analyze data and translate its meaning.

[0572] A program that generates feedback information based on translation results and sends it to the terminal.

[0573] A program that automatically suggests related products and services in a virtual store based on feedback information.

[0574] Operation flow

[0575] 1. The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphone records the frequency and volume of its calls. Cameras and other devices also capture information on the animal's facial expressions and posture.

[0576] 2. The device converts the collected data into JSON format and sends it to the server using a secure protocol (HTTPS), which provides the data to the server in real time.

[0577] 3. The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning and normalizing the data. The machine learning algorithms detect specific patterns in the animals' behavior and vocalizations and infer what they mean.

[0578] 4. The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[0579] 5. The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[0580] 6. The device receives the feedback information sent from the server and notifies the user. For example, the device can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, a voice notification function can be used to notify the user more clearly.

[0581] 7. The server automatically suggests related products and services in the virtual store based on the feedback information. For example, based on the translation result that the dog "wants to play," the server suggests a new toy in the virtual store.

[0582] Specific examples

[0583] 1. The user uses a smartphone to observe the dog's behavior in real time.

[0584] 2. Your pet wags its tail and barks "woof woof."

[0585] 3. The app will notify you, "Your dog seems happy and might want to play with his new toy."

[0586] 4. Virtual stores suggest toys for pets.

[0587] Prompt Sentence Examples

[0588] Input: The smartphone detects the dog's tail wagging and the sound of it barking "woof woof."

[0589] Processing: Data is sent to the server and analyzed by the translation system. Feedback information is generated based on the translation results and sent to the smartphone.

[0590] Output: Your dog seems happy and may be eager to play with his new toy.

[0591] This configuration makes it possible to construct a system within the scope of the invention and realize effective communication between pets and users.

[0592] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0593] Step 1:

[0594] The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. Inputs include location information, speed, and direction of movement from the sensors, as well as the frequency and volume of sounds from the microphones. This data is output as information on the animal's behavior and sounds detected through the sensors and microphones.

[0595] Step 2:

[0596] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. The input is the behavioral and voice data obtained in step 1. These data are packaged in JSON format and sent to the server via a secure communication channel for output.

[0597] Step 3:

[0598] The server preprocesses the data it receives. It takes as input the JSON data sent in step 2. It cleans and normalizes the data, resulting in clean data that can be meaningfully analyzed.

[0599] Step 4:

[0600] The server analyzes the data using machine learning algorithms. The input is the clean, pre-processed data. The machine learning algorithms detect specific patterns in the animal's behavior and vocalizations and infer their meaning. The output is the meaning of the animal's behavior and vocalizations.

[0601] Step 5:

[0602] The server translates the meaning of the animal's actions and voices based on the analysis results and updates the animal-specific language model. The input is the analysis results from step 4. Based on this data, the server generates translations corresponding to the animal's actions and voices and updates the language model. The output is the latest language model.

[0603] Step 6:

[0604] The server generates feedback information for the user based on the updated language model. The inputs are the latest language model and data on the animal's behavior and voice. The feedback information is formed as a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand. The output is the feedback information.

[0605] Step 7:

[0606] The device receives feedback information sent from the server and notifies the user. The input is the feedback information sent from the server. This is displayed on the screen of the user's smartphone or smart glasses, providing feedback to the user. The output is a notification sent to the user.

[0607] Step 8:

[0608] The server automatically suggests related products and services in the virtual store based on the feedback information. The input is the feedback information. For example, based on the feedback information "Your dog wants to play," the virtual store suggests a new toy. The output is information about the suggested products and services.

[0609] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0610] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and voices to facilitate communication with users, and provides more efficient feedback by also analyzing the user's emotions. This system is composed of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and voices, a server that performs analysis and translation, and an emotion engine that recognizes the user's emotions.

[0611] About program processing

[0612] 1. Data collection

[0613] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The sensors capture the animal's location, speed, and direction, the camera captures its facial expressions and posture, and the microphone records the frequency and volume of its calls. At the same time, the user's voice and facial expressions are also collected using an emotion engine.

[0614] 2. Data transmission

[0615] The device converts the collected data into a specific format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS). This allows the data to be provided to the server in real time, along with the user's emotional data.

[0616] 3. Data Analysis

[0617] The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning the data (removing noise and missing data) and normalizing it (aligning numerical data to the same scale). Not only animal behavior and voices, but also user emotion data are analyzed.

[0618] 4. Creating and updating a language model

[0619] The server uses the analysis results to translate the meaning of the animal's behavior and voice, and generates and updates a language model specifically for animals. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. Additionally, the user's emotional data is also reflected in the language model.

[0620] 5. Feedback Generation

[0621] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[0622] 6. User Notices

[0623] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play. You seem busy, but let's take a moment to play with him." If necessary, it can also be provided as a voice message using the voice notification function.

[0624] Specific examples

[0625] Below is a specific example of a dog standing on its hind legs and barking "woof woof."

[0626] 1. The device uses a motion sensor to detect when the dog stands up on its hind legs, and simultaneously records the dog's bark with a microphone. The camera also captures the user's facial expressions, and the audio is analyzed by an emotion engine.

[0627] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol, along with the user's emotional data.

[0628] 3. The server cleans and normalizes the received data and uses machine learning algorithms to determine whether the "reaching on its hind legs" is a "request for something." User sentiment data is also analyzed.

[0629] 4. The server updates the language model based on the results of the classification and translates this behavior as meaning "wanting a snack." The user's emotional data is also reflected in the language model.

[0630] 5. The server generates feedback information such as "The dog wants something. For example, it would be good to give it a treat. It seems you are a little busy, but please make time to respond," and sends this information to the terminal.

[0631] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[0632] In this way, the system of the present invention effectively supports communication between the user and the pet, enabling the system to respond quickly and appropriately to the pet's needs and emotions. Furthermore, by taking the user's emotions into consideration, the system provides more natural and efficient feedback.

[0633] The processing flow will be explained below.

[0634] Step 1:

[0635] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The motion sensor captures the animal's location, speed, and direction of movement. The camera captures the animal's facial expressions and posture, and the microphone records the frequency and volume of its calls. The device also collects the user's voice and facial expressions using an emotion engine.

[0636] Step 2:

[0637] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[0638] Step 3:

[0639] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS) in real time and optimized to minimize latency, along with the user's emotional data.

[0640] Step 4:

[0641] The server preprocesses the received data, which includes removing noise and missing data and normalizing the data (aligning the numerical data to the same scale), thereby improving the accuracy of the analysis.

[0642] Step 5:

[0643] The server then inputs the pre-processed data into a machine learning algorithm, which identifies patterns in the animal's behavior and vocalizations and uses these patterns to analyze the animal's intentions and emotions, while simultaneously analyzing the user's emotional data.

[0644] Step 6:

[0645] The server generates and updates a language model specifically for animals based on the analysis results. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. User emotion data is also reflected in the language model.

[0646] Step 7:

[0647] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[0648] Step 8:

[0649] The server sends the generated feedback information to the device, allowing the device to obtain the latest information in real time. For example, it generates a message such as, "Your dog wants to play. It seems you're a little busy, so please take a moment to play with him."

[0650] Step 9:

[0651] The device notifies the user of the received feedback information. This notification is displayed on the smartphone screen and also provided as a voice message using the voice notification function, allowing the user to receive feedback both visually and audibly.

[0652] Step 10:

[0653] The user can respond appropriately to their pet's needs based on the feedback information provided by the device, for example, by giving them treats or taking them outside to increase their pet's satisfaction. Further data is collected based on the user's actions, improving the accuracy of the system.

[0654] Example 2

[0655] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0656] Conventional animal-human communication systems simply analyze the animal's behavior and voice, and are unable to provide appropriate feedback that takes into account the user's emotions. This makes it difficult to accurately understand the animal's needs and emotions and respond quickly and appropriately, making it difficult to build a trusting relationship between the user and the animal.

[0657] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing data on animal movements, behaviors, and voices and user emotion data, means for analyzing the preprocessed data using a machine learning algorithm, means for generating and updating a language model based on the analysis results, means for generating feedback information for the user and transmitting it to the terminal, and means for notifying the user of the feedback information. This makes it possible to accurately understand the requests and emotions of the animal and provide quick and accurate feedback taking the user's emotions into consideration.

[0658] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[0659] The "server" is a central processing unit that receives and analyzes data sent from the terminal, generates feedback information, and sends it to the terminal.

[0660] "Data preprocessing" refers to the process of removing noise and missing data from the received data and aligning the numerical data to the same scale.

[0661] A "machine learning algorithm" is a mathematical method that analyzes received and preprocessed data to infer the meaning of animal behavior and vocalizations.

[0662] A "language model" is a model for translating the meaning of animal behavior and voices based on analysis results, and is an updatable data structure.

[0663] "Feedback information" is information generated based on a language model that conveys to the user the meaning of the animal's behavior and voice.

[0664] "User emotion data" is data relating to the user's emotional state obtained by analyzing the user's voice and facial expressions.

[0665] "JSON format" is an abbreviation for JavaScript Object Notation, and is a format that represents data in a lightweight text-based format.

[0666] The "HTTPS protocol" is an abbreviation for HyperText Transfer Protocol Secure, and is a protocol for securely communicating data.

[0667] The present invention is a system that collects and analyzes information on animal movements, behavior, and voices, and provides feedback to the user based on that information. This system has the advantage of being able to more accurately understand animal behavior and also analyze the user's emotions in order to communicate with the user more efficiently. Specific embodiments for implementing this system are described below.

[0668] Hardware and Software Configuration

[0669] The device includes the following hardware:

[0670] 1. Motion sensor: Detects animal location, speed, and direction.

[0671] 2. Camera: Captures the animal's facial expressions and posture.

[0672] 3. Microphone: Records the frequency and volume of animal sounds.

[0673] 4. Emotion engine: Analyzes the user's voice and facial expressions to generate emotion data.

[0674] The server includes the following processing capabilities:

[0675] 1. Data preprocessing: Remove noise and missing data, and normalize numerical data.

[0676] 2. Machine learning algorithms: Used to analyze animal movements, behaviors, and sounds, as well as user emotional data.

[0677] 3. Generating and updating a language model: Based on the analysis results, the meaning of the animal's behavior and voice is translated and the model is updated.

[0678] Explanation of program processing

[0679] 1. Data collection

[0680] The device detects the animal's movements, behavior, and voice in real time, and also collects the user's facial expressions and voice. For example, a motion sensor detects the animal's location, a camera captures its posture, and a microphone records its calls. The device also analyzes the user's voice and facial expressions using the user's emotion engine to generate emotional data.

[0681] 2. Data transmission

[0682] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. Data encoded in JSON format represents each data item (e.g., animal location information, call frequency, user facial expression data, etc.) as a key-value pair.

[0683] 3. Data Preprocessing and Analysis

[0684] The server preprocesses the received data, removing noise and missing data and normalizing the numerical data. The preprocessed data is then fed into a machine learning algorithm to analyze the meaning of the animal's behavior and vocalizations. The analysis also includes the user's emotional data, allowing for a more accurate understanding of the animal's needs and emotions.

[0685] 4. Creating and updating language models and providing feedback

[0686] The server generates and updates a language model based on the analysis results, interprets the meaning of the animal's behavior and vocalizations, and generates feedback information for the user based on this. For example, it generates information such as, "If a dog stands on its hind legs and barks, it means it is asking for a treat."

[0687] 5. User Notices

[0688] The device receives the feedback information sent from the server and notifies the user's smartphone, allowing the user to take appropriate action. For example, the smartphone screen might display, "Your dog wants a treat."

[0689] Specific examples

[0690] For example, consider a case where a dog stands on its hind legs and barks "woof woof." The specific processing flow in this case is as follows:

[0691] Example prompt sentence:

[0692] "If a dog stands on its hind legs and barks 'woof woof', explain the processing flow to analyze the meaning of this behavior and provide appropriate feedback to the user."

[0693] In this way, it is possible to comprehensively analyze the animal's movements and the user's emotions and provide prompt and appropriate feedback, which will lead to smoother and deeper communication between the user and the animal.

[0694] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0695] Step 1: Collect data

[0696] The device uses motion sensors, cameras, microphones, and an emotion engine to collect animal movements, behaviors, and voices, as well as the user's facial expressions and voice, in real time. The inputs are the animal's location, posture, and calls, and the user's voice and facial expression data. These data are captured by the device and used for subsequent processing. The output is structured raw data.

[0697] Specifically, the motion sensor captures the animal's position and movement vector, the camera captures its posture as a still image, and the microphone records its cries as an audio file. At the same time, the emotion engine analyzes the user's voice and facial expressions to generate emotion data.

[0698] Step 2: Convert and send data

[0699] The terminal converts the collected data into JSON format. The input is structured raw data, and the data is encoded as key-value pairs based on this. The output is a JSON object. Specifically, each data item, such as collected location information, call frequency, and facial expression data, is converted into a JSON-formatted string.

[0700] The converted JSON data is securely sent to the server using the HTTPS protocol. Here, the input is JSON format data, and the output is safe and secure data transmission to the server.

[0701] Step 3: Preprocessing the data

[0702] The server preprocesses the received JSON data. The input is JSON-formatted data received from the device. Preprocessing removes noise and missing data, and normalizes numerical data to the same scale. For example, background noise is filtered from audio data, and motion data is scaled to the range 0 to 1. The output is a clean, normalized dataset.

[0703] Specifically, the data cleansing function removes noise and runs algorithms to fill in duplicate and missing data.

[0704] Step 4: Analyze the data

[0705] The server inputs the preprocessed data into a machine learning algorithm for analysis. The input is a clean, normalized dataset. This is used to analyze what the animal's movements and sounds mean and how they are affected by the user's emotional state. The output is the analysis results, such as determining that "standing up on one's hind legs" indicates "a request for something."

[0706] Specifically, the machine learning model analyzes the data and infers behavioral patterns and their meaning.

[0707] Step 5: Generate and update a language model

[0708] The server generates and updates a language model for translating the meaning of animal behavior and vocalizations based on the analysis results. The input is the analysis result data, and the output is an updated language model. Specifically, the server uses newly collected data to provide feedback to the existing language model, improving the model's accuracy.

[0709] For example, "standing up on hind legs" is translated as "wanting a treat" and incorporated into the model.

[0710] Step 6: Generate feedback information

[0711] The server generates feedback information for the user based on the updated language model. The input is the updated language model and the analysis results. The output is feedback information. Specifically, it is written in natural language in the form of "The dog is asking for something. For example, it would be good to give it a treat."

[0712] The user's emotional data is also taken into consideration here, and feedback such as "I know you seem busy, but please take a moment to respond" is also provided.

[0713] Step 7: User Notification

[0714] The device receives the feedback information sent from the server and notifies the user. The input is the feedback information, and the output is the feedback displayed or notified on the user's smartphone or audio device. Specifically, the message "The dog wants a treat" is displayed on the smartphone screen, or a voice message is notified.

[0715] This allows users to understand the animal's needs and emotions in real time and respond appropriately.

[0716] (Application example 2)

[0717] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0718] Conventional animal behavior analysis systems can analyze and translate animal behavior and voices, but they cannot take into account the user's emotions. Furthermore, they lack the information necessary for store staff to respond appropriately to customers with pets. This results in a decline in the quality of service for customers with pets, making it difficult to improve customer satisfaction.

[0719] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0720] In this invention, the server includes means for receiving and analyzing data on the animal's movements, behavior, and voice and data on the user's emotion transmitted from the terminal, means for updating a language model that translates the meaning of the animal's behavior and voice based on the analyzed data and the user's emotion data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for displaying the behavior and emotions of the pet and its owner to store staff in real time at physical stores where pets are allowed. This enables translation based on the pet's behavior and voice and feedback that takes into account the user's emotion, making it possible to provide more appropriate and higher quality service to customers with pets.

[0721] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[0722] The "server" is a device that receives and analyzes data on animal movements, actions, and voices and user emotion data transmitted from the terminal, updates the language model, and generates feedback information.

[0723] A "language model" is a data model used to translate animal behavior and vocalizations, and is updated based on analyzed data and user emotion data.

[0724] "User emotion data" is emotion information obtained by analyzing the user's facial expressions and voice.

[0725] "Feedback information" is information that conveys the meaning of an animal's behavior and voice to the user based on the analyzed data and updated language model.

[0726] A "pet-friendly brick-and-mortar store" is a physical store that caters to customers who bring their pets with them.

[0727] "Store associates" are employees who work in physical stores and provide services to customers in general and customers with pets.

[0728] "Real-time display" refers to providing and displaying the results of an analysis of the behavior and emotions of pets and their owners to store staff immediately.

[0729] This invention is a system that collects and analyzes animal movement, behavior, and voice data, as well as user emotion data, to generate feedback information in real time and improve interaction with customers who bring their pets to physical stores.

[0730] System Program

[0731] This system uses the following hardware and software:

[0732] Hardware used

[0733] 1. Smart glasses: A device equipped with a camera, microphone, and motion sensors that collects pet and user data in real time.

[0734] 2. Server: A device that analyzes the received data and generates feedback information.

[0735] Software used

[0736] 1. Analysis engine: Software containing machine learning algorithms for animal behavior analysis and user sentiment analysis.

[0737] 2. Emotion Engine: An AI module for analyzing user emotion data.

[0738] 3. Communication protocol: JSON format and HTTPS protocol are used to send and receive data.

[0739] Data collection and analysis

[0740] The smart glasses, which function as the device, detect animal movements and behaviors using a camera and motion sensors, record voices using a microphone, and collect the user's facial expressions and voice, then analyze the user's emotional data using an analysis engine.

[0741] The collected data is converted into JSON format and sent to the server using the HTTPS protocol. The server receives the received data, first preprocesses it (cleaning and normalizing it), and then applies machine learning algorithms to analyze animal behavior, vocalizations, and user emotions.

[0742] Generating feedback information

[0743] The server updates the animal-specific language model based on the analyzed data and the user's emotional data, which can more accurately translate the meaning of the animal's actions and voices based on the newly acquired data.

[0744] Based on the updated language model, feedback information is generated, including suggestions that take into account the meaning of the animal's actions and sounds, as well as the user's emotions.

[0745] The generated feedback information is displayed in real time on the smart glasses, for example, if the pet is calm, the glasses can provide suggestions to the store clerk such as "The pet is calm, please help the owner to take their time choosing products."

[0746] Specific examples

[0747] For example, if a pet is sitting, wagging its tail, and barking, the smart glasses will detect this and send the data to the server. The server will analyze the data and generate feedback such as, "The pet is excited, but the owner seems relaxed. Please suggest a pet toy." This will enable the store clerk to respond to the customer's needs quickly and appropriately.

[0748] Prompt Sentence Examples

[0749] Pet behavior data: "Sitting, wagging tail, barking"

[0750] User emotion data: "Relaxed"

[0751] In this way, the present invention provides an effective system that analyzes the behavior and emotions of pets and their owners in real time and supports customer service for pets in physical stores.

[0752] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0753] Step 1:

[0754] The device detects animal movements, behaviors, and sounds.

[0755] Input: Real-time animal movements, actions, and sounds, user facial expressions, and voice.

[0756] Processing: The camera built into the smart glasses captures the animal's posture, movements, and behavior, and the microphone records its voice. The analysis engine also captures the user's facial expressions and analyzes the voice with the emotion engine.

[0757] Output: Collected animal and user data (image data, audio data, emotion data).

[0758] Step 2:

[0759] The device sends the data to the server.

[0760] Input: Collected animal movement, behavior, and vocalization data, and user emotion data.

[0761] Processing: The device converts the collected data into JSON format and sends it securely to the server using the HTTPS protocol.

[0762] Output: Data sent in JSON format.

[0763] Step 3:

[0764] The server preprocesses the received data.

[0765] Input: Animal movement, behavior, and vocalization data, and user emotion data, sent in JSON format.

[0766] Processing: The server deserializes the received data, cleans it (removes noise and missing data) and normalizes it (aligns the numerical data to the same scale).

[0767] Output: Preprocessed and clean data.

[0768] Step 4:

[0769] The server analyzes the pre-processed data.

[0770] Input: Preprocessed and clean data.

[0771] Processing: The server uses an analysis engine and an emotion engine to analyze the animal's behavior and voice, as well as the user's emotion data. Specifically, it uses machine learning models to identify patterns in the animal's behavior and voice and infer the user's emotion.

[0772] Output: Analysis results (meaning of animal behavior and vocalizations, emotional state of the user).

[0773] Step 5:

[0774] The server updates the language model.

[0775] Input: Analysis results (meaning of animal behavior and vocalizations, user emotional state).

[0776] Processing: The server updates the animal-specific language model based on the analysis results. Newly collected data is reflected in the model, enabling more accurate translations.

[0777] Output: An updated language model.

[0778] Step 6:

[0779] The server generates the feedback information.

[0780] Input: The updated language model.

[0781] Processing: Based on the updated language model, the server generates feedback information that takes into account the meaning of the animal's actions and sounds, as well as the user's emotions. For example, it might say, "Your pet is making noise, so I suggest a toy."

[0782] Output: Feedback information.

[0783] Step 7:

[0784] The terminal notifies the user of the feedback information.

[0785] Input: Feedback information sent by the server.

[0786] Processing: The device displays feedback information on the smart glasses, such as "Your pet wants something. Please suggest a toy to its owner."

[0787] Output: The displayed feedback information to the user.

[0788] In this way, the system includes a flow that starts with data collection at the terminal, goes through data analysis and feedback information generation at the server, and finally provides real-time feedback to the user.

[0789] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0790] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0791] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0792] [Third embodiment]

[0793] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0794] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0795] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0796] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0797] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0798] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0799] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0800] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0801] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0802] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0803] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0804] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0805] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds, and a server that performs the analysis and translation.

[0806] About program processing

[0807] 1. Data collection

[0808] The device uses sensors and microphones to collect real-time information on the animal's movements, behavior, and voices. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphones record the frequency and volume of its calls. In addition, cameras and other devices are used to capture information on the animal's facial expressions and posture.

[0809] 2. Data transmission

[0810] The device converts the collected data into a certain format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS), which provides the data to the server in real time.

[0811] 3. Data Analysis

[0812] The server pre-processes the data it receives, then analyzes it using machine learning algorithms that detect specific patterns in animal behavior and vocalizations and infer what they mean.

[0813] 4. Creating and updating a language model

[0814] The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[0815] 5. Feedback Generation

[0816] The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[0817] 6. User Notices

[0818] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, it can also use a voice notification function to notify the user more clearly.

[0819] Specific examples

[0820] Below is a concrete example of a dog standing on its hind legs and barking "woof woof."

[0821] 1. The device uses a motion sensor to detect the dog standing up on its hind legs, and simultaneously records the dog's bark "woof woof" with a microphone.

[0822] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol.

[0823] 3. The server cleans and normalizes the data it receives and uses machine learning algorithms to determine that the "reaching sounds" mean "a request for something."

[0824] 4. The server updates the language model based on the results of the classification and interprets this behavior as meaning "wanting a snack."

[0825] 5. The server generates feedback information such as "The dog wants something. For example, it might be a good idea to give it a treat," and sends this information to the terminal.

[0826] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[0827] In this way, the system of the present invention effectively supports communication between the user and the pet, and enables quick and appropriate responses to the pet's requests and emotions.

[0828] The processing flow will be explained below.

[0829] Step 1:

[0830] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time: the motion sensors capture the animal's location, speed, and direction, the camera captures facial expressions and posture, and the microphone records the frequency and volume of the animal's calls.

[0831] Step 2:

[0832] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[0833] Step 3:

[0834] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS), in real time and optimized to minimize latency.

[0835] Step 4:

[0836] The server pre-processes the received data before analyzing it, which involves cleaning the data (removing noise and missing data) and normalizing it (aligning the numerical data to the same scale).

[0837] Step 5:

[0838] The server then feeds the pre-processed data into a machine learning algorithm that identifies patterns in the animal's behavior and vocalizations and uses these patterns to parse the animal's intentions and emotions using pre-trained models.

[0839] Step 6:

[0840] The server generates and updates a language model specifically for animals based on the analysis results, and insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation.

[0841] Step 7:

[0842] The server generates feedback information for the user based on the updated language model, written in natural language, to help the user understand the meaning of the animal's actions and sounds.

[0843] Step 8:

[0844] The server sends the generated feedback information to the terminal, allowing the terminal to obtain the latest information in real time.

[0845] Step 9:

[0846] The device then notifies the user of the received feedback information, which can be displayed on the smartphone screen or provided as a voice message using the voice notification function, such as "Your dog wants to play."

[0847] Step 10:

[0848] Based on the feedback information provided by the device, the user can respond appropriately to the pet's requests, such as giving treats or taking the pet out to play, thereby increasing the pet's satisfaction.

[0849] In this way, the system of the present invention goes through detailed processing steps, allowing the user to understand the intentions and emotions of their pet and take prompt and appropriate action.

[0850] Example 1

[0851] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0852] In modern times, communication between animals and humans remains difficult, particularly due to limited means for properly understanding an animal's intentions and emotions. This makes it difficult for pet owners to understand their pet's requests and emotions and respond quickly and appropriately. Furthermore, conventional systems have struggled to accurately analyze an animal's behavior and voice and provide real-time feedback. Therefore, there is a need for a system that can accurately analyze an animal's movements, behavior, and voice, and then translate the animal's intentions and emotions and notify the user.

[0853] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0854] In this invention, the server includes means for generating and updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and sending it to the terminal, and means for notifying the user of the feedback information to their smartphone. This enables the animal's movements, actions, and voices to be analyzed with high accuracy, allowing the user to quickly and appropriately understand the animal's intentions and emotions.

[0855] A "terminal" is a device equipped with a sensor and microphone that detects the movements, behavior, and sounds of animals, and transmits the data collected from these to a server.

[0856] A "sensor" is a device that detects physical changes and converts them into electrical signals, and in this invention it is used to sense the movement and location of animals.

[0857] A "microphone" is a device that converts sound into an electrical signal, and in this invention it is used to record the frequency and volume of animal cries.

[0858] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[0859] "Data" refers to information about animal movements, behaviors, and sounds collected by sensors and microphones.

[0860] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format for representing data in a structured manner.

[0861] "HTTPS" stands for Hypertext Transfer Protocol Secure, a protocol for encrypting data communications.

[0862] A "machine learning algorithm" is a computer algorithm used to analyze data and recognize patterns.

[0863] A "language model" is a collection of databases and algorithms that translate the meaning of animal behaviors and vocalizations based on analyzed data.

[0864] "Feedback information" is a message in natural language that explains the animal's intentions and emotions and is communicated to the user.

[0865] A "smartphone" is a portable information terminal equipped with mobile communication and computer functions.

[0866] The present invention provides a system that facilitates communication with animals by collecting, analyzing, and translating data on animal movements, behaviors, and voices, and notifying the user of the translation. The following describes in detail the embodiments of the present invention.

[0867] System Overview

[0868] The system consists of a device equipped with sensors and microphones that detect animal movements, behavior, and sounds, and a server that performs analysis and translation. Specifically, the device includes a motion sensor, microphone, and camera. These devices acquire information on the animal's location, speed, direction, frequency, and volume of its calls, as well as its facial expressions and posture. The collected data is converted into JSON format and securely sent to the server using the HTTPS protocol.

[0869] Hardware and Software

[0870] Device: Equipped with motion sensors, a microphone, and a camera to detect animal movements, behavior, and sounds. Includes software to convert this data into JSON format.

[0871] Server: Contains the processing unit and storage for parsing the received data, as well as software for data cleaning, normalization, and analyzing the data using machine learning algorithms, including algorithms for generating and updating language models.

[0872] Analyzing data and generating feedback

[0873] The server preprocesses the received data, for example by removing outliers and normalizing the data. The preprocessed data is then analyzed by machine learning algorithms. These algorithms use neural networks and deep learning models to recognize patterns in animal behavior and vocalizations and infer their meaning. For example, if a dog stands on its hind legs and barks "woof woof," it may determine that this behavior indicates "it wants a treat."

[0874] Based on the analysis results, the server generates and updates a language model specific to the animal. This language model is a collection of databases and algorithms for translating the meaning of the animal's actions and vocalizations. The server then generates feedback information for the user, which explains the animal's intentions and emotions to the user in natural language. For example, a message such as "The dog is asking for a treat" may be generated.

[0875] User Notifications

[0876] The device receives the feedback information sent from the server and notifies the user's smartphone. Notifications are displayed as push notifications, and audio notifications are also possible if necessary. For example, a notification such as "Your dog wants something. Please give it a treat" may appear on the smartphone.

[0877] Specific examples

[0878] A specific example of the system's operation is shown below.

[0879] When a dog stands on its hind legs and barks "woof woof"

[0880] 1. The device's motion sensor detects the dog standing up on its hind legs, and at the same time, the microphone records the dog's bark.

[0881] 2. The device converts the collected data into JSON format and sends it to the server via HTTPS protocol.

[0882] 3. The server cleans and normalizes the data, compares it with historical data, and analyzes it. Based on this analysis, it recognizes that the "standing up on hind legs" and "woof woof" indicate a request for a treat.

[0883] 4. The server updates the language model and generates feedback and sends it to the device.

[0884] 5. The device notifies the user of the feedback information on their smartphone.

[0885] Examples of prompts for generative AI models

[0886] "What does it mean when a dog stands on its hind legs and whines?"

[0887] "When a cat wags its tail and meows, what emotion does it express?"

[0888] "What does the behavior of a bird flapping its wings and chirping loudly indicate?"

[0889] Such prompts can be used to query the generative AI model about the meaning of an animal's behavior, allowing the system to provide appropriate feedback.

[0890] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0891] Step 1: Collect data

[0892] The device uses sensors, microphones, and cameras to collect animal movements, behaviors, and sounds in real time.

[0893] Input: Animal movements, behaviors, and sounds.

[0894] Specific operations: The device's built-in sensors detect the animal's position, speed, and direction, the microphone captures the frequency and volume of the animal's calls, and the camera captures images of its facial expressions and posture.

[0895] Output: Raw data from sensors, microphones, and cameras.

[0896] Step 2: Sending data

[0897] The terminal converts the collected data into JSON format and sends it to the server using the HTTPS protocol.

[0898] Input: Raw data from sensors, microphones, and cameras.

[0899] Specific operation: The terminal software converts the raw data into JSON format and transmits it securely via HTTPS.

[0900] Output: Data converted to JSON format.

[0901] Step 3: Preprocessing the data

[0902] The server cleans and normalizes the received data.

[0903] Input: JSON formatted data.

[0904] Specific operation: The server performs outlier removal, noise filtering, and data scaling.

[0905] Output: Cleaned and normalized data.

[0906] Step 4: Analyze the data

[0907] The server analyzes the data using machine learning algorithms.

[0908] Input: Cleaned and normalized data.

[0909] How it works: The server inputs data into a machine learning model (e.g., a neural network) to analyze the animal's behavior and vocal patterns.

[0910] Output: Analysis results of behavioral and vocal patterns.

[0911] Step 5: Generate and update a language model

[0912] Based on the analysis results, the server translates the meaning of the animal's behavior and voice, and generates and updates a language model specifically for the animal.

[0913] Input: Analysis results.

[0914] Specific operation: The server compares the analysis results with the existing language model and adds new behavioral patterns to the model.

[0915] Output: An updated language model.

[0916] Step 6: Generate feedback information

[0917] The server generates feedback information for the user.

[0918] Input: The updated language model.

[0919] Specific operation: The server converts the analysis results into natural language and generates a message that is easy for the user to understand.

[0920] Output: The generated feedback information.

[0921] Step 7: User Notification

[0922] The terminal notifies the user's smartphone of the feedback information from the server.

[0923] Input: The generated feedback information.

[0924] Specific operation: The device displays a push notification on the smartphone and, if necessary, plays a voice notification.

[0925] Output: The notification message that is displayed to the user.

[0926] This embodied processing allows the system to accurately analyze animal movements, behaviors, and sounds and provide information to the user.

[0927] (Application example 1)

[0928] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0929] While conventional pet communication systems can recognize pets' movements and voices, they are unable to recommend appropriate products and services to users based on that information. This makes it difficult for users to find products and services that meet their pets' needs and to establish effective communication with their pets. Virtual stores, in particular, lack product and service recommendations tailored to pets' specific needs, limiting the user experience.

[0930] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0931] In this invention, the server includes a terminal equipped with a sensor and a microphone that detects the movements, behavior, and voice of the animal, means for receiving and analyzing data on the movements, behavior, and voice of the animal transmitted from the terminal, means for updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for automatically suggesting related products and services in a virtual store based on the feedback information, thereby enabling the user to understand the meaning of their pet's actions and voices and easily find products and services in the virtual store that meet their pet's needs.

[0932] A "terminal" is a device equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[0933] The "server" is a system that receives and analyzes data on animal movements, behavior, and voices sent from the terminal.

[0934] A "language model" refers to the algorithms and datasets used to translate the meaning of animal behavior and vocalizations.

[0935] A "virtual store" is a virtual store operated on the Internet, which is a platform where real products can be purchased and services can be provided.

[0936] "Feedback information" refers to information provided to the user that is generated based on the meaning of the analyzed animal's behavior and voice.

[0937] "JSON format" is a data serialization format, an abbreviation for JavaScript Object Notation, and is a standard format for expressing structured data.

[0938] A "machine learning algorithm" is an algorithm that learns patterns and features from data and makes predictions and classifications for new data.

[0939] A "sensor" is a device that senses physical phenomena or conditions and collects that information as data.

[0940] A "microphone" is a device that converts sound waves into electrical signals and is used to collect animal sounds.

[0941] The "HTTPS protocol" is an abbreviation for Hypertext Transfer Protocol Secure, and is a communication protocol for securely sending and receiving data over the Internet.

[0942] "Animal-specific language models" are specialized algorithms and datasets for analyzing animal behavior and vocal patterns and translating their meaning.

[0943] This invention is a system that collects, analyzes, and translates animal movements, behaviors, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behaviors, and sounds, and a server that performs the analysis and translation.

[0944] System configuration

[0945] Hardware

[0946] Terminal: Equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[0947] Server: Carries out analysis and translation processing. Implements machine learning algorithms.

[0948] software

[0949] On the device:

[0950] A program that collects movements and voices in real time and converts the data into JSON format.

[0951] A communications program that sends data to a server using the HTTPS protocol.

[0952] Server side:

[0953] A program that receives and preprocesses data.

[0954] A program that uses machine learning algorithms to analyze data and translate its meaning.

[0955] A program that generates feedback information based on translation results and sends it to the terminal.

[0956] A program that automatically suggests related products and services in a virtual store based on feedback information.

[0957] Operation flow

[0958] 1. The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphone records the frequency and volume of its calls. Cameras and other devices also capture information on the animal's facial expressions and posture.

[0959] 2. The device converts the collected data into JSON format and sends it to the server using a secure protocol (HTTPS), which provides the data to the server in real time.

[0960] 3. The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning and normalizing the data. The machine learning algorithms detect specific patterns in the animals' behavior and vocalizations and infer what they mean.

[0961] 4. The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[0962] 5. The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[0963] 6. The device receives the feedback information sent from the server and notifies the user. For example, the device can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, a voice notification function can be used to notify the user more clearly.

[0964] 7. The server automatically suggests related products and services in the virtual store based on the feedback information. For example, based on the translation result that the dog "wants to play," the server suggests a new toy in the virtual store.

[0965] Specific examples

[0966] 1. The user uses a smartphone to observe the dog's behavior in real time.

[0967] 2. Your pet wags its tail and barks "woof woof."

[0968] 3. The app will notify you, "Your dog seems happy and might want to play with his new toy."

[0969] 4. Virtual stores suggest toys for pets.

[0970] Prompt Sentence Examples

[0971] Input: The smartphone detects the dog's tail wagging and the sound of it barking "woof woof."

[0972] Processing: Data is sent to the server and analyzed by the translation system. Feedback information is generated based on the translation results and sent to the smartphone.

[0973] Output: Your dog seems happy and may be eager to play with his new toy.

[0974] This configuration makes it possible to construct a system within the scope of the invention and realize effective communication between pets and users.

[0975] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0976] Step 1:

[0977] The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. Inputs include location information, speed, and direction of movement from the sensors, as well as the frequency and volume of sounds from the microphones. This data is output as information on the animal's behavior and sounds detected through the sensors and microphones.

[0978] Step 2:

[0979] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. The input is the behavioral and voice data obtained in step 1. These data are packaged in JSON format and sent to the server via a secure communication channel for output.

[0980] Step 3:

[0981] The server preprocesses the data it receives. It takes as input the JSON data sent in step 2. It cleans and normalizes the data, resulting in clean data that can be meaningfully analyzed.

[0982] Step 4:

[0983] The server analyzes the data using machine learning algorithms. The input is the clean, pre-processed data. The machine learning algorithms detect specific patterns in the animal's behavior and vocalizations and infer their meaning. The output is the meaning of the animal's behavior and vocalizations.

[0984] Step 5:

[0985] The server translates the meaning of the animal's actions and voices based on the analysis results and updates the animal-specific language model. The input is the analysis results from step 4. Based on this data, the server generates translations corresponding to the animal's actions and voices and updates the language model. The output is the latest language model.

[0986] Step 6:

[0987] The server generates feedback information for the user based on the updated language model. The inputs are the latest language model and data on the animal's behavior and voice. The feedback information is formed as a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand. The output is the feedback information.

[0988] Step 7:

[0989] The device receives feedback information sent from the server and notifies the user. The input is the feedback information sent from the server. This is displayed on the screen of the user's smartphone or smart glasses, providing feedback to the user. The output is a notification sent to the user.

[0990] Step 8:

[0991] The server automatically suggests related products and services in the virtual store based on the feedback information. The input is the feedback information. For example, based on the feedback information "Your dog wants to play," the virtual store suggests a new toy. The output is information about the suggested products and services.

[0992] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0993] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and voices to facilitate communication with users, and provides more efficient feedback by also analyzing the user's emotions. This system is composed of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and voices, a server that performs analysis and translation, and an emotion engine that recognizes the user's emotions.

[0994] About program processing

[0995] 1. Data collection

[0996] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The sensors capture the animal's location, speed, and direction, the camera captures its facial expressions and posture, and the microphone records the frequency and volume of its calls. At the same time, the user's voice and facial expressions are also collected using an emotion engine.

[0997] 2. Data transmission

[0998] The device converts the collected data into a specific format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS). This allows the data to be provided to the server in real time, along with the user's emotional data.

[0999] 3. Data Analysis

[1000] The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning the data (removing noise and missing data) and normalizing it (aligning numerical data to the same scale). Not only animal behavior and voices, but also user emotion data are analyzed.

[1001] 4. Creating and updating a language model

[1002] The server uses the analysis results to translate the meaning of the animal's behavior and voice, and generates and updates a language model specifically for animals. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. Additionally, the user's emotional data is also reflected in the language model.

[1003] 5. Feedback Generation

[1004] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[1005] 6. User Notices

[1006] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play. You seem busy, but let's take a moment to play with him." If necessary, it can also be provided as a voice message using the voice notification function.

[1007] Specific examples

[1008] Below is a specific example of a dog standing on its hind legs and barking "woof woof."

[1009] 1. The device uses a motion sensor to detect when the dog stands up on its hind legs, and simultaneously records the dog's bark with a microphone. The camera also captures the user's facial expressions, and the audio is analyzed by an emotion engine.

[1010] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol, along with the user's emotional data.

[1011] 3. The server cleans and normalizes the received data and uses machine learning algorithms to determine whether the "reaching on its hind legs" is a "request for something." User sentiment data is also analyzed.

[1012] 4. The server updates the language model based on the results of the classification and translates this behavior as meaning "wanting a snack." The user's emotional data is also reflected in the language model.

[1013] 5. The server generates feedback information such as "The dog wants something. For example, it would be good to give it a treat. It seems you are a little busy, but please make time to respond," and sends this information to the terminal.

[1014] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[1015] In this way, the system of the present invention effectively supports communication between the user and the pet, enabling the system to respond quickly and appropriately to the pet's needs and emotions. Furthermore, by taking the user's emotions into consideration, the system provides more natural and efficient feedback.

[1016] The processing flow will be explained below.

[1017] Step 1:

[1018] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The motion sensor captures the animal's location, speed, and direction of movement. The camera captures the animal's facial expressions and posture, and the microphone records the frequency and volume of its calls. The device also collects the user's voice and facial expressions using an emotion engine.

[1019] Step 2:

[1020] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[1021] Step 3:

[1022] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS) in real time and optimized to minimize latency, along with the user's emotional data.

[1023] Step 4:

[1024] The server preprocesses the received data, which includes removing noise and missing data and normalizing the data (aligning the numerical data to the same scale), thereby improving the accuracy of the analysis.

[1025] Step 5:

[1026] The server then inputs the pre-processed data into a machine learning algorithm, which identifies patterns in the animal's behavior and vocalizations and uses these patterns to analyze the animal's intentions and emotions, while simultaneously analyzing the user's emotional data.

[1027] Step 6:

[1028] The server generates and updates a language model specifically for animals based on the analysis results. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. User emotion data is also reflected in the language model.

[1029] Step 7:

[1030] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[1031] Step 8:

[1032] The server sends the generated feedback information to the device, allowing the device to obtain the latest information in real time. For example, it generates a message such as, "Your dog wants to play. It seems you're a little busy, so please take a moment to play with him."

[1033] Step 9:

[1034] The device notifies the user of the received feedback information. This notification is displayed on the smartphone screen and also provided as a voice message using the voice notification function, allowing the user to receive feedback both visually and audibly.

[1035] Step 10:

[1036] The user can respond appropriately to their pet's needs based on the feedback information provided by the device, for example, by giving them treats or taking them outside to increase their pet's satisfaction. Further data is collected based on the user's actions, improving the accuracy of the system.

[1037] Example 2

[1038] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1039] Conventional animal-human communication systems simply analyze the animal's behavior and voice, and are unable to provide appropriate feedback that takes into account the user's emotions. This makes it difficult to accurately understand the animal's needs and emotions and respond quickly and appropriately, making it difficult to build a trusting relationship between the user and the animal.

[1040] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing data on animal movements, behaviors, and voices and user emotion data, means for analyzing the preprocessed data using a machine learning algorithm, means for generating and updating a language model based on the analysis results, means for generating feedback information for the user and transmitting it to the terminal, and means for notifying the user of the feedback information. This makes it possible to accurately understand the requests and emotions of the animal and provide quick and accurate feedback taking the user's emotions into consideration.

[1041] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[1042] The "server" is a central processing unit that receives and analyzes data sent from the terminal, generates feedback information, and sends it to the terminal.

[1043] "Data preprocessing" refers to the process of removing noise and missing data from the received data and aligning the numerical data to the same scale.

[1044] A "machine learning algorithm" is a mathematical method that analyzes received and preprocessed data to infer the meaning of animal behavior and vocalizations.

[1045] A "language model" is a model for translating the meaning of animal behavior and voices based on analysis results, and is an updatable data structure.

[1046] "Feedback information" is information generated based on a language model that conveys to the user the meaning of the animal's behavior and voice.

[1047] "User emotion data" is data relating to the user's emotional state obtained by analyzing the user's voice and facial expressions.

[1048] "JSON format" is an abbreviation for JavaScript Object Notation, and is a format that represents data in a lightweight text-based format.

[1049] The "HTTPS protocol" is an abbreviation for HyperText Transfer Protocol Secure, and is a protocol for securely communicating data.

[1050] The present invention is a system that collects and analyzes information on animal movements, behavior, and voices, and provides feedback to the user based on that information. This system has the advantage of being able to more accurately understand animal behavior and also analyze the user's emotions in order to communicate with the user more efficiently. Specific embodiments for implementing this system are described below.

[1051] Hardware and Software Configuration

[1052] The device includes the following hardware:

[1053] 1. Motion sensor: Detects animal location, speed, and direction.

[1054] 2. Camera: Captures the animal's facial expressions and posture.

[1055] 3. Microphone: Records the frequency and volume of animal sounds.

[1056] 4. Emotion engine: Analyzes the user's voice and facial expressions to generate emotion data.

[1057] The server includes the following processing capabilities:

[1058] 1. Data preprocessing: Remove noise and missing data, and normalize numerical data.

[1059] 2. Machine learning algorithms: Used to analyze animal movements, behaviors, and sounds, as well as user emotional data.

[1060] 3. Generating and updating a language model: Based on the analysis results, the meaning of the animal's behavior and voice is translated and the model is updated.

[1061] Explanation of program processing

[1062] 1. Data collection

[1063] The device detects the animal's movements, behavior, and voice in real time, and also collects the user's facial expressions and voice. For example, a motion sensor detects the animal's location, a camera captures its posture, and a microphone records its calls. The device also analyzes the user's voice and facial expressions using the user's emotion engine to generate emotional data.

[1064] 2. Data transmission

[1065] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. Data encoded in JSON format represents each data item (e.g., animal location information, call frequency, user facial expression data, etc.) as a key-value pair.

[1066] 3. Data Preprocessing and Analysis

[1067] The server preprocesses the received data, removing noise and missing data and normalizing the numerical data. The preprocessed data is then fed into a machine learning algorithm to analyze the meaning of the animal's behavior and vocalizations. The analysis also includes the user's emotional data, allowing for a more accurate understanding of the animal's needs and emotions.

[1068] 4. Creating and updating language models and providing feedback

[1069] The server generates and updates a language model based on the analysis results, interprets the meaning of the animal's behavior and vocalizations, and generates feedback information for the user based on this. For example, it generates information such as, "If a dog stands on its hind legs and barks, it means it is asking for a treat."

[1070] 5. User Notices

[1071] The device receives the feedback information sent from the server and notifies the user's smartphone, allowing the user to take appropriate action. For example, the smartphone screen might display, "Your dog wants a treat."

[1072] Specific examples

[1073] For example, consider a case where a dog stands on its hind legs and barks "woof woof." The specific processing flow in this case is as follows:

[1074] Example prompt sentence:

[1075] "If a dog stands on its hind legs and barks 'woof woof', explain the processing flow to analyze the meaning of this behavior and provide appropriate feedback to the user."

[1076] In this way, it is possible to comprehensively analyze the animal's movements and the user's emotions and provide prompt and appropriate feedback, which will lead to smoother and deeper communication between the user and the animal.

[1077] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1078] Step 1: Collect data

[1079] The device uses motion sensors, cameras, microphones, and an emotion engine to collect animal movements, behaviors, and voices, as well as the user's facial expressions and voice, in real time. The inputs are the animal's location, posture, and calls, and the user's voice and facial expression data. These data are captured by the device and used for subsequent processing. The output is structured raw data.

[1080] Specifically, the motion sensor captures the animal's position and movement vector, the camera captures its posture as a still image, and the microphone records its cries as an audio file. At the same time, the emotion engine analyzes the user's voice and facial expressions to generate emotion data.

[1081] Step 2: Convert and send data

[1082] The terminal converts the collected data into JSON format. The input is structured raw data, and the data is encoded as key-value pairs based on this. The output is a JSON object. Specifically, each data item, such as collected location information, call frequency, and facial expression data, is converted into a JSON-formatted string.

[1083] The converted JSON data is securely sent to the server using the HTTPS protocol. Here, the input is JSON format data, and the output is safe and secure data transmission to the server.

[1084] Step 3: Preprocessing the data

[1085] The server preprocesses the received JSON data. The input is JSON-formatted data received from the device. Preprocessing removes noise and missing data, and normalizes numerical data to the same scale. For example, background noise is filtered from audio data, and motion data is scaled to the range 0 to 1. The output is a clean, normalized dataset.

[1086] Specifically, the data cleansing function removes noise and runs algorithms to fill in duplicate and missing data.

[1087] Step 4: Analyze the data

[1088] The server inputs the preprocessed data into a machine learning algorithm for analysis. The input is a clean, normalized dataset. This is used to analyze what the animal's movements and sounds mean and how they are affected by the user's emotional state. The output is the analysis results, such as determining that "standing up on one's hind legs" indicates "a request for something."

[1089] Specifically, the machine learning model analyzes the data and infers behavioral patterns and their meaning.

[1090] Step 5: Generate and update a language model

[1091] The server generates and updates a language model for translating the meaning of animal behavior and vocalizations based on the analysis results. The input is the analysis result data, and the output is an updated language model. Specifically, the server uses newly collected data to provide feedback to the existing language model, improving the model's accuracy.

[1092] For example, "standing up on hind legs" is translated as "wanting a treat" and incorporated into the model.

[1093] Step 6: Generate feedback information

[1094] The server generates feedback information for the user based on the updated language model. The input is the updated language model and the analysis results. The output is feedback information. Specifically, it is written in natural language in the form of "The dog is asking for something. For example, it would be good to give it a treat."

[1095] The user's emotional data is also taken into consideration here, and feedback such as "I know you seem busy, but please take a moment to respond" is also provided.

[1096] Step 7: User Notification

[1097] The device receives the feedback information sent from the server and notifies the user. The input is the feedback information, and the output is the feedback displayed or notified on the user's smartphone or audio device. Specifically, the message "The dog wants a treat" is displayed on the smartphone screen, or a voice message is notified.

[1098] This allows users to understand the animal's needs and emotions in real time and respond appropriately.

[1099] (Application example 2)

[1100] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1101] Conventional animal behavior analysis systems can analyze and translate animal behavior and voices, but they cannot take into account the user's emotions. Furthermore, they lack the information necessary for store staff to respond appropriately to customers with pets. This results in a decline in the quality of service for customers with pets, making it difficult to improve customer satisfaction.

[1102] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1103] In this invention, the server includes means for receiving and analyzing data on the animal's movements, behavior, and voice and data on the user's emotion transmitted from the terminal, means for updating a language model that translates the meaning of the animal's behavior and voice based on the analyzed data and the user's emotion data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for displaying the behavior and emotions of the pet and its owner to store staff in real time at physical stores where pets are allowed. This enables translation based on the pet's behavior and voice and feedback that takes into account the user's emotion, making it possible to provide more appropriate and higher quality service to customers with pets.

[1104] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[1105] The "server" is a device that receives and analyzes data on animal movements, actions, and voices and user emotion data transmitted from the terminal, updates the language model, and generates feedback information.

[1106] A "language model" is a data model used to translate animal behavior and vocalizations, and is updated based on analyzed data and user emotion data.

[1107] "User emotion data" is emotion information obtained by analyzing the user's facial expressions and voice.

[1108] "Feedback information" is information that conveys the meaning of an animal's behavior and voice to the user based on the analyzed data and updated language model.

[1109] A "pet-friendly brick-and-mortar store" is a physical store that caters to customers who bring their pets with them.

[1110] "Store associates" are employees who work in physical stores and provide services to customers in general and customers with pets.

[1111] "Real-time display" refers to providing and displaying the results of an analysis of the behavior and emotions of pets and their owners to store staff immediately.

[1112] This invention is a system that collects and analyzes animal movement, behavior, and voice data, as well as user emotion data, to generate feedback information in real time and improve interaction with customers who bring their pets to physical stores.

[1113] System Program

[1114] This system uses the following hardware and software:

[1115] Hardware used

[1116] 1. Smart glasses: A device equipped with a camera, microphone, and motion sensors that collects pet and user data in real time.

[1117] 2. Server: A device that analyzes the received data and generates feedback information.

[1118] Software used

[1119] 1. Analysis engine: Software containing machine learning algorithms for animal behavior analysis and user sentiment analysis.

[1120] 2. Emotion Engine: An AI module for analyzing user emotion data.

[1121] 3. Communication protocol: JSON format and HTTPS protocol are used to send and receive data.

[1122] Data collection and analysis

[1123] The smart glasses, which function as the device, detect animal movements and behaviors using a camera and motion sensors, record voices using a microphone, and collect the user's facial expressions and voice, then analyze the user's emotional data using an analysis engine.

[1124] The collected data is converted into JSON format and sent to the server using the HTTPS protocol. The server receives the received data, first preprocesses it (cleaning and normalizing it), and then applies machine learning algorithms to analyze animal behavior, vocalizations, and user emotions.

[1125] Generating feedback information

[1126] The server updates the animal-specific language model based on the analyzed data and the user's emotional data, which can more accurately translate the meaning of the animal's actions and voices based on the newly acquired data.

[1127] Based on the updated language model, feedback information is generated, including suggestions that take into account the meaning of the animal's actions and sounds, as well as the user's emotions.

[1128] The generated feedback information is displayed in real time on the smart glasses, for example, if the pet is calm, the glasses can provide suggestions to the store clerk such as "The pet is calm, please help the owner to take their time choosing products."

[1129] Specific examples

[1130] For example, if a pet is sitting, wagging its tail, and barking, the smart glasses will detect this and send the data to the server. The server will analyze the data and generate feedback such as, "The pet is excited, but the owner seems relaxed. Please suggest a pet toy." This will enable the store clerk to respond to the customer's needs quickly and appropriately.

[1131] Prompt Sentence Examples

[1132] Pet behavior data: "Sitting, wagging tail, barking"

[1133] User emotion data: "Relaxed"

[1134] In this way, the present invention provides an effective system that analyzes the behavior and emotions of pets and their owners in real time and supports customer service for pets in physical stores.

[1135] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1136] Step 1:

[1137] The device detects animal movements, behaviors, and sounds.

[1138] Input: Real-time animal movements, actions, and sounds, user facial expressions, and voice.

[1139] Processing: The camera built into the smart glasses captures the animal's posture, movements, and behavior, and the microphone records its voice. The analysis engine also captures the user's facial expressions and analyzes the voice with the emotion engine.

[1140] Output: Collected animal and user data (image data, audio data, emotion data).

[1141] Step 2:

[1142] The device sends the data to the server.

[1143] Input: Collected animal movement, behavior, and vocalization data, and user emotion data.

[1144] Processing: The device converts the collected data into JSON format and sends it securely to the server using the HTTPS protocol.

[1145] Output: Data sent in JSON format.

[1146] Step 3:

[1147] The server preprocesses the received data.

[1148] Input: Animal movement, behavior, and vocalization data, and user emotion data, sent in JSON format.

[1149] Processing: The server deserializes the received data, cleans it (removes noise and missing data) and normalizes it (aligns the numerical data to the same scale).

[1150] Output: Preprocessed and clean data.

[1151] Step 4:

[1152] The server analyzes the pre-processed data.

[1153] Input: Preprocessed and clean data.

[1154] Processing: The server uses an analysis engine and an emotion engine to analyze the animal's behavior and voice, as well as the user's emotion data. Specifically, it uses machine learning models to identify patterns in the animal's behavior and voice and infer the user's emotion.

[1155] Output: Analysis results (meaning of animal behavior and vocalizations, emotional state of the user).

[1156] Step 5:

[1157] The server updates the language model.

[1158] Input: Analysis results (meaning of animal behavior and vocalizations, user emotional state).

[1159] Processing: The server updates the animal-specific language model based on the analysis results. Newly collected data is reflected in the model, enabling more accurate translations.

[1160] Output: An updated language model.

[1161] Step 6:

[1162] The server generates the feedback information.

[1163] Input: The updated language model.

[1164] Processing: Based on the updated language model, the server generates feedback information that takes into account the meaning of the animal's actions and sounds, as well as the user's emotions. For example, it might say, "Your pet is making noise, so I suggest a toy."

[1165] Output: Feedback information.

[1166] Step 7:

[1167] The terminal notifies the user of the feedback information.

[1168] Input: Feedback information sent by the server.

[1169] Processing: The device displays feedback information on the smart glasses, such as "Your pet wants something. Please suggest a toy to its owner."

[1170] Output: The displayed feedback information to the user.

[1171] In this way, the system includes a flow that starts with data collection at the terminal, goes through data analysis and feedback information generation at the server, and finally provides real-time feedback to the user.

[1172] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1173] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1174] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1175] [Fourth embodiment]

[1176] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1177] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1178] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1179] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1180] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1181] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1182] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1183] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1184] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1185] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1186] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1187] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1188] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1189] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds, and a server that performs the analysis and translation.

[1190] About program processing

[1191] 1. Data collection

[1192] The device uses sensors and microphones to collect real-time information on the animal's movements, behavior, and voices. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphones record the frequency and volume of its calls. In addition, cameras and other devices are used to capture information on the animal's facial expressions and posture.

[1193] 2. Data transmission

[1194] The device converts the collected data into a certain format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS), which provides the data to the server in real time.

[1195] 3. Data Analysis

[1196] The server pre-processes the data it receives, then analyzes it using machine learning algorithms that detect specific patterns in animal behavior and vocalizations and infer what they mean.

[1197] 4. Creating and updating a language model

[1198] The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[1199] 5. Feedback Generation

[1200] The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[1201] 6. User Notices

[1202] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, it can also use a voice notification function to notify the user more clearly.

[1203] Specific examples

[1204] Below is a concrete example of a dog standing on its hind legs and barking "woof woof."

[1205] 1. The device uses a motion sensor to detect the dog standing up on its hind legs, and simultaneously records the dog's bark "woof woof" with a microphone.

[1206] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol.

[1207] 3. The server cleans and normalizes the data it receives and uses machine learning algorithms to determine that the "reaching sounds" mean "a request for something."

[1208] 4. The server updates the language model based on the results of the classification and interprets this behavior as meaning "wanting a snack."

[1209] 5. The server generates feedback information such as "The dog wants something. For example, it might be a good idea to give it a treat," and sends this information to the terminal.

[1210] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[1211] In this way, the system of the present invention effectively supports communication between the user and the pet, and enables quick and appropriate responses to the pet's requests and emotions.

[1212] The processing flow will be explained below.

[1213] Step 1:

[1214] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time: the motion sensors capture the animal's location, speed, and direction, the camera captures facial expressions and posture, and the microphone records the frequency and volume of the animal's calls.

[1215] Step 2:

[1216] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[1217] Step 3:

[1218] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS), in real time and optimized to minimize latency.

[1219] Step 4:

[1220] The server pre-processes the received data before analyzing it, which involves cleaning the data (removing noise and missing data) and normalizing it (aligning the numerical data to the same scale).

[1221] Step 5:

[1222] The server then feeds the pre-processed data into a machine learning algorithm that identifies patterns in the animal's behavior and vocalizations and uses these patterns to parse the animal's intentions and emotions using pre-trained models.

[1223] Step 6:

[1224] The server generates and updates a language model specifically for animals based on the analysis results, and insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation.

[1225] Step 7:

[1226] The server generates feedback information for the user based on the updated language model, written in natural language, to help the user understand the meaning of the animal's actions and sounds.

[1227] Step 8:

[1228] The server sends the generated feedback information to the terminal, allowing the terminal to obtain the latest information in real time.

[1229] Step 9:

[1230] The device then notifies the user of the received feedback information, which can be displayed on the smartphone screen or provided as a voice message using the voice notification function, such as "Your dog wants to play."

[1231] Step 10:

[1232] Based on the feedback information provided by the device, the user can respond appropriately to the pet's requests, such as giving treats or taking the pet out to play, thereby increasing the pet's satisfaction.

[1233] In this way, the system of the present invention goes through detailed processing steps, allowing the user to understand the intentions and emotions of their pet and take prompt and appropriate action.

[1234] Example 1

[1235] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1236] In modern times, communication between animals and humans remains difficult, particularly due to limited means for properly understanding an animal's intentions and emotions. This makes it difficult for pet owners to understand their pet's requests and emotions and respond quickly and appropriately. Furthermore, conventional systems have struggled to accurately analyze an animal's behavior and voice and provide real-time feedback. Therefore, there is a need for a system that can accurately analyze an animal's movements, behavior, and voice, and then translate the animal's intentions and emotions and notify the user.

[1237] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1238] In this invention, the server includes means for generating and updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and sending it to the terminal, and means for notifying the user of the feedback information to their smartphone. This enables the animal's movements, actions, and voices to be analyzed with high accuracy, allowing the user to quickly and appropriately understand the animal's intentions and emotions.

[1239] A "terminal" is a device equipped with a sensor and microphone that detects the movements, behavior, and sounds of animals, and transmits the data collected from these to a server.

[1240] A "sensor" is a device that detects physical changes and converts them into electrical signals, and in this invention it is used to sense the movement and location of animals.

[1241] A "microphone" is a device that converts sound into an electrical signal, and in this invention it is used to record the frequency and volume of animal cries.

[1242] A "server" is a computer system that receives data sent from a terminal via a network and analyzes and processes the data.

[1243] "Data" refers to information about animal movements, behaviors, and sounds collected by sensors and microphones.

[1244] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format for representing data in a structured manner.

[1245] "HTTPS" stands for Hypertext Transfer Protocol Secure, a protocol for encrypting data communications.

[1246] A "machine learning algorithm" is a computer algorithm used to analyze data and recognize patterns.

[1247] A "language model" is a collection of databases and algorithms that translate the meaning of animal behaviors and vocalizations based on analyzed data.

[1248] "Feedback information" is a message in natural language that explains the animal's intentions and emotions and is communicated to the user.

[1249] A "smartphone" is a portable information terminal equipped with mobile communication and computer functions.

[1250] The present invention provides a system that facilitates communication with animals by collecting, analyzing, and translating data on animal movements, behaviors, and voices, and notifying the user of the translation. The following describes in detail the embodiments of the present invention.

[1251] System Overview

[1252] The system consists of a device equipped with sensors and microphones that detect animal movements, behavior, and sounds, and a server that performs analysis and translation. Specifically, the device includes a motion sensor, microphone, and camera. These devices acquire information on the animal's location, speed, direction, frequency, and volume of its calls, as well as its facial expressions and posture. The collected data is converted into JSON format and securely sent to the server using the HTTPS protocol.

[1253] Hardware and Software

[1254] Device: Equipped with motion sensors, a microphone, and a camera to detect animal movements, behavior, and sounds. Includes software to convert this data into JSON format.

[1255] Server: Contains the processing unit and storage for parsing the received data, as well as software for data cleaning, normalization, and analyzing the data using machine learning algorithms, including algorithms for generating and updating language models.

[1256] Analyzing data and generating feedback

[1257] The server preprocesses the received data, for example by removing outliers and normalizing the data. The preprocessed data is then analyzed by machine learning algorithms. These algorithms use neural networks and deep learning models to recognize patterns in animal behavior and vocalizations and infer their meaning. For example, if a dog stands on its hind legs and barks "woof woof," it may determine that this behavior indicates "it wants a treat."

[1258] Based on the analysis results, the server generates and updates a language model specific to the animal. This language model is a collection of databases and algorithms for translating the meaning of the animal's actions and vocalizations. The server then generates feedback information for the user, which explains the animal's intentions and emotions to the user in natural language. For example, a message such as "The dog is asking for a treat" may be generated.

[1259] User Notifications

[1260] The device receives the feedback information sent from the server and notifies the user's smartphone. Notifications are displayed as push notifications, and audio notifications are also possible if necessary. For example, a notification such as "Your dog wants something. Please give it a treat" may appear on the smartphone.

[1261] Specific examples

[1262] A specific example of the system's operation is shown below.

[1263] When a dog stands on its hind legs and barks "woof woof"

[1264] 1. The device's motion sensor detects the dog standing up on its hind legs, and at the same time, the microphone records the dog's bark.

[1265] 2. The device converts the collected data into JSON format and sends it to the server via HTTPS protocol.

[1266] 3. The server cleans and normalizes the data, compares it with historical data, and analyzes it. Based on this analysis, it recognizes that the "standing up on hind legs" and "woof woof" indicate a request for a treat.

[1267] 4. The server updates the language model and generates feedback and sends it to the device.

[1268] 5. The device notifies the user of the feedback information on their smartphone.

[1269] Examples of prompts for generative AI models

[1270] "What does it mean when a dog stands on its hind legs and whines?"

[1271] "When a cat wags its tail and meows, what emotion does it express?"

[1272] "What does the behavior of a bird flapping its wings and chirping loudly indicate?"

[1273] Such prompts can be used to query the generative AI model about the meaning of an animal's behavior, allowing the system to provide appropriate feedback.

[1274] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1275] Step 1: Collect data

[1276] The device uses sensors, microphones, and cameras to collect animal movements, behaviors, and sounds in real time.

[1277] Input: Animal movements, behaviors, and sounds.

[1278] Specific operations: The device's built-in sensors detect the animal's position, speed, and direction, the microphone captures the frequency and volume of the animal's calls, and the camera captures images of its facial expressions and posture.

[1279] Output: Raw data from sensors, microphones, and cameras.

[1280] Step 2: Sending data

[1281] The terminal converts the collected data into JSON format and sends it to the server using the HTTPS protocol.

[1282] Input: Raw data from sensors, microphones, and cameras.

[1283] Specific operation: The terminal software converts the raw data into JSON format and transmits it securely via HTTPS.

[1284] Output: Data converted to JSON format.

[1285] Step 3: Preprocessing the data

[1286] The server cleans and normalizes the received data.

[1287] Input: JSON formatted data.

[1288] Specific operation: The server performs outlier removal, noise filtering, and data scaling.

[1289] Output: Cleaned and normalized data.

[1290] Step 4: Analyze the data

[1291] The server analyzes the data using machine learning algorithms.

[1292] Input: Cleaned and normalized data.

[1293] How it works: The server inputs data into a machine learning model (e.g., a neural network) to analyze the animal's behavior and vocal patterns.

[1294] Output: Analysis results of behavioral and vocal patterns.

[1295] Step 5: Generate and update a language model

[1296] Based on the analysis results, the server translates the meaning of the animal's behavior and voice, and generates and updates a language model specifically for the animal.

[1297] Input: Analysis results.

[1298] Specific operation: The server compares the analysis results with the existing language model and adds new behavioral patterns to the model.

[1299] Output: An updated language model.

[1300] Step 6: Generate feedback information

[1301] The server generates feedback information for the user.

[1302] Input: The updated language model.

[1303] Specific operation: The server converts the analysis results into natural language and generates a message that is easy for the user to understand.

[1304] Output: The generated feedback information.

[1305] Step 7: User Notification

[1306] The terminal notifies the user's smartphone of the feedback information from the server.

[1307] Input: The generated feedback information.

[1308] Specific operation: The device displays a push notification on the smartphone and, if necessary, plays a voice notification.

[1309] Output: The notification message that is displayed to the user.

[1310] This embodied processing allows the system to accurately analyze animal movements, behaviors, and sounds and provide information to the user.

[1311] (Application example 1)

[1312] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1313] While conventional pet communication systems can recognize pets' movements and voices, they are unable to recommend appropriate products and services to users based on that information. This makes it difficult for users to find products and services that meet their pets' needs and to establish effective communication with their pets. Virtual stores, in particular, lack product and service recommendations tailored to pets' specific needs, limiting the user experience.

[1314] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1315] In this invention, the server includes a terminal equipped with a sensor and a microphone that detects the movements, behavior, and voice of the animal, means for receiving and analyzing data on the movements, behavior, and voice of the animal transmitted from the terminal, means for updating a language model that translates the meaning of the animal's actions and voices based on the analyzed data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for automatically suggesting related products and services in a virtual store based on the feedback information, thereby enabling the user to understand the meaning of their pet's actions and voices and easily find products and services in the virtual store that meet their pet's needs.

[1316] A "terminal" is a device equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[1317] The "server" is a system that receives and analyzes data on animal movements, behavior, and voices sent from the terminal.

[1318] A "language model" refers to the algorithms and datasets used to translate the meaning of animal behavior and vocalizations.

[1319] A "virtual store" is a virtual store operated on the Internet, which is a platform where real products can be purchased and services can be provided.

[1320] "Feedback information" refers to information provided to the user that is generated based on the meaning of the analyzed animal's behavior and voice.

[1321] "JSON format" is a data serialization format, an abbreviation for JavaScript Object Notation, and is a standard format for expressing structured data.

[1322] A "machine learning algorithm" is an algorithm that learns patterns and features from data and makes predictions and classifications for new data.

[1323] A "sensor" is a device that senses physical phenomena or conditions and collects that information as data.

[1324] A "microphone" is a device that converts sound waves into electrical signals and is used to collect animal sounds.

[1325] The "HTTPS protocol" is an abbreviation for Hypertext Transfer Protocol Secure, and is a communication protocol for securely sending and receiving data over the Internet.

[1326] "Animal-specific language models" are specialized algorithms and datasets for analyzing animal behavior and vocal patterns and translating their meaning.

[1327] This invention is a system that collects, analyzes, and translates animal movements, behaviors, and sounds, facilitating communication with users. This system consists of a terminal equipped with a sensor and microphone that detects animal movements, behaviors, and sounds, and a server that performs the analysis and translation.

[1328] System configuration

[1329] Hardware

[1330] Terminal: Equipped with sensors and a microphone to detect animal movements, behavior, and sounds.

[1331] Server: Carries out analysis and translation processing. Implements machine learning algorithms.

[1332] software

[1333] On the device:

[1334] A program that collects movements and voices in real time and converts the data into JSON format.

[1335] A communications program that sends data to a server using the HTTPS protocol.

[1336] Server side:

[1337] A program that receives and preprocesses data.

[1338] A program that uses machine learning algorithms to analyze data and translate its meaning.

[1339] A program that generates feedback information based on translation results and sends it to the terminal.

[1340] A program that automatically suggests related products and services in a virtual store based on feedback information.

[1341] Operation flow

[1342] 1. The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. For example, the sensors detect the animal's location, speed, and direction of movement, while the microphone records the frequency and volume of its calls. Cameras and other devices also capture information on the animal's facial expressions and posture.

[1343] 2. The device converts the collected data into JSON format and sends it to the server using a secure protocol (HTTPS), which provides the data to the server in real time.

[1344] 3. The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning and normalizing the data. The machine learning algorithms detect specific patterns in the animals' behavior and vocalizations and infer what they mean.

[1345] 4. The server uses the analysis results to translate the meaning of the animal's actions and sounds, and generates and updates a language model specifically for the animal. For example, if a dog stands on its hind legs and barks "woof woof," it will determine that this behavior means "it wants a treat."

[1346] 5. The server generates feedback information for the user based on the updated language model. The feedback information is in the form of a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand.

[1347] 6. The device receives the feedback information sent from the server and notifies the user. For example, the device can display a message on the user's smartphone screen saying, "Your dog wants to play." If necessary, a voice notification function can be used to notify the user more clearly.

[1348] 7. The server automatically suggests related products and services in the virtual store based on the feedback information. For example, based on the translation result that the dog "wants to play," the server suggests a new toy in the virtual store.

[1349] Specific examples

[1350] 1. The user uses a smartphone to observe the dog's behavior in real time.

[1351] 2. Your pet wags its tail and barks "woof woof."

[1352] 3. The app will notify you, "Your dog seems happy and might want to play with his new toy."

[1353] 4. Virtual stores suggest toys for pets.

[1354] Prompt Sentence Examples

[1355] Input: The smartphone detects the dog's tail wagging and the sound of it barking "woof woof."

[1356] Processing: Data is sent to the server and analyzed by the translation system. Feedback information is generated based on the translation results and sent to the smartphone.

[1357] Output: Your dog seems happy and may be eager to play with his new toy.

[1358] This configuration makes it possible to construct a system within the scope of the invention and realize effective communication between pets and users.

[1359] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1360] Step 1:

[1361] The device collects animal movements, behaviors, and sounds in real time using sensors and microphones. Inputs include location information, speed, and direction of movement from the sensors, as well as the frequency and volume of sounds from the microphones. This data is output as information on the animal's behavior and sounds detected through the sensors and microphones.

[1362] Step 2:

[1363] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. The input is the behavioral and voice data obtained in step 1. These data are packaged in JSON format and sent to the server via a secure communication channel for output.

[1364] Step 3:

[1365] The server preprocesses the data it receives. It takes as input the JSON data sent in step 2. It cleans and normalizes the data, resulting in clean data that can be meaningfully analyzed.

[1366] Step 4:

[1367] The server analyzes the data using machine learning algorithms. The input is the clean, pre-processed data. The machine learning algorithms detect specific patterns in the animal's behavior and vocalizations and infer their meaning. The output is the meaning of the animal's behavior and vocalizations.

[1368] Step 5:

[1369] The server translates the meaning of the animal's actions and voices based on the analysis results and updates the animal-specific language model. The input is the analysis results from step 4. Based on this data, the server generates translations corresponding to the animal's actions and voices and updates the language model. The output is the latest language model.

[1370] Step 6:

[1371] The server generates feedback information for the user based on the updated language model. The inputs are the latest language model and data on the animal's behavior and voice. The feedback information is formed as a natural language message that explains the animal's intentions and emotions in a way that is easy for the user to understand. The output is the feedback information.

[1372] Step 7:

[1373] The device receives feedback information sent from the server and notifies the user. The input is the feedback information sent from the server. This is displayed on the screen of the user's smartphone or smart glasses, providing feedback to the user. The output is a notification sent to the user.

[1374] Step 8:

[1375] The server automatically suggests related products and services in the virtual store based on the feedback information. The input is the feedback information. For example, based on the feedback information "Your dog wants to play," the virtual store suggests a new toy. The output is information about the suggested products and services.

[1376] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1377] This invention is a system that collects, analyzes, and translates information on animal movements, behavior, and voices to facilitate communication with users, and provides more efficient feedback by also analyzing the user's emotions. This system is composed of a terminal equipped with a sensor and microphone that detects animal movements, behavior, and voices, a server that performs analysis and translation, and an emotion engine that recognizes the user's emotions.

[1378] About program processing

[1379] 1. Data collection

[1380] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The sensors capture the animal's location, speed, and direction, the camera captures its facial expressions and posture, and the microphone records the frequency and volume of its calls. At the same time, the user's voice and facial expressions are also collected using an emotion engine.

[1381] 2. Data transmission

[1382] The device converts the collected data into a specific format (e.g., JSON format) and sends it to the server using a secure protocol (e.g., HTTPS). This allows the data to be provided to the server in real time, along with the user's emotional data.

[1383] 3. Data Analysis

[1384] The server preprocesses the received data and then analyzes it using machine learning algorithms. Preprocessing involves cleaning the data (removing noise and missing data) and normalizing it (aligning numerical data to the same scale). Not only animal behavior and voices, but also user emotion data are analyzed.

[1385] 4. Creating and updating a language model

[1386] The server uses the analysis results to translate the meaning of the animal's behavior and voice, and generates and updates a language model specifically for animals. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. Additionally, the user's emotional data is also reflected in the language model.

[1387] 5. Feedback Generation

[1388] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[1389] 6. User Notices

[1390] The device receives the feedback information sent from the server and notifies the user. For example, it can display a message on the user's smartphone screen saying, "Your dog wants to play. You seem busy, but let's take a moment to play with him." If necessary, it can also be provided as a voice message using the voice notification function.

[1391] Specific examples

[1392] Below is a specific example of a dog standing on its hind legs and barking "woof woof."

[1393] 1. The device uses a motion sensor to detect when the dog stands up on its hind legs, and simultaneously records the dog's bark with a microphone. The camera also captures the user's facial expressions, and the audio is analyzed by an emotion engine.

[1394] 2. The device converts this data into JSON format and sends it to the server using the HTTPS protocol, along with the user's emotional data.

[1395] 3. The server cleans and normalizes the received data and uses machine learning algorithms to determine whether the "reaching on its hind legs" is a "request for something." User sentiment data is also analyzed.

[1396] 4. The server updates the language model based on the results of the classification and translates this behavior as meaning "wanting a snack." The user's emotional data is also reflected in the language model.

[1397] 5. The server generates feedback information such as "The dog wants something. For example, it would be good to give it a treat. It seems you are a little busy, but please make time to respond," and sends this information to the terminal.

[1398] 6. The device displays the received information on the user's smartphone and notifies them, "Your dog wants something. For example, it might be a good idea to give it a treat."

[1399] In this way, the system of the present invention effectively supports communication between the user and the pet, enabling the system to respond quickly and appropriately to the pet's needs and emotions. Furthermore, by taking the user's emotions into consideration, the system provides more natural and efficient feedback.

[1400] The processing flow will be explained below.

[1401] Step 1:

[1402] The device uses sensors and microphones to detect animal movements, behaviors, and sounds in real time. The motion sensor captures the animal's location, speed, and direction of movement. The camera captures the animal's facial expressions and posture, and the microphone records the frequency and volume of its calls. The device also collects the user's voice and facial expressions using an emotion engine.

[1403] Step 2:

[1404] The device converts the collected data into a certain format (e.g., JSON format), where each data point contains a timestamp and sensor information, which uniquely identifies the data and makes it easier to synchronize later processing.

[1405] Step 3:

[1406] The device transmits the converted data to the server via a secure protocol (e.g., HTTPS) in real time and optimized to minimize latency, along with the user's emotional data.

[1407] Step 4:

[1408] The server preprocesses the received data, which includes removing noise and missing data and normalizing the data (aligning the numerical data to the same scale), thereby improving the accuracy of the analysis.

[1409] Step 5:

[1410] The server then inputs the pre-processed data into a machine learning algorithm, which identifies patterns in the animal's behavior and vocalizations and uses these patterns to analyze the animal's intentions and emotions, while simultaneously analyzing the user's emotional data.

[1411] Step 6:

[1412] The server generates and updates a language model specifically for animals based on the analysis results. Insights gained from newly collected data are incorporated into the existing model, enabling even more accurate translation. User emotion data is also reflected in the language model.

[1413] Step 7:

[1414] The server generates feedback information for the user based on the updated language model. The feedback information is written in natural language and provides the user with an easy-to-understand explanation of the animal's behavior and voice. The feedback also takes into account the user's emotions.

[1415] Step 8:

[1416] The server sends the generated feedback information to the device, allowing the device to obtain the latest information in real time. For example, it generates a message such as, "Your dog wants to play. It seems you're a little busy, so please take a moment to play with him."

[1417] Step 9:

[1418] The device notifies the user of the received feedback information. This notification is displayed on the smartphone screen and also provided as a voice message using the voice notification function, allowing the user to receive feedback both visually and audibly.

[1419] Step 10:

[1420] The user can respond appropriately to their pet's needs based on the feedback information provided by the device, for example, by giving them treats or taking them outside to increase their pet's satisfaction. Further data is collected based on the user's actions, improving the accuracy of the system.

[1421] Example 2

[1422] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1423] Conventional animal-human communication systems simply analyze the animal's behavior and voice, and are unable to provide appropriate feedback that takes into account the user's emotions. This makes it difficult to accurately understand the animal's needs and emotions and respond quickly and appropriately, making it difficult to build a trusting relationship between the user and the animal.

[1424] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing data on animal movements, behaviors, and voices and user emotion data, means for analyzing the preprocessed data using a machine learning algorithm, means for generating and updating a language model based on the analysis results, means for generating feedback information for the user and transmitting it to the terminal, and means for notifying the user of the feedback information. This makes it possible to accurately understand the requests and emotions of the animal and provide quick and accurate feedback taking the user's emotions into consideration.

[1425] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[1426] The "server" is a central processing unit that receives and analyzes data sent from the terminal, generates feedback information, and sends it to the terminal.

[1427] "Data preprocessing" refers to the process of removing noise and missing data from the received data and aligning the numerical data to the same scale.

[1428] A "machine learning algorithm" is a mathematical method that analyzes received and preprocessed data to infer the meaning of animal behavior and vocalizations.

[1429] A "language model" is a model for translating the meaning of animal behavior and voices based on analysis results, and is an updatable data structure.

[1430] "Feedback information" is information generated based on a language model that conveys to the user the meaning of the animal's behavior and voice.

[1431] "User emotion data" is data relating to the user's emotional state obtained by analyzing the user's voice and facial expressions.

[1432] "JSON format" is an abbreviation for JavaScript Object Notation, and is a format that represents data in a lightweight text-based format.

[1433] The "HTTPS protocol" is an abbreviation for HyperText Transfer Protocol Secure, and is a protocol for securely communicating data.

[1434] The present invention is a system that collects and analyzes information on animal movements, behavior, and voices, and provides feedback to the user based on that information. This system has the advantage of being able to more accurately understand animal behavior and also analyze the user's emotions in order to communicate with the user more efficiently. Specific embodiments for implementing this system are described below.

[1435] Hardware and Software Configuration

[1436] The device includes the following hardware:

[1437] 1. Motion sensor: Detects animal location, speed, and direction.

[1438] 2. Camera: Captures the animal's facial expressions and posture.

[1439] 3. Microphone: Records the frequency and volume of animal sounds.

[1440] 4. Emotion engine: Analyzes the user's voice and facial expressions to generate emotion data.

[1441] The server includes the following processing capabilities:

[1442] 1. Data preprocessing: Remove noise and missing data, and normalize numerical data.

[1443] 2. Machine learning algorithms: Used to analyze animal movements, behaviors, and sounds, as well as user emotional data.

[1444] 3. Generating and updating a language model: Based on the analysis results, the meaning of the animal's behavior and voice is translated and the model is updated.

[1445] Explanation of program processing

[1446] 1. Data collection

[1447] The device detects the animal's movements, behavior, and voice in real time, and also collects the user's facial expressions and voice. For example, a motion sensor detects the animal's location, a camera captures its posture, and a microphone records its calls. The device also analyzes the user's voice and facial expressions using the user's emotion engine to generate emotional data.

[1448] 2. Data transmission

[1449] The device converts the collected data into JSON format and sends it to the server using the HTTPS protocol. Data encoded in JSON format represents each data item (e.g., animal location information, call frequency, user facial expression data, etc.) as a key-value pair.

[1450] 3. Data Preprocessing and Analysis

[1451] The server preprocesses the received data, removing noise and missing data and normalizing the numerical data. The preprocessed data is then fed into a machine learning algorithm to analyze the meaning of the animal's behavior and vocalizations. The analysis also includes the user's emotional data, allowing for a more accurate understanding of the animal's needs and emotions.

[1452] 4. Creating and updating language models and providing feedback

[1453] The server generates and updates a language model based on the analysis results, interprets the meaning of the animal's behavior and vocalizations, and generates feedback information for the user based on this. For example, it generates information such as, "If a dog stands on its hind legs and barks, it means it is asking for a treat."

[1454] 5. User Notices

[1455] The device receives the feedback information sent from the server and notifies the user's smartphone, allowing the user to take appropriate action. For example, the smartphone screen might display, "Your dog wants a treat."

[1456] Specific examples

[1457] For example, consider a case where a dog stands on its hind legs and barks "woof woof." The specific processing flow in this case is as follows:

[1458] Example prompt sentence:

[1459] "If a dog stands on its hind legs and barks 'woof woof', explain the processing flow to analyze the meaning of this behavior and provide appropriate feedback to the user."

[1460] In this way, it is possible to comprehensively analyze the animal's movements and the user's emotions and provide prompt and appropriate feedback, which will lead to smoother and deeper communication between the user and the animal.

[1461] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1462] Step 1: Collect data

[1463] The device uses motion sensors, cameras, microphones, and an emotion engine to collect animal movements, behaviors, and voices, as well as the user's facial expressions and voice, in real time. The inputs are the animal's location, posture, and calls, and the user's voice and facial expression data. These data are captured by the device and used for subsequent processing. The output is structured raw data.

[1464] Specifically, the motion sensor captures the animal's position and movement vector, the camera captures its posture as a still image, and the microphone records its cries as an audio file. At the same time, the emotion engine analyzes the user's voice and facial expressions to generate emotion data.

[1465] Step 2: Convert and send data

[1466] The terminal converts the collected data into JSON format. The input is structured raw data, and the data is encoded as key-value pairs based on this. The output is a JSON object. Specifically, each data item, such as collected location information, call frequency, and facial expression data, is converted into a JSON-formatted string.

[1467] The converted JSON data is securely sent to the server using the HTTPS protocol. Here, the input is JSON format data, and the output is safe and secure data transmission to the server.

[1468] Step 3: Preprocessing the data

[1469] The server preprocesses the received JSON data. The input is JSON-formatted data received from the device. Preprocessing removes noise and missing data, and normalizes numerical data to the same scale. For example, background noise is filtered from audio data, and motion data is scaled to the range 0 to 1. The output is a clean, normalized dataset.

[1470] Specifically, the data cleansing function removes noise and runs algorithms to fill in duplicate and missing data.

[1471] Step 4: Analyze the data

[1472] The server inputs the preprocessed data into a machine learning algorithm for analysis. The input is a clean, normalized dataset. This is used to analyze what the animal's movements and sounds mean and how they are affected by the user's emotional state. The output is the analysis results, such as determining that "standing up on one's hind legs" indicates "a request for something."

[1473] Specifically, the machine learning model analyzes the data and infers behavioral patterns and their meaning.

[1474] Step 5: Generate and update a language model

[1475] The server generates and updates a language model for translating the meaning of animal behavior and vocalizations based on the analysis results. The input is the analysis result data, and the output is an updated language model. Specifically, the server uses newly collected data to provide feedback to the existing language model, improving the model's accuracy.

[1476] For example, "standing up on hind legs" is translated as "wanting a treat" and incorporated into the model.

[1477] Step 6: Generate feedback information

[1478] The server generates feedback information for the user based on the updated language model. The input is the updated language model and the analysis results. The output is feedback information. Specifically, it is written in natural language in the form of "The dog is asking for something. For example, it would be good to give it a treat."

[1479] The user's emotional data is also taken into consideration here, and feedback such as "I know you seem busy, but please take a moment to respond" is also provided.

[1480] Step 7: User Notification

[1481] The device receives the feedback information sent from the server and notifies the user. The input is the feedback information, and the output is the feedback displayed or notified on the user's smartphone or audio device. Specifically, the message "The dog wants a treat" is displayed on the smartphone screen, or a voice message is notified.

[1482] This allows users to understand the animal's needs and emotions in real time and respond appropriately.

[1483] (Application example 2)

[1484] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1485] Conventional animal behavior analysis systems can analyze and translate animal behavior and voices, but they cannot take into account the user's emotions. Furthermore, they lack the information necessary for store staff to respond appropriately to customers with pets. This results in a decline in the quality of service for customers with pets, making it difficult to improve customer satisfaction.

[1486] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1487] In this invention, the server includes means for receiving and analyzing data on the animal's movements, behavior, and voice and data on the user's emotion transmitted from the terminal, means for updating a language model that translates the meaning of the animal's behavior and voice based on the analyzed data and the user's emotion data, means for generating feedback information for the user based on the language model and transmitting it to the terminal, and means for displaying the behavior and emotions of the pet and its owner to store staff in real time at physical stores where pets are allowed. This enables translation based on the pet's behavior and voice and feedback that takes into account the user's emotion, making it possible to provide more appropriate and higher quality service to customers with pets.

[1488] A "terminal" is a device equipped with sensors and a microphone that detect the movements, behavior, and sounds of animals, collects data, and transmits it to a server.

[1489] The "server" is a device that receives and analyzes data on animal movements, actions, and voices and user emotion data transmitted from the terminal, updates the language model, and generates feedback information.

[1490] A "language model" is a data model used to translate animal behavior and vocalizations, and is updated based on analyzed data and user emotion data.

[1491] "User emotion data" is emotion information obtained by analyzing the user's facial expressions and voice.

[1492] "Feedback information" is information that conveys the meaning of an animal's behavior and voice to the user based on the analyzed data and updated language model.

[1493] A "pet-friendly brick-and-mortar store" is a physical store that caters to customers who bring their pets with them.

[1494] "Store associates" are employees who work in physical stores and provide services to customers in general and customers with pets.

[1495] "Real-time display" refers to providing and displaying the results of an analysis of the behavior and emotions of pets and their owners to store staff immediately.

[1496] This invention is a system that collects and analyzes animal movement, behavior, and voice data, as well as user emotion data, to generate feedback information in real time and improve interaction with customers who bring their pets to physical stores.

[1497] System Program

[1498] This system uses the following hardware and software:

[1499] Hardware used

[1500] 1. Smart glasses: A device equipped with a camera, microphone, and motion sensors that collects pet and user data in real time.

[1501] 2. Server: A device that analyzes the received data and generates feedback information.

[1502] Software used

[1503] 1. Analysis engine: Software containing machine learning algorithms for animal behavior analysis and user sentiment analysis.

[1504] 2. Emotion Engine: An AI module for analyzing user emotion data.

[1505] 3. Communication protocol: JSON format and HTTPS protocol are used to send and receive data.

[1506] Data collection and analysis

[1507] The smart glasses, which function as the device, detect animal movements and behaviors using a camera and motion sensors, record voices using a microphone, and collect the user's facial expressions and voice, then analyze the user's emotional data using an analysis engine.

[1508] The collected data is converted into JSON format and sent to the server using the HTTPS protocol. The server receives the received data, first preprocesses it (cleaning and normalizing it), and then applies machine learning algorithms to analyze animal behavior, vocalizations, and user emotions.

[1509] Generating feedback information

[1510] The server updates the animal-specific language model based on the analyzed data and the user's emotional data, which can more accurately translate the meaning of the animal's actions and voices based on the newly acquired data.

[1511] Based on the updated language model, feedback information is generated, including suggestions that take into account the meaning of the animal's actions and sounds, as well as the user's emotions.

[1512] The generated feedback information is displayed in real time on the smart glasses, for example, if the pet is calm, the glasses can provide suggestions to the store clerk such as "The pet is calm, please help the owner to take their time choosing products."

[1513] Specific examples

[1514] For example, if a pet is sitting, wagging its tail, and barking, the smart glasses will detect this and send the data to the server. The server will analyze the data and generate feedback such as, "The pet is excited, but the owner seems relaxed. Please suggest a pet toy." This will enable the store clerk to respond to the customer's needs quickly and appropriately.

[1515] Prompt Sentence Examples

[1516] Pet behavior data: "Sitting, wagging tail, barking"

[1517] User emotion data: "Relaxed"

[1518] In this way, the present invention provides an effective system that analyzes the behavior and emotions of pets and their owners in real time and supports customer service for pets in physical stores.

[1519] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1520] Step 1:

[1521] The device detects animal movements, behaviors, and sounds.

[1522] Input: Real-time animal movements, actions, and sounds, user facial expressions, and voice.

[1523] Processing: The camera built into the smart glasses captures the animal's posture, movements, and behavior, and the microphone records its voice. The analysis engine also captures the user's facial expressions and analyzes the voice with the emotion engine.

[1524] Output: Collected animal and user data (image data, audio data, emotion data).

[1525] Step 2:

[1526] The device sends the data to the server.

[1527] Input: Collected animal movement, behavior, and vocalization data, and user emotion data.

[1528] Processing: The device converts the collected data into JSON format and sends it securely to the server using the HTTPS protocol.

[1529] Output: Data sent in JSON format.

[1530] Step 3:

[1531] The server preprocesses the received data.

[1532] Input: Animal movement, behavior, and vocalization data, and user emotion data, sent in JSON format.

[1533] Processing: The server deserializes the received data, cleans it (removes noise and missing data) and normalizes it (aligns the numerical data to the same scale).

[1534] Output: Preprocessed and clean data.

[1535] Step 4:

[1536] The server analyzes the pre-processed data.

[1537] Input: Preprocessed and clean data.

[1538] Processing: The server uses an analysis engine and an emotion engine to analyze the animal's behavior and voice, as well as the user's emotion data. Specifically, it uses machine learning models to identify patterns in the animal's behavior and voice and infer the user's emotion.

[1539] Output: Analysis results (meaning of animal behavior and vocalizations, emotional state of the user).

[1540] Step 5:

[1541] The server updates the language model.

[1542] Input: Analysis results (meaning of animal behavior and vocalizations, user emotional state).

[1543] Processing: The server updates the animal-specific language model based on the analysis results. Newly collected data is reflected in the model, enabling more accurate translations.

[1544] Output: An updated language model.

[1545] Step 6:

[1546] The server generates the feedback information.

[1547] Input: The updated language model.

[1548] Processing: Based on the updated language model, the server generates feedback information that takes into account the meaning of the animal's actions and sounds, as well as the user's emotions. For example, it might say, "Your pet is making noise, so I suggest a toy."

[1549] Output: Feedback information.

[1550] Step 7:

[1551] The terminal notifies the user of the feedback information.

[1552] Input: Feedback information sent by the server.

[1553] Processing: The device displays feedback information on the smart glasses, such as "Your pet wants something. Please suggest a toy to its owner."

[1554] Output: The displayed feedback information to the user.

[1555] In this way, the system includes a flow that starts with data collection at the terminal, goes through data analysis and feedback information generation at the server, and finally provides real-time feedback to the user.

[1556] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1557] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1558] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1559] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1560] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1561] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1562] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1563] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1564] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1565] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1566] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1567] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1568] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1569] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1570] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1571] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1572] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1573] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1574] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1575] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1576] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1577] The following is further disclosed regarding the above embodiment.

[1578] Understood. Below are proposed claims for the described system.

[1579] (Claim 1)

[1580] A terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds;

[1581] a server that receives and analyzes data on animal movements, behaviors, and sounds transmitted from the terminal;

[1582] a means for updating a language model that translates the meaning of animal behavior and voice based on the analyzed data;

[1583] means for generating feedback information for a user based on the language model and transmitting the feedback information to the terminal;

[1584] A system including:

[1585] (Claim 2)

[1586] The system according to claim 1, further comprising a terminal that transmits data on animal movements, behaviors, and sounds to a server in JSON format.

[1587] (Claim 3)

[1588] 2. The system of claim 1, wherein the server uses machine learning algorithms to analyze animal behavior and voices and update a language model specific to the animal.

[1589] "Example 1"

[1590] (Claim 1)

[1591] A terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds;

[1592] a server that receives and analyzes data on animal movements, behaviors, and sounds transmitted from the terminal;

[1593] means for generating and updating a language model that translates the meaning of animal behaviors and sounds based on the analyzed data;

[1594] means for generating feedback information for a user based on the language model and transmitting the feedback information to the terminal;

[1595] means for notifying the user of the feedback information on a smartphone;

[1596] A system including:

[1597] (Claim 2)

[1598] The system according to claim 1, further comprising a terminal that transmits data on animal movements, behaviors, and sounds to a server in JSON format.

[1599] (Claim 3)

[1600] 2. The system of claim 1, wherein the server uses machine learning algorithms to analyze animal behavior and vocalizations and generate and update language models specific to the animals.

[1601] "Application Example 1"

[1602] (Claim 1)

[1603] A terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds;

[1604] a server that receives and analyzes data on animal movements, behaviors, and sounds transmitted from the terminal;

[1605] a means for updating a language model that translates the meaning of animal behavior and voice based on the analyzed data;

[1606] means for generating feedback information for a user based on the language model and transmitting the feedback information to the terminal;

[1607] means for automatically proposing related products and services in a virtual store based on the feedback information;

[1608] A system including:

[1609] (Claim 2)

[1610] The system according to claim 1, further comprising a terminal that transmits data on animal movements, behaviors, and sounds to a server in JSON format.

[1611] (Claim 3)

[1612] 2. The system of claim 1, wherein the server uses machine learning algorithms to analyze animal behavior and voices and update a language model specific to the animal.

[1613] "Example 2: Combining Emotion Engines"

[1614] (Claim 1)

[1615] A terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds;

[1616] a server that receives and analyzes data on animal movements, behaviors, and sounds transmitted from the terminal;

[1617] means for the server to preprocess and analyze data using machine learning algorithms;

[1618] means for generating and updating a language model that translates the meaning of animal behavior and voice based on the analyzed data and user emotion data;

[1619] means for generating feedback information based on the language model while also taking into consideration the user's emotions, and transmitting the feedback information to the terminal;

[1620] means for notifying a user of the feedback information by the terminal;

[1621] A system including:

[1622] (Claim 2)

[1623] 2. The system according to claim 1, further comprising a terminal that transmits data on animal movements, behaviors, and voices and data on user emotions in JSON format to the server.

[1624] (Claim 3)

[1625] The system of claim 1, wherein the server uses machine learning algorithms to analyze animal behavior, voice, and user emotion data to generate and update language models specific to animals.

[1626] "Application example 2 when combining emotion engines"

[1627] (Claim 1)

[1628] A terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds;

[1629] a server that receives and analyzes data on animal movements, behaviors, and sounds transmitted from the terminal;

[1630] means for updating a language model that translates the meaning of animal behavior and voice based on the analyzed data and user emotion data;

[1631] means for generating feedback information for a user based on the language model and transmitting the feedback information to the terminal;

[1632] A means for displaying the behavior and emotions of pets and owners to store staff in real time in a pet-friendly physical store;

[1633] A system including:

[1634] (Claim 2)

[1635] 2. The system according to claim 1, further comprising a terminal that transmits data on animal movements, behaviors, and voices and data on user emotions in JSON format to the server.

[1636] (Claim 3)

[1637] 2. The system of claim 1, wherein the server uses a machine learning algorithm to analyze the animal's behavior and voice and the user's emotional data, and update the animal-specific language model. [Explanation of symbols]

[1638] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A terminal equipped with a sensor and microphone that detects animal movements, behavior, and sounds; a server that receives and analyzes data on animal movements, behaviors, and sounds transmitted from the terminal; a means for updating a language model that translates the meaning of animal behavior and voice based on the analyzed data; means for generating feedback information for a user based on the language model and transmitting the feedback information to the terminal; A system including:

2. The system according to claim 1, further comprising a terminal that transmits data on animal movements, behaviors, and sounds to the server in JSON format.

3. The system of claim 1 , wherein the server uses machine learning algorithms to analyze animal behavior and vocalizations and update a language model specific to the animal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A