system

By recording and analyzing animal movements and sounds, and using machine learning algorithms to infer their intentions and emotions, the problem of human difficulty in understanding animals has been solved, enabling accurate communication and health monitoring.

JP2026069153APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Humans struggle to accurately understand animals' intentions and emotions, and lack effective methods to monitor their health, leading to misunderstandings and failure to address problems in a timely manner.

Method used

The device records the animal's movements and sounds, uses machine learning algorithms to analyze the data to infer its intentions and emotions, and then notifies the user of the results in natural language.

Benefits of technology

It enables accurate communication between humans and animals and continuous monitoring of animal health, providing timely feedback and appropriate action recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069153000001_ABST
    Figure 2026069153000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] An input means for acquiring animal movements and sounds, Analysis means for analyzing acquired motion and sound, An estimation means for inferring the intentions and emotions of an animal based on the analysis results of the aforementioned analysis means, A notification means for notifying the user of the prediction result obtained by the aforementioned prediction means, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] <统一码转义字符: Communication between humans and animals is generally carried out through animal vocalizations and movements, but it is difficult for humans to accurately understand this. For this reason, humans often misunderstand the intentions and emotions of animals and may not be able to take appropriate actions. In addition, the means of grasping the health status of animals are also limited, and it is difficult to respond before the problem worsens. Therefore, there is a need to provide a means to accurately infer the intentions and emotions of animals from their movements and voices and convey them to humans.

Means for Solving the Problems

[0005] This invention includes an input means for acquiring animal movements and sounds, and an inference means for inferring the animal's intentions and emotions using an analysis means for analyzing the acquired data. Furthermore, it includes a notification means for notifying the user of the inference results in natural language. This configuration facilitates communication between humans and animals and enables continuous monitoring of the animal's health status. The analysis means uses a machine learning algorithm to achieve more accurate analysis based on the individual characteristics of the animal.

[0006] "Action" refers to a series of movements or actions that an animal performs using its body.

[0007] "Sound" refers to animal vocalizations and other patterns of sounds.

[0008] "Input means" refers to devices and sensors used to acquire animal movements and sounds.

[0009] "Analysis means" refers to software or hardware used to analyze acquired motion and audio data.

[0010] "Inference methods" refer to processes and algorithms that determine an animal's intentions and emotions based on the results of analytical methods.

[0011] "Notification means" refers to methods or devices used to communicate the predicted results to the user.

[0012] The term "system" refers to the entire framework, including a series of processes and means for analyzing animal behavior and sounds and notifying the user of the results. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the language used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system that recognizes and analyzes the movements and sounds of animals to infer their intentions and emotions, and notifies the user of those intentions. The following processes are necessary to implement this system.

[0035] The device is equipped with a camera and microphone to record the animal's movements and sounds, and uses this to acquire real-time data on the pet. Users record their pet's movements and sounds through this device and send that data to the system.

[0036] The server receives motion and audio data transmitted from the terminal and analyzes this data in detail using an analysis tool. The analysis tool incorporates machine learning algorithms to extract patterns of animal vocalizations and characteristics of their movements, and based on these, infers the animal's emotions and intentions.

[0037] The server then generates a message in natural language based on the inferred intentions and emotions of the animal and sends it to the terminal. This message is provided in a format that the user can intuitively understand, helping to gain a deeper understanding of the animal's condition.

[0038] Users can take appropriate action regarding their pets based on the information they receive. They can also support continuous performance improvements by sending feedback to the system as needed.

[0039] As a concrete example, consider a scenario where a user's dog suddenly starts barking intensely. The user uses their smartphone to record this situation and sends the data to the system. The server analyzes the data, evaluating the barking patterns and the dog's body movements to infer that the dog is feeling anxious. As a result, a message is sent to the user's device stating, "The dog appears anxious. Please move it to a quiet place to calm it down." Based on this information, the user can move the dog to a quiet place and address the problem by providing reassurance.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The device collects animal movements and sounds using a camera and microphone, and stores this data digitally. The collected data includes video and audio data.

[0043] Step 2:

[0044] The terminal performs preprocessing, including noise reduction and data compression, before transmitting the digital data to the server via the internet. Streaming technology is used for stable data transmission.

[0045] Step 3:

[0046] The server processes the received data using analysis tools. For audio data, it applies a speech recognition algorithm to identify patterns in vocalizations, and for behavioral data, it uses image recognition technology to analyze the characteristics of the movements.

[0047] Step 4:

[0048] The server infers the animal's intentions and emotions based on the analysis results. In doing so, it refers to a machine learning model and determines the most likely intentions and emotions based on past learning results.

[0049] Step 5:

[0050] The server generates natural language messages based on inferred intentions and emotions, and formats them in a way that is easy for the user to understand. These messages are then notified to the user in real time.

[0051] Step 6:

[0052] The terminal receives messages sent from the server and displays them to the user. Based on this information, the user can take appropriate action regarding the animal.

[0053] Step 7:

[0054] Users can contribute to improving the system's accuracy by sending feedback to the server as needed. This feedback includes information about the user's observations and the accuracy of the system's predictions.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] In recent years, there has been a growing need to understand animal emotions and intentions. However, conventional methods require specialized knowledge for analysis, making it difficult for ordinary users to intuitively grasp the state of an animal. Furthermore, there is a need for a system that can analyze animal movements and vocalizations in real time and take quick and appropriate action.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes a recording device for recording the movements and sounds of an animal, an analysis device for analyzing the recorded movements and sounds using a machine learning algorithm that utilizes a generative AI model, and an inference device for inferring the animal's intentions and emotions based on the analysis results of the analysis device. This makes it possible to precisely analyze the movements and sounds of an animal and immediately notify the user with an intuitive natural language message.

[0060] "Animals" refer to a group of multicellular organisms belonging to the animal kingdom in biological classification, and are organisms that may exhibit emotions and intentions.

[0061] "Movement" refers to a series of movements or changes in movement that occur when an animal moves its body.

[0062] "Sound" refers to the vocalizations and other sounds made by animals.

[0063] A "recording device" is a device used to collect animal movements and sounds and store them as data.

[0064] A "generative AI model" is a type of algorithm designed to learn the characteristics of data and perform specific tasks, and is used for data analysis and prediction.

[0065] A "machine learning algorithm" is a computational method that learns patterns from data and uses those patterns to make predictions or perform classifications.

[0066] An "analytical device" is a device that performs calculations to analyze recorded data and understand its contents.

[0067] A "prediction device" is a device that identifies an animal's intentions and emotions from analyzed data.

[0068] "Natural language" usually refers to the forms of language that humans use on a daily basis, and is closer to actual conversation and writing than to language generated by computers.

[0069] A "notification device" is a device that communicates the results of analysis and estimation to the user, and its role is to provide information.

[0070] This system provides a comprehensive solution for analyzing animal emotions and intentions and notifying the user. A detailed embodiment of the system is described below.

[0071] The device is equipped with a camera and microphone to record animal movements and sounds. This allows the device to collect animal movements as video and vocalizations and other sounds as audio data in real time. Users can monitor their pets' activities using mobile devices such as smartphones and tablets. This allows users to instantly detect any abnormal behavior or changes in the animal's condition.

[0072] The server uses machine learning algorithms powered by generative AI models to analyze the behavioral and audio data transmitted from the terminal. The server receives this data, processes the audio data using spectrogram analysis techniques, and vectorizes the behavioral features by analyzing the video data frame by frame. This analysis enables highly accurate prediction of the animal's emotions and intentions. The generative AI model associates specific behavioral patterns of the animal with their emotions and intentions.

[0073] The server translates the animal's emotions and intentions into natural language based on the analysis results. This ensures that the generated messages are presented to the user in an intuitively understandable format. This process includes the automatic generation of messages such as, "The dog appears anxious. Please move it to a quiet place to calm it down," using a generative AI model.

[0074] For example, if a user's dog suddenly starts barking violently, the user records this situation via their device. The server analyzes this data, evaluating the barking pattern and the dog's body movements to infer that the dog is feeling anxious. As a result, the user receives a message via their device stating, "Your dog appears anxious. Please move it to a quiet place to calm it down." This allows the user to take intuitive action.

[0075] An example of a prompt for a generative AI model might be, "Analyze this dog's barking and behavioral data to infer its emotions and intentions." Following this prompt, the server analyzes the data and infers its emotions and intentions.

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] The device records animal movements as video and collects vocalizations and sounds as audio data. It uses a camera and microphone to detect animal movements and sounds in real time as input. This input is converted into digital data by the recording device. For example, a dog barking is recorded, and its audio is saved as a file.

[0079] Step 2:

[0080] The terminal sends collected motion data (video files) and audio data (audio files) to the server. This transmission is performed using a secure communication protocol, ensuring data confidentiality. The input is recorded digital data, and the output is the data file sent to the server.

[0081] Step 3:

[0082] The server analyzes the motion and audio data received from the terminal. The input consists of video and audio data, and the server's analysis method is a machine learning algorithm using a generative AI model. Features are extracted from the audio data through spectrogram analysis, and the video data is analyzed frame by frame, with motion features vectorized. The output is the analyzed feature data.

[0083] Step 4:

[0084] The server performs a process of inferring the animal's intentions and emotions based on the analyzed feature data. The input is the feature data obtained in the previous step, and the server's inference method associates specific patterns with emotions and intentions. The output is the inferred emotion and intention data.

[0085] Step 5:

[0086] The server generates natural language messages based on inferred data. The input is inferred intent and emotion data, and the output is an intuitively understandable natural language message. A generative AI model is used to generate messages that help users perceive and respond to their dog's anxiety.

[0087] Step 6:

[0088] The server sends the generated natural language message to the terminal. The input is message data, and the output is the message sent to the terminal. Specifically, the message "The dog appears to be feeling anxious." is generated and a notification is sent to the user's terminal.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] Within factories, there is a need to quickly detect safety risks caused by animal intrusion and strengthen safety management. Conventional systems have difficulty accurately assessing the emotions and intentions of animals, making it difficult to take immediate action. This can lead to unexpected accidents and decreased production efficiency.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes recording means for acquiring animal movements and sounds, analysis means for analyzing the acquired movements and sounds, emotion estimation means for inferring the animal's intentions and emotions based on the analysis results of the analysis means, information provision means for notifying the user of the estimation results by the emotion estimation means, and monitoring and management means for supporting safety management within the factory using the animal emotion estimation results. This makes it possible to improve safety by monitoring animal behavior within the factory in real time and providing appropriate information immediately.

[0094] "Recording means" refers to devices and methods for acquiring animal movements and sounds, and provides a foundation for understanding animal behavior in real time.

[0095] "Analysis means" refers to devices and methods for analyzing acquired animal behavior and sounds, which use machine learning algorithms to extract and analyze data features.

[0096] "Emotion estimation means" refers to a device or method for inferring an animal's intentions and emotions based on the analysis results obtained by the aforementioned analysis means.

[0097] "Information provision means" refers to devices or methods for notifying the user of the inference results obtained by emotion estimation means, and provides the information by converting it into natural language.

[0098] "Monitoring and management means" refers to devices and methods for managing safety within a factory using the results of animal emotion estimation, and are intended to improve safety within the factory.

[0099] To realize this invention, a system is constructed for acquiring and analyzing animal behavior and sounds. The system mainly consists of recording means, analysis means, emotion estimation means, information provision means, and monitoring and management means.

[0100] The server uses cameras and microphones mounted on robots that autonomously patrol the factory to record animal movements and sounds. This functions as a recording tool. The acquired data is sent to the server and analyzed using software such as Python and TENSORFLOW®. The analysis tool uses machine learning algorithms to extract animal movement patterns and vocal characteristics, and based on this, infers the animal's intentions and emotions.

[0101] The analyzed and inferred results are converted into natural language that users can intuitively understand by sentiment estimation tools. This information is then communicated to administrators by information providers, enabling a rapid response on-site. A concrete example of this notification process is when a night security robot at a factory detects the intrusion of a suspicious animal; it generates an alert stating, "A suspicious animal has entered the premises and is exhibiting aggressive behavior. Please be careful," and sends it to the administrator.

[0102] Furthermore, the monitoring and management system integrates the prediction results with the factory's safety management system, automating alarm triggers and ensuring higher safety standards.

[0103] A concrete example of a prompt for a generative AI model is, "Analyze this dataset, identify the intentions and emotions of the animals, and generate a natural language warning message based on the results." This prompt is input to the generative AI model when analyzing the acquired data and is used to obtain appropriate analysis results.

[0104] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0105] Step 1:

[0106] The terminal uses cameras and microphones mounted on robots moving around the factory to record animal movements and sounds in real time. The input consists of camera video data and audio data, and this data is sent to a server as output.

[0107] Step 2:

[0108] The server receives video and audio data provided by the terminal and prepares the data for analysis. The input is the received data, and the output is the formatted data after noise reduction and signal processing.

[0109] Step 3:

[0110] The server utilizes software such as Python and TensorFlow to apply advanced machine learning algorithms to the formatted data and perform analysis. This analysis extracts animal behavior patterns and vocal features, which are then used as input. The output is the feature extraction results.

[0111] Step 4:

[0112] The server infers the animal's intentions and emotions based on the extracted features. It uses a generative AI model to analyze the data through prompt messages. The input at this stage is the feature extraction results, and the output is the inferred intentions and emotions of the animal.

[0113] Step 5:

[0114] The server converts the emotion estimation results into a message expressed in natural language and notifies the user using an information delivery method. An example of a prompt message generated by the system using a generative AI model is: "Analyze this dataset to identify the animal's intentions and emotions, and generate a natural language warning message based on the results." The input is the emotion estimation result, and the output is a natural language message.

[0115] Step 6:

[0116] The user checks the message notified on the terminal and takes appropriate countermeasures as needed, depending on the situation in the factory. The input is a natural language message, and the output is the implementation of the countermeasure.

[0117] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0118] This invention is a system that facilitates communication by simultaneously analyzing the emotions of both the animal and the user. In addition to functions for recognizing and analyzing the animal's movements and sounds, this system integrates an emotion engine that recognizes emotions from the user's facial expressions and voice. The following describes embodiments for carrying out this invention.

[0119] The device is equipped with a camera and microphone to record animal movements and sounds, thereby acquiring real-time data. Similarly, it also has a camera and microphone to record the user's facial expressions and tone of voice. This allows the device to simultaneously acquire data from both the animal and the user.

[0120] The server processes animal data transmitted from the terminal using analysis tools to infer the animal's intentions and emotions. The analysis tools utilize machine learning algorithms and have already learned the individual vocal patterns and behavioral characteristics of each animal. Simultaneously, an emotion engine included in the server analyzes user data to infer the user's emotions. This engine can identify the user's current emotional state through facial recognition and voice analysis technologies.

[0121] The server then uses an analysis combining the animal's intentions and emotions with the user's emotions to generate an appropriate message. This message is adjusted in content and tone according to the user's current emotional state, ensuring the most effective communication.

[0122] For example, if the user is feeling stressed and the dog is showing signs of anxiety, the server will generate a message such as, "It's important for you to relax while also making your dog feel secure," and notify the user's device. Based on this information, the user can take appropriate action that suits both their and their dog's emotional state.

[0123] By receiving these notifications, users can gain a deeper understanding of the animals' conditions and take appropriate actions that are considerate of each other's feelings. They can also provide information to the server through the feedback function to improve the accuracy of the system's analysis. This feedback plays a crucial role in effective communication.

[0124] The following describes the processing flow.

[0125] Step 1:

[0126] The device acquires data using a camera and microphone to record animal movements and sounds. Simultaneously, it collects data on the user's facial expressions and tone of voice. This data is temporarily stored within the device.

[0127] Step 2:

[0128] The terminal preprocesses the acquired data, removing noise and extracting only the necessary information. The preprocessed data is then compressed and sent to the server via communication.

[0129] Step 3:

[0130] The server processes the received animal behavior and audio data using analysis tools. Machine learning algorithms are employed to extract patterns in animal sounds and movements.

[0131] Step 4:

[0132] The server uses an emotion engine to analyze the user's facial expressions and tone of voice data. The algorithm infers the user's emotional state and records that information in a database.

[0133] Step 5:

[0134] The server combines inferred animal intentions and emotions with the user's emotional state to generate the optimal message. This message is then adjusted to suit the user's emotional state.

[0135] Step 6:

[0136] The server translates the message into natural language and sends it to the terminal. This message is designed to be intuitively understandable to the user.

[0137] Step 7:

[0138] The device receives messages sent from the server and notifies the user. Notifications are made via audio alerts or screen displays.

[0139] Step 8:

[0140] Users can respond to the information they receive in accordance with the emotional state of the animal and themselves. They can also use the feedback function to provide additional information to the server, which can be used for future analysis.

[0141] (Example 2)

[0142] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0143] In recent years, there has been a growing demand for technologies that facilitate communication between humans and animals. However, conventional technologies focus solely on analyzing animal movements and vocalizations, lacking consideration for the user's emotional state. Therefore, there is a need for a means to simultaneously analyze the animal's intentions and emotions, as well as the user's emotions, and to achieve appropriate two-way communication based on this analysis.

[0144] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0145] In this invention, the server includes a device for acquiring the animal's movements and sounds, a device for acquiring the user's facial expressions and sounds, and a processing device for analyzing the acquired movements, sounds, and facial expressions. This enables communication that simultaneously analyzes the emotions of the animal and the user, and generates appropriate two-way messages based on the results.

[0146] A "device" is a combination of hardware and software used to perform a specific function.

[0147] "Action" refers to a physical activity or action performed by a living being.

[0148] "Sound" refers to the waveform of sounds emitted by animals and humans, and by analyzing its characteristics, it is possible to infer emotions and intentions.

[0149] "Facial expression" is visual information that conveys emotions and intentions through the movement of facial muscles.

[0150] A "processing device" is a device that takes data as input, analyzes it, and has the computational function to extract or infer specific information.

[0151] "Analysis" is the process of breaking down, comparing, and evaluating data to understand its meaning and relationships.

[0152] A "prediction device" is a device that has a computational function to infer the intentions and emotions of a certain subject based on collected data.

[0153] A "notification device" is a device that has the function of conveying predicted results or generated messages to the user.

[0154] "Natural language" refers to the language that humans use on a daily basis, and the text that is generated in a format suitable for computer processing.

[0155] This invention is a system that simultaneously analyzes the emotions of both animals and their users. The system aims to facilitate communication between the two by combining functions that analyze the animal's movements and sounds with functions that analyze the user's facial expressions and sounds.

[0156] The device is equipped with a camera and microphone to record the animal's movements and sounds. This hardware makes it possible to acquire real-time data on the animal's physical movements and vocalizations. The same device also has a camera and microphone to record the user's facial expressions and voice, allowing for the simultaneous acquisition of the user's emotional data.

[0157] The server analyzes animal and user data transmitted from the terminal. This analysis utilizes machine learning algorithms that analyze animal sounds and movements, as well as the user's facial expressions and tone of voice. This allows the server to accurately predict the animal's emotions and intentions, as well as the user's emotions. Various server-based analysis software and emotion engines are used in this analysis.

[0158] Based on the analysis results, the server generates messages tailored to the animal's and the user's situation. These messages are written in natural language, adjusted to the most effective content and tone, and delivered to the device. Based on this information, users can take appropriate actions that align with the emotions of both the animal and themselves.

[0159] As a concrete example, considering the example prompt, in response to the input, "Please advise how the user should react when the dog appears anxious," the system can provide appropriate advice. This realizes the use of generative AI models to support communication between animals and humans.

[0160] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0161] Step 1:

[0162] The device activates its camera and microphone to capture animal movements and sounds, as well as the user's facial expressions and voice, in real time. The acquired data includes animal movements and vocalizations, as well as the user's facial expressions and voice tone. This data is initially processed within the device for later detailed analysis.

[0163] Step 2:

[0164] The terminal sends the initially processed data to the server. Wireless communication technologies such as Wi-Fi and Bluetooth are used for this communication. The input data includes animal movement data, voice data, user facial expression data, and voice data, which are used as basic data for the next analysis step on the server.

[0165] Step 3:

[0166] The server receives the transmitted data and uses machine learning algorithms to analyze the animal's movements and sounds, as well as the user's facial expressions and voice. Based on the animal data, it infers the animal's intentions and emotions from its behavioral patterns and vocal characteristics. Regarding the user's data, it identifies emotions from changes in facial expressions and tone of voice. As a result of the analysis, the emotional states of both the animal and the user are output.

[0167] Step 4:

[0168] The server uses a generative AI model to generate appropriate messages based on the analysis of the animal's intentions and emotions, as well as the user's emotions. As a prompt, the analysis results are input into the generative AI model, which then outputs an effective message in natural language. This message is structured as specific advice that is beneficial to both the animal and the user.

[0169] Step 5:

[0170] The server sends the generated message to the terminal. The terminal notifies the user of this message either visually or audibly. Based on the received message, the user can take specific actions to adjust their communication with the animal.

[0171] (Application Example 2)

[0172] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0173] In stores that allow pets, there is a challenge in appropriately analyzing the emotional states of both customers and their pets and enabling store staff to take the most appropriate action based on that analysis. Therefore, there is a need for effective support to improve customer satisfaction and pet comfort.

[0174] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0175] In this invention, the server includes acquisition means for acquiring animal behavior and sounds, emotion analysis means for simultaneously analyzing the emotional states of customers and animals, and motion analysis means for analyzing the acquired behavior and sounds. This makes it possible to analyze the emotions of customers and pets with high accuracy and to provide appropriate customer service suggestions to store employees.

[0176] "Acquisition means" refers to equipment and technology for detecting animal behavior and sounds and converting them into data.

[0177] "Emotional analysis means" refers to technology or devices for simultaneously analyzing the emotional states of customers and animals. This means has the function of analyzing data using machine learning and emotion recognition algorithms to identify emotions.

[0178] "Motion analysis means" refers to technology or equipment used to analyze acquired animal behavior and vocalizations. This means has the function of processing data in order to understand the meaning and intent of the behavior.

[0179] "Inference means" refers to techniques or devices used to infer the emotional state or intentions of a subject based on analysis results. These means play a role in inferring the most likely emotional state based on the analyzed data.

[0180] "Proposal provision means" refers to technology or equipment for presenting specific actions or responses to external parties based on the prediction results obtained by prediction means. This means is used to instruct the optimal action based on the information obtained.

[0181] One embodiment of this invention provides a system that understands the emotions of customers accompanied by pets and their pets in a store, and enables optimal customer service. The system mainly consists of terminals and a server.

[0182] The terminal is equipped with a camera and microphone that record the movements of customers and their pets in real time, thereby collecting motion and audio data of customers and pets when they enter the store. The terminal is responsible for transmitting the collected data to a server.

[0183] The server uses machine learning algorithms to analyze animal behavior and customer emotions based on the received data. To analyze animal and customer emotions simultaneously, emotion analysis means are used to identify their respective emotional states. Motion analysis means are used to analyze animal behavior in detail and infer the animal's intentions and emotions.

[0184] Based on the insights gained from these analysis results, the server generates suggestions to facilitate optimal customer service. These suggestions are translated into natural language by a suggestion delivery system and notified to store employees. As a result, employees can provide optimal customer service tailored to the customer and their pet's condition, thereby improving customer satisfaction.

[0185] For example, if the system detects that a customer entering the store appears tense and their pet is frightened, it will generate a natural language suggestion such as "Let's give your pet some time to calm down" and provide it to the store staff via a terminal. This allows employees to provide prompt and appropriate customer service.

[0186] As an example of a prompt to a generative AI model, you could use a sentence like, "In this situation, identify passive-looking behaviors and restless vocal patterns, and suggest specific suggestions for what to do to the store staff. It is important to consider the emotional state of both the customer and the pet."

[0187] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0188] Step 1:

[0189] The terminal uses a camera and microphone to capture the movements and sounds of customers and their pets in real time as they enter the store. During this process, the posture and voices of the customers and pets are captured as input. This data is recorded as image and audio data and sent to a server.

[0190] Step 2:

[0191] The server receives data sent from the terminal and performs analysis using a machine learning model. The input consists of image and audio data, which is used to analyze animal movements and vocalizations. Motion analysis and emotion analysis methods are used for data processing to identify individual emotional states. As a result, emotion prediction data for both animals and customers is obtained.

[0192] Step 3:

[0193] The server processes sentiment inference data and generates optimal customer service suggestions using a generative AI model. Sentiment inference data is used as input, and based on this information, the server calculates what action suggestions are appropriate and outputs suggestions in natural language. This output will be notified to the terminal in a later step.

[0194] Step 4:

[0195] The terminal receives customer service suggestions in natural language from the server and presents them to store employees visually or audibly. Employees use this as a reference to perform specific customer service actions. At this time, the suggested message is displayed or played back on the terminal in a human-readable format.

[0196] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0197] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0198] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0199] [Second Embodiment]

[0200] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0201] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0202] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0203] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0204] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0205] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0206] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0207] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0208] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0209] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0210] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0211] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0212] This invention is a system that recognizes and analyzes the movements and sounds of animals to infer their intentions and emotions, and notifies the user of those intentions. The following processes are necessary to implement this system.

[0213] The device is equipped with a camera and microphone to record the animal's movements and sounds, and uses this to acquire real-time data on the pet. Users record their pet's movements and sounds through this device and send that data to the system.

[0214] The server receives motion and audio data transmitted from the terminal and analyzes this data in detail using an analysis tool. The analysis tool incorporates machine learning algorithms to extract patterns of animal vocalizations and characteristics of their movements, and based on these, infers the animal's emotions and intentions.

[0215] The server then generates a message in natural language based on the inferred intentions and emotions of the animal and sends it to the terminal. This message is provided in a format that the user can intuitively understand, helping to gain a deeper understanding of the animal's condition.

[0216] Users can take appropriate action regarding their pets based on the information they receive. They can also support continuous performance improvements by sending feedback to the system as needed.

[0217] As a concrete example, consider a scenario where a user's dog suddenly starts barking intensely. The user uses their smartphone to record this situation and sends the data to the system. The server analyzes the data, evaluating the barking patterns and the dog's body movements to infer that the dog is feeling anxious. As a result, a message is sent to the user's device stating, "The dog appears anxious. Please move it to a quiet place to calm it down." Based on this information, the user can move the dog to a quiet place and address the problem by providing reassurance.

[0218] The following describes the processing flow.

[0219] Step 1:

[0220] The device collects animal movements and sounds using a camera and microphone, and stores this data digitally. The collected data includes video and audio data.

[0221] Step 2:

[0222] The terminal performs preprocessing, including noise reduction and data compression, before transmitting the digital data to the server via the internet. Streaming technology is used for stable data transmission.

[0223] Step 3:

[0224] The server processes the received data using analysis tools. For audio data, it applies a speech recognition algorithm to identify patterns in vocalizations, and for behavioral data, it uses image recognition technology to analyze the characteristics of the movements.

[0225] Step 4:

[0226] The server infers the animal's intentions and emotions based on the analysis results. In doing so, it refers to a machine learning model and determines the most likely intentions and emotions based on past learning results.

[0227] Step 5:

[0228] The server generates natural language messages based on inferred intentions and emotions, and formats them in a way that is easy for the user to understand. These messages are then notified to the user in real time.

[0229] Step 6:

[0230] The terminal receives messages sent from the server and displays them to the user. Based on this information, the user can take appropriate action regarding the animal.

[0231] Step 7:

[0232] Users can contribute to improving the system's accuracy by sending feedback to the server as needed. This feedback includes information about the user's observations and the accuracy of the system's predictions.

[0233] (Example 1)

[0234] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0235] In recent years, there has been a growing need to understand animal emotions and intentions. However, conventional methods require specialized knowledge for analysis, making it difficult for ordinary users to intuitively grasp the state of an animal. Furthermore, there is a need for a system that can analyze animal movements and vocalizations in real time and take quick and appropriate action.

[0236] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0237] In this invention, the server includes a recording device for recording the movements and sounds of an animal, an analysis device for analyzing the recorded movements and sounds using a machine learning algorithm that utilizes a generative AI model, and an inference device for inferring the animal's intentions and emotions based on the analysis results of the analysis device. This makes it possible to precisely analyze the movements and sounds of an animal and immediately notify the user with an intuitive natural language message.

[0238] "Animals" refer to a group of multicellular organisms belonging to the animal kingdom in biological classification, and are organisms that may exhibit emotions and intentions.

[0239] "Movement" refers to a series of movements or changes in movement that occur when an animal moves its body.

[0240] "Sound" refers to the vocalizations and other sounds made by animals.

[0241] A "recording device" is a device used to collect animal movements and sounds and store them as data.

[0242] A "generative AI model" is a type of algorithm designed to learn the characteristics of data and perform specific tasks, and is used for data analysis and prediction.

[0243] A "machine learning algorithm" is a computational method that learns patterns from data and uses those patterns to make predictions or perform classifications.

[0244] An "analytical device" is a device that performs calculations to analyze recorded data and understand its contents.

[0245] A "prediction device" is a device that identifies an animal's intentions and emotions from analyzed data.

[0246] "Natural language" usually refers to the forms of language that humans use on a daily basis, and is closer to actual conversation and writing than to language generated by computers.

[0247] A "notification device" is a device that communicates the results of analysis and estimation to the user, and its role is to provide information.

[0248] This system provides a comprehensive solution for analyzing animal emotions and intentions and notifying the user. A detailed embodiment of the system is described below.

[0249] The device is equipped with a camera and microphone to record animal movements and sounds. This allows the device to collect animal movements as video and vocalizations and other sounds as audio data in real time. Users can monitor their pets' activities using mobile devices such as smartphones and tablets. This allows users to instantly detect any abnormal behavior or changes in the animal's condition.

[0250] The server uses machine learning algorithms powered by generative AI models to analyze the behavioral and audio data transmitted from the terminal. The server receives this data, processes the audio data using spectrogram analysis techniques, and vectorizes the behavioral features by analyzing the video data frame by frame. This analysis enables highly accurate prediction of the animal's emotions and intentions. The generative AI model associates specific behavioral patterns of the animal with their emotions and intentions.

[0251] The server translates the animal's emotions and intentions into natural language based on the analysis results. This ensures that the generated messages are presented to the user in an intuitively understandable format. This process includes the automatic generation of messages such as, "The dog appears anxious. Please move it to a quiet place to calm it down," using a generative AI model.

[0252] For example, if a user's dog suddenly starts barking violently, the user records this situation via their device. The server analyzes this data, evaluating the barking pattern and the dog's body movements to infer that the dog is feeling anxious. As a result, the user receives a message via their device stating, "Your dog appears anxious. Please move it to a quiet place to calm it down." This allows the user to take intuitive action.

[0253] An example of a prompt for a generative AI model might be, "Analyze this dog's barking and behavioral data to infer its emotions and intentions." Following this prompt, the server analyzes the data and infers its emotions and intentions.

[0254] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0255] Step 1:

[0256] The device records animal movements as video and collects vocalizations and sounds as audio data. It uses a camera and microphone to detect animal movements and sounds in real time as input. This input is converted into digital data by the recording device. For example, a dog barking is recorded, and its audio is saved as a file.

[0257] Step 2:

[0258] The terminal sends collected motion data (video files) and audio data (audio files) to the server. This transmission is performed using a secure communication protocol, ensuring data confidentiality. The input is recorded digital data, and the output is the data file sent to the server.

[0259] Step 3:

[0260] The server analyzes the motion and audio data received from the terminal. The input consists of video and audio data, and the server's analysis method is a machine learning algorithm using a generative AI model. Features are extracted from the audio data through spectrogram analysis, and the video data is analyzed frame by frame, with motion features vectorized. The output is the analyzed feature data.

[0261] Step 4:

[0262] The server performs a process of inferring the animal's intentions and emotions based on the analyzed feature data. The input is the feature data obtained in the previous step, and the server's inference method associates specific patterns with emotions and intentions. The output is the inferred emotion and intention data.

[0263] Step 5:

[0264] The server generates natural language messages based on inferred data. The input is inferred intent and emotion data, and the output is an intuitively understandable natural language message. A generative AI model is used to generate messages that help users perceive and respond to their dog's anxiety.

[0265] Step 6:

[0266] The server sends the generated natural language message to the terminal. The input is message data, and the output is the message sent to the terminal. Specifically, the message "The dog appears to be feeling anxious." is generated and a notification is sent to the user's terminal.

[0267] (Application Example 1)

[0268] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0269] Within factories, there is a need to quickly detect safety risks caused by animal intrusion and strengthen safety management. Conventional systems have difficulty accurately assessing the emotions and intentions of animals, making it difficult to take immediate action. This can lead to unexpected accidents and decreased production efficiency.

[0270] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0271] In this invention, the server includes recording means for acquiring animal movements and sounds, analysis means for analyzing the acquired movements and sounds, emotion estimation means for inferring the animal's intentions and emotions based on the analysis results of the analysis means, information provision means for notifying the user of the estimation results by the emotion estimation means, and monitoring and management means for supporting safety management within the factory using the animal emotion estimation results. This makes it possible to improve safety by monitoring animal behavior within the factory in real time and providing appropriate information immediately.

[0272] "Recording means" refers to devices and methods for acquiring animal movements and sounds, and provides a foundation for understanding animal behavior in real time.

[0273] "Analysis means" refers to devices and methods for analyzing acquired animal behavior and sounds, which use machine learning algorithms to extract and analyze data features.

[0274] "Emotion estimation means" refers to a device or method for inferring an animal's intentions and emotions based on the analysis results obtained by the aforementioned analysis means.

[0275] "Information provision means" refers to devices or methods for notifying the user of the inference results obtained by emotion estimation means, and provides the information by converting it into natural language.

[0276] "Monitoring and management means" refers to devices and methods for managing safety within a factory using the results of animal emotion estimation, and are intended to improve safety within the factory.

[0277] To realize this invention, a system is constructed for acquiring and analyzing animal behavior and sounds. The system mainly consists of recording means, analysis means, emotion estimation means, information provision means, and monitoring and management means.

[0278] The server uses cameras and microphones mounted on robots that autonomously patrol the factory to record animal movements and sounds. This functions as a recording tool. The acquired data is sent to the server and analyzed using software such as Python and TensorFlow. Through this analysis, machine learning algorithms extract animal movement patterns and vocal features, and based on this, infer the animal's intentions and emotions.

[0279] The analyzed and inferred results are converted by the emotion estimation means into natural language that can be intuitively understood by the user. The information providing means notifies the administrator of this, enabling prompt response at the site. As a specific example of this notification process, when a night security robot in a factory detects the intrusion of a suspicious animal, it generates an alert such as "A suspicious animal has intruded and is showing aggressive behavior. Please be careful." and transmits it to the administrator.

[0280] Furthermore, the monitoring and management means coordinates the inferred results with the factory's safety management system to automate the alarm trigger and ensure higher safety standards.

[0281] An example of a prompt sentence for a specific generation AI model is "Please analyze this dataset, identify the intentions and emotions of the animals, and generate a warning sentence in natural language based on the results." This prompt sentence is input to the generation AI model when analyzing the acquired data and is used to obtain appropriate analysis results.

[0282] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0283] Step 1:

[0284] The terminal uses cameras and microphones mounted on robots moving within the factory to record the movements and voices of animals in real time. The input is camera video data and audio data, and these data are transmitted to the server as output.

[0285] Step 2:

[0286] The server receives the video and audio data provided by the terminal and prepares the data for analysis. The input is the received data, and the output is the formatted data after noise removal and signal processing.

[0287] Step 3:

[0288] The server utilizes software such as Python and TensorFlow to apply advanced machine learning algorithms to the formatted data and perform analysis. This analysis extracts animal behavior patterns and vocal features, which are then used as input. The output is the feature extraction results.

[0289] Step 4:

[0290] The server infers the animal's intentions and emotions based on the extracted features. It uses a generative AI model to analyze the data through prompt messages. The input at this stage is the feature extraction results, and the output is the inferred intentions and emotions of the animal.

[0291] Step 5:

[0292] The server converts the emotion estimation results into a message expressed in natural language and notifies the user using an information delivery method. An example of a prompt message generated by the system using a generative AI model is: "Analyze this dataset to identify the animal's intentions and emotions, and generate a natural language warning message based on the results." The input is the emotion estimation result, and the output is a natural language message.

[0293] Step 6:

[0294] The user checks the message notified on the terminal and takes appropriate countermeasures as needed, depending on the situation in the factory. The input is a natural language message, and the output is the implementation of the countermeasure.

[0295] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0296] This invention is a system that facilitates communication by simultaneously analyzing the emotions of both the animal and the user. In addition to functions for recognizing and analyzing the animal's movements and sounds, this system integrates an emotion engine that recognizes emotions from the user's facial expressions and voice. The following describes embodiments for carrying out this invention.

[0297] The device is equipped with a camera and microphone to record animal movements and sounds, thereby acquiring real-time data. Similarly, it also has a camera and microphone to record the user's facial expressions and tone of voice. This allows the device to simultaneously acquire data from both the animal and the user.

[0298] The server processes animal data transmitted from the terminal using analysis tools to infer the animal's intentions and emotions. The analysis tools utilize machine learning algorithms and have already learned the individual vocal patterns and behavioral characteristics of each animal. Simultaneously, an emotion engine included in the server analyzes user data to infer the user's emotions. This engine can identify the user's current emotional state through facial recognition and voice analysis technologies.

[0299] The server then uses an analysis combining the animal's intentions and emotions with the user's emotions to generate an appropriate message. This message is adjusted in content and tone according to the user's current emotional state, ensuring the most effective communication.

[0300] For example, if the user is feeling stressed and the dog is showing signs of anxiety, the server will generate a message such as, "It's important for you to relax while also making your dog feel secure," and notify the user's device. Based on this information, the user can take appropriate action that suits both their and their dog's emotional state.

[0301] By receiving these notifications, users can gain a deeper understanding of the animal's condition and take appropriate actions considering each other's feelings. Also, information for improving the analysis accuracy of the system can be provided to the server through the feedback function. This feedback plays an important role in effective communication.

[0302] The processing flow will be described below.

[0303] Step 1:

[0304] The terminal acquires data using a camera and a microphone that record the animal's movements and sounds. At the same time, the user's facial expression and tone-of-voice data are collected. This data is temporarily stored in the terminal.

[0305] Step 2:

[0306] The terminal preprocesses the acquired data, removes noise, and extracts only the necessary information. The preprocessed data is compressed and transmitted to the server via communication.

[0307] Step 3:

[0308] The server processes the received animal movement and sound data by means of analysis. Here, a machine learning algorithm is utilized to extract the patterns of the animal's cries and movements.

[0309] Step 4:

[0310] The server analyzes the user's facial expression and tone-of-voice data using an emotion engine. The algorithm estimates the user's emotional state and records that information in the database.

[0311] Step 5:

[0312] The server combines inferred animal intentions and emotions with the user's emotional state to generate the optimal message. This message is then adjusted to suit the user's emotional state.

[0313] Step 6:

[0314] The server translates the message into natural language and sends it to the terminal. This message is designed to be intuitively understandable to the user.

[0315] Step 7:

[0316] The device receives messages sent from the server and notifies the user. Notifications are made via audio alerts or screen displays.

[0317] Step 8:

[0318] Users can respond to the information they receive in accordance with the emotional state of the animal and themselves. They can also use the feedback function to provide additional information to the server, which can be used for future analysis.

[0319] (Example 2)

[0320] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0321] In recent years, there has been a growing demand for technologies that facilitate communication between humans and animals. However, conventional technologies focus solely on analyzing animal movements and vocalizations, lacking consideration for the user's emotional state. Therefore, there is a need for a means to simultaneously analyze the animal's intentions and emotions, as well as the user's emotions, and to achieve appropriate two-way communication based on this analysis.

[0322] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0323] In this invention, the server includes a device for acquiring the animal's movements and sounds, a device for acquiring the user's facial expressions and sounds, and a processing device for analyzing the acquired movements, sounds, and facial expressions. This enables communication that simultaneously analyzes the emotions of the animal and the user, and generates appropriate two-way messages based on the results.

[0324] A "device" is a combination of hardware and software used to perform a specific function.

[0325] "Action" refers to a physical activity or action performed by a living being.

[0326] "Sound" refers to the waveform of sounds emitted by animals and humans, and by analyzing its characteristics, it is possible to infer emotions and intentions.

[0327] "Facial expression" is visual information that conveys emotions and intentions through the movement of facial muscles.

[0328] A "processing device" is a device that takes data as input, analyzes it, and has the computational function to extract or infer specific information.

[0329] "Analysis" is the process of breaking down, comparing, and evaluating data to understand its meaning and relationships.

[0330] A "prediction device" is a device that has a computational function to infer the intentions and emotions of a certain subject based on collected data.

[0331] A "notification device" is a device that has the function of conveying predicted results or generated messages to the user.

[0332] "Natural language" refers to the language that humans use on a daily basis, and the text that is generated in a format suitable for computer processing.

[0333] This invention is a system that simultaneously analyzes the emotions of both animals and their users. The system aims to facilitate communication between the two by combining functions that analyze the animal's movements and sounds with functions that analyze the user's facial expressions and sounds.

[0334] The device is equipped with a camera and microphone to record the animal's movements and sounds. This hardware makes it possible to acquire real-time data on the animal's physical movements and vocalizations. The same device also has a camera and microphone to record the user's facial expressions and voice, allowing for the simultaneous acquisition of the user's emotional data.

[0335] The server analyzes animal and user data transmitted from the terminal. This analysis utilizes machine learning algorithms that analyze animal sounds and movements, as well as the user's facial expressions and tone of voice. This allows the server to accurately predict the animal's emotions and intentions, as well as the user's emotions. Various server-based analysis software and emotion engines are used in this analysis.

[0336] Based on the analysis results, the server generates messages tailored to the animal's and the user's situation. These messages are written in natural language, adjusted to the most effective content and tone, and delivered to the device. Based on this information, users can take appropriate actions that align with the emotions of both the animal and themselves.

[0337] As a concrete example, considering the example prompt, in response to the input, "Please advise how the user should react when the dog appears anxious," the system can provide appropriate advice. This realizes the use of generative AI models to support communication between animals and humans.

[0338] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0339] Step 1:

[0340] The device activates its camera and microphone to capture animal movements and sounds, as well as the user's facial expressions and voice, in real time. The acquired data includes animal movements and vocalizations, as well as the user's facial expressions and voice tone. This data is initially processed within the device for later detailed analysis.

[0341] Step 2:

[0342] The terminal sends the initially processed data to the server. Wireless communication technologies such as Wi-Fi and Bluetooth are used for this communication. The input data includes animal movement data, voice data, user facial expression data, and voice data, which are used as basic data for the next analysis step on the server.

[0343] Step 3:

[0344] The server receives the transmitted data and uses machine learning algorithms to analyze the animal's movements and sounds, as well as the user's facial expressions and voice. Based on the animal data, it infers the animal's intentions and emotions from its behavioral patterns and vocal characteristics. Regarding the user's data, it identifies emotions from changes in facial expressions and tone of voice. As a result of the analysis, the emotional states of both the animal and the user are output.

[0345] Step 4:

[0346] The server uses a generative AI model to generate appropriate messages based on the analysis of the animal's intentions and emotions, as well as the user's emotions. As a prompt, the analysis results are input into the generative AI model, which then outputs an effective message in natural language. This message is structured as specific advice that is beneficial to both the animal and the user.

[0347] Step 5:

[0348] The server sends the generated message to the terminal. The terminal notifies the user of this message either visually or audibly. Based on the received message, the user can take specific actions to adjust their communication with the animal.

[0349] (Application Example 2)

[0350] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0351] In stores that allow pets, there is a challenge in appropriately analyzing the emotional states of both customers and their pets and enabling store staff to take the most appropriate action based on that analysis. Therefore, there is a need for effective support to improve customer satisfaction and pet comfort.

[0352] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0353] In this invention, the server includes acquisition means for acquiring animal behavior and sounds, emotion analysis means for simultaneously analyzing the emotional states of customers and animals, and motion analysis means for analyzing the acquired behavior and sounds. This makes it possible to analyze the emotions of customers and pets with high accuracy and to provide appropriate customer service suggestions to store employees.

[0354] "Acquisition means" refers to equipment and technology for detecting animal behavior and sounds and converting them into data.

[0355] "Emotional analysis means" refers to technology or devices for simultaneously analyzing the emotional states of customers and animals. This means has the function of analyzing data using machine learning and emotion recognition algorithms to identify emotions.

[0356] "Motion analysis means" refers to technology or equipment used to analyze acquired animal behavior and vocalizations. This means has the function of processing data in order to understand the meaning and intent of the behavior.

[0357] "Inference means" refers to techniques or devices used to infer the emotional state or intentions of a subject based on analysis results. These means play a role in inferring the most likely emotional state based on the analyzed data.

[0358] "Proposal provision means" refers to technology or equipment for presenting specific actions or responses to external parties based on the prediction results obtained by prediction means. This means is used to instruct the optimal action based on the information obtained.

[0359] One embodiment of this invention provides a system that understands the emotions of customers accompanied by pets and their pets in a store, and enables optimal customer service. The system mainly consists of terminals and a server.

[0360] The terminal is equipped with a camera and microphone that record the movements of customers and their pets in real time, thereby collecting motion and audio data of customers and pets when they enter the store. The terminal is responsible for transmitting the collected data to a server.

[0361] The server uses machine learning algorithms to analyze animal behavior and customer emotions based on the received data. To analyze animal and customer emotions simultaneously, emotion analysis means are used to identify their respective emotional states. Motion analysis means are used to analyze animal behavior in detail and infer the animal's intentions and emotions.

[0362] Based on the insights gained from these analysis results, the server generates suggestions to facilitate optimal customer service. These suggestions are translated into natural language by a suggestion delivery system and notified to store employees. As a result, employees can provide optimal customer service tailored to the customer and their pet's condition, thereby improving customer satisfaction.

[0363] For example, if the system detects that a customer entering the store appears tense and their pet is frightened, it will generate a natural language suggestion such as "Let's give your pet some time to calm down" and provide it to the store staff via a terminal. This allows employees to provide prompt and appropriate customer service.

[0364] As an example of a prompt to a generative AI model, you could use a sentence like, "In this situation, identify passive-looking behaviors and restless vocal patterns, and suggest specific suggestions for what to do to the store staff. It is important to consider the emotional state of both the customer and the pet."

[0365] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0366] Step 1:

[0367] The terminal uses a camera and microphone to capture the movements and sounds of customers and their pets in real time as they enter the store. During this process, the posture and voices of the customers and pets are captured as input. This data is recorded as image and audio data and sent to a server.

[0368] Step 2:

[0369] The server receives data sent from the terminal and performs analysis using a machine learning model. The input consists of image and audio data, which is used to analyze animal movements and vocalizations. Motion analysis and emotion analysis methods are used for data processing to identify individual emotional states. As a result, emotion prediction data for both animals and customers is obtained.

[0370] Step 3:

[0371] The server processes sentiment inference data and generates optimal customer service suggestions using a generative AI model. Sentiment inference data is used as input, and based on this information, the server calculates what action suggestions are appropriate and outputs suggestions in natural language. This output will be notified to the terminal in a later step.

[0372] Step 4:

[0373] The terminal receives customer service suggestions in natural language from the server and presents them to store employees visually or audibly. Employees use this as a reference to perform specific customer service actions. At this time, the suggested message is displayed or played back on the terminal in a human-readable format.

[0374] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0375] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0376] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0377] [Third Embodiment]

[0378] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0379] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0380] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0381] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0382] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0383] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0384] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0385] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0386] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0387] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0388] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0389] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0390] This invention is a system that recognizes and analyzes the movements and sounds of animals to infer their intentions and emotions, and notifies the user of those intentions. The following processes are necessary to implement this system.

[0391] The device is equipped with a camera and microphone to record the animal's movements and sounds, and uses this to acquire real-time data on the pet. Users record their pet's movements and sounds through this device and send that data to the system.

[0392] The server receives motion and audio data transmitted from the terminal and analyzes this data in detail using an analysis tool. The analysis tool incorporates machine learning algorithms to extract patterns of animal vocalizations and characteristics of their movements, and based on these, infers the animal's emotions and intentions.

[0393] The server then generates a message in natural language based on the inferred intentions and emotions of the animal and sends it to the terminal. This message is provided in a format that the user can intuitively understand, helping to gain a deeper understanding of the animal's condition.

[0394] Users can take appropriate action regarding their pets based on the information they receive. They can also support continuous performance improvements by sending feedback to the system as needed.

[0395] As a concrete example, consider a scenario where a user's dog suddenly starts barking intensely. The user uses their smartphone to record this situation and sends the data to the system. The server analyzes the data, evaluating the barking patterns and the dog's body movements to infer that the dog is feeling anxious. As a result, a message is sent to the user's device stating, "The dog appears anxious. Please move it to a quiet place to calm it down." Based on this information, the user can move the dog to a quiet place and address the problem by providing reassurance.

[0396] The following describes the processing flow.

[0397] Step 1:

[0398] The device collects animal movements and sounds using a camera and microphone, and stores this data digitally. The collected data includes video and audio data.

[0399] Step 2:

[0400] The terminal performs preprocessing, including noise reduction and data compression, before transmitting the digital data to the server via the internet. Streaming technology is used for stable data transmission.

[0401] Step 3:

[0402] The server processes the received data using analysis tools. For audio data, it applies a speech recognition algorithm to identify patterns in vocalizations, and for behavioral data, it uses image recognition technology to analyze the characteristics of the movements.

[0403] Step 4:

[0404] The server infers the animal's intentions and emotions based on the analysis results. In doing so, it refers to a machine learning model and determines the most likely intentions and emotions based on past learning results.

[0405] Step 5:

[0406] The server generates natural language messages based on inferred intentions and emotions, and formats them in a way that is easy for the user to understand. These messages are then notified to the user in real time.

[0407] Step 6:

[0408] The terminal receives messages sent from the server and displays them to the user. Based on this information, the user can take appropriate action regarding the animal.

[0409] Step 7:

[0410] Users can contribute to improving the system's accuracy by sending feedback to the server as needed. This feedback includes information about the user's observations and the accuracy of the system's predictions.

[0411] (Example 1)

[0412] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0413] In recent years, there has been a growing need to understand animal emotions and intentions. However, conventional methods require specialized knowledge for analysis, making it difficult for ordinary users to intuitively grasp the state of an animal. Furthermore, there is a need for a system that can analyze animal movements and vocalizations in real time and take quick and appropriate action.

[0414] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0415] In this invention, the server includes a recording device for recording the movements and sounds of an animal, an analysis device for analyzing the recorded movements and sounds using a machine learning algorithm that utilizes a generative AI model, and an inference device for inferring the animal's intentions and emotions based on the analysis results of the analysis device. This makes it possible to precisely analyze the movements and sounds of an animal and immediately notify the user with an intuitive natural language message.

[0416] "Animals" refer to a group of multicellular organisms belonging to the animal kingdom in biological classification, and are organisms that may exhibit emotions and intentions.

[0417] "Movement" refers to a series of movements or changes in movement that occur when an animal moves its body.

[0418] "Sound" refers to the vocalizations and other sounds made by animals.

[0419] A "recording device" is a device used to collect animal movements and sounds and store them as data.

[0420] A "generative AI model" is a type of algorithm designed to learn the characteristics of data and perform specific tasks, and is used for data analysis and prediction.

[0421] A "machine learning algorithm" is a computational method that learns patterns from data and uses those patterns to make predictions or perform classifications.

[0422] An "analytical device" is a device that performs calculations to analyze recorded data and understand its contents.

[0423] A "prediction device" is a device that identifies an animal's intentions and emotions from analyzed data.

[0424] "Natural language" usually refers to the forms of language that humans use on a daily basis, and is closer to actual conversation and writing than to language generated by computers.

[0425] A "notification device" is a device that communicates the results of analysis and estimation to the user, and its role is to provide information.

[0426] This system provides a comprehensive solution for analyzing animal emotions and intentions and notifying the user. A detailed embodiment of the system is described below.

[0427] The device is equipped with a camera and microphone to record animal movements and sounds. This allows the device to collect animal movements as video and vocalizations and other sounds as audio data in real time. Users can monitor their pets' activities using mobile devices such as smartphones and tablets. This allows users to instantly detect any abnormal behavior or changes in the animal's condition.

[0428] The server uses machine learning algorithms powered by generative AI models to analyze the behavioral and audio data transmitted from the terminal. The server receives this data, processes the audio data using spectrogram analysis techniques, and vectorizes the behavioral features by analyzing the video data frame by frame. This analysis enables highly accurate prediction of the animal's emotions and intentions. The generative AI model associates specific behavioral patterns of the animal with their emotions and intentions.

[0429] The server translates the animal's emotions and intentions into natural language based on the analysis results. This ensures that the generated messages are presented to the user in an intuitively understandable format. This process includes the automatic generation of messages such as, "The dog appears anxious. Please move it to a quiet place to calm it down," using a generative AI model.

[0430] For example, if a user's dog suddenly starts barking violently, the user records this situation via their device. The server analyzes this data, evaluating the barking pattern and the dog's body movements to infer that the dog is feeling anxious. As a result, the user receives a message via their device stating, "Your dog appears anxious. Please move it to a quiet place to calm it down." This allows the user to take intuitive action.

[0431] An example of a prompt for a generative AI model might be, "Analyze this dog's barking and behavioral data to infer its emotions and intentions." Following this prompt, the server analyzes the data and infers its emotions and intentions.

[0432] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0433] Step 1:

[0434] The device records animal movements as video and collects vocalizations and sounds as audio data. It uses a camera and microphone to detect animal movements and sounds in real time as input. This input is converted into digital data by the recording device. For example, a dog barking is recorded, and its audio is saved as a file.

[0435] Step 2:

[0436] The terminal sends collected motion data (video files) and audio data (audio files) to the server. This transmission is performed using a secure communication protocol, ensuring data confidentiality. The input is recorded digital data, and the output is the data file sent to the server.

[0437] Step 3:

[0438] The server analyzes the motion and audio data received from the terminal. The input consists of video and audio data, and the server's analysis method is a machine learning algorithm using a generative AI model. Features are extracted from the audio data through spectrogram analysis, and the video data is analyzed frame by frame, with motion features vectorized. The output is the analyzed feature data.

[0439] Step 4:

[0440] The server performs a process of inferring the animal's intentions and emotions based on the analyzed feature data. The input is the feature data obtained in the previous step, and the server's inference method associates specific patterns with emotions and intentions. The output is the inferred emotion and intention data.

[0441] Step 5:

[0442] The server generates natural language messages based on inferred data. The input is inferred intent and emotion data, and the output is an intuitively understandable natural language message. A generative AI model is used to generate messages that help users perceive and respond to their dog's anxiety.

[0443] Step 6:

[0444] The server sends the generated natural language message to the terminal. The input is message data, and the output is the message sent to the terminal. Specifically, the message "The dog appears to be feeling anxious." is generated and a notification is sent to the user's terminal.

[0445] (Application Example 1)

[0446] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0447] Within factories, there is a need to quickly detect safety risks caused by animal intrusion and strengthen safety management. Conventional systems have difficulty accurately assessing the emotions and intentions of animals, making it difficult to take immediate action. This can lead to unexpected accidents and decreased production efficiency.

[0448] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0449] In this invention, the server includes recording means for acquiring animal movements and sounds, analysis means for analyzing the acquired movements and sounds, emotion estimation means for inferring the animal's intentions and emotions based on the analysis results of the analysis means, information provision means for notifying the user of the estimation results by the emotion estimation means, and monitoring and management means for supporting safety management within the factory using the animal emotion estimation results. This makes it possible to improve safety by monitoring animal behavior within the factory in real time and providing appropriate information immediately.

[0450] "Recording means" refers to devices and methods for acquiring animal movements and sounds, and provides a foundation for understanding animal behavior in real time.

[0451] "Analysis means" refers to devices and methods for analyzing acquired animal behavior and sounds, which use machine learning algorithms to extract and analyze data features.

[0452] "Emotion estimation means" refers to a device or method for inferring an animal's intentions and emotions based on the analysis results obtained by the aforementioned analysis means.

[0453] "Information provision means" refers to devices or methods for notifying the user of the inference results obtained by emotion estimation means, and provides the information by converting it into natural language.

[0454] "Monitoring and management means" refers to devices and methods for managing safety within a factory using the results of animal emotion estimation, and are intended to improve safety within the factory.

[0455] To realize this invention, a system is constructed for acquiring and analyzing animal behavior and sounds. The system mainly consists of recording means, analysis means, emotion estimation means, information provision means, and monitoring and management means.

[0456] The server uses cameras and microphones mounted on robots that autonomously patrol the factory to record animal movements and sounds. This functions as a recording tool. The acquired data is sent to the server and analyzed using software such as Python and TensorFlow. Through this analysis, machine learning algorithms extract animal movement patterns and vocal features, and based on this, infer the animal's intentions and emotions.

[0457] The analyzed and inferred results are converted into natural language that users can intuitively understand by sentiment estimation tools. This information is then communicated to administrators by information providers, enabling a rapid response on-site. A concrete example of this notification process is when a night security robot at a factory detects the intrusion of a suspicious animal; it generates an alert stating, "A suspicious animal has entered the premises and is exhibiting aggressive behavior. Please be careful," and sends it to the administrator.

[0458] Furthermore, the monitoring and management system integrates the prediction results with the factory's safety management system, automating alarm triggers and ensuring higher safety standards.

[0459] A concrete example of a prompt for a generative AI model is, "Analyze this dataset, identify the intentions and emotions of the animals, and generate a natural language warning message based on the results." This prompt is input to the generative AI model when analyzing the acquired data and is used to obtain appropriate analysis results.

[0460] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0461] Step 1:

[0462] The terminal uses cameras and microphones mounted on robots moving around the factory to record animal movements and sounds in real time. The input consists of camera video data and audio data, and this data is sent to a server as output.

[0463] Step 2:

[0464] The server receives video and audio data provided by the terminal and prepares the data for analysis. The input is the received data, and the output is the formatted data after noise reduction and signal processing.

[0465] Step 3:

[0466] The server utilizes software such as Python and TensorFlow to apply advanced machine learning algorithms to the formatted data and perform analysis. This analysis extracts animal behavior patterns and vocal features, which are then used as input. The output is the feature extraction results.

[0467] Step 4:

[0468] The server infers the animal's intentions and emotions based on the extracted features. It uses a generative AI model to analyze the data through prompt messages. The input at this stage is the feature extraction results, and the output is the inferred intentions and emotions of the animal.

[0469] Step 5:

[0470] The server converts the emotion estimation results into a message expressed in natural language and notifies the user using an information delivery method. An example of a prompt message generated by the system using a generative AI model is: "Analyze this dataset to identify the animal's intentions and emotions, and generate a natural language warning message based on the results." The input is the emotion estimation result, and the output is a natural language message.

[0471] Step 6:

[0472] The user checks the message notified on the terminal and takes appropriate countermeasures as needed, depending on the situation in the factory. The input is a natural language message, and the output is the implementation of the countermeasure.

[0473] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0474] This invention is a system that facilitates communication by simultaneously analyzing the emotions of both the animal and the user. In addition to functions for recognizing and analyzing the animal's movements and sounds, this system integrates an emotion engine that recognizes emotions from the user's facial expressions and voice. The following describes embodiments for carrying out this invention.

[0475] The device is equipped with a camera and microphone to record animal movements and sounds, thereby acquiring real-time data. Similarly, it also has a camera and microphone to record the user's facial expressions and tone of voice. This allows the device to simultaneously acquire data from both the animal and the user.

[0476] The server processes animal data transmitted from the terminal using analysis tools to infer the animal's intentions and emotions. The analysis tools utilize machine learning algorithms and have already learned the individual vocal patterns and behavioral characteristics of each animal. Simultaneously, an emotion engine included in the server analyzes user data to infer the user's emotions. This engine can identify the user's current emotional state through facial recognition and voice analysis technologies.

[0477] The server then uses an analysis combining the animal's intentions and emotions with the user's emotions to generate an appropriate message. This message is adjusted in content and tone according to the user's current emotional state, ensuring the most effective communication.

[0478] For example, if the user is feeling stressed and the dog is showing signs of anxiety, the server will generate a message such as, "It's important for you to relax while also making your dog feel secure," and notify the user's device. Based on this information, the user can take appropriate action that suits both their and their dog's emotional state.

[0479] By receiving these notifications, users can gain a deeper understanding of the animals' conditions and take appropriate actions that are considerate of each other's feelings. They can also provide information to the server through the feedback function to improve the accuracy of the system's analysis. This feedback plays a crucial role in effective communication.

[0480] The following describes the processing flow.

[0481] Step 1:

[0482] The device acquires data using a camera and microphone to record animal movements and sounds. Simultaneously, it collects data on the user's facial expressions and tone of voice. This data is temporarily stored within the device.

[0483] Step 2:

[0484] The terminal preprocesses the acquired data, removing noise and extracting only the necessary information. The preprocessed data is then compressed and sent to the server via communication.

[0485] Step 3:

[0486] The server processes the received animal behavior and audio data using analysis tools. Machine learning algorithms are employed to extract patterns in animal sounds and movements.

[0487] Step 4:

[0488] The server uses an emotion engine to analyze the user's facial expressions and tone of voice data. The algorithm infers the user's emotional state and records that information in a database.

[0489] Step 5:

[0490] The server combines inferred animal intentions and emotions with the user's emotional state to generate the optimal message. This message is then adjusted to suit the user's emotional state.

[0491] Step 6:

[0492] The server translates the message into natural language and sends it to the terminal. This message is designed to be intuitively understandable to the user.

[0493] Step 7:

[0494] The device receives messages sent from the server and notifies the user. Notifications are made via audio alerts or screen displays.

[0495] Step 8:

[0496] Users can respond to the information they receive in accordance with the emotional state of the animal and themselves. They can also use the feedback function to provide additional information to the server, which can be used for future analysis.

[0497] (Example 2)

[0498] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0499] In recent years, there has been a growing demand for technologies that facilitate communication between humans and animals. However, conventional technologies focus solely on analyzing animal movements and vocalizations, lacking consideration for the user's emotional state. Therefore, there is a need for a means to simultaneously analyze the animal's intentions and emotions, as well as the user's emotions, and to achieve appropriate two-way communication based on this analysis.

[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0501] In this invention, the server includes a device for acquiring the animal's movements and sounds, a device for acquiring the user's facial expressions and sounds, and a processing device for analyzing the acquired movements, sounds, and facial expressions. This enables communication that simultaneously analyzes the emotions of the animal and the user, and generates appropriate two-way messages based on the results.

[0502] A "device" is a combination of hardware and software used to perform a specific function.

[0503] "Action" refers to a physical activity or action performed by a living being.

[0504] "Sound" refers to the waveform of sounds emitted by animals and humans, and by analyzing its characteristics, it is possible to infer emotions and intentions.

[0505] "Facial expression" is visual information that conveys emotions and intentions through the movement of facial muscles.

[0506] A "processing device" is a device that takes data as input, analyzes it, and has the computational function to extract or infer specific information.

[0507] "Analysis" is the process of breaking down, comparing, and evaluating data to understand its meaning and relationships.

[0508] A "prediction device" is a device that has a computational function to infer the intentions and emotions of a certain subject based on collected data.

[0509] A "notification device" is a device that has the function of conveying predicted results or generated messages to the user.

[0510] "Natural language" refers to the language that humans use on a daily basis, and the text that is generated in a format suitable for computer processing.

[0511] This invention is a system that simultaneously analyzes the emotions of both animals and their users. The system aims to facilitate communication between the two by combining functions that analyze the animal's movements and sounds with functions that analyze the user's facial expressions and sounds.

[0512] The device is equipped with a camera and microphone to record the animal's movements and sounds. This hardware makes it possible to acquire real-time data on the animal's physical movements and vocalizations. The same device also has a camera and microphone to record the user's facial expressions and voice, allowing for the simultaneous acquisition of the user's emotional data.

[0513] The server analyzes animal and user data transmitted from the terminal. This analysis utilizes machine learning algorithms that analyze animal sounds and movements, as well as the user's facial expressions and tone of voice. This allows the server to accurately predict the animal's emotions and intentions, as well as the user's emotions. Various server-based analysis software and emotion engines are used in this analysis.

[0514] Based on the analysis results, the server generates messages tailored to the animal's and the user's situation. These messages are written in natural language, adjusted to the most effective content and tone, and delivered to the device. Based on this information, users can take appropriate actions that align with the emotions of both the animal and themselves.

[0515] As a concrete example, considering the example prompt, in response to the input, "Please advise how the user should react when the dog appears anxious," the system can provide appropriate advice. This realizes the use of generative AI models to support communication between animals and humans.

[0516] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0517] Step 1:

[0518] The device activates its camera and microphone to capture animal movements and sounds, as well as the user's facial expressions and voice, in real time. The acquired data includes animal movements and vocalizations, as well as the user's facial expressions and voice tone. This data is initially processed within the device for later detailed analysis.

[0519] Step 2:

[0520] The terminal sends the initially processed data to the server. Wireless communication technologies such as Wi-Fi and Bluetooth are used for this communication. The input data includes animal movement data, voice data, user facial expression data, and voice data, which are used as basic data for the next analysis step on the server.

[0521] Step 3:

[0522] The server receives the transmitted data and uses machine learning algorithms to analyze the animal's movements and sounds, as well as the user's facial expressions and voice. Based on the animal data, it infers the animal's intentions and emotions from its behavioral patterns and vocal characteristics. Regarding the user's data, it identifies emotions from changes in facial expressions and tone of voice. As a result of the analysis, the emotional states of both the animal and the user are output.

[0523] Step 4:

[0524] The server uses a generative AI model to generate appropriate messages based on the analysis of the animal's intentions and emotions, as well as the user's emotions. As a prompt, the analysis results are input into the generative AI model, which then outputs an effective message in natural language. This message is structured as specific advice that is beneficial to both the animal and the user.

[0525] Step 5:

[0526] The server sends the generated message to the terminal. The terminal notifies the user of this message either visually or audibly. Based on the received message, the user can take specific actions to adjust their communication with the animal.

[0527] (Application Example 2)

[0528] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0529] In stores that allow pets, there is a challenge in appropriately analyzing the emotional states of both customers and their pets and enabling store staff to take the most appropriate action based on that analysis. Therefore, there is a need for effective support to improve customer satisfaction and pet comfort.

[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0531] In this invention, the server includes acquisition means for acquiring animal behavior and sounds, emotion analysis means for simultaneously analyzing the emotional states of customers and animals, and motion analysis means for analyzing the acquired behavior and sounds. This makes it possible to analyze the emotions of customers and pets with high accuracy and to provide appropriate customer service suggestions to store employees.

[0532] "Acquisition means" refers to equipment and technology for detecting animal behavior and sounds and converting them into data.

[0533] "Emotional analysis means" refers to technology or devices for simultaneously analyzing the emotional states of customers and animals. This means has the function of analyzing data using machine learning and emotion recognition algorithms to identify emotions.

[0534] "Motion analysis means" refers to technology or equipment used to analyze acquired animal behavior and vocalizations. This means has the function of processing data in order to understand the meaning and intent of the behavior.

[0535] "Inference means" refers to techniques or devices used to infer the emotional state or intentions of a subject based on analysis results. These means play a role in inferring the most likely emotional state based on the analyzed data.

[0536] "Proposal provision means" refers to technology or equipment for presenting specific actions or responses to external parties based on the prediction results obtained by prediction means. This means is used to instruct the optimal action based on the information obtained.

[0537] One embodiment of this invention provides a system that understands the emotions of customers accompanied by pets and their pets in a store, and enables optimal customer service. The system mainly consists of terminals and a server.

[0538] The terminal is equipped with a camera and microphone that record the movements of customers and their pets in real time, thereby collecting motion and audio data of customers and pets when they enter the store. The terminal is responsible for transmitting the collected data to a server.

[0539] The server uses machine learning algorithms to analyze animal behavior and customer emotions based on the received data. To analyze animal and customer emotions simultaneously, emotion analysis means are used to identify their respective emotional states. Motion analysis means are used to analyze animal behavior in detail and infer the animal's intentions and emotions.

[0540] Based on the insights gained from these analysis results, the server generates suggestions to facilitate optimal customer service. These suggestions are translated into natural language by a suggestion delivery system and notified to store employees. As a result, employees can provide optimal customer service tailored to the customer and their pet's condition, thereby improving customer satisfaction.

[0541] For example, if the system detects that a customer entering the store appears tense and their pet is frightened, it will generate a natural language suggestion such as "Let's give your pet some time to calm down" and provide it to the store staff via a terminal. This allows employees to provide prompt and appropriate customer service.

[0542] As an example of a prompt to a generative AI model, you could use a sentence like, "In this situation, identify passive-looking behaviors and restless vocal patterns, and suggest specific suggestions for what to do to the store staff. It is important to consider the emotional state of both the customer and the pet."

[0543] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0544] Step 1:

[0545] The terminal uses a camera and microphone to capture the movements and sounds of customers and their pets in real time as they enter the store. During this process, the posture and voices of the customers and pets are captured as input. This data is recorded as image and audio data and sent to a server.

[0546] Step 2:

[0547] The server receives data sent from the terminal and performs analysis using a machine learning model. The input consists of image and audio data, which is used to analyze animal movements and vocalizations. Motion analysis and emotion analysis methods are used for data processing to identify individual emotional states. As a result, emotion prediction data for both animals and customers is obtained.

[0548] Step 3:

[0549] The server processes sentiment inference data and generates optimal customer service suggestions using a generative AI model. Sentiment inference data is used as input, and based on this information, the server calculates what action suggestions are appropriate and outputs suggestions in natural language. This output will be notified to the terminal in a later step.

[0550] Step 4:

[0551] The terminal receives customer service suggestions in natural language from the server and presents them to store employees visually or audibly. Employees use this as a reference to perform specific customer service actions. At this time, the suggested message is displayed or played back on the terminal in a human-readable format.

[0552] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0553] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0554] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0555] [Fourth Embodiment]

[0556] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0557] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0558] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0559] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0560] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0561] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0562] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0563] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0564] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0565] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0566] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0567] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0568] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0569] This invention is a system that recognizes and analyzes the movements and sounds of animals to infer their intentions and emotions, and notifies the user of those intentions. The following processes are necessary to implement this system.

[0570] The device is equipped with a camera and microphone to record the animal's movements and sounds, and uses this to acquire real-time data on the pet. Users record their pet's movements and sounds through this device and send that data to the system.

[0571] The server receives motion and audio data transmitted from the terminal and analyzes this data in detail using an analysis tool. The analysis tool incorporates machine learning algorithms to extract patterns of animal vocalizations and characteristics of their movements, and based on these, infers the animal's emotions and intentions.

[0572] The server then generates a message in natural language based on the inferred intentions and emotions of the animal and sends it to the terminal. This message is provided in a format that the user can intuitively understand, helping to gain a deeper understanding of the animal's condition.

[0573] Users can take appropriate action regarding their pets based on the information they receive. They can also support continuous performance improvements by sending feedback to the system as needed.

[0574] As a concrete example, consider a scenario where a user's dog suddenly starts barking intensely. The user uses their smartphone to record this situation and sends the data to the system. The server analyzes the data, evaluating the barking patterns and the dog's body movements to infer that the dog is feeling anxious. As a result, a message is sent to the user's device stating, "The dog appears anxious. Please move it to a quiet place to calm it down." Based on this information, the user can move the dog to a quiet place and address the problem by providing reassurance.

[0575] The following describes the processing flow.

[0576] Step 1:

[0577] The device collects animal movements and sounds using a camera and microphone, and stores this data digitally. The collected data includes video and audio data.

[0578] Step 2:

[0579] The terminal performs preprocessing, including noise reduction and data compression, before transmitting the digital data to the server via the internet. Streaming technology is used for stable data transmission.

[0580] Step 3:

[0581] The server processes the received data using analysis tools. For audio data, it applies a speech recognition algorithm to identify patterns in vocalizations, and for behavioral data, it uses image recognition technology to analyze the characteristics of the movements.

[0582] Step 4:

[0583] The server infers the animal's intentions and emotions based on the analysis results. In doing so, it refers to a machine learning model and determines the most likely intentions and emotions based on past learning results.

[0584] Step 5:

[0585] The server generates natural language messages based on inferred intentions and emotions, and formats them in a way that is easy for the user to understand. These messages are then notified to the user in real time.

[0586] Step 6:

[0587] The terminal receives messages sent from the server and displays them to the user. Based on this information, the user can take appropriate action regarding the animal.

[0588] Step 7:

[0589] Users can contribute to improving the system's accuracy by sending feedback to the server as needed. This feedback includes information about the user's observations and the accuracy of the system's predictions.

[0590] (Example 1)

[0591] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0592] In recent years, there has been a growing need to understand animal emotions and intentions. However, conventional methods require specialized knowledge for analysis, making it difficult for ordinary users to intuitively grasp the state of an animal. Furthermore, there is a need for a system that can analyze animal movements and vocalizations in real time and take quick and appropriate action.

[0593] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0594] In this invention, the server includes a recording device for recording the movements and sounds of an animal, an analysis device for analyzing the recorded movements and sounds using a machine learning algorithm that utilizes a generative AI model, and an inference device for inferring the animal's intentions and emotions based on the analysis results of the analysis device. This makes it possible to precisely analyze the movements and sounds of an animal and immediately notify the user with an intuitive natural language message.

[0595] "Animals" refer to a group of multicellular organisms belonging to the animal kingdom in biological classification, and are organisms that may exhibit emotions and intentions.

[0596] "Movement" refers to a series of movements or changes in movement that occur when an animal moves its body.

[0597] "Sound" refers to the vocalizations and other sounds made by animals.

[0598] A "recording device" is a device used to collect animal movements and sounds and store them as data.

[0599] A "generative AI model" is a type of algorithm designed to learn the characteristics of data and perform specific tasks, and is used for data analysis and prediction.

[0600] A "machine learning algorithm" is a computational method that learns patterns from data and uses those patterns to make predictions or perform classifications.

[0601] An "analytical device" is a device that performs calculations to analyze recorded data and understand its contents.

[0602] A "prediction device" is a device that identifies an animal's intentions and emotions from analyzed data.

[0603] "Natural language" usually refers to the forms of language that humans use on a daily basis, and is closer to actual conversation and writing than to language generated by computers.

[0604] A "notification device" is a device that communicates the results of analysis and estimation to the user, and its role is to provide information.

[0605] This system provides a comprehensive solution for analyzing animal emotions and intentions and notifying the user. A detailed embodiment of the system is described below.

[0606] The device is equipped with a camera and microphone to record animal movements and sounds. This allows the device to collect animal movements as video and vocalizations and other sounds as audio data in real time. Users can monitor their pets' activities using mobile devices such as smartphones and tablets. This allows users to instantly detect any abnormal behavior or changes in the animal's condition.

[0607] The server uses machine learning algorithms powered by generative AI models to analyze the behavioral and audio data transmitted from the terminal. The server receives this data, processes the audio data using spectrogram analysis techniques, and vectorizes the behavioral features by analyzing the video data frame by frame. This analysis enables highly accurate prediction of the animal's emotions and intentions. The generative AI model associates specific behavioral patterns of the animal with their emotions and intentions.

[0608] The server translates the animal's emotions and intentions into natural language based on the analysis results. This ensures that the generated messages are presented to the user in an intuitively understandable format. This process includes the automatic generation of messages such as, "The dog appears anxious. Please move it to a quiet place to calm it down," using a generative AI model.

[0609] For example, if a user's dog suddenly starts barking violently, the user records this situation via their device. The server analyzes this data, evaluating the barking pattern and the dog's body movements to infer that the dog is feeling anxious. As a result, the user receives a message via their device stating, "Your dog appears anxious. Please move it to a quiet place to calm it down." This allows the user to take intuitive action.

[0610] An example of a prompt for a generative AI model might be, "Analyze this dog's barking and behavioral data to infer its emotions and intentions." Following this prompt, the server analyzes the data and infers its emotions and intentions.

[0611] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0612] Step 1:

[0613] The device records animal movements as video and collects vocalizations and sounds as audio data. It uses a camera and microphone to detect animal movements and sounds in real time as input. This input is converted into digital data by the recording device. For example, a dog barking is recorded, and its audio is saved as a file.

[0614] Step 2:

[0615] The terminal sends collected motion data (video files) and audio data (audio files) to the server. This transmission is performed using a secure communication protocol, ensuring data confidentiality. The input is recorded digital data, and the output is the data file sent to the server.

[0616] Step 3:

[0617] The server analyzes the motion and audio data received from the terminal. The input consists of video and audio data, and the server's analysis method is a machine learning algorithm using a generative AI model. Features are extracted from the audio data through spectrogram analysis, and the video data is analyzed frame by frame, with motion features vectorized. The output is the analyzed feature data.

[0618] Step 4:

[0619] The server performs a process of inferring the animal's intentions and emotions based on the analyzed feature data. The input is the feature data obtained in the previous step, and the server's inference method associates specific patterns with emotions and intentions. The output is the inferred emotion and intention data.

[0620] Step 5:

[0621] The server generates natural language messages based on inferred data. The input is inferred intent and emotion data, and the output is an intuitively understandable natural language message. A generative AI model is used to generate messages that help users perceive and respond to their dog's anxiety.

[0622] Step 6:

[0623] The server sends the generated natural language message to the terminal. The input is message data, and the output is the message sent to the terminal. Specifically, the message "The dog appears to be feeling anxious." is generated and a notification is sent to the user's terminal.

[0624] (Application Example 1)

[0625] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0626] Within factories, there is a need to quickly detect safety risks caused by animal intrusion and strengthen safety management. Conventional systems have difficulty accurately assessing the emotions and intentions of animals, making it difficult to take immediate action. This can lead to unexpected accidents and decreased production efficiency.

[0627] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0628] In this invention, the server includes recording means for acquiring animal movements and sounds, analysis means for analyzing the acquired movements and sounds, emotion estimation means for inferring the animal's intentions and emotions based on the analysis results of the analysis means, information provision means for notifying the user of the estimation results by the emotion estimation means, and monitoring and management means for supporting safety management within the factory using the animal emotion estimation results. This makes it possible to improve safety by monitoring animal behavior within the factory in real time and providing appropriate information immediately.

[0629] "Recording means" refers to devices and methods for acquiring animal movements and sounds, and provides a foundation for understanding animal behavior in real time.

[0630] "Analysis means" refers to devices and methods for analyzing acquired animal behavior and sounds, which use machine learning algorithms to extract and analyze data features.

[0631] "Emotion estimation means" refers to a device or method for inferring an animal's intentions and emotions based on the analysis results obtained by the aforementioned analysis means.

[0632] "Information provision means" refers to devices or methods for notifying the user of the inference results obtained by emotion estimation means, and provides the information by converting it into natural language.

[0633] "Monitoring and management means" refers to devices and methods for managing safety within a factory using the results of animal emotion estimation, and are intended to improve safety within the factory.

[0634] To realize this invention, a system is constructed for acquiring and analyzing animal behavior and sounds. The system mainly consists of recording means, analysis means, emotion estimation means, information provision means, and monitoring and management means.

[0635] The server uses cameras and microphones mounted on robots that autonomously patrol the factory to record animal movements and sounds. This functions as a recording tool. The acquired data is sent to the server and analyzed using software such as Python and TensorFlow. Through this analysis, machine learning algorithms extract animal movement patterns and vocal features, and based on this, infer the animal's intentions and emotions.

[0636] The analyzed and inferred results are converted into natural language that users can intuitively understand by sentiment estimation tools. This information is then communicated to administrators by information providers, enabling a rapid response on-site. A concrete example of this notification process is when a night security robot at a factory detects the intrusion of a suspicious animal; it generates an alert stating, "A suspicious animal has entered the premises and is exhibiting aggressive behavior. Please be careful," and sends it to the administrator.

[0637] Furthermore, the monitoring and management system integrates the prediction results with the factory's safety management system, automating alarm triggers and ensuring higher safety standards.

[0638] A concrete example of a prompt for a generative AI model is, "Analyze this dataset, identify the intentions and emotions of the animals, and generate a natural language warning message based on the results." This prompt is input to the generative AI model when analyzing the acquired data and is used to obtain appropriate analysis results.

[0639] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0640] Step 1:

[0641] The terminal uses cameras and microphones mounted on robots moving around the factory to record animal movements and sounds in real time. The input consists of camera video data and audio data, and this data is sent to a server as output.

[0642] Step 2:

[0643] The server receives video and audio data provided by the terminal and prepares the data for analysis. The input is the received data, and the output is the formatted data after noise reduction and signal processing.

[0644] Step 3:

[0645] The server utilizes software such as Python and TensorFlow to apply advanced machine learning algorithms to the formatted data and perform analysis. This analysis extracts animal behavior patterns and vocal features, which are then used as input. The output is the feature extraction results.

[0646] Step 4:

[0647] The server infers the animal's intentions and emotions based on the extracted features. It uses a generative AI model to analyze the data through prompt messages. The input at this stage is the feature extraction results, and the output is the inferred intentions and emotions of the animal.

[0648] Step 5:

[0649] The server converts the emotion estimation results into a message expressed in natural language and notifies the user using an information delivery method. An example of a prompt message generated by the system using a generative AI model is: "Analyze this dataset to identify the animal's intentions and emotions, and generate a natural language warning message based on the results." The input is the emotion estimation result, and the output is a natural language message.

[0650] Step 6:

[0651] The user checks the message notified on the terminal and takes appropriate countermeasures as needed, depending on the situation in the factory. The input is a natural language message, and the output is the implementation of the countermeasure.

[0652] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0653] This invention is a system that facilitates communication by simultaneously analyzing the emotions of both the animal and the user. In addition to functions for recognizing and analyzing the animal's movements and sounds, this system integrates an emotion engine that recognizes emotions from the user's facial expressions and voice. The following describes embodiments for carrying out this invention.

[0654] The device is equipped with a camera and microphone to record animal movements and sounds, thereby acquiring real-time data. Similarly, it also has a camera and microphone to record the user's facial expressions and tone of voice. This allows the device to simultaneously acquire data from both the animal and the user.

[0655] The server processes animal data transmitted from the terminal using analysis tools to infer the animal's intentions and emotions. The analysis tools utilize machine learning algorithms and have already learned the individual vocal patterns and behavioral characteristics of each animal. Simultaneously, an emotion engine included in the server analyzes user data to infer the user's emotions. This engine can identify the user's current emotional state through facial recognition and voice analysis technologies.

[0656] The server then uses an analysis combining the animal's intentions and emotions with the user's emotions to generate an appropriate message. This message is adjusted in content and tone according to the user's current emotional state, ensuring the most effective communication.

[0657] For example, if the user is feeling stressed and the dog is showing signs of anxiety, the server will generate a message such as, "It's important for you to relax while also making your dog feel secure," and notify the user's device. Based on this information, the user can take appropriate action that suits both their and their dog's emotional state.

[0658] By receiving these notifications, users can gain a deeper understanding of the animals' conditions and take appropriate actions that are considerate of each other's feelings. They can also provide information to the server through the feedback function to improve the accuracy of the system's analysis. This feedback plays a crucial role in effective communication.

[0659] The following describes the processing flow.

[0660] Step 1:

[0661] The device acquires data using a camera and microphone to record animal movements and sounds. Simultaneously, it collects data on the user's facial expressions and tone of voice. This data is temporarily stored within the device.

[0662] Step 2:

[0663] The terminal preprocesses the acquired data, removing noise and extracting only the necessary information. The preprocessed data is then compressed and sent to the server via communication.

[0664] Step 3:

[0665] The server processes the received animal behavior and audio data using analysis tools. Machine learning algorithms are employed to extract patterns in animal sounds and movements.

[0666] Step 4:

[0667] The server uses an emotion engine to analyze the user's facial expressions and tone of voice data. The algorithm infers the user's emotional state and records that information in a database.

[0668] Step 5:

[0669] The server combines inferred animal intentions and emotions with the user's emotional state to generate the optimal message. This message is then adjusted to suit the user's emotional state.

[0670] Step 6:

[0671] The server translates the message into natural language and sends it to the terminal. This message is designed to be intuitively understandable to the user.

[0672] Step 7:

[0673] The device receives messages sent from the server and notifies the user. Notifications are made via audio alerts or screen displays.

[0674] Step 8:

[0675] Users can respond to the information they receive in accordance with the emotional state of the animal and themselves. They can also use the feedback function to provide additional information to the server, which can be used for future analysis.

[0676] (Example 2)

[0677] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0678] In recent years, there has been a growing demand for technologies that facilitate communication between humans and animals. However, conventional technologies focus solely on analyzing animal movements and vocalizations, lacking consideration for the user's emotional state. Therefore, there is a need for a means to simultaneously analyze the animal's intentions and emotions, as well as the user's emotions, and to achieve appropriate two-way communication based on this analysis.

[0679] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0680] In this invention, the server includes a device for acquiring the animal's movements and sounds, a device for acquiring the user's facial expressions and sounds, and a processing device for analyzing the acquired movements, sounds, and facial expressions. This enables communication that simultaneously analyzes the emotions of the animal and the user, and generates appropriate two-way messages based on the results.

[0681] A "device" is a combination of hardware and software used to perform a specific function.

[0682] "Action" refers to a physical activity or action performed by a living being.

[0683] "Sound" refers to the waveform of sounds emitted by animals and humans, and by analyzing its characteristics, it is possible to infer emotions and intentions.

[0684] "Facial expression" is visual information that conveys emotions and intentions through the movement of facial muscles.

[0685] A "processing device" is a device that takes data as input, analyzes it, and has the computational function to extract or infer specific information.

[0686] "Analysis" is the process of breaking down, comparing, and evaluating data to understand its meaning and relationships.

[0687] A "prediction device" is a device that has a computational function to infer the intentions and emotions of a certain subject based on collected data.

[0688] A "notification device" is a device that has the function of conveying predicted results or generated messages to the user.

[0689] "Natural language" refers to the language that humans use on a daily basis, and the text that is generated in a format suitable for computer processing.

[0690] This invention is a system that simultaneously analyzes the emotions of both animals and their users. The system aims to facilitate communication between the two by combining functions that analyze the animal's movements and sounds with functions that analyze the user's facial expressions and sounds.

[0691] The device is equipped with a camera and microphone to record the animal's movements and sounds. This hardware makes it possible to acquire real-time data on the animal's physical movements and vocalizations. The same device also has a camera and microphone to record the user's facial expressions and voice, allowing for the simultaneous acquisition of the user's emotional data.

[0692] The server analyzes animal and user data transmitted from the terminal. This analysis utilizes machine learning algorithms that analyze animal sounds and movements, as well as the user's facial expressions and tone of voice. This allows the server to accurately predict the animal's emotions and intentions, as well as the user's emotions. Various server-based analysis software and emotion engines are used in this analysis.

[0693] Based on the analysis results, the server generates messages tailored to the animal's and the user's situation. These messages are written in natural language, adjusted to the most effective content and tone, and delivered to the device. Based on this information, users can take appropriate actions that align with the emotions of both the animal and themselves.

[0694] As a concrete example, considering the example prompt, in response to the input, "Please advise how the user should react when the dog appears anxious," the system can provide appropriate advice. This realizes the use of generative AI models to support communication between animals and humans.

[0695] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0696] Step 1:

[0697] The device activates its camera and microphone to capture animal movements and sounds, as well as the user's facial expressions and voice, in real time. The acquired data includes animal movements and vocalizations, as well as the user's facial expressions and voice tone. This data is initially processed within the device for later detailed analysis.

[0698] Step 2:

[0699] The terminal sends the initially processed data to the server. Wireless communication technologies such as Wi-Fi and Bluetooth are used for this communication. The input data includes animal movement data, voice data, user facial expression data, and voice data, which are used as basic data for the next analysis step on the server.

[0700] Step 3:

[0701] The server receives the transmitted data and uses machine learning algorithms to analyze the animal's movements and sounds, as well as the user's facial expressions and voice. Based on the animal data, it infers the animal's intentions and emotions from its behavioral patterns and vocal characteristics. Regarding the user's data, it identifies emotions from changes in facial expressions and tone of voice. As a result of the analysis, the emotional states of both the animal and the user are output.

[0702] Step 4:

[0703] The server uses a generative AI model to generate appropriate messages based on the analysis of the animal's intentions and emotions, as well as the user's emotions. As a prompt, the analysis results are input into the generative AI model, which then outputs an effective message in natural language. This message is structured as specific advice that is beneficial to both the animal and the user.

[0704] Step 5:

[0705] The server sends the generated message to the terminal. The terminal notifies the user of this message either visually or audibly. Based on the received message, the user can take specific actions to adjust their communication with the animal.

[0706] (Application Example 2)

[0707] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0708] In stores that allow pets, there is a challenge in appropriately analyzing the emotional states of both customers and their pets and enabling store staff to take the most appropriate action based on that analysis. Therefore, there is a need for effective support to improve customer satisfaction and pet comfort.

[0709] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0710] In this invention, the server includes acquisition means for acquiring animal behavior and sounds, emotion analysis means for simultaneously analyzing the emotional states of customers and animals, and motion analysis means for analyzing the acquired behavior and sounds. This makes it possible to analyze the emotions of customers and pets with high accuracy and to provide appropriate customer service suggestions to store employees.

[0711] "Acquisition means" refers to equipment and technology for detecting animal behavior and sounds and converting them into data.

[0712] "Emotional analysis means" refers to technology or devices for simultaneously analyzing the emotional states of customers and animals. This means has the function of analyzing data using machine learning and emotion recognition algorithms to identify emotions.

[0713] "Motion analysis means" refers to technology or equipment used to analyze acquired animal behavior and vocalizations. This means has the function of processing data in order to understand the meaning and intent of the behavior.

[0714] "Inference means" refers to techniques or devices used to infer the emotional state or intentions of a subject based on analysis results. These means play a role in inferring the most likely emotional state based on the analyzed data.

[0715] "Proposal provision means" refers to technology or equipment for presenting specific actions or responses to external parties based on the prediction results obtained by prediction means. This means is used to instruct the optimal action based on the information obtained.

[0716] One embodiment of this invention provides a system that understands the emotions of customers accompanied by pets and their pets in a store, and enables optimal customer service. The system mainly consists of terminals and a server.

[0717] The terminal is equipped with a camera and microphone that record the movements of customers and their pets in real time, thereby collecting motion and audio data of customers and pets when they enter the store. The terminal is responsible for transmitting the collected data to a server.

[0718] The server uses machine learning algorithms to analyze animal behavior and customer emotions based on the received data. To analyze animal and customer emotions simultaneously, emotion analysis means are used to identify their respective emotional states. Motion analysis means are used to analyze animal behavior in detail and infer the animal's intentions and emotions.

[0719] Based on the insights gained from these analysis results, the server generates suggestions to facilitate optimal customer service. These suggestions are translated into natural language by a suggestion delivery system and notified to store employees. As a result, employees can provide optimal customer service tailored to the customer and their pet's condition, thereby improving customer satisfaction.

[0720] For example, if the system detects that a customer entering the store appears tense and their pet is frightened, it will generate a natural language suggestion such as "Let's give your pet some time to calm down" and provide it to the store staff via a terminal. This allows employees to provide prompt and appropriate customer service.

[0721] As an example of a prompt to a generative AI model, you could use a sentence like, "In this situation, identify passive-looking behaviors and restless vocal patterns, and suggest specific suggestions for what to do to the store staff. It is important to consider the emotional state of both the customer and the pet."

[0722] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0723] Step 1:

[0724] The terminal uses a camera and microphone to capture the movements and sounds of customers and their pets in real time as they enter the store. During this process, the posture and voices of the customers and pets are captured as input. This data is recorded as image and audio data and sent to a server.

[0725] Step 2:

[0726] The server receives data sent from the terminal and performs analysis using a machine learning model. The input consists of image and audio data, which is used to analyze animal movements and vocalizations. Motion analysis and emotion analysis methods are used for data processing to identify individual emotional states. As a result, emotion prediction data for both animals and customers is obtained.

[0727] Step 3:

[0728] The server processes sentiment inference data and generates optimal customer service suggestions using a generative AI model. Sentiment inference data is used as input, and based on this information, the server calculates what action suggestions are appropriate and outputs suggestions in natural language. This output will be notified to the terminal in a later step.

[0729] Step 4:

[0730] The terminal receives customer service suggestions in natural language from the server and presents them to store employees visually or audibly. Employees use this as a reference to perform specific customer service actions. At this time, the suggested message is displayed or played back on the terminal in a human-readable format.

[0731] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0732] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0733] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0734] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0735] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0736] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0737] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0738] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0739] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0740] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0741] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0742] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0743] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0744] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0745] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0746] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0747] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0748] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0749] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0750] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0751] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0752] The following is further disclosed regarding the embodiments described above.

[0753] (Claim 1)

[0754] An input means for acquiring animal movements and sounds,

[0755] Analysis means for analyzing acquired motion and sound,

[0756] An estimation means for inferring the intentions and emotions of an animal based on the analysis results of the aforementioned analysis means,

[0757] A notification means for notifying the user of the prediction result obtained by the aforementioned prediction means,

[0758] A system that includes this.

[0759] (Claim 2)

[0760] The system according to claim 1, characterized in that the analysis means uses a machine learning algorithm to analyze the operation and sound.

[0761] (Claim 3)

[0762] The system according to claim 1, characterized in that the notification means converts the inference result by the inference means into natural language and provides it to the user.

[0763] "Example 1"

[0764] (Claim 1)

[0765] A recording device for recording the movements and sounds of animals,

[0766] An analysis device that analyzes recorded actions and sounds using a machine learning algorithm that utilizes a generative AI model,

[0767] An inference device that infers the intentions and emotions of an animal based on the analysis results of the aforementioned analytical device,

[0768] A notification device that automatically converts the prediction results of the prediction device into natural language and notifies the user in a format that is intuitively understandable,

[0769] A system that includes this.

[0770] (Claim 2)

[0771] The system according to claim 1, characterized in that the recording device collects the animal's movements as video and its sounds as audio data.

[0772] (Claim 3)

[0773] The system according to claim 1, wherein the notification device transmits the generated natural language message to a terminal and immediately notifies the user.

[0774] "Application Example 1"

[0775] (Claim 1)

[0776] Recording means for acquiring animal movements and sounds,

[0777] Analysis means for analyzing acquired motion and sound,

[0778] An emotion estimation means for inferring the intentions and emotions of an animal based on the analysis results of the aforementioned analysis means,

[0779] Information provision means for notifying the user of the estimation results obtained by the emotion estimation means,

[0780] A monitoring and management system that uses the results of animal emotion estimation to support safety management within a factory,

[0781] A system that includes this.

[0782] (Claim 2)

[0783] The system according to claim 1, characterized in that the analysis means uses a machine learning algorithm to analyze the behavior and sound.

[0784] (Claim 3)

[0785] The system according to claim 1, characterized in that the information providing means converts the inference result by the emotion estimation means into natural language and provides it to the user.

[0786] "Example 2 of combining an emotion engine"

[0787] (Claim 1)

[0788] A device for acquiring animal movements and sounds,

[0789] A device that acquires the user's facial expressions and voice,

[0790] A processing unit that analyzes acquired movements, voice, and facial expressions,

[0791] An estimation device that infers the intentions and emotions of an animal, as well as the emotions of the user, based on the analysis results of the aforementioned processing device,

[0792] A notification device that generates and notifies the user of an appropriate message using the prediction results of the animal's and the user's emotions obtained by the aforementioned prediction device,

[0793] A system that includes this.

[0794] (Claim 2)

[0795] The system according to claim 1, characterized in that the processing device uses a machine learning algorithm to analyze movements, voices, and facial expressions and simultaneously infer the emotions of the animal and the user.

[0796] (Claim 3)

[0797] The system according to claim 1, characterized in that the notification device converts the prediction results of the emotions of the animal and the user by the prediction device into natural language and provides it to the user.

[0798] "Application example 2 of combining emotional engines"

[0799] (Claim 1)

[0800] means for acquiring animal behavior and sounds,

[0801] An emotion analysis means for simultaneously analyzing the emotional state of customers and animals,

[0802] A motion analysis means for analyzing acquired actions and sounds,

[0803] An estimation means for inferring the emotions of animals and customers based on the analysis results of the motion analysis means and emotion analysis means,

[0804] A proposal provision means that provides a proposal to a store employee based on the prediction result of the aforementioned prediction means,

[0805] A system that includes this.

[0806] (Claim 2)

[0807] The system according to claim 1, characterized in that the motion analysis means and emotion analysis means analyze motion and emotion using machine learning algorithms.

[0808] (Claim 3)

[0809] The system according to claim 1, characterized in that the proposal providing means converts the inference result from the inference means into natural language and provides it to the store employee. [Explanation of Symbols]

[0810] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. An input means for acquiring animal movements and sounds, Analysis means for analyzing acquired motion and sound, An estimation means for inferring the intentions and emotions of an animal based on the analysis results of the aforementioned analysis means, A notification means for notifying the user of the prediction result obtained by the aforementioned prediction means, A system that includes this.

2. The system according to claim 1, characterized in that the analysis means uses a machine learning algorithm to analyze the operation and sound.

3. The system according to claim 1, characterized in that the notification means converts the inference result from the inference means into natural language and provides it to the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A