system

A system using deep learning to analyze pet sounds and images predicts emotions, enhancing pet owner understanding and communication by continuously improving model accuracy through feedback.

JP2026070202APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Pet owners face difficulties in accurately understanding their pets' emotions from their cries and expressions, leading to challenges in detecting health status and emergencies.

Method used

A system that records animal sounds and captures images, combining this data with behavioral information into a single packet, preprocesses it, and analyzes it using a deep learning model to predict emotions, with continuous model improvement through user feedback.

Benefits of technology

Enables pet owners to accurately understand and respond to their pets' emotions, improving communication and health detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070202000001_ABST
    Figure 2026070202000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for recording the voice of an animal, Means for taking a picture of an animal, Means for inputting information regarding the behavior of the animal, Means for collecting and transmitting the recorded voice, the taken picture, and the input information in one data packet, Means for preprocessing the voice data and the picture data, Means for analyzing the preprocessed data using a deep learning model and predicting the emotion of the animal, Means for outputting the predicted emotion result, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Pet owners face the problem that it is difficult to accurately understand the emotions of their pets from the pets' cries and expressions, resulting in smooth communication with their pets being impossible. This problem can particularly lead to a serious situation where the health status and emergencies of pets cannot be accurately detected. Therefore, there is a need for a technology to predict the emotions of pets in real time from their behaviors and provide that information to the owners.

Means for Solving the Problems

[0005] To solve this problem, the present invention provides the following system. First, it includes means for recording animal sounds and means for capturing images of animals. This data, along with information about the animal's behavior, is combined into a single data packet and transmitted. Subsequently, the audio and image data are preprocessed on a server in the cloud and analyzed by a deep learning model. This model predicts the animal's emotions and assists in accurate communication by notifying the user of the predicted emotions. Furthermore, the system can improve prediction accuracy by continuously updating the model using feedback from the user.

[0006] "Animals" refers to living beings other than humans, and in this invention, it refers to dogs, cats, and other animals kept as pets.

[0007] "Sound" refers to sound signals such as animal noises, which are converted into data using recording devices.

[0008] An "image" refers to visual data captured by a camera device, such as the appearance and facial expressions of an animal, and is a frame of a photograph or video that is to be analyzed.

[0009] A "data packet" is a collection of information that combines voice data, image data, and behavioral information into a single package, and is configured for simultaneous processing and transmission.

[0010] "Preprocessing" is the process of shaping collected data into a form suitable for analysis, and includes noise reduction and resolution adjustment.

[0011] A "deep learning model" is an artificial intelligence technology that uses multi-layer neural networks to learn from large amounts of data and analyze patterns and features.

[0012] "Means for predicting emotions" refers to a set of algorithms and processes that determine an animal's emotions based on analysis and output the results.

[0013] "Feedback" refers to information provided by system users regarding the accuracy of prediction results and for improving the model, and is used to enhance the learning model. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention relates to a system that records animal sounds and facial expressions using a user's device, analyzes this data, and infers the animal's emotions. The user first records the animal's sounds using a microphone and captures its facial expressions with a camera using a smartphone or dedicated camera device. This data, along with information about the animal's behavior, is input into the device.

[0036] The collected data is integrated into a single data packet at the terminal and sent to a cloud server. The server preprocesses the received data, preparing noise-canceling audio data and image data with optimized resolution. This preprocessed data is then analyzed by a deep learning model to predict the animal's emotions.

[0037] This deep learning model utilizes a multi-layer neural network and is trained on a large amount of training data collected from pet shops and animal shelters. The model includes algorithms that extract features from animal sounds and images and predict corresponding emotions.

[0038] Ultimately, the server sends the animal's emotion prediction results to the user's device, where they are displayed on the user interface via a dedicated application. The user can then decide how to interact with the animal based on the displayed results and send feedback if they feel the prediction is incorrect. This user feedback is recorded in a database and used to improve the model's accuracy.

[0039] As a concrete example, consider a situation where a user owns a dog and the dog is excited and barking. The device records the barking and takes a picture of the dog's facial expression. Next, the user inputs into the app that "the dog likely wants to play." This data is sent to a server, and a deep learning model analyzes it to predict the dog's emotion as "excited," and notifies the user. Based on this result, the user can decide whether to spend time playing with their dog.

[0040] In this way, the present invention enables pet owners to more accurately understand their pets' emotions and needs and respond appropriately by scientifically analyzing their behavior.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] Data collection

[0044] The device uses a smartphone or dedicated device to record animal sounds with a microphone and capture facial expressions with a camera. The recorded audio data is saved in WAV format, and the captured image data is saved in JPEG format, etc. Based on this collected data, the user inputs information related to the pet's behavior and anticipated emotions.

[0045] Step 2:

[0046] Data packet generation and transmission

[0047] The device combines recorded audio data, image data, and user-entered emotional information into a single data packet. This packet also includes necessary metadata such as a timestamp and pet ID. The generated data packet is then sent to a cloud server.

[0048] Step 3:

[0049] Data preprocessing

[0050] The server receives the transmitted data packets and preprocesses the audio and image data before data analysis. The audio data is noise-canceling to convert it into clear audio. The image data is processed using resolution adjustment and facial recognition algorithms to accurately extract animal expressions.

[0051] Step 4:

[0052] Application of emotion analysis models

[0053] The server uses a deep learning model to analyze pre-processed audio and image data as input. This model is trained to infer emotions from the characteristics of pet vocalizations and facial expressions. Using the trained model, the emotions of animals can be predicted with high accuracy.

[0054] Step 5:

[0055] Sending and displaying results

[0056] The server sends the predicted emotion result to the user's device. The device displays the received emotion result on the user interface and adds a visual indicator to inform the user of the pet's current emotional state.

[0057] Step 6:

[0058] Gathering feedback and improving the learning model

[0059] The user checks whether the displayed emotion result matches the pet's actual emotion. If the result is different, they send the correct emotion through the feedback option. The server records this feedback information in a database and uses it as training data for the next model update to improve the accuracy of the analysis model.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] To understand and appropriately respond to animal emotions, it is necessary to accurately analyze animal vocalizations, facial expressions, and behavioral information. However, conventional methods have made it difficult to scientifically predict animal emotions, potentially leading to misunderstandings of animal needs by pet owners and animal care professionals. This invention aims to solve this problem and provide a method for more accurately predicting animal emotions.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes a device for collecting animal vibrations, a device for acquiring visual information of the animal, and a device for inputting data related to the animal's activities. This allows the animal's emotions to be scientifically predicted using a multi-layer artificial intelligence model, enabling the owner to respond appropriately.

[0065] "Animal vibrations" refer to sounds and noises emitted by animals, which are recorded using means of detecting sound waves.

[0066] "Visual information" refers to video data, including the facial expressions and postures of animals, acquired by cameras and image sensors.

[0067] "Activity data" refers to information about the behavior and status of animals, which users input based on their observations and reports.

[0068] An "information unit" is a digital piece of information that encompasses audio, visual, and activity-related data, bundled together as a single data packet.

[0069] "Pre-processing" means applying necessary processing or transformations to raw data to make it easier to analyze.

[0070] A "multilayer artificial intelligence model" is an algorithm that uses multiple neural networks to extract features from input data and perform classification and prediction based on the results.

[0071] "Noise reduction" is a technology that removes unwanted noise from audio data and converts it into clearer audio information.

[0072] "Evaluation information" refers to feedback and opinions provided by users, and is data used to improve the model and enhance its accuracy.

[0073] This invention provides a system that infers an animal's emotions based on vibration and visual information collected by a user. The user acquires animal sounds and images using a terminal such as a smart device. Specifically, it is possible to record the sounds and facial expressions of the animal using a microphone and camera mounted on the device.

[0074] The device applies noise reduction technology to recorded audio data and preprocesses visual information by optimizing image resolution. This results in data that can be efficiently analyzed by multi-layer artificial intelligence models, including deep learning. This analysis is performed by a server in the cloud, which receives the data transmitted from the device and uses a generative AI model implementing a multi-layer neural network to predict the animal's emotions.

[0075] The results of this emotion prediction are sent to the user's device and displayed by a dedicated application. This allows the user to understand the animal's emotional state and take appropriate action accordingly.

[0076] As a concrete example, consider a scenario where a user wants to know the emotions of their pet dog at home. The user uses their device to record the dog's barks and take a picture of its facial expression. This data is sent to a server and analyzed by a deep learning model, which predicts the dog is "excited." In this situation, the prompt could be, "Based on this dog's barking and facial expression data, what emotion does the deep learning model predict?"

[0077] Furthermore, user feedback is recorded as evaluation information and helps improve the accuracy of the generated AI model. This feedback allows the system to continuously learn and improve the accuracy of predicting animal emotions.

[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0079] Step 1:

[0080] The user collects vibration and visual information from the animal. The user records the animal's sounds using the microphone on their smart device and captures its facial expressions using the camera. The input consists of audio and image data, which are then stored on the device.

[0081] Step 2:

[0082] The data collected by the terminal is preprocessed. For audio data, noise reduction technology is used to remove unwanted noise and output clear audio data. For image data, the resolution is optimized and prepared for analysis. This process yields audio and image data that can be analyzed.

[0083] Step 3:

[0084] The terminal packets the pre-processed data into a single information unit. This information unit contains pre-processed audio data, image data, and information about the animal's activity entered by the user. This is then sent to the cloud server as output.

[0085] Step 4:

[0086] The server analyzes the information units it receives. The server inputs pre-processed audio and image data into a generating AI model and uses deep learning technology to predict the animal's emotions. The output is an emotion prediction result.

[0087] Step 5:

[0088] The server sends the prediction results to the user's device. The device displays the received emotion prediction results on a user interface via a dedicated application. Based on these results, the user understands the animal's condition and decides on an appropriate response.

[0089] Step 6:

[0090] Users provide feedback as needed. Users input their opinions on the displayed sentiment predictions and send them to the server. The server records this feedback as evaluation information and uses it to improve the generating AI model.

[0091] (Application Example 1)

[0092] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0093] In developing products using animals, it is desirable to accurately understand the animals' emotions and reactions and use that information to improve the products. Conventional methods have made it difficult to quantitatively evaluate animal emotions, and there has been a lack of means to quickly and accurately measure animals' reactions to products. Therefore, there was a need for an effective system to analyze animal emotions and use that information to advance product development.

[0094] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0095] In this invention, the server includes a device for recording animal sounds, a device for capturing images of animals, and a device for inputting information about the animals' behavior. This makes it possible to analyze the animals' emotions in real time and obtain rapid feedback in the manufacturing process.

[0096] A "device for recording animal sounds" is a device that collects sounds emitted by animals and saves them as digital data.

[0097] A "device for capturing images of animals" is a device that uses a camera to capture the appearance and facial expressions of animals and records that visual information as image data.

[0098] "The device for inputting information about the animal's behavior" refers to a device for manually or automatically inputting information about the animal's movements and behavior.

[0099] A "multi-layered artificial intelligence model" is an algorithm that uses a neural network consisting of multiple layers to extract features from data and obtain analysis results.

[0100] "A device that outputs the predicted emotional results in real time during animal testing in the manufacturing process" refers to a device that immediately displays the analyzed emotional results and provides relevant information during animal testing.

[0101] "Noise reduction" is a technique used to remove unwanted noise from recorded audio data and obtain clear audio information.

[0102] The system for implementing this invention is designed as a comprehensive platform for collecting and analyzing animal sounds and images in real time. The server primarily handles this process, first acquiring animal sound and image data through a sound recording device and image capture device connected to a terminal used by the user.

[0103] The server integrates this data into a single data packet and sends it to the cloud infrastructure. On the cloud, the data is preprocessed using software that removes noise from the audio data and processes the image data to the optimal resolution. This enables more accurate analysis.

[0104] Next, the server uses a multi-layered artificial intelligence model to analyze the pre-processed data. This model is specialized for predicting animal emotions and learns animal vocal and image features based on a large amount of training data.

[0105] The resulting emotion predictions are sent to the user's device and displayed through a dedicated application. Users can use these results to improve products or respond immediately to animal reactions.

[0106] As a concrete example, this system could be used in pet food development to evaluate animals' reactions to new prototypes. For instance, a possible prompt might be, "If the animal appears happy due to the new pet food, please output a record of its emotions." In this way, the present invention provides a tool for scientifically analyzing animal emotions and improving product development and communication with pets.

[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0108] Step 1:

[0109] The device records animal sounds using a microphone and captures images of animals with a camera. Input consists of animal sounds and visual data, while output consists of recorded audio and image data. This data is temporarily stored within the device.

[0110] Step 2:

[0111] The terminal integrates the recorded audio and image data, along with animal behavior information input by the user, into a single data packet. The input is the data set obtained in the previous step, and the output is the integrated data packet. The data is organized and prepared for subsequent analysis.

[0112] Step 3:

[0113] The server receives data packets sent from the terminal and performs noise reduction on the voice data. The input for this step is the integrated data packet, and the output is the voice data with the noise removed. Noise cancellation technology is used to improve voice clarity.

[0114] Step 4:

[0115] The server optimizes the resolution of the image data into a format suitable for analysis. The input is the image data in the data packet obtained in the previous step, and the output is the optimized image data. Image processing software performs resolution adjustment.

[0116] Step 5:

[0117] The server uses a pre-trained, multi-layered artificial intelligence model to analyze pre-processed audio and image data and predict animal emotions. The input is pre-processed audio and image data, and the output is the predicted animal emotion. A generative AI model is used to extract features from the data and determine emotions.

[0118] Step 6:

[0119] The server sends predicted emotion results to the terminal, which displays the results via a dedicated application. The input is the predicted emotion result of the animal, and the output is the display on the application. This provides the user with visual and immediate feedback.

[0120] Step 7:

[0121] The user decides how to interact with the animal based on the displayed emotion results and provides feedback as needed. The input is the emotion prediction result, and the output is the action decision and feedback submission. Information support is provided to help the user adjust their interaction with the animal.

[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0123] This invention relates to a system that incorporates an emotion engine to achieve richer communication by analyzing not only interaction with animals but also the user's emotions. This system analyzes animal sounds and facial expressions in real time and combines the results with the user's emotional information to perform two-way emotion prediction between animals and humans.

[0124] Users use a smartphone or a dedicated device to record animal sounds with a microphone and capture their facial expressions with a camera. The device combines this data into a single data packet and sends it to a cloud server. On the server, the transmitted data is preprocessed to create noise-canceled audio data and optimized image data, which are then analyzed using a deep learning model.

[0125] The server not only predicts the animal's emotions but also recognizes the user's emotions by analyzing the user's facial expressions and voice data acquired using the device's camera or another emotion recognition device. Specifically, it analyzes the user's facial expressions and statements when they are with their pet and performs analysis using an emotion engine. Based on this analysis, it makes a comprehensive judgment about the emotional state of both the animal and the user.

[0126] As a result, the server integrates and presents both the animal's emotions and the user's emotions in a single interface, achieving two-way emotional understanding. For example, if a dog is excited and barking, the server recognizes that the user's expression is smiling and determines that both are having a good time. This information is communicated to the user visually or audibly through the user interface, and pet owners can use it to make better decisions about their time with their pets.

[0127] Users can adjust their interactions with their pets based on the animal and human emotional states presented by the system. They can also contribute to improving the emotion engine by providing feedback when inaccurate results are displayed. This feedback is used to improve the accuracy of the deep learning model in the next learning cycle.

[0128] As described above, the present invention is a significant system that comprehensively analyzes the emotions of animals and humans and provides the results, thereby promoting communication between them.

[0129] The following describes the processing flow.

[0130] Step 1:

[0131] Users record animal sounds using their smartphones or dedicated devices and capture their facial expressions with a camera. The device saves the audio data in WAV format and the image data in JPEG format. Simultaneously, the device uses its camera and emotion recognition device to acquire the user's facial expressions and audio data.

[0132] Step 2:

[0133] The device combines animal audio data, image data, and user emotion data into a single data packet. This data packet includes the pet's ID and time information. The device then sends this data to a cloud server.

[0134] Step 3:

[0135] The server preprocesses the received data packets for analysis. Noise cancellation is applied to animal audio data to improve sound quality. Image data is processed to identify animal expressions using resolution adjustments and face detection algorithms.

[0136] Step 4:

[0137] The server inputs pre-processed animal audio and image data into a deep learning model to predict the animal's emotions. This model is trained on historical data and outputs specific emotion tags (e.g., "excited," "anxious," etc.).

[0138] Step 5:

[0139] The server also analyzes the user's emotional data and uses an emotion engine to recognize the user's current emotions. This engine evaluates the user's facial expressions, tone of voice, and actions to estimate their emotions.

[0140] Step 6:

[0141] The server integrates the emotional data of both the animal and the user and sends it to the terminal as a single emotional result. This allows the user to obtain two-way information that includes not only the animal's emotions but also their own emotional state.

[0142] Step 7:

[0143] The device displays received emotional results on the user interface. If there is a positive emotional match, it provides feedback with emojis or voice messages indicating that state. Users can then use this to adjust their time with their pet.

[0144] Step 8:

[0145] If the results differ from the actual situation, the user provides data to the sentiment engine through feedback. The server uses this feedback as training material to improve accuracy in the next model update.

[0146] (Example 2)

[0147] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0148] Emotional communication between animals and humans is often judged solely on the animal's behavior using conventional technologies, and a challenge lies in the fact that it cannot take into account the user's emotional state. To accurately understand animal emotions, it is necessary to comprehensively analyze both the user's and the animal's emotions and achieve two-way emotional understanding. Furthermore, it is necessary to improve the accuracy of emotion analysis by effectively incorporating feedback.

[0149] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0150] In this invention, the server includes means for inferring the emotions of an animal, means for analyzing the emotions of a user, and means for presenting an integrated emotional state. This makes it possible to understand the emotions of both the animal and the user in a bidirectional manner and to enrich communication.

[0151] A "device for recording animal sounds" is an acoustic collection device that can effectively capture sounds emitted by animals and save them as digital data.

[0152] A "device for capturing images of animals" is an imaging device that captures the appearance and expressions of animals in real time and digitizes them as visual information.

[0153] A "device that analyzes user emotions" is a device that can determine a user's emotional state by analyzing patterns in their facial expressions and voice.

[0154] A "device that bundles and transmits data into data packets" is a device that compresses various forms of digital data into a single, unified format and transmits it to a remote server or other location.

[0155] A "pre-processing and noise removal device" is a device that removes unwanted components from acquired audio and image data to generate high-quality data.

[0156] A "device that analyzes data using deep learning technology" is a device that uses advanced machine learning models, such as neural networks, to analyze input data and extract meaningful information.

[0157] A "device for inferring animal emotions" is a device that can identify and evaluate the psychological state of an animal based on processed data.

[0158] A "device for integrating and presenting information" is a visual or auditory output device that combines multiple pieces of emotional information obtained from analysis and presents them to the user in an easily understandable format.

[0159] This invention provides an interactive system that enables two-way emotional understanding between animals and users. The system collects animal sounds and images and infers the animal's emotions based on this information. Simultaneously, it analyzes the user's facial expressions and voice to determine the user's emotions, thereby constructing an emotional interface between the animal and the user.

[0160] Users can use a dedicated digital device or smartphone to collect animal sounds with a built-in microphone and capture images of animals with a camera. This digital data is combined into a single data packet by the device, encrypted, and sent to a cloud server. On the cloud server, the data is first preprocessed, noise is removed from the audio data, and the image data is optimized into an analyzable format.

[0161] The server analyzes pre-processed audio and image data using a deep learning model to infer the emotions of animals. The deep learning model learns audio and image features and classifies emotions based on them. Next, it determines the user's emotions by analyzing the user's facial expressions and audio data.

[0162] These analysis results are integrated and presented on the user interface. This allows users to understand the emotional state of their animal and their own emotional state at the same time, and adjust their communication with their pet accordingly. Furthermore, users can contribute to improving the deep learning model by submitting feedback, and this feedback will be used to improve the model's accuracy in the next learning cycle.

[0163] For example, if a user records their cat's meows and takes a picture of themselves smiling, the system will infer that the cat is satisfied and sense that the user is also happy. Based on these results, the user can make decisions to improve the quality of their interaction with their cat.

[0164] As an example of a prompt, you could enter: "I filmed my reaction when my dog ​​barked at me. Please analyze this data to tell me the emotions of both my dog ​​and me."

[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0166] Step 1:

[0167] The user initiates interaction with the animal. Specifically, the user uses a smartphone or dedicated device to record the animal's sounds with a microphone and capture the animal's facial expressions with a camera. The input at this time is the animal's sounds and facial expressions, and the output is audio data and image data.

[0168] Step 2:

[0169] The device bundles the acquired animal audio and image data into packets and sends them to the cloud server. This packetization process involves data compression and encryption, ensuring thorough protection of the transmitted data. In this step, the input is the audio and image data, and the output is the data packets sent to the server.

[0170] Step 3:

[0171] The server receives data packets sent from the terminal. It first performs preprocessing on this data. Specific operations include removing noise from audio data to generate clear audio and optimizing image data into a format suitable for analysis. The input is data packets, and the output is preprocessed audio and image data.

[0172] Step 4:

[0173] The server performs analysis by inputting pre-processed data into a deep learning model. In this analysis step, features are extracted from the animal's voice and image, respectively, to infer the animal's emotions. The user's facial expressions and voice are also analyzed to determine the user's emotions. The input is pre-processed audio and image data, and the output is the emotional state of the animal and the user.

[0174] Step 5:

[0175] The server integrates the analyzed animal and user emotional states and presents the summarized results to the user. This process uses visual icons and textual information to make it easy for the user to understand. The input is the emotional states of the animal and the user, and the output is the emotional information displayed in the user interface.

[0176] Step 6:

[0177] The user reviews the presented analysis results and submits feedback as needed. This feedback is used to train the next deep learning model. The input is the analysis results, and the output is the user's feedback information.

[0178] (Application Example 2)

[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0180] In factories, the interface between workers and machines is primarily based on physical operations and simple indicators, lacking emotional communication. This increases worker stress and poses a risk of decreased productivity. Therefore, there is a need for a system that allows workers and robots to understand each other's emotions and collaborate more smoothly.

[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0182] In this invention, the server includes means for recording animal sounds, means for capturing images of animals, and means for acquiring and analyzing the user's facial expressions and voice. This makes it possible to understand the emotions of workers and robots in a two-way manner and improve the work environment.

[0183] "Means for recording animal sounds" refers to a device or mechanism that has the function of acquiring sounds emitted by animals as digital data.

[0184] "Means for photographing images of animals" refers to a device or technology that has the function of acquiring the appearance of an animal as a digital image.

[0185] "Means for inputting information about animal behavior" refers to devices, sensors, or interfaces that have the function of inputting data about the movements and state of animals into a system.

[0186] "Means of sending data in a data packet" refers to a device or technology that has the function of integrating different types of data into a single packet and transmitting it over a network.

[0187] "Means for preprocessing audio and image data" refers to technologies or devices that have functions for filtering or converting acquired audio and images into a format that is easy to analyze.

[0188] "Methods of analysis using deep learning models" refer to computer programs or models that use artificial intelligence technology to analyze large amounts of data and extract specific patterns or features.

[0189] "Methods for predicting animal emotions" refer to technologies and algorithms that have the function of inferring the emotional state of an animal based on acquired data.

[0190] "Means for acquiring and analyzing a user's facial expressions and voice" refers to a device or technology that has the function of acquiring a user's facial movements and speech as digital data, analyzing it, and determining their emotions.

[0191] "Means for achieving two-way emotional understanding" refers to interfaces and analytical technologies that enable the mutual recognition and understanding of emotions between animals and humans.

[0192] "Means for outputting predicted emotional results" refers to a device or interface that has the function of presenting the analyzed emotional state to the user visually or audibly.

[0193] To implement this invention, a device for recording animal sounds and images is used. Specifically, a terminal equipped with a microphone and a camera is used to acquire the sounds and facial expressions emitted by animals. Similarly, the user's facial expressions and voice are also acquired, and the system operates to integrate this data.

[0194] The terminal combines the acquired animal and user audio and image data into a single data packet and sends the data to the cloud server. As a preprocessing step, the server applies noise cancellation to clarify the audio data and converts the image data into a format that is easy to analyze.

[0195] This data is analyzed based on a deep learning model to predict the emotions of both animals and users. This utilizes deep learning frameworks such as TENSORFLOW® or PyTorch. Through the user interface, the server presents the predicted emotional states in an integrated manner, allowing the user to adjust their interaction with the animals based on this information.

[0196] For example, if the user is smiling when the animal is barking, the server will determine that both are having a good time and display that information on the interface. Such a system helps pet owners build better relationships with their pets.

[0197] An example of a prompt message is, "The worker is smiling. Think of a response that would allow the robot to generate a positive comment in this situation." This allows for explicit instructions on understanding emotions in a specific situation.

[0198] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0199] Step 1:

[0200] The device uses a microphone and camera to collect animal sounds and images. The input is real-time captured audio and video data, which is then converted into a digital format.

[0201] Step 2:

[0202] Similarly, the user's facial expressions and voice are collected using the device's camera and microphone. The input consists of the user's voice and facial image data, which are recorded as digital data.

[0203] Step 3:

[0204] The terminal combines the acquired animal and user data into a single data packet. The output is this integrated data packet, which is then sent to the cloud server.

[0205] Step 4:

[0206] The server receives data packets and applies noise cancellation to the audio data. The input is a unified data packet, and the output is the noise-removed audio data. An audio processing algorithm performs this task.

[0207] Step 5:

[0208] Simultaneously, the server optimizes the image data, adjusting its brightness and contrast. The input is the image data within the data packet, and the output is an optimized image that is easy to analyze.

[0209] Step 6:

[0210] The server analyzes pre-processed audio and image data using a deep learning model. The input consists of optimized audio and image data. The model predicts the emotions of animals and users from this data.

[0211] Step 7:

[0212] The server outputs the analysis results to the user interface, visually or audibly notifying the user of the animal's and the user's emotions. The output represents the predicted emotional state. The user can then adjust their interaction with the animal based on this information.

[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0216] [Second Embodiment]

[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0229] This invention relates to a system that records animal sounds and facial expressions using a user's device, analyzes this data, and infers the animal's emotions. The user first records the animal's sounds using a microphone and captures its facial expressions with a camera using a smartphone or dedicated camera device. This data, along with information about the animal's behavior, is input into the device.

[0230] The collected data is integrated into a single data packet at the terminal and sent to a cloud server. The server preprocesses the received data, preparing noise-canceling audio data and image data with optimized resolution. This preprocessed data is then analyzed by a deep learning model to predict the animal's emotions.

[0231] This deep learning model utilizes a multi-layer neural network and is trained on a large amount of training data collected from pet shops and animal shelters. The model includes algorithms that extract features from animal sounds and images and predict corresponding emotions.

[0232] Ultimately, the server sends the animal's emotion prediction results to the user's device, where they are displayed on the user interface via a dedicated application. The user can then decide how to interact with the animal based on the displayed results and send feedback if they feel the prediction is incorrect. This user feedback is recorded in a database and used to improve the model's accuracy.

[0233] As a concrete example, consider a situation where a user owns a dog and the dog is excited and barking. The device records the barking and takes a picture of the dog's facial expression. Next, the user inputs into the app that "the dog likely wants to play." This data is sent to a server, and a deep learning model analyzes it to predict the dog's emotion as "excited," and notifies the user. Based on this result, the user can decide whether to spend time playing with their dog.

[0234] In this way, the present invention enables pet owners to more accurately understand their pets' emotions and needs and respond appropriately by scientifically analyzing their behavior.

[0235] The following describes the processing flow.

[0236] Step 1:

[0237] Data collection

[0238] The device uses a smartphone or dedicated device to record animal sounds with a microphone and capture facial expressions with a camera. The recorded audio data is saved in WAV format, and the captured image data is saved in JPEG format, etc. Based on this collected data, the user inputs information related to the pet's behavior and anticipated emotions.

[0239] Step 2:

[0240] Data packet generation and transmission

[0241] The device combines recorded audio data, image data, and user-entered emotional information into a single data packet. This packet also includes necessary metadata such as a timestamp and pet ID. The generated data packet is then sent to a cloud server.

[0242] Step 3:

[0243] Data preprocessing

[0244] The server receives the transmitted data packets and preprocesses the audio and image data before data analysis. The audio data is noise-canceling to convert it into clear audio. The image data is processed using resolution adjustment and facial recognition algorithms to accurately extract animal expressions.

[0245] Step 4:

[0246] Application of emotion analysis models

[0247] The server uses a deep learning model to analyze pre-processed audio and image data as input. This model is trained to infer emotions from the characteristics of pet vocalizations and facial expressions. Using the trained model, the emotions of animals can be predicted with high accuracy.

[0248] Step 5:

[0249] Sending and displaying results

[0250] The server sends the predicted emotion result to the user's device. The device displays the received emotion result on the user interface and adds a visual indicator to inform the user of the pet's current emotional state.

[0251] Step 6:

[0252] Gathering feedback and improving the learning model

[0253] The user checks whether the displayed emotion result matches the pet's actual emotion. If the result is different, they send the correct emotion through the feedback option. The server records this feedback information in a database and uses it as training data for the next model update to improve the accuracy of the analysis model.

[0254] (Example 1)

[0255] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0256] To understand and appropriately respond to animal emotions, it is necessary to accurately analyze animal vocalizations, facial expressions, and behavioral information. However, conventional methods have made it difficult to scientifically predict animal emotions, potentially leading to misunderstandings of animal needs by pet owners and animal care professionals. This invention aims to solve this problem and provide a method for more accurately predicting animal emotions.

[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0258] In this invention, the server includes a device for collecting animal vibrations, a device for acquiring visual information of the animal, and a device for inputting data related to the animal's activities. This allows the animal's emotions to be scientifically predicted using a multi-layer artificial intelligence model, enabling the owner to respond appropriately.

[0259] "Animal vibrations" refer to sounds and noises emitted by animals, which are recorded using means of detecting sound waves.

[0260] "Visual information" refers to video data, including the facial expressions and postures of animals, acquired by cameras and image sensors.

[0261] "Activity data" refers to information about the behavior and status of animals, which users input based on their observations and reports.

[0262] An "information unit" is a digital piece of information that encompasses audio, visual, and activity-related data, bundled together as a single data packet.

[0263] "Pre-processing" means applying necessary processing or transformations to raw data to make it easier to analyze.

[0264] A "multilayer artificial intelligence model" is an algorithm that uses multiple neural networks to extract features from input data and perform classification and prediction based on the results.

[0265] "Noise reduction" is a technology that removes unwanted noise from audio data and converts it into clearer audio information.

[0266] "Evaluation information" refers to feedback and opinions provided by users, and is data used to improve the model and enhance its accuracy.

[0267] This invention provides a system that infers an animal's emotions based on vibration and visual information collected by a user. The user acquires animal sounds and images using a terminal such as a smart device. Specifically, it is possible to record the sounds and facial expressions of the animal using a microphone and camera mounted on the device.

[0268] The device applies noise reduction technology to recorded audio data and preprocesses visual information by optimizing image resolution. This results in data that can be efficiently analyzed by multi-layer artificial intelligence models, including deep learning. This analysis is performed by a server in the cloud, which receives the data transmitted from the device and uses a generative AI model implementing a multi-layer neural network to predict the animal's emotions.

[0269] The results of this emotion prediction are sent to the user's device and displayed by a dedicated application. This allows the user to understand the animal's emotional state and take appropriate action accordingly.

[0270] As a concrete example, consider a scenario where a user wants to know the emotions of their pet dog at home. The user uses their device to record the dog's barks and take a picture of its facial expression. This data is sent to a server and analyzed by a deep learning model, which predicts the dog is "excited." In this situation, the prompt could be, "Based on this dog's barking and facial expression data, what emotion does the deep learning model predict?"

[0271] Furthermore, user feedback is recorded as evaluation information and helps improve the accuracy of the generated AI model. This feedback allows the system to continuously learn and improve the accuracy of predicting animal emotions.

[0272] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0273] Step 1:

[0274] The user collects vibration and visual information from the animal. The user records the animal's sounds using the microphone on their smart device and captures its facial expressions using the camera. The input consists of audio and image data, which are then stored on the device.

[0275] Step 2:

[0276] The data collected by the terminal is preprocessed. For audio data, noise reduction technology is used to remove unwanted noise and output clear audio data. For image data, the resolution is optimized and prepared for analysis. This process yields audio and image data that can be analyzed.

[0277] Step 3:

[0278] The terminal packets the pre-processed data into a single information unit. This information unit contains pre-processed audio data, image data, and information about the animal's activity entered by the user. This is then sent to the cloud server as output.

[0279] Step 4:

[0280] The server analyzes the information units it receives. The server inputs pre-processed audio and image data into a generating AI model and uses deep learning technology to predict the animal's emotions. The output is an emotion prediction result.

[0281] Step 5:

[0282] The server sends the prediction results to the user's device. The device displays the received emotion prediction results on a user interface via a dedicated application. Based on these results, the user understands the animal's condition and decides on an appropriate response.

[0283] Step 6:

[0284] The user provides feedback as needed. The user inputs their opinion on the displayed emotion prediction and sends it to the server. The server records this feedback as evaluation information and uses it to improve the generated AI model.

[0285] (Application Example 1)

[0286] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0287] In the development of products using animals, it is desired to accurately grasp the emotions and reactions of animals and improve products based on them. In the conventional method, it is difficult to quantitatively evaluate the emotions of animals, and there is a lack of means to quickly and accurately measure the reactions of animals to products. For this reason, an effective system for analyzing the emotions of animals and proceeding with product development based on them has been demanded.

[0288] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0289] In this invention, the server includes a device for recording the voice of an animal, a device for photographing an image of the animal, and a device for inputting information regarding the behavior of the animal. Thereby, it becomes possible to analyze the emotions of animals in real time and obtain quick feedback in the manufacturing process.

[0290] The "device for recording the voice of an animal" is a device for collecting the voice emitted by an animal and storing it as digital data.

[0291] The "device for photographing an image of the animal" is a device for capturing the appearance and expression of an animal with a camera and recording the visual information as image data.

[0292] "The device for inputting information about the animal's behavior" refers to a device for manually or automatically inputting information about the animal's movements and behavior.

[0293] A "multi-layered artificial intelligence model" is an algorithm that uses a neural network consisting of multiple layers to extract features from data and obtain analysis results.

[0294] "A device that outputs the predicted emotional results in real time during animal testing in the manufacturing process" refers to a device that immediately displays the analyzed emotional results and provides relevant information during animal testing.

[0295] "Noise reduction" is a technique used to remove unwanted noise from recorded audio data and obtain clear audio information.

[0296] The system for implementing this invention is designed as a comprehensive platform for collecting and analyzing animal sounds and images in real time. The server primarily handles this process, first acquiring animal sound and image data through a sound recording device and image capture device connected to a terminal used by the user.

[0297] The server integrates this data into a single data packet and sends it to the cloud infrastructure. On the cloud, the data is preprocessed using software that removes noise from the audio data and processes the image data to the optimal resolution. This enables more accurate analysis.

[0298] Next, the server uses a multi-layered artificial intelligence model to analyze the pre-processed data. This model is specialized for predicting animal emotions and learns animal vocal and image features based on a large amount of training data.

[0299] The resulting emotion predictions are sent to the user's device and displayed through a dedicated application. Users can use these results to improve products or respond immediately to animal reactions.

[0300] As a concrete example, this system could be used in pet food development to evaluate animals' reactions to new prototypes. For instance, a possible prompt might be, "If the animal appears happy due to the new pet food, please output a record of its emotions." In this way, the present invention provides a tool for scientifically analyzing animal emotions and improving product development and communication with pets.

[0301] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0302] Step 1:

[0303] The device records animal sounds using a microphone and captures images of animals with a camera. Input consists of animal sounds and visual data, while output consists of recorded audio and image data. This data is temporarily stored within the device.

[0304] Step 2:

[0305] The terminal integrates the recorded audio and image data, along with animal behavior information input by the user, into a single data packet. The input is the data set obtained in the previous step, and the output is the integrated data packet. The data is organized and prepared for subsequent analysis.

[0306] Step 3:

[0307] The server receives data packets sent from the terminal and performs noise reduction on the voice data. The input for this step is the integrated data packet, and the output is the voice data with the noise removed. Noise cancellation technology is used to improve voice clarity.

[0308] Step 4:

[0309] The server optimizes the resolution of the image data into a format suitable for analysis. The input is the image data within the data packet obtained in the previous step, and the output is the optimized image data. Image processing software performs the resolution adjustment.

[0310] Step 5:

[0311] The server uses a pre-trained multi-layer artificial intelligence model to analyze the pre-processed voice and image data and predict the emotions of the animal. The input is the pre-processed voice data and image data, and the output is the emotion prediction result of the animal. A generative AI model is used to extract features from the data to determine the emotion.

[0312] Step 6:

[0313] The server transmits the predicted emotion result to the terminal, and the terminal displays the result via a dedicated application. The input is the emotion prediction result of the animal, and the output is the display on the application. Visual and immediate feedback is provided to the user.

[0314] Step 7:

[0315] The user determines the response to the animal based on the displayed emotion result and provides feedback if necessary. The input is the emotion prediction result, and the output is the action decision and feedback transmission. Information support for the user to adjust the interaction with the animal is provided.

[0316] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion specific model 59 and perform specific processing using the user's emotion

[0317] This invention relates to a system that incorporates an emotion engine to achieve richer communication by analyzing not only interaction with animals but also the user's emotions. This system analyzes animal sounds and facial expressions in real time and combines the results with the user's emotional information to perform two-way emotion prediction between animals and humans.

[0318] Users use a smartphone or a dedicated device to record animal sounds with a microphone and capture their facial expressions with a camera. The device combines this data into a single data packet and sends it to a cloud server. On the server, the transmitted data is preprocessed to create noise-canceled audio data and optimized image data, which are then analyzed using a deep learning model.

[0319] The server not only predicts the animal's emotions but also recognizes the user's emotions by analyzing the user's facial expressions and voice data acquired using the device's camera or another emotion recognition device. Specifically, it analyzes the user's facial expressions and statements when they are with their pet and performs analysis using an emotion engine. Based on this analysis, it makes a comprehensive judgment about the emotional state of both the animal and the user.

[0320] As a result, the server integrates and presents both the animal's emotions and the user's emotions in a single interface, achieving two-way emotional understanding. For example, if a dog is excited and barking, the server recognizes that the user's expression is smiling and determines that both are having a good time. This information is communicated to the user visually or audibly through the user interface, and pet owners can use it to make better decisions about their time with their pets.

[0321] Users can adjust their interactions with their pets based on the animal and human emotional states presented by the system. They can also contribute to improving the emotion engine by providing feedback when inaccurate results are displayed. This feedback is used to improve the accuracy of the deep learning model in the next learning cycle.

[0322] As described above, the present invention is a significant system that comprehensively analyzes the emotions of animals and humans and provides the results, thereby promoting communication between them.

[0323] The following describes the processing flow.

[0324] Step 1:

[0325] Users record animal sounds using their smartphones or dedicated devices and capture their facial expressions with a camera. The device saves the audio data in WAV format and the image data in JPEG format. Simultaneously, the device uses its camera and emotion recognition device to acquire the user's facial expressions and audio data.

[0326] Step 2:

[0327] The device combines animal audio data, image data, and user emotion data into a single data packet. This data packet includes the pet's ID and time information. The device then sends this data to a cloud server.

[0328] Step 3:

[0329] The server preprocesses the received data packets for analysis. Noise cancellation is applied to animal audio data to improve sound quality. Image data is processed to identify animal expressions using resolution adjustments and face detection algorithms.

[0330] Step 4:

[0331] The server inputs pre-processed animal audio and image data into a deep learning model to predict the animal's emotions. This model is trained on historical data and outputs specific emotion tags (e.g., "excited," "anxious," etc.).

[0332] Step 5:

[0333] The server also analyzes the user's emotional data and uses an emotion engine to recognize the user's current emotions. This engine evaluates the user's facial expressions, tone of voice, and actions to estimate their emotions.

[0334] Step 6:

[0335] The server integrates the emotional data of both the animal and the user and sends it to the terminal as a single emotional result. This allows the user to obtain two-way information that includes not only the animal's emotions but also their own emotional state.

[0336] Step 7:

[0337] The device displays received emotional results on the user interface. If there is a positive emotional match, it provides feedback with emojis or voice messages indicating that state. Users can then use this to adjust their time with their pet.

[0338] Step 8:

[0339] If the results differ from the actual situation, the user provides data to the sentiment engine through feedback. The server uses this feedback as training material to improve accuracy in the next model update.

[0340] (Example 2)

[0341] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0342] Emotional communication between animals and humans is often judged solely on the animal's behavior using conventional technologies, and a challenge lies in the fact that it cannot take into account the user's emotional state. To accurately understand animal emotions, it is necessary to comprehensively analyze both the user's and the animal's emotions and achieve two-way emotional understanding. Furthermore, it is necessary to improve the accuracy of emotion analysis by effectively incorporating feedback.

[0343] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0344] In this invention, the server includes means for inferring the emotions of an animal, means for analyzing the emotions of a user, and means for presenting an integrated emotional state. This makes it possible to understand the emotions of both the animal and the user in a bidirectional manner and to enrich communication.

[0345] A "device for recording animal sounds" is an acoustic collection device that can effectively capture sounds emitted by animals and save them as digital data.

[0346] A "device for capturing images of animals" is an imaging device that captures the appearance and expressions of animals in real time and digitizes them as visual information.

[0347] A "device that analyzes user emotions" is a device that can determine a user's emotional state by analyzing patterns in their facial expressions and voice.

[0348] A "device that bundles and transmits data into data packets" is a device that compresses various forms of digital data into a single, unified format and transmits it to a remote server or other location.

[0349] A "pre-processing and noise removal device" is a device that removes unwanted components from acquired audio and image data to generate high-quality data.

[0350] A "device that analyzes data using deep learning technology" is a device that uses advanced machine learning models, such as neural networks, to analyze input data and extract meaningful information.

[0351] A "device for inferring animal emotions" is a device that can identify and evaluate the psychological state of an animal based on processed data.

[0352] A "device for integrating and presenting information" is a visual or auditory output device that combines multiple pieces of emotional information obtained from analysis and presents them to the user in an easily understandable format.

[0353] This invention provides an interactive system that enables two-way emotional understanding between animals and users. The system collects animal sounds and images and infers the animal's emotions based on this information. Simultaneously, it analyzes the user's facial expressions and voice to determine the user's emotions, thereby constructing an emotional interface between the animal and the user.

[0354] Users can use a dedicated digital device or smartphone to collect animal sounds with a built-in microphone and capture images of animals with a camera. This digital data is combined into a single data packet by the device, encrypted, and sent to a cloud server. On the cloud server, the data is first preprocessed, noise is removed from the audio data, and the image data is optimized into an analyzable format.

[0355] The server analyzes pre-processed audio and image data using a deep learning model to infer the emotions of animals. The deep learning model learns audio and image features and classifies emotions based on them. Next, it determines the user's emotions by analyzing the user's facial expressions and audio data.

[0356] These analysis results are integrated and presented on the user interface. This allows users to understand the emotional state of their animal and their own emotional state at the same time, and adjust their communication with their pet accordingly. Furthermore, users can contribute to improving the deep learning model by submitting feedback, and this feedback will be used to improve the model's accuracy in the next learning cycle.

[0357] For example, if a user records their cat's meows and takes a picture of themselves smiling, the system will infer that the cat is satisfied and sense that the user is also happy. Based on these results, the user can make decisions to improve the quality of their interaction with their cat.

[0358] As an example of a prompt, you could enter: "I filmed my reaction when my dog ​​barked at me. Please analyze this data to tell me the emotions of both my dog ​​and me."

[0359] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0360] Step 1:

[0361] The user initiates interaction with the animal. Specifically, the user uses a smartphone or dedicated device to record the animal's sounds with a microphone and capture the animal's facial expressions with a camera. The input at this time is the animal's sounds and facial expressions, and the output is audio data and image data.

[0362] Step 2:

[0363] The device bundles the acquired animal audio and image data into packets and sends them to the cloud server. This packetization process involves data compression and encryption, ensuring thorough protection of the transmitted data. In this step, the input is the audio and image data, and the output is the data packets sent to the server.

[0364] Step 3:

[0365] The server receives data packets sent from the terminal. It first performs preprocessing on this data. Specific operations include removing noise from audio data to generate clear audio and optimizing image data into a format suitable for analysis. The input is data packets, and the output is preprocessed audio and image data.

[0366] Step 4:

[0367] The server performs analysis by inputting pre-processed data into a deep learning model. In this analysis step, features are extracted from the animal's voice and image, respectively, to infer the animal's emotions. The user's facial expressions and voice are also analyzed to determine the user's emotions. The input is pre-processed audio and image data, and the output is the emotional state of the animal and the user.

[0368] Step 5:

[0369] The server integrates the analyzed animal and user emotional states and presents the summarized results to the user. This process uses visual icons and textual information to make it easy for the user to understand. The input is the emotional states of the animal and the user, and the output is the emotional information displayed in the user interface.

[0370] Step 6:

[0371] The user reviews the presented analysis results and submits feedback as needed. This feedback is used to train the next deep learning model. The input is the analysis results, and the output is the user's feedback information.

[0372] (Application Example 2)

[0373] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0374] In factories, the interface between workers and machines is primarily based on physical operations and simple indicators, lacking emotional communication. This increases worker stress and poses a risk of decreased productivity. Therefore, there is a need for a system that allows workers and robots to understand each other's emotions and collaborate more smoothly.

[0375] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0376] In this invention, the server includes means for recording animal sounds, means for capturing images of animals, and means for acquiring and analyzing the user's facial expressions and voice. This makes it possible to understand the emotions of workers and robots in a two-way manner and improve the work environment.

[0377] "Means for recording animal sounds" refers to a device or mechanism that has the function of acquiring sounds emitted by animals as digital data.

[0378] "Means for photographing images of animals" refers to a device or technology that has the function of acquiring the appearance of an animal as a digital image.

[0379] "Means for inputting information about animal behavior" refers to devices, sensors, or interfaces that have the function of inputting data about the movements and state of animals into a system.

[0380] "Means of sending data in a data packet" refers to a device or technology that has the function of integrating different types of data into a single packet and transmitting it over a network.

[0381] "Means for preprocessing audio and image data" refers to technologies or devices that have functions for filtering or converting acquired audio and images into a format that is easy to analyze.

[0382] "Methods of analysis using deep learning models" refer to computer programs or models that use artificial intelligence technology to analyze large amounts of data and extract specific patterns or features.

[0383] "Methods for predicting animal emotions" refer to technologies and algorithms that have the function of inferring the emotional state of an animal based on acquired data.

[0384] "Means for acquiring and analyzing a user's facial expressions and voice" refers to a device or technology that has the function of acquiring a user's facial movements and speech as digital data, analyzing it, and determining their emotions.

[0385] "Means for achieving two-way emotional understanding" refers to interfaces and analytical technologies that enable the mutual recognition and understanding of emotions between animals and humans.

[0386] "Means for outputting predicted emotional results" refers to a device or interface that has the function of presenting the analyzed emotional state to the user visually or audibly.

[0387] To implement this invention, a device for recording animal sounds and images is used. Specifically, a terminal equipped with a microphone and a camera is used to acquire the sounds and facial expressions emitted by animals. Similarly, the user's facial expressions and voice are also acquired, and the system operates to integrate this data.

[0388] The terminal combines the acquired animal and user audio and image data into a single data packet and sends the data to the cloud server. As a preprocessing step, the server applies noise cancellation to clarify the audio data and converts the image data into a format that is easy to analyze.

[0389] This data is analyzed using a deep learning model to predict the emotions of both animals and users. This utilizes deep learning frameworks such as TensorFlow or PyTorch. Through the user interface, the server presents the predicted emotional states in an integrated manner, allowing the user to adjust their interaction with the animals based on this information.

[0390] For example, if the user is smiling when the animal is barking, the server will determine that both are having a good time and display that information on the interface. Such a system helps pet owners build better relationships with their pets.

[0391] An example of a prompt message is, "The worker is smiling. Think of a response that would allow the robot to generate a positive comment in this situation." This allows for explicit instructions on understanding emotions in a specific situation.

[0392] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0393] Step 1:

[0394] The device uses a microphone and camera to collect animal sounds and images. The input is real-time captured audio and video data, which is then converted into a digital format.

[0395] Step 2:

[0396] Similarly, the user's facial expressions and voice are collected using the device's camera and microphone. The input consists of the user's voice and facial image data, which are recorded as digital data.

[0397] Step 3:

[0398] The terminal combines the acquired animal and user data into a single data packet. The output is this integrated data packet, which is then sent to the cloud server.

[0399] Step 4:

[0400] The server receives data packets and applies noise cancellation to the audio data. The input is a unified data packet, and the output is the noise-removed audio data. An audio processing algorithm performs this task.

[0401] Step 5:

[0402] Simultaneously, the server optimizes the image data, adjusting its brightness and contrast. The input is the image data within the data packet, and the output is an optimized image that is easy to analyze.

[0403] Step 6:

[0404] The server analyzes pre-processed audio and image data using a deep learning model. The input consists of optimized audio and image data. The model predicts the emotions of animals and users from this data.

[0405] Step 7:

[0406] The server outputs the analysis results to the user interface, visually or audibly notifying the user of the animal's and the user's emotions. The output represents the predicted emotional state. The user can then adjust their interaction with the animal based on this information.

[0407] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0408] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0409] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0410] [Third Embodiment]

[0411] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0412] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0413] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0414] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0415] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0416] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0417] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0418] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0419] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0420] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0421] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0422] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0423] This invention relates to a system that records animal sounds and facial expressions using a user's device, analyzes this data, and infers the animal's emotions. The user first records the animal's sounds using a microphone and captures its facial expressions with a camera using a smartphone or dedicated camera device. This data, along with information about the animal's behavior, is input into the device.

[0424] The collected data is integrated into a single data packet at the terminal and sent to a cloud server. The server preprocesses the received data, preparing noise-canceling audio data and image data with optimized resolution. This preprocessed data is then analyzed by a deep learning model to predict the animal's emotions.

[0425] This deep learning model utilizes a multi-layer neural network and is trained on a large amount of training data collected from pet shops and animal shelters. The model includes algorithms that extract features from animal sounds and images and predict corresponding emotions.

[0426] Ultimately, the server sends the animal's emotion prediction results to the user's device, where they are displayed on the user interface via a dedicated application. The user can then decide how to interact with the animal based on the displayed results and send feedback if they feel the prediction is incorrect. This user feedback is recorded in a database and used to improve the model's accuracy.

[0427] As a concrete example, consider a situation where a user owns a dog and the dog is excited and barking. The device records the barking and takes a picture of the dog's facial expression. Next, the user inputs into the app that "the dog likely wants to play." This data is sent to a server, and a deep learning model analyzes it to predict the dog's emotion as "excited," and notifies the user. Based on this result, the user can decide whether to spend time playing with their dog.

[0428] In this way, the present invention enables pet owners to more accurately understand their pets' emotions and needs and respond appropriately by scientifically analyzing their behavior.

[0429] The following describes the processing flow.

[0430] Step 1:

[0431] Data collection

[0432] The device uses a smartphone or dedicated device to record animal sounds with a microphone and capture facial expressions with a camera. The recorded audio data is saved in WAV format, and the captured image data is saved in JPEG format, etc. Based on this collected data, the user inputs information related to the pet's behavior and anticipated emotions.

[0433] Step 2:

[0434] Data packet generation and transmission

[0435] The device combines recorded audio data, image data, and user-entered emotional information into a single data packet. This packet also includes necessary metadata such as a timestamp and pet ID. The generated data packet is then sent to a cloud server.

[0436] Step 3:

[0437] Data preprocessing

[0438] The server receives the transmitted data packets and preprocesses the audio and image data before data analysis. The audio data is noise-canceling to convert it into clear audio. The image data is processed using resolution adjustment and facial recognition algorithms to accurately extract animal expressions.

[0439] Step 4:

[0440] Application of emotion analysis models

[0441] The server uses a deep learning model to analyze pre-processed audio and image data as input. This model is trained to infer emotions from the characteristics of pet vocalizations and facial expressions. Using the trained model, the emotions of animals can be predicted with high accuracy.

[0442] Step 5:

[0443] Sending and displaying results

[0444] The server sends the predicted emotion result to the user's device. The device displays the received emotion result on the user interface and adds a visual indicator to inform the user of the pet's current emotional state.

[0445] Step 6:

[0446] Gathering feedback and improving the learning model

[0447] The user checks whether the displayed emotion result matches the pet's actual emotion. If the result is different, they send the correct emotion through the feedback option. The server records this feedback information in a database and uses it as training data for the next model update to improve the accuracy of the analysis model.

[0448] (Example 1)

[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0450] To understand and appropriately respond to animal emotions, it is necessary to accurately analyze animal vocalizations, facial expressions, and behavioral information. However, conventional methods have made it difficult to scientifically predict animal emotions, potentially leading to misunderstandings of animal needs by pet owners and animal care professionals. This invention aims to solve this problem and provide a method for more accurately predicting animal emotions.

[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0452] In this invention, the server includes a device for collecting animal vibrations, a device for acquiring visual information of the animal, and a device for inputting data related to the animal's activities. This allows the animal's emotions to be scientifically predicted using a multi-layer artificial intelligence model, enabling the owner to respond appropriately.

[0453] "Animal vibrations" refer to sounds and noises emitted by animals, which are recorded using means of detecting sound waves.

[0454] "Visual information" refers to video data, including the facial expressions and postures of animals, acquired by cameras and image sensors.

[0455] "Activity data" refers to information about the behavior and status of animals, which users input based on their observations and reports.

[0456] An "information unit" is a digital piece of information that encompasses audio, visual, and activity-related data, bundled together as a single data packet.

[0457] "Pre-processing" means applying necessary processing or transformations to raw data to make it easier to analyze.

[0458] A "multilayer artificial intelligence model" is an algorithm that uses multiple neural networks to extract features from input data and perform classification and prediction based on the results.

[0459] "Noise reduction" is a technology that removes unwanted noise from audio data and converts it into clearer audio information.

[0460] "Evaluation information" refers to feedback and opinions provided by users, and is data used to improve the model and enhance its accuracy.

[0461] This invention provides a system that infers an animal's emotions based on vibration and visual information collected by a user. The user acquires animal sounds and images using a terminal such as a smart device. Specifically, it is possible to record the sounds and facial expressions of the animal using a microphone and camera mounted on the device.

[0462] The device applies noise reduction technology to recorded audio data and preprocesses visual information by optimizing image resolution. This results in data that can be efficiently analyzed by multi-layer artificial intelligence models, including deep learning. This analysis is performed by a server in the cloud, which receives the data transmitted from the device and uses a generative AI model implementing a multi-layer neural network to predict the animal's emotions.

[0463] The results of this emotion prediction are sent to the user's device and displayed by a dedicated application. This allows the user to understand the animal's emotional state and take appropriate action accordingly.

[0464] As a concrete example, consider a scenario where a user wants to know the emotions of their pet dog at home. The user uses their device to record the dog's barks and take a picture of its facial expression. This data is sent to a server and analyzed by a deep learning model, which predicts the dog is "excited." In this situation, the prompt could be, "Based on this dog's barking and facial expression data, what emotion does the deep learning model predict?"

[0465] Furthermore, user feedback is recorded as evaluation information and helps improve the accuracy of the generated AI model. This feedback allows the system to continuously learn and improve the accuracy of predicting animal emotions.

[0466] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0467] Step 1:

[0468] The user collects vibration and visual information from the animal. The user records the animal's sounds using the microphone on their smart device and captures its facial expressions using the camera. The input consists of audio and image data, which are then stored on the device.

[0469] Step 2:

[0470] The data collected by the terminal is preprocessed. For audio data, noise reduction technology is used to remove unwanted noise and output clear audio data. For image data, the resolution is optimized and prepared for analysis. This process yields audio and image data that can be analyzed.

[0471] Step 3:

[0472] The terminal packets the pre-processed data into a single information unit. This information unit contains pre-processed audio data, image data, and information about the animal's activity entered by the user. This is then sent to the cloud server as output.

[0473] Step 4:

[0474] The server analyzes the information units it receives. The server inputs pre-processed audio and image data into a generating AI model and uses deep learning technology to predict the animal's emotions. The output is an emotion prediction result.

[0475] Step 5:

[0476] The server sends the prediction results to the user's device. The device displays the received emotion prediction results on a user interface via a dedicated application. Based on these results, the user understands the animal's condition and decides on an appropriate response.

[0477] Step 6:

[0478] Users provide feedback as needed. Users input their opinions on the displayed sentiment predictions and send them to the server. The server records this feedback as evaluation information and uses it to improve the generating AI model.

[0479] (Application Example 1)

[0480] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0481] In developing products using animals, it is desirable to accurately understand the animals' emotions and reactions and use that information to improve the products. Conventional methods have made it difficult to quantitatively evaluate animal emotions, and there has been a lack of means to quickly and accurately measure animals' reactions to products. Therefore, there was a need for an effective system to analyze animal emotions and use that information to advance product development.

[0482] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0483] In this invention, the server includes a device for recording animal sounds, a device for capturing images of animals, and a device for inputting information about the animals' behavior. This makes it possible to analyze the animals' emotions in real time and obtain rapid feedback in the manufacturing process.

[0484] A "device for recording animal sounds" is a device that collects sounds emitted by animals and saves them as digital data.

[0485] A "device for capturing images of animals" is a device that uses a camera to capture the appearance and facial expressions of animals and records that visual information as image data.

[0486] "The device for inputting information about the animal's behavior" refers to a device for manually or automatically inputting information about the animal's movements and behavior.

[0487] A "multi-layered artificial intelligence model" is an algorithm that uses a neural network consisting of multiple layers to extract features from data and obtain analysis results.

[0488] "A device that outputs the predicted emotional results in real time during animal testing in the manufacturing process" refers to a device that immediately displays the analyzed emotional results and provides relevant information during animal testing.

[0489] "Noise reduction" is a technique used to remove unwanted noise from recorded audio data and obtain clear audio information.

[0490] The system for implementing this invention is designed as a comprehensive platform for collecting and analyzing animal sounds and images in real time. The server primarily handles this process, first acquiring animal sound and image data through a sound recording device and image capture device connected to a terminal used by the user.

[0491] The server integrates this data into a single data packet and sends it to the cloud infrastructure. On the cloud, the data is preprocessed using software that removes noise from the audio data and processes the image data to the optimal resolution. This enables more accurate analysis.

[0492] Next, the server uses a multi-layered artificial intelligence model to analyze the pre-processed data. This model is specialized for predicting animal emotions and learns animal vocal and image features based on a large amount of training data.

[0493] The resulting emotion predictions are sent to the user's device and displayed through a dedicated application. Users can use these results to improve products or respond immediately to animal reactions.

[0494] As a concrete example, this system could be used in pet food development to evaluate animals' reactions to new prototypes. For instance, a possible prompt might be, "If the animal appears happy due to the new pet food, please output a record of its emotions." In this way, the present invention provides a tool for scientifically analyzing animal emotions and improving product development and communication with pets.

[0495] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0496] Step 1:

[0497] The device records animal sounds using a microphone and captures images of animals with a camera. Input consists of animal sounds and visual data, while output consists of recorded audio and image data. This data is temporarily stored within the device.

[0498] Step 2:

[0499] The terminal integrates the recorded audio and image data, along with animal behavior information input by the user, into a single data packet. The input is the data set obtained in the previous step, and the output is the integrated data packet. The data is organized and prepared for subsequent analysis.

[0500] Step 3:

[0501] The server receives data packets sent from the terminal and performs noise reduction on the voice data. The input for this step is the integrated data packet, and the output is the voice data with the noise removed. Noise cancellation technology is used to improve voice clarity.

[0502] Step 4:

[0503] The server optimizes the resolution of the image data into a format suitable for analysis. The input is the image data in the data packet obtained in the previous step, and the output is the optimized image data. Image processing software performs resolution adjustment.

[0504] Step 5:

[0505] The server uses a pre-trained, multi-layered artificial intelligence model to analyze pre-processed audio and image data and predict animal emotions. The input is pre-processed audio and image data, and the output is the predicted animal emotion. A generative AI model is used to extract features from the data and determine emotions.

[0506] Step 6:

[0507] The server sends predicted emotion results to the terminal, which displays the results via a dedicated application. The input is the predicted emotion result of the animal, and the output is the display on the application. This provides the user with visual and immediate feedback.

[0508] Step 7:

[0509] The user decides how to interact with the animal based on the displayed emotion results and provides feedback as needed. The input is the emotion prediction result, and the output is the action decision and feedback submission. Information support is provided to help the user adjust their interaction with the animal.

[0510] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0511] This invention relates to a system that incorporates an emotion engine to achieve richer communication by analyzing not only interaction with animals but also the user's emotions. This system analyzes animal sounds and facial expressions in real time and combines the results with the user's emotional information to perform two-way emotion prediction between animals and humans.

[0512] Users use a smartphone or a dedicated device to record animal sounds with a microphone and capture their facial expressions with a camera. The device combines this data into a single data packet and sends it to a cloud server. On the server, the transmitted data is preprocessed to create noise-canceled audio data and optimized image data, which are then analyzed using a deep learning model.

[0513] The server not only predicts the animal's emotions but also recognizes the user's emotions by analyzing the user's facial expressions and voice data acquired using the device's camera or another emotion recognition device. Specifically, it analyzes the user's facial expressions and statements when they are with their pet and performs analysis using an emotion engine. Based on this analysis, it makes a comprehensive judgment about the emotional state of both the animal and the user.

[0514] As a result, the server integrates and presents both the animal's emotions and the user's emotions in a single interface, achieving two-way emotional understanding. For example, if a dog is excited and barking, the server recognizes that the user's expression is smiling and determines that both are having a good time. This information is communicated to the user visually or audibly through the user interface, and pet owners can use it to make better decisions about their time with their pets.

[0515] Users can adjust their interactions with their pets based on the animal and human emotional states presented by the system. They can also contribute to improving the emotion engine by providing feedback when inaccurate results are displayed. This feedback is used to improve the accuracy of the deep learning model in the next learning cycle.

[0516] As described above, the present invention is a significant system that comprehensively analyzes the emotions of animals and humans and provides the results, thereby promoting communication between them.

[0517] The following describes the processing flow.

[0518] Step 1:

[0519] Users record animal sounds using their smartphones or dedicated devices and capture their facial expressions with a camera. The device saves the audio data in WAV format and the image data in JPEG format. Simultaneously, the device uses its camera and emotion recognition device to acquire the user's facial expressions and audio data.

[0520] Step 2:

[0521] The device combines animal audio data, image data, and user emotion data into a single data packet. This data packet includes the pet's ID and time information. The device then sends this data to a cloud server.

[0522] Step 3:

[0523] The server preprocesses the received data packets for analysis. Noise cancellation is applied to animal audio data to improve sound quality. Image data is processed to identify animal expressions using resolution adjustments and face detection algorithms.

[0524] Step 4:

[0525] The server inputs pre-processed animal audio and image data into a deep learning model to predict the animal's emotions. This model is trained on historical data and outputs specific emotion tags (e.g., "excited," "anxious," etc.).

[0526] Step 5:

[0527] The server also analyzes the user's emotional data and uses an emotion engine to recognize the user's current emotions. This engine evaluates the user's facial expressions, tone of voice, and actions to estimate their emotions.

[0528] Step 6:

[0529] The server integrates the emotional data of both the animal and the user and sends it to the terminal as a single emotional result. This allows the user to obtain two-way information that includes not only the animal's emotions but also their own emotional state.

[0530] Step 7:

[0531] The device displays received emotional results on the user interface. If there is a positive emotional match, it provides feedback with emojis or voice messages indicating that state. Users can then use this to adjust their time with their pet.

[0532] Step 8:

[0533] If the results differ from the actual situation, the user provides data to the sentiment engine through feedback. The server uses this feedback as training material to improve accuracy in the next model update.

[0534] (Example 2)

[0535] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0536] Emotional communication between animals and humans is often judged solely on the animal's behavior using conventional technologies, and a challenge lies in the fact that it cannot take into account the user's emotional state. To accurately understand animal emotions, it is necessary to comprehensively analyze both the user's and the animal's emotions and achieve two-way emotional understanding. Furthermore, it is necessary to improve the accuracy of emotion analysis by effectively incorporating feedback.

[0537] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0538] In this invention, the server includes means for inferring the emotions of an animal, means for analyzing the emotions of a user, and means for presenting an integrated emotional state. This makes it possible to understand the emotions of both the animal and the user in a bidirectional manner and to enrich communication.

[0539] A "device for recording animal sounds" is an acoustic collection device that can effectively capture sounds emitted by animals and save them as digital data.

[0540] A "device for capturing images of animals" is an imaging device that captures the appearance and expressions of animals in real time and digitizes them as visual information.

[0541] A "device that analyzes user emotions" is a device that can determine a user's emotional state by analyzing patterns in their facial expressions and voice.

[0542] A "device that bundles and transmits data into data packets" is a device that compresses various forms of digital data into a single, unified format and transmits it to a remote server or other location.

[0543] A "pre-processing and noise removal device" is a device that removes unwanted components from acquired audio and image data to generate high-quality data.

[0544] A "device that analyzes data using deep learning technology" is a device that uses advanced machine learning models, such as neural networks, to analyze input data and extract meaningful information.

[0545] A "device for inferring animal emotions" is a device that can identify and evaluate the psychological state of an animal based on processed data.

[0546] A "device for integrating and presenting information" is a visual or auditory output device that combines multiple pieces of emotional information obtained from analysis and presents them to the user in an easily understandable format.

[0547] This invention provides an interactive system that enables two-way emotional understanding between animals and users. The system collects animal sounds and images and infers the animal's emotions based on this information. Simultaneously, it analyzes the user's facial expressions and voice to determine the user's emotions, thereby constructing an emotional interface between the animal and the user.

[0548] Users can use a dedicated digital device or smartphone to collect animal sounds with a built-in microphone and capture images of animals with a camera. This digital data is combined into a single data packet by the device, encrypted, and sent to a cloud server. On the cloud server, the data is first preprocessed, noise is removed from the audio data, and the image data is optimized into an analyzable format.

[0549] The server analyzes pre-processed audio and image data using a deep learning model to infer the emotions of animals. The deep learning model learns audio and image features and classifies emotions based on them. Next, it determines the user's emotions by analyzing the user's facial expressions and audio data.

[0550] These analysis results are integrated and presented on the user interface. This allows users to understand the emotional state of their animal and their own emotional state at the same time, and adjust their communication with their pet accordingly. Furthermore, users can contribute to improving the deep learning model by submitting feedback, and this feedback will be used to improve the model's accuracy in the next learning cycle.

[0551] For example, if a user records their cat's meows and takes a picture of themselves smiling, the system will infer that the cat is satisfied and sense that the user is also happy. Based on these results, the user can make decisions to improve the quality of their interaction with their cat.

[0552] As an example of a prompt, you could enter: "I filmed my reaction when my dog ​​barked at me. Please analyze this data to tell me the emotions of both my dog ​​and me."

[0553] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0554] Step 1:

[0555] The user initiates interaction with the animal. Specifically, the user uses a smartphone or dedicated device to record the animal's sounds with a microphone and capture the animal's facial expressions with a camera. The input at this time is the animal's sounds and facial expressions, and the output is audio data and image data.

[0556] Step 2:

[0557] The device bundles the acquired animal audio and image data into packets and sends them to the cloud server. This packetization process involves data compression and encryption, ensuring thorough protection of the transmitted data. In this step, the input is the audio and image data, and the output is the data packets sent to the server.

[0558] Step 3:

[0559] The server receives data packets sent from the terminal. It first performs preprocessing on this data. Specific operations include removing noise from audio data to generate clear audio and optimizing image data into a format suitable for analysis. The input is data packets, and the output is preprocessed audio and image data.

[0560] Step 4:

[0561] The server performs analysis by inputting pre-processed data into a deep learning model. In this analysis step, features are extracted from the animal's voice and image, respectively, to infer the animal's emotions. The user's facial expressions and voice are also analyzed to determine the user's emotions. The input is pre-processed audio and image data, and the output is the emotional state of the animal and the user.

[0562] Step 5:

[0563] The server integrates the analyzed animal and user emotional states and presents the summarized results to the user. This process uses visual icons and textual information to make it easy for the user to understand. The input is the emotional states of the animal and the user, and the output is the emotional information displayed in the user interface.

[0564] Step 6:

[0565] The user reviews the presented analysis results and submits feedback as needed. This feedback is used to train the next deep learning model. The input is the analysis results, and the output is the user's feedback information.

[0566] (Application Example 2)

[0567] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0568] In factories, the interface between workers and machines is primarily based on physical operations and simple indicators, lacking emotional communication. This increases worker stress and poses a risk of decreased productivity. Therefore, there is a need for a system that allows workers and robots to understand each other's emotions and collaborate more smoothly.

[0569] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0570] In this invention, the server includes means for recording animal sounds, means for capturing images of animals, and means for acquiring and analyzing the user's facial expressions and voice. This makes it possible to understand the emotions of workers and robots in a two-way manner and improve the work environment.

[0571] "Means for recording animal sounds" refers to a device or mechanism that has the function of acquiring sounds emitted by animals as digital data.

[0572] "Means for photographing images of animals" refers to a device or technology that has the function of acquiring the appearance of an animal as a digital image.

[0573] "Means for inputting information about animal behavior" refers to devices, sensors, or interfaces that have the function of inputting data about the movements and state of animals into a system.

[0574] "Means of sending data in a data packet" refers to a device or technology that has the function of integrating different types of data into a single packet and transmitting it over a network.

[0575] "Means for preprocessing audio and image data" refers to technologies or devices that have functions for filtering or converting acquired audio and images into a format that is easy to analyze.

[0576] "Methods of analysis using deep learning models" refer to computer programs or models that use artificial intelligence technology to analyze large amounts of data and extract specific patterns or features.

[0577] "Methods for predicting animal emotions" refer to technologies and algorithms that have the function of inferring the emotional state of an animal based on acquired data.

[0578] "Means for acquiring and analyzing a user's facial expressions and voice" refers to a device or technology that has the function of acquiring a user's facial movements and speech as digital data, analyzing it, and determining their emotions.

[0579] "Means for achieving two-way emotional understanding" refers to interfaces and analytical technologies that enable the mutual recognition and understanding of emotions between animals and humans.

[0580] "Means for outputting predicted emotional results" refers to a device or interface that has the function of presenting the analyzed emotional state to the user visually or audibly.

[0581] To implement this invention, a device for recording animal sounds and images is used. Specifically, a terminal equipped with a microphone and a camera is used to acquire the sounds and facial expressions emitted by animals. Similarly, the user's facial expressions and voice are also acquired, and the system operates to integrate this data.

[0582] The terminal combines the acquired animal and user audio and image data into a single data packet and sends the data to the cloud server. As a preprocessing step, the server applies noise cancellation to clarify the audio data and converts the image data into a format that is easy to analyze.

[0583] This data is analyzed using a deep learning model to predict the emotions of both animals and users. This utilizes deep learning frameworks such as TensorFlow or PyTorch. Through the user interface, the server presents the predicted emotional states in an integrated manner, allowing the user to adjust their interaction with the animals based on this information.

[0584] For example, if the user is smiling when the animal is barking, the server will determine that both are having a good time and display that information on the interface. Such a system helps pet owners build better relationships with their pets.

[0585] An example of a prompt message is, "The worker is smiling. Think of a response that would allow the robot to generate a positive comment in this situation." This allows for explicit instructions on understanding emotions in a specific situation.

[0586] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0587] Step 1:

[0588] The device uses a microphone and camera to collect animal sounds and images. The input is real-time captured audio and video data, which is then converted into a digital format.

[0589] Step 2:

[0590] Similarly, the user's facial expressions and voice are collected using the device's camera and microphone. The input consists of the user's voice and facial image data, which are recorded as digital data.

[0591] Step 3:

[0592] The terminal combines the acquired animal and user data into a single data packet. The output is this integrated data packet, which is then sent to the cloud server.

[0593] Step 4:

[0594] The server receives data packets and applies noise cancellation to the audio data. The input is a unified data packet, and the output is the noise-removed audio data. An audio processing algorithm performs this task.

[0595] Step 5:

[0596] Simultaneously, the server optimizes the image data, adjusting its brightness and contrast. The input is the image data within the data packet, and the output is an optimized image that is easy to analyze.

[0597] Step 6:

[0598] The server analyzes pre-processed audio and image data using a deep learning model. The input consists of optimized audio and image data. The model predicts the emotions of animals and users from this data.

[0599] Step 7:

[0600] The server outputs the analysis results to the user interface, visually or audibly notifying the user of the animal's and the user's emotions. The output represents the predicted emotional state. The user can then adjust their interaction with the animal based on this information.

[0601] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0602] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0603] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0604] [Fourth Embodiment]

[0605] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0606] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0607] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0608] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0609] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0610] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0611] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0612] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0613] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0614] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0615] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0616] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0617] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0618] This invention relates to a system that records animal sounds and facial expressions using a user's device, analyzes this data, and infers the animal's emotions. The user first records the animal's sounds using a microphone and captures its facial expressions with a camera using a smartphone or dedicated camera device. This data, along with information about the animal's behavior, is input into the device.

[0619] The collected data is integrated into a single data packet at the terminal and sent to a cloud server. The server preprocesses the received data, preparing noise-canceling audio data and image data with optimized resolution. This preprocessed data is then analyzed by a deep learning model to predict the animal's emotions.

[0620] This deep learning model utilizes a multi-layer neural network and is trained on a large amount of training data collected from pet shops and animal shelters. The model includes algorithms that extract features from animal sounds and images and predict corresponding emotions.

[0621] Ultimately, the server sends the animal's emotion prediction results to the user's device, where they are displayed on the user interface via a dedicated application. The user can then decide how to interact with the animal based on the displayed results and send feedback if they feel the prediction is incorrect. This user feedback is recorded in a database and used to improve the model's accuracy.

[0622] As a concrete example, consider a situation where a user owns a dog and the dog is excited and barking. The device records the barking and takes a picture of the dog's facial expression. Next, the user inputs into the app that "the dog likely wants to play." This data is sent to a server, and a deep learning model analyzes it to predict the dog's emotion as "excited," and notifies the user. Based on this result, the user can decide whether to spend time playing with their dog.

[0623] In this way, the present invention enables pet owners to more accurately understand their pets' emotions and needs and respond appropriately by scientifically analyzing their behavior.

[0624] The following describes the processing flow.

[0625] Step 1:

[0626] Data collection

[0627] The device uses a smartphone or dedicated device to record animal sounds with a microphone and capture facial expressions with a camera. The recorded audio data is saved in WAV format, and the captured image data is saved in JPEG format, etc. Based on this collected data, the user inputs information related to the pet's behavior and anticipated emotions.

[0628] Step 2:

[0629] Data packet generation and transmission

[0630] The device combines recorded audio data, image data, and user-entered emotional information into a single data packet. This packet also includes necessary metadata such as a timestamp and pet ID. The generated data packet is then sent to a cloud server.

[0631] Step 3:

[0632] Data preprocessing

[0633] The server receives the transmitted data packets and preprocesses the audio and image data before data analysis. The audio data is noise-canceling to convert it into clear audio. The image data is processed using resolution adjustment and facial recognition algorithms to accurately extract animal expressions.

[0634] Step 4:

[0635] Application of emotion analysis models

[0636] The server uses a deep learning model to analyze pre-processed audio and image data as input. This model is trained to infer emotions from the characteristics of pet vocalizations and facial expressions. Using the trained model, the emotions of animals can be predicted with high accuracy.

[0637] Step 5:

[0638] Sending and displaying results

[0639] The server sends the predicted emotion result to the user's device. The device displays the received emotion result on the user interface and adds a visual indicator to inform the user of the pet's current emotional state.

[0640] Step 6:

[0641] Gathering feedback and improving the learning model

[0642] The user checks whether the displayed emotion result matches the pet's actual emotion. If the result is different, they send the correct emotion through the feedback option. The server records this feedback information in a database and uses it as training data for the next model update to improve the accuracy of the analysis model.

[0643] (Example 1)

[0644] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0645] To understand and appropriately respond to animal emotions, it is necessary to accurately analyze animal vocalizations, facial expressions, and behavioral information. However, conventional methods have made it difficult to scientifically predict animal emotions, potentially leading to misunderstandings of animal needs by pet owners and animal care professionals. This invention aims to solve this problem and provide a method for more accurately predicting animal emotions.

[0646] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0647] In this invention, the server includes a device for collecting animal vibrations, a device for acquiring visual information of the animal, and a device for inputting data related to the animal's activities. This allows the animal's emotions to be scientifically predicted using a multi-layer artificial intelligence model, enabling the owner to respond appropriately.

[0648] "Animal vibrations" refer to sounds and noises emitted by animals, which are recorded using means of detecting sound waves.

[0649] "Visual information" refers to video data, including the facial expressions and postures of animals, acquired by cameras and image sensors.

[0650] "Activity data" refers to information about the behavior and status of animals, which users input based on their observations and reports.

[0651] An "information unit" is a digital piece of information that encompasses audio, visual, and activity-related data, bundled together as a single data packet.

[0652] "Pre-processing" means applying necessary processing or transformations to raw data to make it easier to analyze.

[0653] A "multilayer artificial intelligence model" is an algorithm that uses multiple neural networks to extract features from input data and perform classification and prediction based on the results.

[0654] "Noise reduction" is a technology that removes unwanted noise from audio data and converts it into clearer audio information.

[0655] "Evaluation information" refers to feedback and opinions provided by users, and is data used to improve the model and enhance its accuracy.

[0656] This invention provides a system that infers an animal's emotions based on vibration and visual information collected by a user. The user acquires animal sounds and images using a terminal such as a smart device. Specifically, it is possible to record the sounds and facial expressions of the animal using a microphone and camera mounted on the device.

[0657] The device applies noise reduction technology to recorded audio data and preprocesses visual information by optimizing image resolution. This results in data that can be efficiently analyzed by multi-layer artificial intelligence models, including deep learning. This analysis is performed by a server in the cloud, which receives the data transmitted from the device and uses a generative AI model implementing a multi-layer neural network to predict the animal's emotions.

[0658] The results of this emotion prediction are sent to the user's device and displayed by a dedicated application. This allows the user to understand the animal's emotional state and take appropriate action accordingly.

[0659] As a concrete example, consider a scenario where a user wants to know the emotions of their pet dog at home. The user uses their device to record the dog's barks and take a picture of its facial expression. This data is sent to a server and analyzed by a deep learning model, which predicts the dog is "excited." In this situation, the prompt could be, "Based on this dog's barking and facial expression data, what emotion does the deep learning model predict?"

[0660] Furthermore, user feedback is recorded as evaluation information and helps improve the accuracy of the generated AI model. This feedback allows the system to continuously learn and improve the accuracy of predicting animal emotions.

[0661] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0662] Step 1:

[0663] The user collects vibration and visual information from the animal. The user records the animal's sounds using the microphone on their smart device and captures its facial expressions using the camera. The input consists of audio and image data, which are then stored on the device.

[0664] Step 2:

[0665] The data collected by the terminal is preprocessed. For audio data, noise reduction technology is used to remove unwanted noise and output clear audio data. For image data, the resolution is optimized and prepared for analysis. This process yields audio and image data that can be analyzed.

[0666] Step 3:

[0667] The terminal packets the pre-processed data into a single information unit. This information unit contains pre-processed audio data, image data, and information about the animal's activity entered by the user. This is then sent to the cloud server as output.

[0668] Step 4:

[0669] The server analyzes the information units it receives. The server inputs pre-processed audio and image data into a generating AI model and uses deep learning technology to predict the animal's emotions. The output is an emotion prediction result.

[0670] Step 5:

[0671] The server sends the prediction results to the user's device. The device displays the received emotion prediction results on a user interface via a dedicated application. Based on these results, the user understands the animal's condition and decides on an appropriate response.

[0672] Step 6:

[0673] Users provide feedback as needed. Users input their opinions on the displayed sentiment predictions and send them to the server. The server records this feedback as evaluation information and uses it to improve the generating AI model.

[0674] (Application Example 1)

[0675] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0676] In developing products using animals, it is desirable to accurately understand the animals' emotions and reactions and use that information to improve the products. Conventional methods have made it difficult to quantitatively evaluate animal emotions, and there has been a lack of means to quickly and accurately measure animals' reactions to products. Therefore, there was a need for an effective system to analyze animal emotions and use that information to advance product development.

[0677] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0678] In this invention, the server includes a device for recording animal sounds, a device for capturing images of animals, and a device for inputting information about the animals' behavior. This makes it possible to analyze the animals' emotions in real time and obtain rapid feedback in the manufacturing process.

[0679] A "device for recording animal sounds" is a device that collects sounds emitted by animals and saves them as digital data.

[0680] A "device for capturing images of animals" is a device that uses a camera to capture the appearance and facial expressions of animals and records that visual information as image data.

[0681] "The device for inputting information about the animal's behavior" refers to a device for manually or automatically inputting information about the animal's movements and behavior.

[0682] A "multi-layered artificial intelligence model" is an algorithm that uses a neural network consisting of multiple layers to extract features from data and obtain analysis results.

[0683] "A device that outputs the predicted emotional results in real time during animal testing in the manufacturing process" refers to a device that immediately displays the analyzed emotional results and provides relevant information during animal testing.

[0684] "Noise reduction" is a technique used to remove unwanted noise from recorded audio data and obtain clear audio information.

[0685] The system for implementing this invention is designed as a comprehensive platform for collecting and analyzing animal sounds and images in real time. The server primarily handles this process, first acquiring animal sound and image data through a sound recording device and image capture device connected to a terminal used by the user.

[0686] The server integrates this data into a single data packet and sends it to the cloud infrastructure. On the cloud, the data is preprocessed using software that removes noise from the audio data and processes the image data to the optimal resolution. This enables more accurate analysis.

[0687] Next, the server uses a multi-layered artificial intelligence model to analyze the pre-processed data. This model is specialized for predicting animal emotions and learns animal vocal and image features based on a large amount of training data.

[0688] The resulting emotion predictions are sent to the user's device and displayed through a dedicated application. Users can use these results to improve products or respond immediately to animal reactions.

[0689] As a concrete example, this system could be used in pet food development to evaluate animals' reactions to new prototypes. For instance, a possible prompt might be, "If the animal appears happy due to the new pet food, please output a record of its emotions." In this way, the present invention provides a tool for scientifically analyzing animal emotions and improving product development and communication with pets.

[0690] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0691] Step 1:

[0692] The device records animal sounds using a microphone and captures images of animals with a camera. Input consists of animal sounds and visual data, while output consists of recorded audio and image data. This data is temporarily stored within the device.

[0693] Step 2:

[0694] The terminal integrates the recorded audio and image data, along with animal behavior information input by the user, into a single data packet. The input is the data set obtained in the previous step, and the output is the integrated data packet. The data is organized and prepared for subsequent analysis.

[0695] Step 3:

[0696] The server receives data packets sent from the terminal and performs noise reduction on the voice data. The input for this step is the integrated data packet, and the output is the voice data with the noise removed. Noise cancellation technology is used to improve voice clarity.

[0697] Step 4:

[0698] The server optimizes the resolution of the image data into a format suitable for analysis. The input is the image data in the data packet obtained in the previous step, and the output is the optimized image data. Image processing software performs resolution adjustment.

[0699] Step 5:

[0700] The server uses a pre-trained, multi-layered artificial intelligence model to analyze pre-processed audio and image data and predict animal emotions. The input is pre-processed audio and image data, and the output is the predicted animal emotion. A generative AI model is used to extract features from the data and determine emotions.

[0701] Step 6:

[0702] The server sends predicted emotion results to the terminal, which displays the results via a dedicated application. The input is the predicted emotion result of the animal, and the output is the display on the application. This provides the user with visual and immediate feedback.

[0703] Step 7:

[0704] The user decides how to interact with the animal based on the displayed emotion results and provides feedback as needed. The input is the emotion prediction result, and the output is the action decision and feedback submission. Information support is provided to help the user adjust their interaction with the animal.

[0705] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0706] This invention relates to a system that incorporates an emotion engine to achieve richer communication by analyzing not only interaction with animals but also the user's emotions. This system analyzes animal sounds and facial expressions in real time and combines the results with the user's emotional information to perform two-way emotion prediction between animals and humans.

[0707] Users use a smartphone or a dedicated device to record animal sounds with a microphone and capture their facial expressions with a camera. The device combines this data into a single data packet and sends it to a cloud server. On the server, the transmitted data is preprocessed to create noise-canceled audio data and optimized image data, which are then analyzed using a deep learning model.

[0708] The server not only predicts the animal's emotions but also recognizes the user's emotions by analyzing the user's facial expressions and voice data acquired using the device's camera or another emotion recognition device. Specifically, it analyzes the user's facial expressions and statements when they are with their pet and performs analysis using an emotion engine. Based on this analysis, it makes a comprehensive judgment about the emotional state of both the animal and the user.

[0709] As a result, the server integrates and presents both the animal's emotions and the user's emotions in a single interface, achieving two-way emotional understanding. For example, if a dog is excited and barking, the server recognizes that the user's expression is smiling and determines that both are having a good time. This information is communicated to the user visually or audibly through the user interface, and pet owners can use it to make better decisions about their time with their pets.

[0710] Users can adjust their interactions with their pets based on the animal and human emotional states presented by the system. They can also contribute to improving the emotion engine by providing feedback when inaccurate results are displayed. This feedback is used to improve the accuracy of the deep learning model in the next learning cycle.

[0711] As described above, the present invention is a significant system that comprehensively analyzes the emotions of animals and humans and provides the results, thereby promoting communication between them.

[0712] The following describes the processing flow.

[0713] Step 1:

[0714] Users record animal sounds using their smartphones or dedicated devices and capture their facial expressions with a camera. The device saves the audio data in WAV format and the image data in JPEG format. Simultaneously, the device uses its camera and emotion recognition device to acquire the user's facial expressions and audio data.

[0715] Step 2:

[0716] The device combines animal audio data, image data, and user emotion data into a single data packet. This data packet includes the pet's ID and time information. The device then sends this data to a cloud server.

[0717] Step 3:

[0718] The server preprocesses the received data packets for analysis. Noise cancellation is applied to animal audio data to improve sound quality. Image data is processed to identify animal expressions using resolution adjustments and face detection algorithms.

[0719] Step 4:

[0720] The server inputs pre-processed animal audio and image data into a deep learning model to predict the animal's emotions. This model is trained on historical data and outputs specific emotion tags (e.g., "excited," "anxious," etc.).

[0721] Step 5:

[0722] The server also analyzes the user's emotional data and uses an emotion engine to recognize the user's current emotions. This engine evaluates the user's facial expressions, tone of voice, and actions to estimate their emotions.

[0723] Step 6:

[0724] The server integrates the emotional data of both the animal and the user and sends it to the terminal as a single emotional result. This allows the user to obtain two-way information that includes not only the animal's emotions but also their own emotional state.

[0725] Step 7:

[0726] The device displays received emotional results on the user interface. If there is a positive emotional match, it provides feedback with emojis or voice messages indicating that state. Users can then use this to adjust their time with their pet.

[0727] Step 8:

[0728] If the results differ from the actual situation, the user provides data to the sentiment engine through feedback. The server uses this feedback as training material to improve accuracy in the next model update.

[0729] (Example 2)

[0730] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0731] Emotional communication between animals and humans is often judged solely on the animal's behavior using conventional technologies, and a challenge lies in the fact that it cannot take into account the user's emotional state. To accurately understand animal emotions, it is necessary to comprehensively analyze both the user's and the animal's emotions and achieve two-way emotional understanding. Furthermore, it is necessary to improve the accuracy of emotion analysis by effectively incorporating feedback.

[0732] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0733] In this invention, the server includes means for inferring the emotions of an animal, means for analyzing the emotions of a user, and means for presenting an integrated emotional state. This makes it possible to understand the emotions of both the animal and the user in a bidirectional manner and to enrich communication.

[0734] A "device for recording animal sounds" is an acoustic collection device that can effectively capture sounds emitted by animals and save them as digital data.

[0735] A "device for capturing images of animals" is an imaging device that captures the appearance and expressions of animals in real time and digitizes them as visual information.

[0736] A "device that analyzes user emotions" is a device that can determine a user's emotional state by analyzing patterns in their facial expressions and voice.

[0737] A "device that bundles and transmits data into data packets" is a device that compresses various forms of digital data into a single, unified format and transmits it to a remote server or other location.

[0738] A "pre-processing and noise removal device" is a device that removes unwanted components from acquired audio and image data to generate high-quality data.

[0739] A "device that analyzes data using deep learning technology" is a device that uses advanced machine learning models, such as neural networks, to analyze input data and extract meaningful information.

[0740] A "device for inferring animal emotions" is a device that can identify and evaluate the psychological state of an animal based on processed data.

[0741] A "device for integrating and presenting information" is a visual or auditory output device that combines multiple pieces of emotional information obtained from analysis and presents them to the user in an easily understandable format.

[0742] This invention provides an interactive system that enables two-way emotional understanding between animals and users. The system collects animal sounds and images and infers the animal's emotions based on this information. Simultaneously, it analyzes the user's facial expressions and voice to determine the user's emotions, thereby constructing an emotional interface between the animal and the user.

[0743] Users can use a dedicated digital device or smartphone to collect animal sounds with a built-in microphone and capture images of animals with a camera. This digital data is combined into a single data packet by the device, encrypted, and sent to a cloud server. On the cloud server, the data is first preprocessed, noise is removed from the audio data, and the image data is optimized into an analyzable format.

[0744] The server analyzes pre-processed audio and image data using a deep learning model to infer the emotions of animals. The deep learning model learns audio and image features and classifies emotions based on them. Next, it determines the user's emotions by analyzing the user's facial expressions and audio data.

[0745] These analysis results are integrated and presented on the user interface. This allows users to understand the emotional state of their animal and their own emotional state at the same time, and adjust their communication with their pet accordingly. Furthermore, users can contribute to improving the deep learning model by submitting feedback, and this feedback will be used to improve the model's accuracy in the next learning cycle.

[0746] For example, if a user records their cat's meows and takes a picture of themselves smiling, the system will infer that the cat is satisfied and sense that the user is also happy. Based on these results, the user can make decisions to improve the quality of their interaction with their cat.

[0747] As an example of a prompt, you could enter: "I filmed my reaction when my dog ​​barked at me. Please analyze this data to tell me the emotions of both my dog ​​and me."

[0748] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0749] Step 1:

[0750] The user initiates interaction with the animal. Specifically, the user uses a smartphone or dedicated device to record the animal's sounds with a microphone and capture the animal's facial expressions with a camera. The input at this time is the animal's sounds and facial expressions, and the output is audio data and image data.

[0751] Step 2:

[0752] The device bundles the acquired animal audio and image data into packets and sends them to the cloud server. This packetization process involves data compression and encryption, ensuring thorough protection of the transmitted data. In this step, the input is the audio and image data, and the output is the data packets sent to the server.

[0753] Step 3:

[0754] The server receives data packets sent from the terminal. It first performs preprocessing on this data. Specific operations include removing noise from audio data to generate clear audio and optimizing image data into a format suitable for analysis. The input is data packets, and the output is preprocessed audio and image data.

[0755] Step 4:

[0756] The server performs analysis by inputting pre-processed data into a deep learning model. In this analysis step, features are extracted from the animal's voice and image, respectively, to infer the animal's emotions. The user's facial expressions and voice are also analyzed to determine the user's emotions. The input is pre-processed audio and image data, and the output is the emotional state of the animal and the user.

[0757] Step 5:

[0758] The server integrates the analyzed animal and user emotional states and presents the summarized results to the user. This process uses visual icons and textual information to make it easy for the user to understand. The input is the emotional states of the animal and the user, and the output is the emotional information displayed in the user interface.

[0759] Step 6:

[0760] The user reviews the presented analysis results and submits feedback as needed. This feedback is used to train the next deep learning model. The input is the analysis results, and the output is the user's feedback information.

[0761] (Application Example 2)

[0762] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0763] In factories, the interface between workers and machines is primarily based on physical operations and simple indicators, lacking emotional communication. This increases worker stress and poses a risk of decreased productivity. Therefore, there is a need for a system that allows workers and robots to understand each other's emotions and collaborate more smoothly.

[0764] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0765] In this invention, the server includes means for recording animal sounds, means for capturing images of animals, and means for acquiring and analyzing the user's facial expressions and voice. This makes it possible to understand the emotions of workers and robots in a two-way manner and improve the work environment.

[0766] "Means for recording animal sounds" refers to a device or mechanism that has the function of acquiring sounds emitted by animals as digital data.

[0767] "Means for photographing images of animals" refers to a device or technology that has the function of acquiring the appearance of an animal as a digital image.

[0768] "Means for inputting information about animal behavior" refers to devices, sensors, or interfaces that have the function of inputting data about the movements and state of animals into a system.

[0769] "Means of sending data in a data packet" refers to a device or technology that has the function of integrating different types of data into a single packet and transmitting it over a network.

[0770] "Means for preprocessing audio and image data" refers to technologies or devices that have functions for filtering or converting acquired audio and images into a format that is easy to analyze.

[0771] "Methods of analysis using deep learning models" refer to computer programs or models that use artificial intelligence technology to analyze large amounts of data and extract specific patterns or features.

[0772] "Methods for predicting animal emotions" refer to technologies and algorithms that have the function of inferring the emotional state of an animal based on acquired data.

[0773] "Means for acquiring and analyzing a user's facial expressions and voice" refers to a device or technology that has the function of acquiring a user's facial movements and speech as digital data, analyzing it, and determining their emotions.

[0774] "Means for achieving two-way emotional understanding" refers to interfaces and analytical technologies that enable the mutual recognition and understanding of emotions between animals and humans.

[0775] "Means for outputting predicted emotional results" refers to a device or interface that has the function of presenting the analyzed emotional state to the user visually or audibly.

[0776] To implement this invention, a device for recording animal sounds and images is used. Specifically, a terminal equipped with a microphone and a camera is used to acquire the sounds and facial expressions emitted by animals. Similarly, the user's facial expressions and voice are also acquired, and the system operates to integrate this data.

[0777] The terminal combines the acquired animal and user audio and image data into a single data packet and sends the data to the cloud server. As a preprocessing step, the server applies noise cancellation to clarify the audio data and converts the image data into a format that is easy to analyze.

[0778] This data is analyzed using a deep learning model to predict the emotions of both animals and users. This utilizes deep learning frameworks such as TensorFlow or PyTorch. Through the user interface, the server presents the predicted emotional states in an integrated manner, allowing the user to adjust their interaction with the animals based on this information.

[0779] For example, if the user is smiling when the animal is barking, the server will determine that both are having a good time and display that information on the interface. Such a system helps pet owners build better relationships with their pets.

[0780] An example of a prompt message is, "The worker is smiling. Think of a response that would allow the robot to generate a positive comment in this situation." This allows for explicit instructions on understanding emotions in a specific situation.

[0781] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0782] Step 1:

[0783] The device uses a microphone and camera to collect animal sounds and images. The input is real-time captured audio and video data, which is then converted into a digital format.

[0784] Step 2:

[0785] Similarly, the user's facial expressions and voice are collected using the device's camera and microphone. The input consists of the user's voice and facial image data, which are recorded as digital data.

[0786] Step 3:

[0787] The terminal combines the acquired animal and user data into a single data packet. The output is this integrated data packet, which is then sent to the cloud server.

[0788] Step 4:

[0789] The server receives data packets and applies noise cancellation to the audio data. The input is a unified data packet, and the output is the noise-removed audio data. An audio processing algorithm performs this task.

[0790] Step 5:

[0791] Simultaneously, the server optimizes the image data, adjusting its brightness and contrast. The input is the image data within the data packet, and the output is an optimized image that is easy to analyze.

[0792] Step 6:

[0793] The server analyzes pre-processed audio and image data using a deep learning model. The input consists of optimized audio and image data. The model predicts the emotions of animals and users from this data.

[0794] Step 7:

[0795] The server outputs the analysis results to the user interface, visually or audibly notifying the user of the animal's and the user's emotions. The output represents the predicted emotional state. The user can then adjust their interaction with the animal based on this information.

[0796] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0797] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0798] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0799] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0800] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0801] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0802] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0803] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0804] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0805] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0806] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0807] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0808] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0809] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0810] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0811] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0812] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0813] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0814] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0815] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0816] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0817] The following is further disclosed regarding the embodiments described above.

[0818] (Claim 1)

[0819] Means for recording animal sounds,

[0820] Means of taking pictures of animals,

[0821] Means for inputting information regarding the behavior of the aforementioned animals,

[0822] A means for transmitting the recorded audio, captured images, and input information together in a single data packet,

[0823] Means for preprocessing the aforementioned audio data and image data,

[0824] A method for predicting animal emotions by analyzing data preprocessed using a deep learning model,

[0825] means for outputting the predicted emotion result,

[0826] A system that includes this.

[0827] (Claim 2)

[0828] The system according to claim 1, which uses user feedback to update the deep learning model.

[0829] (Claim 3)

[0830] The system according to claim 1, wherein noise cancellation is applied to the aforementioned audio data.

[0831] "Example 1"

[0832] (Claim 1)

[0833] A device for collecting animal vibrations,

[0834] A device for acquiring visual information from animals,

[0835] A device for inputting data related to the activities of the aforementioned animals,

[0836] A device that combines the recorded vibrations, acquired visual information, and input data into a single information unit and communicates them,

[0837] A device for pre-processing the aforementioned vibration data and visual information data,

[0838] A device that analyzes pre-processed information using a multi-layer artificial intelligence model to predict animal emotions,

[0839] A device for displaying the predicted emotion result,

[0840] A system that includes this.

[0841] (Claim 2)

[0842] The system according to claim 1, which uses user evaluation information to improve the aforementioned multi-layer artificial intelligence model.

[0843] (Claim 3)

[0844] The system according to claim 1, wherein noise reduction is applied to the vibration data.

[0845] "Application Example 1"

[0846] (Claim 1)

[0847] A device for recording animal sounds,

[0848] A device for taking pictures of animals,

[0849] A device for inputting information about the behavior of the aforementioned animals,

[0850] A device that combines the recorded audio, captured images, and input information into a single data packet and transmits it,

[0851] A device for preprocessing the aforementioned audio data and image data,

[0852] A device that analyzes preprocessed data using a multi-layered artificial intelligence model to predict animal emotions,

[0853] A device that outputs the predicted emotion results in real time during animal testing in the manufacturing process,

[0854] A system that includes this.

[0855] (Claim 2)

[0856] The system according to claim 1, which uses user feedback to update the aforementioned multi-layered artificial intelligence model.

[0857] (Claim 3)

[0858] The system according to claim 1, wherein noise reduction is applied to the aforementioned audio data.

[0859] "Example 2 of combining an emotion engine"

[0860] (Claim 1)

[0861] A device for recording animal sounds,

[0862] A device for taking pictures of animals,

[0863] A device that analyzes user emotions,

[0864] A device that combines the recorded animal sounds and the photographed animal images into a single data packet and transmits it,

[0865] A device for preprocessing the aforementioned audio data and image data and removing noise,

[0866] A device that analyzes pre-processed data using deep learning technology to predict animal emotions,

[0867] A device that integrates and presents the emotions of the estimated animal and the emotions of the analyzed user,

[0868] A system that includes this.

[0869] (Claim 2)

[0870] The system according to claim 1, which uses user feedback to update the deep learning technology.

[0871] (Claim 3)

[0872] The system according to claim 1, further comprising a device for analyzing the user's facial expressions and voice in order to recognize the user's emotional state.

[0873] "Application example 2 when combining with an emotional engine"

[0874] (Claim 1)

[0875] Means for recording animal sounds,

[0876] Means of taking pictures of animals,

[0877] Means for inputting information regarding the behavior of the aforementioned animals,

[0878] A means for transmitting the recorded audio, captured images, and input information together in a single data packet,

[0879] Means for preprocessing the aforementioned audio data and image data,

[0880] A method for predicting animal emotions by analyzing data preprocessed using a deep learning model,

[0881] A means for acquiring and analyzing the user's facial expressions and voice,

[0882] A means to achieve two-way emotional understanding by comprehensively analyzing the emotions of animals and users,

[0883] A means of outputting predicted emotional results,

[0884] A system that includes this.

[0885] (Claim 2)

[0886] The system according to claim 1, which uses user feedback to update the deep learning model.

[0887] (Claim 3)

[0888] The system according to claim 1, wherein noise cancellation is applied to the aforementioned audio data. [Explanation of Symbols]

[0889] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for recording animal sounds, Means of taking pictures of animals, Means for inputting information regarding the behavior of the aforementioned animals, A means for transmitting the recorded audio, captured images, and input information together in a single data packet, Means for preprocessing the aforementioned audio data and image data, A method for predicting animal emotions by analyzing data preprocessed using a deep learning model, means for outputting the predicted emotion result, A system that includes this.

2. The system according to claim 1, wherein user feedback is used to update the deep learning model.

3. The system according to claim 1, wherein noise cancellation is applied to the aforementioned audio data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A