System

The system addresses the challenge of providing real-time, accurate responses to visitor questions about animal status by integrating data collection, video analysis, and generative AI for immediate and tailored answers, improving visitor experience.

JP2026022341APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123858
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Current systems in zoos and aquariums struggle to provide real-time, accurate, and tailored responses to visitor questions about animal status, especially for children, due to delays and lack of technology that can monitor animal conditions efficiently.

Method used

A system that includes data collection, real-time video analysis, voice input conversion, and generative AI to generate and provide voice responses based on animal conditions, using cameras, microphones, and AI models for immediate and appropriate answers.

Benefits of technology

Enables zoos and aquariums to provide quick and accurate responses to user questions, enhancing visitor satisfaction and interaction by leveraging real-time animal monitoring and AI-generated answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022341000001_ABST
    Figure 2026022341000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system includes means for collecting and storing rearing data, means for acquiring an image of an animal in real time, means for analyzing the acquired image to determine a state of the animal, means for receiving a question from a user as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined state data of the animal, and means for providing the generated response to the user by voice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] At zoos and aquariums, visitors often ask questions about the current status and activity of animals out of curiosity. However, current systems often experience delays or difficulty in accurately answering questions. This reduces visitor satisfaction and limits the zoo or aquarium experience. Furthermore, especially when children ask questions, immediate responses that are tailored to their interests are important. However, there is still a lack of technology that can grasp the latest status of animals in real time and automatically provide appropriate answers based on that information. Thus, there is a need for a system that can efficiently and accurately monitor the status of animals and quickly and appropriately respond to user questions. [Means for solving the problem]

[0005] The present invention provides a system that includes means for collecting and storing animal care data, means for capturing video of animals in real time, means for analyzing the captured video to determine the animal's condition, means for receiving user questions as voice input, means for converting the voice input into text, means for generating responses using the converted text and the determined animal condition data, and means for providing the generated responses to the user via voice. This enables zoos and aquariums to grasp the latest animal conditions in real time when users ask questions and instantly provide appropriate responses based on that information, significantly improving user satisfaction. Furthermore, the use of generative AI models can improve the accuracy and flexibility of responses to questions, making the user experience richer and more interactive.

[0006] "Breeding data" refers to information about the growth of individual and group animals, such as their health condition, food intake, exercise, and sleep time.

[0007] "Means for acquiring" refers to a method or apparatus for capturing real-time video of an animal using a device such as a camera and inputting the video data into the system.

[0008] "Means for analyzing video footage" refers to a method or device for analyzing the behavior and state of an animal based on the acquired video data, and determining whether the animal is resting, active, etc., using, for example, a deep learning model.

[0009] "User questions" refer to questions that zoo and aquarium visitors have about the condition or activity of animals and ask the system via voice.

[0010] "Means for receiving voice input" refers to a method or apparatus for collecting a user's voice questions using a device such as a microphone and incorporating that voice data into the system.

[0011] "Means for converting to text" refers to a method or device for converting acquired voice data into character data using voice recognition technology.

[0012] "Means for generating a response" refers to a method or device for creating an appropriate response based on the user's question text and the animal's condition data, using a generative AI model or similar.

[0013] "Means for providing by voice" refers to a method or device for reproducing the generated text data response using speech synthesis technology and conveying it to the user through a device such as a speaker. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] The present invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for providing prompt and appropriate responses to questions from users. This system is composed of a server, terminals, and users.

[0036] Overall system overview

[0037] The system consists of the following main parts:

[0038] 1. Server:

[0039] Collect and store breeding data.

[0040] A generative AI model is used to generate appropriate responses to user questions.

[0041] 2. Terminal:

[0042] It uses a camera and microphone to capture footage of animals and user voice questions.

[0043] The acquired data is analyzed and sent to the server.

[0044] The response received from the server is synthesized into voice and provided to the user.

[0045] 3. User:

[0046] Ask questions about animals by speaking them into the device.

[0047] Program processing

[0048] Collection and storage of breeding data

[0049] server

[0050] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[0051] Examples:

[0052] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[0053] Animal imaging and analysis

[0054] Terminal

[0055] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[0056] Examples:

[0057] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[0058] Analysis of voice questions from users

[0059] User

[0060] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[0061] Terminal

[0062] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[0063] Examples:

[0064] A child asks the device, "What is this lion doing right now?" and the device converts the voice into text data.

[0065] Generating and serving the response

[0066] server

[0067] The server uses a generative AI model to generate a response based on the user's question text and the animal's status data. For example, if it determines that the lion is currently resting, it generates the response "The lion is currently resting. Please keep quiet."

[0068] Terminal

[0069] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[0070] Examples:

[0071] The device receives the response generated by the server, "The lion is currently resting. Please be quiet," converts it into voice and conveys it to the user.

[0072] Processing flow through concrete examples

[0073] 1. The server stores the lion's health data in a database.

[0074] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[0075] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[0076] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[0077] 5. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please keep quiet."

[0078] 6. The device converts the response into audio and transmits it to the child through the speaker.

[0079] In this way, the system of the present invention improves the user experience at zoos and aquariums, providing quick and appropriate responses based on the animals' current status.

[0080] The processing flow will be explained below.

[0081] Step 1:

[0082] server

[0083] Collects animal care data. The server receives input from keepers and data from sensors, and stores information about the animals' health status (e.g., amount of food eaten, amount of exercise, amount of sleep, etc.) in a database. For example, keepers enter the amount of food eaten each day using a web form, and the amount of exercise automatically recorded by sensors is stored in a database such as MySQL.

[0084] Step 2:

[0085] Terminal

[0086] Acquire video of the animal. The device acquires video in real time using a camera installed in the animal's cage, etc. The camera can be an IP camera or a USB-connected webcam. The device processes the video stream using libraries such as OpenCV and formats the data into a format that can be analyzed in real time.

[0087] Step 3:

[0088] Terminal

[0089] The device analyzes the video to determine the animal's state. The device analyzes the animal's movements and state based on the acquired video data. For example, it uses deep learning models (such as YOLO or PoseNet) to detect the animal's posture and movement and determine its state, such as "active," "resting," or "sleeping."

[0090] Step 4:

[0091] User

[0092] Users (especially children) can speak questions into the device by talking to the microphone and asking questions about the animal's condition or activity, such as "What is this penguin doing right now?"

[0093] Step 5:

[0094] Terminal

[0095] Converts voice to text. The device converts the user's voice questions captured through the microphone into text data using voice recognition technology such as the Google Speech-to-Text API.

[0096] Step 6:

[0097] Terminal

[0098] Send the text and animal status data to the server. The converted text data and the real-time animal status information are compiled and sent to the server using an HTTP POST request.

[0099] Step 7:

[0100] server

[0101] Generate a response: The server uses a generative AI model to generate an appropriate response based on the received text data and animal status data. For example, it generates a response such as "This penguin is currently resting."

[0102] Step 8:

[0103] server

[0104] The generated response is sent to the terminal. The generated response in text data format is returned to the terminal as an HTTP response.

[0105] Step 9:

[0106] Terminal

[0107] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[0108] Step 10:

[0109] Terminal

[0110] Providing a voice response to the user: The terminal plays the voice-converted response to the user through a speaker, providing an answer to the question.

[0111] For example, if a user asks the device, "What is this lion doing right now?", the system will analyze the lion's status in real time, the server will generate a response such as "The lion is currently resting," and the device will play back the response aloud. In this way, the user can obtain accurate information based on the animal's latest status.

[0112] Example 1

[0113] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0114] In conventional zoos and aquariums, it has been difficult to grasp the condition of animals in real time and respond quickly and accurately to visitor questions. In addition, it has not been possible to integrate and utilize breeding data and real-time animal behavior data, which has limited the ability to improve visitor satisfaction and streamline breeding management.

[0115] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0116] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the condition of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined animal condition data, and means for providing the generated response to the user by voice, thereby enabling zoos and aquariums to provide quick and appropriate responses based on the latest condition of animals.

[0117] "Breeding data" refers to data that includes information on the health, food intake, and exercise of animals kept in zoos and aquariums.

[0118] "Real-time" refers to the instantaneous acquisition and processing of information and situations at the current time.

[0119] "Animal condition" is a concept that includes information such as the animal's health, behavior, and activity status.

[0120] "Audio input" refers to speech produced by a user through a microphone or other audio capture device.

[0121] "Means for converting voice input to text" refers to technologies or systems that convert voice data into character data.

[0122] "Means for generating a response" refers to the technology or mechanism for generating an appropriate response to a user's question.

[0123] A "generative AI model" refers to an artificial intelligence model that uses natural language processing technology to understand and generate human language.

[0124] "Analytical algorithm" refers to a computational procedure or method for analyzing data and deriving a particular result or information.

[0125] System Overview

[0126] The present invention is a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate answers to questions from users. The system is composed of a server, terminals, and users, and has the following functions:

[0127] server

[0128] The server has the following functions:

[0129] 1. Collection and storage of rearing data:

[0130] The server receives breeding data entered by keepers using a dedicated application or web interface. For example, it periodically collects information such as the lions' diet, exercise, and weight. It also collects animal health data (heart rate, body temperature, etc.) obtained from sensors. The collected data is stored in a relational database such as MySQL or PostgreSQL.

[0131] Terminal

[0132] The device has the following features:

[0133] 1. Animal footage acquisition:

[0134] The device captures real-time images of animals using cameras installed in cages or aquariums, and the camera video data is temporarily stored in the device's local storage.

[0135] 2. Animal video analysis:

[0136] The device analyzes the stored video data to determine the animal's behavior using OpenCV and deep learning models (such as YOLO and PoseNet), allowing it to determine in real time whether the animal is currently active, resting, or performing a specific behavior.

[0137] 3. Acquire voice questions from users and convert them into text:

[0138] The user speaks to the device to ask a question about an animal. For example, "What is this lion doing right now?" The device then converts the spoken question picked up by the microphone into text. This process uses the Google Speech-to-Text API or IBM Watson Speech to Text.

[0139] 4. Speech synthesis and delivery of responses:

[0140] The response sent from the server is converted into speech using the Google Text-to-Speech API or Amazon Polly, and the generated speech data is then provided to the user through the speaker.

[0141] Specific examples of programs

[0142] Example of user interaction:

[0143] For example, if a user asks the terminal, "What is this lion doing right now?", the process is as follows:

[0144] 1. Acquiring a voice question: The user speaks into the device's microphone, "What is this lion doing right now?"

[0145] 2. Converting voice questions into text: The device sends the captured voice to the Google Speech-to-Text API and receives the text "What is this lion doing right now?"

[0146] 3. Determining the animal's state: The device analyzes the video data using OpenCV and determines that the lion is "resting."

[0147] 4. Generating and providing a response: The server inputs the data "The lion is currently resting" into the generative AI model, generates a response "The lion is currently resting. Please be quiet," and the device converts the response into audio and plays it through the speaker.

[0148] Prompt Sentence Examples

[0149] An example of a prompt sentence might be:

[0150] 1. "What is this lion doing now?"

[0151] 2. "When is the bear's meal time?"

[0152] 3. "Where are the penguins?"

[0153] This system allows zoos and aquariums to provide quick and appropriate responses based on the latest status of animals, improving visitor satisfaction and streamlining animal management.

[0154] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0155] Step 1: Collecting rearing data

[0156] The server receives the care data entered by the zookeepers using a dedicated application or web interface. When the zookeepers enter information such as the lions' food intake, exercise, and weight, it is sent to the server. The input is data entered manually by the zookeepers, and the server receives it and processes it into an appropriate format. The output is structured care data.

[0157] Specific behavior:

[0158] A zookeeper enters "Today's lion's food intake is 10kg," and the data is sent to the server, which structures the data and processes it into a format that can be stored in the database.

[0159] Step 2: Saving breeding data

[0160] The server stores the received breeding data in a database. The database used is MySQL or PostgreSQL. The input is structured data sent by the zookeepers, which the server inserts into the database. The output is the breeding data stored in the database.

[0161] Specific behavior:

[0162] The server inserts the data "Today's lion's food intake is 10 kg" into a MySQL database, which stores information such as the date, animal species, and food intake.

[0163] Step 3: Acquire footage of the animal

[0164] The device captures real-time video of the animals using a camera installed in the cage or aquarium. The input is video data from the camera, which the device temporarily stores in local storage. The output is the video data stored in local storage.

[0165] Specific behavior:

[0166] A camera inside the cage captures footage of the lion and sends it to the device, which stores the video data in its local storage.

[0167] Step 4: Animal footage analysis

[0168] The device analyzes the stored video data to determine the animal's behavior. OpenCV and deep learning models (YOLO and PoseNet) are used for the analysis. The input is the video data stored in local storage, which the device applies to the analysis algorithm. The output is the animal's behavior data.

[0169] Specific behavior:

[0170] The device analyzes the video data using OpenCV and determines whether the lion is "resting," "active," or "eating," and outputs the results as animal behavior data.

[0171] Step 5: Obtaining voice questions from the user

[0172] The user speaks a question about an animal into the device. The input is the user's voice question, which the device picks up with a microphone. The output is the captured voice data.

[0173] Specific behavior:

[0174] The user speaks into the device's microphone, saying, "What is this lion doing now?" The device captures this voice and processes it.

[0175] Step 6: Convert spoken questions to text

[0176] The device converts the acquired voice questions into text data using the Google Speech-to-Text API or IBM Watson Speech to Text. The input is the acquired voice data, and the output is the converted text data.

[0177] Specific behavior:

[0178] The device sends the voice data to the Google Speech-to-Text API and receives the text data, "What is this lion doing right now?"

[0179] Step 7: Generate a response

[0180] The server inputs the converted text question into a generative AI model to generate a response. The input is the user's text question and the animal's behavior data, and the output is the generated response text. Natural language processing technology is used to generate the response.

[0181] Specific behavior:

[0182] The server inputs the data "The lion is currently resting" into the generated AI model and generates the response "The lion is currently resting. Please be quiet."

[0183] Step 8: Text-to-speech response

[0184] The device converts the response sent from the server into speech using the Google Text-to-Speech API or Amazon Polly. The input is the response text data, and the output is the generated speech data.

[0185] Specific behavior:

[0186] The device converts the text response "The lion is currently resting. Please be quiet" into audio using the Google Text-to-Speech API.

[0187] Step 9: Providing a response

[0188] The terminal provides the generated voice data to the user through a speaker: the input is the voice data, and the output is a voice response that the user hears.

[0189] Specific behavior:

[0190] A voice will play from the device's speaker saying, "The lion is currently resting. Please be quiet."

[0191] In this way, the system can quickly provide appropriate responses based on the current status of animals at zoos and aquariums through each processing step.

[0192] (Application example 1)

[0193] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0194] In the past, visitors to zoos and aquariums needed information from display panels or explanations from keepers to know the current status of animals. However, this method made it difficult to obtain detailed information about the animals' behavior and health status in real time, and it was also difficult to provide immediate responses to questions. Furthermore, there was a lack of a system that could accurately grasp the animals' condition and provide appropriate explanations to visitors.

[0195] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0196] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the status of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined status data of the animals, means for providing the generated response to the user by voice, and means for the user to use at the zoo or aquarium via an application installed on their smartphone. This allows visitors to obtain detailed information about the current status and behavior of animals in real time and receive immediate responses to their voice questions, thereby improving the visitor experience.

[0197] "Breeding data" refers to information related to animal care management, such as the animal's health, amount of food eaten, amount of exercise, and amount of sleep.

[0198] "Animal footage" refers to real-time footage of animals on display at zoos and aquariums.

[0199] "Means for analyzing video footage to determine the condition of an animal" refers to technology for analyzing acquired video data of an animal and automatically determining the animal's behavior and health condition.

[0200] "Means for receiving questions from the user as voice input" refers to a mechanism for receiving and recording questions posed by the user via a device.

[0201] "Means for converting voice input to text" refers to technology that converts received voice data into text data.

[0202] The "means for generating a response" refers to a mechanism for generating an appropriate answer to a user's question using the converted text data and the animal's condition data.

[0203] "Means for providing a generated response to a user audibly" refers to technology that converts a generated text-based response into audio and provides it to a user.

[0204] "Applications installed on smartphones" refers to software that is downloaded and installed on the mobile device used by the user.

[0205] A "generative AI model" is a type of machine learning model that learns from large amounts of data and responds to and generates information such as text, images, and voice.

[0206] This invention is a system developed to improve the user experience at physical stores of zoos and aquariums. The system consists of a server, a terminal, and a user.

[0207] server

[0208] The server performs the following main functions:

[0209] 1. Collection and storage of rearing data:

[0210] The server periodically collects data on the animals' health and behavior from keepers and sensors and stores it in a database.

[0211] For example, this includes data on food intake and exercise volume entered by keepers, and data on sleep duration recorded by sensors.

[0212] 2. Generate a response:

[0213] A generative AI model is used to generate appropriate responses based on the user's question text and the animal's condition data.

[0214] For example, if it determines that the animal is currently resting, it generates the response "The lion is currently resting."

[0215] Examples of technologies used:

[0216] Database: MySQL

[0217] Generative AI models: large-scale language models such as GPT-3

[0218] Terminal

[0219] The terminal performs the following main functions:

[0220] 1. Animal image acquisition and analysis:

[0221] The system uses a camera to capture real-time video of the animals and analyzes the video data to determine their behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[0222] For example, cameras installed in lion cages record the lion's movements in real time and use the footage to determine whether the lion is currently active or resting.

[0223] 2. Voice input and conversion of user questions:

[0224] The system uses a microphone to capture the user's voice question and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[0225] For example, a child might ask, "What is this lion doing right now?" and the device would convert the voice into text data.

[0226] 3. Provide audio response:

[0227] The response sent from the server is synthesized and provided to the user using the Google Text-to-Speech API or Amazon Polly to convert the generated text into speech and play it back through the speaker.

[0228] For example, the device receives a response generated by the server saying "The lion is currently resting," converts it into voice, and conveys it to the user.

[0229] User

[0230] A user (e.g., a zoo visitor) accesses the system in the following ways:

[0231] 1. Speak your question:

[0232] The user speaks a question about an animal, such as "What is this lion doing right now?"

[0233] 2. Receiving the response:

[0234] The device responds with audio, giving users real-time information about the animal's condition.

[0235] Specific examples of processing

[0236] 1. The server stores the lion's health data in a database.

[0237] 2. The device uses its camera to capture footage of the lion and analyzes the footage to determine whether the lion is resting.

[0238] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[0239] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[0240] 5. The server passes the question to the generative AI model, which generates the response "The lion is currently resting."

[0241] 6. The device converts the response into audio and transmits it to the child through the speaker.

[0242] Prompt Sentence Examples

[0243] "What is the current state of this animal?"

[0244] What activities are lions currently engaged in?

[0245] "How is this penguin's health?"

[0246] In this way, the system of the present invention can quickly provide appropriate responses based on the real-time status of animals to enhance user experience at zoos and aquariums.

[0247] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0248] Step 1:

[0249] The server collects and stores the husbandry data. It receives data provided by the zookeepers and sensors and stores it in a MySQL database. For example, the lions' diet, exercise, and sleep duration are entered and stored in the database. Inputs include data entered manually by the zookeepers and data automatically sent from the sensors. The output is a stored database entry.

[0250] Step 2:

[0251] The device uses a camera to capture real-time video of the animals. The camera is installed in the animal's exhibition area and captures video in real time. The input is the video data captured by the camera, and the output is video frames for analysis.

[0252] Step 3:

[0253] The device analyzes the captured video to determine the animal's state. It analyzes the video data using technologies such as OpenCV and the YOLO model to identify the animal's behavior. For example, it determines whether a lion is resting or active. The input is the video frame obtained in step 2, and the output is the animal's state determination result (e.g., resting, active).

[0254] Step 4:

[0255] The user speaks a question about an animal into the terminal, for example, "What is this lion doing now?" The input is the voice data from the user, and the output is a recorded voice file.

[0256] Step 5:

[0257] The device converts the user's voice question into text using speech recognition technology, such as the Google Speech-to-Text API. The input is the audio file obtained in step 4, and the output is the converted text.

[0258] Step 6:

[0259] The server generates a response using the converted text data and the animal's state data. It uses a generative AI model (e.g., GPT-3) to create an appropriate response to the user's question. The input is the animal's state data obtained in step 3 and the text data obtained in step 5, and the output is the generated response text.

[0260] Step 7:

[0261] The device converts the response sent from the server into speech and provides it to the user. It uses the Google Text-to-Speech API or Amazon Polly to convert text to speech. The input is the response text generated in step 6, and the output is the synthesized speech data.

[0262] Step 8:

[0263] The terminal plays the generated audio data to the user through the speaker, so that the user can hear information about the animal's condition in real time. The input is the audio data obtained in step 7, and the output is the audio response provided to the user.

[0264] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0265] The present invention relates to a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate responses to questions from users. This system is composed of a server, terminals, users, and a sentiment analysis engine.

[0266] Overall system overview

[0267] The system consists of the following main parts:

[0268] 1. Server:

[0269] Collect and store breeding data.

[0270] A generative AI model is used to generate appropriate responses to user questions.

[0271] 2. Terminal:

[0272] It uses a camera and microphone to capture footage of animals and user voice questions.

[0273] The acquired data is analyzed and sent to the server.

[0274] The response received from the server is synthesized into voice and provided to the user.

[0275] 3. User:

[0276] Ask questions about animals by speaking them into the device.

[0277] 4. Sentiment Analysis Engine:

[0278] Recognizes emotions by analyzing the user's voice and text data.

[0279] Program processing

[0280] Collection and storage of breeding data

[0281] server

[0282] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[0283] Examples:

[0284] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[0285] Animal imaging and analysis

[0286] Terminal

[0287] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[0288] Examples:

[0289] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[0290] Acquiring and analyzing voice questions from users

[0291] User

[0292] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[0293] Terminal

[0294] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[0295] Sentiment Analysis Engine

[0296] The emotion analysis engine analyzes the user's emotions based on the acquired voice and text data, and determines whether the user is excited, calm, happy, etc.

[0297] Examples:

[0298] If a user excitedly asks, "What is this lion doing right now?", the sentiment analysis engine will determine from the tone of voice and the use of words that the user is excited.

[0299] Generating and serving the response

[0300] server

[0301] The server uses a generative AI model to generate a response based on the user's question text and the animal's state data. The response is also tailored to the user's emotions. For example, if the user is excited, the server might generate a response like, "The lion is currently resting. Please watch quietly!"

[0302] Terminal

[0303] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[0304] Examples:

[0305] The device receives the response generated by the server, "The lion is currently resting. Please watch over it quietly!", converts it into voice and conveys it to the user. In this way, the user can receive accurate information based on the animal's latest condition and feedback according to its individual emotions.

[0306] Processing flow through concrete examples

[0307] 1. The server stores the lion's health data in a database.

[0308] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[0309] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[0310] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[0311] 5. The sentiment analysis engine analyzes the user's tone of voice and text to determine if they are excited.

[0312] 6. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please watch over it quietly!"

[0313] 7. The device converts the response into audio and transmits it to the child through the speaker.

[0314] In this way, the system of the present invention improves the user experience at zoos and aquariums, allowing for quick and appropriate responses based on the animals' current status and the user's emotions.

[0315] The processing flow will be explained below.

[0316] Step 1:

[0317] server

[0318] Collects breeding data. The server receives data entered by keepers and data sent from sensors, and stores the animal's health status (e.g., amount of food eaten, amount of exercise, hours of sleep, etc.) in a database such as MySQL.

[0319] Specific operation: Keepers enter data such as food intake and exercise data via a web form and store it in a database using a PHP script. Data automatically sent by sensors is also received via a REST API and stored in the database in the same way.

[0320] Step 2:

[0321] Terminal

[0322] Obtaining images of animals. The terminal obtains real-time images of animals using IP cameras installed in cages, etc.

[0323] Specific operation: The terminal receives the video stream from the IP camera using the RTSP protocol and processes the video data in real time using the OpenCV library.

[0324] Step 3:

[0325] Terminal

[0326] Analyzing the video to determine the animal's state: The device analyzes the acquired video data to determine the animal's behavior and state (e.g., active, resting, sleeping).

[0327] How it works: The device loads a deep learning model (such as YOLO or PoseNet) and analyzes the animal's pose and movement for each video frame. For example, it uses the YOLO model to detect the animal's outline and PoseNet to estimate its pose.

[0328] Step 4:

[0329] User

[0330] Users speak questions into the device by voice. Users ask specific questions into the microphone about the animal's condition and activity.

[0331] Specific action: The user speaks into the microphone, saying something like, "What is this penguin doing right now?"

[0332] Step 5:

[0333] Terminal

[0334] Converts voice to text. The device converts the acquired voice data into text data using the Google Speech-to-Text API.

[0335] What it does: The device records audio data collected by the microphone in WAV format and sends the data to the Google Speech-to-Text API to convert it into text.

[0336] Step 6:

[0337] Terminal

[0338] Send the text and animal status data to the server. The converted text data and the determined animal status information are sent to the server using an HTTP POST request.

[0339] Specific operation: The device compiles the user's question text and the analyzed animal's condition data into JSON format and sends it to the server as an HTTP POST request.

[0340] Step 7:

[0341] server

[0342] Recognizing user emotions: The server uses a sentiment analysis engine to determine the user's emotions (e.g., excitement, surprise, joy, anger, etc.) from the received text data.

[0343] Specific operation: The server passes the received question text to a sentiment analysis library (e.g., Azure Emotion API or NLTK library) to obtain the user's sentiment.

[0344] Step 8:

[0345] server

[0346] Generate a response: The server uses a generative AI model to generate an appropriate response based on the user's question text, the animal's condition data, and the results of sentiment analysis.

[0347] Specific operation: The server inputs the question text, animal state data, and user emotion data into a generative AI model (e.g., GPT-3), and generates a response such as "The lion is currently resting. Please watch over it quietly."

[0348] Step 9:

[0349] server

[0350] Send the generated response to the terminal. Send the generated text response back to the terminal as an HTTP response.

[0351] Specific operation: The server compiles the generated response text into JSON format and sends it to the terminal using an HTTP response.

[0352] Step 10:

[0353] Terminal

[0354] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[0355] Specific operation: The device passes the received response text to the Google Text-to-Speech API to generate audio data and saves it as an audio file.

[0356] Step 11:

[0357] Terminal

[0358] Providing a voice response to the user: The terminal plays the generated voice data through a speaker to provide the user with an appropriate answer.

[0359] Specific operation: The device uses the playback function (for example, the playback function of the Pygame library) to play the generated audio file from the speaker.

[0360] For example, if a user asks the device, "What is this lion doing now?", the system analyzes the lion's state in real time, the server generates a response such as, "The lion is currently resting. Please watch over it quietly!", and the device plays back the response aloud. In this way, the user can obtain highly accurate information based on the animal's latest state and their own emotions in real time.

[0361] Example 2

[0362] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0363] Conventional animal monitoring and user information provision systems in zoos and aquariums are required not only to grasp the status of animals in real time, but also to respond quickly and appropriately to user questions. However, it is technically difficult to not only monitor animal behavior but also to generate responses based on the user's emotions. Furthermore, a comprehensive system is required, as advanced processing is required to respond appropriately to the diverse questions and emotions of users.

[0364] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0365] In this invention, the server includes means for collecting and saving breeding information, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the behavior of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for analyzing the user's emotions from the voice input, means for generating a response using the converted text, the determined animal status data, and the analyzed emotion data, and means for providing the generated response to the user by voice, thereby enabling the server to quickly provide an appropriate response based on both the animal's behavior and the user's emotions.

[0366] "Breeding information" refers to data about the animal's health, behavioral patterns, diet, amount of exercise, etc.

[0367] "Footage" refers to real-time video data of animals captured using a camera.

[0368] "Behavior" refers to the type of movement or activity an animal is engaged in at a particular time.

[0369] "Voice input" refers to voice data that a user utters to the system through a voice acquisition device such as a microphone.

[0370] "Text" refers to data that has been converted from voice data into text using voice recognition technology.

[0371] "Emotion" refers to the user's psychological state (e.g., excitement, calmness, joy) based on an analysis of voice tone and text content.

[0372] "Response" refers to the answer data generated by the system in response to a user's question.

[0373] A "generative AI model" refers to an algorithm that uses machine learning and artificial intelligence techniques to generate appropriate responses based on user questions and other data.

[0374] This invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for responding to user questions promptly and appropriately. This system is mainly composed of a server, terminals, users, and a sentiment analysis engine.

[0375] server

[0376] The server is primarily responsible for collecting and storing animal care information, including the animals' health, behavior, food intake, and exercise. This data is collected periodically from keepers and sensors and stored in a database such as MySQL. The server also uses generative AI models to generate appropriate responses to user questions.

[0377] Technology used

[0378] Database Management System (DBMS): MySQL

[0379] Generative AI models: Machine learning and artificial intelligence techniques

[0380] Terminal

[0381] The device captures real-time video footage using a camera installed in the animal cage and analyzes it to determine the animal's behavior. Deep learning models such as OpenCV, YOLO, and PoseNet are used for the analysis. It also receives voice questions from users and converts them into text using speech recognition technology. The acquired data is sent to a server, and the responses received from the server are provided to the user via voice synthesis.

[0382] Technology used

[0383] Video analysis: OpenCV, YOLO, PoseNet

[0384] Speech Recognition: Google Speech-to-Text API

[0385] Speech synthesis: Google Text-to-Speech API, Amazon Polly

[0386] User

[0387] Users can ask the system questions about the animal's condition and behavior by voice. For example, a child might ask the device, "What is this lion doing right now?" This voice data is captured by the device and converted into text using voice recognition technology.

[0388] Sentiment Analysis Engine

[0389] The emotion analysis engine analyzes the user's voice and text data to recognize their emotions, allowing it to determine whether they are excited, calm, or happy.

[0390] Technology used

[0391] Sentiment analysis: Natural language processing techniques, pitch analysis, text analysis

[0392] Specific examples

[0393] For example, imagine the following scenario at a zoo one day. A database connected to a server stores daily updated data on the lion's diet, exercise volume, and sleep duration. A device captures real-time footage using a camera installed in the lion's cage and uses OpenCV and the YOLO model to determine that the lion is currently resting. A child asks into the device's microphone, "What is this lion doing right now?" The device converts the voice into text using the Google Speech-to-Text API. An emotion analysis engine determines from the tone of the voice that the child is excited. The server uses a generative AI model to generate a response based on the prompt: "The lion is currently resting. Please watch over him quietly!" The device converts this response into speech using the Google Text-to-Speech API and relays it to the child through the speaker.

[0394] Prompt Sentence Examples

[0395] The user is asking about the current status of the lion. The lion is currently resting and the user is excited.

[0396] This allows the system to quickly provide appropriate responses based on the animal's condition and the user's emotions in zoos and aquariums.

[0397] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0398] Step 1: Collect and store breeding information

[0399] server

[0400] The server collects animal care data. Specifically, keepers enter data such as food intake and exercise volume into an input form, and this data is sent to the server as an HTTP request. In addition, sensors (e.g., vital sign monitors, activity meters) send animal health status data to the server in real time. The server receives this data and inserts it into the corresponding tables in the MySQL database.

[0401] Input: Keeper input data on food intake and exercise, and health status data from sensors

[0402] Data processing: Database insert processing, data consistency check

[0403] Output: Breeding information stored in the database

[0404] Step 2: Video acquisition and behavioral analysis

[0405] Terminal

[0406] The device captures real-time video using a camera installed in the animal cage. The video data is processed frame by frame and the animal's behavior is analyzed using deep learning models such as OpenCV, YOLO, and PoseNet. The behavioral results are sent to the server in JSON format.

[0407] Input: Real-time video captured from a camera

[0408] Data processing: video frame analysis, animal behavior recognition

[0409] Output: Action result (e.g., lion is resting)

[0410] Step 3: Capture voice questions and convert them to text

[0411] User

[0412] Users can speak into the device's microphone to ask questions about animals, such as, "What is this lion doing right now?"

[0413] Terminal

[0414] The device captures the user's voice with a microphone, stores the audio data in a buffer, and converts the captured audio data into text in real time using the Google Speech-to-Text API. The converted text data is then sent to the next step.

[0415] Input: User voice question

[0416] Data processing: converting voice data to text

[0417] Output: Question text (e.g., "What is this lion doing right now?")

[0418] Step 4: Sentiment Analysis

[0419] Sentiment Analysis Engine

[0420] The sentiment analysis engine analyzes the user's emotions based on voice and text data, determining the user's emotional state from the voice tone, speed, and text content, and sending the results to the response generation step.

[0421] Input: Text data of voice questions, voice tone data

[0422] Data processing: Natural language processing and speech tone analysis

[0423] Output: Emotion determination result (e.g., user is excited)

[0424] Step 5: Generate a response

[0425] server

[0426] The server uses a generative AI model to generate a response based on the user's question text, the animal's behavioral assessment results, and the analyzed emotion data. The server uses a prompt sentence to generate an appropriate response text.

[0427] Input: Question text, behavioral judgment result, emotion judgment result, prompt text

[0428] Data processing: Response generation using generative AI models

[0429] Output: Response text (e.g. "The lion is currently resting. Please keep quiet and watch over it!")

[0430] Step 6: Provide a voice response

[0431] Terminal

[0432] The device stores the response text received from the server in a buffer and converts it into audio data using the Text-to-Speech API, which is then provided to the user through the speaker.

[0433] Input: Response text

[0434] Data processing: Converting text data into speech

[0435] Output: Audio response (e.g. "The lion is currently resting. Please keep quiet and watch over it!" played from the speaker)

[0436] In this way, it is possible to quickly provide an appropriate response based on the animal's condition and the user's emotions.

[0437] (Application example 2)

[0438] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0439] Modern factories require systems that can monitor robots in real time to ensure they are operating efficiently and provide appropriate responses to staff questions. However, current systems struggle to not only monitor the robot's status in real time, but also to respond quickly and accurately to staff questions. To solve this problem, a system is needed that can analyze the robot's behavior and automatically generate responses that take staff emotions into account.

[0440] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0441] In this invention, the server includes a means for collecting and storing breeding data and behavior data, a means for acquiring images of animals and equipment in real time, and a means for analyzing the acquired images to determine the status of the animals and equipment. This makes it possible to monitor the behavior of robots in the factory in real time, and furthermore, to generate and provide quick and accurate responses to staff questions using a generative AI model.

[0442] "Animal and operation data" refers to information relating to the health, behavior, and operating status of animals and equipment.

[0443] "Means for acquiring images of animals or devices in real time" refers to devices that use cameras or sensors to observe the activities of animals or devices in real time and acquire the image data.

[0444] The "means for analyzing the acquired video and determining the state of the animal or device" is a system that analyzes the acquired video data and determines the current state of the animal or device based on the results.

[0445] The "means for receiving a question from a user as a voice input" is a device that acquires a voice question uttered by a user via a microphone or the like.

[0446] The "means for converting speech input to text" is a speech recognition system that converts acquired speech data into text data.

[0447] The "means for generating a response using the converted text and the determined state data of the animal or device" refers to a generative AI model that generates an appropriate response using the text converted from the voice data and the determined state data of the animal or device.

[0448] The "means for providing a generated response to a user by voice" is a system that synthesizes a generated text-based response into speech and provides it to a user by voice.

[0449] A "system for analyzing the behavior of animals and equipment" is a system that uses video analysis technology to understand the behavior of animals and equipment and determine their condition based on that information.

[0450] A "generative AI model" is an artificial intelligence model that uses machine learning technology to automatically generate responses from text.

[0451] A "prompt sentence" is an input sentence given to a generative AI model, and is a document based on which the generative AI model generates a response.

[0452] MODE FOR CARRYING OUT THE INVENTION

[0453] Overall system overview

[0454] This invention relates to a system that monitors the operation of robots and equipment in a factory in real time and responds quickly and appropriately to questions from staff. The system consists of a server, terminals, users, and a sentiment analysis engine. The system consists of the following main parts:

[0455] server

[0456] The server has the following roles:

[0457] Data collection and storage:

[0458] The server periodically collects operational and other status data from the robots and equipment in the factory and stores it in a database, including health data, operational data, etc. For example, it uses SQLAlchemy to connect to a MySQL database and stores this data in tables.

[0459] Terminal

[0460] The terminals are used by staff in the factory and have the following roles:

[0461] Video acquisition and analysis:

[0462] The device uses a camera to capture real-time video of the robot or equipment. The captured video is analyzed and the current status of the robot or equipment is determined based on the results. The analysis is performed using OpenCV and deep learning models (such as YOLO and PoseNet).

[0463] Voice Input and Recognition:

[0464] The user speaks a question into the device, and the voice is picked up through the microphone and converted into text using voice recognition technology, such as the Google Speech-to-Text API.

[0465] User

[0466] Users, especially factory staff, use the terminals as follows:

[0467] Submit a question:

[0468] The user can ask a voice question to the device, such as, "What is this robot doing right now?"

[0469] Sentiment Analysis Engine

[0470] The sentiment analysis engine has the following functions:

[0471] Emotion analysis:

[0472] The system analyzes the captured voice and text data to recognize the user's emotions, determining whether the user is excited or calm. The analysis is performed using the Google Cloud Natural Language API.

[0473] Generating and serving the response

[0474] The server uses a generative AI model to generate a response based on the user's question text and the status data of the robot and equipment. The generated response is adjusted according to the user's emotions. For example, if the user is excited, the response generated will be something like, "The robot is currently packaging. Please be careful."

[0475] Specific examples

[0476] Data collection:

[0477] The robot is monitored in real time through a camera installed on the staff's terminal, which captures images of the robot and analyzes them using OpenCV.

[0478] Voice questions:

[0479] A factory worker asks into the device's microphone, "What is this robot doing now?" The device converts the voice to text and uses the Google Speech-to-Text API to translate it.

[0480] Response generation:

[0481] Using the converted text and the analyzed robot state data, the server sends prompt sentences to the generative AI model to generate a response.

[0482] Response provided:

[0483] The generated response is converted into audio using the Google Text-to-Speech API and provided to staff via the terminal.

[0484] Prompt Sentence Examples

[0485] Question: What is this robot doing now?

[0486] Robot status: Packaging in progress

[0487] User sentiment score: 0.8

[0488] response:

[0489] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0490] Step 1:

[0491] The server collects operational and other status data from the robots and equipment in the factory and stores it in a database. The input is data from the sensors on the robots and equipment, and the output is organized and stored database entries. Specifically, it uses SQLAlchemy to connect to a MySQL database and periodically updates the operational data.

[0492] Step 2:

[0493] The device uses a camera to capture real-time video of the robot or equipment. The input is the video data captured by the camera, and the output is the current status of the robot or equipment as an analysis result. The device analyzes the captured video using OpenCV or a deep learning model (such as YOLO or PoseNet) to determine the work the robot or equipment is currently performing.

[0494] Step 3:

[0495] The user speaks a question into the terminal. The input is the user's voice question, and the output is voice data. For example, the user speaks a question such as "What is this robot doing now?" into the microphone of the terminal.

[0496] Step 4:

[0497] The device receives a voice question from the user and converts it into text. The input is the captured voice data, and the output is the converted text data. The device converts the voice data into text using the Google Speech-to-Text API.

[0498] Step 5:

[0499] The server generates a response using the converted text and the determined robot or equipment status data. The input is the user's question text and the robot or equipment status data, and the output is the generated response text. The server provides the prompt text to the generative AI model, which then generates an appropriate response. Examples of prompt text are as follows:

[0500] Question: What is this robot doing now?

[0501] Robot status: Packaging in progress

[0502] User sentiment score: 0.8

[0503] response:

[0504] Step 6:

[0505] The sentiment analysis engine analyzes the user's voice and text data to recognize emotions. The input is the user's voice data and converted text data, and the output is an emotion score. It uses the Google Cloud Natural Language API to analyze the user's emotions and provide the emotion score to the server.

[0506] Step 7:

[0507] The device synthesizes the response sent from the server and provides it to the user. The input is the response text received from the server, and the output is the response in audio format. The device uses the Google Text-to-Speech API to convert the text into audio and transmit it to the user through the speaker.

[0508] This allows for real-time monitoring of robots within the factory and appropriate responses to staff questions.

[0509] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0510] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0511] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0512] [Second embodiment]

[0513] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0514] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0515] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0516] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0517] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0518] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0519] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0520] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0521] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0522] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0523] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0524] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0525] The present invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for providing prompt and appropriate responses to questions from users. This system is composed of a server, terminals, and users.

[0526] Overall system overview

[0527] The system consists of the following main parts:

[0528] 1. Server:

[0529] Collect and store breeding data.

[0530] A generative AI model is used to generate appropriate responses to user questions.

[0531] 2. Terminal:

[0532] It uses a camera and microphone to capture footage of animals and user voice questions.

[0533] The acquired data is analyzed and sent to the server.

[0534] The response received from the server is synthesized into voice and provided to the user.

[0535] 3. User:

[0536] Ask questions about animals by speaking them into the device.

[0537] Program processing

[0538] Collection and storage of breeding data

[0539] server

[0540] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[0541] Examples:

[0542] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[0543] Animal imaging and analysis

[0544] Terminal

[0545] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[0546] Examples:

[0547] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[0548] Analysis of voice questions from users

[0549] User

[0550] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[0551] Terminal

[0552] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[0553] Examples:

[0554] A child asks the device, "What is this lion doing right now?" and the device converts the voice into text data.

[0555] Generating and serving the response

[0556] server

[0557] The server uses a generative AI model to generate a response based on the user's question text and the animal's status data. For example, if it determines that the lion is currently resting, it generates the response "The lion is currently resting. Please keep quiet."

[0558] Terminal

[0559] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[0560] Examples:

[0561] The device receives the response generated by the server, "The lion is currently resting. Please be quiet," converts it into voice and conveys it to the user.

[0562] Processing flow through concrete examples

[0563] 1. The server stores the lion's health data in a database.

[0564] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[0565] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[0566] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[0567] 5. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please keep quiet."

[0568] 6. The device converts the response into audio and transmits it to the child through the speaker.

[0569] In this way, the system of the present invention improves the user experience at zoos and aquariums, providing quick and appropriate responses based on the animals' current status.

[0570] The processing flow will be explained below.

[0571] Step 1:

[0572] server

[0573] Collects animal care data. The server receives input from keepers and data from sensors, and stores information about the animals' health status (e.g., amount of food eaten, amount of exercise, amount of sleep, etc.) in a database. For example, keepers enter the amount of food eaten each day using a web form, and the amount of exercise automatically recorded by sensors is stored in a database such as MySQL.

[0574] Step 2:

[0575] Terminal

[0576] Acquire video of the animal. The device acquires video in real time using a camera installed in the animal's cage, etc. The camera can be an IP camera or a USB-connected webcam. The device processes the video stream using libraries such as OpenCV and formats the data into a format that can be analyzed in real time.

[0577] Step 3:

[0578] Terminal

[0579] The device analyzes the video to determine the animal's state. The device analyzes the animal's movements and state based on the acquired video data. For example, it uses deep learning models (such as YOLO or PoseNet) to detect the animal's posture and movement and determine its state, such as "active," "resting," or "sleeping."

[0580] Step 4:

[0581] User

[0582] Users (especially children) can speak questions into the device by talking to the microphone and asking questions about the animal's condition or activity, such as "What is this penguin doing right now?"

[0583] Step 5:

[0584] Terminal

[0585] Converts voice to text. The device converts the user's voice questions captured through the microphone into text data using voice recognition technology such as the Google Speech-to-Text API.

[0586] Step 6:

[0587] Terminal

[0588] Send the text and animal status data to the server. The converted text data and the real-time animal status information are compiled and sent to the server using an HTTP POST request.

[0589] Step 7:

[0590] server

[0591] Generate a response: The server uses a generative AI model to generate an appropriate response based on the received text data and animal status data. For example, it generates a response such as "This penguin is currently resting."

[0592] Step 8:

[0593] server

[0594] The generated response is sent to the terminal. The generated response in text data format is returned to the terminal as an HTTP response.

[0595] Step 9:

[0596] Terminal

[0597] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[0598] Step 10:

[0599] Terminal

[0600] Providing a voice response to the user: The terminal plays the voice-converted response to the user through a speaker, providing an answer to the question.

[0601] For example, if a user asks the device, "What is this lion doing right now?", the system will analyze the lion's status in real time, the server will generate a response such as "The lion is currently resting," and the device will play back the response aloud. In this way, the user can obtain accurate information based on the animal's latest status.

[0602] Example 1

[0603] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0604] In conventional zoos and aquariums, it has been difficult to grasp the condition of animals in real time and respond quickly and accurately to visitor questions. In addition, it has not been possible to integrate and utilize breeding data and real-time animal behavior data, which has limited the ability to improve visitor satisfaction and streamline breeding management.

[0605] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0606] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the condition of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined animal condition data, and means for providing the generated response to the user by voice, thereby enabling zoos and aquariums to provide quick and appropriate responses based on the latest condition of animals.

[0607] "Breeding data" refers to data that includes information on the health, food intake, and exercise of animals kept in zoos and aquariums.

[0608] "Real-time" refers to the instantaneous acquisition and processing of information and situations at the current time.

[0609] "Animal condition" is a concept that includes information such as the animal's health, behavior, and activity status.

[0610] "Audio input" refers to speech produced by a user through a microphone or other audio capture device.

[0611] "Means for converting voice input to text" refers to technologies or systems that convert voice data into character data.

[0612] "Means for generating a response" refers to the technology or mechanism for generating an appropriate response to a user's question.

[0613] A "generative AI model" refers to an artificial intelligence model that uses natural language processing technology to understand and generate human language.

[0614] "Analytical algorithm" refers to a computational procedure or method for analyzing data and deriving a particular result or information.

[0615] System Overview

[0616] The present invention is a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate answers to questions from users. The system is composed of a server, terminals, and users, and has the following functions:

[0617] server

[0618] The server has the following functions:

[0619] 1. Collection and storage of rearing data:

[0620] The server receives breeding data entered by keepers using a dedicated application or web interface. For example, it periodically collects information such as the lions' diet, exercise, and weight. It also collects animal health data (heart rate, body temperature, etc.) obtained from sensors. The collected data is stored in a relational database such as MySQL or PostgreSQL.

[0621] Terminal

[0622] The device has the following features:

[0623] 1. Animal footage acquisition:

[0624] The device captures real-time images of animals using cameras installed in cages or aquariums, and the camera video data is temporarily stored in the device's local storage.

[0625] 2. Animal video analysis:

[0626] The device analyzes the stored video data to determine the animal's behavior using OpenCV and deep learning models (such as YOLO and PoseNet), allowing it to determine in real time whether the animal is currently active, resting, or performing a specific behavior.

[0627] 3. Acquire voice questions from users and convert them into text:

[0628] The user speaks to the device to ask a question about an animal. For example, "What is this lion doing right now?" The device then converts the spoken question picked up by the microphone into text. This process uses the Google Speech-to-Text API or IBM Watson Speech to Text.

[0629] 4. Speech synthesis and delivery of responses:

[0630] The response sent from the server is converted into speech using the Google Text-to-Speech API or Amazon Polly, and the generated speech data is then provided to the user through the speaker.

[0631] Specific examples of programs

[0632] Example of user interaction:

[0633] For example, if a user asks the terminal, "What is this lion doing right now?", the process is as follows:

[0634] 1. Acquiring a voice question: The user speaks into the device's microphone, "What is this lion doing right now?"

[0635] 2. Converting voice questions into text: The device sends the captured voice to the Google Speech-to-Text API and receives the text "What is this lion doing right now?"

[0636] 3. Determining the animal's state: The device analyzes the video data using OpenCV and determines that the lion is "resting."

[0637] 4. Generating and providing a response: The server inputs the data "The lion is currently resting" into the generative AI model, generates a response "The lion is currently resting. Please be quiet," and the device converts the response into audio and plays it through the speaker.

[0638] Prompt Sentence Examples

[0639] An example of a prompt sentence might be:

[0640] 1. "What is this lion doing now?"

[0641] 2. "When is the bear's meal time?"

[0642] 3. "Where are the penguins?"

[0643] This system allows zoos and aquariums to provide quick and appropriate responses based on the latest status of animals, improving visitor satisfaction and streamlining animal management.

[0644] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0645] Step 1: Collecting rearing data

[0646] The server receives the care data entered by the zookeepers using a dedicated application or web interface. When the zookeepers enter information such as the lions' food intake, exercise, and weight, it is sent to the server. The input is data entered manually by the zookeepers, and the server receives it and processes it into an appropriate format. The output is structured care data.

[0647] Specific behavior:

[0648] A zookeeper enters "Today's lion's food intake is 10kg," and the data is sent to the server, which structures the data and processes it into a format that can be stored in the database.

[0649] Step 2: Saving breeding data

[0650] The server stores the received breeding data in a database. The database used is MySQL or PostgreSQL. The input is structured data sent by the zookeepers, which the server inserts into the database. The output is the breeding data stored in the database.

[0651] Specific behavior:

[0652] The server inserts the data "Today's lion's food intake is 10 kg" into a MySQL database, which stores information such as the date, animal species, and food intake.

[0653] Step 3: Acquire footage of the animal

[0654] The device captures real-time video of the animals using a camera installed in the cage or aquarium. The input is video data from the camera, which the device temporarily stores in local storage. The output is the video data stored in local storage.

[0655] Specific behavior:

[0656] A camera inside the cage captures footage of the lion and sends it to the device, which stores the video data in its local storage.

[0657] Step 4: Animal footage analysis

[0658] The device analyzes the stored video data to determine the animal's behavior. OpenCV and deep learning models (YOLO and PoseNet) are used for the analysis. The input is the video data stored in local storage, which the device applies to the analysis algorithm. The output is the animal's behavior data.

[0659] Specific behavior:

[0660] The device analyzes the video data using OpenCV and determines whether the lion is "resting," "active," or "eating," and outputs the results as animal behavior data.

[0661] Step 5: Obtaining voice questions from the user

[0662] The user speaks a question about an animal into the device. The input is the user's voice question, which the device picks up with a microphone. The output is the captured voice data.

[0663] Specific behavior:

[0664] The user speaks into the device's microphone, saying, "What is this lion doing now?" The device captures this voice and processes it.

[0665] Step 6: Convert spoken questions to text

[0666] The device converts the acquired voice questions into text data using the Google Speech-to-Text API or IBM Watson Speech to Text. The input is the acquired voice data, and the output is the converted text data.

[0667] Specific behavior:

[0668] The device sends the voice data to the Google Speech-to-Text API and receives the text data, "What is this lion doing right now?"

[0669] Step 7: Generate a response

[0670] The server inputs the converted text question into a generative AI model to generate a response. The input is the user's text question and the animal's behavior data, and the output is the generated response text. Natural language processing technology is used to generate the response.

[0671] Specific behavior:

[0672] The server inputs the data "The lion is currently resting" into the generated AI model and generates the response "The lion is currently resting. Please be quiet."

[0673] Step 8: Text-to-speech response

[0674] The device converts the response sent from the server into speech using the Google Text-to-Speech API or Amazon Polly. The input is the response text data, and the output is the generated speech data.

[0675] Specific behavior:

[0676] The device converts the text response "The lion is currently resting. Please be quiet" into audio using the Google Text-to-Speech API.

[0677] Step 9: Providing a response

[0678] The terminal provides the generated voice data to the user through a speaker: the input is the voice data, and the output is a voice response that the user hears.

[0679] Specific behavior:

[0680] A voice will play from the device's speaker saying, "The lion is currently resting. Please be quiet."

[0681] In this way, the system can quickly provide appropriate responses based on the current status of animals at zoos and aquariums through each processing step.

[0682] (Application example 1)

[0683] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0684] In the past, visitors to zoos and aquariums needed information from display panels or explanations from keepers to know the current status of animals. However, this method made it difficult to obtain detailed information about the animals' behavior and health status in real time, and it was also difficult to provide immediate responses to questions. Furthermore, there was a lack of a system that could accurately grasp the animals' condition and provide appropriate explanations to visitors.

[0685] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0686] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the status of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined status data of the animals, means for providing the generated response to the user by voice, and means for the user to use at the zoo or aquarium via an application installed on their smartphone. This allows visitors to obtain detailed information about the current status and behavior of animals in real time and receive immediate responses to their voice questions, thereby improving the visitor experience.

[0687] "Breeding data" refers to information related to animal care management, such as the animal's health, amount of food eaten, amount of exercise, and amount of sleep.

[0688] "Animal footage" refers to real-time footage of animals on display at zoos and aquariums.

[0689] "Means for analyzing video footage to determine the condition of an animal" refers to technology for analyzing acquired video data of an animal and automatically determining the animal's behavior and health condition.

[0690] "Means for receiving questions from the user as voice input" refers to a mechanism for receiving and recording questions posed by the user via a device.

[0691] "Means for converting voice input to text" refers to technology that converts received voice data into text data.

[0692] The "means for generating a response" refers to a mechanism for generating an appropriate answer to a user's question using the converted text data and the animal's condition data.

[0693] "Means for providing a generated response to a user audibly" refers to technology that converts a generated text-based response into audio and provides it to a user.

[0694] "Applications installed on smartphones" refers to software that is downloaded and installed on the mobile device used by the user.

[0695] A "generative AI model" is a type of machine learning model that learns from large amounts of data and responds to and generates information such as text, images, and voice.

[0696] This invention is a system developed to improve the user experience at physical stores of zoos and aquariums. The system consists of a server, a terminal, and a user.

[0697] server

[0698] The server performs the following main functions:

[0699] 1. Collection and storage of rearing data:

[0700] The server periodically collects data on the animals' health and behavior from keepers and sensors and stores it in a database.

[0701] For example, this includes data on food intake and exercise volume entered by keepers, and data on sleep duration recorded by sensors.

[0702] 2. Generate a response:

[0703] A generative AI model is used to generate appropriate responses based on the user's question text and the animal's condition data.

[0704] For example, if it determines that the animal is currently resting, it generates the response "The lion is currently resting."

[0705] Examples of technologies used:

[0706] Database: MySQL

[0707] Generative AI models: large-scale language models such as GPT-3

[0708] Terminal

[0709] The terminal performs the following main functions:

[0710] 1. Animal image acquisition and analysis:

[0711] The system uses a camera to capture real-time video of the animals and analyzes the video data to determine their behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[0712] For example, cameras installed in lion cages record the lion's movements in real time and use the footage to determine whether the lion is currently active or resting.

[0713] 2. Voice input and conversion of user questions:

[0714] The system uses a microphone to capture the user's voice question and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[0715] For example, a child might ask, "What is this lion doing right now?" and the device would convert the voice into text data.

[0716] 3. Provide audio response:

[0717] The response sent from the server is synthesized and provided to the user using the Google Text-to-Speech API or Amazon Polly to convert the generated text into speech and play it back through the speaker.

[0718] For example, the device receives a response generated by the server saying "The lion is currently resting," converts it into voice, and conveys it to the user.

[0719] User

[0720] A user (e.g., a zoo visitor) accesses the system in the following ways:

[0721] 1. Speak your question:

[0722] The user speaks a question about an animal, such as "What is this lion doing right now?"

[0723] 2. Receiving the response:

[0724] The device responds with audio, giving users real-time information about the animal's condition.

[0725] Specific examples of processing

[0726] 1. The server stores the lion's health data in a database.

[0727] 2. The device uses its camera to capture footage of the lion and analyzes the footage to determine whether the lion is resting.

[0728] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[0729] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[0730] 5. The server passes the question to the generative AI model, which generates the response "The lion is currently resting."

[0731] 6. The device converts the response into audio and transmits it to the child through the speaker.

[0732] Prompt Sentence Examples

[0733] "What is the current state of this animal?"

[0734] What activities are lions currently engaged in?

[0735] "How is this penguin's health?"

[0736] In this way, the system of the present invention can quickly provide appropriate responses based on the real-time status of animals to enhance user experience at zoos and aquariums.

[0737] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0738] Step 1:

[0739] The server collects and stores the husbandry data. It receives data provided by the zookeepers and sensors and stores it in a MySQL database. For example, the lions' diet, exercise, and sleep duration are entered and stored in the database. Inputs include data entered manually by the zookeepers and data automatically sent from the sensors. The output is a stored database entry.

[0740] Step 2:

[0741] The device uses a camera to capture real-time video of the animals. The camera is installed in the animal's exhibition area and captures video in real time. The input is the video data captured by the camera, and the output is video frames for analysis.

[0742] Step 3:

[0743] The device analyzes the captured video to determine the animal's state. It analyzes the video data using technologies such as OpenCV and the YOLO model to identify the animal's behavior. For example, it determines whether a lion is resting or active. The input is the video frame obtained in step 2, and the output is the animal's state determination result (e.g., resting, active).

[0744] Step 4:

[0745] The user speaks a question about an animal into the terminal, for example, "What is this lion doing now?" The input is the voice data from the user, and the output is a recorded voice file.

[0746] Step 5:

[0747] The device converts the user's voice question into text using speech recognition technology, such as the Google Speech-to-Text API. The input is the audio file obtained in step 4, and the output is the converted text.

[0748] Step 6:

[0749] The server generates a response using the converted text data and the animal's state data. It uses a generative AI model (e.g., GPT-3) to create an appropriate response to the user's question. The input is the animal's state data obtained in step 3 and the text data obtained in step 5, and the output is the generated response text.

[0750] Step 7:

[0751] The device converts the response sent from the server into speech and provides it to the user. It uses the Google Text-to-Speech API or Amazon Polly to convert text to speech. The input is the response text generated in step 6, and the output is the synthesized speech data.

[0752] Step 8:

[0753] The terminal plays the generated audio data to the user through the speaker, so that the user can hear information about the animal's condition in real time. The input is the audio data obtained in step 7, and the output is the audio response provided to the user.

[0754] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0755] The present invention relates to a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate responses to questions from users. This system is composed of a server, terminals, users, and a sentiment analysis engine.

[0756] Overall system overview

[0757] The system consists of the following main parts:

[0758] 1. Server:

[0759] Collect and store breeding data.

[0760] A generative AI model is used to generate appropriate responses to user questions.

[0761] 2. Terminal:

[0762] It uses a camera and microphone to capture footage of animals and user voice questions.

[0763] The acquired data is analyzed and sent to the server.

[0764] The response received from the server is synthesized into voice and provided to the user.

[0765] 3. User:

[0766] Ask questions about animals by speaking them into the device.

[0767] 4. Sentiment Analysis Engine:

[0768] Recognizes emotions by analyzing the user's voice and text data.

[0769] Program processing

[0770] Collection and storage of breeding data

[0771] server

[0772] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[0773] Examples:

[0774] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[0775] Animal imaging and analysis

[0776] Terminal

[0777] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[0778] Examples:

[0779] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[0780] Acquiring and analyzing voice questions from users

[0781] User

[0782] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[0783] Terminal

[0784] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[0785] Sentiment Analysis Engine

[0786] The emotion analysis engine analyzes the user's emotions based on the acquired voice and text data, and determines whether the user is excited, calm, happy, etc.

[0787] Examples:

[0788] If a user excitedly asks, "What is this lion doing right now?", the sentiment analysis engine will determine from the tone of voice and the use of words that the user is excited.

[0789] Generating and serving the response

[0790] server

[0791] The server uses a generative AI model to generate a response based on the user's question text and the animal's state data. The response is also tailored to the user's emotions. For example, if the user is excited, the server might generate a response like, "The lion is currently resting. Please watch quietly!"

[0792] Terminal

[0793] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[0794] Examples:

[0795] The device receives the response generated by the server, "The lion is currently resting. Please watch over it quietly!", converts it into voice and conveys it to the user. In this way, the user can receive accurate information based on the animal's latest condition and feedback according to its individual emotions.

[0796] Processing flow through concrete examples

[0797] 1. The server stores the lion's health data in a database.

[0798] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[0799] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[0800] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[0801] 5. The sentiment analysis engine analyzes the user's tone of voice and text to determine if they are excited.

[0802] 6. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please watch over it quietly!"

[0803] 7. The device converts the response into audio and transmits it to the child through the speaker.

[0804] In this way, the system of the present invention improves the user experience at zoos and aquariums, allowing for quick and appropriate responses based on the animals' current status and the user's emotions.

[0805] The processing flow will be explained below.

[0806] Step 1:

[0807] server

[0808] Collects breeding data. The server receives data entered by keepers and data sent from sensors, and stores the animal's health status (e.g., amount of food eaten, amount of exercise, hours of sleep, etc.) in a database such as MySQL.

[0809] Specific operation: Keepers enter data such as food intake and exercise data via a web form and store it in a database using a PHP script. Data automatically sent by sensors is also received via a REST API and stored in the database in the same way.

[0810] Step 2:

[0811] Terminal

[0812] Obtaining images of animals. The terminal obtains real-time images of animals using IP cameras installed in cages, etc.

[0813] Specific operation: The terminal receives the video stream from the IP camera using the RTSP protocol and processes the video data in real time using the OpenCV library.

[0814] Step 3:

[0815] Terminal

[0816] Analyzing the video to determine the animal's state: The device analyzes the acquired video data to determine the animal's behavior and state (e.g., active, resting, sleeping).

[0817] How it works: The device loads a deep learning model (such as YOLO or PoseNet) and analyzes the animal's pose and movement for each video frame. For example, it uses the YOLO model to detect the animal's outline and PoseNet to estimate its pose.

[0818] Step 4:

[0819] User

[0820] Users speak questions into the device by voice. Users ask specific questions into the microphone about the animal's condition and activity.

[0821] Specific action: The user speaks into the microphone, saying something like, "What is this penguin doing right now?"

[0822] Step 5:

[0823] Terminal

[0824] Converts voice to text. The device converts the acquired voice data into text data using the Google Speech-to-Text API.

[0825] What it does: The device records audio data collected by the microphone in WAV format and sends the data to the Google Speech-to-Text API to convert it into text.

[0826] Step 6:

[0827] Terminal

[0828] Send the text and animal status data to the server. The converted text data and the determined animal status information are sent to the server using an HTTP POST request.

[0829] Specific operation: The device compiles the user's question text and the analyzed animal's condition data into JSON format and sends it to the server as an HTTP POST request.

[0830] Step 7:

[0831] server

[0832] Recognizing user emotions: The server uses a sentiment analysis engine to determine the user's emotions (e.g., excitement, surprise, joy, anger, etc.) from the received text data.

[0833] Specific operation: The server passes the received question text to a sentiment analysis library (e.g., Azure Emotion API or NLTK library) to obtain the user's sentiment.

[0834] Step 8:

[0835] server

[0836] Generate a response: The server uses a generative AI model to generate an appropriate response based on the user's question text, the animal's condition data, and the results of sentiment analysis.

[0837] Specific operation: The server inputs the question text, animal state data, and user emotion data into a generative AI model (e.g., GPT-3), and generates a response such as "The lion is currently resting. Please watch over it quietly."

[0838] Step 9:

[0839] server

[0840] Send the generated response to the terminal. Send the generated text response back to the terminal as an HTTP response.

[0841] Specific operation: The server compiles the generated response text into JSON format and sends it to the terminal using an HTTP response.

[0842] Step 10:

[0843] Terminal

[0844] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[0845] Specific operation: The device passes the received response text to the Google Text-to-Speech API to generate audio data and saves it as an audio file.

[0846] Step 11:

[0847] Terminal

[0848] Providing a voice response to the user: The terminal plays the generated voice data through a speaker to provide the user with an appropriate answer.

[0849] Specific operation: The device uses the playback function (for example, the playback function of the Pygame library) to play the generated audio file from the speaker.

[0850] For example, if a user asks the device, "What is this lion doing now?", the system analyzes the lion's state in real time, the server generates a response such as, "The lion is currently resting. Please watch over it quietly!", and the device plays back the response aloud. In this way, the user can obtain highly accurate information based on the animal's latest state and their own emotions in real time.

[0851] Example 2

[0852] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0853] Conventional animal monitoring and user information provision systems in zoos and aquariums are required not only to grasp the status of animals in real time, but also to respond quickly and appropriately to user questions. However, it is technically difficult to not only monitor animal behavior but also to generate responses based on the user's emotions. Furthermore, a comprehensive system is required, as advanced processing is required to respond appropriately to the diverse questions and emotions of users.

[0854] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0855] In this invention, the server includes means for collecting and saving breeding information, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the behavior of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for analyzing the user's emotions from the voice input, means for generating a response using the converted text, the determined animal status data, and the analyzed emotion data, and means for providing the generated response to the user by voice, thereby enabling the server to quickly provide an appropriate response based on both the animal's behavior and the user's emotions.

[0856] "Breeding information" refers to data about the animal's health, behavioral patterns, diet, amount of exercise, etc.

[0857] "Footage" refers to real-time video data of animals captured using a camera.

[0858] "Behavior" refers to the type of movement or activity an animal is engaged in at a particular time.

[0859] "Voice input" refers to voice data that a user utters to the system through a voice acquisition device such as a microphone.

[0860] "Text" refers to data that has been converted from voice data into text using voice recognition technology.

[0861] "Emotion" refers to the user's psychological state (e.g., excitement, calmness, joy) based on an analysis of voice tone and text content.

[0862] "Response" refers to the answer data generated by the system in response to a user's question.

[0863] A "generative AI model" refers to an algorithm that uses machine learning and artificial intelligence techniques to generate appropriate responses based on user questions and other data.

[0864] This invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for responding to user questions promptly and appropriately. This system is mainly composed of a server, terminals, users, and a sentiment analysis engine.

[0865] server

[0866] The server is primarily responsible for collecting and storing animal care information, including the animals' health, behavior, food intake, and exercise. This data is collected periodically from keepers and sensors and stored in a database such as MySQL. The server also uses generative AI models to generate appropriate responses to user questions.

[0867] Technology used

[0868] Database Management System (DBMS): MySQL

[0869] Generative AI models: Machine learning and artificial intelligence techniques

[0870] Terminal

[0871] The device captures real-time video footage using a camera installed in the animal cage and analyzes it to determine the animal's behavior. Deep learning models such as OpenCV, YOLO, and PoseNet are used for the analysis. It also receives voice questions from users and converts them into text using speech recognition technology. The acquired data is sent to a server, and the responses received from the server are provided to the user via voice synthesis.

[0872] Technology used

[0873] Video analysis: OpenCV, YOLO, PoseNet

[0874] Speech Recognition: Google Speech-to-Text API

[0875] Speech synthesis: Google Text-to-Speech API, Amazon Polly

[0876] User

[0877] Users can ask the system questions about the animal's condition and behavior by voice. For example, a child might ask the device, "What is this lion doing right now?" This voice data is captured by the device and converted into text using voice recognition technology.

[0878] Sentiment Analysis Engine

[0879] The emotion analysis engine analyzes the user's voice and text data to recognize their emotions, allowing it to determine whether they are excited, calm, or happy.

[0880] Technology used

[0881] Sentiment analysis: Natural language processing techniques, pitch analysis, text analysis

[0882] Specific examples

[0883] For example, imagine the following scenario at a zoo one day. A database connected to a server stores daily updated data on the lion's diet, exercise volume, and sleep duration. A device captures real-time footage using a camera installed in the lion's cage and uses OpenCV and the YOLO model to determine that the lion is currently resting. A child asks into the device's microphone, "What is this lion doing right now?" The device converts the voice into text using the Google Speech-to-Text API. An emotion analysis engine determines from the tone of the voice that the child is excited. The server uses a generative AI model to generate a response based on the prompt: "The lion is currently resting. Please watch over him quietly!" The device converts this response into speech using the Google Text-to-Speech API and relays it to the child through the speaker.

[0884] Prompt Sentence Examples

[0885] The user is asking about the current status of the lion. The lion is currently resting and the user is excited.

[0886] This allows the system to quickly provide appropriate responses based on the animal's condition and the user's emotions in zoos and aquariums.

[0887] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0888] Step 1: Collect and store breeding information

[0889] server

[0890] The server collects animal care data. Specifically, keepers enter data such as food intake and exercise volume into an input form, and this data is sent to the server as an HTTP request. In addition, sensors (e.g., vital sign monitors, activity meters) send animal health status data to the server in real time. The server receives this data and inserts it into the corresponding tables in the MySQL database.

[0891] Input: Keeper input data on food intake and exercise, and health status data from sensors

[0892] Data processing: Database insert processing, data consistency check

[0893] Output: Breeding information stored in the database

[0894] Step 2: Video acquisition and behavioral analysis

[0895] Terminal

[0896] The device captures real-time video using a camera installed in the animal cage. The video data is processed frame by frame and the animal's behavior is analyzed using deep learning models such as OpenCV, YOLO, and PoseNet. The behavioral results are sent to the server in JSON format.

[0897] Input: Real-time video captured from a camera

[0898] Data processing: video frame analysis, animal behavior recognition

[0899] Output: Action result (e.g., lion is resting)

[0900] Step 3: Capture voice questions and convert them to text

[0901] User

[0902] Users can speak into the device's microphone to ask questions about animals, such as, "What is this lion doing right now?"

[0903] Terminal

[0904] The device captures the user's voice with a microphone, stores the audio data in a buffer, and converts the captured audio data into text in real time using the Google Speech-to-Text API. The converted text data is then sent to the next step.

[0905] Input: User voice question

[0906] Data processing: converting voice data to text

[0907] Output: Question text (e.g., "What is this lion doing right now?")

[0908] Step 4: Sentiment Analysis

[0909] Sentiment Analysis Engine

[0910] The sentiment analysis engine analyzes the user's emotions based on voice and text data, determining the user's emotional state from the voice tone, speed, and text content, and sending the results to the response generation step.

[0911] Input: Text data of voice questions, voice tone data

[0912] Data processing: Natural language processing and speech tone analysis

[0913] Output: Emotion determination result (e.g., user is excited)

[0914] Step 5: Generate a response

[0915] server

[0916] The server uses a generative AI model to generate a response based on the user's question text, the animal's behavioral assessment results, and the analyzed emotion data. The server uses a prompt sentence to generate an appropriate response text.

[0917] Input: Question text, behavioral judgment result, emotion judgment result, prompt text

[0918] Data processing: Response generation using generative AI models

[0919] Output: Response text (e.g. "The lion is currently resting. Please keep quiet and watch over it!")

[0920] Step 6: Provide a voice response

[0921] Terminal

[0922] The device stores the response text received from the server in a buffer and converts it into audio data using the Text-to-Speech API, which is then provided to the user through the speaker.

[0923] Input: Response text

[0924] Data processing: Converting text data into speech

[0925] Output: Audio response (e.g. "The lion is currently resting. Please keep quiet and watch over it!" played from the speaker)

[0926] In this way, it is possible to quickly provide an appropriate response based on the animal's condition and the user's emotions.

[0927] (Application example 2)

[0928] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0929] Modern factories require systems that can monitor robots in real time to ensure they are operating efficiently and provide appropriate responses to staff questions. However, current systems struggle to not only monitor the robot's status in real time, but also to respond quickly and accurately to staff questions. To solve this problem, a system is needed that can analyze the robot's behavior and automatically generate responses that take staff emotions into account.

[0930] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0931] In this invention, the server includes a means for collecting and storing breeding data and behavior data, a means for acquiring images of animals and equipment in real time, and a means for analyzing the acquired images to determine the status of the animals and equipment. This makes it possible to monitor the behavior of robots in the factory in real time, and furthermore, to generate and provide quick and accurate responses to staff questions using a generative AI model.

[0932] "Animal and operation data" refers to information relating to the health, behavior, and operating status of animals and equipment.

[0933] "Means for acquiring images of animals or devices in real time" refers to devices that use cameras or sensors to observe the activities of animals or devices in real time and acquire the image data.

[0934] The "means for analyzing the acquired video and determining the state of the animal or device" is a system that analyzes the acquired video data and determines the current state of the animal or device based on the results.

[0935] The "means for receiving a question from a user as a voice input" is a device that acquires a voice question uttered by a user via a microphone or the like.

[0936] The "means for converting speech input to text" is a speech recognition system that converts acquired speech data into text data.

[0937] The "means for generating a response using the converted text and the determined state data of the animal or device" refers to a generative AI model that generates an appropriate response using the text converted from the voice data and the determined state data of the animal or device.

[0938] The "means for providing a generated response to a user by voice" is a system that synthesizes a generated text-based response into speech and provides it to a user by voice.

[0939] A "system for analyzing the behavior of animals and equipment" is a system that uses video analysis technology to understand the behavior of animals and equipment and determine their condition based on that information.

[0940] A "generative AI model" is an artificial intelligence model that uses machine learning technology to automatically generate responses from text.

[0941] A "prompt sentence" is an input sentence given to a generative AI model, and is a document based on which the generative AI model generates a response.

[0942] MODE FOR CARRYING OUT THE INVENTION

[0943] Overall system overview

[0944] This invention relates to a system that monitors the operation of robots and equipment in a factory in real time and responds quickly and appropriately to questions from staff. The system consists of a server, terminals, users, and a sentiment analysis engine. The system consists of the following main parts:

[0945] server

[0946] The server has the following roles:

[0947] Data collection and storage:

[0948] The server periodically collects operational and other status data from the robots and equipment in the factory and stores it in a database, including health data, operational data, etc. For example, it uses SQLAlchemy to connect to a MySQL database and stores this data in tables.

[0949] Terminal

[0950] The terminals are used by staff in the factory and have the following roles:

[0951] Video acquisition and analysis:

[0952] The device uses a camera to capture real-time video of the robot or equipment. The captured video is analyzed and the current status of the robot or equipment is determined based on the results. The analysis is performed using OpenCV and deep learning models (such as YOLO and PoseNet).

[0953] Voice Input and Recognition:

[0954] The user speaks a question into the device, and the voice is picked up through the microphone and converted into text using voice recognition technology, such as the Google Speech-to-Text API.

[0955] User

[0956] Users, especially factory staff, use the terminals as follows:

[0957] Submit a question:

[0958] The user can ask a voice question to the device, such as, "What is this robot doing right now?"

[0959] Sentiment Analysis Engine

[0960] The sentiment analysis engine has the following functions:

[0961] Emotion analysis:

[0962] The system analyzes the captured voice and text data to recognize the user's emotions, determining whether the user is excited or calm. The analysis is performed using the Google Cloud Natural Language API.

[0963] Generating and serving the response

[0964] The server uses a generative AI model to generate a response based on the user's question text and the status data of the robot and equipment. The generated response is adjusted according to the user's emotions. For example, if the user is excited, the response generated will be something like, "The robot is currently packaging. Please be careful."

[0965] Specific examples

[0966] Data collection:

[0967] The robot is monitored in real time through a camera installed on the staff's terminal, which captures images of the robot and analyzes them using OpenCV.

[0968] Voice questions:

[0969] A factory worker asks into the device's microphone, "What is this robot doing now?" The device converts the voice to text and uses the Google Speech-to-Text API to translate it.

[0970] Response generation:

[0971] Using the converted text and the analyzed robot state data, the server sends prompt sentences to the generative AI model to generate a response.

[0972] Response provided:

[0973] The generated response is converted into audio using the Google Text-to-Speech API and provided to staff via the terminal.

[0974] Prompt Sentence Examples

[0975] Question: What is this robot doing now?

[0976] Robot status: Packaging in progress

[0977] User sentiment score: 0.8

[0978] response:

[0979] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0980] Step 1:

[0981] The server collects operational and other status data from the robots and equipment in the factory and stores it in a database. The input is data from the sensors on the robots and equipment, and the output is organized and stored database entries. Specifically, it uses SQLAlchemy to connect to a MySQL database and periodically updates the operational data.

[0982] Step 2:

[0983] The device uses a camera to capture real-time video of the robot or equipment. The input is the video data captured by the camera, and the output is the current status of the robot or equipment as an analysis result. The device analyzes the captured video using OpenCV or a deep learning model (such as YOLO or PoseNet) to determine the work the robot or equipment is currently performing.

[0984] Step 3:

[0985] The user speaks a question into the terminal. The input is the user's voice question, and the output is voice data. For example, the user speaks a question such as "What is this robot doing now?" into the microphone of the terminal.

[0986] Step 4:

[0987] The device receives a voice question from the user and converts it into text. The input is the captured voice data, and the output is the converted text data. The device converts the voice data into text using the Google Speech-to-Text API.

[0988] Step 5:

[0989] The server generates a response using the converted text and the determined robot or equipment status data. The input is the user's question text and the robot or equipment status data, and the output is the generated response text. The server provides the prompt text to the generative AI model, which then generates an appropriate response. Examples of prompt text are as follows:

[0990] Question: What is this robot doing now?

[0991] Robot status: Packaging in progress

[0992] User sentiment score: 0.8

[0993] response:

[0994] Step 6:

[0995] The sentiment analysis engine analyzes the user's voice and text data to recognize emotions. The input is the user's voice data and converted text data, and the output is an emotion score. It uses the Google Cloud Natural Language API to analyze the user's emotions and provide the emotion score to the server.

[0996] Step 7:

[0997] The device synthesizes the response sent from the server and provides it to the user. The input is the response text received from the server, and the output is the response in audio format. The device uses the Google Text-to-Speech API to convert the text into audio and transmit it to the user through the speaker.

[0998] This allows for real-time monitoring of robots within the factory and appropriate responses to staff questions.

[0999] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1000] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1001] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1002] [Third embodiment]

[1003] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1004] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1005] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1006] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1007] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1008] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1009] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1010] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1011] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1012] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1013] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1014] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1015] The present invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for providing prompt and appropriate responses to questions from users. This system is composed of a server, terminals, and users.

[1016] Overall system overview

[1017] The system consists of the following main parts:

[1018] 1. Server:

[1019] Collect and store breeding data.

[1020] A generative AI model is used to generate appropriate responses to user questions.

[1021] 2. Terminal:

[1022] It uses a camera and microphone to capture footage of animals and user voice questions.

[1023] The acquired data is analyzed and sent to the server.

[1024] The response received from the server is synthesized into voice and provided to the user.

[1025] 3. User:

[1026] Ask questions about animals by speaking them into the device.

[1027] Program processing

[1028] Collection and storage of breeding data

[1029] server

[1030] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[1031] Examples:

[1032] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[1033] Animal imaging and analysis

[1034] Terminal

[1035] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[1036] Examples:

[1037] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[1038] Analysis of voice questions from users

[1039] User

[1040] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[1041] Terminal

[1042] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[1043] Examples:

[1044] A child asks the device, "What is this lion doing right now?" and the device converts the voice into text data.

[1045] Generating and serving the response

[1046] server

[1047] The server uses a generative AI model to generate a response based on the user's question text and the animal's status data. For example, if it determines that the lion is currently resting, it generates the response "The lion is currently resting. Please keep quiet."

[1048] Terminal

[1049] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[1050] Examples:

[1051] The device receives the response generated by the server, "The lion is currently resting. Please be quiet," converts it into voice and conveys it to the user.

[1052] Processing flow through concrete examples

[1053] 1. The server stores the lion's health data in a database.

[1054] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[1055] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[1056] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[1057] 5. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please keep quiet."

[1058] 6. The device converts the response into audio and transmits it to the child through the speaker.

[1059] In this way, the system of the present invention improves the user experience at zoos and aquariums, providing quick and appropriate responses based on the animals' current status.

[1060] The processing flow will be explained below.

[1061] Step 1:

[1062] server

[1063] Collects animal care data. The server receives input from keepers and data from sensors, and stores information about the animals' health status (e.g., amount of food eaten, amount of exercise, amount of sleep, etc.) in a database. For example, keepers enter the amount of food eaten each day using a web form, and the amount of exercise automatically recorded by sensors is stored in a database such as MySQL.

[1064] Step 2:

[1065] Terminal

[1066] Acquire video of the animal. The device acquires video in real time using a camera installed in the animal's cage, etc. The camera can be an IP camera or a USB-connected webcam. The device processes the video stream using libraries such as OpenCV and formats the data into a format that can be analyzed in real time.

[1067] Step 3:

[1068] Terminal

[1069] The device analyzes the video to determine the animal's state. The device analyzes the animal's movements and state based on the acquired video data. For example, it uses deep learning models (such as YOLO or PoseNet) to detect the animal's posture and movement and determine its state, such as "active," "resting," or "sleeping."

[1070] Step 4:

[1071] User

[1072] Users (especially children) can speak questions into the device by talking to the microphone and asking questions about the animal's condition or activity, such as "What is this penguin doing right now?"

[1073] Step 5:

[1074] Terminal

[1075] Converts voice to text. The device converts the user's voice questions captured through the microphone into text data using voice recognition technology such as the Google Speech-to-Text API.

[1076] Step 6:

[1077] Terminal

[1078] Send the text and animal status data to the server. The converted text data and the real-time animal status information are compiled and sent to the server using an HTTP POST request.

[1079] Step 7:

[1080] server

[1081] Generate a response: The server uses a generative AI model to generate an appropriate response based on the received text data and animal status data. For example, it generates a response such as "This penguin is currently resting."

[1082] Step 8:

[1083] server

[1084] The generated response is sent to the terminal. The generated response in text data format is returned to the terminal as an HTTP response.

[1085] Step 9:

[1086] Terminal

[1087] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[1088] Step 10:

[1089] Terminal

[1090] Providing a voice response to the user: The terminal plays the voice-converted response to the user through a speaker, providing an answer to the question.

[1091] For example, if a user asks the device, "What is this lion doing right now?", the system will analyze the lion's status in real time, the server will generate a response such as "The lion is currently resting," and the device will play back the response aloud. In this way, the user can obtain accurate information based on the animal's latest status.

[1092] Example 1

[1093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1094] In conventional zoos and aquariums, it has been difficult to grasp the condition of animals in real time and respond quickly and accurately to visitor questions. In addition, it has not been possible to integrate and utilize breeding data and real-time animal behavior data, which has limited the ability to improve visitor satisfaction and streamline breeding management.

[1095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1096] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the condition of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined animal condition data, and means for providing the generated response to the user by voice, thereby enabling zoos and aquariums to provide quick and appropriate responses based on the latest condition of animals.

[1097] "Breeding data" refers to data that includes information on the health, food intake, and exercise of animals kept in zoos and aquariums.

[1098] "Real-time" refers to the instantaneous acquisition and processing of information and situations at the current time.

[1099] "Animal condition" is a concept that includes information such as the animal's health, behavior, and activity status.

[1100] "Audio input" refers to speech produced by a user through a microphone or other audio capture device.

[1101] "Means for converting voice input to text" refers to technologies or systems that convert voice data into character data.

[1102] "Means for generating a response" refers to the technology or mechanism for generating an appropriate response to a user's question.

[1103] A "generative AI model" refers to an artificial intelligence model that uses natural language processing technology to understand and generate human language.

[1104] "Analytical algorithm" refers to a computational procedure or method for analyzing data and deriving a particular result or information.

[1105] System Overview

[1106] The present invention is a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate answers to questions from users. The system is composed of a server, terminals, and users, and has the following functions:

[1107] server

[1108] The server has the following functions:

[1109] 1. Collection and storage of rearing data:

[1110] The server receives breeding data entered by keepers using a dedicated application or web interface. For example, it periodically collects information such as the lions' diet, exercise, and weight. It also collects animal health data (heart rate, body temperature, etc.) obtained from sensors. The collected data is stored in a relational database such as MySQL or PostgreSQL.

[1111] Terminal

[1112] The device has the following features:

[1113] 1. Animal footage acquisition:

[1114] The device captures real-time images of animals using cameras installed in cages or aquariums, and the camera video data is temporarily stored in the device's local storage.

[1115] 2. Animal video analysis:

[1116] The device analyzes the stored video data to determine the animal's behavior using OpenCV and deep learning models (such as YOLO and PoseNet), allowing it to determine in real time whether the animal is currently active, resting, or performing a specific behavior.

[1117] 3. Acquire voice questions from users and convert them into text:

[1118] The user speaks to the device to ask a question about an animal. For example, "What is this lion doing right now?" The device then converts the spoken question picked up by the microphone into text. This process uses the Google Speech-to-Text API or IBM Watson Speech to Text.

[1119] 4. Speech synthesis and delivery of responses:

[1120] The response sent from the server is converted into speech using the Google Text-to-Speech API or Amazon Polly, and the generated speech data is then provided to the user through the speaker.

[1121] Specific examples of programs

[1122] Example of user interaction:

[1123] For example, if a user asks the terminal, "What is this lion doing right now?", the process is as follows:

[1124] 1. Acquiring a voice question: The user speaks into the device's microphone, "What is this lion doing right now?"

[1125] 2. Converting voice questions into text: The device sends the captured voice to the Google Speech-to-Text API and receives the text "What is this lion doing right now?"

[1126] 3. Determining the animal's state: The device analyzes the video data using OpenCV and determines that the lion is "resting."

[1127] 4. Generating and providing a response: The server inputs the data "The lion is currently resting" into the generative AI model, generates a response "The lion is currently resting. Please be quiet," and the device converts the response into audio and plays it through the speaker.

[1128] Prompt Sentence Examples

[1129] An example of a prompt sentence might be:

[1130] 1. "What is this lion doing now?"

[1131] 2. "When is the bear's meal time?"

[1132] 3. "Where are the penguins?"

[1133] This system allows zoos and aquariums to provide quick and appropriate responses based on the latest status of animals, improving visitor satisfaction and streamlining animal management.

[1134] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1135] Step 1: Collecting rearing data

[1136] The server receives the care data entered by the zookeepers using a dedicated application or web interface. When the zookeepers enter information such as the lions' food intake, exercise, and weight, it is sent to the server. The input is data entered manually by the zookeepers, and the server receives it and processes it into an appropriate format. The output is structured care data.

[1137] Specific behavior:

[1138] A zookeeper enters "Today's lion's food intake is 10kg," and the data is sent to the server, which structures the data and processes it into a format that can be stored in the database.

[1139] Step 2: Saving breeding data

[1140] The server stores the received breeding data in a database. The database used is MySQL or PostgreSQL. The input is structured data sent by the zookeepers, which the server inserts into the database. The output is the breeding data stored in the database.

[1141] Specific behavior:

[1142] The server inserts the data "Today's lion's food intake is 10 kg" into a MySQL database, which stores information such as the date, animal species, and food intake.

[1143] Step 3: Acquire footage of the animal

[1144] The device captures real-time video of the animals using a camera installed in the cage or aquarium. The input is video data from the camera, which the device temporarily stores in local storage. The output is the video data stored in local storage.

[1145] Specific behavior:

[1146] A camera inside the cage captures footage of the lion and sends it to the device, which stores the video data in its local storage.

[1147] Step 4: Animal footage analysis

[1148] The device analyzes the stored video data to determine the animal's behavior. OpenCV and deep learning models (YOLO and PoseNet) are used for the analysis. The input is the video data stored in local storage, which the device applies to the analysis algorithm. The output is the animal's behavior data.

[1149] Specific behavior:

[1150] The device analyzes the video data using OpenCV and determines whether the lion is "resting," "active," or "eating," and outputs the results as animal behavior data.

[1151] Step 5: Obtaining voice questions from the user

[1152] The user speaks a question about an animal into the device. The input is the user's voice question, which the device picks up with a microphone. The output is the captured voice data.

[1153] Specific behavior:

[1154] The user speaks into the device's microphone, saying, "What is this lion doing now?" The device captures this voice and processes it.

[1155] Step 6: Convert spoken questions to text

[1156] The device converts the acquired voice questions into text data using the Google Speech-to-Text API or IBM Watson Speech to Text. The input is the acquired voice data, and the output is the converted text data.

[1157] Specific behavior:

[1158] The device sends the voice data to the Google Speech-to-Text API and receives the text data, "What is this lion doing right now?"

[1159] Step 7: Generate a response

[1160] The server inputs the converted text question into a generative AI model to generate a response. The input is the user's text question and the animal's behavior data, and the output is the generated response text. Natural language processing technology is used to generate the response.

[1161] Specific behavior:

[1162] The server inputs the data "The lion is currently resting" into the generated AI model and generates the response "The lion is currently resting. Please be quiet."

[1163] Step 8: Text-to-speech response

[1164] The device converts the response sent from the server into speech using the Google Text-to-Speech API or Amazon Polly. The input is the response text data, and the output is the generated speech data.

[1165] Specific behavior:

[1166] The device converts the text response "The lion is currently resting. Please be quiet" into audio using the Google Text-to-Speech API.

[1167] Step 9: Providing a response

[1168] The terminal provides the generated voice data to the user through a speaker: the input is the voice data, and the output is a voice response that the user hears.

[1169] Specific behavior:

[1170] A voice will play from the device's speaker saying, "The lion is currently resting. Please be quiet."

[1171] In this way, the system can quickly provide appropriate responses based on the current status of animals at zoos and aquariums through each processing step.

[1172] (Application example 1)

[1173] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1174] In the past, visitors to zoos and aquariums needed information from display panels or explanations from keepers to know the current status of animals. However, this method made it difficult to obtain detailed information about the animals' behavior and health status in real time, and it was also difficult to provide immediate responses to questions. Furthermore, there was a lack of a system that could accurately grasp the animals' condition and provide appropriate explanations to visitors.

[1175] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1176] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the status of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined status data of the animals, means for providing the generated response to the user by voice, and means for the user to use at the zoo or aquarium via an application installed on their smartphone. This allows visitors to obtain detailed information about the current status and behavior of animals in real time and receive immediate responses to their voice questions, thereby improving the visitor experience.

[1177] "Breeding data" refers to information related to animal care management, such as the animal's health, amount of food eaten, amount of exercise, and amount of sleep.

[1178] "Animal footage" refers to real-time footage of animals on display at zoos and aquariums.

[1179] "Means for analyzing video footage to determine the condition of an animal" refers to technology for analyzing acquired video data of an animal and automatically determining the animal's behavior and health condition.

[1180] "Means for receiving questions from the user as voice input" refers to a mechanism for receiving and recording questions posed by the user via a device.

[1181] "Means for converting voice input to text" refers to technology that converts received voice data into text data.

[1182] The "means for generating a response" refers to a mechanism for generating an appropriate answer to a user's question using the converted text data and the animal's condition data.

[1183] "Means for providing a generated response to a user audibly" refers to technology that converts a generated text-based response into audio and provides it to a user.

[1184] "Applications installed on smartphones" refers to software that is downloaded and installed on the mobile device used by the user.

[1185] A "generative AI model" is a type of machine learning model that learns from large amounts of data and responds to and generates information such as text, images, and voice.

[1186] This invention is a system developed to improve the user experience at physical stores of zoos and aquariums. The system consists of a server, a terminal, and a user.

[1187] server

[1188] The server performs the following main functions:

[1189] 1. Collection and storage of rearing data:

[1190] The server periodically collects data on the animals' health and behavior from keepers and sensors and stores it in a database.

[1191] For example, this includes data on food intake and exercise volume entered by keepers, and data on sleep duration recorded by sensors.

[1192] 2. Generate a response:

[1193] A generative AI model is used to generate appropriate responses based on the user's question text and the animal's condition data.

[1194] For example, if it determines that the animal is currently resting, it generates the response "The lion is currently resting."

[1195] Examples of technologies used:

[1196] Database: MySQL

[1197] Generative AI models: large-scale language models such as GPT-3

[1198] Terminal

[1199] The terminal performs the following main functions:

[1200] 1. Animal image acquisition and analysis:

[1201] The system uses a camera to capture real-time video of the animals and analyzes the video data to determine their behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[1202] For example, cameras installed in lion cages record the lion's movements in real time and use the footage to determine whether the lion is currently active or resting.

[1203] 2. Voice input and conversion of user questions:

[1204] The system uses a microphone to capture the user's voice question and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[1205] For example, a child might ask, "What is this lion doing right now?" and the device would convert the voice into text data.

[1206] 3. Provide audio response:

[1207] The response sent from the server is synthesized and provided to the user using the Google Text-to-Speech API or Amazon Polly to convert the generated text into speech and play it back through the speaker.

[1208] For example, the device receives a response generated by the server saying "The lion is currently resting," converts it into voice, and conveys it to the user.

[1209] User

[1210] A user (e.g., a zoo visitor) accesses the system in the following ways:

[1211] 1. Speak your question:

[1212] The user speaks a question about an animal, such as "What is this lion doing right now?"

[1213] 2. Receiving the response:

[1214] The device responds with audio, giving users real-time information about the animal's condition.

[1215] Specific examples of processing

[1216] 1. The server stores the lion's health data in a database.

[1217] 2. The device uses its camera to capture footage of the lion and analyzes the footage to determine whether the lion is resting.

[1218] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[1219] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[1220] 5. The server passes the question to the generative AI model, which generates the response "The lion is currently resting."

[1221] 6. The device converts the response into audio and transmits it to the child through the speaker.

[1222] Prompt Sentence Examples

[1223] "What is the current state of this animal?"

[1224] What activities are lions currently engaged in?

[1225] "How is this penguin's health?"

[1226] In this way, the system of the present invention can quickly provide appropriate responses based on the real-time status of animals to enhance user experience at zoos and aquariums.

[1227] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1228] Step 1:

[1229] The server collects and stores the husbandry data. It receives data provided by the zookeepers and sensors and stores it in a MySQL database. For example, the lions' diet, exercise, and sleep duration are entered and stored in the database. Inputs include data entered manually by the zookeepers and data automatically sent from the sensors. The output is a stored database entry.

[1230] Step 2:

[1231] The device uses a camera to capture real-time video of the animals. The camera is installed in the animal's exhibition area and captures video in real time. The input is the video data captured by the camera, and the output is video frames for analysis.

[1232] Step 3:

[1233] The device analyzes the captured video to determine the animal's state. It analyzes the video data using technologies such as OpenCV and the YOLO model to identify the animal's behavior. For example, it determines whether a lion is resting or active. The input is the video frame obtained in step 2, and the output is the animal's state determination result (e.g., resting, active).

[1234] Step 4:

[1235] The user speaks a question about an animal into the terminal, for example, "What is this lion doing now?" The input is the voice data from the user, and the output is a recorded voice file.

[1236] Step 5:

[1237] The device converts the user's voice question into text using speech recognition technology, such as the Google Speech-to-Text API. The input is the audio file obtained in step 4, and the output is the converted text.

[1238] Step 6:

[1239] The server generates a response using the converted text data and the animal's state data. It uses a generative AI model (e.g., GPT-3) to create an appropriate response to the user's question. The input is the animal's state data obtained in step 3 and the text data obtained in step 5, and the output is the generated response text.

[1240] Step 7:

[1241] The device converts the response sent from the server into speech and provides it to the user. It uses the Google Text-to-Speech API or Amazon Polly to convert text to speech. The input is the response text generated in step 6, and the output is the synthesized speech data.

[1242] Step 8:

[1243] The terminal plays the generated audio data to the user through the speaker, so that the user can hear information about the animal's condition in real time. The input is the audio data obtained in step 7, and the output is the audio response provided to the user.

[1244] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1245] The present invention relates to a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate responses to questions from users. This system is composed of a server, terminals, users, and a sentiment analysis engine.

[1246] Overall system overview

[1247] The system consists of the following main parts:

[1248] 1. Server:

[1249] Collect and store breeding data.

[1250] A generative AI model is used to generate appropriate responses to user questions.

[1251] 2. Terminal:

[1252] It uses a camera and microphone to capture footage of animals and user voice questions.

[1253] The acquired data is analyzed and sent to the server.

[1254] The response received from the server is synthesized into voice and provided to the user.

[1255] 3. User:

[1256] Ask questions about animals by speaking them into the device.

[1257] 4. Sentiment Analysis Engine:

[1258] Recognizes emotions by analyzing the user's voice and text data.

[1259] Program processing

[1260] Collection and storage of breeding data

[1261] server

[1262] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[1263] Examples:

[1264] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[1265] Animal imaging and analysis

[1266] Terminal

[1267] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[1268] Examples:

[1269] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[1270] Acquiring and analyzing voice questions from users

[1271] User

[1272] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[1273] Terminal

[1274] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[1275] Sentiment Analysis Engine

[1276] The emotion analysis engine analyzes the user's emotions based on the acquired voice and text data, and determines whether the user is excited, calm, happy, etc.

[1277] Examples:

[1278] If a user excitedly asks, "What is this lion doing right now?", the sentiment analysis engine will determine from the tone of voice and the use of words that the user is excited.

[1279] Generating and serving the response

[1280] server

[1281] The server uses a generative AI model to generate a response based on the user's question text and the animal's state data. The response is also tailored to the user's emotions. For example, if the user is excited, the server might generate a response like, "The lion is currently resting. Please watch quietly!"

[1282] Terminal

[1283] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[1284] Examples:

[1285] The device receives the response generated by the server, "The lion is currently resting. Please watch over it quietly!", converts it into voice and conveys it to the user. In this way, the user can receive accurate information based on the animal's latest condition and feedback according to its individual emotions.

[1286] Processing flow through concrete examples

[1287] 1. The server stores the lion's health data in a database.

[1288] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[1289] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[1290] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[1291] 5. The sentiment analysis engine analyzes the user's tone of voice and text to determine if they are excited.

[1292] 6. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please watch over it quietly!"

[1293] 7. The device converts the response into audio and transmits it to the child through the speaker.

[1294] In this way, the system of the present invention improves the user experience at zoos and aquariums, allowing for quick and appropriate responses based on the animals' current status and the user's emotions.

[1295] The processing flow will be explained below.

[1296] Step 1:

[1297] server

[1298] Collects breeding data. The server receives data entered by keepers and data sent from sensors, and stores the animal's health status (e.g., amount of food eaten, amount of exercise, hours of sleep, etc.) in a database such as MySQL.

[1299] Specific operation: Keepers enter data such as food intake and exercise data via a web form and store it in a database using a PHP script. Data automatically sent by sensors is also received via a REST API and stored in the database in the same way.

[1300] Step 2:

[1301] Terminal

[1302] Obtaining images of animals. The terminal obtains real-time images of animals using IP cameras installed in cages, etc.

[1303] Specific operation: The terminal receives the video stream from the IP camera using the RTSP protocol and processes the video data in real time using the OpenCV library.

[1304] Step 3:

[1305] Terminal

[1306] Analyzing the video to determine the animal's state: The device analyzes the acquired video data to determine the animal's behavior and state (e.g., active, resting, sleeping).

[1307] How it works: The device loads a deep learning model (such as YOLO or PoseNet) and analyzes the animal's pose and movement for each video frame. For example, it uses the YOLO model to detect the animal's outline and PoseNet to estimate its pose.

[1308] Step 4:

[1309] User

[1310] Users speak questions into the device by voice. Users ask specific questions into the microphone about the animal's condition and activity.

[1311] Specific action: The user speaks into the microphone, saying something like, "What is this penguin doing right now?"

[1312] Step 5:

[1313] Terminal

[1314] Converts voice to text. The device converts the acquired voice data into text data using the Google Speech-to-Text API.

[1315] What it does: The device records audio data collected by the microphone in WAV format and sends the data to the Google Speech-to-Text API to convert it into text.

[1316] Step 6:

[1317] Terminal

[1318] Send the text and animal status data to the server. The converted text data and the determined animal status information are sent to the server using an HTTP POST request.

[1319] Specific operation: The device compiles the user's question text and the analyzed animal's condition data into JSON format and sends it to the server as an HTTP POST request.

[1320] Step 7:

[1321] server

[1322] Recognizing user emotions: The server uses a sentiment analysis engine to determine the user's emotions (e.g., excitement, surprise, joy, anger, etc.) from the received text data.

[1323] Specific operation: The server passes the received question text to a sentiment analysis library (e.g., Azure Emotion API or NLTK library) to obtain the user's sentiment.

[1324] Step 8:

[1325] server

[1326] Generate a response: The server uses a generative AI model to generate an appropriate response based on the user's question text, the animal's condition data, and the results of sentiment analysis.

[1327] Specific operation: The server inputs the question text, animal state data, and user emotion data into a generative AI model (e.g., GPT-3), and generates a response such as "The lion is currently resting. Please watch over it quietly."

[1328] Step 9:

[1329] server

[1330] Send the generated response to the terminal. Send the generated text response back to the terminal as an HTTP response.

[1331] Specific operation: The server compiles the generated response text into JSON format and sends it to the terminal using an HTTP response.

[1332] Step 10:

[1333] Terminal

[1334] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[1335] Specific operation: The device passes the received response text to the Google Text-to-Speech API to generate audio data and saves it as an audio file.

[1336] Step 11:

[1337] Terminal

[1338] Providing a voice response to the user: The terminal plays the generated voice data through a speaker to provide the user with an appropriate answer.

[1339] Specific operation: The device uses the playback function (for example, the playback function of the Pygame library) to play the generated audio file from the speaker.

[1340] For example, if a user asks the device, "What is this lion doing now?", the system analyzes the lion's state in real time, the server generates a response such as, "The lion is currently resting. Please watch over it quietly!", and the device plays back the response aloud. In this way, the user can obtain highly accurate information based on the animal's latest state and their own emotions in real time.

[1341] Example 2

[1342] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1343] Conventional animal monitoring and user information provision systems in zoos and aquariums are required not only to grasp the status of animals in real time, but also to respond quickly and appropriately to user questions. However, it is technically difficult to not only monitor animal behavior but also to generate responses based on the user's emotions. Furthermore, a comprehensive system is required, as advanced processing is required to respond appropriately to the diverse questions and emotions of users.

[1344] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1345] In this invention, the server includes means for collecting and saving breeding information, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the behavior of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for analyzing the user's emotions from the voice input, means for generating a response using the converted text, the determined animal status data, and the analyzed emotion data, and means for providing the generated response to the user by voice, thereby enabling the server to quickly provide an appropriate response based on both the animal's behavior and the user's emotions.

[1346] "Breeding information" refers to data about the animal's health, behavioral patterns, diet, amount of exercise, etc.

[1347] "Footage" refers to real-time video data of animals captured using a camera.

[1348] "Behavior" refers to the type of movement or activity an animal is engaged in at a particular time.

[1349] "Voice input" refers to voice data that a user utters to the system through a voice acquisition device such as a microphone.

[1350] "Text" refers to data that has been converted from voice data into text using voice recognition technology.

[1351] "Emotion" refers to the user's psychological state (e.g., excitement, calmness, joy) based on an analysis of voice tone and text content.

[1352] "Response" refers to the answer data generated by the system in response to a user's question.

[1353] A "generative AI model" refers to an algorithm that uses machine learning and artificial intelligence techniques to generate appropriate responses based on user questions and other data.

[1354] This invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for responding to user questions promptly and appropriately. This system is mainly composed of a server, terminals, users, and a sentiment analysis engine.

[1355] server

[1356] The server is primarily responsible for collecting and storing animal care information, including the animals' health, behavior, food intake, and exercise. This data is collected periodically from keepers and sensors and stored in a database such as MySQL. The server also uses generative AI models to generate appropriate responses to user questions.

[1357] Technology used

[1358] Database Management System (DBMS): MySQL

[1359] Generative AI models: Machine learning and artificial intelligence techniques

[1360] Terminal

[1361] The device captures real-time video footage using a camera installed in the animal cage and analyzes it to determine the animal's behavior. Deep learning models such as OpenCV, YOLO, and PoseNet are used for the analysis. It also receives voice questions from users and converts them into text using speech recognition technology. The acquired data is sent to a server, and the responses received from the server are provided to the user via voice synthesis.

[1362] Technology used

[1363] Video analysis: OpenCV, YOLO, PoseNet

[1364] Speech Recognition: Google Speech-to-Text API

[1365] Speech synthesis: Google Text-to-Speech API, Amazon Polly

[1366] User

[1367] Users can ask the system questions about the animal's condition and behavior by voice. For example, a child might ask the device, "What is this lion doing right now?" This voice data is captured by the device and converted into text using voice recognition technology.

[1368] Sentiment Analysis Engine

[1369] The emotion analysis engine analyzes the user's voice and text data to recognize their emotions, allowing it to determine whether they are excited, calm, or happy.

[1370] Technology used

[1371] Sentiment analysis: Natural language processing techniques, pitch analysis, text analysis

[1372] Specific examples

[1373] For example, imagine the following scenario at a zoo one day. A database connected to a server stores daily updated data on the lion's diet, exercise volume, and sleep duration. A device captures real-time footage using a camera installed in the lion's cage and uses OpenCV and the YOLO model to determine that the lion is currently resting. A child asks into the device's microphone, "What is this lion doing right now?" The device converts the voice into text using the Google Speech-to-Text API. An emotion analysis engine determines from the tone of the voice that the child is excited. The server uses a generative AI model to generate a response based on the prompt: "The lion is currently resting. Please watch over him quietly!" The device converts this response into speech using the Google Text-to-Speech API and relays it to the child through the speaker.

[1374] Prompt Sentence Examples

[1375] The user is asking about the current status of the lion. The lion is currently resting and the user is excited.

[1376] This allows the system to quickly provide appropriate responses based on the animal's condition and the user's emotions in zoos and aquariums.

[1377] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1378] Step 1: Collect and store breeding information

[1379] server

[1380] The server collects animal care data. Specifically, keepers enter data such as food intake and exercise volume into an input form, and this data is sent to the server as an HTTP request. In addition, sensors (e.g., vital sign monitors, activity meters) send animal health status data to the server in real time. The server receives this data and inserts it into the corresponding tables in the MySQL database.

[1381] Input: Keeper input data on food intake and exercise, and health status data from sensors

[1382] Data processing: Database insert processing, data consistency check

[1383] Output: Breeding information stored in the database

[1384] Step 2: Video acquisition and behavioral analysis

[1385] Terminal

[1386] The device captures real-time video using a camera installed in the animal cage. The video data is processed frame by frame and the animal's behavior is analyzed using deep learning models such as OpenCV, YOLO, and PoseNet. The behavioral results are sent to the server in JSON format.

[1387] Input: Real-time video captured from a camera

[1388] Data processing: video frame analysis, animal behavior recognition

[1389] Output: Action result (e.g., lion is resting)

[1390] Step 3: Capture voice questions and convert them to text

[1391] User

[1392] Users can speak into the device's microphone to ask questions about animals, such as, "What is this lion doing right now?"

[1393] Terminal

[1394] The device captures the user's voice with a microphone, stores the audio data in a buffer, and converts the captured audio data into text in real time using the Google Speech-to-Text API. The converted text data is then sent to the next step.

[1395] Input: User voice question

[1396] Data processing: converting voice data to text

[1397] Output: Question text (e.g., "What is this lion doing right now?")

[1398] Step 4: Sentiment Analysis

[1399] Sentiment Analysis Engine

[1400] The sentiment analysis engine analyzes the user's emotions based on voice and text data, determining the user's emotional state from the voice tone, speed, and text content, and sending the results to the response generation step.

[1401] Input: Text data of voice questions, voice tone data

[1402] Data processing: Natural language processing and speech tone analysis

[1403] Output: Emotion determination result (e.g., user is excited)

[1404] Step 5: Generate a response

[1405] server

[1406] The server uses a generative AI model to generate a response based on the user's question text, the animal's behavioral assessment results, and the analyzed emotion data. The server uses a prompt sentence to generate an appropriate response text.

[1407] Input: Question text, behavioral judgment result, emotion judgment result, prompt text

[1408] Data processing: Response generation using generative AI models

[1409] Output: Response text (e.g. "The lion is currently resting. Please keep quiet and watch over it!")

[1410] Step 6: Provide a voice response

[1411] Terminal

[1412] The device stores the response text received from the server in a buffer and converts it into audio data using the Text-to-Speech API, which is then provided to the user through the speaker.

[1413] Input: Response text

[1414] Data processing: Converting text data into speech

[1415] Output: Audio response (e.g. "The lion is currently resting. Please keep quiet and watch over it!" played from the speaker)

[1416] In this way, it is possible to quickly provide an appropriate response based on the animal's condition and the user's emotions.

[1417] (Application example 2)

[1418] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1419] Modern factories require systems that can monitor robots in real time to ensure they are operating efficiently and provide appropriate responses to staff questions. However, current systems struggle to not only monitor the robot's status in real time, but also to respond quickly and accurately to staff questions. To solve this problem, a system is needed that can analyze the robot's behavior and automatically generate responses that take staff emotions into account.

[1420] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1421] In this invention, the server includes a means for collecting and storing breeding data and behavior data, a means for acquiring images of animals and equipment in real time, and a means for analyzing the acquired images to determine the status of the animals and equipment. This makes it possible to monitor the behavior of robots in the factory in real time, and furthermore, to generate and provide quick and accurate responses to staff questions using a generative AI model.

[1422] "Animal and operation data" refers to information relating to the health, behavior, and operating status of animals and equipment.

[1423] "Means for acquiring images of animals or devices in real time" refers to devices that use cameras or sensors to observe the activities of animals or devices in real time and acquire the image data.

[1424] The "means for analyzing the acquired video and determining the state of the animal or device" is a system that analyzes the acquired video data and determines the current state of the animal or device based on the results.

[1425] The "means for receiving a question from a user as a voice input" is a device that acquires a voice question uttered by a user via a microphone or the like.

[1426] The "means for converting speech input to text" is a speech recognition system that converts acquired speech data into text data.

[1427] The "means for generating a response using the converted text and the determined state data of the animal or device" refers to a generative AI model that generates an appropriate response using the text converted from the voice data and the determined state data of the animal or device.

[1428] The "means for providing a generated response to a user by voice" is a system that synthesizes a generated text-based response into speech and provides it to a user by voice.

[1429] A "system for analyzing the behavior of animals and equipment" is a system that uses video analysis technology to understand the behavior of animals and equipment and determine their condition based on that information.

[1430] A "generative AI model" is an artificial intelligence model that uses machine learning technology to automatically generate responses from text.

[1431] A "prompt sentence" is an input sentence given to a generative AI model, and is a document based on which the generative AI model generates a response.

[1432] MODE FOR CARRYING OUT THE INVENTION

[1433] Overall system overview

[1434] This invention relates to a system that monitors the operation of robots and equipment in a factory in real time and responds quickly and appropriately to questions from staff. The system consists of a server, terminals, users, and a sentiment analysis engine. The system consists of the following main parts:

[1435] server

[1436] The server has the following roles:

[1437] Data collection and storage:

[1438] The server periodically collects operational and other status data from the robots and equipment in the factory and stores it in a database, including health data, operational data, etc. For example, it uses SQLAlchemy to connect to a MySQL database and stores this data in tables.

[1439] Terminal

[1440] The terminals are used by staff in the factory and have the following roles:

[1441] Video acquisition and analysis:

[1442] The device uses a camera to capture real-time video of the robot or equipment. The captured video is analyzed and the current status of the robot or equipment is determined based on the results. The analysis is performed using OpenCV and deep learning models (such as YOLO and PoseNet).

[1443] Voice Input and Recognition:

[1444] The user speaks a question into the device, and the voice is picked up through the microphone and converted into text using voice recognition technology, such as the Google Speech-to-Text API.

[1445] User

[1446] Users, especially factory staff, use the terminals as follows:

[1447] Submit a question:

[1448] The user can ask a voice question to the device, such as, "What is this robot doing right now?"

[1449] Sentiment Analysis Engine

[1450] The sentiment analysis engine has the following functions:

[1451] Emotion analysis:

[1452] The system analyzes the captured voice and text data to recognize the user's emotions, determining whether the user is excited or calm. The analysis is performed using the Google Cloud Natural Language API.

[1453] Generating and serving the response

[1454] The server uses a generative AI model to generate a response based on the user's question text and the status data of the robot and equipment. The generated response is adjusted according to the user's emotions. For example, if the user is excited, the response generated will be something like, "The robot is currently packaging. Please be careful."

[1455] Specific examples

[1456] Data collection:

[1457] The robot is monitored in real time through a camera installed on the staff's terminal, which captures images of the robot and analyzes them using OpenCV.

[1458] Voice questions:

[1459] A factory worker asks into the device's microphone, "What is this robot doing now?" The device converts the voice to text and uses the Google Speech-to-Text API to translate it.

[1460] Response generation:

[1461] Using the converted text and the analyzed robot state data, the server sends prompt sentences to the generative AI model to generate a response.

[1462] Response provided:

[1463] The generated response is converted into audio using the Google Text-to-Speech API and provided to staff via the terminal.

[1464] Prompt Sentence Examples

[1465] Question: What is this robot doing now?

[1466] Robot status: Packaging in progress

[1467] User sentiment score: 0.8

[1468] response:

[1469] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1470] Step 1:

[1471] The server collects operational and other status data from the robots and equipment in the factory and stores it in a database. The input is data from the sensors on the robots and equipment, and the output is organized and stored database entries. Specifically, it uses SQLAlchemy to connect to a MySQL database and periodically updates the operational data.

[1472] Step 2:

[1473] The device uses a camera to capture real-time video of the robot or equipment. The input is the video data captured by the camera, and the output is the current status of the robot or equipment as an analysis result. The device analyzes the captured video using OpenCV or a deep learning model (such as YOLO or PoseNet) to determine the work the robot or equipment is currently performing.

[1474] Step 3:

[1475] The user speaks a question into the terminal. The input is the user's voice question, and the output is voice data. For example, the user speaks a question such as "What is this robot doing now?" into the microphone of the terminal.

[1476] Step 4:

[1477] The device receives a voice question from the user and converts it into text. The input is the captured voice data, and the output is the converted text data. The device converts the voice data into text using the Google Speech-to-Text API.

[1478] Step 5:

[1479] The server generates a response using the converted text and the determined robot or equipment status data. The input is the user's question text and the robot or equipment status data, and the output is the generated response text. The server provides the prompt text to the generative AI model, which then generates an appropriate response. Examples of prompt text are as follows:

[1480] Question: What is this robot doing now?

[1481] Robot status: Packaging in progress

[1482] User sentiment score: 0.8

[1483] response:

[1484] Step 6:

[1485] The sentiment analysis engine analyzes the user's voice and text data to recognize emotions. The input is the user's voice data and converted text data, and the output is an emotion score. It uses the Google Cloud Natural Language API to analyze the user's emotions and provide the emotion score to the server.

[1486] Step 7:

[1487] The device synthesizes the response sent from the server and provides it to the user. The input is the response text received from the server, and the output is the response in audio format. The device uses the Google Text-to-Speech API to convert the text into audio and transmit it to the user through the speaker.

[1488] This allows for real-time monitoring of robots within the factory and appropriate responses to staff questions.

[1489] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1490] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1491] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1492] [Fourth embodiment]

[1493] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1494] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1495] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1496] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1497] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1498] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1499] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1500] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1501] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1502] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1503] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1504] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1505] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1506] The present invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for providing prompt and appropriate responses to questions from users. This system is composed of a server, terminals, and users.

[1507] Overall system overview

[1508] The system consists of the following main parts:

[1509] 1. Server:

[1510] Collect and store breeding data.

[1511] A generative AI model is used to generate appropriate responses to user questions.

[1512] 2. Terminal:

[1513] It uses a camera and microphone to capture footage of animals and user voice questions.

[1514] The acquired data is analyzed and sent to the server.

[1515] The response received from the server is synthesized into voice and provided to the user.

[1516] 3. User:

[1517] Ask questions about animals by speaking them into the device.

[1518] Program processing

[1519] Collection and storage of breeding data

[1520] server

[1521] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[1522] Examples:

[1523] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[1524] Animal imaging and analysis

[1525] Terminal

[1526] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[1527] Examples:

[1528] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[1529] Analysis of voice questions from users

[1530] User

[1531] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[1532] Terminal

[1533] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[1534] Examples:

[1535] A child asks the device, "What is this lion doing right now?" and the device converts the voice into text data.

[1536] Generating and serving the response

[1537] server

[1538] The server uses a generative AI model to generate a response based on the user's question text and the animal's status data. For example, if it determines that the lion is currently resting, it generates the response "The lion is currently resting. Please keep quiet."

[1539] Terminal

[1540] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[1541] Examples:

[1542] The device receives the response generated by the server, "The lion is currently resting. Please be quiet," converts it into voice and conveys it to the user.

[1543] Processing flow through concrete examples

[1544] 1. The server stores the lion's health data in a database.

[1545] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[1546] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[1547] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[1548] 5. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please keep quiet."

[1549] 6. The device converts the response into audio and transmits it to the child through the speaker.

[1550] In this way, the system of the present invention improves the user experience at zoos and aquariums, providing quick and appropriate responses based on the animals' current status.

[1551] The processing flow will be explained below.

[1552] Step 1:

[1553] server

[1554] Collects animal care data. The server receives input from keepers and data from sensors, and stores information about the animals' health status (e.g., amount of food eaten, amount of exercise, amount of sleep, etc.) in a database. For example, keepers enter the amount of food eaten each day using a web form, and the amount of exercise automatically recorded by sensors is stored in a database such as MySQL.

[1555] Step 2:

[1556] Terminal

[1557] Acquire video of the animal. The device acquires video in real time using a camera installed in the animal's cage, etc. The camera can be an IP camera or a USB-connected webcam. The device processes the video stream using libraries such as OpenCV and formats the data into a format that can be analyzed in real time.

[1558] Step 3:

[1559] Terminal

[1560] The device analyzes the video to determine the animal's state. The device analyzes the animal's movements and state based on the acquired video data. For example, it uses deep learning models (such as YOLO or PoseNet) to detect the animal's posture and movement and determine its state, such as "active," "resting," or "sleeping."

[1561] Step 4:

[1562] User

[1563] Users (especially children) can speak questions into the device by talking to the microphone and asking questions about the animal's condition or activity, such as "What is this penguin doing right now?"

[1564] Step 5:

[1565] Terminal

[1566] Converts voice to text. The device converts the user's voice questions captured through the microphone into text data using voice recognition technology such as the Google Speech-to-Text API.

[1567] Step 6:

[1568] Terminal

[1569] Send the text and animal status data to the server. The converted text data and the real-time animal status information are compiled and sent to the server using an HTTP POST request.

[1570] Step 7:

[1571] server

[1572] Generate a response: The server uses a generative AI model to generate an appropriate response based on the received text data and animal status data. For example, it generates a response such as "This penguin is currently resting."

[1573] Step 8:

[1574] server

[1575] The generated response is sent to the terminal. The generated response in text data format is returned to the terminal as an HTTP response.

[1576] Step 9:

[1577] Terminal

[1578] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[1579] Step 10:

[1580] Terminal

[1581] Providing a voice response to the user: The terminal plays the voice-converted response to the user through a speaker, providing an answer to the question.

[1582] For example, if a user asks the device, "What is this lion doing right now?", the system will analyze the lion's status in real time, the server will generate a response such as "The lion is currently resting," and the device will play back the response aloud. In this way, the user can obtain accurate information based on the animal's latest status.

[1583] Example 1

[1584] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1585] In conventional zoos and aquariums, it has been difficult to grasp the condition of animals in real time and respond quickly and accurately to visitor questions. In addition, it has not been possible to integrate and utilize breeding data and real-time animal behavior data, which has limited the ability to improve visitor satisfaction and streamline breeding management.

[1586] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1587] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the condition of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined animal condition data, and means for providing the generated response to the user by voice, thereby enabling zoos and aquariums to provide quick and appropriate responses based on the latest condition of animals.

[1588] "Breeding data" refers to data that includes information on the health, food intake, and exercise of animals kept in zoos and aquariums.

[1589] "Real-time" refers to the instantaneous acquisition and processing of information and situations at the current time.

[1590] "Animal condition" is a concept that includes information such as the animal's health, behavior, and activity status.

[1591] "Audio input" refers to speech produced by a user through a microphone or other audio capture device.

[1592] "Means for converting voice input to text" refers to technologies or systems that convert voice data into character data.

[1593] "Means for generating a response" refers to the technology or mechanism for generating an appropriate response to a user's question.

[1594] A "generative AI model" refers to an artificial intelligence model that uses natural language processing technology to understand and generate human language.

[1595] "Analytical algorithm" refers to a computational procedure or method for analyzing data and deriving a particular result or information.

[1596] System Overview

[1597] The present invention is a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate answers to questions from users. The system is composed of a server, terminals, and users, and has the following functions:

[1598] server

[1599] The server has the following functions:

[1600] 1. Collection and storage of rearing data:

[1601] The server receives breeding data entered by keepers using a dedicated application or web interface. For example, it periodically collects information such as the lions' diet, exercise, and weight. It also collects animal health data (heart rate, body temperature, etc.) obtained from sensors. The collected data is stored in a relational database such as MySQL or PostgreSQL.

[1602] Terminal

[1603] The device has the following features:

[1604] 1. Animal footage acquisition:

[1605] The device captures real-time images of animals using cameras installed in cages or aquariums, and the camera video data is temporarily stored in the device's local storage.

[1606] 2. Animal video analysis:

[1607] The device analyzes the stored video data to determine the animal's behavior using OpenCV and deep learning models (such as YOLO and PoseNet), allowing it to determine in real time whether the animal is currently active, resting, or performing a specific behavior.

[1608] 3. Acquire voice questions from users and convert them into text:

[1609] The user speaks to the device to ask a question about an animal. For example, "What is this lion doing right now?" The device then converts the spoken question picked up by the microphone into text. This process uses the Google Speech-to-Text API or IBM Watson Speech to Text.

[1610] 4. Speech synthesis and delivery of responses:

[1611] The response sent from the server is converted into speech using the Google Text-to-Speech API or Amazon Polly, and the generated speech data is then provided to the user through the speaker.

[1612] Specific examples of programs

[1613] Example of user interaction:

[1614] For example, if a user asks the terminal, "What is this lion doing right now?", the process is as follows:

[1615] 1. Acquiring a voice question: The user speaks into the device's microphone, "What is this lion doing right now?"

[1616] 2. Converting voice questions into text: The device sends the captured voice to the Google Speech-to-Text API and receives the text "What is this lion doing right now?"

[1617] 3. Determining the animal's state: The device analyzes the video data using OpenCV and determines that the lion is "resting."

[1618] 4. Generating and providing a response: The server inputs the data "The lion is currently resting" into the generative AI model, generates a response "The lion is currently resting. Please be quiet," and the device converts the response into audio and plays it through the speaker.

[1619] Prompt Sentence Examples

[1620] An example of a prompt sentence might be:

[1621] 1. "What is this lion doing now?"

[1622] 2. "When is the bear's meal time?"

[1623] 3. "Where are the penguins?"

[1624] This system allows zoos and aquariums to provide quick and appropriate responses based on the latest status of animals, improving visitor satisfaction and streamlining animal management.

[1625] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1626] Step 1: Collecting rearing data

[1627] The server receives the care data entered by the zookeepers using a dedicated application or web interface. When the zookeepers enter information such as the lions' food intake, exercise, and weight, it is sent to the server. The input is data entered manually by the zookeepers, and the server receives it and processes it into an appropriate format. The output is structured care data.

[1628] Specific behavior:

[1629] A zookeeper enters "Today's lion's food intake is 10kg," and the data is sent to the server, which structures the data and processes it into a format that can be stored in the database.

[1630] Step 2: Saving breeding data

[1631] The server stores the received breeding data in a database. The database used is MySQL or PostgreSQL. The input is structured data sent by the zookeepers, which the server inserts into the database. The output is the breeding data stored in the database.

[1632] Specific behavior:

[1633] The server inserts the data "Today's lion's food intake is 10 kg" into a MySQL database, which stores information such as the date, animal species, and food intake.

[1634] Step 3: Acquire footage of the animal

[1635] The device captures real-time video of the animals using a camera installed in the cage or aquarium. The input is video data from the camera, which the device temporarily stores in local storage. The output is the video data stored in local storage.

[1636] Specific behavior:

[1637] A camera inside the cage captures footage of the lion and sends it to the device, which stores the video data in its local storage.

[1638] Step 4: Animal footage analysis

[1639] The device analyzes the stored video data to determine the animal's behavior. OpenCV and deep learning models (YOLO and PoseNet) are used for the analysis. The input is the video data stored in local storage, which the device applies to the analysis algorithm. The output is the animal's behavior data.

[1640] Specific behavior:

[1641] The device analyzes the video data using OpenCV and determines whether the lion is "resting," "active," or "eating," and outputs the results as animal behavior data.

[1642] Step 5: Obtaining voice questions from the user

[1643] The user speaks a question about an animal into the device. The input is the user's voice question, which the device picks up with a microphone. The output is the captured voice data.

[1644] Specific behavior:

[1645] The user speaks into the device's microphone, saying, "What is this lion doing now?" The device captures this voice and processes it.

[1646] Step 6: Convert spoken questions to text

[1647] The device converts the acquired voice questions into text data using the Google Speech-to-Text API or IBM Watson Speech to Text. The input is the acquired voice data, and the output is the converted text data.

[1648] Specific behavior:

[1649] The device sends the voice data to the Google Speech-to-Text API and receives the text data, "What is this lion doing right now?"

[1650] Step 7: Generate a response

[1651] The server inputs the converted text question into a generative AI model to generate a response. The input is the user's text question and the animal's behavior data, and the output is the generated response text. Natural language processing technology is used to generate the response.

[1652] Specific behavior:

[1653] The server inputs the data "The lion is currently resting" into the generated AI model and generates the response "The lion is currently resting. Please be quiet."

[1654] Step 8: Text-to-speech response

[1655] The device converts the response sent from the server into speech using the Google Text-to-Speech API or Amazon Polly. The input is the response text data, and the output is the generated speech data.

[1656] Specific behavior:

[1657] The device converts the text response "The lion is currently resting. Please be quiet" into audio using the Google Text-to-Speech API.

[1658] Step 9: Providing a response

[1659] The terminal provides the generated voice data to the user through a speaker: the input is the voice data, and the output is a voice response that the user hears.

[1660] Specific behavior:

[1661] A voice will play from the device's speaker saying, "The lion is currently resting. Please be quiet."

[1662] In this way, the system can quickly provide appropriate responses based on the current status of animals at zoos and aquariums through each processing step.

[1663] (Application example 1)

[1664] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1665] In the past, visitors to zoos and aquariums needed information from display panels or explanations from keepers to know the current status of animals. However, this method made it difficult to obtain detailed information about the animals' behavior and health status in real time, and it was also difficult to provide immediate responses to questions. Furthermore, there was a lack of a system that could accurately grasp the animals' condition and provide appropriate explanations to visitors.

[1666] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1667] In this invention, the server includes means for collecting and storing breeding data, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the status of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for generating a response using the converted text and the determined status data of the animals, means for providing the generated response to the user by voice, and means for the user to use at the zoo or aquarium via an application installed on their smartphone. This allows visitors to obtain detailed information about the current status and behavior of animals in real time and receive immediate responses to their voice questions, thereby improving the visitor experience.

[1668] "Breeding data" refers to information related to animal care management, such as the animal's health, amount of food eaten, amount of exercise, and amount of sleep.

[1669] "Animal footage" refers to real-time footage of animals on display at zoos and aquariums.

[1670] "Means for analyzing video footage to determine the condition of an animal" refers to technology for analyzing acquired video data of an animal and automatically determining the animal's behavior and health condition.

[1671] "Means for receiving questions from the user as voice input" refers to a mechanism for receiving and recording questions posed by the user via a device.

[1672] "Means for converting voice input to text" refers to technology that converts received voice data into text data.

[1673] The "means for generating a response" refers to a mechanism for generating an appropriate answer to a user's question using the converted text data and the animal's condition data.

[1674] "Means for providing a generated response to a user audibly" refers to technology that converts a generated text-based response into audio and provides it to a user.

[1675] "Applications installed on smartphones" refers to software that is downloaded and installed on the mobile device used by the user.

[1676] A "generative AI model" is a type of machine learning model that learns from large amounts of data and responds to and generates information such as text, images, and voice.

[1677] This invention is a system developed to improve the user experience at physical stores of zoos and aquariums. The system consists of a server, a terminal, and a user.

[1678] server

[1679] The server performs the following main functions:

[1680] 1. Collection and storage of rearing data:

[1681] The server periodically collects data on the animals' health and behavior from keepers and sensors and stores it in a database.

[1682] For example, this includes data on food intake and exercise volume entered by keepers, and data on sleep duration recorded by sensors.

[1683] 2. Generate a response:

[1684] A generative AI model is used to generate appropriate responses based on the user's question text and the animal's condition data.

[1685] For example, if it determines that the animal is currently resting, it generates the response "The lion is currently resting."

[1686] Examples of technologies used:

[1687] Database: MySQL

[1688] Generative AI models: large-scale language models such as GPT-3

[1689] Terminal

[1690] The terminal performs the following main functions:

[1691] 1. Animal image acquisition and analysis:

[1692] The system uses a camera to capture real-time video of the animals and analyzes the video data to determine their behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[1693] For example, cameras installed in lion cages record the lion's movements in real time and use the footage to determine whether the lion is currently active or resting.

[1694] 2. Voice input and conversion of user questions:

[1695] The system uses a microphone to capture the user's voice question and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[1696] For example, a child might ask, "What is this lion doing right now?" and the device would convert the voice into text data.

[1697] 3. Provide audio response:

[1698] The response sent from the server is synthesized and provided to the user using the Google Text-to-Speech API or Amazon Polly to convert the generated text into speech and play it back through the speaker.

[1699] For example, the device receives a response generated by the server saying "The lion is currently resting," converts it into voice, and conveys it to the user.

[1700] User

[1701] A user (e.g., a zoo visitor) accesses the system in the following ways:

[1702] 1. Speak your question:

[1703] The user speaks a question about an animal, such as "What is this lion doing right now?"

[1704] 2. Receiving the response:

[1705] The device responds with audio, giving users real-time information about the animal's condition.

[1706] Specific examples of processing

[1707] 1. The server stores the lion's health data in a database.

[1708] 2. The device uses its camera to capture footage of the lion and analyzes the footage to determine whether the lion is resting.

[1709] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[1710] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[1711] 5. The server passes the question to the generative AI model, which generates the response "The lion is currently resting."

[1712] 6. The device converts the response into audio and transmits it to the child through the speaker.

[1713] Prompt Sentence Examples

[1714] "What is the current state of this animal?"

[1715] What activities are lions currently engaged in?

[1716] "How is this penguin's health?"

[1717] In this way, the system of the present invention can quickly provide appropriate responses based on the real-time status of animals to enhance user experience at zoos and aquariums.

[1718] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1719] Step 1:

[1720] The server collects and stores the husbandry data. It receives data provided by the zookeepers and sensors and stores it in a MySQL database. For example, the lions' diet, exercise, and sleep duration are entered and stored in the database. Inputs include data entered manually by the zookeepers and data automatically sent from the sensors. The output is a stored database entry.

[1721] Step 2:

[1722] The device uses a camera to capture real-time video of the animals. The camera is installed in the animal's exhibition area and captures video in real time. The input is the video data captured by the camera, and the output is video frames for analysis.

[1723] Step 3:

[1724] The device analyzes the captured video to determine the animal's state. It analyzes the video data using technologies such as OpenCV and the YOLO model to identify the animal's behavior. For example, it determines whether a lion is resting or active. The input is the video frame obtained in step 2, and the output is the animal's state determination result (e.g., resting, active).

[1725] Step 4:

[1726] The user speaks a question about an animal into the terminal, for example, "What is this lion doing now?" The input is the voice data from the user, and the output is a recorded voice file.

[1727] Step 5:

[1728] The device converts the user's voice question into text using speech recognition technology, such as the Google Speech-to-Text API. The input is the audio file obtained in step 4, and the output is the converted text.

[1729] Step 6:

[1730] The server generates a response using the converted text data and the animal's state data. It uses a generative AI model (e.g., GPT-3) to create an appropriate response to the user's question. The input is the animal's state data obtained in step 3 and the text data obtained in step 5, and the output is the generated response text.

[1731] Step 7:

[1732] The device converts the response sent from the server into speech and provides it to the user. It uses the Google Text-to-Speech API or Amazon Polly to convert text to speech. The input is the response text generated in step 6, and the output is the synthesized speech data.

[1733] Step 8:

[1734] The terminal plays the generated audio data to the user through the speaker, so that the user can hear information about the animal's condition in real time. The input is the audio data obtained in step 7, and the output is the audio response provided to the user.

[1735] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1736] The present invention relates to a system that monitors the status of animals in zoos and aquariums in real time and provides prompt and appropriate responses to questions from users. This system is composed of a server, terminals, users, and a sentiment analysis engine.

[1737] Overall system overview

[1738] The system consists of the following main parts:

[1739] 1. Server:

[1740] Collect and store breeding data.

[1741] A generative AI model is used to generate appropriate responses to user questions.

[1742] 2. Terminal:

[1743] It uses a camera and microphone to capture footage of animals and user voice questions.

[1744] The acquired data is analyzed and sent to the server.

[1745] The response received from the server is synthesized into voice and provided to the user.

[1746] 3. User:

[1747] Ask questions about animals by speaking them into the device.

[1748] 4. Sentiment Analysis Engine:

[1749] Recognizes emotions by analyzing the user's voice and text data.

[1750] Program processing

[1751] Collection and storage of breeding data

[1752] server

[1753] The server periodically collects data on the animals' health and behavior from keepers and sensors, and stores it in a database, including data on food intake and exercise volume entered by keepers and data on sleep duration recorded by sensors.

[1754] Examples:

[1755] The server stores daily updated data on the lions' diet, exercise levels, and sleep duration in a database such as MySQL.

[1756] Animal imaging and analysis

[1757] Terminal

[1758] The device uses a camera to capture real-time video of the animal and analyzes the video data to determine the animal's behavior, using OpenCV and deep learning models (such as YOLO and PoseNet).

[1759] Examples:

[1760] Cameras installed in the lions' cages record their movements in real time and use the footage to determine whether the lions are currently active or resting.

[1761] Acquiring and analyzing voice questions from users

[1762] User

[1763] Users (especially children) can ask questions about animals by voice into the device, such as "What is this lion doing right now?"

[1764] Terminal

[1765] The device uses a microphone to capture the user's voice and converts it into text using voice recognition technology, such as the Google Speech-to-Text API.

[1766] Sentiment Analysis Engine

[1767] The emotion analysis engine analyzes the user's emotions based on the acquired voice and text data, and determines whether the user is excited, calm, happy, etc.

[1768] Examples:

[1769] If a user excitedly asks, "What is this lion doing right now?", the sentiment analysis engine will determine from the tone of voice and the use of words that the user is excited.

[1770] Generating and serving the response

[1771] server

[1772] The server uses a generative AI model to generate a response based on the user's question text and the animal's state data. The response is also tailored to the user's emotions. For example, if the user is excited, the server might generate a response like, "The lion is currently resting. Please watch quietly!"

[1773] Terminal

[1774] The device receives the response from the server, synthesizes it, and provides it to the user. For example, the generated text is converted into speech using the Google Text-to-Speech API or Amazon Polly, and played back through the speaker.

[1775] Examples:

[1776] The device receives the response generated by the server, "The lion is currently resting. Please watch over it quietly!", converts it into voice and conveys it to the user. In this way, the user can receive accurate information based on the animal's latest condition and feedback according to its individual emotions.

[1777] Processing flow through concrete examples

[1778] 1. The server stores the lion's health data in a database.

[1779] 2. The device uses a camera to capture footage of the lion, analyzes it, and determines whether the lion is resting.

[1780] 3. The child asks into the device's microphone, "What is this lion doing right now?"

[1781] 4. The device converts the speech into text and sends the text "What is this lion doing right now?" to the server.

[1782] 5. The sentiment analysis engine analyzes the user's tone of voice and text to determine if they are excited.

[1783] 6. The server passes the question to the generative AI model, which generates the response, "The lion is currently resting. Please watch over it quietly!"

[1784] 7. The device converts the response into audio and transmits it to the child through the speaker.

[1785] In this way, the system of the present invention improves the user experience at zoos and aquariums, allowing for quick and appropriate responses based on the animals' current status and the user's emotions.

[1786] The processing flow will be explained below.

[1787] Step 1:

[1788] server

[1789] Collects breeding data. The server receives data entered by keepers and data sent from sensors, and stores the animal's health status (e.g., amount of food eaten, amount of exercise, hours of sleep, etc.) in a database such as MySQL.

[1790] Specific operation: Keepers enter data such as food intake and exercise data via a web form and store it in a database using a PHP script. Data automatically sent by sensors is also received via a REST API and stored in the database in the same way.

[1791] Step 2:

[1792] Terminal

[1793] Obtaining images of animals. The terminal obtains real-time images of animals using IP cameras installed in cages, etc.

[1794] Specific operation: The terminal receives the video stream from the IP camera using the RTSP protocol and processes the video data in real time using the OpenCV library.

[1795] Step 3:

[1796] Terminal

[1797] Analyzing the video to determine the animal's state: The device analyzes the acquired video data to determine the animal's behavior and state (e.g., active, resting, sleeping).

[1798] How it works: The device loads a deep learning model (such as YOLO or PoseNet) and analyzes the animal's pose and movement for each video frame. For example, it uses the YOLO model to detect the animal's outline and PoseNet to estimate its pose.

[1799] Step 4:

[1800] User

[1801] Users speak questions into the device by voice. Users ask specific questions into the microphone about the animal's condition and activity.

[1802] Specific action: The user speaks into the microphone, saying something like, "What is this penguin doing right now?"

[1803] Step 5:

[1804] Terminal

[1805] Converts voice to text. The device converts the acquired voice data into text data using the Google Speech-to-Text API.

[1806] What it does: The device records audio data collected by the microphone in WAV format and sends the data to the Google Speech-to-Text API to convert it into text.

[1807] Step 6:

[1808] Terminal

[1809] Send the text and animal status data to the server. The converted text data and the determined animal status information are sent to the server using an HTTP POST request.

[1810] Specific operation: The device compiles the user's question text and the analyzed animal's condition data into JSON format and sends it to the server as an HTTP POST request.

[1811] Step 7:

[1812] server

[1813] Recognizing user emotions: The server uses a sentiment analysis engine to determine the user's emotions (e.g., excitement, surprise, joy, anger, etc.) from the received text data.

[1814] Specific operation: The server passes the received question text to a sentiment analysis library (e.g., Azure Emotion API or NLTK library) to obtain the user's sentiment.

[1815] Step 8:

[1816] server

[1817] Generate a response: The server uses a generative AI model to generate an appropriate response based on the user's question text, the animal's condition data, and the results of sentiment analysis.

[1818] Specific operation: The server inputs the question text, animal state data, and user emotion data into a generative AI model (e.g., GPT-3), and generates a response such as "The lion is currently resting. Please watch over it quietly."

[1819] Step 9:

[1820] server

[1821] Send the generated response to the terminal. Send the generated text response back to the terminal as an HTTP response.

[1822] Specific operation: The server compiles the generated response text into JSON format and sends it to the terminal using an HTTP response.

[1823] Step 10:

[1824] Terminal

[1825] Synthesize the response: The device converts the received text response into speech using speech synthesis technology such as the Google Text-to-Speech API or Amazon Polly.

[1826] Specific operation: The device passes the received response text to the Google Text-to-Speech API to generate audio data and saves it as an audio file.

[1827] Step 11:

[1828] Terminal

[1829] Providing a voice response to the user: The terminal plays the generated voice data through a speaker to provide the user with an appropriate answer.

[1830] Specific operation: The device uses the playback function (for example, the playback function of the Pygame library) to play the generated audio file from the speaker.

[1831] For example, if a user asks the device, "What is this lion doing now?", the system analyzes the lion's state in real time, the server generates a response such as, "The lion is currently resting. Please watch over it quietly!", and the device plays back the response aloud. In this way, the user can obtain highly accurate information based on the animal's latest state and their own emotions in real time.

[1832] Example 2

[1833] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1834] Conventional animal monitoring and user information provision systems in zoos and aquariums are required not only to grasp the status of animals in real time, but also to respond quickly and appropriately to user questions. However, it is technically difficult to not only monitor animal behavior but also to generate responses based on the user's emotions. Furthermore, a comprehensive system is required, as advanced processing is required to respond appropriately to the diverse questions and emotions of users.

[1835] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1836] In this invention, the server includes means for collecting and saving breeding information, means for acquiring video of animals in real time, means for analyzing the acquired video to determine the behavior of the animals, means for receiving questions from users as voice input, means for converting the voice input into text, means for analyzing the user's emotions from the voice input, means for generating a response using the converted text, the determined animal status data, and the analyzed emotion data, and means for providing the generated response to the user by voice, thereby enabling the server to quickly provide an appropriate response based on both the animal's behavior and the user's emotions.

[1837] "Breeding information" refers to data about the animal's health, behavioral patterns, diet, amount of exercise, etc.

[1838] "Footage" refers to real-time video data of animals captured using a camera.

[1839] "Behavior" refers to the type of movement or activity an animal is engaged in at a particular time.

[1840] "Voice input" refers to voice data that a user utters to the system through a voice acquisition device such as a microphone.

[1841] "Text" refers to data that has been converted from voice data into text using voice recognition technology.

[1842] "Emotion" refers to the user's psychological state (e.g., excitement, calmness, joy) based on an analysis of voice tone and text content.

[1843] "Response" refers to the answer data generated by the system in response to a user's question.

[1844] A "generative AI model" refers to an algorithm that uses machine learning and artificial intelligence techniques to generate appropriate responses based on user questions and other data.

[1845] This invention relates to a system for monitoring the status of animals in real time at zoos and aquariums, and for responding to user questions promptly and appropriately. This system is mainly composed of a server, terminals, users, and a sentiment analysis engine.

[1846] server

[1847] The server is primarily responsible for collecting and storing animal care information, including the animals' health, behavior, food intake, and exercise. This data is collected periodically from keepers and sensors and stored in a database such as MySQL. The server also uses generative AI models to generate appropriate responses to user questions.

[1848] Technology used

[1849] Database Management System (DBMS): MySQL

[1850] Generative AI models: Machine learning and artificial intelligence techniques

[1851] Terminal

[1852] The device captures real-time video footage using a camera installed in the animal cage and analyzes it to determine the animal's behavior. Deep learning models such as OpenCV, YOLO, and PoseNet are used for the analysis. It also receives voice questions from users and converts them into text using speech recognition technology. The acquired data is sent to a server, and the responses received from the server are provided to the user via voice synthesis.

[1853] Technology used

[1854] Video analysis: OpenCV, YOLO, PoseNet

[1855] Speech Recognition: Google Speech-to-Text API

[1856] Speech synthesis: Google Text-to-Speech API, Amazon Polly

[1857] User

[1858] Users can ask the system questions about the animal's condition and behavior by voice. For example, a child might ask the device, "What is this lion doing right now?" This voice data is captured by the device and converted into text using voice recognition technology.

[1859] Sentiment Analysis Engine

[1860] The emotion analysis engine analyzes the user's voice and text data to recognize their emotions, allowing it to determine whether they are excited, calm, or happy.

[1861] Technology used

[1862] Sentiment analysis: Natural language processing techniques, pitch analysis, text analysis

[1863] Specific examples

[1864] For example, imagine the following scenario at a zoo one day. A database connected to a server stores daily updated data on the lion's diet, exercise volume, and sleep duration. A device captures real-time footage using a camera installed in the lion's cage and uses OpenCV and the YOLO model to determine that the lion is currently resting. A child asks into the device's microphone, "What is this lion doing right now?" The device converts the voice into text using the Google Speech-to-Text API. An emotion analysis engine determines from the tone of the voice that the child is excited. The server uses a generative AI model to generate a response based on the prompt: "The lion is currently resting. Please watch over him quietly!" The device converts this response into speech using the Google Text-to-Speech API and relays it to the child through the speaker.

[1865] Prompt Sentence Examples

[1866] The user is asking about the current status of the lion. The lion is currently resting and the user is excited.

[1867] This allows the system to quickly provide appropriate responses based on the animal's condition and the user's emotions in zoos and aquariums.

[1868] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1869] Step 1: Collect and store breeding information

[1870] server

[1871] The server collects animal care data. Specifically, keepers enter data such as food intake and exercise volume into an input form, and this data is sent to the server as an HTTP request. In addition, sensors (e.g., vital sign monitors, activity meters) send animal health status data to the server in real time. The server receives this data and inserts it into the corresponding tables in the MySQL database.

[1872] Input: Keeper input data on food intake and exercise, and health status data from sensors

[1873] Data processing: Database insert processing, data consistency check

[1874] Output: Breeding information stored in the database

[1875] Step 2: Video acquisition and behavioral analysis

[1876] Terminal

[1877] The device captures real-time video using a camera installed in the animal cage. The video data is processed frame by frame and the animal's behavior is analyzed using deep learning models such as OpenCV, YOLO, and PoseNet. The behavioral results are sent to the server in JSON format.

[1878] Input: Real-time video captured from a camera

[1879] Data processing: video frame analysis, animal behavior recognition

[1880] Output: Action result (e.g., lion is resting)

[1881] Step 3: Capture voice questions and convert them to text

[1882] User

[1883] Users can speak into the device's microphone to ask questions about animals, such as, "What is this lion doing right now?"

[1884] Terminal

[1885] The device captures the user's voice with a microphone, stores the audio data in a buffer, and converts the captured audio data into text in real time using the Google Speech-to-Text API. The converted text data is then sent to the next step.

[1886] Input: User voice question

[1887] Data processing: converting voice data to text

[1888] Output: Question text (e.g., "What is this lion doing right now?")

[1889] Step 4: Sentiment Analysis

[1890] Sentiment Analysis Engine

[1891] The sentiment analysis engine analyzes the user's emotions based on voice and text data, determining the user's emotional state from the voice tone, speed, and text content, and sending the results to the response generation step.

[1892] Input: Text data of voice questions, voice tone data

[1893] Data processing: Natural language processing and speech tone analysis

[1894] Output: Emotion determination result (e.g., user is excited)

[1895] Step 5: Generate a response

[1896] server

[1897] The server uses a generative AI model to generate a response based on the user's question text, the animal's behavioral assessment results, and the analyzed emotion data. The server uses a prompt sentence to generate an appropriate response text.

[1898] Input: Question text, behavioral judgment result, emotion judgment result, prompt text

[1899] Data processing: Response generation using generative AI models

[1900] Output: Response text (e.g. "The lion is currently resting. Please keep quiet and watch over it!")

[1901] Step 6: Provide a voice response

[1902] Terminal

[1903] The device stores the response text received from the server in a buffer and converts it into audio data using the Text-to-Speech API, which is then provided to the user through the speaker.

[1904] Input: Response text

[1905] Data processing: Converting text data into speech

[1906] Output: Audio response (e.g. "The lion is currently resting. Please keep quiet and watch over it!" played from the speaker)

[1907] In this way, it is possible to quickly provide an appropriate response based on the animal's condition and the user's emotions.

[1908] (Application example 2)

[1909] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1910] Modern factories require systems that can monitor robots in real time to ensure they are operating efficiently and provide appropriate responses to staff questions. However, current systems struggle to not only monitor the robot's status in real time, but also to respond quickly and accurately to staff questions. To solve this problem, a system is needed that can analyze the robot's behavior and automatically generate responses that take staff emotions into account.

[1911] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1912] In this invention, the server includes a means for collecting and storing breeding data and behavior data, a means for acquiring images of animals and equipment in real time, and a means for analyzing the acquired images to determine the status of the animals and equipment. This makes it possible to monitor the behavior of robots in the factory in real time, and furthermore, to generate and provide quick and accurate responses to staff questions using a generative AI model.

[1913] "Animal and operation data" refers to information relating to the health, behavior, and operating status of animals and equipment.

[1914] "Means for acquiring images of animals or devices in real time" refers to devices that use cameras or sensors to observe the activities of animals or devices in real time and acquire the image data.

[1915] The "means for analyzing the acquired video and determining the state of the animal or device" is a system that analyzes the acquired video data and determines the current state of the animal or device based on the results.

[1916] The "means for receiving a question from a user as a voice input" is a device that acquires a voice question uttered by a user via a microphone or the like.

[1917] The "means for converting speech input to text" is a speech recognition system that converts acquired speech data into text data.

[1918] The "means for generating a response using the converted text and the determined state data of the animal or device" refers to a generative AI model that generates an appropriate response using the text converted from the voice data and the determined state data of the animal or device.

[1919] The "means for providing a generated response to a user by voice" is a system that synthesizes a generated text-based response into speech and provides it to a user by voice.

[1920] A "system for analyzing the behavior of animals and equipment" is a system that uses video analysis technology to understand the behavior of animals and equipment and determine their condition based on that information.

[1921] A "generative AI model" is an artificial intelligence model that uses machine learning technology to automatically generate responses from text.

[1922] A "prompt sentence" is an input sentence given to a generative AI model, and is a document based on which the generative AI model generates a response.

[1923] MODE FOR CARRYING OUT THE INVENTION

[1924] Overall system overview

[1925] This invention relates to a system that monitors the operation of robots and equipment in a factory in real time and responds quickly and appropriately to questions from staff. The system consists of a server, terminals, users, and a sentiment analysis engine. The system consists of the following main parts:

[1926] server

[1927] The server has the following roles:

[1928] Data collection and storage:

[1929] The server periodically collects operational and other status data from the robots and equipment in the factory and stores it in a database, including health data, operational data, etc. For example, it uses SQLAlchemy to connect to a MySQL database and stores this data in tables.

[1930] Terminal

[1931] The terminals are used by staff in the factory and have the following roles:

[1932] Video acquisition and analysis:

[1933] The device uses a camera to capture real-time video of the robot or equipment. The captured video is analyzed and the current status of the robot or equipment is determined based on the results. The analysis is performed using OpenCV and deep learning models (such as YOLO and PoseNet).

[1934] Voice Input and Recognition:

[1935] The user speaks a question into the device, and the voice is picked up through the microphone and converted into text using voice recognition technology, such as the Google Speech-to-Text API.

[1936] User

[1937] Users, especially factory staff, use the terminals as follows:

[1938] Submit a question:

[1939] The user can ask a voice question to the device, such as, "What is this robot doing right now?"

[1940] Sentiment Analysis Engine

[1941] The sentiment analysis engine has the following functions:

[1942] Emotion analysis:

[1943] The system analyzes the captured voice and text data to recognize the user's emotions, determining whether the user is excited or calm. The analysis is performed using the Google Cloud Natural Language API.

[1944] Generating and serving the response

[1945] The server uses a generative AI model to generate a response based on the user's question text and the status data of the robot and equipment. The generated response is adjusted according to the user's emotions. For example, if the user is excited, the response generated will be something like, "The robot is currently packaging. Please be careful."

[1946] Specific examples

[1947] Data collection:

[1948] The robot is monitored in real time through a camera installed on the staff's terminal, which captures images of the robot and analyzes them using OpenCV.

[1949] Voice questions:

[1950] A factory worker asks into the device's microphone, "What is this robot doing now?" The device converts the voice to text and uses the Google Speech-to-Text API to translate it.

[1951] Response generation:

[1952] Using the converted text and the analyzed robot state data, the server sends prompt sentences to the generative AI model to generate a response.

[1953] Response provided:

[1954] The generated response is converted into audio using the Google Text-to-Speech API and provided to staff via the terminal.

[1955] Prompt Sentence Examples

[1956] Question: What is this robot doing now?

[1957] Robot status: Packaging in progress

[1958] User sentiment score: 0.8

[1959] response:

[1960] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1961] Step 1:

[1962] The server collects operational and other status data from the robots and equipment in the factory and stores it in a database. The input is data from the sensors on the robots and equipment, and the output is organized and stored database entries. Specifically, it uses SQLAlchemy to connect to a MySQL database and periodically updates the operational data.

[1963] Step 2:

[1964] The device uses a camera to capture real-time video of the robot or equipment. The input is the video data captured by the camera, and the output is the current status of the robot or equipment as an analysis result. The device analyzes the captured video using OpenCV or a deep learning model (such as YOLO or PoseNet) to determine the work the robot or equipment is currently performing.

[1965] Step 3:

[1966] The user speaks a question into the terminal. The input is the user's voice question, and the output is voice data. For example, the user speaks a question such as "What is this robot doing now?" into the microphone of the terminal.

[1967] Step 4:

[1968] The device receives a voice question from the user and converts it into text. The input is the captured voice data, and the output is the converted text data. The device converts the voice data into text using the Google Speech-to-Text API.

[1969] Step 5:

[1970] The server generates a response using the converted text and the determined robot or equipment status data. The input is the user's question text and the robot or equipment status data, and the output is the generated response text. The server provides the prompt text to the generative AI model, which then generates an appropriate response. Examples of prompt text are as follows:

[1971] Question: What is this robot doing now?

[1972] Robot status: Packaging in progress

[1973] User sentiment score: 0.8

[1974] response:

[1975] Step 6:

[1976] The sentiment analysis engine analyzes the user's voice and text data to recognize emotions. The input is the user's voice data and converted text data, and the output is an emotion score. It uses the Google Cloud Natural Language API to analyze the user's emotions and provide the emotion score to the server.

[1977] Step 7:

[1978] The device synthesizes the response sent from the server and provides it to the user. The input is the response text received from the server, and the output is the response in audio format. The device uses the Google Text-to-Speech API to convert the text into audio and transmit it to the user through the speaker.

[1979] This allows for real-time monitoring of robots within the factory and appropriate responses to staff questions.

[1980] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1981] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1982] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1983] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1984] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1985] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1986] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1987] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1988] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1989] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1990] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1991] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1992] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1993] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1994] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1995] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1996] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1997] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1998] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1999] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2000] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2001] The following is further disclosed regarding the above embodiment.

[2002] (Claim 1)

[2003] A means of collecting and storing breeding data;

[2004] a means for acquiring real-time footage of the animal;

[2005] A means for analyzing the acquired video to determine the state of the animal;

[2006] a means for receiving a question from a user as voice input;

[2007] a means for converting voice input into text;

[2008] means for generating a response using the converted text and the determined animal status data;

[2009] a means for providing the generated response audibly to the user;

[2010] A system including:

[2011] (Claim 2)

[2012] 2. The system according to claim 1, wherein the means for determining the state of the animal is a system for analyzing the behavior of the animal.

[2013] (Claim 3)

[2014] 10. The system of claim 1, wherein the means for generating a response uses a generative AI model.

[2015] "Example 1"

[2016] (Claim 1)

[2017] A means of collecting and storing breeding data;

[2018] a means for acquiring real-time footage of the animal;

[2019] A means for analyzing the acquired video to determine the state of the animal;

[2020] a means for receiving a question from a user as voice input;

[2021] a means for converting voice input into text;

[2022] means for generating a response using the converted text and the determined animal status data;

[2023] a means for providing the generated response audibly to the user;

[2024] A system including:

[2025] (Claim 2)

[2026] 2. The system of claim 1, wherein the means for determining the state of the animal is a system that uses an analysis algorithm to determine the behavior of the animal.

[2027] (Claim 3)

[2028] 2. The system of claim 1, wherein the means for generating a response uses a generative model that employs natural language processing technology.

[2029] "Application Example 1"

[2030] (Claim 1)

[2031] A means of collecting and storing breeding data;

[2032] a means for acquiring real-time footage of the animal;

[2033] A means for analyzing the acquired video to determine the state of the animal;

[2034] a means for receiving a question from a user as voice input;

[2035] a means for converting voice input into text;

[2036] means for generating a response using the converted text and the determined animal status data;

[2037] a means for providing the generated response audibly to the user;

[2038] The means by which users can use the zoo or aquarium through an application installed on their smartphone;

[2039] A system including:

[2040] (Claim 2)

[2041] 2. The system according to claim 1, wherein the means for determining the state of the animal is a system for analyzing the behavior of the animal.

[2042] (Claim 3)

[2043] 10. The system of claim 1, wherein the means for generating a response uses a generative AI model.

[2044] "Example 2: Combining Emotion Engines"

[2045] (Claim 1)

[2046] A means of collecting and storing breeding information,

[2047] a means for acquiring real-time footage of the animal;

[2048] A means for analyzing the acquired video to determine the behavior of the animal;

[2049] a means for receiving a question from a user as voice input;

[2050] a means for converting voice input into text;

[2051] A means of analyzing user emotions from voice input;

[2052] means for generating a response using the converted text, the determined animal state data, and the analyzed emotion data;

[2053] a means for providing the generated response audibly to the user;

[2054] A system including:

[2055] (Claim 2)

[2056] 10. The system of claim 1, wherein the means for determining the animal's behavior comprises techniques for analyzing animal behavior patterns.

[2057] (Claim 3)

[2058] 2. The system of claim 1, wherein the means for generating a response uses a generative AI model using prompt sentences.

[2059] "Application example 2 when combining emotion engines"

[2060] (Claim 1)

[2061] A means for collecting and storing rearing data and behavior data;

[2062] a means for acquiring real-time video of the animal or device;

[2063] A means for analyzing the acquired video to determine the status of the animal or device;

[2064] a means for receiving a question from a user as voice input;

[2065] a means for converting voice input into text;

[2066] means for generating a response using the converted text and the determined animal or device status data;

[2067] a means for providing the generated response audibly to the user;

[2068] A system including:

[2069] (Claim 2)

[2070] 2. The system according to claim 1, wherein the means for determining the state of the animal or device is a system for analyzing the behavior of the animal or device.

[2071] (Claim 3)

[2072] 2. The system of claim 1, wherein the means for generating a response is a system that uses a generative AI model to generate a prompt sentence. [Explanation of symbols]

[2073] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of collecting and storing breeding data; a means for acquiring real-time footage of the animal; A means for analyzing the acquired video to determine the state of the animal; a means for receiving a question from a user as voice input; a means for converting voice input into text; means for generating a response using the converted text and the determined animal status data; a means for providing the generated response audibly to the user; A system including:

2. 2. The system of claim 1, wherein the means for determining the state of the animal is a system for analyzing the behavior of the animal.

3. 10. The system of claim 1, wherein the means for generating a response uses a generative AI model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A