System

The system enhances animal-assisted therapy by accurately converting animal behavior and vocalizations into natural language, enabling effective communication and healing for elderly individuals.

JP2026017273APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118055
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

Smart Images

  • Figure 2026017273000001_ABST
    Figure 2026017273000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring a behavior and a cry of a dog or cat; means for transmitting the acquired data to a server; analysis means for analyzing the acquired data in the server and converting the data into a natural language; means for notifying a user of a result of the conversion into the private language; and means for converting a response of the user into a signal understandable by the dog or cat and transmitting the signal to the dog or cat.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, elderly people face problems such as loneliness and the progression of dementia. Animal-assisted therapy involving interaction with dogs and cats is considered effective in addressing these issues. However, current animal-assisted therapy often lacks sufficient communication with animals, preventing elderly people from deriving sufficient psychological benefits. Therefore, the present invention aims to provide a system that enables elderly people to converse with dogs and cats with high accuracy, thereby achieving physical and mental healing and improving the symptoms of dementia. [Means for solving the problem]

[0005] The present invention provides a system including: means for acquiring the behavior and meows of dogs and cats; means for transmitting the acquired data to a server; analysis means in the server for analyzing the acquired data and converting it into natural language; means for notifying a user of the results of the conversion into natural language; and means for converting the user's responses into signals understandable by dogs and cats and communicating them to the dogs and cats. The analysis means uses voice recognition technology and a behavior analysis algorithm to accurately understand the intentions and requests of dogs and cats. The data acquisition means comprises a camera and microphone mounted on the user's device. This allows for natural and smooth communication between elderly people and dogs and cats, providing physical and mental healing for the elderly and potentially improving the symptoms of dementia.

[0006] "Behavior" refers to the physical movements and reactions exhibited by dogs and cats, including walking, jumping, wagging their tails, etc.

[0007] "Meow" refers to the sounds made by dogs and cats, including barking, meowing, and growling.

[0008] "Means of acquisition" refers to devices and technologies that detect the behavior and meows of dogs and cats and collect them as data.

[0009] A "server" is a central processing unit that processes and analyzes data over a network, and its role is to analyze the behavior and cries of dogs and cats.

[0010] "Analysis means" refers to the technology or algorithms used to analyze acquired data, interpret its content, and convert it into human language.

[0011] "Natural language" refers to the language used by humans on a daily basis, and refers to the result of converting the behavior and cries of dogs and cats into a form that humans can understand.

[0012] "Notification means" refers to devices or technologies for communicating the analysis results to the user, including displays and audio output devices.

[0013] "Responding" refers to the user's reaction to the dog or cat's requests or actions, including vocal and behavioral responses.

[0014] A "signal" is an audio or visual instruction that a dog or cat can understand, and is conveyed to the dog or cat as a result of converting the user's response.

[0015] A "terminal" is a device operated by a user, including smartphones and dedicated hardware. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention is an animal therapy system that allows elderly people to converse with dogs and cats with high accuracy. The program processing of this system is explained in natural language below. Specific examples are also provided for further details.

[0038] Overall system configuration

[0039] The system consists of the following components:

[0040] 1. Device: A device operated by a user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[0041] 2. Server: The central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[0042] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0043] Program processing

[0044] Initial settings and information entry

[0045] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[0046] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0047] Voice and behavioral data acquisition

[0048] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[0049] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0050] Data analysis and transformation

[0051] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's calls and behavioral data are converted into natural language and text and voice data are generated to be communicated to the user.

[0052] Example: The server interprets Pochi's bark "woof" as meaning "I want to play," and the generative AI model converts it into natural language.

[0053] User Feedback

[0054] The device receives the analysis results and notifies the user visually and audibly, by displaying a message on the screen and announcing it using a text-to-speech function.

[0055] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0056] User response input

[0057] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[0058] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0059] Speech Transcription and Feedback

[0060] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[0061] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0062] Effects and Applications

[0063] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[0064] The processing flow will be explained below.

[0065] Step 1: Initial setup and information entry

[0066] The user installs and launches the application. After launching the application, they enter basic information about their dog or cat (name, age, sex, breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[0067] Step 2: Acquire audio and behavioral data

[0068] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[0069] Step 3: Send data

[0070] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[0071] Step 4: Receiving Data

[0072] The server receives the data sent from the terminal and waits for analysis.

[0073] Step 5: Data analysis

[0074] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[0075] Step 6: Change your mindset

[0076] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[0077] Step 7: User feedback

[0078] The device receives the analysis results sent from the server. The device notifies the user of the analysis results by text and voice. A message is displayed on the screen and voice is played from the speaker.

[0079] Step 8: User response input

[0080] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[0081] Step 9: Analyze user voice

[0082] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[0083] Step 10: Generate the signal

[0084] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0085] Step 11: Feedback to your dog or cat

[0086] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[0087] Step 12: Iterating

[0088] By repeating this series of processes, natural and smooth communication is maintained between the user and the dog or cat.

[0089] Example 1

[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0091] It is known that communicating with animals has a positive effect on the physical and mental health of elderly people. However, conventional technology has difficulty accurately understanding animal behavior and sounds and converting them into natural language. Furthermore, technology for converting user responses into a form that animals can understand has not yet been fully developed. This has made it difficult for elderly people to deepen their communication with animals.

[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0093] In this invention, the server includes a means for acquiring the behavior and sounds of animals, a means for transmitting the acquired data to the server, and an analysis means for analyzing the acquired data in the server and converting it into natural language, thereby enabling elderly people to communicate smoothly with animals in real time.

[0094] The "means for acquiring animal behavior and sounds" refers to a device that collects animal movements and sounds using a sensor or input device.

[0095] The "means for transmitting the acquired data to the server" is a device that transfers the collected data as a digital signal to the server via a network.

[0096] The "analysis means for analyzing the acquired data in the server and converting it into natural language" is a device that includes algorithms and programs for analyzing the collected data and converting it into language that the user can understand.

[0097] The "means for notifying the user of the converted results" is a device that has the function of visually or audibly conveying the analysis results to the user.

[0098] "Means for converting the user's response into a signal that the animal can understand and transmitting it to the animal" refers to a device that converts the user's instructions and responses into sounds or signals that the animal can recognize and transmits them to the animal.

[0099] "Voice recognition technology" is a technology that analyzes input voice data and converts it into characters or commands.

[0100] A "behavioral analysis algorithm" is a computational method for analyzing animal behavioral data and inferring their intentions and motivations.

[0101] "Filming equipment" refers to cameras or video equipment used to record animal behavior.

[0102] The "voice input device" is a microphone for collecting animal sounds and the user's voice.

[0103] MODE FOR CARRYING OUT THE INVENTION

[0104] The animal therapy system of the present invention is designed to assist elderly people in advanced communication with animals, and each component of the system is configured as follows.

[0105] Overall system configuration

[0106] The system consists of the following main components:

[0107] 1. Device: A device operated by the user (elderly person), such as a smartphone or dedicated hardware device. This device is equipped with a camera, microphone, display, and speaker.

[0108] 2. Server: This is the central analysis unit where the generative AI model runs. The server analyzes animal behavior and vocalization data and converts it into natural language.

[0109] 3. Network: The infrastructure that connects terminals and servers and enables data communication.

[0110] Initial settings and information entry

[0111] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, species) on the initial setup screen, and the system then sets the optimal analysis parameters for the animal.

[0112] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0113] Voice and behavioral data acquisition

[0114] The device activates the camera and microphone to record and record animal sounds and behavior in real time, and this data is saved as audio and video files.

[0115] Example: The device records and films Pochi's barking and tail wagging, and saves the data.

[0116] Data transmission

[0117] The device divides the recorded data into data packets and sends them to the server over the network, a process that allows the server to receive the data in real time.

[0118] Example: The device splits Pochi's audio and video files into small packets and sends them to a server via Wi-Fi.

[0119] Data Analysis and Natural Language Translation

[0120] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[0121] Example: The server analyzes Pochi's bark "woof" and recognizes it as an intention to "want to play." The generative AI model then converts it into the message "Pochi wants to play."

[0122] User Feedback

[0123] The device receives the analysis results and notifies the user visually and audibly. A message is displayed on the screen and the results are simultaneously announced using a text-to-speech function.

[0124] Example: The device display will show "Pochi wants to play" and at the same time it will read out loud "Pochi wants to play."

[0125] User response input

[0126] The user responds to the animal by speaking into the device, and the voice is recorded and sent to the server.

[0127] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[0128] Response transcription and feedback

[0129] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand. The converted data is then sent to the device, which then plays it back and communicates it to the animals.

[0130] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0131] Example prompts for generative AI models

[0132] "Please convert the dog's bark 'woof' into natural language."

[0133] "Convert the user's response, 'Stay tuned, I'll bring you a toy,' into a signal that the dog can understand."

[0134] This system enables advanced communication between users and animals, allowing elderly people to experience physical and mental healing through contact with animals. It is also expected to improve the symptoms of dementia.

[0135] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0136] Step 1: Initial setup and information entry

[0137] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, and species) on the initial setup screen.

[0138] Input: Animal name, age, sex, species.

[0139] Data processing: The input data is applied to the system settings to generate analysis parameters specific to the animal.

[0140] Output: The basic information of the animal is saved in the system.

[0141] How it works: After the user downloads and launches the app on their smartphone, they enter details about the animal into the form that appears on the screen.

[0142] Step 2: Acquire voice and behavioral data

[0143] The device activates the camera and microphone to record and record animal sounds and behavior in real time, and this data is saved as audio and video files.

[0144] Input: Animal sounds and behavior.

[0145] Data processing: Convert audio into an audio file and save video as a video file.

[0146] Output: Audio and video files.

[0147] What it does: When Pochi barks in the living room, the device's microphone picks up the sound and the camera records its movements.

[0148] Step 3: Send data

[0149] The device divides the recorded data into data packets and sends them to the server over the network, allowing the server to receive the data in real time.

[0150] Input: Audio and video files.

[0151] Data processing: Splitting audio and video files into data packets.

[0152] Output: The data packet is sent to the server.

[0153] Specific operation: The device splits Pochi's audio and video files into small packets and sends them to the server via Wi-Fi.

[0154] Step 4: Data analysis and natural language translation

[0155] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[0156] Input: Data packets (audio and video files).

[0157] Data processing: Analysis of audio and video data, and conversion of analysis results into natural language.

[0158] Output: A natural language message.

[0159] Specific operation: The server analyzes Pochi's bark "woof", recognizes it as an intention to "want to play", and converts this into a message in natural language saying "Pochi wants to play."

[0160] Step 5: User feedback

[0161] The device receives the analysis results and notifies the user visually and audibly, with a message displayed on the screen and a voice readout announcing the results.

[0162] Input: Natural language message (analysis result).

[0163] Data processing: Converting natural language messages into display and voice messages.

[0164] Output: Visual and audio notification.

[0165] Specific operation: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0166] Step 6: User response

[0167] The user responds to the animal by speaking into the device, which records the voice and sends it to the server.

[0168] Input: The user's spoken response.

[0169] Data processing: Convert the user's voice into an audio file.

[0170] Output: The audio file is sent to the server.

[0171] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[0172] Step 7: Response transcription and feedback

[0173] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand. The converted data is then sent to the device, which then plays it back and communicates it to the animals.

[0174] Input: The user's audio file.

[0175] Data processing: Converting the user's voice into signals that animals can understand.

[0176] Output: A signal that can be understood by animals is sent to the terminal and played.

[0177] Specific operation: The server converts the user's voice saying "Wait, I'll bring the toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0178] (Application example 1)

[0179] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0180] There is a need for systems that can enrich the time that elderly people spend with their pets and facilitate smooth communication. However, current systems have difficulty accurately analyzing pet behavior and sounds, and lack the means to provide appropriate feedback to users. Furthermore, technology to convert user input and responses into signals that pets can easily understand is still in development. This makes communication with pets ineffective, and does not sufficiently enhance the physical and mental healing and satisfaction of elderly people.

[0181] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0182] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and transmitting them to the dogs and cats, and means for performing initial settings based on the user's input. This makes it possible to analyze the pet's intentions with high accuracy and convey them to the user, and convert the user's responses into signals easily understandable by the pet and convey them to the pet. This improves communication between the elderly and their pets, increasing physical and mental healing and satisfaction.

[0183] "Means for capturing behavior and sounds" refers to devices that use cameras and microphones to record the behavior and sounds of dogs and cats.

[0184] The "means for transmitting acquired data to a server" is a system that transfers the captured behavioral and vocalization data to a server via the Internet or other communication network.

[0185] The "analysis means" is a combination of software and hardware that runs on a server, analyzes acquired data using voice recognition technology and behavior analysis algorithms, and converts the intentions of dogs and cats into natural language.

[0186] The "means for notifying the user" refers to a device for visually or audibly notifying the user of the analyzed results, such as a system including a display or speaker.

[0187] "Means for converting the user's response into signals that dogs and cats can understand and transmitting them to dogs and cats" is a system that analyzes the user's speech and input, converts them into voice and movement signals that dogs and cats can easily understand, and transmits them to them.

[0188] The "means for performing initial settings" refers to an interface and processing means for having the user input necessary information when using the system for the first time, and optimizing analysis parameters based on that information.

[0189] This invention provides an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following components:

[0190] Overall system configuration

[0191] 1. Device: A device operated by the user (elderly person), including a smartphone or dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[0192] 2. Server: This is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[0193] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0194] Initial settings and information entry

[0195] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen, and the system will then set up optimal analysis parameters.

[0196] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," age "5 years old," and gender "male."

[0197] Voice and behavioral data acquisition

[0198] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time. This data is then sent to the server. The video is captured using the opencv library, and the audio is recorded using the pyaudio library.

[0199] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0200] Data analysis and transformation

[0201] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's vocalizations and behavioral data are converted into natural language, using a generative AI model.

[0202] Example: The server analyzes Pochi's bark "woof" as the intention to "want to play," and the generative AI model converts it into natural language.

[0203] User Feedback

[0204] The device receives the analysis results and notifies the user visually and audibly, including by displaying a message on the screen and announcing it using a text-to-speech function.

[0205] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0206] User response input

[0207] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[0208] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0209] Speech Transcription and Feedback

[0210] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[0211] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0212] Effects and Applications

[0213] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[0214] Example prompt sentence:

[0215] "Analyze your pet's cry 'woof' and translate it into natural language to mean 'I want to play'."

[0216] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0217] Step 1: Initial setup and information entry

[0218] Specific operation: The user installs and launches the application. After launching the application, the user enters basic information about the pet (name, age, sex, and species) on the initial setup screen. This information is sent from the device to the server.

[0219] Input: Pet's name, age, sex, type

[0220] Data processing and calculation: Based on the pet information entered, the server sets the optimal analysis parameters.

[0221] Output: Initial setup completed and analysis parameters set

[0222] Step 2: Acquire audio and behavioral data

[0223] Specific operation: The device's camera and microphone are activated to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[0224] Input: Dog and cat meows and behavior

[0225] Data processing and data calculation: Capture video with the camera (using opencv) and record audio with the microphone (using pyaudio). Divide this data into packets and send them to the server.

[0226] Output: Audio and video data sent to the server

[0227] Step 3: Data analysis and transformation

[0228] Specific operation: The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The analyzed results are converted into natural language.

[0229] Input: Audio and video data

[0230] Data processing and data calculation: Generative AI models are used to analyze vocalizations and behavioral patterns and generate prompts that are translated into natural language.

[0231] Output: Translated natural language text

[0232] Step 4: User feedback

[0233] Specific operation: The device notifies the user of the analysis results received from the server visually and audibly, showing a message on the display and announcing it using the text-to-speech function.

[0234] Input: Natural language text sent from the server

[0235] Data processing and data calculation: The text data is displayed on the screen and simultaneously read aloud to the user as audio data.

[0236] Output: User notification (display message and voice announcement)

[0237] Step 5: User response input

[0238] Specific operation: The user responds to the dog or cat by speaking into the device, and the recorded voice is sent to the server.

[0239] Input: User's voice response

[0240] Data processing and data calculation: The device records the user's voice and sends the data to the server.

[0241] Output: Audio data sent to the server

[0242] Step 6: Voice conversion and feedback

[0243] How it works: The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is sent to the device, which then plays back the signals and transmits them to the dog or cat.

[0244] Input: Voice data sent by the user

[0245] Data Processing and Data Computation: Using a generative AI model, we convert the user's voice into an audio signal that is easy for the pet to understand.

[0246] Output: Audio signal to communicate with your pet

[0247] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0248] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it realizes interactions based on the user's psychological state. The program processing of this system is explained in natural language and detailed with concrete examples.

[0249] Overall system configuration

[0250] The system consists of the following components:

[0251] 1. Device: A device operated by the user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, speaker, and emotion engine.

[0252] 2. Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[0253] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0254] Program processing

[0255] Initial settings and information entry

[0256] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[0257] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0258] Voice and behavioral data acquisition

[0259] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[0260] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0261] Data transmission

[0262] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[0263] Data reception

[0264] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[0265] Data analysis

[0266] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[0267] Emotion analysis

[0268] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[0269] Example: The emotion engine detects when a user is smiling, and the system provides positive feedback to the user.

[0270] Change of mind

[0271] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[0272] User Feedback

[0273] The device receives the analysis results sent from the server and the emotion engine. The device notifies the user of the analysis results by text and voice. A message is displayed on the display and voice is played from the speaker.

[0274] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0275] User response input

[0276] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[0277] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0278] User voice analysis

[0279] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[0280] Signal Generation

[0281] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0282] Feedback for dogs and cats

[0283] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[0284] Effects and Applications

[0285] This system allows for advanced communication between the user and the dog or cat. By taking into account the user's emotional state, it is possible to provide more effective animal therapy, providing physical and mental healing for the elderly and improving the symptoms of dementia.

[0286] The processing flow will be explained below.

[0287] Step 1: Initial setup and information entry

[0288] The user installs and launches the application. On the application's initial setup screen, they enter basic information about their dog or cat (name, age, sex, breed) to complete the setup.

[0289] Example: A user launches the app, enters the dog's name "Pochi," its age "5 years old," and its gender "male," and completes the initial setup.

[0290] Step 2: Acquire audio and behavioral data

[0291] The device will activate the camera and microphone to record and record the sounds and behavior of your dog or cat in real time.

[0292] Example: The device uses the camera and microphone to record Pochi's barking and tail wagging.

[0293] Step 3: Obtaining emotion data

[0294] The device captures the user's facial expressions and voice using a camera and microphone, and requests the emotion engine to analyze them, thereby recognizing the user's emotional state in real time.

[0295] Example: A camera and microphone capture the user's facial expressions and voice, and the emotion engine analyzes that the user is smiling.

[0296] Step 4: Send data

[0297] The terminal transmits the acquired dog or cat voice data, behavioral data, and user emotion data to the server.

[0298] Example: Pochi's barking, behavioral data, and the user's smile data are divided into packets and sent to the server.

[0299] Step 5: Receiving Data

[0300] The server receives the data from the terminal and waits for analysis.

[0301] Example: A server receives dog and cat meow data, behavior data, and user emotion data.

[0302] Step 6: Data analysis

[0303] The server converts the voice data into text using speech recognition technology, analyzes the behavioral data using a behavioral analysis algorithm, and integrates the analysis results into a generative AI model.

[0304] Example: The server converts Pochi's bark "woof" into text and analyzes its tail-wagging behavior to recognize that it wants to play.

[0305] Step 7: Natural Language Translation

[0306] The server uses the generative AI model to convert the dog or cat's intentions into natural language and generate text and voice data to communicate with the user.

[0307] Example: The server translates a dog's barking and tail wagging into "Pochi wants to play."

[0308] Step 8: Reflecting on your emotional state

[0309] The server adjusts the messages and voices it outputs based on the user's emotional data. If the user is in a positive emotional state, it adds reassuring feedback.

[0310] Example: Because the user is smiling, the server generates positive feedback such as "Pochi wants to play. He looks very happy."

[0311] Step 9: User Feedback

[0312] The device notifies the user of the analysis results received from the server in text and audio, showing a message on the display and playing audio through the speaker.

[0313] Example: The device displays "Pochi wants to play. He looks very happy" on the display and reads it out loud.

[0314] Step 10: User response input

[0315] The user responds to the dog or cat. The device's microphone records the user's voice and sends it to the server.

[0316] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0317] Step 11: Analyze user voice

[0318] The server receives the user's voice data and converts it into text using speech analysis technology, which is then input into a generative AI model that converts it into signals that dogs and cats can understand.

[0319] Example: The server analyzes the user's voice saying "Wait, I'll bring you a toy" and converts it into signals that a dog can understand.

[0320] Step 12: Generate and transmit a signal

[0321] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0322] Example: The server generates a signal and sends it to the terminal, which then plays it back.

[0323] Step 13: Feedback to your dog or cat

[0324] The device transmits signals from the server to the dog or cat as audio and visual instructions.

[0325] Example: The device plays an "attention-grabbing audio signal" and Pochi responds to that signal.

[0326] Step 14: Iterating

[0327] By repeating a series of processes, the system maintains natural and smooth communication between the user and the dog or cat.

[0328] Example 2

[0329] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0330] Conventional animal therapy systems have limited means for elderly people to have smooth and sophisticated communication with dogs and cats, and in particular, they do not adequately consider the user's emotional state. The effectiveness of animal therapy is limited due to the lack of responses and feedback according to the elderly person's psychological state. There is a need to solve this issue and provide a system that enhances psychological and emotional support for the elderly.

[0331] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0332] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, and means for recognizing the user's emotions and providing interaction based on the user's psychological state. This enables elderly people to have advanced communication with dogs and cats, and provides effective animal therapy based on the user's emotional state.

[0333] "Behavior" refers to the movements and gestures of dogs and cats.

[0334] "Meow" is the sound made by dogs and cats.

[0335] "Users" refer to the elderly and other users of the system.

[0336] "Device" refers to a device operated by a user, including a smartphone or dedicated hardware, equipped with a camera, microphone, display, speaker, and emotion engine.

[0337] The "server" is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[0338] The "camera" is a photographic device that records the behavior of dogs and cats and detects the user's facial expressions.

[0339] A "microphone" is a sound collection device that records the sounds of dogs and cats, as well as the user's voice.

[0340] An "emotion engine" is software that analyzes a user's voice and facial expression data to determine the user's emotional state.

[0341] "Analysis means" refers to the process by which the server analyzes the data it receives using voice recognition technology and behavioral analysis algorithms.

[0342] "Natural language" refers to the language that humans use on a daily basis.

[0343] "Interaction" refers to the interaction between the user and the dog or cat.

[0344] "Feedback" refers to the process of providing responses or instructions to the user or dog or cat based on the analysis results.

[0345] A "signal" is a command or response from the user that has been translated into a format that a dog or cat can understand.

[0346] MODE FOR CARRYING OUT THE INVENTION

[0347] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. This system realizes interaction based on the user's psychological state by combining an emotion engine that recognizes the user's emotions. The system of this invention uses the following hardware and software.

[0348] Hardware Configuration

[0349] Device: A device operated by a user. This can be a smartphone or dedicated hardware. A device is equipped with a camera, microphone, display, speaker, and emotion engine.

[0350] Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[0351] Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0352] Software Configuration

[0353] Emotion engine: Software that analyzes the user's voice and facial expression data to determine the user's emotional state.

[0354] Generative AI model: An AI that converts the actions and meows of dogs and cats into text and translates their intentions into natural language.

[0355] Voice recognition technology and behavior analysis algorithm: Technology that analyzes the meows and behavior of dogs and cats and processes them as text data on a server.

[0356] Overall system configuration

[0357] When a user uses the system, the system operates as follows.

[0358] 1. Initial setup: When a user uses the system for the first time, they install the application on their device and enter basic information such as the dog or cat's name, age, sex, and breed on the initial setup screen.

[0359] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0360] 2. Data collection and transmission: The device activates the camera and microphone to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[0361] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0362] 3. Data analysis: The server analyzes the received voice data and converts the barks into text using voice recognition technology. At the same time, it analyzes the video data and recognizes the behavior of the dog or cat using a behavior analysis algorithm. The emotion engine also analyzes the user's voice and facial expression data to determine their emotional state.

[0363] Example: The emotion engine analyzes the user's smile as a "positive emotion."

[0364] 4. Intention Conversion and Feedback: The server combines the voice text and behavior analysis results and inputs them into a generative AI model. This converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate to the user. The device then notifies the user of these analysis results in text and voice.

[0365] Example: The device notifies the user via text and voice, "Pochi wants to play."

[0366] 5. User response: When the user speaks to the dog or cat, the device's microphone records the user's voice and sends it to the server. The server analyzes this voice data, converts it into a format that the dog or cat can understand, and sends it to the device. The device then conveys the converted instructions to the dog or cat as audio or visual instructions.

[0367] Example: The user says, "Pochi, wait, I'll bring you a toy," and the device helps the dog understand the command.

[0368] Examples of prompt statements

[0369] "Write a program that describes a framework that analyzes a dog's barks and behavior and translates its intentions into natural language using a generative AI model."

[0370] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0371] Step 1:

[0372] Initial settings and information entry

[0373] The user installs and launches the application on their device. When the application launches, an initial setup screen appears, and the user enters basic information about the dog or cat, such as its name, age, sex, and breed. The entered information is used by the system to set optimal analysis parameters.

[0374] Input: Basic information such as the name, age, sex, and breed of your dog or cat

[0375] Output: Optimal analysis parameters are set

[0376] Specific actions: The user operates the device's touch screen and enters "Pochi," "5 years old," and "male" into the text boxes.

[0377] Step 2:

[0378] Voice and behavioral data acquisition

[0379] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and the collected data is sent to the server.

[0380] Input: Real-time dog and cat sounds and behavior

[0381] Output: Recorded and video data is collected

[0382] Specific actions: The device's camera captures Pochi's barking and tail wagging, and the microphone records the sounds.

[0383] Step 3:

[0384] Data transmission

[0385] The audio and video data collected by the device is divided into packets and sent to the server. This transmission is done in real time, so there is little delay.

[0386] Input: Collected audio and video data

[0387] Output: Data packet is sent to the server

[0388] How it works: The device splits the recording data into small packets and sends them to the server via Wi-Fi.

[0389] Step 4:

[0390] Data reception

[0391] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[0392] Input: Data packets sent from the device

[0393] Output: Data waiting to be analyzed is saved on the server

[0394] Specific operation: The server's network adapter receives the data packet and adds the data to the analysis queue.

[0395] Step 5:

[0396] Data analysis

[0397] The server analyzes the received audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[0398] Input: Received audio and video data

[0399] Output: Call data converted to text and recognized behavior data

[0400] Specific operation: The server analyzes the audio data, converts the "woof woof" sound into text as a "dog bark," and analyzes the video data to recognize the behavior of "tail wagging."

[0401] Step 6:

[0402] Emotion analysis

[0403] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[0404] Input: User's voice and facial expression data

[0405] Output: Parsed user's emotional state

[0406] Specific operation: The device's camera captures the user's smiling expression, and the emotion engine analyzes it as a "positive state."

[0407] Step 7:

[0408] Change of mind

[0409] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[0410] Input: Speech text and behavioral analysis results

[0411] Output: Natural language text and audio data to communicate to the user

[0412] Specific operation: The server's generated AI model analyzes behavioral patterns such as "barking" and "wagging its tail," generates the intention of "I want to play," and converts this into natural language.

[0413] Step 8:

[0414] User Feedback

[0415] The device receives the analysis results sent from the server and the emotion engine and notifies the user of the results. The device notifies the user of the analysis results by text and voice.

[0416] Input: Analysis results sent from the server, analysis results from the emotion engine

[0417] Output: Notification message and audio to the user

[0418] Specific operation: The device display will show "Pochi wants to play" and the speaker will play "Pochi wants to play."

[0419] Step 9:

[0420] User response input

[0421] When a user speaks to a dog or cat, the device's microphone records the user's voice and sends the audio data to the server.

[0422] Input: Voice data as the user's response

[0423] Output: The recorded audio data is sent to the server.

[0424] Specific operation: The user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the microphone records the voice.

[0425] Step 10:

[0426] User voice analysis

[0427] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the dog or cat can understand.

[0428] Input: Received user voice data

[0429] Output: Instruction data converted to text

[0430] Specific operation: The server converts the speech "Pochi, wait for me, I'll bring you a toy" into text, and the generative AI model analyzes the text and converts it into the instruction "wait."

[0431] Step 11:

[0432] Signal Generation

[0433] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0434] Input: Transformed instruction data

[0435] Output: Instruction data sent to the terminal

[0436] Specific operation: The server sends "waiting" instruction data to the terminal, and the terminal prepares to play the data.

[0437] Step 12:

[0438] Feedback for dogs and cats

[0439] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[0440] Input: Signal received from the server

[0441] Output: Instructions given to dogs and cats

[0442] Specific operation: The device's speaker will play the voice command "wait" and the dog or cat will follow the command.

[0443] (Application example 2)

[0444] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0445] When elderly people communicate with dogs or cats, it is difficult to accurately understand their pets' needs and emotions, making it difficult to provide effective animal therapy. Furthermore, because it is not possible to respond to elderly people with pets that take their psychological state into consideration, it is difficult to provide them with further psychological comfort and healing. To solve these issues, a system is needed that not only analyzes pet behavior and meows, but also analyzes the user's psychological state in real time and provides feedback based on that analysis.

[0446] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0447] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and communicating them to the dogs and cats, and means for analyzing the user's psychological state and providing feedback based on the psychological state. This enables accurate and effective communication between elderly people and their pets, improving the provision of psychological comfort and healing, and thereby enhancing the effectiveness of animal therapy.

[0448] "Behavior" refers to the general actions and behaviors of dogs and cats.

[0449] "Meow" refers to the sound made by dogs and cats.

[0450] "Data" refers to information including the behavior and sounds of dogs and cats.

[0451] "Server" refers to a centralized computer system that performs data analysis.

[0452] "Analysis means" refers to the technology and algorithms used to analyze acquired data and extract intent and emotion.

[0453] "Conversion to natural language" refers to the process of converting information extracted by analytical means into language that can be understood by humans.

[0454] The "means for notifying" refers to a means for notifying the user of the result of the conversion into natural language.

[0455] "Response" refers to the words or actions that the user responds to the dog or cat.

[0456] "Means for converting into signals" refers to technology that converts the user's response into a format that dogs and cats can understand.

[0457] "Mental state" refers to the user's psychological and emotional state.

[0458] "Means of providing feedback" refers to means for making appropriate responses or taking appropriate actions based on the analyzed results.

[0459] The present invention relates to an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following main components:

[0460] 1. Device:

[0461] A device operated by the user, including a smartphone or a robot installed in a store. The device is equipped with a camera, microphone, display, speaker, and emotion engine. The camera and microphone are used to capture data on the behavior and meows of dogs and cats, and the data is sent to a server.

[0462] 2. Server:

[0463] It is a centralized computer system that analyzes data. The server uses a generative AI model to analyze data on dog and cat behavior and meows, converting it into natural language. It also has an emotion engine that analyzes the user's psychological state and provides feedback based on this.

[0464] 3. Network:

[0465] This is a communications infrastructure that connects devices and servers and transmits data, enabling real-time analysis and minimizing time lags.

[0466] A natural language description of the program's operation

[0467] The server analyzes the behavior and meow data of the cat or dog sent from the device and converts it into natural language. It also uses an emotion engine to analyze the user's psychological state based on the user's voice and facial expression data collected from the device. To provide appropriate feedback to the user based on the analysis results, the server uses high-performance voice recognition technology (e.g., Google Speech-to-Text API) and behavior analysis algorithms (e.g., OpenCV).

[0468] The device notifies the user of the analysis results and feedback sent from the server via text and voice, with messages displayed on the display and voice played through the speaker.

[0469] For example, if an analysis of the barking behavior data of a dog named "Pochi" determines that Pochi wants to play, the device's display will show "Pochi wants to play" and a voice will read out "Pochi wants to play." This feedback allows the user to understand Pochi's request and respond appropriately.

[0470] Specific examples

[0471] As a concrete example, consider the use of the system when an elderly person visits a physical store and spends time with their pet. The elderly person visits the store with their dog, Pochi, and launches the app on their smartphone. The device uses a camera and microphone to collect Pochi's barking, jumping, and other behaviors, and sends the collected information to a server. The server analyzes this data and generates a result, such as "Pochi likes his new toy," which is sent back to the device. The device then notifies the elderly of this result via a display and speaker.

[0472] Prompt Sentence Examples

[0473] As a concrete example, here is an example of an input prompt for a generative AI model:

[0474] "Tell me what Pochi is thinking."

[0475] "Please analyze Pochi's current emotional state."

[0476] This will enable accurate and effective communication between elderly people and pets, improving psychological comfort and healing, and maximizing the effects of animal-assisted therapy.

[0477] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0478] Step 1:

[0479] Initial Setup

[0480] The user launches the application on their device and enters basic information about their pet (dog or cat), such as its name, age, sex, and breed, which the system uses to set optimal analysis parameters.

[0481] Input: Pet's name, age, sex, type

[0482] Output: Basic information about the pet registered in the system

[0483] Specific operation: The user launches the smartphone app and enters "Pochi," "5 years old," "male," and "dog" into the input form.

[0484] Step 2:

[0485] Voice and behavioral data acquisition

[0486] The device captures images of your pet's behavior with a camera and records its cries with a microphone, and this data is collected in real time and sent to a server.

[0487] Input: Camera video data, microphone audio data

[0488] Output: Captured video and audio data

[0489] Specific operation: The device's camera takes a picture of Pochi barking and the microphone records the sound.

[0490] Step 3:

[0491] Data transmission

[0492] The video and audio data collected by the device is divided into packets and sent to the server. Since the transmission is done in real time, the time lag of the information can be minimized.

[0493] Input: Video data, audio data

[0494] Output: Data packets sent to the server

[0495] Specific operation: The device divides the recording data into packets and sends them to the server via Wi-Fi or mobile communication.

[0496] Step 4:

[0497] Data analysis

[0498] The server analyzes the received data. The voice data is converted into text using voice recognition technology (e.g., Google Speech-to-Text API), and the video data is analyzed using a behavioral analysis algorithm (e.g., OpenCV).

[0499] Input: Data packet sent to the server

[0500] Output: Textualized call data, recognized behavior data

[0501] Specific operation: The server analyzes the audio data and converts it into text that Pochi is barking "woof woof." It also analyzes the video data and recognizes that Pochi is wagging its tail.

[0502] Step 5:

[0503] Emotion analysis

[0504] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their psychological state, enabling interactions based on the user's psychological state.

[0505] Input: User's voice data, facial expression data

[0506] Output: Analyzed user's mental state

[0507] Specific operation: The device collects the user's tone of voice and facial expressions (e.g., smile) using a camera and microphone, and analyzes them using an emotion engine. It then recognizes that the user is smiling.

[0508] Step 6:

[0509] Change of mind

[0510] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[0511] Input: Textualized call data, recognized behavior data

[0512] Output: Text and audio data to be communicated to the user

[0513] Specific operation: The server generates a message saying "Pochi wants to play" and sends it to the device as text and voice data.

[0514] Step 7:

[0515] User Feedback

[0516] The device receives the analysis results sent from the server and from the emotion engine, and notifies the user through the display and speaker.

[0517] Input: Text data and audio data sent from the server

[0518] Output: Messages displayed on the display, audio played from the speaker

[0519] Specific operation: The device's display will show "Pochi wants to play" and a voice notification will be played at the same time.

[0520] Step 8:

[0521] User response input

[0522] The user speaks to the pet, the voice is recorded by the microphone on the terminal, and the voice data is sent to the server.

[0523] Input: User's voice data

[0524] Output: Audio data sent to the server

[0525] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the device's microphone records the voice.

[0526] Step 9:

[0527] User voice analysis

[0528] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the pet can understand.

[0529] Input: Audio data sent to the server

[0530] Output: Instruction data that your pet can understand

[0531] Specific operation: The server converts the voice data "Pochi, wait while I bring you a toy" into text and converts that instruction into a signal for the pet.

[0532] Step 10:

[0533] Signal Generation

[0534] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0535] Input: Instruction data that your pet can understand

[0536] Output: Signal data ready for playback

[0537] Specific operation: The server sends instruction data to the terminal, and the terminal receives the signal data and prepares for playback.

[0538] Step 11:

[0539] Pet Feedback

[0540] The device receives signals from the server and transmits them to the pet as audio and visual instructions, allowing the pet to understand the elderly person's responses.

[0541] Input: Signal data ready to be played

[0542] Output: Commands (audio or visual) that your pet will understand

[0543] Specific operation: The device plays a voice command saying "Please wait" and Pochi follows the command.

[0544] This will enable accurate and effective communication between elderly people and their pets, improving psychological comfort and healing.

[0545] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0546] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0547] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0548] [Second embodiment]

[0549] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0550] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0551] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0552] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0553] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0554] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0555] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0556] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0557] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0558] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0559] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0560] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0561] The present invention is an animal therapy system that allows elderly people to converse with dogs and cats with high accuracy. The program processing of this system is explained in natural language below. Specific examples are also provided for further details.

[0562] Overall system configuration

[0563] The system consists of the following components:

[0564] 1. Device: A device operated by a user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[0565] 2. Server: The central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[0566] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0567] Program processing

[0568] Initial settings and information entry

[0569] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[0570] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0571] Voice and behavioral data acquisition

[0572] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[0573] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0574] Data analysis and transformation

[0575] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's calls and behavioral data are converted into natural language and text and voice data are generated to be communicated to the user.

[0576] Example: The server interprets Pochi's bark "woof" as meaning "I want to play," and the generative AI model converts it into natural language.

[0577] User Feedback

[0578] The device receives the analysis results and notifies the user visually and audibly, by displaying a message on the screen and announcing it using a text-to-speech function.

[0579] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0580] User response input

[0581] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[0582] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0583] Speech Transcription and Feedback

[0584] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[0585] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0586] Effects and Applications

[0587] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[0588] The processing flow will be explained below.

[0589] Step 1: Initial setup and information entry

[0590] The user installs and launches the application. After launching the application, they enter basic information about their dog or cat (name, age, sex, breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[0591] Step 2: Acquire audio and behavioral data

[0592] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[0593] Step 3: Send data

[0594] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[0595] Step 4: Receiving Data

[0596] The server receives the data sent from the terminal and waits for analysis.

[0597] Step 5: Data analysis

[0598] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[0599] Step 6: Change your mindset

[0600] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[0601] Step 7: User feedback

[0602] The device receives the analysis results sent from the server. The device notifies the user of the analysis results by text and voice. A message is displayed on the screen and voice is played from the speaker.

[0603] Step 8: User response input

[0604] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[0605] Step 9: Analyze user voice

[0606] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[0607] Step 10: Generate the signal

[0608] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0609] Step 11: Feedback to your dog or cat

[0610] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[0611] Step 12: Iterating

[0612] By repeating this series of processes, natural and smooth communication is maintained between the user and the dog or cat.

[0613] Example 1

[0614] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0615] It is known that communicating with animals has a positive effect on the physical and mental health of elderly people. However, conventional technology has difficulty accurately understanding animal behavior and sounds and converting them into natural language. Furthermore, technology for converting user responses into a form that animals can understand has not yet been fully developed. This has made it difficult for elderly people to deepen their communication with animals.

[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0617] In this invention, the server includes a means for acquiring the behavior and sounds of animals, a means for transmitting the acquired data to the server, and an analysis means for analyzing the acquired data in the server and converting it into natural language, thereby enabling elderly people to communicate smoothly with animals in real time.

[0618] The "means for acquiring animal behavior and sounds" refers to a device that collects animal movements and sounds using a sensor or input device.

[0619] The "means for transmitting the acquired data to the server" is a device that transfers the collected data as a digital signal to the server via a network.

[0620] The "analysis means for analyzing the acquired data in the server and converting it into natural language" is a device that includes algorithms and programs for analyzing the collected data and converting it into language that the user can understand.

[0621] The "means for notifying the user of the converted results" is a device that has the function of visually or audibly conveying the analysis results to the user.

[0622] "Means for converting the user's response into a signal that the animal can understand and transmitting it to the animal" refers to a device that converts the user's instructions and responses into sounds or signals that the animal can recognize and transmits them to the animal.

[0623] "Voice recognition technology" is a technology that analyzes input voice data and converts it into characters or commands.

[0624] A "behavioral analysis algorithm" is a computational method for analyzing animal behavioral data and inferring their intentions and motivations.

[0625] "Filming equipment" refers to cameras or video equipment used to record animal behavior.

[0626] The "voice input device" is a microphone for collecting animal sounds and the user's voice.

[0627] MODE FOR CARRYING OUT THE INVENTION

[0628] The animal therapy system of the present invention is designed to assist elderly people in advanced communication with animals, and each component of the system is configured as follows.

[0629] Overall system configuration

[0630] The system consists of the following main components:

[0631] 1. Device: A device operated by the user (elderly person), such as a smartphone or dedicated hardware device. This device is equipped with a camera, microphone, display, and speaker.

[0632] 2. Server: This is the central analysis unit where the generative AI model runs. The server analyzes animal behavior and vocalization data and converts it into natural language.

[0633] 3. Network: The infrastructure that connects terminals and servers and enables data communication.

[0634] Initial settings and information entry

[0635] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, species) on the initial setup screen, and the system then sets the optimal analysis parameters for the animal.

[0636] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0637] Voice and behavioral data acquisition

[0638] The device activates the camera and microphone to record animal sounds and behavior in real time, and this data is saved as audio and video files.

[0639] Example: The device records and films Pochi's barking and tail wagging, and saves the data.

[0640] Data transmission

[0641] The device divides the recorded data into data packets and sends them to the server over the network, a process that allows the server to receive the data in real time.

[0642] Example: The device splits Pochi's audio and video files into small packets and sends them to a server via Wi-Fi.

[0643] Data Analysis and Natural Language Translation

[0644] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[0645] Example: The server analyzes Pochi's bark "woof" and recognizes it as an intention to "want to play." The generative AI model then converts it into the message "Pochi wants to play."

[0646] User Feedback

[0647] The device receives the analysis results and notifies the user visually and audibly. A message is displayed on the screen and the results are simultaneously announced using a text-to-speech function.

[0648] Example: The device display will show "Pochi wants to play" and at the same time it will read out loud "Pochi wants to play."

[0649] User response input

[0650] The user responds to the animal by speaking into the device, and the voice is recorded and sent to the server.

[0651] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[0652] Response transcription and feedback

[0653] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand, then sends the converted data to the device, which plays it back and communicates it to the animals.

[0654] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0655] Example prompts for generative AI models

[0656] "Please convert the dog's bark 'woof' into natural language."

[0657] "Convert the user's response, 'Stay tuned, I'll bring you a toy,' into a signal that the dog can understand."

[0658] This system enables advanced communication between users and animals, allowing elderly people to experience physical and mental healing through contact with animals. It is also expected to improve the symptoms of dementia.

[0659] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0660] Step 1: Initial setup and information entry

[0661] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, and species) on the initial setup screen.

[0662] Input: Animal name, age, sex, species.

[0663] Data processing: The input data is applied to the system settings to generate analysis parameters specific to the animal.

[0664] Output: The basic information of the animal is saved in the system.

[0665] How it works: After the user downloads and launches the app on their smartphone, they enter details about the animal into the form that appears on the screen.

[0666] Step 2: Acquire voice and behavioral data

[0667] The device activates the camera and microphone to record and record animal sounds and behavior in real time, and this data is saved as audio and video files.

[0668] Input: Animal sounds and behavior.

[0669] Data processing: Convert audio into an audio file and save video as a video file.

[0670] Output: Audio and video files.

[0671] What it does: When Pochi barks in the living room, the device's microphone picks up the sound and the camera records its movements.

[0672] Step 3: Send data

[0673] The device divides the recorded data into data packets and sends them to the server over the network, allowing the server to receive the data in real time.

[0674] Input: Audio and video files.

[0675] Data processing: Splitting audio and video files into data packets.

[0676] Output: The data packet is sent to the server.

[0677] Specific operation: The device splits Pochi's audio and video files into small packets and sends them to the server via Wi-Fi.

[0678] Step 4: Data analysis and natural language translation

[0679] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[0680] Input: Data packets (audio and video files).

[0681] Data processing: Analysis of audio and video data, and conversion of analysis results into natural language.

[0682] Output: A natural language message.

[0683] Specific operation: The server analyzes Pochi's bark "woof", recognizes it as an intention to "want to play", and converts this into a message in natural language saying "Pochi wants to play."

[0684] Step 5: User feedback

[0685] The device receives the analysis results and notifies the user visually and audibly, with a message displayed on the screen and a voice readout announcing the results.

[0686] Input: Natural language message (analysis result).

[0687] Data processing: Converting natural language messages into display and voice messages.

[0688] Output: Visual and audio notification.

[0689] Specific operation: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0690] Step 6: User response

[0691] The user responds to the animal by speaking into the device, which records the voice and sends it to the server.

[0692] Input: The user's spoken response.

[0693] Data processing: Convert the user's voice into an audio file.

[0694] Output: The audio file is sent to the server.

[0695] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[0696] Step 7: Response transcription and feedback

[0697] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand. The converted data is then sent to the device, which then plays it back and communicates it to the animals.

[0698] Input: The user's audio file.

[0699] Data processing: Converting the user's voice into signals that animals can understand.

[0700] Output: A signal that can be understood by animals is sent to the terminal and played.

[0701] Specific operation: The server converts the user's voice saying "Wait, I'll bring the toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0702] (Application example 1)

[0703] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0704] There is a need for systems that can enrich the time that elderly people spend with their pets and facilitate smooth communication. However, current systems have difficulty accurately analyzing pet behavior and sounds, and lack the means to provide appropriate feedback to users. Furthermore, technology to convert user input and responses into signals that pets can easily understand is still in development. This makes communication with pets ineffective, and does not sufficiently enhance the physical and mental healing and satisfaction of elderly people.

[0705] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0706] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and transmitting them to the dogs and cats, and means for performing initial settings based on the user's input. This makes it possible to analyze the pet's intentions with high accuracy and convey them to the user, and convert the user's responses into signals easily understandable by the pet and convey them to the pet. This improves communication between the elderly and their pets, increasing physical and mental healing and satisfaction.

[0707] "Means for capturing behavior and sounds" refers to devices that use cameras and microphones to record the behavior and sounds of dogs and cats.

[0708] The "means for transmitting acquired data to a server" is a system that transfers the captured behavioral and vocalization data to a server via the Internet or other communication network.

[0709] The "analysis means" is a combination of software and hardware that runs on a server, analyzes acquired data using voice recognition technology and behavior analysis algorithms, and converts the intentions of dogs and cats into natural language.

[0710] The "means for notifying the user" refers to a device for visually or audibly notifying the user of the analyzed results, such as a system including a display or speaker.

[0711] "Means for converting the user's response into signals that dogs and cats can understand and transmitting them to dogs and cats" is a system that analyzes the user's speech and input, converts them into voice and movement signals that dogs and cats can easily understand, and transmits them to them.

[0712] The "means for performing initial settings" refers to an interface and processing means for having the user input necessary information when using the system for the first time, and optimizing analysis parameters based on that information.

[0713] This invention provides an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following components:

[0714] Overall system configuration

[0715] 1. Device: A device operated by the user (elderly person), including a smartphone or dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[0716] 2. Server: This is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[0717] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0718] Initial settings and information entry

[0719] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen, and the system will then set up optimal analysis parameters.

[0720] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," age "5 years old," and gender "male."

[0721] Voice and behavioral data acquisition

[0722] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time. This data is then sent to the server. The video is captured using the opencv library, and the audio is recorded using the pyaudio library.

[0723] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0724] Data analysis and transformation

[0725] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's vocalizations and behavioral data are converted into natural language, using a generative AI model.

[0726] Example: The server analyzes Pochi's bark "woof" as the intention to "want to play," and the generative AI model converts it into natural language.

[0727] User Feedback

[0728] The device receives the analysis results and notifies the user visually and audibly, including by displaying a message on the screen and announcing it using a text-to-speech function.

[0729] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0730] User response input

[0731] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[0732] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0733] Speech Transcription and Feedback

[0734] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[0735] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[0736] Effects and Applications

[0737] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[0738] Example prompt sentence:

[0739] "Analyze your pet's cry 'woof' and translate it into natural language to mean 'I want to play'."

[0740] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0741] Step 1: Initial setup and information entry

[0742] Specific operation: The user installs and launches the application. After launching the application, the user enters basic information about the pet (name, age, sex, and species) on the initial setup screen. This information is sent from the device to the server.

[0743] Input: Pet's name, age, sex, type

[0744] Data processing and calculation: Based on the pet information entered, the server sets the optimal analysis parameters.

[0745] Output: Initial setup completed and analysis parameters set

[0746] Step 2: Acquire audio and behavioral data

[0747] Specific operation: The device's camera and microphone are activated to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[0748] Input: Dog and cat meows and behavior

[0749] Data processing and data calculation: Capture video with the camera (using opencv) and record audio with the microphone (using pyaudio). Divide this data into packets and send them to the server.

[0750] Output: Audio and video data sent to the server

[0751] Step 3: Data analysis and transformation

[0752] Specific operation: The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The analyzed results are converted into natural language.

[0753] Input: Audio and video data

[0754] Data processing and data calculation: Generative AI models are used to analyze vocalizations and behavioral patterns and generate prompts that are translated into natural language.

[0755] Output: Translated natural language text

[0756] Step 4: User feedback

[0757] Specific operation: The device notifies the user of the analysis results received from the server visually and audibly, showing a message on the display and announcing it using the text-to-speech function.

[0758] Input: Natural language text sent from the server

[0759] Data processing and data calculation: The text data is displayed on the screen and simultaneously read aloud to the user as audio data.

[0760] Output: User notification (display message and voice announcement)

[0761] Step 5: User response input

[0762] Specific operation: The user responds to the dog or cat by speaking into the device, and the recorded voice is sent to the server.

[0763] Input: User's voice response

[0764] Data processing and data calculation: The device records the user's voice and sends the data to the server.

[0765] Output: Audio data sent to the server

[0766] Step 6: Voice conversion and feedback

[0767] How it works: The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is sent to the device, which then plays back the signals and transmits them to the dog or cat.

[0768] Input: Voice data sent by the user

[0769] Data Processing and Data Computation: Using a generative AI model, we convert the user's voice into an audio signal that is easy for the pet to understand.

[0770] Output: Audio signal to communicate with your pet

[0771] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0772] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it realizes interactions based on the user's psychological state. The program processing of this system is explained in natural language and detailed with concrete examples.

[0773] Overall system configuration

[0774] The system consists of the following components:

[0775] 1. Device: A device operated by the user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, speaker, and emotion engine.

[0776] 2. Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[0777] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0778] Program processing

[0779] Initial settings and information entry

[0780] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[0781] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0782] Voice and behavioral data acquisition

[0783] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[0784] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0785] Data transmission

[0786] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[0787] Data reception

[0788] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[0789] Data analysis

[0790] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[0791] Emotion analysis

[0792] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[0793] Example: The emotion engine detects when a user is smiling, and the system provides positive feedback to the user.

[0794] Change of mind

[0795] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[0796] User Feedback

[0797] The device receives the analysis results sent from the server and the emotion engine. The device notifies the user of the analysis results by text and voice. A message is displayed on the display and voice is played from the speaker.

[0798] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[0799] User response input

[0800] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[0801] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0802] User voice analysis

[0803] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[0804] Signal Generation

[0805] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0806] Feedback for dogs and cats

[0807] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[0808] Effects and Applications

[0809] This system allows for advanced communication between the user and the dog or cat. By taking into account the user's emotional state, it is possible to provide more effective animal therapy, providing physical and mental healing for the elderly and improving the symptoms of dementia.

[0810] The processing flow will be explained below.

[0811] Step 1: Initial setup and information entry

[0812] The user installs and launches the application. On the application's initial setup screen, they enter basic information about their dog or cat (name, age, sex, breed) to complete the setup.

[0813] Example: A user launches the app, enters the dog's name "Pochi," its age "5 years old," and its gender "male," and completes the initial setup.

[0814] Step 2: Acquire audio and behavioral data

[0815] The device will activate the camera and microphone to record and record the sounds and behavior of your dog or cat in real time.

[0816] Example: The device uses the camera and microphone to record Pochi's barking and tail wagging.

[0817] Step 3: Obtaining emotion data

[0818] The device captures the user's facial expressions and voice using a camera and microphone, and requests the emotion engine to analyze them, thereby recognizing the user's emotional state in real time.

[0819] Example: A camera and microphone capture the user's facial expressions and voice, and the emotion engine analyzes that the user is smiling.

[0820] Step 4: Send data

[0821] The terminal transmits the acquired dog or cat voice data, behavioral data, and user emotion data to the server.

[0822] Example: Pochi's barking, behavioral data, and the user's smile data are divided into packets and sent to the server.

[0823] Step 5: Receiving Data

[0824] The server receives the data from the terminal and waits for analysis.

[0825] Example: A server receives dog and cat meow data, behavior data, and user emotion data.

[0826] Step 6: Data analysis

[0827] The server converts the voice data into text using speech recognition technology, analyzes the behavioral data using a behavioral analysis algorithm, and integrates the analysis results into a generative AI model.

[0828] Example: The server converts Pochi's bark "woof" into text and analyzes its tail-wagging behavior to recognize that it wants to play.

[0829] Step 7: Natural Language Translation

[0830] The server uses the generative AI model to convert the dog or cat's intentions into natural language and generate text and voice data to communicate with the user.

[0831] Example: The server translates a dog's barking and tail wagging into "Pochi wants to play."

[0832] Step 8: Reflecting on your emotional state

[0833] The server adjusts the messages and voices it outputs based on the user's emotional data. If the user is in a positive emotional state, it adds reassuring feedback.

[0834] Example: Because the user is smiling, the server generates positive feedback such as "Pochi wants to play. He looks very happy."

[0835] Step 9: User Feedback

[0836] The device notifies the user of the analysis results received from the server in text and audio, showing a message on the display and playing audio through the speaker.

[0837] Example: The device displays "Pochi wants to play. He looks very happy" on the display and reads it out loud.

[0838] Step 10: User response input

[0839] The user responds to the dog or cat. The device's microphone records the user's voice and sends it to the server.

[0840] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[0841] Step 11: Analyze user voice

[0842] The server receives the user's voice data and converts it into text using speech analysis technology, which is then input into a generative AI model that converts it into signals that dogs and cats can understand.

[0843] Example: The server analyzes the user's voice saying "Wait, I'll bring you a toy" and converts it into signals that a dog can understand.

[0844] Step 12: Generate and transmit a signal

[0845] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0846] Example: The server generates a signal and sends it to the terminal, which then plays it back.

[0847] Step 13: Feedback to your dog or cat

[0848] The device transmits signals from the server to the dog or cat as audio and visual instructions.

[0849] Example: The device plays an "attention-grabbing audio signal" and Pochi responds to that signal.

[0850] Step 14: Iterating

[0851] By repeating a series of processes, the system maintains natural and smooth communication between the user and the dog or cat.

[0852] Example 2

[0853] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0854] Conventional animal therapy systems have limited means for elderly people to have smooth and sophisticated communication with dogs and cats, and in particular, they do not adequately consider the user's emotional state. The effectiveness of animal therapy is limited due to the lack of responses and feedback according to the elderly person's psychological state. There is a need to solve this issue and provide a system that enhances psychological and emotional support for the elderly.

[0855] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0856] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, and means for recognizing the user's emotions and providing interaction based on the user's psychological state. This enables elderly people to have advanced communication with dogs and cats, and provides effective animal therapy based on the user's emotional state.

[0857] "Behavior" refers to the movements and gestures of dogs and cats.

[0858] "Meow" is the sound made by dogs and cats.

[0859] "Users" refer to the elderly and other users of the system.

[0860] "Device" refers to a device operated by a user, including a smartphone or dedicated hardware, equipped with a camera, microphone, display, speaker, and emotion engine.

[0861] The "server" is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[0862] The "camera" is a photographic device that records the behavior of dogs and cats and detects the user's facial expressions.

[0863] A "microphone" is a sound collection device that records the sounds of dogs and cats, as well as the user's voice.

[0864] An "emotion engine" is software that analyzes a user's voice and facial expression data to determine the user's emotional state.

[0865] "Analysis means" refers to the process by which the server analyzes the data it receives using voice recognition technology and behavioral analysis algorithms.

[0866] "Natural language" refers to the language that humans use on a daily basis.

[0867] "Interaction" refers to the interaction between the user and the dog or cat.

[0868] "Feedback" refers to the process of providing responses or instructions to the user or dog or cat based on the analysis results.

[0869] A "signal" is a command or response from the user that has been translated into a format that a dog or cat can understand.

[0870] MODE FOR CARRYING OUT THE INVENTION

[0871] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. This system realizes interaction based on the user's psychological state by combining an emotion engine that recognizes the user's emotions. The system of this invention uses the following hardware and software.

[0872] Hardware Configuration

[0873] Device: A device operated by a user. This can be a smartphone or dedicated hardware. A device is equipped with a camera, microphone, display, speaker, and emotion engine.

[0874] Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[0875] Network: A communications infrastructure that connects terminals and servers and enables data communication.

[0876] Software Configuration

[0877] Emotion engine: Software that analyzes the user's voice and facial expression data to determine the user's emotional state.

[0878] Generative AI model: An AI that converts the actions and meows of dogs and cats into text and translates their intentions into natural language.

[0879] Voice recognition technology and behavior analysis algorithm: Technology that analyzes the meows and behavior of dogs and cats and processes them as text data on a server.

[0880] Overall system configuration

[0881] When a user uses the system, the system operates as follows.

[0882] 1. Initial setup: When a user uses the system for the first time, they install the application on their device and enter basic information such as the dog or cat's name, age, sex, and breed on the initial setup screen.

[0883] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[0884] 2. Data collection and transmission: The device activates the camera and microphone to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[0885] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[0886] 3. Data analysis: The server analyzes the received voice data and converts the barks into text using voice recognition technology. At the same time, it analyzes the video data and recognizes the behavior of the dog or cat using a behavior analysis algorithm. The emotion engine also analyzes the user's voice and facial expression data to determine their emotional state.

[0887] Example: The emotion engine analyzes the user's smile as a "positive emotion."

[0888] 4. Intention Conversion and Feedback: The server combines the voice text and behavior analysis results and inputs them into a generative AI model. This converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate to the user. The device then notifies the user of these analysis results in text and voice.

[0889] Example: The device notifies the user via text and voice, "Pochi wants to play."

[0890] 5. User response: When the user speaks to the dog or cat, the device's microphone records the user's voice and sends it to the server. The server analyzes this voice data, converts it into a format that the dog or cat can understand, and sends it to the device. The device then conveys the converted instructions to the dog or cat as audio or visual instructions.

[0891] Example: The user says, "Pochi, wait, I'll bring you a toy," and the device helps the dog understand the command.

[0892] Examples of prompt statements

[0893] "Write a program that describes a framework that analyzes a dog's barks and behavior and translates its intentions into natural language using a generative AI model."

[0894] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0895] Step 1:

[0896] Initial settings and information entry

[0897] The user installs and launches the application on their device. When the application launches, an initial setup screen appears, and the user enters basic information about the dog or cat, such as its name, age, sex, and breed. The information entered is used by the system to set optimal analysis parameters.

[0898] Input: Basic information such as the name, age, sex, and breed of your dog or cat

[0899] Output: Optimal analysis parameters are set

[0900] Specific actions: The user operates the device's touch screen and enters "Pochi," "5 years old," and "male" into the text boxes.

[0901] Step 2:

[0902] Voice and behavioral data acquisition

[0903] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and the collected data is sent to the server.

[0904] Input: Real-time dog and cat sounds and behavior

[0905] Output: Recorded and video data is collected

[0906] Specific actions: The device's camera captures Pochi's barking and tail wagging, and the microphone records the sounds.

[0907] Step 3:

[0908] Data transmission

[0909] The audio and video data collected by the device is divided into packets and sent to the server. This transmission is done in real time, so there is little delay.

[0910] Input: Collected audio and video data

[0911] Output: Data packet is sent to the server

[0912] How it works: The device splits the recording data into small packets and sends them to the server via Wi-Fi.

[0913] Step 4:

[0914] Data reception

[0915] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[0916] Input: Data packets sent from the device

[0917] Output: Data waiting to be analyzed is saved on the server

[0918] Specific operation: The server's network adapter receives the data packet and adds the data to the analysis queue.

[0919] Step 5:

[0920] Data analysis

[0921] The server analyzes the received audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[0922] Input: Received audio and video data

[0923] Output: Call data converted to text and recognized behavior data

[0924] Specific operation: The server analyzes the audio data, converts the "woof woof" sound into text as a "dog bark," and analyzes the video data to recognize the behavior of "tail wagging."

[0925] Step 6:

[0926] Emotion analysis

[0927] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[0928] Input: User's voice and facial expression data

[0929] Output: Parsed user's emotional state

[0930] Specific operation: The device's camera captures the user's smiling expression, and the emotion engine analyzes it as a "positive state."

[0931] Step 7:

[0932] Change of mind

[0933] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[0934] Input: Speech text and behavioral analysis results

[0935] Output: Natural language text and audio data to communicate to the user

[0936] Specific operation: The server's generated AI model analyzes behavioral patterns such as "barking" and "wagging its tail," generates the intention of "I want to play," and converts this into natural language.

[0937] Step 8:

[0938] User Feedback

[0939] The device receives the analysis results sent from the server and the emotion engine and notifies the user of the results. The device notifies the user of the analysis results by text and voice.

[0940] Input: Analysis results sent from the server, analysis results from the emotion engine

[0941] Output: Notification message and audio to the user

[0942] Specific operation: The device display will show "Pochi wants to play" and the speaker will play "Pochi wants to play."

[0943] Step 9:

[0944] User response input

[0945] When a user speaks to a dog or cat, the device's microphone records the user's voice and sends the audio data to the server.

[0946] Input: Voice data as the user's response

[0947] Output: The recorded audio data is sent to the server.

[0948] Specific operation: The user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the microphone records the voice.

[0949] Step 10:

[0950] User voice analysis

[0951] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the dog or cat can understand.

[0952] Input: Received user voice data

[0953] Output: Instruction data converted to text

[0954] Specific operation: The server converts the speech "Pochi, wait for me, I'll bring you a toy" into text, and the generative AI model analyzes the text and converts it into the instruction "wait."

[0955] Step 11:

[0956] Signal Generation

[0957] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[0958] Input: Transformed instruction data

[0959] Output: Instruction data sent to the terminal

[0960] Specific operation: The server sends "waiting" instruction data to the terminal, and the terminal prepares to play the data.

[0961] Step 12:

[0962] Feedback for dogs and cats

[0963] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[0964] Input: Signal received from the server

[0965] Output: Instructions given to dogs and cats

[0966] Specific operation: The device's speaker will play the voice command "wait" and the dog or cat will follow the command.

[0967] (Application example 2)

[0968] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0969] When elderly people communicate with dogs or cats, it is difficult to accurately understand their pets' needs and emotions, making it difficult to provide effective animal therapy. Furthermore, because it is not possible to respond to elderly people with pets that take their psychological state into consideration, it is difficult to provide them with further psychological comfort and healing. To solve these issues, a system is needed that not only analyzes pet behavior and meows, but also analyzes the user's psychological state in real time and provides feedback based on that analysis.

[0970] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0971] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and communicating them to the dogs and cats, and means for analyzing the user's psychological state and providing feedback based on the psychological state. This enables accurate and effective communication between elderly people and their pets, improving the provision of psychological comfort and healing, and thereby enhancing the effectiveness of animal therapy.

[0972] "Behavior" refers to the general actions and behaviors of dogs and cats.

[0973] "Meow" refers to the sound made by dogs and cats.

[0974] "Data" refers to information including the behavior and sounds of dogs and cats.

[0975] "Server" refers to a centralized computer system that performs data analysis.

[0976] "Analysis means" refers to the technology and algorithms used to analyze acquired data and extract intent and emotion.

[0977] "Conversion to natural language" refers to the process of converting information extracted by analytical means into language that can be understood by humans.

[0978] The "means for notifying" refers to a means for notifying the user of the result of the conversion into natural language.

[0979] "Response" refers to the words or actions that the user responds to the dog or cat.

[0980] "Means for converting into signals" refers to technology that converts the user's response into a format that dogs and cats can understand.

[0981] "Mental state" refers to the user's psychological and emotional state.

[0982] "Means of providing feedback" refers to means for making appropriate responses or taking appropriate actions based on the analyzed results.

[0983] The present invention relates to an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following main components:

[0984] 1. Device:

[0985] A device operated by the user, including a smartphone or a robot installed in a store. The device is equipped with a camera, microphone, display, speaker, and emotion engine. The camera and microphone are used to capture data on the behavior and meows of dogs and cats, and the data is sent to a server.

[0986] 2. Server:

[0987] It is a centralized computer system that analyzes data. The server uses a generative AI model to analyze data on dog and cat behavior and meows, converting it into natural language. It also has an emotion engine that analyzes the user's psychological state and provides feedback based on this.

[0988] 3. Network:

[0989] This is a communications infrastructure that connects devices and servers and transmits data, enabling real-time analysis and minimizing time lags.

[0990] A natural language description of the program's operation

[0991] The server analyzes the behavior and meow data of the cat or dog sent from the device and converts it into natural language. It also uses an emotion engine to analyze the user's psychological state based on the user's voice and facial expression data collected from the device. To provide appropriate feedback to the user based on the analysis results, the server uses high-performance voice recognition technology (e.g., Google Speech-to-Text API) and behavior analysis algorithms (e.g., OpenCV).

[0992] The device notifies the user of the analysis results and feedback sent from the server via text and voice, with messages displayed on the display and voice played through the speaker.

[0993] For example, if an analysis of the barking behavior data of a dog named "Pochi" determines that Pochi wants to play, the device's display will show "Pochi wants to play" and a voice will read out "Pochi wants to play." This feedback allows the user to understand Pochi's request and respond appropriately.

[0994] Specific examples

[0995] As a concrete example, consider the use of the system when an elderly person visits a physical store and spends time with their pet. The elderly person visits the store with their dog, Pochi, and launches the app on their smartphone. The device uses a camera and microphone to collect Pochi's barking, jumping, and other behaviors, and sends the collected information to a server. The server analyzes this data and generates a result, such as "Pochi likes his new toy," which is sent back to the device. The device then notifies the elderly of this result via a display and speaker.

[0996] Prompt Sentence Examples

[0997] As a concrete example, here is an example of an input prompt for a generative AI model:

[0998] "Tell me what Pochi is thinking."

[0999] "Please analyze Pochi's current emotional state."

[1000] This will enable accurate and effective communication between elderly people and pets, improving psychological comfort and healing, and maximizing the effects of animal-assisted therapy.

[1001] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1002] Step 1:

[1003] Initial Setup

[1004] The user launches the application on their device and enters basic information about their pet (dog or cat), such as its name, age, sex, and breed, which the system uses to set optimal analysis parameters.

[1005] Input: Pet's name, age, sex, type

[1006] Output: Basic information about the pet registered in the system

[1007] Specific operation: The user launches the smartphone app and enters "Pochi," "5 years old," "male," and "dog" into the input form.

[1008] Step 2:

[1009] Voice and behavioral data acquisition

[1010] The device captures images of your pet's behavior with a camera and records its cries with a microphone, and this data is collected in real time and sent to a server.

[1011] Input: Camera video data, microphone audio data

[1012] Output: Captured video and audio data

[1013] Specific operation: The device's camera takes a picture of Pochi barking and the microphone records the sound.

[1014] Step 3:

[1015] Data transmission

[1016] The video and audio data collected by the device is divided into packets and sent to the server. Since the transmission is done in real time, the time lag of the information can be minimized.

[1017] Input: Video data, audio data

[1018] Output: Data packets sent to the server

[1019] Specific operation: The device divides the recording data into packets and sends them to the server via Wi-Fi or mobile communication.

[1020] Step 4:

[1021] Data analysis

[1022] The server analyzes the received data. The voice data is converted into text using voice recognition technology (e.g., Google Speech-to-Text API), and the video data is analyzed using a behavioral analysis algorithm (e.g., OpenCV).

[1023] Input: Data packet sent to the server

[1024] Output: Textualized call data, recognized behavior data

[1025] Specific operation: The server analyzes the audio data and converts it into text that Pochi is barking "woof woof." It also analyzes the video data and recognizes that Pochi is wagging its tail.

[1026] Step 5:

[1027] Emotion analysis

[1028] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their psychological state, enabling interactions based on the user's psychological state.

[1029] Input: User's voice data, facial expression data

[1030] Output: Analyzed user's mental state

[1031] Specific operation: The device collects the user's tone of voice and facial expressions (e.g., smile) using a camera and microphone, and analyzes them using an emotion engine. It then recognizes that the user is smiling.

[1032] Step 6:

[1033] Change of mind

[1034] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1035] Input: Textualized call data, recognized behavior data

[1036] Output: Text and audio data to be communicated to the user

[1037] Specific operation: The server generates a message saying "Pochi wants to play" and sends it to the device as text and voice data.

[1038] Step 7:

[1039] User Feedback

[1040] The device receives the analysis results sent from the server and from the emotion engine, and notifies the user through the display and speaker.

[1041] Input: Text data and audio data sent from the server

[1042] Output: Messages displayed on the display, audio played from the speaker

[1043] Specific operation: The device's display will show "Pochi wants to play" and a voice notification will be played at the same time.

[1044] Step 8:

[1045] User response input

[1046] The user speaks to the pet, the voice is recorded by the microphone on the terminal, and the voice data is sent to the server.

[1047] Input: User's voice data

[1048] Output: Audio data sent to the server

[1049] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the device's microphone records the voice.

[1050] Step 9:

[1051] User voice analysis

[1052] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the pet can understand.

[1053] Input: Audio data sent to the server

[1054] Output: Instruction data that your pet can understand

[1055] Specific operation: The server converts the voice data "Pochi, wait while I bring you a toy" into text and converts that instruction into a signal for the pet.

[1056] Step 10:

[1057] Signal Generation

[1058] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1059] Input: Instruction data that your pet can understand

[1060] Output: Signal data ready for playback

[1061] Specific operation: The server sends instruction data to the terminal, and the terminal receives the signal data and prepares for playback.

[1062] Step 11:

[1063] Pet Feedback

[1064] The device receives signals from the server and transmits them to the pet as audio and visual instructions, allowing the pet to understand the elderly person's responses.

[1065] Input: Signal data ready to be played

[1066] Output: Commands (audio or visual) that your pet will understand

[1067] Specific operation: The device plays a voice command saying "Please wait" and Pochi follows the command.

[1068] This will enable accurate and effective communication between elderly people and their pets, improving psychological comfort and healing.

[1069] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1070] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1071] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1072] [Third embodiment]

[1073] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1074] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1075] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1076] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1077] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1078] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1079] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1080] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1081] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1082] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1083] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1084] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1085] The present invention is an animal therapy system that allows elderly people to converse with dogs and cats with high accuracy. The program processing of this system is explained in natural language below. Specific examples are also provided for further details.

[1086] Overall system configuration

[1087] The system consists of the following components:

[1088] 1. Device: A device operated by a user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[1089] 2. Server: The central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[1090] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1091] Program processing

[1092] Initial settings and information entry

[1093] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[1094] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1095] Voice and behavioral data acquisition

[1096] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[1097] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1098] Data analysis and transformation

[1099] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's calls and behavioral data are converted into natural language and text and voice data are generated to be communicated to the user.

[1100] Example: The server interprets Pochi's bark "woof" as meaning "I want to play," and the generative AI model converts it into natural language.

[1101] User Feedback

[1102] The device receives the analysis results and notifies the user visually and audibly, by displaying a message on the screen and announcing it using a text-to-speech function.

[1103] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1104] User response input

[1105] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[1106] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1107] Speech Transcription and Feedback

[1108] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[1109] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1110] Effects and Applications

[1111] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[1112] The processing flow will be explained below.

[1113] Step 1: Initial setup and information entry

[1114] The user installs and launches the application. After launching the application, they enter basic information about their dog or cat (name, age, sex, breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[1115] Step 2: Acquire audio and behavioral data

[1116] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[1117] Step 3: Send data

[1118] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[1119] Step 4: Receiving Data

[1120] The server receives the data sent from the terminal and waits for analysis.

[1121] Step 5: Data analysis

[1122] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[1123] Step 6: Change your mindset

[1124] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1125] Step 7: User feedback

[1126] The device receives the analysis results sent from the server. The device notifies the user of the analysis results by text and voice. A message is displayed on the screen and voice is played from the speaker.

[1127] Step 8: User response input

[1128] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[1129] Step 9: Analyze user voice

[1130] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[1131] Step 10: Generate the signal

[1132] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1133] Step 11: Feedback to your dog or cat

[1134] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[1135] Step 12: Iterating

[1136] By repeating this series of processes, natural and smooth communication is maintained between the user and the dog or cat.

[1137] Example 1

[1138] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1139] It is known that communicating with animals has a positive effect on the physical and mental health of elderly people. However, conventional technology has difficulty accurately understanding animal behavior and sounds and converting them into natural language. Furthermore, technology for converting user responses into a form that animals can understand has not yet been fully developed. This has made it difficult for elderly people to deepen their communication with animals.

[1140] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1141] In this invention, the server includes a means for acquiring the behavior and sounds of animals, a means for transmitting the acquired data to the server, and an analysis means for analyzing the acquired data in the server and converting it into natural language, thereby enabling elderly people to communicate smoothly with animals in real time.

[1142] The "means for acquiring animal behavior and sounds" refers to a device that collects animal movements and sounds using a sensor or input device.

[1143] The "means for transmitting the acquired data to the server" is a device that transfers the collected data as a digital signal to the server via a network.

[1144] The "analysis means for analyzing the acquired data in the server and converting it into natural language" is a device that includes algorithms and programs for analyzing the collected data and converting it into language that the user can understand.

[1145] The "means for notifying the user of the converted results" is a device that has the function of visually or audibly conveying the analysis results to the user.

[1146] "Means for converting the user's response into a signal that the animal can understand and transmitting it to the animal" refers to a device that converts the user's instructions and responses into sounds or signals that the animal can recognize and transmits them to the animal.

[1147] "Voice recognition technology" is a technology that analyzes input voice data and converts it into characters or commands.

[1148] A "behavioral analysis algorithm" is a computational method for analyzing animal behavioral data and inferring their intentions and motivations.

[1149] "Filming equipment" refers to cameras or video equipment used to record animal behavior.

[1150] The "voice input device" is a microphone for collecting animal sounds and the user's voice.

[1151] MODE FOR CARRYING OUT THE INVENTION

[1152] The animal therapy system of the present invention is designed to assist elderly people in advanced communication with animals, and each component of the system is configured as follows.

[1153] Overall system configuration

[1154] The system consists of the following main components:

[1155] 1. Device: A device operated by the user (elderly person), such as a smartphone or dedicated hardware device. This device is equipped with a camera, microphone, display, and speaker.

[1156] 2. Server: This is the central analysis unit where the generative AI model runs. The server analyzes animal behavior and vocalization data and converts it into natural language.

[1157] 3. Network: The infrastructure that connects terminals and servers and enables data communication.

[1158] Initial settings and information entry

[1159] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, species) on the initial setup screen, and the system then sets the optimal analysis parameters for the animal.

[1160] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1161] Voice and behavioral data acquisition

[1162] The device activates the camera and microphone to record animal sounds and behavior in real time, and this data is saved as audio and video files.

[1163] Example: The device records and films Pochi's barking and tail wagging, and saves the data.

[1164] Data transmission

[1165] The device divides the recorded data into data packets and sends them to the server over the network, a process that allows the server to receive the data in real time.

[1166] Example: The device splits Pochi's audio and video files into small packets and sends them to a server via Wi-Fi.

[1167] Data Analysis and Natural Language Translation

[1168] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[1169] Example: The server analyzes Pochi's bark "woof" and recognizes it as an intention to "want to play." The generative AI model then converts it into the message "Pochi wants to play."

[1170] User Feedback

[1171] The device receives the analysis results and notifies the user visually and audibly. A message is displayed on the screen and the results are simultaneously announced using a text-to-speech function.

[1172] Example: The device display will show "Pochi wants to play" and at the same time it will read out loud "Pochi wants to play."

[1173] User response input

[1174] The user responds to the animal by speaking into the device, and the voice is recorded and sent to the server.

[1175] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[1176] Response transcription and feedback

[1177] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand. The converted data is then sent to the device, which then plays it back and communicates it to the animals.

[1178] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1179] Example prompts for generative AI models

[1180] "Please convert the dog's bark 'woof' into natural language."

[1181] "Convert the user's response, 'Stay tuned, I'll bring you a toy,' into a signal that the dog can understand."

[1182] This system enables advanced communication between users and animals, allowing elderly people to experience physical and mental healing through contact with animals. It is also expected to improve the symptoms of dementia.

[1183] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1184] Step 1: Initial setup and information entry

[1185] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, and species) on the initial setup screen.

[1186] Input: Animal name, age, sex, species.

[1187] Data processing: The input data is applied to the system settings to generate analysis parameters specific to the animal.

[1188] Output: The basic information of the animal is saved in the system.

[1189] How it works: After the user downloads and launches the app on their smartphone, they enter details about the animal into the form that appears on the screen.

[1190] Step 2: Acquire voice and behavioral data

[1191] The device activates the camera and microphone to record and record animal sounds and behavior in real time, and this data is saved as audio and video files.

[1192] Input: Animal sounds and behavior.

[1193] Data processing: Convert audio into an audio file and save video as a video file.

[1194] Output: Audio and video files.

[1195] What it does: When Pochi barks in the living room, the device's microphone picks up the sound and the camera records its movements.

[1196] Step 3: Send data

[1197] The device divides the recorded data into data packets and sends them to the server over the network, allowing the server to receive the data in real time.

[1198] Input: Audio and video files.

[1199] Data processing: Splitting audio and video files into data packets.

[1200] Output: The data packet is sent to the server.

[1201] Specific operation: The device splits Pochi's audio and video files into small packets and sends them to the server via Wi-Fi.

[1202] Step 4: Data analysis and natural language translation

[1203] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[1204] Input: Data packets (audio and video files).

[1205] Data processing: Analysis of audio and video data, and conversion of analysis results into natural language.

[1206] Output: A natural language message.

[1207] Specific operation: The server analyzes Pochi's bark "woof", recognizes it as an intention to "want to play", and converts this into a message in natural language saying "Pochi wants to play."

[1208] Step 5: User feedback

[1209] The device receives the analysis results and notifies the user visually and audibly, with a message displayed on the screen and a voice readout announcing the results.

[1210] Input: Natural language message (analysis result).

[1211] Data processing: Converting natural language messages into display and voice messages.

[1212] Output: Visual and audio notification.

[1213] Specific operation: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1214] Step 6: User response

[1215] The user responds to the animal by speaking into the device, which records the voice and sends it to the server.

[1216] Input: The user's spoken response.

[1217] Data processing: Convert the user's voice into an audio file.

[1218] Output: The audio file is sent to the server.

[1219] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[1220] Step 7: Response transcription and feedback

[1221] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand. The converted data is then sent to the device, which then plays it back and communicates it to the animals.

[1222] Input: The user's audio file.

[1223] Data processing: Converting the user's voice into signals that animals can understand.

[1224] Output: A signal that can be understood by animals is sent to the terminal and played.

[1225] Specific operation: The server converts the user's voice saying "Wait, I'll bring the toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1226] (Application example 1)

[1227] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1228] There is a need for systems that can enrich the time that elderly people spend with their pets and facilitate smooth communication. However, current systems have difficulty accurately analyzing pet behavior and sounds, and lack the means to provide appropriate feedback to users. Furthermore, technology to convert user input and responses into signals that pets can easily understand is still in development. This makes communication with pets ineffective, and does not sufficiently enhance the physical and mental healing and satisfaction of elderly people.

[1229] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1230] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and transmitting them to the dogs and cats, and means for performing initial settings based on the user's input. This makes it possible to analyze the pet's intentions with high accuracy and convey them to the user, and convert the user's responses into signals easily understandable by the pet and convey them to the pet. This improves communication between the elderly and their pets, increasing physical and mental healing and satisfaction.

[1231] "Means for capturing behavior and sounds" refers to devices that use cameras and microphones to record the behavior and sounds of dogs and cats.

[1232] The "means for transmitting acquired data to a server" is a system that transfers the captured behavioral and vocalization data to a server via the Internet or other communication network.

[1233] The "analysis means" is a combination of software and hardware that runs on a server, analyzes acquired data using voice recognition technology and behavior analysis algorithms, and converts the intentions of dogs and cats into natural language.

[1234] The "means for notifying the user" refers to a device for visually or audibly notifying the user of the analyzed results, such as a system including a display or speaker.

[1235] "Means for converting the user's response into signals that dogs and cats can understand and transmitting them to dogs and cats" is a system that analyzes the user's speech and input, converts them into voice and movement signals that dogs and cats can easily understand, and transmits them to them.

[1236] The "means for performing initial settings" refers to an interface and processing means for having the user input necessary information when using the system for the first time, and optimizing analysis parameters based on that information.

[1237] This invention provides an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following components:

[1238] Overall system configuration

[1239] 1. Device: A device operated by the user (elderly person), including a smartphone or dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[1240] 2. Server: This is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[1241] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1242] Initial settings and information entry

[1243] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen, and the system will then set up optimal analysis parameters.

[1244] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," age "5 years old," and gender "male."

[1245] Voice and behavioral data acquisition

[1246] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time. This data is then sent to the server. The video is captured using the opencv library, and the audio is recorded using the pyaudio library.

[1247] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1248] Data analysis and transformation

[1249] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's vocalizations and behavioral data are converted into natural language, using a generative AI model.

[1250] Example: The server analyzes Pochi's bark "woof" as the intention to "want to play," and the generative AI model converts it into natural language.

[1251] User Feedback

[1252] The device receives the analysis results and notifies the user visually and audibly, including by displaying a message on the screen and announcing it using a text-to-speech function.

[1253] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1254] User response input

[1255] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[1256] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1257] Speech Transcription and Feedback

[1258] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[1259] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1260] Effects and Applications

[1261] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[1262] Example prompt sentence:

[1263] "Analyze your pet's cry 'woof' and translate it into natural language to mean 'I want to play'."

[1264] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1265] Step 1: Initial setup and information entry

[1266] Specific operation: The user installs and launches the application. After launching the application, the user enters basic information about the pet (name, age, sex, and species) on the initial setup screen. This information is sent from the device to the server.

[1267] Input: Pet's name, age, sex, type

[1268] Data processing and calculation: Based on the pet information entered, the server sets the optimal analysis parameters.

[1269] Output: Initial setup completed and analysis parameters set

[1270] Step 2: Acquire audio and behavioral data

[1271] Specific operation: The device's camera and microphone are activated to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[1272] Input: Dog and cat meows and behavior

[1273] Data processing and data calculation: Capture video with the camera (using opencv) and record audio with the microphone (using pyaudio). Divide this data into packets and send them to the server.

[1274] Output: Audio and video data sent to the server

[1275] Step 3: Data analysis and transformation

[1276] Specific operation: The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The analyzed results are converted into natural language.

[1277] Input: Audio and video data

[1278] Data processing and data calculation: Generative AI models are used to analyze vocalizations and behavioral patterns and generate prompts that are translated into natural language.

[1279] Output: Translated natural language text

[1280] Step 4: User feedback

[1281] Specific operation: The device notifies the user of the analysis results received from the server visually and audibly, showing a message on the display and announcing it using the text-to-speech function.

[1282] Input: Natural language text sent from the server

[1283] Data processing and data calculation: The text data is displayed on the screen and simultaneously read aloud to the user as audio data.

[1284] Output: User notification (display message and voice announcement)

[1285] Step 5: User response input

[1286] Specific operation: The user responds to the dog or cat by speaking into the device, and the recorded voice is sent to the server.

[1287] Input: User's voice response

[1288] Data processing and data calculation: The device records the user's voice and sends the data to the server.

[1289] Output: Audio data sent to the server

[1290] Step 6: Voice conversion and feedback

[1291] How it works: The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is sent to the device, which then plays back the signals and transmits them to the dog or cat.

[1292] Input: Voice data sent by the user

[1293] Data Processing and Data Computation: Using a generative AI model, we convert the user's voice into an audio signal that is easy for the pet to understand.

[1294] Output: Audio signal to communicate with your pet

[1295] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1296] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it realizes interactions based on the user's psychological state. The program processing of this system is explained in natural language and detailed with concrete examples.

[1297] Overall system configuration

[1298] The system consists of the following components:

[1299] 1. Device: A device operated by the user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, speaker, and emotion engine.

[1300] 2. Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[1301] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1302] Program processing

[1303] Initial settings and information entry

[1304] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[1305] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1306] Voice and behavioral data acquisition

[1307] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[1308] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1309] Data transmission

[1310] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[1311] Data reception

[1312] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[1313] Data analysis

[1314] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[1315] Emotion analysis

[1316] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[1317] Example: The emotion engine detects when a user is smiling, and the system provides positive feedback to the user.

[1318] Change of mind

[1319] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1320] User Feedback

[1321] The device receives the analysis results sent from the server and the emotion engine. The device notifies the user of the analysis results by text and voice. A message is displayed on the display and voice is played from the speaker.

[1322] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1323] User response input

[1324] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[1325] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1326] User voice analysis

[1327] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[1328] Signal Generation

[1329] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1330] Feedback for dogs and cats

[1331] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[1332] Effects and Applications

[1333] This system allows for advanced communication between the user and the dog or cat. By taking into account the user's emotional state, it is possible to provide more effective animal therapy, providing physical and mental healing for the elderly and improving the symptoms of dementia.

[1334] The processing flow will be explained below.

[1335] Step 1: Initial setup and information entry

[1336] The user installs and launches the application. On the application's initial setup screen, they enter basic information about their dog or cat (name, age, sex, breed) to complete the setup.

[1337] Example: A user launches the app, enters the dog's name "Pochi," its age "5 years old," and its gender "male," and completes the initial setup.

[1338] Step 2: Acquire audio and behavioral data

[1339] The device will activate the camera and microphone to record and record the sounds and behavior of your dog or cat in real time.

[1340] Example: The device uses the camera and microphone to record Pochi's barking and tail wagging.

[1341] Step 3: Obtaining emotion data

[1342] The device captures the user's facial expressions and voice using a camera and microphone, and requests analysis from the emotion engine, thereby recognizing the user's emotional state in real time.

[1343] Example: A camera and microphone capture the user's facial expressions and voice, and the emotion engine analyzes that the user is smiling.

[1344] Step 4: Send data

[1345] The terminal transmits the acquired dog or cat voice data, behavioral data, and user emotion data to the server.

[1346] Example: Pochi's barking, behavioral data, and the user's smile data are divided into packets and sent to the server.

[1347] Step 5: Receiving Data

[1348] The server receives the data from the terminal and waits for analysis.

[1349] Example: A server receives dog and cat meow data, behavior data, and user emotion data.

[1350] Step 6: Data analysis

[1351] The server converts the voice data into text using speech recognition technology, analyzes the behavioral data using a behavioral analysis algorithm, and integrates the analysis results into a generative AI model.

[1352] Example: The server converts Pochi's bark "woof" into text and analyzes its tail-wagging behavior to recognize that it wants to play.

[1353] Step 7: Natural Language Translation

[1354] The server uses the generative AI model to convert the dog or cat's intentions into natural language and generate text and voice data to communicate with the user.

[1355] Example: The server translates a dog's barking and tail wagging into "Pochi wants to play."

[1356] Step 8: Reflecting on your emotional state

[1357] The server adjusts the messages and voices it outputs based on the user's emotional data. If the user is in a positive emotional state, it adds reassuring feedback.

[1358] Example: Because the user is smiling, the server generates positive feedback such as "Pochi wants to play. He looks very happy."

[1359] Step 9: User Feedback

[1360] The device notifies the user of the analysis results received from the server in text and audio, showing a message on the display and playing audio through the speaker.

[1361] Example: The device displays "Pochi wants to play. He looks very happy" on the display and reads it out loud.

[1362] Step 10: User response input

[1363] The user responds to the dog or cat. The device's microphone records the user's voice and sends it to the server.

[1364] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1365] Step 11: Analyze user voice

[1366] The server receives the user's voice data and converts it into text using speech analysis technology, which is then input into a generative AI model that converts it into signals that dogs and cats can understand.

[1367] Example: The server analyzes the user's voice saying "Wait, I'll bring you a toy" and converts it into signals that a dog can understand.

[1368] Step 12: Generate and transmit a signal

[1369] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1370] Example: The server generates a signal and sends it to the terminal, which then plays it back.

[1371] Step 13: Feedback to your dog or cat

[1372] The device transmits signals from the server to the dog or cat as audio and visual instructions.

[1373] Example: The device plays an "attention-grabbing audio signal" and Pochi responds to that signal.

[1374] Step 14: Iterating

[1375] By repeating a series of processes, the system maintains natural and smooth communication between the user and the dog or cat.

[1376] Example 2

[1377] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1378] Conventional animal therapy systems have limited means for elderly people to have smooth and sophisticated communication with dogs and cats, and in particular, they do not adequately consider the user's emotional state. The effectiveness of animal therapy is limited due to the lack of responses and feedback according to the elderly person's psychological state. There is a need to solve this issue and provide a system that enhances psychological and emotional support for the elderly.

[1379] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1380] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, and means for recognizing the user's emotions and providing interaction based on the user's psychological state. This enables elderly people to have advanced communication with dogs and cats, and provides effective animal therapy based on the user's emotional state.

[1381] "Behavior" refers to the movements and gestures of dogs and cats.

[1382] "Meow" is the sound made by dogs and cats.

[1383] "Users" refer to the elderly and other users of the system.

[1384] "Device" refers to a device operated by a user, including a smartphone or dedicated hardware, equipped with a camera, microphone, display, speaker, and emotion engine.

[1385] The "server" is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[1386] The "camera" is a photographic device that records the behavior of dogs and cats and detects the user's facial expressions.

[1387] A "microphone" is a sound collection device that records the sounds of dogs and cats, as well as the user's voice.

[1388] An "emotion engine" is software that analyzes a user's voice and facial expression data to determine the user's emotional state.

[1389] "Analysis means" refers to the process by which the server analyzes the data it receives using voice recognition technology and behavioral analysis algorithms.

[1390] "Natural language" refers to the language that humans use on a daily basis.

[1391] "Interaction" refers to the interaction between the user and the dog or cat.

[1392] "Feedback" refers to the process of providing responses or instructions to the user or dog or cat based on the analysis results.

[1393] A "signal" is a command or response from the user that has been translated into a format that a dog or cat can understand.

[1394] MODE FOR CARRYING OUT THE INVENTION

[1395] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. This system realizes interaction based on the user's psychological state by combining an emotion engine that recognizes the user's emotions. The system of this invention uses the following hardware and software.

[1396] Hardware Configuration

[1397] Device: A device operated by a user. This can be a smartphone or dedicated hardware. A device is equipped with a camera, microphone, display, speaker, and emotion engine.

[1398] Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[1399] Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1400] Software Configuration

[1401] Emotion engine: Software that analyzes the user's voice and facial expression data to determine the user's emotional state.

[1402] Generative AI model: An AI that converts the actions and meows of dogs and cats into text and translates their intentions into natural language.

[1403] Voice recognition technology and behavior analysis algorithm: Technology that analyzes the meows and behavior of dogs and cats and processes them as text data on a server.

[1404] Overall system configuration

[1405] When a user uses the system, the system operates as follows.

[1406] 1. Initial setup: When a user uses the system for the first time, they install the application on their device and enter basic information such as the dog or cat's name, age, sex, and breed on the initial setup screen.

[1407] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1408] 2. Data collection and transmission: The device activates the camera and microphone to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[1409] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1410] 3. Data analysis: The server analyzes the received voice data and converts the barks into text using voice recognition technology. At the same time, it analyzes the video data and recognizes the behavior of the dog or cat using a behavior analysis algorithm. The emotion engine also analyzes the user's voice and facial expression data to determine their emotional state.

[1411] Example: The emotion engine analyzes the user's smile as a "positive emotion."

[1412] 4. Intention Conversion and Feedback: The server combines the voice text and behavior analysis results and inputs them into a generative AI model. This converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate to the user. The device then notifies the user of these analysis results in text and voice.

[1413] Example: The device notifies the user via text and voice, "Pochi wants to play."

[1414] 5. User response: When the user speaks to the dog or cat, the device's microphone records the user's voice and sends it to the server. The server analyzes this voice data, converts it into a format that the dog or cat can understand, and sends it to the device. The device then conveys the converted instructions to the dog or cat as audio or visual instructions.

[1415] Example: The user says, "Pochi, wait, I'll bring you a toy," and the device helps the dog understand the command.

[1416] Examples of prompt statements

[1417] "Write a program that describes a framework that analyzes a dog's barks and behavior and translates its intentions into natural language using a generative AI model."

[1418] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1419] Step 1:

[1420] Initial settings and information entry

[1421] The user installs and launches the application on their device. When the application launches, an initial setup screen appears, and the user enters basic information about the dog or cat, such as its name, age, sex, and breed. The information entered is used by the system to set optimal analysis parameters.

[1422] Input: Basic information such as the name, age, sex, and breed of your dog or cat

[1423] Output: Optimal analysis parameters are set

[1424] Specific actions: The user operates the device's touch screen and enters "Pochi," "5 years old," and "male" into the text boxes.

[1425] Step 2:

[1426] Voice and behavioral data acquisition

[1427] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and the collected data is sent to the server.

[1428] Input: Real-time dog and cat sounds and behavior

[1429] Output: Recorded and video data is collected

[1430] Specific actions: The device's camera captures Pochi's barking and tail wagging, and the microphone records the sounds.

[1431] Step 3:

[1432] Data transmission

[1433] The audio and video data collected by the device is divided into packets and sent to the server. This transmission is done in real time, so there is little delay.

[1434] Input: Collected audio and video data

[1435] Output: Data packet is sent to the server

[1436] How it works: The device splits the recording data into small packets and sends them to the server via Wi-Fi.

[1437] Step 4:

[1438] Data reception

[1439] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[1440] Input: Data packets sent from the device

[1441] Output: Data waiting to be analyzed is saved on the server

[1442] Specific operation: The server's network adapter receives the data packet and adds the data to the analysis queue.

[1443] Step 5:

[1444] Data analysis

[1445] The server analyzes the received audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[1446] Input: Received audio and video data

[1447] Output: Call data converted to text and recognized behavior data

[1448] Specific operation: The server analyzes the audio data, converts the "woof woof" sound into text as a "dog bark," and analyzes the video data to recognize the behavior of "tail wagging."

[1449] Step 6:

[1450] Emotion analysis

[1451] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[1452] Input: User's voice and facial expression data

[1453] Output: Parsed user's emotional state

[1454] Specific operation: The device's camera captures the user's smiling expression, and the emotion engine analyzes it as a "positive state."

[1455] Step 7:

[1456] Change of mind

[1457] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1458] Input: Speech text and behavioral analysis results

[1459] Output: Natural language text and audio data to communicate to the user

[1460] Specific operation: The server's generated AI model analyzes behavioral patterns such as "barking" and "wagging its tail," generates the intention of "I want to play," and converts this into natural language.

[1461] Step 8:

[1462] User Feedback

[1463] The device receives the analysis results sent from the server and the emotion engine and notifies the user of the results. The device notifies the user of the analysis results by text and voice.

[1464] Input: Analysis results sent from the server, analysis results from the emotion engine

[1465] Output: Notification message and audio to the user

[1466] Specific operation: The device display will show "Pochi wants to play" and the speaker will play "Pochi wants to play."

[1467] Step 9:

[1468] User response input

[1469] When a user speaks to a dog or cat, the device's microphone records the user's voice and sends the audio data to the server.

[1470] Input: Voice data as the user's response

[1471] Output: The recorded audio data is sent to the server.

[1472] Specific operation: The user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the microphone records the voice.

[1473] Step 10:

[1474] User voice analysis

[1475] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the dog or cat can understand.

[1476] Input: Received user voice data

[1477] Output: Instruction data converted to text

[1478] Specific operation: The server converts the speech "Pochi, wait for me, I'll bring you a toy" into text, and the generative AI model analyzes the text and converts it into the instruction "wait."

[1479] Step 11:

[1480] Signal Generation

[1481] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1482] Input: Transformed instruction data

[1483] Output: Instruction data sent to the terminal

[1484] Specific operation: The server sends "waiting" instruction data to the terminal, and the terminal prepares to play the data.

[1485] Step 12:

[1486] Feedback for dogs and cats

[1487] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[1488] Input: Signal received from the server

[1489] Output: Instructions given to dogs and cats

[1490] Specific operation: The device's speaker will play the voice command "wait" and the dog or cat will follow the command.

[1491] (Application example 2)

[1492] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1493] When elderly people communicate with dogs or cats, it is difficult to accurately understand their pets' needs and emotions, making it difficult to provide effective animal therapy. Furthermore, because it is not possible to respond to elderly people with pets that take their psychological state into consideration, it is difficult to provide them with further psychological comfort and healing. To solve these issues, a system is needed that not only analyzes pet behavior and meows, but also analyzes the user's psychological state in real time and provides feedback based on that analysis.

[1494] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1495] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and communicating them to the dogs and cats, and means for analyzing the user's psychological state and providing feedback based on the psychological state. This enables accurate and effective communication between elderly people and their pets, improving the provision of psychological comfort and healing, and thereby enhancing the effectiveness of animal therapy.

[1496] "Behavior" refers to the general actions and behaviors of dogs and cats.

[1497] "Meow" refers to the sound made by dogs and cats.

[1498] "Data" refers to information including the behavior and sounds of dogs and cats.

[1499] "Server" refers to a centralized computer system that performs data analysis.

[1500] "Analysis means" refers to the technology and algorithms used to analyze acquired data and extract intent and emotion.

[1501] "Conversion to natural language" refers to the process of converting information extracted by analytical means into language that can be understood by humans.

[1502] The "means for notifying" refers to a means for notifying the user of the result of the conversion into natural language.

[1503] "Response" refers to the words or actions that the user responds to the dog or cat.

[1504] "Means for converting into signals" refers to technology that converts the user's response into a format that dogs and cats can understand.

[1505] "Mental state" refers to the user's psychological and emotional state.

[1506] "Means of providing feedback" refers to means for making appropriate responses or taking appropriate actions based on the analyzed results.

[1507] The present invention relates to an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following main components:

[1508] 1. Device:

[1509] A device operated by the user, including a smartphone or a robot installed in a store. The device is equipped with a camera, microphone, display, speaker, and emotion engine. The camera and microphone are used to capture data on the behavior and meows of dogs and cats, and the data is sent to a server.

[1510] 2. Server:

[1511] It is a centralized computer system that analyzes data. The server uses a generative AI model to analyze data on dog and cat behavior and meows, converting it into natural language. It also has an emotion engine that analyzes the user's psychological state and provides feedback based on this.

[1512] 3. Network:

[1513] This is a communications infrastructure that connects devices and servers and transmits data, enabling real-time analysis and minimizing time lags.

[1514] A natural language description of the program's operation

[1515] The server analyzes the behavior and meow data of the cat or dog sent from the device and converts it into natural language. It also uses an emotion engine to analyze the user's psychological state based on the user's voice and facial expression data collected from the device. To provide appropriate feedback to the user based on the analysis results, the server uses high-performance voice recognition technology (e.g., Google Speech-to-Text API) and behavior analysis algorithms (e.g., OpenCV).

[1516] The device notifies the user of the analysis results and feedback sent from the server via text and voice, with messages displayed on the display and voice played through the speaker.

[1517] For example, if an analysis of the barking behavior data of a dog named "Pochi" determines that Pochi wants to play, the device's display will show "Pochi wants to play" and a voice will read out "Pochi wants to play." This feedback allows the user to understand Pochi's request and respond appropriately.

[1518] Specific examples

[1519] As a concrete example, consider the use of the system when an elderly person visits a physical store and spends time with their pet. The elderly person visits the store with their dog, Pochi, and launches the app on their smartphone. The device uses a camera and microphone to collect Pochi's barking, jumping, and other behaviors, and sends the collected information to a server. The server analyzes this data and generates a result, such as "Pochi likes his new toy," which is sent back to the device. The device then notifies the elderly of this result via a display and speaker.

[1520] Prompt Sentence Examples

[1521] As a concrete example, here is an example of an input prompt for a generative AI model:

[1522] "Tell me what Pochi is thinking."

[1523] "Please analyze Pochi's current emotional state."

[1524] This will enable accurate and effective communication between elderly people and pets, improving psychological comfort and healing, and maximizing the effects of animal-assisted therapy.

[1525] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1526] Step 1:

[1527] Initial Setup

[1528] The user launches the application on their device and enters basic information about their pet (dog or cat), such as its name, age, sex, and breed, which the system uses to set optimal analysis parameters.

[1529] Input: Pet's name, age, sex, type

[1530] Output: Basic information about the pet registered in the system

[1531] Specific operation: The user launches the smartphone app and enters "Pochi," "5 years old," "male," and "dog" into the input form.

[1532] Step 2:

[1533] Voice and behavioral data acquisition

[1534] The device captures images of your pet's behavior with a camera and records its cries with a microphone, and this data is collected in real time and sent to a server.

[1535] Input: Camera video data, microphone audio data

[1536] Output: Captured video and audio data

[1537] Specific operation: The device's camera takes a picture of Pochi barking and the microphone records the sound.

[1538] Step 3:

[1539] Data transmission

[1540] The video and audio data collected by the device is divided into packets and sent to the server. Since the transmission is done in real time, the time lag of the information can be minimized.

[1541] Input: Video data, audio data

[1542] Output: Data packets sent to the server

[1543] Specific operation: The device divides the recording data into packets and sends them to the server via Wi-Fi or mobile communication.

[1544] Step 4:

[1545] Data analysis

[1546] The server analyzes the received data: the voice data is converted into text using speech recognition technology (e.g., Google Speech-to-Text API), and the video data is analyzed using a behavioral analysis algorithm (e.g., OpenCV).

[1547] Input: Data packet sent to the server

[1548] Output: Textualized call data, recognized behavior data

[1549] Specific operation: The server analyzes the audio data and converts it into text that Pochi is barking "woof woof." It also analyzes the video data and recognizes that Pochi is wagging its tail.

[1550] Step 5:

[1551] Emotion analysis

[1552] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their psychological state, enabling interactions based on the user's psychological state.

[1553] Input: User's voice data, facial expression data

[1554] Output: Analyzed user's mental state

[1555] Specific operation: The device collects the user's tone of voice and facial expressions (e.g., smile) using a camera and microphone, and analyzes them using an emotion engine. It then recognizes that the user is smiling.

[1556] Step 6:

[1557] Change of mind

[1558] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1559] Input: Textualized call data, recognized behavior data

[1560] Output: Text and audio data to be communicated to the user

[1561] Specific operation: The server generates a message saying "Pochi wants to play" and sends it to the device as text and voice data.

[1562] Step 7:

[1563] User Feedback

[1564] The device receives the analysis results sent from the server and from the emotion engine, and notifies the user through the display and speaker.

[1565] Input: Text data and audio data sent from the server

[1566] Output: Messages displayed on the display, audio played from the speaker

[1567] Specific operation: The device's display will show "Pochi wants to play" and a voice notification will be played at the same time.

[1568] Step 8:

[1569] User response input

[1570] The user speaks to the pet, the voice is recorded by the microphone on the terminal, and the voice data is sent to the server.

[1571] Input: User's voice data

[1572] Output: Audio data sent to the server

[1573] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the device's microphone records the voice.

[1574] Step 9:

[1575] User voice analysis

[1576] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the pet can understand.

[1577] Input: Audio data sent to the server

[1578] Output: Instruction data that your pet can understand

[1579] Specific operation: The server converts the voice data "Pochi, wait while I bring you a toy" into text and converts that instruction into a signal for the pet.

[1580] Step 10:

[1581] Signal Generation

[1582] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1583] Input: Instruction data that your pet can understand

[1584] Output: Signal data ready for playback

[1585] Specific operation: The server sends instruction data to the terminal, and the terminal receives the signal data and prepares for playback.

[1586] Step 11:

[1587] Pet Feedback

[1588] The device receives signals from the server and transmits them to the pet as audio and visual instructions, allowing the pet to understand the elderly person's responses.

[1589] Input: Signal data ready to be played

[1590] Output: Commands (audio or visual) that your pet will understand

[1591] Specific operation: The device plays a voice command saying "Please wait" and Pochi follows the command.

[1592] This will enable accurate and effective communication between elderly people and their pets, improving psychological comfort and healing.

[1593] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1594] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1595] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1596] [Fourth embodiment]

[1597] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1598] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1599] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1600] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1601] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1602] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1603] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1604] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1605] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1606] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1607] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1608] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1609] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1610] The present invention is an animal therapy system that allows elderly people to converse with dogs and cats with high accuracy. The program processing of this system is explained in natural language below. Specific examples are also provided for further details.

[1611] Overall system configuration

[1612] The system consists of the following components:

[1613] 1. Device: A device operated by a user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[1614] 2. Server: The central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[1615] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1616] Program processing

[1617] Initial settings and information entry

[1618] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[1619] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1620] Voice and behavioral data acquisition

[1621] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[1622] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1623] Data analysis and transformation

[1624] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's calls and behavioral data are converted into natural language and text and voice data are generated to be communicated to the user.

[1625] Example: The server interprets Pochi's bark "woof" as meaning "I want to play," and the generative AI model converts it into natural language.

[1626] User Feedback

[1627] The device receives the analysis results and notifies the user visually and audibly, by displaying a message on the screen and announcing it using a text-to-speech function.

[1628] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1629] User response input

[1630] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[1631] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1632] Speech Transcription and Feedback

[1633] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[1634] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1635] Effects and Applications

[1636] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[1637] The processing flow will be explained below.

[1638] Step 1: Initial setup and information entry

[1639] The user installs and launches the application. After launching the application, they enter basic information about their dog or cat (name, age, sex, breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[1640] Step 2: Acquire audio and behavioral data

[1641] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[1642] Step 3: Send data

[1643] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[1644] Step 4: Receiving Data

[1645] The server receives the data sent from the terminal and waits for analysis.

[1646] Step 5: Data analysis

[1647] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[1648] Step 6: Change your mindset

[1649] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1650] Step 7: User feedback

[1651] The device receives the analysis results sent from the server. The device notifies the user of the analysis results by text and voice. A message is displayed on the screen and voice is played from the speaker.

[1652] Step 8: User response input

[1653] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[1654] Step 9: Analyze user voice

[1655] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[1656] Step 10: Generate the signal

[1657] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1658] Step 11: Feedback to your dog or cat

[1659] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[1660] Step 12: Iterating

[1661] By repeating this series of processes, natural and smooth communication is maintained between the user and the dog or cat.

[1662] Example 1

[1663] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1664] It is known that communicating with animals has a positive effect on the physical and mental health of elderly people. However, conventional technology has difficulty accurately understanding animal behavior and sounds and converting them into natural language. Furthermore, technology for converting user responses into a form that animals can understand has not yet been fully developed. This has made it difficult for elderly people to deepen their communication with animals.

[1665] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1666] In this invention, the server includes a means for acquiring the behavior and sounds of animals, a means for transmitting the acquired data to the server, and an analysis means for analyzing the acquired data in the server and converting it into natural language, thereby enabling elderly people to communicate smoothly with animals in real time.

[1667] The "means for acquiring animal behavior and sounds" refers to a device that collects animal movements and sounds using a sensor or input device.

[1668] The "means for transmitting the acquired data to the server" is a device that transfers the collected data as a digital signal to the server via a network.

[1669] The "analysis means for analyzing the acquired data in the server and converting it into natural language" is a device that includes algorithms and programs for analyzing the collected data and converting it into language that the user can understand.

[1670] The "means for notifying the user of the converted results" is a device that has the function of visually or audibly conveying the analysis results to the user.

[1671] "Means for converting the user's response into a signal that the animal can understand and transmitting it to the animal" refers to a device that converts the user's instructions and responses into sounds or signals that the animal can recognize and transmits them to the animal.

[1672] "Voice recognition technology" is a technology that analyzes input voice data and converts it into characters or commands.

[1673] A "behavioral analysis algorithm" is a computational method for analyzing animal behavioral data and inferring their intentions and motivations.

[1674] "Filming equipment" refers to cameras or video equipment used to record animal behavior.

[1675] The "voice input device" is a microphone for collecting animal sounds and the user's voice.

[1676] MODE FOR CARRYING OUT THE INVENTION

[1677] The animal therapy system of the present invention is designed to assist elderly people in advanced communication with animals, and each component of the system is configured as follows.

[1678] Overall system configuration

[1679] The system consists of the following main components:

[1680] 1. Device: A device operated by the user (elderly person), such as a smartphone or dedicated hardware device. This device is equipped with a camera, microphone, display, and speaker.

[1681] 2. Server: This is the central analysis unit where the generative AI model runs. The server analyzes animal behavior and vocalization data and converts it into natural language.

[1682] 3. Network: The infrastructure that connects terminals and servers and enables data communication.

[1683] Initial settings and information entry

[1684] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, species) on the initial setup screen, and the system then sets the optimal analysis parameters for the animal.

[1685] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1686] Voice and behavioral data acquisition

[1687] The device activates the camera and microphone to record and record animal sounds and behavior in real time, and this data is saved as audio and video files.

[1688] Example: The device records and films Pochi's barking and tail wagging, and saves the data.

[1689] Data transmission

[1690] The device divides the recorded data into data packets and sends them to the server over the network, a process that allows the server to receive the data in real time.

[1691] Example: The device splits Pochi's audio and video files into small packets and sends them to a server via Wi-Fi.

[1692] Data Analysis and Natural Language Translation

[1693] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[1694] Example: The server analyzes Pochi's bark "woof" and recognizes it as an intention to "want to play." The generative AI model then converts it into the message "Pochi wants to play."

[1695] User Feedback

[1696] The device receives the analysis results and notifies the user visually and audibly. A message is displayed on the screen and the results are simultaneously announced using a text-to-speech function.

[1697] Example: The device display will show "Pochi wants to play" and at the same time it will read out loud "Pochi wants to play."

[1698] User response input

[1699] The user responds to the animal by speaking into the device, and the voice is recorded and sent to the server.

[1700] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[1701] Response transcription and feedback

[1702] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand. The converted data is then sent to the device, which then plays it back and communicates it to the animals.

[1703] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1704] Example prompts for generative AI models

[1705] "Please convert the dog's bark 'woof' into natural language."

[1706] "Convert the user's response, 'Stay tuned, I'll bring you a toy,' into a signal that the dog can understand."

[1707] This system enables advanced communication between users and animals, allowing elderly people to experience physical and mental healing through contact with animals. It is also expected to improve the symptoms of dementia.

[1708] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1709] Step 1: Initial setup and information entry

[1710] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about the animal (name, age, sex, and species) on the initial setup screen.

[1711] Input: Animal name, age, sex, species.

[1712] Data processing: The input data is applied to the system settings to generate analysis parameters specific to the animal.

[1713] Output: The basic information of the animal is saved in the system.

[1714] How it works: After the user downloads and launches the app on their smartphone, they enter details about the animal into the form that appears on the screen.

[1715] Step 2: Acquire voice and behavioral data

[1716] The device activates the camera and microphone to record and record animal sounds and behavior in real time, and this data is saved as audio and video files.

[1717] Input: Animal sounds and behavior.

[1718] Data processing: Convert audio into an audio file and save video as a video file.

[1719] Output: Audio and video files.

[1720] What it does: When Pochi barks in the living room, the device's microphone picks up the sound and the camera records its movements.

[1721] Step 3: Send data

[1722] The device divides the recorded data into data packets and sends them to the server over the network, allowing the server to receive the data in real time.

[1723] Input: Audio and video files.

[1724] Data processing: Splitting audio and video files into data packets.

[1725] Output: The data packet is sent to the server.

[1726] Specific operation: The device splits Pochi's audio and video files into small packets and sends them to the server via Wi-Fi.

[1727] Step 4: Data analysis and natural language translation

[1728] The server analyzes the received data, using voice recognition technology to analyze the bird's calls and behavioral analysis algorithms to analyze the video. The generative AI model then converts the analysis results into natural language.

[1729] Input: Data packets (audio and video files).

[1730] Data processing: Analysis of audio and video data, and conversion of analysis results into natural language.

[1731] Output: A natural language message.

[1732] Specific operation: The server analyzes Pochi's bark "woof", recognizes it as an intention to "want to play", and converts this into a message in natural language saying "Pochi wants to play."

[1733] Step 5: User feedback

[1734] The device receives the analysis results and notifies the user visually and audibly, with a message displayed on the screen and a voice readout announcing the results.

[1735] Input: Natural language message (analysis result).

[1736] Data processing: Converting natural language messages into display and voice messages.

[1737] Output: Visual and audio notification.

[1738] Specific operation: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1739] Step 6: User response

[1740] The user responds to the animal by speaking into the device, which records the voice and sends it to the server.

[1741] Input: The user's spoken response.

[1742] Data processing: Convert the user's voice into an audio file.

[1743] Output: The audio file is sent to the server.

[1744] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the voice is recorded.

[1745] Step 7: Response transcription and feedback

[1746] The server analyzes the user's voice and converts it into sounds and signals that the animals can understand, then sends the converted data to the device, which plays it back and communicates it to the animals.

[1747] Input: The user's audio file.

[1748] Data processing: Converting the user's voice into signals that animals can understand.

[1749] Output: A signal that can be understood by animals is sent to the terminal and played.

[1750] Specific operation: The server converts the user's voice saying "Wait, I'll bring the toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1751] (Application example 1)

[1752] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1753] There is a need for systems that can enrich the time that elderly people spend with their pets and facilitate smooth communication. However, current systems have difficulty accurately analyzing pet behavior and sounds, and lack the means to provide appropriate feedback to users. Furthermore, technology to convert user input and responses into signals that pets can easily understand is still in development. This makes communication with pets ineffective, and does not sufficiently enhance the physical and mental healing and satisfaction of elderly people.

[1754] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1755] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and transmitting them to the dogs and cats, and means for performing initial settings based on the user's input. This makes it possible to analyze the pet's intentions with high accuracy and convey them to the user, and convert the user's responses into signals easily understandable by the pet and convey them to the pet. This improves communication between the elderly and their pets, increasing physical and mental healing and satisfaction.

[1756] "Means for capturing behavior and sounds" refers to devices that use cameras and microphones to record the behavior and sounds of dogs and cats.

[1757] The "means for transmitting acquired data to a server" is a system that transfers the captured behavioral and vocalization data to a server via the Internet or other communication network.

[1758] The "analysis means" is a combination of software and hardware that runs on a server, analyzes acquired data using voice recognition technology and behavior analysis algorithms, and converts the intentions of dogs and cats into natural language.

[1759] The "means for notifying the user" refers to a device for visually or audibly notifying the user of the analyzed results, such as a system including a display or speaker.

[1760] "Means for converting the user's response into signals that dogs and cats can understand and transmitting them to dogs and cats" is a system that analyzes the user's speech and input, converts them into voice and movement signals that dogs and cats can easily understand, and transmits them to them.

[1761] The "means for performing initial settings" refers to an interface and processing means for having the user input necessary information when using the system for the first time, and optimizing analysis parameters based on that information.

[1762] This invention provides an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following components:

[1763] Overall system configuration

[1764] 1. Device: A device operated by the user (elderly person), including a smartphone or dedicated hardware. It is equipped with a camera, microphone, display, and speaker.

[1765] 2. Server: This is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[1766] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1767] Initial settings and information entry

[1768] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen, and the system will then set up optimal analysis parameters.

[1769] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," age "5 years old," and gender "male."

[1770] Voice and behavioral data acquisition

[1771] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time. This data is then sent to the server. The video is captured using the opencv library, and the audio is recorded using the pyaudio library.

[1772] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1773] Data analysis and transformation

[1774] The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The results of the analysis of the bird's vocalizations and behavioral data are converted into natural language, using a generative AI model.

[1775] Example: The server analyzes Pochi's bark "woof" as the intention to "want to play," and the generative AI model converts it into natural language.

[1776] User Feedback

[1777] The device receives the analysis results and notifies the user visually and audibly, including by displaying a message on the screen and announcing it using a text-to-speech function.

[1778] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1779] User response input

[1780] The user responds to the dog or cat by speaking into the device, and the voice is recorded and sent to the server.

[1781] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1782] Speech Transcription and Feedback

[1783] The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is then sent to the device, which then transmits the signals to the dog or cat.

[1784] Example: The server converts the user's voice saying "Wait, I'll bring you a toy" into an "attention-grabbing audio signal," and the device plays that signal.

[1785] Effects and Applications

[1786] This system allows for advanced communication between users and cats and dogs, allowing elderly people to experience physical and mental healing through contact with cats and dogs, and is also expected to improve the symptoms of dementia.

[1787] Example prompt sentence:

[1788] "Analyze your pet's cry 'woof' and translate it into natural language to mean 'I want to play'."

[1789] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1790] Step 1: Initial setup and information entry

[1791] Specific operation: The user installs and launches the application. After launching the application, the user enters basic information about the pet (name, age, sex, and species) on the initial setup screen. This information is sent from the device to the server.

[1792] Input: Pet's name, age, sex, type

[1793] Data processing and calculation: Based on the pet information entered, the server sets the optimal analysis parameters.

[1794] Output: Initial setup completed and analysis parameters set

[1795] Step 2: Acquire audio and behavioral data

[1796] Specific operation: The device's camera and microphone are activated to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[1797] Input: Dog and cat meows and behavior

[1798] Data processing and data calculation: Capture video with the camera (using opencv) and record audio with the microphone (using pyaudio). Divide this data into packets and send them to the server.

[1799] Output: Audio and video data sent to the server

[1800] Step 3: Data analysis and transformation

[1801] Specific operation: The server receives the data sent from the device and analyzes it using voice recognition technology and behavioral analysis algorithms. The analyzed results are converted into natural language.

[1802] Input: Audio and video data

[1803] Data processing and data calculation: Generative AI models are used to analyze vocalizations and behavioral patterns and generate prompts that are translated into natural language.

[1804] Output: Translated natural language text

[1805] Step 4: User feedback

[1806] Specific operation: The device notifies the user of the analysis results received from the server visually and audibly, showing a message on the display and announcing it using the text-to-speech function.

[1807] Input: Natural language text sent from the server

[1808] Data processing and data calculation: The text data is displayed on the screen and simultaneously read aloud to the user as audio data.

[1809] Output: User notification (display message and voice announcement)

[1810] Step 5: User response input

[1811] Specific operation: The user responds to the dog or cat by speaking into the device, and the recorded voice is sent to the server.

[1812] Input: User's voice response

[1813] Data processing and data calculation: The device records the user's voice and sends the data to the server.

[1814] Output: Audio data sent to the server

[1815] Step 6: Voice conversion and feedback

[1816] How it works: The server analyzes the user's voice and converts it into sounds and signals that the dog or cat can understand. The converted data is sent to the device, which then plays back the signals and transmits them to the dog or cat.

[1817] Input: Voice data sent by the user

[1818] Data Processing and Data Computation: Using a generative AI model, we convert the user's voice into an audio signal that is easy for the pet to understand.

[1819] Output: Audio signal to communicate with your pet

[1820] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1821] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it realizes interactions based on the user's psychological state. The program processing of this system is explained in natural language and detailed with concrete examples.

[1822] Overall system configuration

[1823] The system consists of the following components:

[1824] 1. Device: A device operated by the user (elderly person). This includes smartphones and dedicated hardware. It is equipped with a camera, microphone, display, speaker, and emotion engine.

[1825] 2. Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[1826] 3. Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1827] Program processing

[1828] Initial settings and information entry

[1829] When a user uses the system for the first time, they install and launch the application on their device. After launching the application, they enter basic information about their dog or cat (name, age, sex, and breed) on the initial setup screen. This allows the system to set optimal analysis parameters.

[1830] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1831] Voice and behavioral data acquisition

[1832] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and this data is then sent to the server.

[1833] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1834] Data transmission

[1835] The audio and video data collected by the device is divided into packets and sent to the server. The transmission is done in real time, so there is little time lag.

[1836] Data reception

[1837] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[1838] Data analysis

[1839] The server analyzes the audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[1840] Emotion analysis

[1841] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[1842] Example: The emotion engine detects when a user is smiling, and the system provides positive feedback to the user.

[1843] Change of mind

[1844] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1845] User Feedback

[1846] The device receives the analysis results sent from the server and the emotion engine. The device notifies the user of the analysis results by text and voice. A message is displayed on the display and voice is played from the speaker.

[1847] Example: The device display will show "Pochi wants to play" and at the same time a voice will read out "Pochi wants to play."

[1848] User response input

[1849] The user speaks to the dog or cat. The device's microphone records the user's voice and sends the audio data to the server.

[1850] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1851] User voice analysis

[1852] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which converts it into instructions that the dog or cat can understand.

[1853] Signal Generation

[1854] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1855] Feedback for dogs and cats

[1856] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[1857] Effects and Applications

[1858] This system allows for advanced communication between the user and the dog or cat. By taking into account the user's emotional state, it is possible to provide more effective animal therapy, providing physical and mental healing for the elderly and improving the symptoms of dementia.

[1859] The processing flow will be explained below.

[1860] Step 1: Initial setup and information entry

[1861] The user installs and launches the application. On the application's initial setup screen, they enter basic information about their dog or cat (name, age, sex, breed) to complete the setup.

[1862] Example: A user launches the app, enters the dog's name "Pochi," its age "5 years old," and its gender "male," and completes the initial setup.

[1863] Step 2: Acquire audio and behavioral data

[1864] The device will activate the camera and microphone to record and record the sounds and behavior of your dog or cat in real time.

[1865] Example: The device uses the camera and microphone to record Pochi's barking and tail wagging.

[1866] Step 3: Obtaining emotion data

[1867] The device captures the user's facial expressions and voice using a camera and microphone, and requests the emotion engine to analyze them, thereby recognizing the user's emotional state in real time.

[1868] Example: A camera and microphone capture the user's facial expressions and voice, and the emotion engine analyzes that the user is smiling.

[1869] Step 4: Send data

[1870] The terminal transmits the acquired dog or cat voice data, behavioral data, and user emotion data to the server.

[1871] Example: Pochi's barking, behavioral data, and the user's smile data are divided into packets and sent to the server.

[1872] Step 5: Receiving Data

[1873] The server receives the data from the terminal and waits for analysis.

[1874] Example: A server receives dog and cat meow data, behavior data, and user emotion data.

[1875] Step 6: Data analysis

[1876] The server converts the voice data into text using speech recognition technology, analyzes the behavioral data using a behavioral analysis algorithm, and integrates the analysis results into a generative AI model.

[1877] Example: The server converts Pochi's bark "woof" into text and analyzes its tail-wagging behavior to recognize that it wants to play.

[1878] Step 7: Natural Language Translation

[1879] The server uses the generative AI model to convert the dog or cat's intentions into natural language and generate text and voice data to communicate with the user.

[1880] Example: The server translates a dog's barking and tail wagging into "Pochi wants to play."

[1881] Step 8: Reflecting on your emotional state

[1882] The server adjusts the messages and voices it outputs based on the user's emotional data. If the user is in a positive emotional state, it adds reassuring feedback.

[1883] Example: Because the user is smiling, the server generates positive feedback such as "Pochi wants to play. He looks very happy."

[1884] Step 9: User Feedback

[1885] The device notifies the user of the analysis results received from the server in text and audio, showing a message on the display and playing audio through the speaker.

[1886] Example: The device displays "Pochi wants to play. He looks very happy" on the display and reads it out loud.

[1887] Step 10: User response input

[1888] The user responds to the dog or cat. The device's microphone records the user's voice and sends it to the server.

[1889] Example: A user speaks to the device, "Pochi, wait here, I'll bring you a toy."

[1890] Step 11: Analyze user voice

[1891] The server receives the user's voice data and converts it into text using speech analysis technology, which is then input into a generative AI model that converts it into signals that dogs and cats can understand.

[1892] Example: The server analyzes the user's voice saying "Wait, I'll bring you a toy" and converts it into signals that a dog can understand.

[1893] Step 12: Generate and transmit a signal

[1894] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[1895] Example: The server generates a signal and sends it to the terminal, which then plays it back.

[1896] Step 13: Feedback to your dog or cat

[1897] The device transmits signals from the server to the dog or cat as audio and visual instructions.

[1898] Example: The device plays an "attention-grabbing audio signal" and Pochi responds to that signal.

[1899] Step 14: Iterating

[1900] By repeating a series of processes, the system maintains natural and smooth communication between the user and the dog or cat.

[1901] Example 2

[1902] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1903] Conventional animal therapy systems have limited means for elderly people to have smooth and sophisticated communication with dogs and cats, and in particular, they do not adequately consider the user's emotional state. The effectiveness of animal therapy is limited due to the lack of responses and feedback according to the elderly person's psychological state. There is a need to solve this issue and provide a system that enhances psychological and emotional support for the elderly.

[1904] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1905] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, and means for recognizing the user's emotions and providing interaction based on the user's psychological state. This enables elderly people to have advanced communication with dogs and cats, and provides effective animal therapy based on the user's emotional state.

[1906] "Behavior" refers to the movements and gestures of dogs and cats.

[1907] "Meow" is the sound made by dogs and cats.

[1908] "Users" refer to the elderly and other users of the system.

[1909] "Device" refers to a device operated by a user, including a smartphone or dedicated hardware, equipped with a camera, microphone, display, speaker, and emotion engine.

[1910] The "server" is the central analysis device where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts them into natural language.

[1911] The "camera" is a photographic device that records the behavior of dogs and cats and detects the user's facial expressions.

[1912] A "microphone" is a sound collection device that records the sounds of dogs and cats, as well as the user's voice.

[1913] An "emotion engine" is software that analyzes a user's voice and facial expression data to determine the user's emotional state.

[1914] "Analysis means" refers to the process by which the server analyzes the data it receives using voice recognition technology and behavioral analysis algorithms.

[1915] "Natural language" refers to the language that humans use on a daily basis.

[1916] "Interaction" refers to the interaction between the user and the dog or cat.

[1917] "Feedback" refers to the process of providing responses or instructions to the user or dog or cat based on the analysis results.

[1918] A "signal" is a command or response from the user that has been translated into a format that a dog or cat can understand.

[1919] MODE FOR CARRYING OUT THE INVENTION

[1920] This invention relates to an animal therapy system that enables elderly people to converse with dogs and cats with high accuracy. This system realizes interaction based on the user's psychological state by combining an emotion engine that recognizes the user's emotions. The system of this invention uses the following hardware and software:

[1921] Hardware Configuration

[1922] Device: A device operated by a user. This can be a smartphone or dedicated hardware. A device is equipped with a camera, microphone, display, speaker, and emotion engine.

[1923] Server: The central analysis unit where the generative AI model runs. It analyzes data on dog and cat behavior and meows and converts it into natural language. It also processes data from the emotion engine.

[1924] Network: A communications infrastructure that connects terminals and servers and enables data communication.

[1925] Software Configuration

[1926] Emotion engine: Software that analyzes the user's voice and facial expression data to determine the user's emotional state.

[1927] Generative AI model: An AI that converts the actions and meows of dogs and cats into text and translates their intentions into natural language.

[1928] Voice recognition technology and behavior analysis algorithm: Technology that analyzes the meows and behavior of dogs and cats and processes them as text data on a server.

[1929] Overall system configuration

[1930] When a user uses the system, the system operates as follows.

[1931] 1. Initial setup: When a user uses the system for the first time, they install the application on their device and enter basic information such as the dog or cat's name, age, sex, and breed on the initial setup screen.

[1932] Example: The user launches the app and completes the initial setup by entering the dog's name "Pochi," its age "5 years old," and its gender "male."

[1933] 2. Data collection and transmission: The device activates the camera and microphone to record and record the sounds and behavior of dogs and cats in real time. This data is then sent to the server.

[1934] Example: The device records and films Pochi's barking and tail wagging, and then divides the data into packets and sends them to the server.

[1935] 3. Data analysis: The server analyzes the received voice data and converts the barks into text using voice recognition technology. At the same time, it analyzes the video data and recognizes the behavior of the dog or cat using a behavior analysis algorithm. The emotion engine also analyzes the user's voice and facial expression data to determine their emotional state.

[1936] Example: The emotion engine analyzes the user's smile as a "positive emotion."

[1937] 4. Intention Conversion and Feedback: The server combines the voice text and behavior analysis results and inputs them into a generative AI model. This converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate to the user. The device then notifies the user of these analysis results in text and voice.

[1938] Example: The device notifies the user via text and voice, "Pochi wants to play."

[1939] 5. User response: When the user speaks to the dog or cat, the device's microphone records the user's voice and sends it to the server. The server analyzes this voice data, converts it into a format that the dog or cat can understand, and sends it to the device. The device then conveys the converted instructions to the dog or cat as audio or visual instructions.

[1940] Example: The user says, "Pochi, wait, I'll bring you a toy," and the device helps the dog understand the command.

[1941] Examples of prompt statements

[1942] "Write a program that describes a framework that analyzes a dog's barks and behavior and translates its intentions into natural language using a generative AI model."

[1943] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1944] Step 1:

[1945] Initial settings and information entry

[1946] The user installs and launches the application on their device. When the application launches, an initial setup screen appears, and the user enters basic information about the dog or cat, such as its name, age, sex, and breed. The information entered is used by the system to set optimal analysis parameters.

[1947] Input: Basic information such as the name, age, sex, and breed of your dog or cat

[1948] Output: Optimal analysis parameters are set

[1949] Specific actions: The user operates the device's touch screen and enters "Pochi," "5 years old," and "male" into the text boxes.

[1950] Step 2:

[1951] Voice and behavioral data acquisition

[1952] The device activates the camera and microphone to record the sounds and behavior of dogs and cats in real time, and the collected data is sent to the server.

[1953] Input: Real-time dog and cat sounds and behavior

[1954] Output: Recorded and video data is collected

[1955] Specific actions: The device's camera captures Pochi's barking and tail wagging, and the microphone records the sounds.

[1956] Step 3:

[1957] Data transmission

[1958] The audio and video data collected by the device is divided into packets and sent to the server. This transmission is done in real time, so there is little delay.

[1959] Input: Collected audio and video data

[1960] Output: Data packet is sent to the server

[1961] How it works: The device splits the recording data into small packets and sends them to the server via Wi-Fi.

[1962] Step 4:

[1963] Data reception

[1964] The server receives the data sent from the device and waits for analysis. The received data includes information about the behavior and meowing of dogs and cats.

[1965] Input: Data packets sent from the device

[1966] Output: Data waiting to be analyzed is saved on the server

[1967] Specific operation: The server's network adapter receives the data packet and adds the data to the analysis queue.

[1968] Step 5:

[1969] Data analysis

[1970] The server analyzes the received audio data and converts the barks into text using voice recognition technology, while simultaneously analyzing the video data and recognizing the behavior of dogs and cats using a behavior analysis algorithm.

[1971] Input: Received audio and video data

[1972] Output: Call data converted to text and recognized behavior data

[1973] Specific operation: The server analyzes the audio data, converts the "woof woof" sound into text as a "dog bark," and analyzes the video data to recognize the behavior of "tail wagging."

[1974] Step 6:

[1975] Emotion analysis

[1976] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their emotional state, enabling interactions based on the user's psychological state.

[1977] Input: User's voice and facial expression data

[1978] Output: Parsed user's emotional state

[1979] Specific operation: The device's camera captures the user's smiling expression, and the emotion engine analyzes it as a "positive state."

[1980] Step 7:

[1981] Change of mind

[1982] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[1983] Input: Speech text and behavioral analysis results

[1984] Output: Natural language text and audio data to communicate to the user

[1985] Specific operation: The server's generated AI model analyzes behavioral patterns such as "barking" and "wagging its tail," generates the intention of "I want to play," and converts this into natural language.

[1986] Step 8:

[1987] User Feedback

[1988] The device receives the analysis results sent from the server and the emotion engine and notifies the user of the results. The device notifies the user of the analysis results by text and voice.

[1989] Input: Analysis results sent from the server, analysis results from the emotion engine

[1990] Output: Notification message and audio to the user

[1991] Specific operation: The device display will show "Pochi wants to play" and the speaker will play "Pochi wants to play."

[1992] Step 9:

[1993] User response input

[1994] When a user speaks to a dog or cat, the device's microphone records the user's voice and sends the audio data to the server.

[1995] Input: Voice data as the user's response

[1996] Output: The recorded audio data is sent to the server.

[1997] Specific operation: The user speaks to the device, "Pochi, wait here, I'll bring you a toy," and the microphone records the voice.

[1998] Step 10:

[1999] User voice analysis

[2000] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the dog or cat can understand.

[2001] Input: Received user voice data

[2002] Output: Instruction data converted to text

[2003] Specific operation: The server converts the speech "Pochi, wait for me, I'll bring you a toy" into text, and the generative AI model analyzes the text and converts it into the instruction "wait."

[2004] Step 11:

[2005] Signal Generation

[2006] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[2007] Input: Transformed instruction data

[2008] Output: Instruction data sent to the terminal

[2009] Specific operation: The server sends "waiting" instruction data to the terminal, and the terminal prepares to play the data.

[2010] Step 12:

[2011] Feedback for dogs and cats

[2012] The device receives signals from the server and transmits them to the dog or cat as audio or visual instructions, allowing the elderly person's responses to be understood by the dog or cat.

[2013] Input: Signal received from the server

[2014] Output: Instructions given to dogs and cats

[2015] Specific operation: The device's speaker will play the voice command "wait" and the dog or cat will follow the command.

[2016] (Application example 2)

[2017] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2018] When elderly people communicate with dogs or cats, it is difficult to accurately understand their pets' needs and emotions, making it difficult to provide effective animal therapy. Furthermore, because it is not possible to respond to elderly people with pets that take their psychological state into consideration, it is difficult to provide them with further psychological comfort and healing. To solve these issues, a system is needed that not only analyzes pet behavior and meows, but also analyzes the user's psychological state in real time and provides feedback based on that analysis.

[2019] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2020] In this invention, the server includes means for acquiring the behavior and meows of dogs and cats, means for transmitting the acquired data to the server, means for analyzing the acquired data in the server and converting it into natural language, means for notifying the user of the results of the conversion into natural language, means for converting the user's responses into signals understandable by dogs and cats and communicating them to the dogs and cats, and means for analyzing the user's psychological state and providing feedback based on the psychological state. This enables accurate and effective communication between elderly people and their pets, improving the provision of psychological comfort and healing, and thereby enhancing the effectiveness of animal therapy.

[2021] "Behavior" refers to the general actions and behaviors of dogs and cats.

[2022] "Meow" refers to the sound made by dogs and cats.

[2023] "Data" refers to information including the behavior and sounds of dogs and cats.

[2024] "Server" refers to a centralized computer system that performs data analysis.

[2025] "Analysis means" refers to the technology and algorithms used to analyze acquired data and extract intent and emotion.

[2026] "Conversion to natural language" refers to the process of converting information extracted by analytical means into language that can be understood by humans.

[2027] The "means for notifying" refers to a means for notifying the user of the result of the conversion into natural language.

[2028] "Response" refers to the words or actions that the user responds to the dog or cat.

[2029] "Means for converting into signals" refers to technology that converts the user's response into a format that dogs and cats can understand.

[2030] "Mental state" refers to the user's psychological and emotional state.

[2031] "Means of providing feedback" refers to means for making appropriate responses or taking appropriate actions based on the analyzed results.

[2032] The present invention relates to an animal therapy system that allows elderly people to communicate with dogs and cats with high accuracy. The system consists of the following main components:

[2033] 1. Device:

[2034] A device operated by the user, including a smartphone or a robot installed in a store. The device is equipped with a camera, microphone, display, speaker, and emotion engine. The camera and microphone are used to capture data on the behavior and meows of dogs and cats, and the data is sent to a server.

[2035] 2. Server:

[2036] It is a centralized computer system that analyzes data. The server uses a generative AI model to analyze data on dog and cat behavior and meows, converting it into natural language. It also has an emotion engine that analyzes the user's psychological state and provides feedback based on this.

[2037] 3. Network:

[2038] This is a communications infrastructure that connects devices and servers and transmits data, enabling real-time analysis and minimizing time lags.

[2039] A natural language description of the program's operation

[2040] The server analyzes the behavior and meow data of the cat or dog sent from the device and converts it into natural language. It also uses an emotion engine to analyze the user's psychological state based on the user's voice and facial expression data collected from the device. To provide appropriate feedback to the user based on the analysis results, the server uses high-performance voice recognition technology (e.g., Google Speech-to-Text API) and behavior analysis algorithms (e.g., OpenCV).

[2041] The device notifies the user of the analysis results and feedback sent from the server via text and voice, with messages displayed on the display and voice played through the speaker.

[2042] For example, if an analysis of the barking behavior data of a dog named "Pochi" determines that Pochi wants to play, the device's display will show "Pochi wants to play" and a voice will read out "Pochi wants to play." This feedback allows the user to understand Pochi's request and respond appropriately.

[2043] Specific examples

[2044] As a concrete example, consider the use of the system when an elderly person visits a physical store and spends time with their pet. The elderly person visits the store with their dog, Pochi, and launches the app on their smartphone. The device uses a camera and microphone to collect Pochi's barking, jumping, and other behaviors, and sends the collected information to a server. The server analyzes this data and generates a result, such as "Pochi likes his new toy," which is sent back to the device. The device then notifies the elderly of this result via a display and speaker.

[2045] Prompt Sentence Examples

[2046] As a concrete example, here is an example of an input prompt for a generative AI model:

[2047] "Tell me what Pochi is thinking."

[2048] "Please analyze Pochi's current emotional state."

[2049] This will enable accurate and effective communication between elderly people and pets, improving psychological comfort and healing, and maximizing the effects of animal-assisted therapy.

[2050] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2051] Step 1:

[2052] Initial Setup

[2053] The user launches the application on their device and enters basic information about their pet (dog or cat), such as its name, age, sex, and breed, which the system uses to set optimal analysis parameters.

[2054] Input: Pet's name, age, sex, type

[2055] Output: Basic information about the pet registered in the system

[2056] Specific operation: The user launches the smartphone app and enters "Pochi," "5 years old," "male," and "dog" into the input form.

[2057] Step 2:

[2058] Voice and behavioral data acquisition

[2059] The device captures images of your pet's behavior with a camera and records its cries with a microphone, and this data is collected in real time and sent to a server.

[2060] Input: Camera video data, microphone audio data

[2061] Output: Captured video and audio data

[2062] Specific operation: The device's camera takes a picture of Pochi barking and the microphone records the sound.

[2063] Step 3:

[2064] Data transmission

[2065] The video and audio data collected by the device is divided into packets and sent to the server. Since the transmission is done in real time, the time lag of the information can be minimized.

[2066] Input: Video data, audio data

[2067] Output: Data packets sent to the server

[2068] Specific operation: The device divides the recording data into packets and sends them to the server via Wi-Fi or mobile communication.

[2069] Step 4:

[2070] Data analysis

[2071] The server analyzes the received data: the voice data is converted into text using speech recognition technology (e.g., Google Speech-to-Text API), and the video data is analyzed using a behavioral analysis algorithm (e.g., OpenCV).

[2072] Input: Data packet sent to the server

[2073] Output: Textualized call data, recognized behavior data

[2074] Specific operation: The server analyzes the audio data and converts it into text that Pochi is barking "woof woof." It also analyzes the video data and recognizes that Pochi is wagging its tail.

[2075] Step 5:

[2076] Emotion analysis

[2077] The device uses an emotion engine to analyze the user's voice and facial expression data and determine their psychological state, enabling interactions based on the user's psychological state.

[2078] Input: User's voice data, facial expression data

[2079] Output: Analyzed user's mental state

[2080] Specific operation: The device collects the user's tone of voice and facial expressions (e.g., smile) using a camera and microphone, and analyzes them using an emotion engine. It then recognizes that the user is smiling.

[2081] Step 6:

[2082] Change of mind

[2083] The server combines the voice text and behavioral analysis results and inputs them into a generative AI model, which converts the dog or cat's intentions and requests into natural language and generates text and voice data to communicate with the user.

[2084] Input: Textualized call data, recognized behavior data

[2085] Output: Text and audio data to be communicated to the user

[2086] Specific operation: The server generates a message saying "Pochi wants to play" and sends it to the device as text and voice data.

[2087] Step 7:

[2088] User Feedback

[2089] The device receives the analysis results sent from the server and from the emotion engine, and notifies the user through the display and speaker.

[2090] Input: Text data and audio data sent from the server

[2091] Output: Messages displayed on the display, audio played from the speaker

[2092] Specific operation: The device's display will show "Pochi wants to play" and a voice notification will be played at the same time.

[2093] Step 8:

[2094] User response input

[2095] The user speaks to the pet, the voice is recorded by the microphone on the terminal, and the voice data is sent to the server.

[2096] Input: User's voice data

[2097] Output: Audio data sent to the server

[2098] Specific operation: The user speaks to the device, saying, "Pochi, wait here, I'll bring you a toy," and the device's microphone records the voice.

[2099] Step 9:

[2100] User voice analysis

[2101] The server receives the user's voice data and converts it into text using speech analysis technology. The converted text data is then input into a generative AI model, which then converts it into instructions that the pet can understand.

[2102] Input: Audio data sent to the server

[2103] Output: Instruction data that your pet can understand

[2104] Specific operation: The server converts the voice data "Pochi, wait while I bring you a toy" into text and converts that instruction into a signal for the pet.

[2105] Step 10:

[2106] Signal Generation

[2107] The server sends the converted instruction data to the terminal, which prepares to play the signal.

[2108] Input: Instruction data that your pet can understand

[2109] Output: Signal data ready for playback

[2110] Specific operation: The server sends instruction data to the terminal, and the terminal receives the signal data and prepares for playback.

[2111] Step 11:

[2112] Pet Feedback

[2113] The device receives signals from the server and transmits them to the pet as audio and visual instructions, allowing the pet to understand the elderly person's responses.

[2114] Input: Signal data ready to be played

[2115] Output: Commands (audio or visual) that your pet will understand

[2116] Specific operation: The device plays a voice command saying "Please wait" and Pochi follows the command.

[2117] This will enable accurate and effective communication between elderly people and their pets, improving psychological comfort and healing.

[2118] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2119] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2120] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2121] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2122] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2123] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2124] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2125] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2126] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2127] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2128] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2129] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2130] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2131] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2132] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2133] The hardware resource for executing a specific pr...

Claims

1. A means for acquiring dog and cat behavior and meow sounds; means for transmitting the acquired data to a server; an analysis means in the server for analyzing the acquired data and converting it into natural language; means for notifying a user of the result of the conversion into the private language; A means for converting the user's response into a signal that can be understood by dogs and cats and transmitting the signal to the dogs and cats; A system including:

2. 2. The system of claim 1, wherein said analysis means uses voice recognition techniques and behavioral analysis algorithms.

3. 2. The system according to claim 1, wherein the means for acquiring the data is a camera and a microphone mounted on the user's terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A