system

The system addresses the challenge of human-animal communication by recording, processing, and translating animal sounds into natural language, facilitating effective interaction.

JP2026062129APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional methods struggle to facilitate direct communication between humans and animals, particularly in understanding the intentions of animals through sounds like those of dolphins, dogs, and birds, requiring high expertise and leading to inaccurate understanding of their behavior and state.

Method used

A system that records animal sounds, converts them into digital signals, removes noise, analyzes and maps the features to animal-specific communication formats, translates the data into natural language, and plays back the data to enable two-way communication.

Benefits of technology

Enables smooth communication between humans and animals by accurately analyzing and translating animal vocalizations into understandable language, allowing for better understanding of animal intentions and behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062129000001_ABST
    Figure 2026062129000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】Means for recording the vocalizations of animals, Means for converting the recorded vocalizations into digital signals, Means for removing noise from the digital signals, Means for transmitting the preprocessed audio data to a server, Means for analyzing the audio data and extracting features, Means for mapping the extracted features to a communication format for each animal, Means for translating the mapped data into natural language, Means for transmitting the translated data and the generated vocalization data to a terminal, Means for playing back the data transmitted to the terminal, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventionally, it has been difficult to directly communicate between humans and animals. In particular, it has been difficult to convey intentions to animals through sounds such as those of dolphins, dogs, and birds, and a high level of expertise has been required to understand those intentions. As a result, in many situations, the behavior and state of animals could not be accurately understood, and the lack of communication has been a problem. The present invention aims to solve such conventional problems and provide a system that enables smooth communication between humans and animals.

Means for Solving the Problems

[0005] The present invention is a system that includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting the pre-processed audio data to a server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format specific to each animal, means for translating the mapped data into natural language, means for transmitting the translated data and the generated sound data to a terminal, and means for playing back the data transmitted to the terminal. It further includes means for estimating the animal's intentions and translating the estimated intentions into natural language, and means for analyzing user input and converting the user's natural language into an animal sound format. This enables smooth communication between humans and animals.

[0006] "Means for recording animal sounds" refers to a device equipped with equipment and software for capturing sounds emitted by animals.

[0007] "Means for converting recorded animal sounds into digital signals" refers to converters and processing technologies for converting analog audio data into a digital format.

[0008] "Means for removing noise from digital signals" refer to algorithms and processing techniques for filtering unwanted background noise from digital audio data.

[0009] "Means for transmitting pre-processed audio data to a server" refers to communication technologies and data transmission protocols for transmitting pre-processed digital audio data to a remote server.

[0010] "Means for analyzing audio data and extracting features" refer to analysis algorithms and processing techniques for identifying specific features or patterns from audio data.

[0011] "Means for mapping extracted features to animal-specific communication formats" refers to data mapping algorithms and processing techniques for converting analyzed features into animal-specific communication formats.

[0012] "Means for translating mapped data into natural language" refers to translation algorithms and processing techniques for converting animal vocalization patterns into human natural language.

[0013] "Means for transmitting translated data and generated sound data to a terminal" refers to communication technology and data transmission protocols for transmitting translation results and sound data generated on a server to a terminal.

[0014] "Means for playing back data transmitted to a terminal" refers to hardware and software for playing back and displaying data received on a terminal as audio or text.

[0015] "Means for estimating animal intentions and translating those intentions into natural language" refers to algorithms and processing technologies for identifying an animal's intentions from its vocalizations and converting them into natural language that humans can understand.

[0016] "Means for analyzing user input and converting the user's natural language into animal sound format" refers to algorithms and processing technologies for analyzing natural language instructions entered by a user and converting them into the corresponding animal sound format. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the language used in the following description will be explained.

[0020] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0021] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] The present invention's system analyzes and translates animal sounds, enabling communication between humans and animals. Specific embodiments of the system are described below.

[0039] System Configuration

[0040] 1. Terminal

[0041] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, and a communication function for sending data to a server.

[0042] 2. Server

[0043] The server is a computer system that performs data analysis and translation in a central location. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication format of each animal, and a mechanism for translating into natural language.

[0044] Processing flow

[0045] 1. Input

[0046] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[0047] 2. Preprocessing

[0048] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[0049] 3. Data transmission

[0050] The terminal sends pre-processed audio data to the server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[0051] 4. Voice Analysis

[0052] The server analyzes the received audio data and extracts features. Specifically, this involves spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs).

[0053] 5. Feature Mapping

[0054] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[0055] 6. Translation

[0056] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[0057] 7. Response generation

[0058] The server receives input from the user and converts it into an animal sound format. If the user asks "Tell me where to go," the corresponding animal sound is generated.

[0059] 8. Output

[0060] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[0061] Specific example

[0062] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[0063] The above describes a specific embodiment of the system of the present invention. This system enables smooth communication with animals.

[0064] The following describes the processing flow.

[0065] Step 1:

[0066] The user speaks into the device and says, "Please record dolphin sounds." The device uses its voice recognition function to detect the user's command and enters recording mode.

[0067] Step 2:

[0068] The device uses its built-in microphone to record ambient sounds. Recording continues for a period set by the user, or until manually stopped.

[0069] Step 3:

[0070] The device converts the recorded analog audio data into a digital signal. This is done using an AD converter (analog-to-digital converter).

[0071] Step 4:

[0072] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[0073] Step 5:

[0074] The terminal compresses the pre-processed audio data and packages it into data packets for efficient transmission to the server.

[0075] Step 6:

[0076] The terminal sends packaged data to the server via the internet. Stable communication is required.

[0077] Step 7:

[0078] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[0079] Step 8:

[0080] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0081] Step 9:

[0082] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[0083] Step 10:

[0084] The server estimates the animal's intentions from the mapped data and translates those intentions into natural language that the user can understand. For example, it might generate a message like, "There are lots of fish around."

[0085] Step 11:

[0086] The user speaks to the device saying, "Tell me where I should go." The device converts this voice command into text and sends it to the server.

[0087] Step 12:

[0088] The server analyzes the user's question and generates appropriate animal sounds based on its content. For example, it converts information like "There are many fish if you head east" into a dolphin sound format.

[0089] Step 13:

[0090] The server sends the generated sound data to the terminal. The terminal analyzes the received data and prepares for playback.

[0091] Step 14:

[0092] The device plays generated animal sound data and simultaneously displays translated text information to the user. The user can communicate directly with the animal through the device.

[0093] (Example 1)

[0094] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] Traditional methods of communication with animals have been limited, making it difficult to accurately analyze and translate animal calls. Furthermore, the lack of a method to analyze animal calls and translate them into human natural language, and vice versa, prevented two-way communication between animals and humans. This resulted in insufficient communication between animals in captivity and research environments, posing a serious problem.

[0096] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0097] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the features to communication formats specific to each animal, and means for translating the mapped data into natural language. This makes it possible to accurately analyze animal calls and translate their intent into natural language. Furthermore, it becomes possible to analyze user input and convert natural language into animal call formats, enabling two-way communication between animals and humans.

[0098] "Means for recording animal sounds" refers to devices or systems that input the sounds of animals specified by the user as audio.

[0099] "Means for converting recorded animal sounds into digital signals" refers to devices or systems for converting analog audio data into a digital format.

[0100] "Methods for removing noise from digital signals" refer to devices and systems that improve the quality of digital audio data by removing unwanted background noise and interference sounds.

[0101] "Means for sending pre-processed audio data to a server" refers to a device or system for sending processed audio data to a remote server.

[0102] "Means for analyzing audio data and extracting features" refers to devices or systems for analyzing audio data and extracting important feature patterns.

[0103] "Means for mapping extracted features to animal-specific communication formats" refers to devices or systems that map the features of analyzed audio data to the communication formats of specific animals.

[0104] "Means for translating mapped data into natural language" refers to devices or systems that convert data corresponding to animal communication patterns into natural language that humans can understand.

[0105] "Means for transmitting translated data and generated sound data to a terminal" refers to devices or systems for transmitting data translated on a server and animal sound data to a user's terminal.

[0106] "Means for playing back data sent to a terminal" refers to devices or systems for playing back data received on a user's terminal in audio or text format.

[0107] "Means for users to input voice commands" refers to devices or systems that allow users to give instructions to a system using their voice.

[0108] "Means for a server to analyze audio data and extract features" refers to devices or systems that allow a server to analyze audio data and extract important feature patterns.

[0109] "Means by which a server maps animal vocalization patterns" refers to devices or systems that allow a server to map the characteristics of analyzed audio data to the communication patterns of specific animals.

[0110] "Means for a server to translate intentions into natural language" refers to a device or system in which a server estimates an animal's intentions and converts them into natural language that humans can understand.

[0111] "Means by which a server generates a response" refers to a device or system that allows a server to generate corresponding animal sounds based on user input.

[0112] "Means by which a terminal receives response data and presents the results to the user" refers to a device or system that allows a terminal to receive response data from a server and display or play it aloud for the user.

[0113] "Means for a terminal to reproduce animal sounds" refers to a device or system that allows a terminal to reproduce generated animal sounds as audio.

[0114] The system of the present invention analyzes and translates animal sounds to enable smooth communication between humans and animals. Specific embodiments of the system are described below.

[0115] System Configuration

[0116] 1. Terminal

[0117] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, and communication functions for sending data to a server. Specific hardware includes the smartphone's built-in microphone, an external microphone, and Wi-Fi or mobile data communication for communication. Software used includes audio recording applications and noise reduction software such as Audacity.

[0118] 2. Server

[0119] The server is a central computer system that performs data analysis and translation. This server implements an audio data analysis engine, feature extraction algorithms, a system for mapping to animal-specific communication formats, and a mechanism for translating into natural language. Specific software includes spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculation using the Python Librosa library. Furthermore, a generative AI model (e.g., GPT-4®) is responsible for natural language translation.

[0120] Processing flow

[0121] The system operates using the following steps:

[0122] 1. Input: The user issues a voice command to the device, such as, "Please record dolphin sounds." For example, this could be done while observing dolphins and recording their sounds.

[0123] 2. Preprocessing: The audio data recorded by the terminal is converted into a digital signal and subjected to noise reduction processing. This improves the accuracy of the analysis.

[0124] 3. Data Transmission: The terminal sends pre-processed audio data to the server via the internet. Since stable communication is required, Wi-Fi or mobile data communication is used.

[0125] 4. Speech Analysis: The server analyzes the received speech data and extracts features. Specifically, the Librosa library is used to perform spectrogram analysis and MFCC calculations.

[0126] 5. Feature Mapping: The server uses the extracted features to map them to animal-specific vocal patterns. This allows for the estimation of the animal's intended message.

[0127] 6. Translation: The server infers the intent and translates it into natural language. For example, a message like "There are lots of fish" is generated.

[0128] 7. Response Generation: The server generates a response based on user input and converts it into an animal sound format. For example, if the user asks "Where should I go?", the corresponding animal sound will be generated.

[0129] 8. Output: The terminal receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[0130] Specific example

[0131] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[0132] Examples of specific prompt statements are as follows:

[0133] The user instructed the device to "record dolphin sounds."

[0134] The user asks the server, "Tell me where I should go."

[0135] This system will enable smoother communication with animals.

[0136] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0137] Step 1:

[0138] The user enters a voice command.

[0139] Specific action: The user speaks into the smartphone's microphone and says, "Please record dolphin sounds."

[0140] Input: User voice command.

[0141] Output: Audio input data.

[0142] Step 2:

[0143] The device records the animal's sounds.

[0144] Specific operation: The microphone built into the device records dolphin sounds according to the user's instructions.

[0145] Input: Voice input data.

[0146] Output: Recorded animal sound data.

[0147] Step 3:

[0148] The device converts the recorded data into a digital signal and removes noise.

[0149] Specific operation: The device converts the recorded data into a digital signal and performs noise reduction processing using Audacity or the built-in DSP (Digital Signal Processor).

[0150] Input: Recorded animal sound data.

[0151] Output: De-noised digital audio data.

[0152] Step 4:

[0153] The terminal sends the pre-processed data to the server.

[0154] Specific operation: The device uses Wi-Fi or mobile data communication to send pre-processed audio data to the server via the internet.

[0155] Input: Denoised digital audio data.

[0156] Output: Audio data sent to the server.

[0157] Step 5:

[0158] The server analyzes the audio data and extracts features.

[0159] Specific operation: The server uses Python libraries such as Librosa to perform spectrogram analysis of audio data and calculate Mel-frequency cepstrum coefficients (MFCCs) to extract data features.

[0160] Input: Audio data sent to the server.

[0161] Output: Extracted speech feature data.

[0162] Step 6:

[0163] The server maps the extracted features to the communication format of each animal.

[0164] Specific operation: The server identifies what the animal sounds mean based on features that match animal sound patterns it has learned in advance.

[0165] Input: Extracted speech feature data.

[0166] Output: Data mapped to a communication format.

[0167] Step 7:

[0168] The server infers the intent and translates it into natural language.

[0169] Specific operation: The server estimates the intent and translates it into natural language using a generative AI model (e.g., GPT-4).

[0170] Input: Data mapped to a communication format.

[0171] Output: Translated natural language message.

[0172] Step 8:

[0173] The server generates a response based on the user's input.

[0174] Specific operation: When the user types "Tell me where I should go," the system generates the corresponding animal sound.

[0175] Input: User's question data.

[0176] Output: Generated animal sound data.

[0177] Step 9:

[0178] The terminal receives the response data and presents the result to the user.

[0179] Specific operation: The terminal receives the translation results and sound data sent from the server, and displays them on the screen or communicates them to the user via text or voice.

[0180] Input: Translated natural language message and generated sound data.

[0181] Output: Result data presented to the user.

[0182] Step 10:

[0183] The device plays the generated animal sounds back to life.

[0184] Specific operation: The device's speaker is used to play the generated animal sounds towards the animal.

[0185] Input: Generated sound data.

[0186] Output: Playback of animal sounds.

[0187] (Application Example 1)

[0188] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0189] Traditional pet owners often struggled to accurately understand their pets' needs, particularly in selecting the right timing and type of pet food. Furthermore, responding quickly to a pet's health issues and needs relied heavily on the owner's experience and observation skills, leading to misunderstandings and delays. This created a growing demand for efficient and reliable methods to support a comfortable life for pets.

[0190] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0191] In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format specific to each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, and information communication processing means for ordering food based on the translated data. This enables pet owners to accurately understand their pets' needs and order the right food at the right time.

[0192] "Means for recording animal sounds" refers to a device or system for capturing sounds emitted by animals and recording them as electrical signals.

[0193] "Means for converting recorded vocalizations into digital signals" refers to a device or process for converting recorded analog audio signals into a digital format.

[0194] "Means of removing noise from digital signals" refers to a process or device that removes unwanted background noise and interference from digitized audio signals, thereby improving the clarity of the audio.

[0195] "Means for transmitting pre-processed audio data to a server" refers to a communication device or module for transferring noise-removed audio data to a central server via a network.

[0196] "Means for analyzing audio data and extracting features" refers to algorithms or software for analyzing transmitted audio data and identifying specific patterns or characteristics.

[0197] "Means for mapping extracted features to animal-specific communication formats" refers to a system or process that associates identified vocal patterns with communication formats determined for each animal.

[0198] "Means for translating mapped data into natural language" refers to software or algorithms for converting data mapped to a communication format into a natural language that is understandable to humans.

[0199] "Means for transmitting translated data and generated animal sounds to a terminal" refers to a communication function for transferring data translated into natural language and generated animal sounds to a corresponding terminal.

[0200] "Means for playing back data sent to a terminal" refers to a device or application for outputting received data as audio or text.

[0201] "Information and communication processing means for ordering food based on translated data" refers to a system or program for executing the process of ordering pet food from an online store or the like based on the analyzed pet's requests.

[0202] This invention provides a system that enables communication between humans and animals by analyzing and translating animal sounds. Specifically, it realizes a system that allows pet owners to accurately understand their pets' needs and order appropriate pet food accordingly.

[0203] System Configuration

[0204] 1. Terminal

[0205] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, noise reduction functions, and communication functions for sending data to the server. It also has playback functions for playing back data sent from the server and information communication processing functions for ordering food.

[0206] 2. Server

[0207] The server is a computer system that performs data analysis and translation in a central location. This server is equipped with an audio data analysis engine, a function to map identified features to the communication format of each animal, a mechanism for translating into natural language, and an information and communication processing device for ordering pet food as needed.

[0208] Specific processing flow

[0209] 1. Recording

[0210] The user gives a voice command to the device saying, "Please record my pet's barking."

[0211] The device records pet sounds and converts the analog audio signal into a digital format.

[0212] 2. Preprocessing

[0213] The terminal performs a process to remove noise from the recorded digital audio signal.

[0214] 3. Data transmission

[0215] The pre-processed audio data is sent to the server via Wi-Fi or mobile data communication.

[0216] 4. Voice Analysis

[0217] The server analyzes the received audio data and extracts specific patterns and features. For example, spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs) are performed.

[0218] 5. Feature Mapping

[0219] Using the extracted features, we map animal-specific vocal patterns to communication formats.

[0220] 6. Translation

[0221] The server translates the mapped data into natural language. For example, a message like "I'm hungry" is generated.

[0222] 7. Response generation

[0223] When a user types "Order food," the server processes the order for pet food and sends an order completion notification to the device.

[0224] 8. Output

[0225] The terminal displays the translation results received from the server and the order processing results to the user, and plays them back according to the necessary instructions.

[0226] Hardware and software used

[0227] Smartphone microphone and voice processing function: Record, convert, and remove noise from animal sounds.

[0228] Wi-Fi or mobile data: Send pre-processed data to the server.

[0229] Server analysis engine: Performs speech data analysis, feature extraction, mapping, and translation.

[0230] API communication: Based on the translated data, it performs information communication processing to order food.

[0231] Specific example

[0232] When a user records their pet's barking, the device starts recording, converts it to a digital signal, and removes noise. The data is then sent to a server where voice analysis and feature extraction are performed. The extracted features are mapped as the animal's request and translated as "I'm hungry." When the user instructs "Order food," the server processes the order, and an order completion notification is displayed on the user's device.

[0233] Example of a prompt

[0234] "Create a smartphone application that records pet sounds and orders pet food from a delivery service based on the analysis results. The user interface will consist of three buttons: a recording button, a display of analysis results, and a food order button. The recorded data should be sent to a server, and the analysis results should be received and displayed. Order processing should be done using an API, and the success / failure status should be displayed. Please provide a specific code example along with the output."

[0235] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0236] Step 1:

[0237] The user instructs the device to "record my pet's barking." The device then captures the barking and inputs it as an analog audio signal. The device's audio processing function then activates to convert the input barking into a digital signal. As a result, the analog audio is output as a digital signal.

[0238] Step 2:

[0239] The terminal removes noise from the converted digital audio signal. It takes a digital signal as input and applies a noise reduction algorithm. As a result, clear audio data with noise removed is output. Specifically, the process filters out unwanted information such as high frequencies and background noise.

[0240] Step 3:

[0241] The terminal sends pre-processed audio data to the server via the internet. It takes clear audio data as input and transfers it to the server using Wi-Fi or mobile data communication. The server receives this data and is ready for audio analysis.

[0242] Step 4:

[0243] The server analyzes the received audio data and extracts specific patterns and features. It takes audio data as input and performs feature extraction using algorithms such as spectrogram analysis and Mel-frequency cepstrum coefficients (MFCC). The extracted features are output as data, and the process proceeds to the next step.

[0244] Step 5:

[0245] The server uses extracted features to map animal-specific vocal patterns to communication formats. It takes feature data as input and performs mapping according to the communication format rules for each animal. This results in outputting data corresponding to the pet's intentions.

[0246] Step 6:

[0247] The server translates the mapped data into natural language. It takes formatted data as input and applies a natural language processing algorithm to perform the translation. As a result, it outputs a natural language message such as "I'm hungry."

[0248] Step 7:

[0249] The user enters "Order food" into the device. The device takes the user's instruction as input and sends it to the server. The server receives the instruction and prepares to place the pet food order via API.

[0250] Step 8:

[0251] The server performs information and communication processing to order food based on the translated data. It takes order data as input and sends an order request to the online store using an API. If the order is successful, confirmation data is output and sent to the terminal.

[0252] Step 9:

[0253] The terminal receives confirmation data sent from the server and presents the user with an order completion notification. It takes the confirmation data as input and notifies the user through display and audio notifications. This allows the user to confirm that the order has been successfully completed.

[0254] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0255] The system of the present invention enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. Specific embodiments of the system are shown below.

[0256] System Configuration

[0257] 1. Terminal

[0258] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, a communication function for sending data to a server, and an emotion engine that recognizes the user's emotions.

[0259] 2. Server

[0260] The server is a computer system that centrally performs data analysis, translation, and emotional data analysis. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication formats of each animal, a mechanism for translating into natural language, and a system for receiving and analyzing emotional data transmitted from the emotional engine.

[0261] Processing flow

[0262] 1. Input

[0263] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[0264] 2. Preprocessing

[0265] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[0266] 3. Data transmission

[0267] The device sends pre-processed audio data and user emotion data analyzed by the emotion engine to a server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[0268] 4. Voice Analysis

[0269] The server analyzes the received audio data and extracts features. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0270] 5. Feature Mapping

[0271] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[0272] 6. Translation

[0273] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[0274] 7. Response generation

[0275] The server receives input and emotion data from the user and converts it into an animal sound format. If the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates a relaxing sound data such as "There are lots of fish if you head east."

[0276] 8. Output

[0277] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[0278] Specific example

[0279] When analyzing the sound of a dolphin, the user gives an instruction to the terminal, saying "Please record the sound of a dolphin." The terminal records the sound, converts it into a digital signal, and removes noise. Then, the preprocessed data and the user's emotion data analyzed by the emotion engine are sent to the server, and the voice data is analyzed. The server extracts the characteristics of the dolphin's sound and translates the message "There are many fish" into natural language based on it. Also, when the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates an appropriate dolphin sound for that question and provides information to relax the user.

[0280] The above is a specific embodiment of the system of the present invention. With this system, not only does the communication with animals become smooth, but also an appropriate interaction considering the user's emotional state can be realized.

[0281] The processing flow will be described below.

[0282] Step 1:

[0283] The user speaks to the terminal, saying "Please record the sound of a dolphin." The terminal uses the voice recognition function to detect the user's command and enters the recording mode.

[0284] Step 2:

[0285] The terminal uses the built-in microphone to record the surrounding sounds. The recording time continues for the period set by the user or until manually stopped.

[0286] Step 3:

[0287] The terminal converts the analog voice data recorded into a digital signal. An AD converter (analog-digital converter) is used for this.

[0288] Step 4:

[0289] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[0290] Step 5:

[0291] Simultaneously, the device analyzes the user's voice, facial expressions, and gestures using an emotion engine to recognize their emotional state. It analyzes tone and speed from the voice, facial movements from the expressions, and body movements from the gestures.

[0292] Step 6:

[0293] The terminal packages pre-processed audio data and user emotion data analyzed by the emotion engine into data packets.

[0294] Step 7:

[0295] The terminal sends packaged data to the server via the internet. Stable communication is required.

[0296] Step 8:

[0297] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[0298] Step 9:

[0299] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0300] Step 10:

[0301] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[0302] Step 11:

[0303] The server estimates the intention of the animal from the mapped data and translates the intention into natural language that can be understood by the user. For example, a message such as "There are a lot of fish around" is generated.

[0304] Step 12:

[0305] The server receives the input from the user and the emotion data, and converts it into the format of the animal's vocalization based on this. For example, when the user asks excitedly "Where should I go?", a relaxing message "There are many fish when you go east" is generated using the vocalization of the dolphin.

[0306] Step 13:

[0307] The server transmits the generated vocalization data and the translated text data to the terminal. The terminal analyzes the received data and prepares for playback.

[0308] Step 14:

[0309] The terminal plays the generated vocalization data and simultaneously displays the translated text information to the user. The user can communicate directly with the animal through the terminal.

[0310] (Example 2)

[0311] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".

[0312] Traditional methods of communicating with animals have the problem that it is difficult for users to understand the meaning of animal sounds. Furthermore, because communication is conducted without considering the user's emotional state, it can be stressful for the user. In addition, the ability to properly analyze and translate animal sounds is not yet fully realized. Therefore, there is a need to develop a system that accurately grasps the intentions of animals and conveys them to the user in an easily understandable way.

[0313] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0314] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the extracted features to the communication format of each animal, and means for the server to analyze the user's emotional data. This makes it possible to analyze animal sounds and generate appropriate responses based on the user's emotional state.

[0315] "Means for recording animal sounds" refers to devices that collect sounds emitted by animals using acoustic sensors such as microphones.

[0316] "Means for converting recorded animal sounds into digital signals" refers to devices or software that convert collected animal sound audio data from analog to digital format through digital signal processing.

[0317] "Methods for removing noise from digital signals" refer to filtering algorithms and software used to remove unwanted background noise and other unwanted sounds from converted digital signals.

[0318] "Means for transmitting pre-processed audio data to a server" refers to communication devices and protocols for transferring noise-reduced audio data to a server using internet communication or wireless communication.

[0319] "Methods for analyzing audio data and extracting features" refer to software that performs analysis on audio data on a server to extract important features such as the frequency characteristics and temporal characteristics of the audio.

[0320] "Means for mapping extracted features to animal-specific communication formats" refers to algorithms or software that associate extracted vocal features with animal-specific vocalization patterns.

[0321] "Means for translating mapped data into natural language" refers to software that converts data obtained based on animal communication formats into natural language that humans can understand.

[0322] "Means for a server to analyze user emotional data" refers to software or algorithms that allow a server to analyze emotional data received from a user and understand the user's emotional state.

[0323] "Means for transmitting generated sound data to a terminal" refers to communication devices and protocols for transferring animal sound data generated by a server to a terminal using internet communication or wireless communication.

[0324] "Means for playing back data transmitted to a terminal" refers to output devices such as speakers or playback software that play back the sound data and translation data received by the terminal to the user or animal.

[0325] This invention describes a system that enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. This system mainly includes the following components.

[0326] System Configuration

[0327] 1. Terminal

[0328] It consists of the user's smartphone or dedicated device. This device has the following functions:

[0329] Microphone: An acoustic sensor used to record animal sounds.

[0330] Audio processing function: A digital signal processing (DSP) algorithm that converts recorded audio data into a digital signal and removes noise.

[0331] Communication functions: Wi-Fi and mobile data communication functions for sending pre-processed voice data and user emotion data to the server.

[0332] Emotion engine: Software used to recognize and analyze a user's emotions.

[0333] 2. Server

[0334] A central computer system that performs data analysis, translation, and sentiment data analysis. It includes the following functions:

[0335] Speech analysis engine: Software that analyzes transmitted speech data and performs spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculations to extract features.

[0336] Feature extraction algorithm: An algorithm that associates extracted features with the vocalization patterns of different animals.

[0337] Natural language translation system: Software that estimates the intentions of animals and translates those intentions into natural language.

[0338] Emotional Data Analysis System: A system that analyzes user emotional data transmitted from an emotional engine and generates responses based on that analysis.

[0339] System processing flow

[0340] The system's processing flow is as follows: First, the user issues a voice command to the terminal, initiating the recording of animal sounds. The recorded voice data is converted into a digital signal and subjected to noise reduction processing. Subsequently, the pre-processed voice data and the user's emotional data are sent from the terminal to the server. The server analyzes the voice data and maps it to animal sound patterns. Next, these patterns are translated into natural language, and the content is presented to the user. In addition, an appropriate response is generated according to the user's emotional state and communicated to the animal using the playback function.

[0341] Specific example

[0342] When analyzing dolphin sounds, the user instructs the device to "record dolphin sounds." The device records the sounds and removes noise using a DSP algorithm. The cleared audio data and the user's emotional data, analyzed by the emotion engine, are then sent to the server. The server analyzes the audio data using spectrogram analysis and MFCC calculations to extract dolphin-specific sound patterns. Based on these patterns, it translates the message "There are lots of fish" into natural language. If the user asks "Where should I go?" and the emotion engine detects an excited state, the server generates appropriate dolphin sounds and provides information to help the user relax.

[0343] Example of a prompt

[0344] The following are specific examples of voice commands that users can issue to their devices.

[0345] "Please record the sounds of dolphins."

[0346] "Please tell me the next destination."

[0347] "Tell me what the dolphins are feeling right now."

[0348] The above describes a specific embodiment of the system of the present invention. This system facilitates smooth communication with animals and enables appropriate interaction that takes into account the user's emotional state.

[0349] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0350] Step 1:

[0351] Input: The user issues a voice command to the device saying, "Please record the sound of a dolphin."

[0352] Operation: The user speaks voice commands into their smartphone or a dedicated device.

[0353] Output: The terminal receives the recording command and begins preparing to record the animal sounds.

[0354] Step 2:

[0355] Input: The terminal receives a recording command from the user.

[0356] Operation: The device's microphone activates and records dolphin sounds. The acoustic sensor inside the smartphone is used during this process.

[0357] Output: Generates recorded analog audio data.

[0358] Step 3:

[0359] Input: Recorded analog audio data.

[0360] Operation: The terminal converts this analog audio data into a digital signal. Digital signal processing (DSP) algorithms are used in this process.

[0361] Output: Generates audio data converted into a digital signal.

[0362] Step 4:

[0363] Input: Audio data converted to a digital signal.

[0364] Operation: The terminal performs noise reduction processing on this digital signal. Unnecessary background noise and static are filtered out.

[0365] Output: Generates clear audio data with noise reduction.

[0366] Step 5:

[0367] Input: Denoised audio data and user emotion data analyzed by the emotion engine.

[0368] Operation: The device collects this data and sends it to the server via internet communication. Wi-Fi or mobile data communication is used for this purpose.

[0369] Output: Emotional data and audio data sent to the server.

[0370] Step 6:

[0371] Input: Audio data sent to the server.

[0372] Operation: The server analyzes audio data using an audio analysis engine. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0373] Output: Extracted speech feature data.

[0374] Step 7:

[0375] Input: Extracted speech feature data.

[0376] Operation: The server maps this feature data to animal-specific vocal patterns. The data is then transformed into dolphin-specific communication patterns.

[0377] Output: Mapped communication data.

[0378] Step 8:

[0379] Input: Mapped communication data.

[0380] Operation: The server translates this into natural language. For example, it can generate a specific message like "There are lots of fish" from the sound of dolphins.

[0381] Output: Translated natural language message.

[0382] Step 9:

[0383] Input: Server-generated natural language messages and user sentiment data.

[0384] Operation: The server generates appropriate animal sound formats based on user input and emotional data. For example, if it detects that the user is agitated, it will generate sound data to help them relax.

[0385] Output: Audio data converted to animal sound format.

[0386] Step 10:

[0387] Input: Translation results and sound format data sent from the server.

[0388] Operation: The device receives this data and notifies the user. Furthermore, it transmits the sounds generated using the device's playback function to the animals.

[0389] Output: Translated message presented to the user and played sound data.

[0390] The above outlines the specific program processing flow of this system. Through the specific actions at each step, smooth communication between animals and humans is achieved, and interaction that takes into account the user's emotional state becomes possible.

[0391] (Application Example 2)

[0392] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0393] Traditional animal communication systems focus on analyzing animal sounds and translating them into natural language, but this has limitations in real-time interaction between humans and animals and in ensuring safety. Furthermore, there is a lack of means for users to detect their pet's emotional state or abnormal behavior. This makes it difficult to respond quickly when a pet suddenly becomes agitated or makes suspicious noises, potentially leading to safety issues within the home.

[0394] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format for each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, means for analyzing the emotional state of the animal using the recorded audio data, and means for sending an alert to the user based on the analyzed emotional state. This makes it possible not only to communicate smoothly with animals, but also to quickly detect the emotional state and abnormal behavior of pets, thereby enhancing safety in the home.

[0395] Definitions of important words

[0396] "Means for recording animal sounds" refers to a device or system for electronically capturing animal sounds along with ambient noise and saving them as digital audio files.

[0397] "Means for converting recorded vocalizations into digital signals" refers to a device or technology for converting recorded analog audio data into a digital format.

[0398] "Means for removing noise from digital signals" refers to devices or algorithms for removing unwanted background noise and other unwanted sounds contained in digital audio data, thereby improving the accuracy of the analysis.

[0399] "Means for transmitting pre-processed audio data to a server" refers to a device or system for transmitting audio data that has undergone pre-processing, such as noise reduction, to a central server via the internet or other communication means.

[0400] "Means for analyzing audio data and extracting features" refers to technologies or devices for analyzing audio data and extracting features such as specific frequencies or sound patterns.

[0401] "Means for mapping extracted features to animal-specific communication formats" refers to a device or algorithm for converting extracted vocal features into data corresponding to the communication format specific to that animal.

[0402] "Means for translating mapped data into natural language" refers to devices or technologies for converting data representing animal intentions and emotions into natural language that humans can understand.

[0403] "Means for transmitting translated data and generated sound data to a terminal" refers to a device or system for transmitting data translated into natural language and, if necessary, generated animal sound data to a user's terminal.

[0404] "Means for playing back data transmitted to a terminal" refers to a device or application that allows the user to visually or aurally confirm transmitted audio data or translation results.

[0405] "Means for analyzing an animal's emotional state using recorded audio data" refers to a device or algorithm for analyzing recorded audio data to estimate the animal's current emotional state (e.g., excitement, tension, relaxation).

[0406] "Means for sending alerts to users based on analyzed emotional states" refers to a device or system for sending a warning notification to a user when an abnormal emotional state of an animal is detected.

[0407] Modes for carrying out the invention

[0408] System Configuration

[0409] The system of this invention mainly consists of terminals and servers. This system makes it possible to analyze animal sounds in real time and provide useful information to the user.

[0410] 1. Terminal

[0411] The device consists of the user's smartphone or dedicated device. The device has the following functions:

[0412] Microphone for recording animal sounds

[0413] Audio processing function that converts recorded animal sounds into digital signals.

[0414] A processing function that removes noise from digital signals.

[0415] A communication function that sends pre-processed audio data to the server.

[0416] These features allow for the real-time capture of animal sounds, pre-processing, and transmission of the data to a server.

[0417] 2. Server

[0418] A server is a computer system that centrally performs data analysis, translation, and sentiment data analysis. The server has the following functions:

[0419] Voice data analysis engine

[0420] Feature extraction algorithm

[0421] A system that maps to the communication format of each animal.

[0422] A mechanism for translating into natural language

[0423] Generative AI models for developing emotional engines

[0424] This allows the server to analyze the received audio data and extract its features. Based on these extracted features, it classifies the sounds as animal calls and translates them into natural language. It can also analyze user input and convert their intent into animal calls.

[0425] Analysis and alert function for animal emotional states

[0426] The system also has a function to analyze the emotional state of animals using recorded audio data. This function estimates the animal's current emotions (e.g., excitement, tension, relaxation) from the analyzed vocalization data. It then sends alerts to the user's device, notifying them in real time of abnormal behavior or suspicious vocalizations. This allows the user to take necessary actions quickly. For example, if a pet suddenly becomes agitated or makes a suspicious vocalization, the device can send an alert saying "Suspicious vocalization detected," prompting the user to take immediate action.

[0427] Specific example

[0428] Suppose a user is checking on their pet at home via their smartphone, and the pet suddenly makes an unusual noise. In this case, the device records the noise, converts it to a digital signal, removes noise, and sends it to a server. The server analyzes the audio data, extracts its characteristics, and estimates the pet's emotional state (e.g., fear, excitement). If an abnormality is detected, an alert is sent to the user. The user receives the alert and can take appropriate action quickly.

[0429] Example of a prompt

[0430] Design a model that extracts pet vocalization characteristics from audio data and predicts suspicious vocalizations and emotional states. Use MFCC for audio data preprocessing and build the model using TENSORFLOW® / Keras.

[0431] The above describes the embodiments for carrying out this invention. This not only facilitates smooth communication with animals but also enhances safety within the home.

[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0433] Program processing flow

[0434] Processing flow of the system program that implements the application example

[0435] Step 1:

[0436] The device records animal sounds based on user instructions. The input is an audio signal obtained through a recording device, and the output is a digital audio file. This step includes specific actions to capture animal sounds along with ambient sounds.

[0437] Step 2:

[0438] The terminal converts the recorded audio signal into a digital signal. The input is an analog audio signal, and the output is audio data converted into a digital format (e.g., PCM format). This conversion facilitates subsequent data processing.

[0439] Step 3:

[0440] The terminal removes noise from the digital audio data. The input is digital audio data, and the output is audio data that has been denoised. The specific action performed in this step is to improve the accuracy of the analysis by removing unwanted background noise using a filtering algorithm.

[0441] Step 4:

[0442] The terminal sends pre-processed audio data to the server. The input is denoised digital audio data, and the output is audio data transmitted over the internet. The specific actions in this step include transferring the data to the server via a communication protocol.

[0443] Step 5:

[0444] The server analyzes the received audio data and extracts features. The input is audio data transmitted from the terminal, and the output is audio feature quantities (e.g., MFCC). The analysis engine is used to process the audio waveform and perform the actual calculations to extract features.

[0445] Step 6:

[0446] The server uses extracted features to map them to the communication formats of each animal. The input is the features, and the output is data mapped to the animal's unique vocal patterns. Based on the features, a specific algorithm is implemented to estimate the animal's intentions and emotions.

[0447] Step 7:

[0448] The server translates mapped data into natural language. The input is animal vocalization pattern data, and the output is a text message translated into natural language. This translation process uses a generative AI model to convert what the animals are trying to communicate into human language.

[0449] Step 8:

[0450] The server sends the translated data and generated sound data to the terminal. The input is natural language text and sound data, and the output is data sent to the user's terminal via the internet. This allows the user to see what the animals are trying to communicate in real time.

[0451] Step 9:

[0452] The device plays back data received from the server. Input is natural language text and animal sound data, and output is audio and text presented to the user. Specific actions include displaying text messages and playing animal sound data through the speaker.

[0453] Step 10:

[0454] The server analyzes the emotional state of an animal using recorded audio data. The input is recorded audio data, and the output is an estimated result of the animal's emotional state. An emotion engine is used to perform specific data calculations to estimate emotional states such as excitement and relaxation.

[0455] Step 11:

[0456] Based on the analyzed emotional state, the server sends an alert to the user. The input is the estimated result of the animal's emotional state, and the output is the alert message sent to the user. This message allows the user to take action early.

[0457] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0458] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0459] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0460] [Second Embodiment]

[0461] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0462] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0463] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0464] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0465] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0467] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0468] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0469] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0470] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0471] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0472] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0473] The present invention's system analyzes and translates animal sounds, enabling communication between humans and animals. Specific embodiments of the system are described below.

[0474] System Configuration

[0475] 1. Terminal

[0476] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, and a communication function for sending data to a server.

[0477] 2. Server

[0478] The server is a computer system that performs data analysis and translation in a central location. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication format of each animal, and a mechanism for translating into natural language.

[0479] Processing flow

[0480] 1. Input

[0481] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[0482] 2. Preprocessing

[0483] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[0484] 3. Data transmission

[0485] The terminal sends pre-processed audio data to the server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[0486] 4. Voice Analysis

[0487] The server analyzes the received audio data and extracts features. Specifically, this involves spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs).

[0488] 5. Feature Mapping

[0489] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[0490] 6. Translation

[0491] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[0492] 7. Response generation

[0493] The server receives input from the user and converts it into an animal sound format. If the user asks "Tell me where to go," the corresponding animal sound is generated.

[0494] 8. Output

[0495] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[0496] Specific example

[0497] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[0498] The above describes a specific embodiment of the system of the present invention. This system enables smooth communication with animals.

[0499] The following describes the processing flow.

[0500] Step 1:

[0501] The user speaks into the device and says, "Please record dolphin sounds." The device uses its voice recognition function to detect the user's command and enters recording mode.

[0502] Step 2:

[0503] The device uses its built-in microphone to record ambient sounds. Recording continues for a period set by the user, or until manually stopped.

[0504] Step 3:

[0505] The device converts the recorded analog audio data into a digital signal. This is done using an AD converter (analog-to-digital converter).

[0506] Step 4:

[0507] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[0508] Step 5:

[0509] The terminal compresses the pre-processed audio data and packages it into data packets for efficient transmission to the server.

[0510] Step 6:

[0511] The terminal sends packaged data to the server via the internet. Stable communication is required.

[0512] Step 7:

[0513] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[0514] Step 8:

[0515] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0516] Step 9:

[0517] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[0518] Step 10:

[0519] The server estimates the animal's intentions from the mapped data and translates those intentions into natural language that the user can understand. For example, it might generate a message like, "There are lots of fish around."

[0520] Step 11:

[0521] The user speaks to the device saying, "Tell me where I should go." The device converts this voice command into text and sends it to the server.

[0522] Step 12:

[0523] The server analyzes the user's question and generates appropriate animal sounds based on its content. For example, it converts information like "There are many fish if you head east" into a dolphin sound format.

[0524] Step 13:

[0525] The server sends the generated sound data to the terminal. The terminal analyzes the received data and prepares for playback.

[0526] Step 14:

[0527] The device plays generated animal sound data and simultaneously displays translated text information to the user. The user can communicate directly with the animal through the device.

[0528] (Example 1)

[0529] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0530] Traditional methods of communication with animals have been limited, making it difficult to accurately analyze and translate animal calls. Furthermore, the lack of a method to analyze animal calls and translate them into human natural language, and vice versa, prevented two-way communication between animals and humans. This resulted in insufficient communication between animals in captivity and research environments, posing a serious problem.

[0531] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0532] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the features to communication formats specific to each animal, and means for translating the mapped data into natural language. This makes it possible to accurately analyze animal calls and translate their intent into natural language. Furthermore, it becomes possible to analyze user input and convert natural language into animal call formats, enabling two-way communication between animals and humans.

[0533] "Means for recording animal sounds" refers to devices or systems that input the sounds of animals specified by the user as audio.

[0534] "Means for converting recorded animal sounds into digital signals" refers to devices or systems for converting analog audio data into a digital format.

[0535] "Methods for removing noise from digital signals" refer to devices and systems that improve the quality of digital audio data by removing unwanted background noise and interference sounds.

[0536] "Means for sending pre-processed audio data to a server" refers to a device or system for sending processed audio data to a remote server.

[0537] "Means for analyzing audio data and extracting features" refers to devices or systems for analyzing audio data and extracting important feature patterns.

[0538] "Means for mapping extracted features to animal-specific communication formats" refers to devices or systems that map the features of analyzed audio data to the communication formats of specific animals.

[0539] "Means for translating mapped data into natural language" refers to devices or systems that convert data corresponding to animal communication patterns into natural language that humans can understand.

[0540] "Means for transmitting translated data and generated sound data to a terminal" refers to devices or systems for transmitting data translated on a server and animal sound data to a user's terminal.

[0541] "Means for playing back data sent to a terminal" refers to devices or systems for playing back data received on a user's terminal in audio or text format.

[0542] "Means for users to input voice commands" refers to devices or systems that allow users to give instructions to a system using their voice.

[0543] "Means for a server to analyze audio data and extract features" refers to devices or systems that allow a server to analyze audio data and extract important feature patterns.

[0544] "Means by which a server maps animal vocalization patterns" refers to devices or systems that allow a server to map the characteristics of analyzed audio data to the communication patterns of specific animals.

[0545] "Means for a server to translate intentions into natural language" refers to a device or system in which a server estimates an animal's intentions and converts them into natural language that humans can understand.

[0546] "Means by which a server generates a response" refers to a device or system that allows a server to generate corresponding animal sounds based on user input.

[0547] "Means by which a terminal receives response data and presents the results to the user" refers to a device or system that allows a terminal to receive response data from a server and display or play it aloud for the user.

[0548] "Means for a terminal to reproduce animal sounds" refers to a device or system that allows a terminal to reproduce generated animal sounds as audio.

[0549] The system of the present invention analyzes and translates animal sounds to enable smooth communication between humans and animals. Specific embodiments of the system are described below.

[0550] System Configuration

[0551] 1. Terminal

[0552] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, and communication functions for sending data to a server. Specific hardware includes the smartphone's built-in microphone, an external microphone, and Wi-Fi or mobile data communication for communication. Software used includes audio recording applications and noise reduction software such as Audacity.

[0553] 2. Server

[0554] The server is a central computer system that performs data analysis and translation. This server implements an audio data analysis engine, feature extraction algorithms, a system for mapping to animal-specific communication formats, and a mechanism for translating into natural language. Specific software includes spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculation using the Python Librosa library. Furthermore, a generative AI model (e.g., GPT-4) is responsible for natural language translation.

[0555] Processing flow

[0556] The system operates using the following steps:

[0557] 1. Input: The user issues a voice command to the device, such as, "Please record dolphin sounds." For example, this could be done while observing dolphins and recording their sounds.

[0558] 2. Preprocessing: The audio data recorded by the terminal is converted into a digital signal and subjected to noise reduction processing. This improves the accuracy of the analysis.

[0559] 3. Data Transmission: The terminal sends pre-processed audio data to the server via the internet. Since stable communication is required, Wi-Fi or mobile data communication is used.

[0560] 4. Speech Analysis: The server analyzes the received speech data and extracts features. Specifically, the Librosa library is used to perform spectrogram analysis and MFCC calculations.

[0561] 5. Feature Mapping: The server uses the extracted features to map them to animal-specific vocal patterns. This allows for the estimation of the animal's intended message.

[0562] 6. Translation: The server infers the intent and translates it into natural language. For example, a message like "There are lots of fish" is generated.

[0563] 7. Response Generation: The server generates a response based on user input and converts it into an animal sound format. For example, if the user asks "Where should I go?", the corresponding animal sound will be generated.

[0564] 8. Output: The terminal receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[0565] Specific example

[0566] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[0567] Examples of specific prompt statements are as follows:

[0568] The user instructed the device to "record dolphin sounds."

[0569] The user asks the server, "Tell me where I should go."

[0570] This system will enable smoother communication with animals.

[0571] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0572] Step 1:

[0573] The user enters a voice command.

[0574] Specific action: The user speaks into the smartphone's microphone and says, "Please record dolphin sounds."

[0575] Input: User voice command.

[0576] Output: Audio input data.

[0577] Step 2:

[0578] The device records the animal's sounds.

[0579] Specific operation: The microphone built into the device records dolphin sounds according to the user's instructions.

[0580] Input: Voice input data.

[0581] Output: Recorded animal sound data.

[0582] Step 3:

[0583] The device converts the recorded data into a digital signal and removes noise.

[0584] Specific operation: The device converts the recorded data into a digital signal and performs noise reduction processing using Audacity or the built-in DSP (Digital Signal Processor).

[0585] Input: Recorded animal sound data.

[0586] Output: De-noised digital audio data.

[0587] Step 4:

[0588] The terminal sends the pre-processed data to the server.

[0589] Specific operation: The device uses Wi-Fi or mobile data communication to send pre-processed audio data to the server via the internet.

[0590] Input: Denoised digital audio data.

[0591] Output: Audio data sent to the server.

[0592] Step 5:

[0593] The server analyzes the audio data and extracts features.

[0594] Specific operation: The server uses Python libraries such as Librosa to perform spectrogram analysis of audio data and calculate Mel-frequency cepstrum coefficients (MFCCs) to extract data features.

[0595] Input: Audio data sent to the server.

[0596] Output: Extracted speech feature data.

[0597] Step 6:

[0598] The server maps the extracted features to the communication format of each animal.

[0599] Specific operation: The server identifies what the animal sounds mean based on features that match animal sound patterns it has learned in advance.

[0600] Input: Extracted speech feature data.

[0601] Output: Data mapped to a communication format.

[0602] Step 7:

[0603] The server infers the intent and translates it into natural language.

[0604] Specific operation: The server estimates the intent and translates it into natural language using a generative AI model (e.g., GPT-4).

[0605] Input: Data mapped to a communication format.

[0606] Output: Translated natural language message.

[0607] Step 8:

[0608] The server generates a response based on the user's input.

[0609] Specific operation: When the user types "Tell me where I should go," the system generates the corresponding animal sound.

[0610] Input: User's question data.

[0611] Output: Generated animal sound data.

[0612] Step 9:

[0613] The terminal receives the response data and presents the result to the user.

[0614] Specific operation: The terminal receives the translation results and sound data sent from the server, and displays them on the screen or communicates them to the user via text or voice.

[0615] Input: Translated natural language message and generated sound data.

[0616] Output: Result data presented to the user.

[0617] Step 10:

[0618] The device plays the generated animal sounds back to life.

[0619] Specific operation: The device's speaker is used to play the generated animal sounds towards the animal.

[0620] Input: Generated sound data.

[0621] Output: Playback of animal sounds.

[0622] (Application Example 1)

[0623] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0624] Traditional pet owners often struggled to accurately understand their pets' needs, particularly in selecting the right timing and type of pet food. Furthermore, responding quickly to a pet's health issues and needs relied heavily on the owner's experience and observation skills, leading to misunderstandings and delays. This created a growing demand for efficient and reliable methods to support a comfortable life for pets.

[0625] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0626] In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format specific to each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, and information communication processing means for ordering food based on the translated data. This enables pet owners to accurately understand their pets' needs and order the right food at the right time.

[0627] "Means for recording animal sounds" refers to a device or system for capturing sounds emitted by animals and recording them as electrical signals.

[0628] "Means for converting recorded vocalizations into digital signals" refers to a device or process for converting recorded analog audio signals into a digital format.

[0629] "Means of removing noise from digital signals" refers to a process or device that removes unwanted background noise and interference from digitized audio signals, thereby improving the clarity of the audio.

[0630] "Means for transmitting pre-processed audio data to a server" refers to a communication device or module for transferring noise-removed audio data to a central server via a network.

[0631] "Means for analyzing audio data and extracting features" refers to algorithms or software for analyzing transmitted audio data and identifying specific patterns or characteristics.

[0632] "Means for mapping extracted features to animal-specific communication formats" refers to a system or process that associates identified vocal patterns with communication formats determined for each animal.

[0633] "Means for translating mapped data into natural language" refers to software or algorithms for converting data mapped to a communication format into a natural language that is understandable to humans.

[0634] "Means for transmitting translated data and generated animal sounds to a terminal" refers to a communication function for transferring data translated into natural language and generated animal sounds to a corresponding terminal.

[0635] "Means for playing back data sent to a terminal" refers to a device or application for outputting received data as audio or text.

[0636] "Information and communication processing means for ordering food based on translated data" refers to a system or program for executing the process of ordering pet food from an online store or the like based on the analyzed pet's requests.

[0637] This invention provides a system that enables communication between humans and animals by analyzing and translating animal sounds. Specifically, it realizes a system that allows pet owners to accurately understand their pets' needs and order appropriate pet food accordingly.

[0638] System Configuration

[0639] 1. Terminal

[0640] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, noise reduction functions, and communication functions for sending data to the server. It also has playback functions for playing back data sent from the server and information communication processing functions for ordering food.

[0641] 2. Server

[0642] The server is a computer system that performs data analysis and translation in a central location. This server is equipped with an audio data analysis engine, a function to map identified features to the communication format of each animal, a mechanism for translating into natural language, and an information and communication processing device for ordering pet food as needed.

[0643] Specific processing flow

[0644] 1. Recording

[0645] The user gives a voice command to the device saying, "Please record my pet's barking."

[0646] The device records pet sounds and converts the analog audio signal into a digital format.

[0647] 2. Preprocessing

[0648] The terminal performs a process to remove noise from the recorded digital audio signal.

[0649] 3. Data transmission

[0650] The pre-processed audio data is sent to the server via Wi-Fi or mobile data communication.

[0651] 4. Voice Analysis

[0652] The server analyzes the received audio data and extracts specific patterns and features. For example, spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs) are performed.

[0653] 5. Feature Mapping

[0654] Using the extracted features, we map animal-specific vocal patterns to communication formats.

[0655] 6. Translation

[0656] The server translates the mapped data into natural language. For example, a message like "I'm hungry" is generated.

[0657] 7. Response generation

[0658] When a user types "Order food," the server processes the order for pet food and sends an order completion notification to the device.

[0659] 8. Output

[0660] The terminal displays the translation results received from the server and the order processing results to the user, and plays them back according to the necessary instructions.

[0661] Hardware and software used

[0662] Smartphone microphone and voice processing function: Record, convert, and remove noise from animal sounds.

[0663] Wi-Fi or mobile data: Send pre-processed data to the server.

[0664] Server analysis engine: Performs speech data analysis, feature extraction, mapping, and translation.

[0665] API communication: Based on the translated data, it performs information communication processing to order food.

[0666] Specific example

[0667] When a user records their pet's barking, the device starts recording, converts it to a digital signal, and removes noise. The data is then sent to a server where voice analysis and feature extraction are performed. The extracted features are mapped as the animal's request and translated as "I'm hungry." When the user instructs "Order food," the server processes the order, and an order completion notification is displayed on the user's device.

[0668] Example of a prompt

[0669] "Create a smartphone application that records pet sounds and orders pet food from a delivery service based on the analysis results. The user interface will consist of three buttons: a recording button, a display of analysis results, and a food order button. The recorded data should be sent to a server, and the analysis results should be received and displayed. Order processing should be done using an API, and the success / failure status should be displayed. Please provide a specific code example along with the output."

[0670] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0671] Step 1:

[0672] The user instructs the device to "record my pet's barking." The device then captures the barking and inputs it as an analog audio signal. The device's audio processing function then activates to convert the input barking into a digital signal. As a result, the analog audio is output as a digital signal.

[0673] Step 2:

[0674] The terminal removes noise from the converted digital audio signal. It takes a digital signal as input and applies a noise reduction algorithm. As a result, clear audio data with noise removed is output. Specifically, the process filters out unwanted information such as high frequencies and background noise.

[0675] Step 3:

[0676] The terminal sends pre-processed audio data to the server via the internet. It takes clear audio data as input and transfers it to the server using Wi-Fi or mobile data communication. The server receives this data and is ready for audio analysis.

[0677] Step 4:

[0678] The server analyzes the received audio data and extracts specific patterns and features. It takes audio data as input and performs feature extraction using algorithms such as spectrogram analysis and Mel-frequency cepstrum coefficients (MFCC). The extracted features are output as data, and the process proceeds to the next step.

[0679] Step 5:

[0680] The server uses extracted features to map animal-specific vocal patterns to communication formats. It takes feature data as input and performs mapping according to the communication format rules for each animal. This results in outputting data corresponding to the pet's intentions.

[0681] Step 6:

[0682] The server translates the mapped data into natural language. It takes formatted data as input and applies a natural language processing algorithm to perform the translation. As a result, it outputs a natural language message such as "I'm hungry."

[0683] Step 7:

[0684] The user enters "Order food" into the device. The device takes the user's instruction as input and sends it to the server. The server receives the instruction and prepares to place the pet food order via API.

[0685] Step 8:

[0686] The server performs information and communication processing to order food based on the translated data. It takes order data as input and sends an order request to the online store using an API. If the order is successful, confirmation data is output and sent to the terminal.

[0687] Step 9:

[0688] The terminal receives confirmation data sent from the server and presents the user with an order completion notification. It takes the confirmation data as input and notifies the user through display and audio notifications. This allows the user to confirm that the order has been successfully completed.

[0689] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0690] The system of the present invention enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. Specific embodiments of the system are shown below.

[0691] System Configuration

[0692] 1. Terminal

[0693] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, a communication function for sending data to a server, and an emotion engine that recognizes the user's emotions.

[0694] 2. Server

[0695] The server is a computer system that centrally performs data analysis, translation, and emotional data analysis. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication formats of each animal, a mechanism for translating into natural language, and a system for receiving and analyzing emotional data transmitted from the emotional engine.

[0696] Processing flow

[0697] 1. Input

[0698] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[0699] 2. Preprocessing

[0700] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[0701] 3. Data transmission

[0702] The device sends pre-processed audio data and user emotion data analyzed by the emotion engine to a server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[0703] 4. Voice Analysis

[0704] The server analyzes the received audio data and extracts features. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0705] 5. Feature Mapping

[0706] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[0707] 6. Translation

[0708] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[0709] 7. Response generation

[0710] The server receives input and emotion data from the user and converts it into an animal sound format. If the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates a relaxing sound data such as "There are lots of fish if you head east."

[0711] 8. Output

[0712] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[0713] Specific example

[0714] When analyzing dolphin sounds, the user instructs the device to "record dolphin sounds." The device records the sounds, converts them into digital signals, and removes noise. The pre-processed data and the user's emotional data analyzed by the emotion engine are then sent to the server, where the audio data is analyzed. The server extracts the characteristics of the dolphin sounds and uses them to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates an appropriate dolphin sound in response to that question, providing information to help the user relax.

[0715] The above describes a specific embodiment of the system of the present invention. This system not only facilitates smooth communication with animals but also enables appropriate interaction that takes into account the user's emotional state.

[0716] The following describes the processing flow.

[0717] Step 1:

[0718] The user speaks into the device and says, "Please record dolphin sounds." The device uses its voice recognition function to detect the user's command and enters recording mode.

[0719] Step 2:

[0720] The device uses its built-in microphone to record ambient sounds. Recording continues for a period set by the user, or until manually stopped.

[0721] Step 3:

[0722] The device converts the recorded analog audio data into a digital signal. This is done using an AD converter (analog-to-digital converter).

[0723] Step 4:

[0724] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[0725] Step 5:

[0726] Simultaneously, the device analyzes the user's voice, facial expressions, and gestures using an emotion engine to recognize their emotional state. It analyzes tone and speed from the voice, facial movements from the expressions, and body movements from the gestures.

[0727] Step 6:

[0728] The terminal packages pre-processed audio data and user emotion data analyzed by the emotion engine into data packets.

[0729] Step 7:

[0730] The terminal sends packaged data to the server via the internet. Stable communication is required.

[0731] Step 8:

[0732] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[0733] Step 9:

[0734] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0735] Step 10:

[0736] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[0737] Step 11:

[0738] The server estimates the animal's intentions from the mapped data and translates those intentions into natural language that the user can understand. For example, it might generate a message like, "There are lots of fish around."

[0739] Step 12:

[0740] The server receives user input and emotional data, and converts it into animal sound formats. For example, if a user asks excitedly, "Tell me where I should go," the server will use dolphin sounds to generate a calming message such as, "There are lots of fish if you head east."

[0741] Step 13:

[0742] The server sends the generated sound data and translated text data to the terminal. The terminal analyzes the received data and prepares for playback.

[0743] Step 14:

[0744] The device plays generated animal sound data and simultaneously displays translated text information to the user. The user can communicate directly with the animal through the device.

[0745] (Example 2)

[0746] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0747] Traditional methods of communicating with animals have the problem that it is difficult for users to understand the meaning of animal sounds. Furthermore, because communication is conducted without considering the user's emotional state, it can be stressful for the user. In addition, the ability to properly analyze and translate animal sounds is not yet fully realized. Therefore, there is a need to develop a system that accurately grasps the intentions of animals and conveys them to the user in an easily understandable way.

[0748] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0749] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the extracted features to the communication format of each animal, and means for the server to analyze the user's emotional data. This makes it possible to analyze animal sounds and generate appropriate responses based on the user's emotional state.

[0750] "Means for recording animal sounds" refers to devices that collect sounds emitted by animals using acoustic sensors such as microphones.

[0751] "Means for converting recorded animal sounds into digital signals" refers to devices or software that convert collected animal sound audio data from analog to digital format through digital signal processing.

[0752] "Methods for removing noise from digital signals" refer to filtering algorithms and software used to remove unwanted background noise and other unwanted sounds from converted digital signals.

[0753] "Means for transmitting pre-processed audio data to a server" refers to communication devices and protocols for transferring noise-reduced audio data to a server using internet communication or wireless communication.

[0754] "Methods for analyzing audio data and extracting features" refer to software that performs analysis on audio data on a server to extract important features such as the frequency characteristics and temporal characteristics of the audio.

[0755] "Means for mapping extracted features to animal-specific communication formats" refers to algorithms or software that associate extracted vocal features with animal-specific vocalization patterns.

[0756] "Means for translating mapped data into natural language" refers to software that converts data obtained based on animal communication formats into natural language that humans can understand.

[0757] "Means for a server to analyze user emotional data" refers to software or algorithms that allow a server to analyze emotional data received from a user and understand the user's emotional state.

[0758] "Means for transmitting generated sound data to a terminal" refers to communication devices and protocols for transferring animal sound data generated by a server to a terminal using internet communication or wireless communication.

[0759] "Means for playing back data transmitted to a terminal" refers to output devices such as speakers or playback software that play back the sound data and translation data received by the terminal to the user or animal.

[0760] This invention describes a system that enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. This system mainly includes the following components.

[0761] System Configuration

[0762] 1. Terminal

[0763] It consists of the user's smartphone or dedicated device. This device has the following functions:

[0764] Microphone: An acoustic sensor used to record animal sounds.

[0765] Audio processing function: A digital signal processing (DSP) algorithm that converts recorded audio data into a digital signal and removes noise.

[0766] Communication functions: Wi-Fi and mobile data communication functions for sending pre-processed voice data and user emotion data to the server.

[0767] Emotion engine: Software used to recognize and analyze a user's emotions.

[0768] 2. Server

[0769] A central computer system that performs data analysis, translation, and sentiment data analysis. It includes the following functions:

[0770] Speech analysis engine: Software that analyzes transmitted speech data and performs spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculations to extract features.

[0771] Feature extraction algorithm: An algorithm that associates extracted features with the vocalization patterns of different animals.

[0772] Natural language translation system: Software that estimates the intentions of animals and translates those intentions into natural language.

[0773] Emotional Data Analysis System: A system that analyzes user emotional data transmitted from an emotional engine and generates responses based on that analysis.

[0774] System processing flow

[0775] The system's processing flow is as follows: First, the user issues a voice command to the terminal, initiating the recording of animal sounds. The recorded voice data is converted into a digital signal and subjected to noise reduction processing. Subsequently, the pre-processed voice data and the user's emotional data are sent from the terminal to the server. The server analyzes the voice data and maps it to animal sound patterns. Next, these patterns are translated into natural language, and the content is presented to the user. In addition, an appropriate response is generated according to the user's emotional state and communicated to the animal using the playback function.

[0776] Specific example

[0777] When analyzing dolphin sounds, the user instructs the device to "record dolphin sounds." The device records the sounds and removes noise using a DSP algorithm. The cleared audio data and the user's emotional data, analyzed by the emotion engine, are then sent to the server. The server analyzes the audio data using spectrogram analysis and MFCC calculations to extract dolphin-specific sound patterns. Based on these patterns, it translates the message "There are lots of fish" into natural language. If the user asks "Where should I go?" and the emotion engine detects an excited state, the server generates appropriate dolphin sounds and provides information to help the user relax.

[0778] Example of a prompt

[0779] The following are specific examples of voice commands that users can issue to their devices.

[0780] "Please record the sounds of dolphins."

[0781] "Please tell me the next destination."

[0782] "Tell me what the dolphins are feeling right now."

[0783] The above describes a specific embodiment of the system of the present invention. This system facilitates smooth communication with animals and enables appropriate interaction that takes into account the user's emotional state.

[0784] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0785] Step 1:

[0786] Input: The user issues a voice command to the device saying, "Please record the sound of a dolphin."

[0787] Operation: The user speaks voice commands into their smartphone or a dedicated device.

[0788] Output: The terminal receives the recording command and begins preparing to record the animal sounds.

[0789] Step 2:

[0790] Input: The terminal receives a recording command from the user.

[0791] Operation: The device's microphone activates and records dolphin sounds. The acoustic sensor inside the smartphone is used during this process.

[0792] Output: Generates recorded analog audio data.

[0793] Step 3:

[0794] Input: Recorded analog audio data.

[0795] Operation: The terminal converts this analog audio data into a digital signal. Digital signal processing (DSP) algorithms are used in this process.

[0796] Output: Generates audio data converted into a digital signal.

[0797] Step 4:

[0798] Input: Audio data converted to a digital signal.

[0799] Operation: The terminal performs noise reduction processing on this digital signal. Unnecessary background noise and static are filtered out.

[0800] Output: Generates clear audio data with noise reduction.

[0801] Step 5:

[0802] Input: Denoised audio data and user emotion data analyzed by the emotion engine.

[0803] Operation: The device collects this data and sends it to the server via internet communication. Wi-Fi or mobile data communication is used for this purpose.

[0804] Output: Emotional data and audio data sent to the server.

[0805] Step 6:

[0806] Input: Audio data sent to the server.

[0807] Operation: The server analyzes audio data using an audio analysis engine. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0808] Output: Extracted speech feature data.

[0809] Step 7:

[0810] Input: Extracted speech feature data.

[0811] Operation: The server maps this feature data to animal-specific vocal patterns. The data is then transformed into dolphin-specific communication patterns.

[0812] Output: Mapped communication data.

[0813] Step 8:

[0814] Input: Mapped communication data.

[0815] Operation: The server translates this into natural language. For example, it can generate a specific message like "There are lots of fish" from the sound of dolphins.

[0816] Output: Translated natural language message.

[0817] Step 9:

[0818] Input: Server-generated natural language messages and user sentiment data.

[0819] Operation: The server generates appropriate animal sound formats based on user input and emotional data. For example, if it detects that the user is agitated, it will generate sound data to help them relax.

[0820] Output: Audio data converted to animal sound format.

[0821] Step 10:

[0822] Input: Translation results and sound format data sent from the server.

[0823] Operation: The device receives this data and notifies the user. Furthermore, it transmits the sounds generated using the device's playback function to the animals.

[0824] Output: Translated message presented to the user and played sound data.

[0825] The above outlines the specific program processing flow of this system. Through the specific actions at each step, smooth communication between animals and humans is achieved, and interaction that takes into account the user's emotional state becomes possible.

[0826] (Application Example 2)

[0827] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0828] Traditional animal communication systems focus on analyzing animal sounds and translating them into natural language, but this has limitations in real-time interaction between humans and animals and in ensuring safety. Furthermore, there is a lack of means for users to detect their pet's emotional state or abnormal behavior. This makes it difficult to respond quickly when a pet suddenly becomes agitated or makes suspicious noises, potentially leading to safety issues within the home.

[0829] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format for each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, means for analyzing the emotional state of the animal using the recorded audio data, and means for sending an alert to the user based on the analyzed emotional state. This makes it possible not only to communicate smoothly with animals, but also to quickly detect the emotional state and abnormal behavior of pets, thereby enhancing safety in the home.

[0830] Definitions of important words

[0831] "Means for recording animal sounds" refers to a device or system for electronically capturing animal sounds along with ambient noise and saving them as digital audio files.

[0832] "Means for converting recorded vocalizations into digital signals" refers to a device or technology for converting recorded analog audio data into a digital format.

[0833] "Means for removing noise from digital signals" refers to devices or algorithms for removing unwanted background noise and other unwanted sounds contained in digital audio data, thereby improving the accuracy of the analysis.

[0834] "Means for transmitting pre-processed audio data to a server" refers to a device or system for transmitting audio data that has undergone pre-processing, such as noise reduction, to a central server via the internet or other communication means.

[0835] "Means for analyzing audio data and extracting features" refers to technologies or devices for analyzing audio data and extracting features such as specific frequencies or sound patterns.

[0836] "Means for mapping extracted features to animal-specific communication formats" refers to a device or algorithm for converting extracted vocal features into data corresponding to the communication format specific to that animal.

[0837] "Means for translating mapped data into natural language" refers to devices or technologies for converting data representing animal intentions and emotions into natural language that humans can understand.

[0838] "Means for transmitting translated data and generated sound data to a terminal" refers to a device or system for transmitting data translated into natural language and, if necessary, generated animal sound data to a user's terminal.

[0839] "Means for playing back data transmitted to a terminal" refers to a device or application that allows the user to visually or aurally confirm transmitted audio data or translation results.

[0840] "Means for analyzing an animal's emotional state using recorded audio data" refers to a device or algorithm for analyzing recorded audio data to estimate the animal's current emotional state (e.g., excitement, tension, relaxation).

[0841] "Means for sending alerts to users based on analyzed emotional states" refers to a device or system for sending a warning notification to a user when an abnormal emotional state of an animal is detected.

[0842] Modes for carrying out the invention

[0843] System Configuration

[0844] The system of this invention mainly consists of terminals and servers. This system makes it possible to analyze animal sounds in real time and provide useful information to the user.

[0845] 1. Terminal

[0846] The device consists of the user's smartphone or dedicated device. The device has the following functions:

[0847] Microphone for recording animal sounds

[0848] Audio processing function that converts recorded animal sounds into digital signals.

[0849] A processing function that removes noise from digital signals.

[0850] A communication function that sends pre-processed audio data to the server.

[0851] These features allow for the real-time capture of animal sounds, pre-processing, and transmission of the data to a server.

[0852] 2. Server

[0853] A server is a computer system that centrally performs data analysis, translation, and sentiment data analysis. The server has the following functions:

[0854] Voice data analysis engine

[0855] Feature extraction algorithm

[0856] A system that maps to the communication format of each animal.

[0857] A mechanism for translating into natural language

[0858] Generative AI models for developing emotional engines

[0859] This allows the server to analyze the received audio data and extract its features. Based on these extracted features, it classifies the sounds as animal calls and translates them into natural language. It can also analyze user input and convert their intent into animal calls.

[0860] Analysis and alert function for animal emotional states

[0861] The system also has a function to analyze the emotional state of animals using recorded audio data. This function estimates the animal's current emotions (e.g., excitement, tension, relaxation) from the analyzed vocalization data. It then sends alerts to the user's device, notifying them in real time of abnormal behavior or suspicious vocalizations. This allows the user to take necessary actions quickly. For example, if a pet suddenly becomes agitated or makes a suspicious vocalization, the device can send an alert saying "Suspicious vocalization detected," prompting the user to take immediate action.

[0862] Specific example

[0863] Suppose a user is checking on their pet at home via their smartphone, and the pet suddenly makes an unusual noise. In this case, the device records the noise, converts it to a digital signal, removes noise, and sends it to a server. The server analyzes the audio data, extracts its characteristics, and estimates the pet's emotional state (e.g., fear, excitement). If an abnormality is detected, an alert is sent to the user. The user receives the alert and can take appropriate action quickly.

[0864] Example of a prompt

[0865] Design a model that extracts pet vocalization characteristics from audio data and predicts suspicious vocalizations and emotional states. Use MFCC for audio data preprocessing and build the model using TensorFlow / Keras.

[0866] The above describes the embodiments for carrying out this invention. This not only facilitates smooth communication with animals but also enhances safety within the home.

[0867] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0868] Program processing flow

[0869] Processing flow of the system program that implements the application example

[0870] Step 1:

[0871] The device records animal sounds based on user instructions. The input is an audio signal obtained through a recording device, and the output is a digital audio file. This step includes specific actions to capture animal sounds along with ambient sounds.

[0872] Step 2:

[0873] The terminal converts the recorded audio signal into a digital signal. The input is an analog audio signal, and the output is audio data converted into a digital format (e.g., PCM format). This conversion facilitates subsequent data processing.

[0874] Step 3:

[0875] The terminal removes noise from the digital audio data. The input is digital audio data, and the output is audio data that has been denoised. The specific action performed in this step is to improve the accuracy of the analysis by removing unwanted background noise using a filtering algorithm.

[0876] Step 4:

[0877] The terminal sends pre-processed audio data to the server. The input is denoised digital audio data, and the output is audio data transmitted over the internet. The specific actions in this step include transferring the data to the server via a communication protocol.

[0878] Step 5:

[0879] The server analyzes the received audio data and extracts features. The input is audio data transmitted from the terminal, and the output is audio feature quantities (e.g., MFCC). The analysis engine is used to process the audio waveform and perform the actual calculations to extract features.

[0880] Step 6:

[0881] The server uses extracted features to map them to the communication formats of each animal. The input is the features, and the output is data mapped to the animal's unique vocal patterns. Based on the features, a specific algorithm is implemented to estimate the animal's intentions and emotions.

[0882] Step 7:

[0883] The server translates mapped data into natural language. The input is animal vocalization pattern data, and the output is a text message translated into natural language. This translation process uses a generative AI model to convert what the animals are trying to communicate into human language.

[0884] Step 8:

[0885] The server sends the translated data and generated sound data to the terminal. The input is natural language text and sound data, and the output is data sent to the user's terminal via the internet. This allows the user to see what the animals are trying to communicate in real time.

[0886] Step 9:

[0887] The device plays back data received from the server. Input is natural language text and animal sound data, and output is audio and text presented to the user. Specific actions include displaying text messages and playing animal sound data through the speaker.

[0888] Step 10:

[0889] The server analyzes the emotional state of an animal using recorded audio data. The input is recorded audio data, and the output is an estimated result of the animal's emotional state. An emotion engine is used to perform specific data calculations to estimate emotional states such as excitement and relaxation.

[0890] Step 11:

[0891] Based on the analyzed emotional state, the server sends an alert to the user. The input is the estimated result of the animal's emotional state, and the output is the alert message sent to the user. This message allows the user to take action early.

[0892] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0893] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0894] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0895] [Third Embodiment]

[0896] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0897] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0898] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0899] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0900] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0901] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0902] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0903] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0904] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0905] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0906] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0907] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0908] The present invention's system analyzes and translates animal sounds, enabling communication between humans and animals. Specific embodiments of the system are described below.

[0909] System Configuration

[0910] 1. Terminal

[0911] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, and a communication function for sending data to a server.

[0912] 2. Server

[0913] The server is a computer system that performs data analysis and translation in a central location. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication format of each animal, and a mechanism for translating into natural language.

[0914] Processing flow

[0915] 1. Input

[0916] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[0917] 2. Preprocessing

[0918] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[0919] 3. Data transmission

[0920] The terminal sends pre-processed audio data to the server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[0921] 4. Voice Analysis

[0922] The server analyzes the received audio data and extracts features. Specifically, this involves spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs).

[0923] 5. Feature Mapping

[0924] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[0925] 6. Translation

[0926] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[0927] 7. Response generation

[0928] The server receives input from the user and converts it into an animal sound format. If the user asks "Tell me where to go," the corresponding animal sound is generated.

[0929] 8. Output

[0930] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[0931] Specific example

[0932] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[0933] The above describes a specific embodiment of the system of the present invention. This system enables smooth communication with animals.

[0934] The following describes the processing flow.

[0935] Step 1:

[0936] The user speaks into the device and says, "Please record dolphin sounds." The device uses its voice recognition function to detect the user's command and enters recording mode.

[0937] Step 2:

[0938] The device uses its built-in microphone to record ambient sounds. Recording continues for a period set by the user, or until manually stopped.

[0939] Step 3:

[0940] The device converts the recorded analog audio data into a digital signal. This is done using an AD converter (analog-to-digital converter).

[0941] Step 4:

[0942] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[0943] Step 5:

[0944] The terminal compresses the pre-processed audio data and packages it into data packets for efficient transmission to the server.

[0945] Step 6:

[0946] The terminal sends packaged data to the server via the internet. Stable communication is required.

[0947] Step 7:

[0948] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[0949] Step 8:

[0950] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[0951] Step 9:

[0952] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[0953] Step 10:

[0954] The server estimates the animal's intentions from the mapped data and translates those intentions into natural language that the user can understand. For example, it might generate a message like, "There are lots of fish around."

[0955] Step 11:

[0956] The user speaks to the device saying, "Tell me where I should go." The device converts this voice command into text and sends it to the server.

[0957] Step 12:

[0958] The server analyzes the user's question and generates appropriate animal sounds based on its content. For example, it converts information like "There are many fish if you head east" into a dolphin sound format.

[0959] Step 13:

[0960] The server sends the generated sound data to the terminal. The terminal analyzes the received data and prepares for playback.

[0961] Step 14:

[0962] The device plays generated animal sound data and simultaneously displays translated text information to the user. The user can communicate directly with the animal through the device.

[0963] (Example 1)

[0964] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0965] Traditional methods of communication with animals have been limited, making it difficult to accurately analyze and translate animal calls. Furthermore, the lack of a method to analyze animal calls and translate them into human natural language, and vice versa, prevented two-way communication between animals and humans. This resulted in insufficient communication between animals in captivity and research environments, posing a serious problem.

[0966] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0967] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the features to communication formats specific to each animal, and means for translating the mapped data into natural language. This makes it possible to accurately analyze animal calls and translate their intent into natural language. Furthermore, it becomes possible to analyze user input and convert natural language into animal call formats, enabling two-way communication between animals and humans.

[0968] "Means for recording animal sounds" refers to devices or systems that input the sounds of animals specified by the user as audio.

[0969] "Means for converting recorded animal sounds into digital signals" refers to devices or systems for converting analog audio data into a digital format.

[0970] "Methods for removing noise from digital signals" refer to devices and systems that improve the quality of digital audio data by removing unwanted background noise and interference sounds.

[0971] "Means for sending pre-processed audio data to a server" refers to a device or system for sending processed audio data to a remote server.

[0972] "Means for analyzing audio data and extracting features" refers to devices or systems for analyzing audio data and extracting important feature patterns.

[0973] "Means for mapping extracted features to animal-specific communication formats" refers to devices or systems that map the features of analyzed audio data to the communication formats of specific animals.

[0974] "Means for translating mapped data into natural language" refers to devices or systems that convert data corresponding to animal communication patterns into natural language that humans can understand.

[0975] "Means for transmitting translated data and generated sound data to a terminal" refers to devices or systems for transmitting data translated on a server and animal sound data to a user's terminal.

[0976] "Means for playing back data sent to a terminal" refers to devices or systems for playing back data received on a user's terminal in audio or text format.

[0977] "Means for users to input voice commands" refers to devices or systems that allow users to give instructions to a system using their voice.

[0978] "Means for a server to analyze audio data and extract features" refers to devices or systems that allow a server to analyze audio data and extract important feature patterns.

[0979] "Means by which a server maps animal vocalization patterns" refers to devices or systems that allow a server to map the characteristics of analyzed audio data to the communication patterns of specific animals.

[0980] "Means for a server to translate intentions into natural language" refers to a device or system in which a server estimates an animal's intentions and converts them into natural language that humans can understand.

[0981] "Means by which a server generates a response" refers to a device or system that allows a server to generate corresponding animal sounds based on user input.

[0982] "Means by which a terminal receives response data and presents the results to the user" refers to a device or system that allows a terminal to receive response data from a server and display or play it aloud for the user.

[0983] "Means for a terminal to reproduce animal sounds" refers to a device or system that allows a terminal to reproduce generated animal sounds as audio.

[0984] The system of the present invention analyzes and translates animal sounds to enable smooth communication between humans and animals. Specific embodiments of the system are described below.

[0985] System Configuration

[0986] 1. Terminal

[0987] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, and communication functions for sending data to a server. Specific hardware includes the smartphone's built-in microphone, an external microphone, and Wi-Fi or mobile data communication for communication. Software used includes audio recording applications and noise reduction software such as Audacity.

[0988] 2. Server

[0989] The server is a central computer system that performs data analysis and translation. This server implements an audio data analysis engine, feature extraction algorithms, a system for mapping to animal-specific communication formats, and a mechanism for translating into natural language. Specific software includes spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculation using the Python Librosa library. Furthermore, a generative AI model (e.g., GPT-4) is responsible for natural language translation.

[0990] Processing flow

[0991] The system operates using the following steps:

[0992] 1. Input: The user issues a voice command to the device, such as, "Please record dolphin sounds." For example, this could be done while observing dolphins and recording their sounds.

[0993] 2. Preprocessing: The audio data recorded by the terminal is converted into a digital signal and subjected to noise reduction processing. This improves the accuracy of the analysis.

[0994] 3. Data Transmission: The terminal sends pre-processed audio data to the server via the internet. Since stable communication is required, Wi-Fi or mobile data communication is used.

[0995] 4. Speech Analysis: The server analyzes the received speech data and extracts features. Specifically, the Librosa library is used to perform spectrogram analysis and MFCC calculations.

[0996] 5. Feature Mapping: The server uses the extracted features to map them to animal-specific vocal patterns. This allows for the estimation of the animal's intended message.

[0997] 6. Translation: The server infers the intent and translates it into natural language. For example, a message like "There are lots of fish" is generated.

[0998] 7. Response Generation: The server generates a response based on user input and converts it into an animal sound format. For example, if the user asks "Where should I go?", the corresponding animal sound will be generated.

[0999] 8. Output: The terminal receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[1000] Specific example

[1001] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[1002] Examples of specific prompt statements are as follows:

[1003] The user instructed the device to "record dolphin sounds."

[1004] The user asks the server, "Tell me where I should go."

[1005] This system will enable smoother communication with animals.

[1006] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1007] Step 1:

[1008] The user enters a voice command.

[1009] Specific action: The user speaks into the smartphone's microphone and says, "Please record dolphin sounds."

[1010] Input: User voice command.

[1011] Output: Audio input data.

[1012] Step 2:

[1013] The device records the animal's sounds.

[1014] Specific operation: The microphone built into the device records dolphin sounds according to the user's instructions.

[1015] Input: Voice input data.

[1016] Output: Recorded animal sound data.

[1017] Step 3:

[1018] The device converts the recorded data into a digital signal and removes noise.

[1019] Specific operation: The device converts the recorded data into a digital signal and performs noise reduction processing using Audacity or the built-in DSP (Digital Signal Processor).

[1020] Input: Recorded animal sound data.

[1021] Output: De-noised digital audio data.

[1022] Step 4:

[1023] The terminal sends the pre-processed data to the server.

[1024] Specific operation: The device uses Wi-Fi or mobile data communication to send pre-processed audio data to the server via the internet.

[1025] Input: Denoised digital audio data.

[1026] Output: Audio data sent to the server.

[1027] Step 5:

[1028] The server analyzes the audio data and extracts features.

[1029] Specific operation: The server uses Python libraries such as Librosa to perform spectrogram analysis of audio data and calculate Mel-frequency cepstrum coefficients (MFCCs) to extract data features.

[1030] Input: Audio data sent to the server.

[1031] Output: Extracted speech feature data.

[1032] Step 6:

[1033] The server maps the extracted features to the communication format of each animal.

[1034] Specific operation: The server identifies what the animal sounds mean based on features that match animal sound patterns it has learned in advance.

[1035] Input: Extracted speech feature data.

[1036] Output: Data mapped to a communication format.

[1037] Step 7:

[1038] The server infers the intent and translates it into natural language.

[1039] Specific operation: The server estimates the intent and translates it into natural language using a generative AI model (e.g., GPT-4).

[1040] Input: Data mapped to a communication format.

[1041] Output: Translated natural language message.

[1042] Step 8:

[1043] The server generates a response based on the user's input.

[1044] Specific operation: When the user types "Tell me where I should go," the system generates the corresponding animal sound.

[1045] Input: User's question data.

[1046] Output: Generated animal sound data.

[1047] Step 9:

[1048] The terminal receives the response data and presents the result to the user.

[1049] Specific operation: The terminal receives the translation results and sound data sent from the server, and displays them on the screen or communicates them to the user via text or voice.

[1050] Input: Translated natural language message and generated sound data.

[1051] Output: Result data presented to the user.

[1052] Step 10:

[1053] The device plays the generated animal sounds back to life.

[1054] Specific operation: The device's speaker is used to play the generated animal sounds towards the animal.

[1055] Input: Generated sound data.

[1056] Output: Playback of animal sounds.

[1057] (Application Example 1)

[1058] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1059] Traditional pet owners often struggled to accurately understand their pets' needs, particularly in selecting the right timing and type of pet food. Furthermore, responding quickly to a pet's health issues and needs relied heavily on the owner's experience and observation skills, leading to misunderstandings and delays. This created a growing demand for efficient and reliable methods to support a comfortable life for pets.

[1060] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1061] In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format specific to each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, and information communication processing means for ordering food based on the translated data. This enables pet owners to accurately understand their pets' needs and order the right food at the right time.

[1062] "Means for recording animal sounds" refers to a device or system for capturing sounds emitted by animals and recording them as electrical signals.

[1063] "Means for converting recorded vocalizations into digital signals" refers to a device or process for converting recorded analog audio signals into a digital format.

[1064] "Means of removing noise from digital signals" refers to a process or device that removes unwanted background noise and interference from digitized audio signals, thereby improving the clarity of the audio.

[1065] "Means for transmitting pre-processed audio data to a server" refers to a communication device or module for transferring noise-removed audio data to a central server via a network.

[1066] "Means for analyzing audio data and extracting features" refers to algorithms or software for analyzing transmitted audio data and identifying specific patterns or characteristics.

[1067] "Means for mapping extracted features to animal-specific communication formats" refers to a system or process that associates identified vocal patterns with communication formats determined for each animal.

[1068] "Means for translating mapped data into natural language" refers to software or algorithms for converting data mapped to a communication format into a natural language that is understandable to humans.

[1069] "Means for transmitting translated data and generated animal sounds to a terminal" refers to a communication function for transferring data translated into natural language and generated animal sounds to a corresponding terminal.

[1070] "Means for playing back data sent to a terminal" refers to a device or application for outputting received data as audio or text.

[1071] "Information and communication processing means for ordering food based on translated data" refers to a system or program for executing the process of ordering pet food from an online store or the like based on the analyzed pet's requests.

[1072] This invention provides a system that enables communication between humans and animals by analyzing and translating animal sounds. Specifically, it realizes a system that allows pet owners to accurately understand their pets' needs and order appropriate pet food accordingly.

[1073] System Configuration

[1074] 1. Terminal

[1075] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, noise reduction functions, and communication functions for sending data to the server. It also has playback functions for playing back data sent from the server and information communication processing functions for ordering food.

[1076] 2. Server

[1077] The server is a computer system that performs data analysis and translation in a central location. This server is equipped with an audio data analysis engine, a function to map identified features to the communication format of each animal, a mechanism for translating into natural language, and an information and communication processing device for ordering pet food as needed.

[1078] Specific processing flow

[1079] 1. Recording

[1080] The user gives a voice command to the device saying, "Please record my pet's barking."

[1081] The device records pet sounds and converts the analog audio signal into a digital format.

[1082] 2. Preprocessing

[1083] The terminal performs a process to remove noise from the recorded digital audio signal.

[1084] 3. Data transmission

[1085] The pre-processed audio data is sent to the server via Wi-Fi or mobile data communication.

[1086] 4. Voice Analysis

[1087] The server analyzes the received audio data and extracts specific patterns and features. For example, spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs) are performed.

[1088] 5. Feature Mapping

[1089] Using the extracted features, we map animal-specific vocal patterns to communication formats.

[1090] 6. Translation

[1091] The server translates the mapped data into natural language. For example, a message like "I'm hungry" is generated.

[1092] 7. Response generation

[1093] When a user types "Order food," the server processes the order for pet food and sends an order completion notification to the device.

[1094] 8. Output

[1095] The terminal displays the translation results received from the server and the order processing results to the user, and plays them back according to the necessary instructions.

[1096] Hardware and software used

[1097] Smartphone microphone and voice processing function: Record, convert, and remove noise from animal sounds.

[1098] Wi-Fi or mobile data: Send pre-processed data to the server.

[1099] Server analysis engine: Performs speech data analysis, feature extraction, mapping, and translation.

[1100] API communication: Based on the translated data, it performs information communication processing to order food.

[1101] Specific example

[1102] When a user records their pet's barking, the device starts recording, converts it to a digital signal, and removes noise. The data is then sent to a server where voice analysis and feature extraction are performed. The extracted features are mapped as the animal's request and translated as "I'm hungry." When the user instructs "Order food," the server processes the order, and an order completion notification is displayed on the user's device.

[1103] Example of a prompt

[1104] "Create a smartphone application that records pet sounds and orders pet food from a delivery service based on the analysis results. The user interface will consist of three buttons: a recording button, a display of analysis results, and a food order button. The recorded data should be sent to a server, and the analysis results should be received and displayed. Order processing should be done using an API, and the success / failure status should be displayed. Please provide a specific code example along with the output."

[1105] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1106] Step 1:

[1107] The user instructs the device to "record my pet's barking." The device then captures the barking and inputs it as an analog audio signal. The device's audio processing function then activates to convert the input barking into a digital signal. As a result, the analog audio is output as a digital signal.

[1108] Step 2:

[1109] The terminal removes noise from the converted digital audio signal. It takes a digital signal as input and applies a noise reduction algorithm. As a result, clear audio data with noise removed is output. Specifically, the process filters out unwanted information such as high frequencies and background noise.

[1110] Step 3:

[1111] The terminal sends pre-processed audio data to the server via the internet. It takes clear audio data as input and transfers it to the server using Wi-Fi or mobile data communication. The server receives this data and is ready for audio analysis.

[1112] Step 4:

[1113] The server analyzes the received audio data and extracts specific patterns and features. It takes audio data as input and performs feature extraction using algorithms such as spectrogram analysis and Mel-frequency cepstrum coefficients (MFCC). The extracted features are output as data, and the process proceeds to the next step.

[1114] Step 5:

[1115] The server uses extracted features to map animal-specific vocal patterns to communication formats. It takes feature data as input and performs mapping according to the communication format rules for each animal. This results in outputting data corresponding to the pet's intentions.

[1116] Step 6:

[1117] The server translates the mapped data into natural language. It takes formatted data as input and applies a natural language processing algorithm to perform the translation. As a result, it outputs a natural language message such as "I'm hungry."

[1118] Step 7:

[1119] The user enters "Order food" into the device. The device takes the user's instruction as input and sends it to the server. The server receives the instruction and prepares to place the pet food order via API.

[1120] Step 8:

[1121] The server performs information and communication processing to order food based on the translated data. It takes order data as input and sends an order request to the online store using an API. If the order is successful, confirmation data is output and sent to the terminal.

[1122] Step 9:

[1123] The terminal receives confirmation data sent from the server and presents the user with an order completion notification. It takes the confirmation data as input and notifies the user through display and audio notifications. This allows the user to confirm that the order has been successfully completed.

[1124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1125] The system of the present invention enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. Specific embodiments of the system are shown below.

[1126] System Configuration

[1127] 1. Terminal

[1128] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, a communication function for sending data to a server, and an emotion engine that recognizes the user's emotions.

[1129] 2. Server

[1130] The server is a computer system that centrally performs data analysis, translation, and emotional data analysis. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication formats of each animal, a mechanism for translating into natural language, and a system for receiving and analyzing emotional data transmitted from the emotional engine.

[1131] Processing flow

[1132] 1. Input

[1133] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[1134] 2. Preprocessing

[1135] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[1136] 3. Data transmission

[1137] The device sends pre-processed audio data and user emotion data analyzed by the emotion engine to a server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[1138] 4. Voice Analysis

[1139] The server analyzes the received audio data and extracts features. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[1140] 5. Feature Mapping

[1141] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[1142] 6. Translation

[1143] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[1144] 7. Response generation

[1145] The server receives input and emotion data from the user and converts it into an animal sound format. If the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates a relaxing sound data such as "There are lots of fish if you head east."

[1146] 8. Output

[1147] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[1148] Specific example

[1149] When analyzing dolphin sounds, the user instructs the device to "record dolphin sounds." The device records the sounds, converts them into digital signals, and removes noise. The pre-processed data and the user's emotional data analyzed by the emotion engine are then sent to the server, where the audio data is analyzed. The server extracts the characteristics of the dolphin sounds and uses them to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates an appropriate dolphin sound in response to that question, providing information to help the user relax.

[1150] The above describes a specific embodiment of the system of the present invention. This system not only facilitates smooth communication with animals but also enables appropriate interaction that takes into account the user's emotional state.

[1151] The following describes the processing flow.

[1152] Step 1:

[1153] The user speaks into the device and says, "Please record dolphin sounds." The device uses its voice recognition function to detect the user's command and enters recording mode.

[1154] Step 2:

[1155] The device uses its built-in microphone to record ambient sounds. Recording continues for a period set by the user, or until manually stopped.

[1156] Step 3:

[1157] The device converts the recorded analog audio data into a digital signal. This is done using an AD converter (analog-to-digital converter).

[1158] Step 4:

[1159] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[1160] Step 5:

[1161] Simultaneously, the device analyzes the user's voice, facial expressions, and gestures using an emotion engine to recognize their emotional state. It analyzes tone and speed from the voice, facial movements from the expressions, and body movements from the gestures.

[1162] Step 6:

[1163] The terminal packages pre-processed audio data and user emotion data analyzed by the emotion engine into data packets.

[1164] Step 7:

[1165] The terminal sends packaged data to the server via the internet. Stable communication is required.

[1166] Step 8:

[1167] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[1168] Step 9:

[1169] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[1170] Step 10:

[1171] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[1172] Step 11:

[1173] The server estimates the animal's intentions from the mapped data and translates those intentions into natural language that the user can understand. For example, it might generate a message like, "There are lots of fish around."

[1174] Step 12:

[1175] The server receives user input and emotional data, and converts it into animal sound formats. For example, if a user asks excitedly, "Tell me where I should go," the server will use dolphin sounds to generate a calming message such as, "There are lots of fish if you head east."

[1176] Step 13:

[1177] The server sends the generated sound data and translated text data to the terminal. The terminal analyzes the received data and prepares for playback.

[1178] Step 14:

[1179] The device plays generated animal sound data and simultaneously displays translated text information to the user. The user can communicate directly with the animal through the device.

[1180] (Example 2)

[1181] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1182] Traditional methods of communicating with animals have the problem that it is difficult for users to understand the meaning of animal sounds. Furthermore, because communication is conducted without considering the user's emotional state, it can be stressful for the user. In addition, the ability to properly analyze and translate animal sounds is not yet fully realized. Therefore, there is a need to develop a system that accurately grasps the intentions of animals and conveys them to the user in an easily understandable way.

[1183] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1184] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the extracted features to the communication format of each animal, and means for the server to analyze the user's emotional data. This makes it possible to analyze animal sounds and generate appropriate responses based on the user's emotional state.

[1185] "Means for recording animal sounds" refers to devices that collect sounds emitted by animals using acoustic sensors such as microphones.

[1186] "Means for converting recorded animal sounds into digital signals" refers to devices or software that convert collected animal sound audio data from analog to digital format through digital signal processing.

[1187] "Methods for removing noise from digital signals" refer to filtering algorithms and software used to remove unwanted background noise and other unwanted sounds from converted digital signals.

[1188] "Means for transmitting pre-processed audio data to a server" refers to communication devices and protocols for transferring noise-reduced audio data to a server using internet communication or wireless communication.

[1189] "Methods for analyzing audio data and extracting features" refer to software that performs analysis on audio data on a server to extract important features such as the frequency characteristics and temporal characteristics of the audio.

[1190] "Means for mapping extracted features to animal-specific communication formats" refers to algorithms or software that associate extracted vocal features with animal-specific vocalization patterns.

[1191] "Means for translating mapped data into natural language" refers to software that converts data obtained based on animal communication formats into natural language that humans can understand.

[1192] "Means for a server to analyze user emotional data" refers to software or algorithms that allow a server to analyze emotional data received from a user and understand the user's emotional state.

[1193] "Means for transmitting generated sound data to a terminal" refers to communication devices and protocols for transferring animal sound data generated by a server to a terminal using internet communication or wireless communication.

[1194] "Means for playing back data transmitted to a terminal" refers to output devices such as speakers or playback software that play back the sound data and translation data received by the terminal to the user or animal.

[1195] This invention describes a system that enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. This system mainly includes the following components.

[1196] System Configuration

[1197] 1. Terminal

[1198] It consists of the user's smartphone or dedicated device. This device has the following functions:

[1199] Microphone: An acoustic sensor used to record animal sounds.

[1200] Audio processing function: A digital signal processing (DSP) algorithm that converts recorded audio data into a digital signal and removes noise.

[1201] Communication functions: Wi-Fi and mobile data communication functions for sending pre-processed voice data and user emotion data to the server.

[1202] Emotion engine: Software used to recognize and analyze a user's emotions.

[1203] 2. Server

[1204] A central computer system that performs data analysis, translation, and sentiment data analysis. It includes the following functions:

[1205] Speech analysis engine: Software that analyzes transmitted speech data and performs spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculations to extract features.

[1206] Feature extraction algorithm: An algorithm that associates extracted features with the vocalization patterns of different animals.

[1207] Natural language translation system: Software that estimates the intentions of animals and translates those intentions into natural language.

[1208] Emotional Data Analysis System: A system that analyzes user emotional data transmitted from an emotional engine and generates responses based on that analysis.

[1209] System processing flow

[1210] The system's processing flow is as follows: First, the user issues a voice command to the terminal, initiating the recording of animal sounds. The recorded voice data is converted into a digital signal and subjected to noise reduction processing. Subsequently, the pre-processed voice data and the user's emotional data are sent from the terminal to the server. The server analyzes the voice data and maps it to animal sound patterns. Next, these patterns are translated into natural language, and the content is presented to the user. In addition, an appropriate response is generated according to the user's emotional state and communicated to the animal using the playback function.

[1211] Specific example

[1212] When analyzing dolphin sounds, the user instructs the device to "record dolphin sounds." The device records the sounds and removes noise using a DSP algorithm. The cleared audio data and the user's emotional data, analyzed by the emotion engine, are then sent to the server. The server analyzes the audio data using spectrogram analysis and MFCC calculations to extract dolphin-specific sound patterns. Based on these patterns, it translates the message "There are lots of fish" into natural language. If the user asks "Where should I go?" and the emotion engine detects an excited state, the server generates appropriate dolphin sounds and provides information to help the user relax.

[1213] Example of a prompt

[1214] The following are specific examples of voice commands that users can issue to their devices.

[1215] "Please record the sounds of dolphins."

[1216] "Please tell me the next destination."

[1217] "Tell me what the dolphins are feeling right now."

[1218] The above describes a specific embodiment of the system of the present invention. This system facilitates smooth communication with animals and enables appropriate interaction that takes into account the user's emotional state.

[1219] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1220] Step 1:

[1221] Input: The user issues a voice command to the device saying, "Please record the sound of a dolphin."

[1222] Operation: The user speaks voice commands into their smartphone or a dedicated device.

[1223] Output: The terminal receives the recording command and begins preparing to record the animal sounds.

[1224] Step 2:

[1225] Input: The terminal receives a recording command from the user.

[1226] Operation: The device's microphone activates and records dolphin sounds. The acoustic sensor inside the smartphone is used during this process.

[1227] Output: Generates recorded analog audio data.

[1228] Step 3:

[1229] Input: Recorded analog audio data.

[1230] Operation: The terminal converts this analog audio data into a digital signal. Digital signal processing (DSP) algorithms are used in this process.

[1231] Output: Generates audio data converted into a digital signal.

[1232] Step 4:

[1233] Input: Audio data converted to a digital signal.

[1234] Operation: The terminal performs noise reduction processing on this digital signal. Unnecessary background noise and static are filtered out.

[1235] Output: Generates clear audio data with noise reduction.

[1236] Step 5:

[1237] Input: Denoised audio data and user emotion data analyzed by the emotion engine.

[1238] Operation: The device collects this data and sends it to the server via internet communication. Wi-Fi or mobile data communication is used for this purpose.

[1239] Output: Emotional data and audio data sent to the server.

[1240] Step 6:

[1241] Input: Audio data sent to the server.

[1242] Operation: The server analyzes audio data using an audio analysis engine. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[1243] Output: Extracted speech feature data.

[1244] Step 7:

[1245] Input: Extracted speech feature data.

[1246] Operation: The server maps this feature data to animal-specific vocal patterns. The data is then transformed into dolphin-specific communication patterns.

[1247] Output: Mapped communication data.

[1248] Step 8:

[1249] Input: Mapped communication data.

[1250] Operation: The server translates this into natural language. For example, it can generate a specific message like "There are lots of fish" from the sound of dolphins.

[1251] Output: Translated natural language message.

[1252] Step 9:

[1253] Input: Server-generated natural language messages and user sentiment data.

[1254] Operation: The server generates appropriate animal sound formats based on user input and emotional data. For example, if it detects that the user is agitated, it will generate sound data to help them relax.

[1255] Output: Audio data converted to animal sound format.

[1256] Step 10:

[1257] Input: Translation results and sound format data sent from the server.

[1258] Operation: The device receives this data and notifies the user. Furthermore, it transmits the sounds generated using the device's playback function to the animals.

[1259] Output: Translated message presented to the user and played sound data.

[1260] The above outlines the specific program processing flow of this system. Through the specific actions at each step, smooth communication between animals and humans is achieved, and interaction that takes into account the user's emotional state becomes possible.

[1261] (Application Example 2)

[1262] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1263] Traditional animal communication systems focus on analyzing animal sounds and translating them into natural language, but this has limitations in real-time interaction between humans and animals and in ensuring safety. Furthermore, there is a lack of means for users to detect their pet's emotional state or abnormal behavior. This makes it difficult to respond quickly when a pet suddenly becomes agitated or makes suspicious noises, potentially leading to safety issues within the home.

[1264] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format for each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, means for analyzing the emotional state of the animal using the recorded audio data, and means for sending an alert to the user based on the analyzed emotional state. This makes it possible not only to communicate smoothly with animals, but also to quickly detect the emotional state and abnormal behavior of pets, thereby enhancing safety in the home.

[1265] Definitions of important words

[1266] "Means for recording animal sounds" refers to a device or system for electronically capturing animal sounds along with ambient noise and saving them as digital audio files.

[1267] "Means for converting recorded vocalizations into digital signals" refers to a device or technology for converting recorded analog audio data into a digital format.

[1268] "Means for removing noise from digital signals" refers to devices or algorithms for removing unwanted background noise and other unwanted sounds contained in digital audio data, thereby improving the accuracy of the analysis.

[1269] "Means for transmitting pre-processed audio data to a server" refers to a device or system for transmitting audio data that has undergone pre-processing, such as noise reduction, to a central server via the internet or other communication means.

[1270] "Means for analyzing audio data and extracting features" refers to technologies or devices for analyzing audio data and extracting features such as specific frequencies or sound patterns.

[1271] "Means for mapping extracted features to animal-specific communication formats" refers to a device or algorithm for converting extracted vocal features into data corresponding to the communication format specific to that animal.

[1272] "Means for translating mapped data into natural language" refers to devices or technologies for converting data representing animal intentions and emotions into natural language that humans can understand.

[1273] "Means for transmitting translated data and generated sound data to a terminal" refers to a device or system for transmitting data translated into natural language and, if necessary, generated animal sound data to a user's terminal.

[1274] "Means for playing back data transmitted to a terminal" refers to a device or application that allows the user to visually or aurally confirm transmitted audio data or translation results.

[1275] "Means for analyzing an animal's emotional state using recorded audio data" refers to a device or algorithm for analyzing recorded audio data to estimate the animal's current emotional state (e.g., excitement, tension, relaxation).

[1276] "Means for sending alerts to users based on analyzed emotional states" refers to a device or system for sending a warning notification to a user when an abnormal emotional state of an animal is detected.

[1277] Modes for carrying out the invention

[1278] System Configuration

[1279] The system of this invention mainly consists of terminals and servers. This system makes it possible to analyze animal sounds in real time and provide useful information to the user.

[1280] 1. Terminal

[1281] The device consists of the user's smartphone or dedicated device. The device has the following functions:

[1282] Microphone for recording animal sounds

[1283] Audio processing function that converts recorded animal sounds into digital signals.

[1284] A processing function that removes noise from digital signals.

[1285] A communication function that sends pre-processed audio data to the server.

[1286] These features allow for the real-time capture of animal sounds, pre-processing, and transmission of the data to a server.

[1287] 2. Server

[1288] A server is a computer system that centrally performs data analysis, translation, and sentiment data analysis. The server has the following functions:

[1289] Voice data analysis engine

[1290] Feature extraction algorithm

[1291] A system that maps to the communication format of each animal.

[1292] A mechanism for translating into natural language

[1293] Generative AI models for developing emotional engines

[1294] This allows the server to analyze the received audio data and extract its features. Based on these extracted features, it classifies the sounds as animal calls and translates them into natural language. It can also analyze user input and convert their intent into animal calls.

[1295] Analysis and alert function for animal emotional states

[1296] The system also has a function to analyze the emotional state of animals using recorded audio data. This function estimates the animal's current emotions (e.g., excitement, tension, relaxation) from the analyzed vocalization data. It then sends alerts to the user's device, notifying them in real time of abnormal behavior or suspicious vocalizations. This allows the user to take necessary actions quickly. For example, if a pet suddenly becomes agitated or makes a suspicious vocalization, the device can send an alert saying "Suspicious vocalization detected," prompting the user to take immediate action.

[1297] Specific example

[1298] Suppose a user is checking on their pet at home via their smartphone, and the pet suddenly makes an unusual noise. In this case, the device records the noise, converts it to a digital signal, removes noise, and sends it to a server. The server analyzes the audio data, extracts its characteristics, and estimates the pet's emotional state (e.g., fear, excitement). If an abnormality is detected, an alert is sent to the user. The user receives the alert and can take appropriate action quickly.

[1299] Example of a prompt

[1300] Design a model that extracts pet vocalization characteristics from audio data and predicts suspicious vocalizations and emotional states. Use MFCC for audio data preprocessing and build the model using TensorFlow / Keras.

[1301] The above describes the embodiments for carrying out this invention. This not only facilitates smooth communication with animals but also enhances safety within the home.

[1302] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1303] Program processing flow

[1304] Processing flow of the system program that implements the application example

[1305] Step 1:

[1306] The device records animal sounds based on user instructions. The input is an audio signal obtained through a recording device, and the output is a digital audio file. This step includes specific actions to capture animal sounds along with ambient sounds.

[1307] Step 2:

[1308] The terminal converts the recorded audio signal into a digital signal. The input is an analog audio signal, and the output is audio data converted into a digital format (e.g., PCM format). This conversion facilitates subsequent data processing.

[1309] Step 3:

[1310] The terminal removes noise from the digital audio data. The input is digital audio data, and the output is audio data that has been denoised. The specific action performed in this step is to improve the accuracy of the analysis by removing unwanted background noise using a filtering algorithm.

[1311] Step 4:

[1312] The terminal sends pre-processed audio data to the server. The input is denoised digital audio data, and the output is audio data transmitted over the internet. The specific actions in this step include transferring the data to the server via a communication protocol.

[1313] Step 5:

[1314] The server analyzes the received audio data and extracts features. The input is audio data transmitted from the terminal, and the output is audio feature quantities (e.g., MFCC). The analysis engine is used to process the audio waveform and perform the actual calculations to extract features.

[1315] Step 6:

[1316] The server uses extracted features to map them to the communication formats of each animal. The input is the features, and the output is data mapped to the animal's unique vocal patterns. Based on the features, a specific algorithm is implemented to estimate the animal's intentions and emotions.

[1317] Step 7:

[1318] The server translates mapped data into natural language. The input is animal vocalization pattern data, and the output is a text message translated into natural language. This translation process uses a generative AI model to convert what the animals are trying to communicate into human language.

[1319] Step 8:

[1320] The server sends the translated data and generated sound data to the terminal. The input is natural language text and sound data, and the output is data sent to the user's terminal via the internet. This allows the user to see what the animals are trying to communicate in real time.

[1321] Step 9:

[1322] The device plays back data received from the server. Input is natural language text and animal sound data, and output is audio and text presented to the user. Specific actions include displaying text messages and playing animal sound data through the speaker.

[1323] Step 10:

[1324] The server analyzes the emotional state of an animal using recorded audio data. The input is recorded audio data, and the output is an estimated result of the animal's emotional state. An emotion engine is used to perform specific data calculations to estimate emotional states such as excitement and relaxation.

[1325] Step 11:

[1326] Based on the analyzed emotional state, the server sends an alert to the user. The input is the estimated result of the animal's emotional state, and the output is the alert message sent to the user. This message allows the user to take action early.

[1327] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1328] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1329] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1330] [Fourth Embodiment]

[1331] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1332] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1333] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1334] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1335] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1336] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1337] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1338] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1339] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1340] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1341] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1342] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1343] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1344] The present invention's system analyzes and translates animal sounds, enabling communication between humans and animals. Specific embodiments of the system are described below.

[1345] System Configuration

[1346] 1. Terminal

[1347] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, and a communication function for sending data to a server.

[1348] 2. Server

[1349] The server is a computer system that performs data analysis and translation in a central location. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication format of each animal, and a mechanism for translating into natural language.

[1350] Processing flow

[1351] 1. Input

[1352] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[1353] 2. Preprocessing

[1354] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[1355] 3. Data transmission

[1356] The terminal sends pre-processed audio data to the server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[1357] 4. Voice Analysis

[1358] The server analyzes the received audio data and extracts features. Specifically, this involves spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs).

[1359] 5. Feature Mapping

[1360] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[1361] 6. Translation

[1362] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[1363] 7. Response generation

[1364] The server receives input from the user and converts it into an animal sound format. If the user asks "Tell me where to go," the corresponding animal sound is generated.

[1365] 8. Output

[1366] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[1367] Specific example

[1368] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[1369] The above describes a specific embodiment of the system of the present invention. This system enables smooth communication with animals.

[1370] The following describes the processing flow.

[1371] Step 1:

[1372] The user speaks into the device and says, "Please record dolphin sounds." The device uses its voice recognition function to detect the user's command and enters recording mode.

[1373] Step 2:

[1374] The device uses its built-in microphone to record ambient sounds. Recording continues for a period set by the user, or until manually stopped.

[1375] Step 3:

[1376] The device converts the recorded analog audio data into a digital signal. This is done using an AD converter (analog-to-digital converter).

[1377] Step 4:

[1378] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[1379] Step 5:

[1380] The terminal compresses the pre-processed audio data and packages it into data packets for efficient transmission to the server.

[1381] Step 6:

[1382] The terminal sends packaged data to the server via the internet. Stable communication is required.

[1383] Step 7:

[1384] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[1385] Step 8:

[1386] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[1387] Step 9:

[1388] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[1389] Step 10:

[1390] The server estimates the animal's intentions from the mapped data and translates those intentions into natural language that the user can understand. For example, it might generate a message like, "There are lots of fish around."

[1391] Step 11:

[1392] The user speaks to the device saying, "Tell me where I should go." The device converts this voice command into text and sends it to the server.

[1393] Step 12:

[1394] The server analyzes the user's question and generates appropriate animal sounds based on its content. For example, it converts information like "There are many fish if you head east" into a dolphin sound format.

[1395] Step 13:

[1396] The server sends the generated sound data to the terminal. The terminal analyzes the received data and prepares for playback.

[1397] Step 14:

[1398] The device plays generated animal sound data and simultaneously displays translated text information to the user. The user can communicate directly with the animal through the device.

[1399] (Example 1)

[1400] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1401] Traditional methods of communication with animals have been limited, making it difficult to accurately analyze and translate animal calls. Furthermore, the lack of a method to analyze animal calls and translate them into human natural language, and vice versa, prevented two-way communication between animals and humans. This resulted in insufficient communication between animals in captivity and research environments, posing a serious problem.

[1402] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1403] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the features to communication formats specific to each animal, and means for translating the mapped data into natural language. This makes it possible to accurately analyze animal calls and translate their intent into natural language. Furthermore, it becomes possible to analyze user input and convert natural language into animal call formats, enabling two-way communication between animals and humans.

[1404] "Means for recording animal sounds" refers to devices or systems that input the sounds of animals specified by the user as audio.

[1405] "Means for converting recorded animal sounds into digital signals" refers to devices or systems for converting analog audio data into a digital format.

[1406] "Methods for removing noise from digital signals" refer to devices and systems that improve the quality of digital audio data by removing unwanted background noise and interference sounds.

[1407] "Means for sending pre-processed audio data to a server" refers to a device or system for sending processed audio data to a remote server.

[1408] "Means for analyzing audio data and extracting features" refers to devices or systems for analyzing audio data and extracting important feature patterns.

[1409] "Means for mapping extracted features to animal-specific communication formats" refers to devices or systems that map the features of analyzed audio data to the communication formats of specific animals.

[1410] "Means for translating mapped data into natural language" refers to devices or systems that convert data corresponding to animal communication patterns into natural language that humans can understand.

[1411] "Means for transmitting translated data and generated sound data to a terminal" refers to devices or systems for transmitting data translated on a server and animal sound data to a user's terminal.

[1412] "Means for playing back data sent to a terminal" refers to devices or systems for playing back data received on a user's terminal in audio or text format.

[1413] "Means for users to input voice commands" refers to devices or systems that allow users to give instructions to a system using their voice.

[1414] "Means for a server to analyze audio data and extract features" refers to devices or systems that allow a server to analyze audio data and extract important feature patterns.

[1415] "Means by which a server maps animal vocalization patterns" refers to devices or systems that allow a server to map the characteristics of analyzed audio data to the communication patterns of specific animals.

[1416] "Means for a server to translate intentions into natural language" refers to a device or system in which a server estimates an animal's intentions and converts them into natural language that humans can understand.

[1417] "Means by which a server generates a response" refers to a device or system that allows a server to generate corresponding animal sounds based on user input.

[1418] "Means by which a terminal receives response data and presents the results to the user" refers to a device or system that allows a terminal to receive response data from a server and display or play it aloud for the user.

[1419] "Means for a terminal to reproduce animal sounds" refers to a device or system that allows a terminal to reproduce generated animal sounds as audio.

[1420] The system of the present invention analyzes and translates animal sounds to enable smooth communication between humans and animals. Specific embodiments of the system are described below.

[1421] System Configuration

[1422] 1. Terminal

[1423] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, and communication functions for sending data to a server. Specific hardware includes the smartphone's built-in microphone, an external microphone, and Wi-Fi or mobile data communication for communication. Software used includes audio recording applications and noise reduction software such as Audacity.

[1424] 2. Server

[1425] The server is a central computer system that performs data analysis and translation. This server implements an audio data analysis engine, feature extraction algorithms, a system for mapping to animal-specific communication formats, and a mechanism for translating into natural language. Specific software includes spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculation using the Python Librosa library. Furthermore, a generative AI model (e.g., GPT-4) is responsible for natural language translation.

[1426] Processing flow

[1427] The system operates using the following steps:

[1428] 1. Input: The user issues a voice command to the device, such as, "Please record dolphin sounds." For example, this could be done while observing dolphins and recording their sounds.

[1429] 2. Preprocessing: The audio data recorded by the terminal is converted into a digital signal and subjected to noise reduction processing. This improves the accuracy of the analysis.

[1430] 3. Data Transmission: The terminal sends pre-processed audio data to the server via the internet. Since stable communication is required, Wi-Fi or mobile data communication is used.

[1431] 4. Speech Analysis: The server analyzes the received speech data and extracts features. Specifically, the Librosa library is used to perform spectrogram analysis and MFCC calculations.

[1432] 5. Feature Mapping: The server uses the extracted features to map them to animal-specific vocal patterns. This allows for the estimation of the animal's intended message.

[1433] 6. Translation: The server infers the intent and translates it into natural language. For example, a message like "There are lots of fish" is generated.

[1434] 7. Response Generation: The server generates a response based on user input and converts it into an animal sound format. For example, if the user asks "Where should I go?", the corresponding animal sound will be generated.

[1435] 8. Output: The terminal receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[1436] Specific example

[1437] When analyzing dolphin sounds, the user instructs the device to "record a dolphin sound." The device records the sound, converts it to a digital signal, and removes noise. The pre-processed data is then sent to a server where the audio data is analyzed. The server extracts the characteristics of the dolphin sound and uses that to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?", the server generates an appropriate dolphin sound in response, which the device plays.

[1438] Examples of specific prompt statements are as follows:

[1439] The user instructed the device to "record dolphin sounds."

[1440] The user asks the server, "Tell me where I should go."

[1441] This system will enable smoother communication with animals.

[1442] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1443] Step 1:

[1444] The user enters a voice command.

[1445] Specific action: The user speaks into the smartphone's microphone and says, "Please record dolphin sounds."

[1446] Input: User voice command.

[1447] Output: Audio input data.

[1448] Step 2:

[1449] The device records the animal's sounds.

[1450] Specific operation: The microphone built into the device records dolphin sounds according to the user's instructions.

[1451] Input: Voice input data.

[1452] Output: Recorded animal sound data.

[1453] Step 3:

[1454] The device converts the recorded data into a digital signal and removes noise.

[1455] Specific operation: The device converts the recorded data into a digital signal and performs noise reduction processing using Audacity or the built-in DSP (Digital Signal Processor).

[1456] Input: Recorded animal sound data.

[1457] Output: De-noised digital audio data.

[1458] Step 4:

[1459] The terminal sends the pre-processed data to the server.

[1460] Specific operation: The device uses Wi-Fi or mobile data communication to send pre-processed audio data to the server via the internet.

[1461] Input: Denoised digital audio data.

[1462] Output: Audio data sent to the server.

[1463] Step 5:

[1464] The server analyzes the audio data and extracts features.

[1465] Specific operation: The server uses Python libraries such as Librosa to perform spectrogram analysis of audio data and calculate Mel-frequency cepstrum coefficients (MFCCs) to extract data features.

[1466] Input: Audio data sent to the server.

[1467] Output: Extracted speech feature data.

[1468] Step 6:

[1469] The server maps the extracted features to the communication format of each animal.

[1470] Specific operation: The server identifies what the animal sounds mean based on features that match animal sound patterns it has learned in advance.

[1471] Input: Extracted speech feature data.

[1472] Output: Data mapped to a communication format.

[1473] Step 7:

[1474] The server infers the intent and translates it into natural language.

[1475] Specific operation: The server estimates the intent and translates it into natural language using a generative AI model (e.g., GPT-4).

[1476] Input: Data mapped to a communication format.

[1477] Output: Translated natural language message.

[1478] Step 8:

[1479] The server generates a response based on the user's input.

[1480] Specific operation: When the user types "Tell me where I should go," the system generates the corresponding animal sound.

[1481] Input: User's question data.

[1482] Output: Generated animal sound data.

[1483] Step 9:

[1484] The terminal receives the response data and presents the result to the user.

[1485] Specific operation: The terminal receives the translation results and sound data sent from the server, and displays them on the screen or communicates them to the user via text or voice.

[1486] Input: Translated natural language message and generated sound data.

[1487] Output: Result data presented to the user.

[1488] Step 10:

[1489] The device plays the generated animal sounds back to life.

[1490] Specific operation: The device's speaker is used to play the generated animal sounds towards the animal.

[1491] Input: Generated sound data.

[1492] Output: Playback of animal sounds.

[1493] (Application Example 1)

[1494] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1495] Traditional pet owners often struggled to accurately understand their pets' needs, particularly in selecting the right timing and type of pet food. Furthermore, responding quickly to a pet's health issues and needs relied heavily on the owner's experience and observation skills, leading to misunderstandings and delays. This created a growing demand for efficient and reliable methods to support a comfortable life for pets.

[1496] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1497] In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format specific to each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, and information communication processing means for ordering food based on the translated data. This enables pet owners to accurately understand their pets' needs and order the right food at the right time.

[1498] "Means for recording animal sounds" refers to a device or system for capturing sounds emitted by animals and recording them as electrical signals.

[1499] "Means for converting recorded vocalizations into digital signals" refers to a device or process for converting recorded analog audio signals into a digital format.

[1500] "Means of removing noise from digital signals" refers to a process or device that removes unwanted background noise and interference from digitized audio signals, thereby improving the clarity of the audio.

[1501] "Means for transmitting pre-processed audio data to a server" refers to a communication device or module for transferring noise-removed audio data to a central server via a network.

[1502] "Means for analyzing audio data and extracting features" refers to algorithms or software for analyzing transmitted audio data and identifying specific patterns or characteristics.

[1503] "Means for mapping extracted features to animal-specific communication formats" refers to a system or process that associates identified vocal patterns with communication formats determined for each animal.

[1504] "Means for translating mapped data into natural language" refers to software or algorithms for converting data mapped to a communication format into a natural language that is understandable to humans.

[1505] "Means for transmitting translated data and generated animal sounds to a terminal" refers to a communication function for transferring data translated into natural language and generated animal sounds to a corresponding terminal.

[1506] "Means for playing back data sent to a terminal" refers to a device or application for outputting received data as audio or text.

[1507] "Information and communication processing means for ordering food based on translated data" refers to a system or program for executing the process of ordering pet food from an online store or the like based on the analyzed pet's requests.

[1508] This invention provides a system that enables communication between humans and animals by analyzing and translating animal sounds. Specifically, it realizes a system that allows pet owners to accurately understand their pets' needs and order appropriate pet food accordingly.

[1509] System Configuration

[1510] 1. Terminal

[1511] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, audio processing functions for converting them into digital signals, noise reduction functions, and communication functions for sending data to the server. It also has playback functions for playing back data sent from the server and information communication processing functions for ordering food.

[1512] 2. Server

[1513] The server is a computer system that performs data analysis and translation in a central location. This server is equipped with an audio data analysis engine, a function to map identified features to the communication format of each animal, a mechanism for translating into natural language, and an information and communication processing device for ordering pet food as needed.

[1514] Specific processing flow

[1515] 1. Recording

[1516] The user gives a voice command to the device saying, "Please record my pet's barking."

[1517] The device records pet sounds and converts the analog audio signal into a digital format.

[1518] 2. Preprocessing

[1519] The terminal performs a process to remove noise from the recorded digital audio signal.

[1520] 3. Data transmission

[1521] The pre-processed audio data is sent to the server via Wi-Fi or mobile data communication.

[1522] 4. Voice Analysis

[1523] The server analyzes the received audio data and extracts specific patterns and features. For example, spectrogram analysis and calculation of Mel-frequency cepstrum coefficients (MFCCs) are performed.

[1524] 5. Feature Mapping

[1525] Using the extracted features, we map animal-specific vocal patterns to communication formats.

[1526] 6. Translation

[1527] The server translates the mapped data into natural language. For example, a message like "I'm hungry" is generated.

[1528] 7. Response generation

[1529] When a user types "Order food," the server processes the order for pet food and sends an order completion notification to the device.

[1530] 8. Output

[1531] The terminal displays the translation results received from the server and the order processing results to the user, and plays them back according to the necessary instructions.

[1532] Hardware and software used

[1533] Smartphone microphone and voice processing function: Record, convert, and remove noise from animal sounds.

[1534] Wi-Fi or mobile data: Send pre-processed data to the server.

[1535] Server analysis engine: Performs speech data analysis, feature extraction, mapping, and translation.

[1536] API communication: Based on the translated data, it performs information communication processing to order food.

[1537] Specific example

[1538] When a user records their pet's barking, the device starts recording, converts it to a digital signal, and removes noise. The data is then sent to a server where voice analysis and feature extraction are performed. The extracted features are mapped as the animal's request and translated as "I'm hungry." When the user instructs "Order food," the server processes the order, and an order completion notification is displayed on the user's device.

[1539] Example of a prompt

[1540] "Create a smartphone application that records pet sounds and orders pet food from a delivery service based on the analysis results. The user interface will consist of three buttons: a recording button, a display of analysis results, and a food order button. The recorded data should be sent to a server, and the analysis results should be received and displayed. Order processing should be done using an API, and the success / failure status should be displayed. Please provide a specific code example along with the output."

[1541] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1542] Step 1:

[1543] The user instructs the device to "record my pet's barking." The device then captures the barking and inputs it as an analog audio signal. The device's audio processing function then activates to convert the input barking into a digital signal. As a result, the analog audio is output as a digital signal.

[1544] Step 2:

[1545] The terminal removes noise from the converted digital audio signal. It takes a digital signal as input and applies a noise reduction algorithm. As a result, clear audio data with noise removed is output. Specifically, the process filters out unwanted information such as high frequencies and background noise.

[1546] Step 3:

[1547] The terminal sends pre-processed audio data to the server via the internet. It takes clear audio data as input and transfers it to the server using Wi-Fi or mobile data communication. The server receives this data and is ready for audio analysis.

[1548] Step 4:

[1549] The server analyzes the received audio data and extracts specific patterns and features. It takes audio data as input and performs feature extraction using algorithms such as spectrogram analysis and Mel-frequency cepstrum coefficients (MFCC). The extracted features are output as data, and the process proceeds to the next step.

[1550] Step 5:

[1551] The server uses extracted features to map animal-specific vocal patterns to communication formats. It takes feature data as input and performs mapping according to the communication format rules for each animal. This results in outputting data corresponding to the pet's intentions.

[1552] Step 6:

[1553] The server translates the mapped data into natural language. It takes formatted data as input and applies a natural language processing algorithm to perform the translation. As a result, it outputs a natural language message such as "I'm hungry."

[1554] Step 7:

[1555] The user enters "Order food" into the device. The device takes the user's instruction as input and sends it to the server. The server receives the instruction and prepares to place the pet food order via API.

[1556] Step 8:

[1557] The server performs information and communication processing to order food based on the translated data. It takes order data as input and sends an order request to the online store using an API. If the order is successful, confirmation data is output and sent to the terminal.

[1558] Step 9:

[1559] The terminal receives confirmation data sent from the server and presents the user with an order completion notification. It takes the confirmation data as input and notifies the user through display and audio notifications. This allows the user to confirm that the order has been successfully completed.

[1560] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1561] The system of the present invention enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. Specific embodiments of the system are shown below.

[1562] System Configuration

[1563] 1. Terminal

[1564] The terminal consists of the user's smartphone or a dedicated device. This terminal is equipped with a microphone for recording animal sounds, an audio processing function for converting them into digital signals, a communication function for sending data to a server, and an emotion engine that recognizes the user's emotions.

[1565] 2. Server

[1566] The server is a computer system that centrally performs data analysis, translation, and emotional data analysis. This server implements an audio data analysis engine, a feature extraction algorithm, a system for mapping to the communication formats of each animal, a mechanism for translating into natural language, and a system for receiving and analyzing emotional data transmitted from the emotional engine.

[1567] Processing flow

[1568] 1. Input

[1569] The user issues a voice command to the device saying, "Please record dolphin sounds." For example, this could be used to record sounds while observing dolphins.

[1570] 2. Preprocessing

[1571] The device converts the recorded audio data into a digital signal and performs noise reduction processing. This improves the accuracy of the analysis.

[1572] 3. Data transmission

[1573] The device sends pre-processed audio data and user emotion data analyzed by the emotion engine to a server via the internet. Because stable communication is required, Wi-Fi or mobile data communication is used.

[1574] 4. Voice Analysis

[1575] The server analyzes the received audio data and extracts features. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[1576] 5. Feature Mapping

[1577] The server uses the extracted features to map them to animal-specific vocal patterns. This allows it to estimate the animal's intended message.

[1578] 6. Translation

[1579] The server infers the intent and translates it into natural language. For example, it might generate a message like, "There are lots of fish."

[1580] 7. Response generation

[1581] The server receives input and emotion data from the user and converts it into an animal sound format. If the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates a relaxing sound data such as "There are lots of fish if you head east."

[1582] 8. Output

[1583] The device receives the translation results and generated sound data from the server and presents them to the user. The playback function allows the sounds to be transmitted to animals.

[1584] Specific example

[1585] When analyzing dolphin sounds, the user instructs the device to "record dolphin sounds." The device records the sounds, converts them into digital signals, and removes noise. The pre-processed data and the user's emotional data analyzed by the emotion engine are then sent to the server, where the audio data is analyzed. The server extracts the characteristics of the dolphin sounds and uses them to translate a message like "There are lots of fish" into natural language. Also, if the user asks "Where should I go?" and the emotion engine detects the user's excited state, the server generates an appropriate dolphin sound in response to that question, providing information to help the user relax.

[1586] The above describes a specific embodiment of the system of the present invention. This system not only facilitates smooth communication with animals but also enables appropriate interaction that takes into account the user's emotional state.

[1587] The following describes the processing flow.

[1588] Step 1:

[1589] The user speaks into the device and says, "Please record dolphin sounds." The device uses its voice recognition function to detect the user's command and enters recording mode.

[1590] Step 2:

[1591] The device uses its built-in microphone to record ambient sounds. Recording continues for a period set by the user, or until manually stopped.

[1592] Step 3:

[1593] The device converts the recorded analog audio data into a digital signal. This is done using an AD converter (analog-to-digital converter).

[1594] Step 4:

[1595] The terminal performs noise reduction processing on the converted digital audio data. A filtering algorithm is used to reduce background noise.

[1596] Step 5:

[1597] Simultaneously, the device analyzes the user's voice, facial expressions, and gestures using an emotion engine to recognize their emotional state. It analyzes tone and speed from the voice, facial movements from the expressions, and body movements from the gestures.

[1598] Step 6:

[1599] The terminal packages pre-processed audio data and user emotion data analyzed by the emotion engine into data packets.

[1600] Step 7:

[1601] The terminal sends packaged data to the server via the internet. Stable communication is required.

[1602] Step 8:

[1603] The server receives data packets sent from the terminal. After receiving the data, it reconstructs it and prepares it for analysis.

[1604] Step 9:

[1605] The server uses a speech analysis engine to extract features from the received speech data. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[1606] Step 10:

[1607] The server maps the extracted data to animal-specific communication patterns. This is done using a pre-trained machine learning model.

[1608] Step 11:

[1609] The server estimates the animal's intentions from the mapped data and translates those intentions into natural language that the user can understand. For example, it might generate a message like, "There are lots of fish around."

[1610] Step 12:

[1611] The server receives user input and emotional data, and converts it into animal sound formats. For example, if a user asks excitedly, "Tell me where I should go," the server will use dolphin sounds to generate a calming message such as, "There are lots of fish if you head east."

[1612] Step 13:

[1613] The server sends the generated sound data and translated text data to the terminal. The terminal analyzes the received data and prepares for playback.

[1614] Step 14:

[1615] The device plays generated animal sound data and simultaneously displays translated text information to the user. The user can communicate directly with the animal through the device.

[1616] (Example 2)

[1617] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1618] Traditional methods of communicating with animals have the problem that it is difficult for users to understand the meaning of animal sounds. Furthermore, because communication is conducted without considering the user's emotional state, it can be stressful for the user. In addition, the ability to properly analyze and translate animal sounds is not yet fully realized. Therefore, there is a need to develop a system that accurately grasps the intentions of animals and conveys them to the user in an easily understandable way.

[1619] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1620] In this invention, the server includes means for analyzing audio data and extracting features, means for mapping the extracted features to the communication format of each animal, and means for the server to analyze the user's emotional data. This makes it possible to analyze animal sounds and generate appropriate responses based on the user's emotional state.

[1621] "Means for recording animal sounds" refers to devices that collect sounds emitted by animals using acoustic sensors such as microphones.

[1622] "Means for converting recorded animal sounds into digital signals" refers to devices or software that convert collected animal sound audio data from analog to digital format through digital signal processing.

[1623] "Methods for removing noise from digital signals" refer to filtering algorithms and software used to remove unwanted background noise and other unwanted sounds from converted digital signals.

[1624] "Means for transmitting pre-processed audio data to a server" refers to communication devices and protocols for transferring noise-reduced audio data to a server using internet communication or wireless communication.

[1625] "Methods for analyzing audio data and extracting features" refer to software that performs analysis on audio data on a server to extract important features such as the frequency characteristics and temporal characteristics of the audio.

[1626] "Means for mapping extracted features to animal-specific communication formats" refers to algorithms or software that associate extracted vocal features with animal-specific vocalization patterns.

[1627] "Means for translating mapped data into natural language" refers to software that converts data obtained based on animal communication formats into natural language that humans can understand.

[1628] "Means for a server to analyze user emotional data" refers to software or algorithms that allow a server to analyze emotional data received from a user and understand the user's emotional state.

[1629] "Means for transmitting generated sound data to a terminal" refers to communication devices and protocols for transferring animal sound data generated by a server to a terminal using internet communication or wireless communication.

[1630] "Means for playing back data transmitted to a terminal" refers to output devices such as speakers or playback software that play back the sound data and translation data received by the terminal to the user or animal.

[1631] This invention describes a system that enables smoother communication between humans and animals by analyzing and translating animal sounds, and further recognizing and adjusting to the user's emotions. This system mainly includes the following components.

[1632] System Configuration

[1633] 1. Terminal

[1634] It consists of the user's smartphone or dedicated device. This device has the following functions:

[1635] Microphone: An acoustic sensor used to record animal sounds.

[1636] Audio processing function: A digital signal processing (DSP) algorithm that converts recorded audio data into a digital signal and removes noise.

[1637] Communication functions: Wi-Fi and mobile data communication functions for sending pre-processed voice data and user emotion data to the server.

[1638] Emotion engine: Software used to recognize and analyze a user's emotions.

[1639] 2. Server

[1640] A central computer system that performs data analysis, translation, and sentiment data analysis. It includes the following functions:

[1641] Speech analysis engine: Software that analyzes transmitted speech data and performs spectrogram analysis and Mel-frequency cepstrum coefficient (MFCC) calculations to extract features.

[1642] Feature extraction algorithm: An algorithm that associates extracted features with the vocalization patterns of different animals.

[1643] Natural language translation system: Software that estimates the intentions of animals and translates those intentions into natural language.

[1644] Emotional Data Analysis System: A system that analyzes user emotional data transmitted from an emotional engine and generates responses based on that analysis.

[1645] System processing flow

[1646] The system's processing flow is as follows: First, the user issues a voice command to the terminal, initiating the recording of animal sounds. The recorded voice data is converted into a digital signal and subjected to noise reduction processing. Subsequently, the pre-processed voice data and the user's emotional data are sent from the terminal to the server. The server analyzes the voice data and maps it to animal sound patterns. Next, these patterns are translated into natural language, and the content is presented to the user. In addition, an appropriate response is generated according to the user's emotional state and communicated to the animal using the playback function.

[1647] Specific example

[1648] When analyzing dolphin sounds, the user instructs the device to "record dolphin sounds." The device records the sounds and removes noise using a DSP algorithm. The cleared audio data and the user's emotional data, analyzed by the emotion engine, are then sent to the server. The server analyzes the audio data using spectrogram analysis and MFCC calculations to extract dolphin-specific sound patterns. Based on these patterns, it translates the message "There are lots of fish" into natural language. If the user asks "Where should I go?" and the emotion engine detects an excited state, the server generates appropriate dolphin sounds and provides information to help the user relax.

[1649] Example of a prompt

[1650] The following are specific examples of voice commands that users can issue to their devices.

[1651] "Please record the sounds of dolphins."

[1652] "Please tell me the next destination."

[1653] "Tell me what the dolphins are feeling right now."

[1654] The above describes a specific embodiment of the system of the present invention. This system facilitates smooth communication with animals and enables appropriate interaction that takes into account the user's emotional state.

[1655] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1656] Step 1:

[1657] Input: The user issues a voice command to the device saying, "Please record the sound of a dolphin."

[1658] Operation: The user speaks voice commands into their smartphone or a dedicated device.

[1659] Output: The terminal receives the recording command and begins preparing to record the animal sounds.

[1660] Step 2:

[1661] Input: The terminal receives a recording command from the user.

[1662] Operation: The device's microphone activates and records dolphin sounds. The acoustic sensor inside the smartphone is used during this process.

[1663] Output: Generates recorded analog audio data.

[1664] Step 3:

[1665] Input: Recorded analog audio data.

[1666] Operation: The terminal converts this analog audio data into a digital signal. Digital signal processing (DSP) algorithms are used in this process.

[1667] Output: Generates audio data converted into a digital signal.

[1668] Step 4:

[1669] Input: Audio data converted to a digital signal.

[1670] Operation: The terminal performs noise reduction processing on this digital signal. Unnecessary background noise and static are filtered out.

[1671] Output: Generates clear audio data with noise reduction.

[1672] Step 5:

[1673] Input: Denoised audio data and user emotion data analyzed by the emotion engine.

[1674] Operation: The device collects this data and sends it to the server via internet communication. Wi-Fi or mobile data communication is used for this purpose.

[1675] Output: Emotional data and audio data sent to the server.

[1676] Step 6:

[1677] Input: Audio data sent to the server.

[1678] Operation: The server analyzes audio data using an audio analysis engine. Specifically, it performs spectrogram analysis and calculates Mel-frequency cepstrum coefficients (MFCCs).

[1679] Output: Extracted speech feature data.

[1680] Step 7:

[1681] Input: Extracted speech feature data.

[1682] Operation: The server maps this feature data to animal-specific vocal patterns. The data is then transformed into dolphin-specific communication patterns.

[1683] Output: Mapped communication data.

[1684] Step 8:

[1685] Input: Mapped communication data.

[1686] Operation: The server translates this into natural language. For example, it can generate a specific message like "There are lots of fish" from the sound of dolphins.

[1687] Output: Translated natural language message.

[1688] Step 9:

[1689] Input: Server-generated natural language messages and user sentiment data.

[1690] Operation: The server generates appropriate animal sound formats based on user input and emotional data. For example, if it detects that the user is agitated, it will generate sound data to help them relax.

[1691] Output: Audio data converted to animal sound format.

[1692] Step 10:

[1693] Input: Translation results and sound format data sent from the server.

[1694] Operation: The device receives this data and notifies the user. Furthermore, it transmits the sounds generated using the device's playback function to the animals.

[1695] Output: Translated message presented to the user and played sound data.

[1696] The above outlines the specific program processing flow of this system. Through the specific actions at each step, smooth communication between animals and humans is achieved, and interaction that takes into account the user's emotional state becomes possible.

[1697] (Application Example 2)

[1698] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1699] Traditional animal communication systems focus on analyzing animal sounds and translating them into natural language, but this has limitations in real-time interaction between humans and animals and in ensuring safety. Furthermore, there is a lack of means for users to detect their pet's emotional state or abnormal behavior. This makes it difficult to respond quickly when a pet suddenly becomes agitated or makes suspicious noises, potentially leading to safety issues within the home.

[1700] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording animal sounds, means for converting the recorded sounds into digital signals, means for removing noise from the digital signals, means for transmitting pre-processed audio data to the server, means for analyzing the audio data and extracting features, means for mapping the extracted features to a communication format for each animal, means for translating the mapped data into natural language, means for transmitting the translated data and generated sound data to a terminal, means for playing back the data transmitted to the terminal, means for analyzing the emotional state of the animal using the recorded audio data, and means for sending an alert to the user based on the analyzed emotional state. This makes it possible not only to communicate smoothly with animals, but also to quickly detect the emotional state and abnormal behavior of pets, thereby enhancing safety in the home.

[1701] Definitions of important words

[1702] "Means for recording animal sounds" refers to a device or system for electronically capturing animal sounds along with ambient noise and saving them as digital audio files.

[1703] "Means for converting recorded vocalizations into digital signals" refers to a device or technology for converting recorded analog audio data into a digital format.

[1704] "Means for removing noise from digital signals" refers to devices or algorithms for removing unwanted background noise and other unwanted sounds contained in digital audio data, thereby improving the accuracy of the analysis.

[1705] "Means for transmitting pre-processed audio data to a server" refers to a device or system for transmitting audio data that has undergone pre-processing, such as noise reduction, to a central server via the internet or other communication means.

[1706] "Means for analyzing audio data and extracting features" refers to technologies or devices for analyzing audio data and extracting features such as specific frequencies or sound patterns.

[1707] "Means for mapping extracted features to animal-specific communication formats" refers to a device or algorithm for converting extracted vocal features into data corresponding to the communication format specific to that animal.

[1708] "Means for translating mapped data into natural language" refers to devices or technologies for converting data representing animal intentions and emotions into natural language that humans can understand.

[1709] "Means for transmitting translated data and generated sound data to a terminal" refers to a device or system for transmitting data translated into natural language and, if necessary, generated animal sound data to a user's terminal.

[1710] "Means for playing back data transmitted to a terminal" refers to a device or application that allows the user to visually or aurally confirm transmitted audio data or translation results.

[1711] "Means for analyzing an animal's emotional state using recorded audio data" refers to a device or algorithm for analyzing recorded audio data to estimate the animal's current emotional state (e.g., excitement, tension, relaxation).

[1712] "Means for sending alerts to users based on analyzed emotional states" refers to a device or system for sending a warning notification to a user when an abnormal emotional state of an animal is detected.

[1713] Modes for carrying out the invention

[1714] System Configuration

[1715] The system of this invention mainly consists of terminals and servers. This system makes it possible to analyze animal sounds in real time and provide useful information to the user.

[1716] 1. Terminal

[1717] The device consists of the user's smartphone or dedicated device. The device has the following functions:

[1718] Microphone for recording animal sounds

[1719] Audio processing function that converts recorded animal sounds into digital signals.

[1720] A processing function that removes noise from digital signals.

[1721] A communication function that sends pre-processed audio data to the server.

[1722] These features allow for the real-time capture of animal sounds, pre-processing, and transmission of the data to a server.

[1723] 2. Server

[1724] A server is a computer system that centrally performs data analysis, translation, and sentiment data analysis. The server has the following functions:

[1725] Voice data analysis engine

[1726] Feature extraction algorithm

[1727] A system that maps to the communication format of each animal.

[1728] A mechanism for translating into natural language

[1729] Generative AI models for developing emotional engines

[1730] This allows the server to analyze the received audio data and extract its features. Based on these extracted features, it classifies the sounds as animal calls and translates them into natural language. It can also analyze user input and convert their intent into animal calls.

[1731] Analysis and alert function for animal emotional states

[1732] The system also has a function to analyze the emotional state of animals using recorded audio data. This function estimates the animal's current emotions (e.g., excitement, tension, relaxation) from the analyzed vocalization data. It then sends alerts to the user's device, notifying them in real time of abnormal behavior or suspicious vocalizations. This allows the user to take necessary actions quickly. For example, if a pet suddenly becomes agitated or makes a suspicious vocalization, the device can send an alert saying "Suspicious vocalization detected," prompting the user to take immediate action.

[1733] Specific example

[1734] Suppose a user is checking on their pet at home via their smartphone, and the pet suddenly makes an unusual noise. In this case, the device records the noise, converts it to a digital signal, removes noise, and sends it to a server. The server analyzes the audio data, extracts its characteristics, and estimates the pet's emotional state (e.g., fear, excitement). If an abnormality is detected, an alert is sent to the user. The user receives the alert and can take appropriate action quickly.

[1735] Example of a prompt

[1736] Design a model that extracts pet vocalization characteristics from audio data and predicts suspicious vocalizations and emotional states. Use MFCC for audio data preprocessing and build the model using TensorFlow / Keras.

[1737] The above describes the embodiments for carrying out this invention. This not only facilitates smooth communication with animals but also enhances safety within the home.

[1738] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1739] Program processing flow

[1740] Processing flow of the system program that implements the application example

[1741] Step 1:

[1742] The device records animal sounds based on user instructions. The input is an audio signal obtained through a recording device, and the output is a digital audio file. This step includes specific actions to capture animal sounds along with ambient sounds.

[1743] Step 2:

[1744] The terminal converts the recorded audio signal into a digital signal. The input is an analog audio signal, and the output is audio data converted into a digital format (e.g., PCM format). This conversion facilitates subsequent data processing.

[1745] Step 3:

[1746] The terminal removes noise from the digital audio data. The input is digital audio data, and the output is audio data that has been denoised. The specific action performed in this step is to improve the accuracy of the analysis by removing unwanted background noise using a filtering algorithm.

[1747] Step 4:

[1748] The terminal sends pre-processed audio data to the server. The input is denoised digital audio data, and the output is audio data transmitted over the internet. The specific actions in this step include transferring the data to the server via a communication protocol.

[1749] Step 5:

[1750] The server analyzes the received audio data and extracts features. The input is audio data transmitted from the terminal, and the output is audio feature quantities (e.g., MFCC). The analysis engine is used to process the audio waveform and perform the actual calculations to extract features.

[1751] Step 6:

[1752] The server uses extracted features to map them to the communication formats of each animal. The input is the features, and the output is data mapped to the animal's unique vocal patterns. Based on the features, a specific algorithm is implemented to estimate the animal's intentions and emotions.

[1753] Step 7:

[1754] The server translates mapped data into natural language. The input is animal vocalization pattern data, and the output is a text message translated into natural language. This translation process uses a generative AI model to convert what the animals are trying to communicate into human language.

[1755] Step 8:

[1756] The server sends the translated data and generated sound data to the terminal. The input is natural language text and sound data, and the output is data sent to the user's terminal via the internet. This allows the user to see what the animals are trying to communicate in real time.

[1757] Step 9:

[1758] The device plays back data received from the server. Input is natural language text and animal sound data, and output is audio and text presented to the user. Specific actions include displaying text messages and playing animal sound data through the speaker.

[1759] Step 10:

[1760] The server analyzes the emotional state of an animal using recorded audio data. The input is recorded audio data, and the output is an estimated result of the animal's emotional state. An emotion engine is used to perform specific data calculations to estimate emotional states such as excitement and relaxation.

[1761] Step 11:

[1762] Based on the analyzed emotional state, the server sends an alert to the user. The input is the estimated result of the animal's emotional state, and the output is the alert message sent to the user. This message allows the user to take action early.

[1763] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1764] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1765] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1766] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1767] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1768] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1769] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1770] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1771] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1772] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1773] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1774] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1775] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1776] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1777] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1778] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1779] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1780] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1781] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1782] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1783] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1784] The following is further disclosed regarding the embodiments described above.

[1785] (Claim 1)

[1786] Methods for recording animal sounds,

[1787] A means of converting recorded animal sounds into digital signals,

[1788] A means of removing noise from digital signals,

[1789] A means for sending pre-processed audio data to a server,

[1790] A means of analyzing audio data and extracting features,

[1791] A means of mapping extracted features to the communication formats of each animal,

[1792] A means of translating mapped data into natural language,

[1793] A means for transmitting translated data and generated sound data to a terminal,

[1794] A means of playing back data sent to the terminal,

[1795] A system that includes this.

[1796] (Claim 2)

[1797] The system according to claim 1, further comprising means for estimating the intent of an animal and translating the estimated intent into natural language.

[1798] (Claim 3)

[1799] The system according to claim 1, further comprising means for analyzing user input and converting the user's natural language into an animal sound format.

[1800] "Example 1"

[1801] (Claim 1)

[1802] Methods for recording animal sounds,

[1803] A means of converting recorded animal sounds into digital signals,

[1804] A means of removing noise from digital signals,

[1805] A means for sending pre-processed audio data to a server,

[1806] A means of analyzing audio data and extracting features,

[1807] A means of mapping extracted features to the communication formats of each animal,

[1808] A means of translating mapped data into natural language,

[1809] A means for transmitting translated data and generated sound data to a terminal,

[1810] A means of playing back data sent to the terminal,

[1811] A means for the user to input voice commands,

[1812] A means by which the terminal sends pre-processed data to the server,

[1813] The server analyzes the audio data and extracts features,

[1814] A means by which the server maps animal sound patterns,

[1815] A means for the server to translate intent into natural language,

[1816] The means by which the server generates a response,

[1817] A means by which the terminal receives response data and presents the result to the user,

[1818] The means by which the device plays sounds,

[1819] A system that includes this.

[1820] (Claim 2)

[1821] The system according to claim 1, further comprising means for estimating the intent of an animal and translating the estimated intent into natural language.

[1822] (Claim 3)

[1823] The system according to claim 1, further comprising means for analyzing user input and converting the user's natural language into an animal sound format.

[1824] "Application Example 1"

[1825] (Claim 1)

[1826] Methods for recording animal sounds,

[1827] A means of converting recorded animal sounds into digital signals,

[1828] A means of removing noise from digital signals,

[1829] A means for sending pre-processed audio data to a server,

[1830] A means of analyzing audio data and extracting features,

[1831] A means of mapping extracted features to the communication formats of each animal,

[1832] A means of translating mapped data into natural language,

[1833] A means for transmitting translated data and generated sound data to a terminal,

[1834] A means of playing back data sent to the terminal,

[1835] Information and communication processing means for ordering food based on translated data,

[1836] A system that includes this.

[1837] (Claim 2)

[1838] The system according to claim 1, further comprising means for estimating the intent of an animal and translating the estimated intent into natural language.

[1839] (Claim 3)

[1840] The system according to claim 1, further comprising means for analyzing user input and converting the user's natural language into an animal sound format.

[1841] "Example 2 of combining an emotion engine"

[1842] (Claim 1)

[1843] Methods for recording animal sounds,

[1844] A means of converting recorded animal sounds into digital signals,

[1845] A means of removing noise from digital signals,

[1846] A means for sending pre-processed audio data to a server,

[1847] A means of analyzing audio data and extracting features,

[1848] A means of mapping extracted features to the communication formats of each animal,

[1849] A means of translating mapped data into natural language,

[1850] A means for the server to analyze user sentiment data,

[1851] A means of transmitting the generated sound data to the terminal,

[1852] A means of playing back data sent to the terminal,

[1853] A system that includes this.

[1854] (Claim 2)

[1855] The system according to claim 1, further comprising means for estimating the intent of an animal and translating the estimated intent into natural language.

[1856] (Claim 3)

[1857] The system according to claim 1, further comprising means for analyzing user input and sentiment data and converting the user's natural language into an animal sound format.

[1858] "Application example 2 when combining with an emotional engine"

[1859] (Claim 1)

[1860] Methods for recording animal sounds,

[1861] A means of converting recorded animal sounds into digital signals,

[1862] A means of removing noise from digital signals,

[1863] A means for sending pre-processed audio data to a server,

[1864] A means of analyzing audio data and extracting features,

[1865] A means of mapping extracted features to the communication formats of each animal,

[1866] A means of translating mapped data into natural language,

[1867] A means for transmitting translated data and generated sound data to a terminal,

[1868] A means of playing back data sent to the terminal,

[1869] A method for analyzing the emotional state of animals using recorded audio data,

[1870] A means of sending alerts to the user based on the analyzed emotional state,

[1871] A system that includes this.

[1872] (Claim 2)

[1873] The system according to claim 1, further comprising means for estimating the intent of an animal and translating the estimated intent into natural language.

[1874] (Claim 3)

[1875] The system according to claim 1, further comprising means for analyzing user input and converting the user's natural language into an animal sound format. [Explanation of Symbols]

[1876] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Methods for recording animal sounds, A means of converting recorded animal sounds into digital signals, A means of removing noise from digital signals, A means for sending pre-processed audio data to a server, A means of analyzing audio data and extracting features, A means of mapping extracted features to the communication formats of each animal, A means of translating mapped data into natural language, A means for transmitting translated data and generated sound data to a terminal, A means of playing back data sent to the terminal, A system that includes this.

2. The system according to claim 1, further comprising means for estimating the intentions of an animal and translating the estimated intentions into natural language.

3. The system according to claim 1, further comprising means for analyzing user input and converting the user's natural language into an animal sound format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A