System
The system uses a terminal and generative AI model to analyze pet photos and videos, offering quick and accurate health management and emotional understanding, addressing the challenge of managing pet health and emotions.
Patent Information
- Application Number
- JP2024126298
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2026-02-13
AI Technical Summary
Managing pet health and understanding emotions is challenging for animals that do not make sounds, as it is difficult to accurately grasp their feelings and physical condition from gestures and facial expressions, and there is a need for quick and appropriate feedback and solutions when they become ill.
A system comprising a terminal for taking and analyzing photos and videos of pets, a server with a generative AI model to analyze the data, and means for providing feedback and appropriate measures, including information on nearby veterinary clinics when necessary.
Enables quick and accurate analysis of pet feelings and health conditions, providing users with appropriate feedback and advice, and facilitating prompt action when pets are unwell.
Smart Images

Figure 2026023977000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The difficulty of managing pet health and understanding emotions is particularly pronounced for animals that do not make sounds. It is difficult to accurately grasp a pet's feelings and physical condition from their gestures and facial expressions. It is also difficult to find a quick and appropriate way to deal with a pet when it suddenly becomes ill. Therefore, there is a need for technology that can accurately and quickly analyze a pet's feelings and health condition and provide appropriate feedback and solutions. [Means for solving the problem]
[0005] In order to solve the above problems, the present invention provides the following means.
[0006] The system includes a terminal that takes and analyzes photos and videos of pets, a server that receives the photo and video data sent from the terminal and includes a generative AI model that analyzes the data, and means for transmitting feedback on the pet's feelings and health status analyzed by the server to the terminal. Furthermore, the server has the function of providing appropriate measures when the pet is unwell and generating and transmitting information on the nearest veterinary clinic, and the terminal has means for inputting the pet's symptoms, taking videos, and transmitting them to the server, thereby enabling quick and accurate health management and emotional understanding of pets.
[0007] A "terminal" is a device that has the function of taking photos and videos of pets and sending the data to a server.
[0008] A "server" is a device that has the function of receiving photo and video data sent from a terminal, analyzing the data using a generative AI model, and sending the analysis results to the terminal.
[0009] The "generative AI model" is an artificial intelligence model that analyzes received photo and video data to accurately predict your pet's feelings and health condition.
[0010] "Feedback" is information about your pet's mood and health that is generated by the server as an analysis result and sent to your device.
[0011] "How to deal with illness" refers to appropriate measures and responses provided by the server based on the analysis results when a pet shows signs of illness.
[0012] "Veterinary clinic information" is information that enables prompt action when a pet shows signs of illness, such as the location and contact information of the nearest veterinary clinic.
[0013] "Inputting symptoms" refers to the act of a user using a terminal to record specific details of symptoms when a pet is unwell and sending the details to a server. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system mainly includes the following elements: a device that takes photos and videos of pets, a server that contains a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[0036] System configuration and operation
[0037] 1. Photo / video shooting and transmission phase
[0038] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[0039] 2. Data analysis phase
[0040] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[0041] 3. Analysis results feedback phase
[0042] Based on the analysis results of the generative AI model, the server generates feedback information such as "your pet is relaxed," "your pet is excited," or "your pet looks unwell." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen and notifies the user of the pet's status.
[0043] 4. Phase of responding when you are unwell
[0044] Furthermore, if a pet shows signs of illness, the user can use the device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model, which analyzes the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[0045] Specific examples
[0046] Typical feedback examples
[0047] 1. The user takes a video of their cat with their smartphone.
[0048] 2. The device sends the captured video to the server via the app.
[0049] 3. The server receives the video data and passes it to the generative AI model.
[0050] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0051] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0052] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0053] Examples of what to do when you are feeling unwell
[0054] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0055] 2. The device sends this data to the server.
[0056] 3. The server receives the data and inputs it into the generative AI model.
[0057] 4. The server's AI model analyzes that the pet may have indigestion.
[0058] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0059] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0060] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[0061] The processing flow will be explained below.
[0062] Step 1:
[0063] Users use their smartphone camera to take photos or videos of their pets.
[0064] Step 2:
[0065] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[0066] Step 3:
[0067] The terminal uploads the converted data to the server.
[0068] Step 4:
[0069] The server receives the photo and video data sent from the terminal.
[0070] Step 5:
[0071] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[0072] Step 6:
[0073] The server feeds the complete data into a generative AI model.
[0074] Step 7:
[0075] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[0076] Step 8:
[0077] The server generates feedback information based on the analysis results of the generative AI model.
[0078] Step 9:
[0079] The server transmits the generated feedback information to the user's terminal.
[0080] Step 10:
[0081] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[0082] Step 11:
[0083] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[0084] Step 12:
[0085] The device sends the input symptom information and captured video data to the server.
[0086] Step 13:
[0087] The server receives the input symptom information and video data.
[0088] Step 14:
[0089] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[0090] Step 15:
[0091] Based on the analysis results of the generated AI model, the server generates appropriate countermeasures and information on the nearest veterinary clinic.
[0092] Step 16:
[0093] The server sends the generated countermeasure information to the user's terminal.
[0094] Step 17:
[0095] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] There is a need to quickly and accurately grasp the mood and health status of pets and provide users with appropriate feedback and solutions. However, conventional systems lack analytical accuracy and data completeness, which often results in delays in providing information to users. Another problem is the lack of a means to quickly provide appropriate solutions when pets are unwell or information about nearby veterinary clinics.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes means for taking and transmitting photos and videos of the pet using a photographing device, means for checking the integrity of the photo and video data transmitted from the photographing device and requesting retransmission if there is a problem, means for receiving the data whose integrity has been confirmed and analyzing the pet's facial expressions and movements using a generative AI model, and means for generating feedback information based on the analysis results and transmitting it to the user's device. This makes it possible to quickly and accurately analyze the pet's mood and health condition and provide the user with appropriate feedback and advice.
[0101] "Photography device" refers to equipment that can take photos and videos of pets and transmit the data.
[0102] "Means" refers to a method, device, function, etc. used to achieve a specific purpose.
[0103] "Checking the integrity of photo and video data" refers to the process of verifying that the received data is not lost or corrupted and is accurately received from the sender.
[0104] A "retry request" refers to a request to the sender to resend data when the data is incomplete or corrupted.
[0105] A "generative AI model" is a type of artificial intelligence model that uses algorithms or networks to generate results for specific tasks or data analysis.
[0106] "Analyzing facial expressions and movements" refers to the process of analyzing the pet's facial expressions and body movements from the received image and video data, and inferring their meaning and state.
[0107] "Generating feedback information" refers to the process of creating specific information to provide to the user based on the analysis results, such as whether the pet is relaxed or not feeling well.
[0108] "User device" refers to a device that receives analysis results and feedback and notifies the user of them. Specifically, this applies to smartphones and tablets.
[0109] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. This system mainly includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[0110] System configuration and operation
[0111] Photo and video shooting
[0112] Users use devices such as smartphones and tablets to take photos and videos of their pets. This is done using the device's built-in camera function. For example, if a user wants to take a video of their cat, they launch the device's camera app and capture the desired scene.
[0113] Data submission and preprocessing
[0114] Once the photo or video data is captured, the device converts it into the appropriate format within the application and performs pre-processing such as compression. Once pre-processing is complete, the device sends the data to the server, typically using the HTTPS protocol to ensure data protection and security.
[0115] Data reception and integrity check
[0116] The server receives the photo and video data sent from the device. Upon receiving the data, it checks its integrity. For example, it calculates a hash value (such as MD5 or SHA-256) to ensure the data is not corrupted. If the data is incomplete or corrupted, the server issues a resend request and repeats this process until the correct data is sent.
[0117] Data analysis
[0118] Once complete data is obtained, the server analyzes the data using a generative AI model, often based on TensorFlow or PyTorch models. The generative AI model analyzes the pet's facial expressions, gestures, movements, and other features contained in the received data to infer the pet's mood and health. Specifically, it processes images using libraries such as OpenCV and utilizes deep learning-based object detection algorithms (such as YOLO and SSD).
[0119] Generate feedback
[0120] The server generates feedback information based on the analysis results of the generative AI model. For example, the analysis results may include "The cat is relaxed" or "The dog is excited." This information is sent to the user's device via a RESTful API.
[0121] Providing feedback information
[0122] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. For example, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the application.
[0123] What to do when you are unwell
[0124] In particular, if a pet's condition worsens, the user uses their device to enter the pet's symptoms into the application and send additional video data. The server receives this data and again inputs it into the generative AI model to analyze the cause of the illness and specific countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for several hours") and information about nearby veterinary clinics. The analysis results are then sent back to the user's device and displayed on the application screen.
[0125] Specific examples
[0126] Typical feedback examples
[0127] 1. The user takes a video of their cat with their smartphone.
[0128] 2. The device sends the captured video to the server via the app.
[0129] 3. The server receives the video data and passes it to the generative AI model.
[0130] 4. The server's generated AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0131] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0132] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0133] Examples of what to do when you are feeling unwell
[0134] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0135] 2. The device sends this data to the server.
[0136] 3. The server receives the data and inputs it into the generative AI model.
[0137] 4. The server's generative AI model analyzes that "your pet may have indigestion."
[0138] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0139] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0141] Step 1:
[0142] Users take photos and videos of their pets using devices such as smartphones or tablets. Once the photos and videos are taken, the device converts the photo and video data into the appropriate format within the application and performs preprocessing such as compression. This preprocessing reduces the data size and makes transmission more efficient. The input is the pet's photo and video data, and the output is compressed, format-converted data. Specifically, the user takes a video of their cat, converts it to MP4 format, and then compresses it using H.264 encoding. The converted data is then ready to be transmitted.
[0143] Step 2:
[0144] Once preprocessing is complete, the device sends the data to the server. At this time, the HTTPS protocol is used to ensure data protection and security. The input is the format-converted compressed data, and the output is the data sent to the server. Specifically, the data is uploaded to the server by pressing the send button in the application.
[0145] Step 3:
[0146] The server receives photo and video data sent from the device. To check the integrity of the received data, it calculates a hash value (such as MD5 or SHA-256) to see if the data is corrupted. If the data is incomplete or corrupted, the server sends a resend request to the device. The input is the data sent from the device, and the output is data whose integrity has been confirmed or a resend request. Specifically, the server calculates the hash value of the received data, and if it does not match, it sends an HTTP request for resend to the device.
[0147] Step 4:
[0148] Once complete data is obtained, the server analyzes the data using a generative AI model. This model uses a TensorFlow model or a PyTorch model. The generative AI model uses the received data as input and analyzes the pet's facial expressions and movements. The output is the analysis result, which includes information such as "The pet is relaxed" or "The pet is excited." Specific operations include image processing for each frame and running face recognition using OpenCV and deep learning object detection algorithms (such as YOLO and SSD).
[0149] Step 5:
[0150] The server generates feedback information based on the analysis results of the generative AI model. Based on the analysis results, it generates feedback in text format, such as "The cat is relaxed" or "The dog is excited." The input is the analysis results of the generative AI model, and the output is the feedback information. This feedback information is again sent from the server to the user's device. Specifically, the feedback information is sent to the user's device via a RESTful API, and receipt is confirmed.
[0151] Step 6:
[0152] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. The input is the feedback information sent from the server, and the output is the notification displayed on the application screen. Specifically, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the app.
[0153] Step 7:
[0154] If a pet's condition worsens, the user uses the device to enter the pet's symptoms into the application and shoot and send additional video data. The input is text data and video data about the pet's symptoms, and the output is data sent to the server. Specifically, the user enters the symptoms in the text box, shoots additional video, and presses the send button from the application.
[0155] Step 8:
[0156] The server receives the data sent by the user and analyzes it again using the generative AI model. The input is text and video data about the symptoms, and the output is the analysis results regarding the cause of the illness and countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for two hours") and information about nearby veterinary clinics.
[0157] Step 9:
[0158] After the analysis results are generated, the server sends them to the user's device. The input is the analysis result for the poor health condition, and the output is information on countermeasures provided to the user. Specifically, data such as "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____" is sent via a RESTful API.
[0159] Step 10:
[0160] The terminal displays the countermeasure information received from the server on the application screen and notifies the user of specific countermeasures. The input is the countermeasure information sent from the server, and the output is a display of specific countermeasures provided to the user. Specifically, the message displayed is, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0161] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[0162] (Application example 1)
[0163] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0164] Conventional machine monitoring systems in factories typically check the machine's condition visually or through regular inspections. However, these methods make it difficult to respond quickly when an abnormality occurs, making it difficult to prevent machine breakdowns and accidents. In extreme cases, this could lead to a serious accident or the shutdown of the production line, significantly impacting factory operations. To solve this problem, a system is needed that can constantly monitor the machine's condition and quickly detect and notify when an abnormality occurs.
[0165] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0166] In this invention, the server includes a means including a generative AI model that analyzes photos and videos of pets, a means including a generative AI model that analyzes video data from cameras installed in the factory, and a means for transmitting the analyzed machine status to an operator terminal, thereby enabling constant monitoring of the machine status in the factory and immediate detection and notification of any abnormalities that occur.
[0167] A "terminal" is a device that takes photos and videos of pets and machines and sends the data to a server.
[0168] The "server" is a device that receives photo and video data sent from a device, analyzes the data using a generative AI model, and feeds back the results.
[0169] A "generative AI model" is an artificial intelligence model that analyzes received photo and video data and infers the state and feelings of the subject.
[0170] "Pets" refer to animals kept at home and are the subject of analysis in the present invention.
[0171] The term "machine" refers to equipment and devices installed in a factory, and is the subject of analysis in the present invention.
[0172] A "camera" is a device for taking still or video images, and is often built into a terminal.
[0173] "Analysis" is a process carried out to understand the state and feelings of the target object based on the received data.
[0174] "Feedback" refers to the exchange of information to notify the user of the analysis results.
[0175] The "worker terminal" is a terminal used in a factory, and is a device for receiving and displaying analysis results.
[0176] An "abnormality" refers to a state in which the machine is in a different state from its normal state, and early detection is required.
[0177] This invention is a system that analyzes the status of pets and machinery in a factory and immediately notifies you if an abnormality occurs. This system is composed of the following elements.
[0178] System configuration and operation
[0179] 1. Photo / video shooting and transmission phase
[0180] The device takes photos and videos of pets and factory machinery. In the case of pets, the target is animals kept at home. In the case of factory machinery, the target is equipment and devices. Once the photos are taken, the device converts the format of the data (compresses it, etc.) and efficiently sends it to the server. The device has a built-in camera, so the photos and transmission are done automatically. The hardware used includes smartphones and factory surveillance cameras.
[0181] 2. Data analysis phase
[0182] The server receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server. This model is built using deep learning frameworks such as Keras. The generative AI model analyzes the pet's facial expressions and behavior, as well as the operation of machinery in the factory, to infer the pet's feelings and health condition, and any abnormalities in the machinery.
[0183] 3. Analysis results feedback phase
[0184] Based on the analysis results of the generative AI model, the server generates feedback information about the pet's mood and health status, or about abnormal machine conditions. For example, the information may be, "The pet is relaxed," or "There is an abnormality in the machine." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen, informing the user of the pet's condition or the status of the machinery in the factory.
[0185] 4. Phase of responding when you are unwell
[0186] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. As with the previous procedure, the server inputs this data into a generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Avoid feeding certain foods," and "The nearest veterinary clinic is here." In the case of machinery in a factory, the system also suggests appropriate measures to take when an abnormality occurs.
[0187] Specific examples
[0188] For example, a video of a cat is taken using a smartphone and sent to a server. The server receives the video data and passes it to a generative AI model. The resulting analysis concludes that "the cat is relaxed." Furthermore, if a camera on a machine in a factory detects abnormal movement, the server immediately interprets this as "an abnormality has occurred in the machine" and notifies the worker's terminal. In this case, an example of a prompt sentence to be input to the generative AI model is as follows:
[0189] "Analyze the current video footage obtained from the robot's camera and detect if there are any abnormal conditions."
[0190] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0191] Step 1:
[0192] The device takes photos and videos of pets and factory machinery. Specifically, it uses the device's built-in camera to capture continuous or still image data. This data is compressed and pre-processed for efficiency. The input is the video of the pet or machinery, and the output is compressed video data.
[0193] Step 2:
[0194] The terminal transmits the captured video data to the server. Specifically, the video data is uploaded to the server via a network. The input is the preprocessed video data, and the output is the data transmitted to the server.
[0195] Step 3:
[0196] The server receives the video data sent from the terminal. Here, the data integrity is checked. If incomplete data is detected, a retransmission is requested. The input is the video data sent from the terminal, and the output is the complete video data.
[0197] Step 4:
[0198] The server inputs the received video data into a generative AI model for analysis. Specifically, it uses deep learning frameworks such as Keras to perform data preprocessing, feature extraction, and analysis. The input is the complete video data, and the output is the analysis results.
[0199] Step 5:
[0200] The server generates feedback information about the pet's mood and health, or about abnormal conditions in the machine, based on the analysis results of the generative AI model. Specifically, it converts the analysis results into text format and prepares the feedback in a format that is easy for the user to understand. The input is the analysis results, and the output is the feedback information.
[0201] Step 6:
[0202] The server sends the generated feedback information to the user's terminal. Specifically, the server uploads the feedback information to the terminal through the network. The input is the feedback information, and the output is the information sent to the terminal.
[0203] Step 7:
[0204] The device displays the received feedback information on the application screen and notifies the user of the status of the pet or machine. Specific operations include adjusting the layout and arranging the content for display on the screen. The input is the feedback information sent from the server, and the output is the feedback content displayed to the user.
[0205] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0206] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[0207] System configuration and operation
[0208] 1. Photo / video shooting and transmission phase
[0209] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[0210] 2. Data analysis phase
[0211] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[0212] 3. User Emotion Recognition Phase
[0213] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the pet's situation. This data is obtained from the user's device.
[0214] 4. Analysis results feedback phase
[0215] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device. The device then displays the received feedback on the application screen and notifies the user of the pet's status.
[0216] 5. Phase of responding when you are unwell
[0217] Furthermore, if a pet shows signs of illness, the user uses their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____." The server then uses an emotion engine to customize this information according to the user's emotional state and generate recommended action information. Finally, the server sends the recommended action information to the user's device, which displays it on the app screen.
[0218] Specific examples
[0219] Typical feedback examples
[0220] 1. The user takes a video of their cat with their smartphone.
[0221] 2. The device sends the captured video to the server via the app.
[0222] 3. The server receives the video data and passes it to the generative AI model.
[0223] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0224] 5. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is relaxed.
[0225] 6. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0226] 7. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0227] Examples of what to do when you are feeling unwell
[0228] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0229] 2. The device sends this data to the server.
[0230] 3. The server receives the data and inputs it into the generative AI model.
[0231] 4. The server's AI model analyzes that the pet may have indigestion.
[0232] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0233] 6. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is worried.
[0234] 7. The server customizes countermeasure information based on the user's emotional state and sends it to the user's device.
[0235] 8. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0236] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[0237] The processing flow will be explained below.
[0238] Step 1:
[0239] Users use their smartphone camera to take photos or videos of their pets.
[0240] Step 2:
[0241] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[0242] Step 3:
[0243] The terminal uploads the converted data to the server.
[0244] Step 4:
[0245] The server receives the photo and video data sent from the terminal.
[0246] Step 5:
[0247] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[0248] Step 6:
[0249] The server feeds the complete data into a generative AI model.
[0250] Step 7:
[0251] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[0252] Step 8:
[0253] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[0254] Step 9:
[0255] The server combines the analysis results of the generative AI model and the emotion engine to generate feedback information.
[0256] Step 10:
[0257] The server transmits the generated feedback information to the user's terminal.
[0258] Step 11:
[0259] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[0260] Step 12:
[0261] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[0262] Step 13:
[0263] The device sends the input symptom information and captured video data to the server.
[0264] Step 14:
[0265] The server receives the input symptom information and video data.
[0266] Step 15:
[0267] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[0268] Step 16:
[0269] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[0270] Step 17:
[0271] The server generates customized countermeasure information according to the user's emotional state based on the countermeasure information and the results of the emotion engine.
[0272] Step 18:
[0273] The server sends the generated countermeasure information to the user's terminal.
[0274] Step 19:
[0275] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[0276] Example 2
[0277] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0278] It is difficult to quickly and accurately grasp the health condition and mood of a pet, and there is a lack of technology to provide appropriate feedback according to the user's emotional state. Furthermore, when a pet becomes ill, there is a need for a means to quickly provide appropriate measures and alleviate the user's anxiety.
[0279] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0280] In this invention, the server includes a device for a user to take photos and videos of their pet, a means including a generative AI model that receives the photo and video data sent from the device and checks and analyzes the data for completeness, a means for transmitting feedback information on the pet's mood and health status analyzed by the server to the device, and a means for the server to analyze the user's emotional state and provide customized feedback based on that state. This makes it possible to quickly and accurately grasp the pet's health status and mood. Furthermore, appropriate feedback can be provided according to the user's emotional state, reducing the user's anxiety.
[0281] A "user" is someone who uses the system to take photos and videos of their pet and receive the analysis results.
[0282] A "terminal" is a photographic device used by a user, such as a smartphone or tablet.
[0283] "Photo and video data" refers to visual data of pets taken using a device.
[0284] A "server" is a computer system that receives and analyzes photo and video data sent from a terminal.
[0285] The "generative AI model" is an artificial intelligence model that resides on a server and analyzes the received data to predict the pet's feelings and health condition.
[0286] "Data integrity" means that the data received is in the correct format and is not missing or corrupted.
[0287] "Analysis" is the process of using a generative AI model to analyze a pet's facial expressions and movements to infer its feelings and health.
[0288] "Feedback information" is information about the pet's feelings and health condition that is generated by the server based on the analysis results and provided to the user.
[0289] The "emotion engine" is an algorithm that analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[0290] "Customized feedback" is feedback information that is tailored to the user's emotional state.
[0291] "What to do when your pet is unwell" refers to instructions and advice on how to respond appropriately when your pet becomes ill.
[0292] "Information about the nearest veterinary clinic" is information about the location and contact details of the nearest veterinary clinic where you can get your pet examined.
[0293] MODE FOR CARRYING OUT THE INVENTION
[0294] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[0295] Photo / video shooting and transmission phase
[0296] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos or videos are taken, the device converts the data format (compresses it, etc.) before sending it to the server via the application, efficiently uploading the data to the server.
[0297] A concrete example is a process in which a user takes a video of their cat on their smartphone, the device compresses the video, and sends it to a server using an HTTP POST request.
[0298] Data Analysis Phase
[0299] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[0300] For example, the generative AI model in the server uses deep learning frameworks such as TensorFlow to analyze a video of a cat and determine that the cat is relaxed.
[0301] User emotion recognition phase
[0302] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the situation of their pet. This data is obtained from the user's device.
[0303] As a specific example, the server collects facial expression data and tone of voice data from the user's device, inputs it into the emotion engine for analysis, and evaluates the user's emotional state by analyzing whether the user is relaxed while watching the video.
[0304] Analysis results feedback phase
[0305] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device, which then displays the received feedback on the application screen and notifies the user of the pet's status.
[0306] For example, a message might appear saying, "Your cat is relaxing. Come enjoy some relaxation time with him."
[0307] Phases of response when feeling unwell
[0308] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[0309] The server uses an emotion engine to customize this information according to the user's emotional state and generate countermeasure information. Finally, the server sends the countermeasure information to the user's device, which then displays it on the app screen.
[0310] For example, you might see feedback like, "Your pet may have indigestion. Keep your pet hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0311] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[0312] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0313] Step 1:
[0314] The user takes photos and videos of their pet using the device. The user launches the camera app on their smartphone and captures the pet's behavior. The photos and videos are then saved as digital data on the device.
[0315] Input: Pet photos and videos
[0316] Output: Data stored in the device
[0317] Step 2:
[0318] The device saves the captured data in the application. The device saves the captured photos and videos in a specific folder so that they can be sent to the server later.
[0319] Input: Photo and video data
[0320] Output: Data stored within the application
[0321] Step 3:
[0322] Before the device sends the captured data to the server, it performs format conversion, compressing and encoding the data to convert it into a format that can be transmitted efficiently.
[0323] Input: Data stored within the application
[0324] Output: Compressed and encoded data
[0325] Step 4:
[0326] The device sends the converted data to the server, using an HTTP POST request to upload the data to the server.
[0327] Input: Compressed and encoded data
[0328] Output: Data sent to the server
[0329] Step 5:
[0330] The server receives the data sent from the device, takes in the data, and stores it in storage.
[0331] Input: Data sent from the terminal
[0332] Output: Data stored in the server
[0333] Step 6:
[0334] The server checks the integrity of the data, ensuring that the data received is not corrupted and is in the correct format, and if there is a problem it issues a request to resend it.
[0335] Input: Data stored on the server
[0336] Output: Data integrity check result (normal or resend request)
[0337] Step 7:
[0338] The server inputs the data into the generative AI model, passing the data to be analyzed to the AI model and starting the analysis process.
[0339] Input: Normal data
[0340] Output: The data fed into the generative AI model
[0341] Step 8:
[0342] A generative AI model analyzes the data, analyzing your pet's facial expressions and movements to infer its mood and health.
[0343] Input: Data fed into a generative AI model
[0344] Output: Analysis results (predictions of pet's mood and health)
[0345] Step 9:
[0346] The server recognizes the user's emotional state based on the analysis results, using an emotion engine to analyze the user's facial expressions, tone of voice, etc.
[0347] Input: Analysis results, and user facial expression and tone of voice data
[0348] Output: User's emotional state analysis results
[0349] Step 10:
[0350] The server integrates the analysis results and the emotional state, and generates feedback information based on the pet's state and the user's emotional state.
[0351] Input: Analysis results, user emotional state analysis results
[0352] Output: Integrated feedback information
[0353] Step 11:
[0354] The server sends feedback information to the user's device, and information based on the analysis results is sent to the device and notified to the user.
[0355] Input: Integrated feedback information
[0356] Output: Feedback information sent to the terminal
[0357] Step 12:
[0358] The terminal displays the feedback information to the user. The sent feedback information is displayed on the application screen to notify the user.
[0359] Input: Feedback information sent to the device
[0360] Output: Feedback information displayed on the application screen
[0361] (Application example 2)
[0362] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0363] When pets are left alone at home, it is difficult for owners to monitor their pets' health and emotions in real time and ensure their safety. Furthermore, if a pet shows signs of anxiety or poor health while the owner is away for an extended period of time, there is a lack of means to take prompt and appropriate action. This increases the risk of overlooking abnormal behavior or health problems in pets, and also reduces the sense of security for pet owners.
[0364] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal that takes and analyzes photos and videos of the pet, means including a generative AI model that receives photo and video data transmitted from the terminal and analyzes the data, means for transmitting feedback on the pet's feelings and health status analyzed by the server to the terminal, means including an emotion engine that collects and analyzes user emotion data, and means for the server to generate and transmit customized feedback information based on the analysis results and the user's emotional state. This makes it possible to monitor the pet's health status and emotions in real time and to respond quickly if an abnormality occurs. Furthermore, providing feedback according to the user's emotional state can provide a sense of security to the owner.
[0365] A "terminal" is a device that takes photos and videos of pets and sends the data to a server.
[0366] The "generative AI model" is an artificial intelligence model that analyzes received pet photos and video data to predict the pet's feelings and health condition.
[0367] A "server" is a central computer system that receives and analyzes data sent from terminals.
[0368] The "emotion engine" is an analytical engine that analyzes the user's facial expressions and tone of voice to understand the user's emotional state.
[0369] "Feedback information" is information about the pet's status that is generated by the server based on the analysis results and the user's emotional state.
[0370] "Customized feedback information" is feedback information whose content is adjusted according to the user's emotional state.
[0371] A specific system for implementing this invention will be described. The entire system is composed of multiple means, including a "terminal," a "server," and an "emotion engine." The purpose of this system is to analyze photos and videos of pets and provide feedback to users in real time.
[0372] 1. Photo / video shooting and transmission phase
[0373] Users take photos and videos of their pets using a device such as a smartphone. The captured data is converted into an appropriate format (JPEG, MP4, etc.) on the device. The device then sends the data to a server. This transmission is performed using the HTTP or RTSP protocol.
[0374] 2. Data analysis phase
[0375] The server receives the photo and video data sent from the device. It checks the integrity of the received data and requests a resend if it is incomplete. Once the data is verified, it is input into a generative AI model. This generative AI model uses machine learning algorithms to analyze the pet's facial expressions and movements, and performs data calculations to infer the pet's mood and health. Specifically, libraries such as TensorFlow and PyTorch are used.
[0376] 3. User emotion recognition phase
[0377] The server also incorporates an emotion engine that analyzes facial images and voice data acquired from the user's device. This analysis uses facial expression recognition and voice analysis technologies (OpenCV, DeepFace, SpeechRecognition, etc.). Based on this data, the user's emotional state is understood.
[0378] 4. Analysis results feedback phase
[0379] The server generates customized feedback information based on the analysis results of the generative AI model and the emotion engine. For example, if the user is worried, detailed analysis results and reassuring advice are provided. The generated feedback information is sent from the server to the device, which then displays it on the application screen.
[0380] 5. Response phase when you feel unwell
[0381] If a pet shows signs of illness, the user uses the device to input the pet's symptoms, take additional video footage, and send it to the server. The server then inputs the received data back into the generative AI model to analyze the cause of the pet's illness and the appropriate course of action. The analysis results include a specific action plan (e.g., "keep the pet hydrated and rest for several hours," "avoid feeding certain foods," etc.). Information about the nearest veterinary clinic is also provided. Using an emotion engine, customized information is generated according to the user's emotional state and sent to the device.
[0382] Specific examples
[0383] The following is an example of a prompt that the user enters:
[0384] Example prompt sentence:
[0385] "It analyzes pet behavior to determine if the pet is relaxed or not, and if the user is feeling anxious, it provides reassuring advice about the pet's situation."
[0386] In this way, this system allows users to safely monitor their pet's condition in real time and take prompt and appropriate action when necessary.
[0387] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0388] Step 1:
[0389] A user takes photos and videos of their pet using a device such as a smartphone.
[0390] Specific behavior:
[0391] The user launches the camera app on their smartphone and takes photos or videos of their pet. The captured data is converted into a format (e.g., JPEG or MP4).
[0392] Input: Pet photos and videos
[0393] Output: Format-converted photo and video data
[0394] Step 2:
[0395] The device sends the captured data to the server.
[0396] Specific behavior:
[0397] The terminal compresses the data and sends it to the server using the HTTP or RTSP protocol.
[0398] Input: Format-converted photo and video data
[0399] Output: Data sent to the server
[0400] Step 3:
[0401] The server checks the integrity of the photo and video data received.
[0402] Specific behavior:
[0403] The server calculates a checksum for the data and compares it with the received data, and if any incomplete data is found, it requests a retransmission.
[0404] Input: Data sent to the server
[0405] Output: Data with integrity checked
[0406] Step 4:
[0407] The server analyzes the data using the generative AI model.
[0408] Specific behavior:
[0409] The server then inputs the verified data into a generative AI model to analyze the pet's facial expressions and movements, using TensorFlow or PyTorch, for example, to perform data calculations to infer the pet's emotions and health status.
[0410] Input: Integrity checked data
[0411] Output: Analysis of pet's mood and health condition
[0412] Step 5:
[0413] The server compares the analysis results with the user's emotion engine.
[0414] Specific behavior:
[0415] The server inputs facial images and voice data acquired from the user's device into the emotion engine and analyzes the user's emotional state. The analysis uses libraries such as OpenCV, DeepFace, and SpeechRecognition.
[0416] Input: User's facial image and voice data
[0417] Output: User's emotional state
[0418] Step 6:
[0419] The server generates customized feedback information based on the analysis results and the user's emotional state.
[0420] Specific behavior:
[0421] The server generates feedback information based on the analysis results of the generative AI model and the emotion engine, depending on the user's emotional state. For example, if the user is worried, the server provides reassuring advice such as, "Your pet seems a little anxious, but there is nothing serious going on."
[0422] Input: Analysis results of pet's feelings and health condition, user's emotional state
[0423] Output: Customized feedback information
[0424] Step 7:
[0425] The server transmits the generated feedback information to the terminal.
[0426] Specific behavior:
[0427] The server transmits the generated feedback information to the user's terminal using the HTTP protocol or the like.
[0428] Input:Customized feedback information
[0429] Output: Feedback information sent to the device
[0430] Step 8:
[0431] The feedback information received by the terminal is displayed on the application screen.
[0432] Specific behavior:
[0433] The device analyzes the feedback information received from the server and displays it on the application screen, allowing the user to check it and understand the status of their pet.
[0434] Input: Feedback information sent to the device
[0435] Output: Feedback information displayed on the screen
[0436] The above is the processing flow of the system of the present invention. Through the specific operations performed at each step, it is possible to monitor the health and emotions of pets in real time and take prompt action if necessary.
[0437] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0438] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0439] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0440] [Second embodiment]
[0441] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0442] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0443] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0444] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0445] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0446] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0447] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0448] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0449] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0450] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0451] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0452] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0453] This invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system mainly includes the following elements: a device that takes photos and videos of pets, a server that contains a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[0454] System configuration and operation
[0455] 1. Photo / video shooting and transmission phase
[0456] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[0457] 2. Data analysis phase
[0458] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[0459] 3. Analysis results feedback phase
[0460] Based on the analysis results of the generative AI model, the server generates feedback information such as "your pet is relaxed," "your pet is excited," or "your pet looks unwell." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen and notifies the user of the pet's status.
[0461] 4. Phase of responding when you are unwell
[0462] Furthermore, if a pet shows signs of illness, the user can use the device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model, which analyzes the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[0463] Specific examples
[0464] Typical feedback examples
[0465] 1. The user takes a video of their cat with their smartphone.
[0466] 2. The device sends the captured video to the server via the app.
[0467] 3. The server receives the video data and passes it to the generative AI model.
[0468] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0469] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0470] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0471] Examples of what to do when you are feeling unwell
[0472] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0473] 2. The device sends this data to the server.
[0474] 3. The server receives the data and inputs it into the generative AI model.
[0475] 4. The server's AI model analyzes that the pet may have indigestion.
[0476] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0477] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0478] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[0479] The processing flow will be explained below.
[0480] Step 1:
[0481] Users use their smartphone camera to take photos or videos of their pets.
[0482] Step 2:
[0483] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[0484] Step 3:
[0485] The terminal uploads the converted data to the server.
[0486] Step 4:
[0487] The server receives the photo and video data sent from the terminal.
[0488] Step 5:
[0489] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[0490] Step 6:
[0491] The server feeds the complete data into a generative AI model.
[0492] Step 7:
[0493] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[0494] Step 8:
[0495] The server generates feedback information based on the analysis results of the generative AI model.
[0496] Step 9:
[0497] The server transmits the generated feedback information to the user's terminal.
[0498] Step 10:
[0499] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[0500] Step 11:
[0501] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[0502] Step 12:
[0503] The device sends the input symptom information and captured video data to the server.
[0504] Step 13:
[0505] The server receives the input symptom information and video data.
[0506] Step 14:
[0507] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[0508] Step 15:
[0509] Based on the analysis results of the generated AI model, the server generates appropriate countermeasures and information on the nearest veterinary clinic.
[0510] Step 16:
[0511] The server sends the generated countermeasure information to the user's terminal.
[0512] Step 17:
[0513] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[0514] Example 1
[0515] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0516] There is a need to quickly and accurately grasp the mood and health status of pets and provide users with appropriate feedback and solutions. However, conventional systems lack analytical accuracy and data completeness, which often results in delays in providing information to users. Another problem is the lack of a means to quickly provide appropriate solutions when pets are unwell or information about nearby veterinary clinics.
[0517] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0518] In this invention, the server includes means for taking and transmitting photos and videos of the pet using a photographing device, means for checking the integrity of the photo and video data transmitted from the photographing device and requesting retransmission if there is a problem, means for receiving the data whose integrity has been confirmed and analyzing the pet's facial expressions and movements using a generative AI model, and means for generating feedback information based on the analysis results and transmitting it to the user's device. This makes it possible to quickly and accurately analyze the pet's mood and health condition and provide the user with appropriate feedback and advice.
[0519] "Photography device" refers to equipment that can take photos and videos of pets and transmit the data.
[0520] "Means" refers to a method, device, function, etc. used to achieve a specific purpose.
[0521] "Checking the integrity of photo and video data" refers to the process of verifying that the received data is not lost or corrupted and is accurately received from the sender.
[0522] A "retry request" refers to a request to the sender to resend data when the data is incomplete or corrupted.
[0523] A "generative AI model" is a type of artificial intelligence model that uses algorithms or networks to generate results for specific tasks or data analysis.
[0524] "Analyzing facial expressions and movements" refers to the process of analyzing the pet's facial expressions and body movements from the received image and video data, and inferring their meaning and state.
[0525] "Generating feedback information" refers to the process of creating specific information to provide to the user based on the analysis results, such as whether the pet is relaxed or not feeling well.
[0526] "User device" refers to a device that receives analysis results and feedback and notifies the user of them. Specifically, this applies to smartphones and tablets.
[0527] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. This system mainly includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[0528] System configuration and operation
[0529] Photo and video shooting
[0530] Users use devices such as smartphones and tablets to take photos and videos of their pets. This is done using the device's built-in camera function. For example, if a user wants to take a video of their cat, they launch the device's camera app and capture the desired scene.
[0531] Data submission and preprocessing
[0532] Once the photo or video data is captured, the device converts it into the appropriate format within the application and performs pre-processing such as compression. Once pre-processing is complete, the device sends the data to the server, typically using the HTTPS protocol to ensure data protection and security.
[0533] Data reception and integrity check
[0534] The server receives the photo and video data sent from the device. Upon receiving the data, it checks its integrity. For example, it calculates a hash value (such as MD5 or SHA-256) to ensure the data is not corrupted. If the data is incomplete or corrupted, the server issues a resend request and repeats this process until the correct data is sent.
[0535] Data analysis
[0536] Once complete data is obtained, the server analyzes the data using a generative AI model, often based on TensorFlow or PyTorch models. The generative AI model analyzes the pet's facial expressions, gestures, movements, and other features contained in the received data to infer the pet's mood and health. Specifically, it processes images using libraries such as OpenCV and utilizes deep learning-based object detection algorithms (such as YOLO and SSD).
[0537] Generate feedback
[0538] The server generates feedback information based on the analysis results of the generative AI model. For example, the analysis results may include "The cat is relaxed" or "The dog is excited." This information is sent to the user's device via a RESTful API.
[0539] Providing feedback information
[0540] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. For example, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the application.
[0541] What to do when you are unwell
[0542] In particular, if a pet's condition worsens, the user uses their device to enter the pet's symptoms into the application and send additional video data. The server receives this data and again inputs it into the generative AI model to analyze the cause of the illness and specific countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for several hours") and information about nearby veterinary clinics. The analysis results are then sent back to the user's device and displayed on the application screen.
[0543] Specific examples
[0544] Typical feedback examples
[0545] 1. The user takes a video of their cat with their smartphone.
[0546] 2. The device sends the captured video to the server via the app.
[0547] 3. The server receives the video data and passes it to the generative AI model.
[0548] 4. The server's generated AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0549] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0550] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0551] Examples of what to do when you are feeling unwell
[0552] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0553] 2. The device sends this data to the server.
[0554] 3. The server receives the data and inputs it into the generative AI model.
[0555] 4. The server's generative AI model analyzes that "your pet may have indigestion."
[0556] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0557] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0558] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0559] Step 1:
[0560] Users take photos and videos of their pets using devices such as smartphones or tablets. Once the photos and videos are taken, the device converts the photo and video data into the appropriate format within the application and performs preprocessing such as compression. This preprocessing reduces the data size and makes transmission more efficient. The input is the pet's photo and video data, and the output is compressed, format-converted data. Specifically, the user takes a video of their cat, converts it to MP4 format, and then compresses it using H.264 encoding. The converted data is then ready to be transmitted.
[0561] Step 2:
[0562] Once preprocessing is complete, the device sends the data to the server. At this time, the HTTPS protocol is used to ensure data protection and security. The input is the format-converted compressed data, and the output is the data sent to the server. Specifically, the data is uploaded to the server by pressing the send button in the application.
[0563] Step 3:
[0564] The server receives photo and video data sent from the device. To check the integrity of the received data, it calculates a hash value (such as MD5 or SHA-256) to see if the data is corrupted. If the data is incomplete or corrupted, the server sends a resend request to the device. The input is the data sent from the device, and the output is data whose integrity has been confirmed or a resend request. Specifically, the server calculates the hash value of the received data, and if it does not match, it sends an HTTP request for resend to the device.
[0565] Step 4:
[0566] Once complete data is obtained, the server analyzes the data using a generative AI model. This model uses a TensorFlow model or a PyTorch model. The generative AI model uses the received data as input and analyzes the pet's facial expressions and movements. The output is the analysis result, which includes information such as "The pet is relaxed" or "The pet is excited." Specific operations include image processing for each frame and running face recognition using OpenCV and deep learning object detection algorithms (such as YOLO and SSD).
[0567] Step 5:
[0568] The server generates feedback information based on the analysis results of the generative AI model. Based on the analysis results, it generates feedback in text format, such as "The cat is relaxed" or "The dog is excited." The input is the analysis results of the generative AI model, and the output is the feedback information. This feedback information is again sent from the server to the user's device. Specifically, the feedback information is sent to the user's device via a RESTful API, and receipt is confirmed.
[0569] Step 6:
[0570] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. The input is the feedback information sent from the server, and the output is the notification displayed on the application screen. Specifically, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the app.
[0571] Step 7:
[0572] If a pet's condition worsens, the user uses the device to enter the pet's symptoms into the application and shoot and send additional video data. The input is text data and video data about the pet's symptoms, and the output is data sent to the server. Specifically, the user enters the symptoms in the text box, shoots additional video, and presses the send button from the application.
[0573] Step 8:
[0574] The server receives the data sent by the user and analyzes it again using the generative AI model. The input is text and video data about the symptoms, and the output is the analysis results regarding the cause of the illness and countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for two hours") and information about nearby veterinary clinics.
[0575] Step 9:
[0576] After the analysis results are generated, the server sends them to the user's device. The input is the analysis result for the poor health condition, and the output is information on countermeasures provided to the user. Specifically, data such as "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____" is sent via a RESTful API.
[0577] Step 10:
[0578] The terminal displays the countermeasure information received from the server on the application screen and notifies the user of specific countermeasures. The input is the countermeasure information sent from the server, and the output is a display of specific countermeasures provided to the user. Specifically, the message displayed is, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0579] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[0580] (Application example 1)
[0581] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0582] Conventional machine monitoring systems in factories typically check the machine's condition visually or through regular inspections. However, these methods make it difficult to respond quickly when an abnormality occurs, making it difficult to prevent machine breakdowns and accidents. In extreme cases, this could lead to a serious accident or the shutdown of the production line, significantly impacting factory operations. To solve this problem, a system is needed that can constantly monitor the machine's condition and quickly detect and notify when an abnormality occurs.
[0583] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0584] In this invention, the server includes a means including a generative AI model that analyzes photos and videos of pets, a means including a generative AI model that analyzes video data from cameras installed in the factory, and a means for transmitting the analyzed machine status to an operator terminal, thereby enabling constant monitoring of the machine status in the factory and immediate detection and notification of any abnormalities that occur.
[0585] A "terminal" is a device that takes photos and videos of pets and machines and sends the data to a server.
[0586] The "server" is a device that receives photo and video data sent from a device, analyzes the data using a generative AI model, and feeds back the results.
[0587] A "generative AI model" is an artificial intelligence model that analyzes received photo and video data and infers the state and feelings of the subject.
[0588] "Pets" refer to animals kept at home and are the subject of analysis in the present invention.
[0589] The term "machine" refers to equipment and devices installed in a factory, and is the subject of analysis in the present invention.
[0590] A "camera" is a device for taking still or video images, and is often built into a terminal.
[0591] "Analysis" is a process carried out to understand the state and feelings of the target object based on the received data.
[0592] "Feedback" refers to the exchange of information to notify the user of the analysis results.
[0593] The "worker terminal" is a terminal used in a factory, and is a device for receiving and displaying analysis results.
[0594] An "abnormality" refers to a state in which the machine is in a different state from its normal state, and early detection is required.
[0595] This invention is a system that analyzes the status of pets and machinery in a factory and immediately notifies you if an abnormality occurs. This system is composed of the following elements.
[0596] System configuration and operation
[0597] 1. Photo / video shooting and transmission phase
[0598] The device takes photos and videos of pets and factory machinery. In the case of pets, the target is animals kept at home. In the case of factory machinery, the target is equipment and devices. Once the photos are taken, the device converts the format of the data (compresses it, etc.) and efficiently sends it to the server. The device has a built-in camera, so the photos and transmission are done automatically. The hardware used includes smartphones and factory surveillance cameras.
[0599] 2. Data analysis phase
[0600] The server receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server. This model is built using deep learning frameworks such as Keras. The generative AI model analyzes the pet's facial expressions and behavior, as well as the operation of machinery in the factory, to infer the pet's feelings and health condition, and any abnormalities in the machinery.
[0601] 3. Analysis results feedback phase
[0602] Based on the analysis results of the generative AI model, the server generates feedback information about the pet's mood and health status, or about abnormal machine conditions. For example, the information may be, "The pet is relaxed," or "There is an abnormality in the machine." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen, informing the user of the pet's condition or the status of the machinery in the factory.
[0603] 4. Phase of responding when you are unwell
[0604] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. As with the previous procedure, the server inputs this data into a generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Avoid feeding certain foods," and "The nearest veterinary clinic is here." In the case of machinery in a factory, the system also suggests appropriate measures to take when an abnormality occurs.
[0605] Specific examples
[0606] For example, a video of a cat is taken using a smartphone and sent to a server. The server receives the video data and passes it to a generative AI model. The resulting analysis concludes that "the cat is relaxed." Furthermore, if a camera on a machine in a factory detects abnormal movement, the server immediately interprets this as "an abnormality has occurred in the machine" and notifies the worker's terminal. In this case, an example of a prompt sentence to be input to the generative AI model is as follows:
[0607] "Analyze the current video footage obtained from the robot's camera and detect if there are any abnormal conditions."
[0608] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0609] Step 1:
[0610] The device takes photos and videos of pets and factory machinery. Specifically, it uses the device's built-in camera to capture continuous or still image data. This data is compressed and pre-processed for efficiency. The input is the video of the pet or machinery, and the output is compressed video data.
[0611] Step 2:
[0612] The terminal transmits the captured video data to the server. Specifically, the video data is uploaded to the server via a network. The input is the preprocessed video data, and the output is the data transmitted to the server.
[0613] Step 3:
[0614] The server receives the video data sent from the terminal. Here, the data integrity is checked. If incomplete data is detected, a retransmission is requested. The input is the video data sent from the terminal, and the output is the complete video data.
[0615] Step 4:
[0616] The server inputs the received video data into a generative AI model for analysis. Specifically, it uses deep learning frameworks such as Keras to perform data preprocessing, feature extraction, and analysis. The input is the complete video data, and the output is the analysis results.
[0617] Step 5:
[0618] The server generates feedback information about the pet's mood and health, or about abnormal conditions in the machine, based on the analysis results of the generative AI model. Specifically, it converts the analysis results into text format and prepares the feedback in a format that is easy for the user to understand. The input is the analysis results, and the output is the feedback information.
[0619] Step 6:
[0620] The server sends the generated feedback information to the user's terminal. Specifically, the server uploads the feedback information to the terminal through the network. The input is the feedback information, and the output is the information sent to the terminal.
[0621] Step 7:
[0622] The device displays the received feedback information on the application screen and notifies the user of the status of the pet or machine. Specific operations include adjusting the layout and arranging the content for display on the screen. The input is the feedback information sent from the server, and the output is the feedback content displayed to the user.
[0623] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0624] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[0625] System configuration and operation
[0626] 1. Photo / video shooting and transmission phase
[0627] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[0628] 2. Data analysis phase
[0629] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[0630] 3. User Emotion Recognition Phase
[0631] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the pet's situation. This data is obtained from the user's device.
[0632] 4. Analysis results feedback phase
[0633] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device. The device then displays the received feedback on the application screen and notifies the user of the pet's status.
[0634] 5. Phase of responding when you are unwell
[0635] Furthermore, if a pet shows signs of illness, the user uses their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____." The server then uses an emotion engine to customize this information according to the user's emotional state and generate recommended action information. Finally, the server sends the recommended action information to the user's device, which displays it on the app screen.
[0636] Specific examples
[0637] Typical feedback examples
[0638] 1. The user takes a video of their cat with their smartphone.
[0639] 2. The device sends the captured video to the server via the app.
[0640] 3. The server receives the video data and passes it to the generative AI model.
[0641] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0642] 5. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is relaxed.
[0643] 6. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0644] 7. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0645] Examples of what to do when you are feeling unwell
[0646] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0647] 2. The device sends this data to the server.
[0648] 3. The server receives the data and inputs it into the generative AI model.
[0649] 4. The server's AI model analyzes that the pet may have indigestion.
[0650] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0651] 6. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is worried.
[0652] 7. The server customizes countermeasure information based on the user's emotional state and sends it to the user's device.
[0653] 8. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0654] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[0655] The processing flow will be explained below.
[0656] Step 1:
[0657] Users use their smartphone camera to take photos or videos of their pets.
[0658] Step 2:
[0659] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[0660] Step 3:
[0661] The terminal uploads the converted data to the server.
[0662] Step 4:
[0663] The server receives the photo and video data sent from the terminal.
[0664] Step 5:
[0665] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[0666] Step 6:
[0667] The server feeds the complete data into a generative AI model.
[0668] Step 7:
[0669] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[0670] Step 8:
[0671] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[0672] Step 9:
[0673] The server combines the analysis results of the generative AI model and the emotion engine to generate feedback information.
[0674] Step 10:
[0675] The server transmits the generated feedback information to the user's terminal.
[0676] Step 11:
[0677] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[0678] Step 12:
[0679] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[0680] Step 13:
[0681] The device sends the input symptom information and captured video data to the server.
[0682] Step 14:
[0683] The server receives the input symptom information and video data.
[0684] Step 15:
[0685] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[0686] Step 16:
[0687] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[0688] Step 17:
[0689] The server generates customized countermeasure information according to the user's emotional state based on the countermeasure information and the results of the emotion engine.
[0690] Step 18:
[0691] The server sends the generated countermeasure information to the user's terminal.
[0692] Step 19:
[0693] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[0694] Example 2
[0695] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0696] It is difficult to quickly and accurately grasp the health condition and mood of a pet, and there is a lack of technology to provide appropriate feedback according to the user's emotional state. Furthermore, when a pet becomes ill, there is a need for a means to quickly provide appropriate measures and alleviate the user's anxiety.
[0697] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0698] In this invention, the server includes a device for a user to take photos and videos of their pet, a means including a generative AI model that receives the photo and video data sent from the device and checks and analyzes the data for completeness, a means for transmitting feedback information on the pet's mood and health status analyzed by the server to the device, and a means for the server to analyze the user's emotional state and provide customized feedback based on that state. This makes it possible to quickly and accurately grasp the pet's health status and mood. Furthermore, appropriate feedback can be provided according to the user's emotional state, reducing the user's anxiety.
[0699] A "user" is someone who uses the system to take photos and videos of their pet and receive the analysis results.
[0700] A "terminal" is a photographic device used by a user, such as a smartphone or tablet.
[0701] "Photo and video data" refers to visual data of pets taken using a device.
[0702] A "server" is a computer system that receives and analyzes photo and video data sent from a terminal.
[0703] The "generative AI model" is an artificial intelligence model that resides on a server and analyzes the received data to predict the pet's feelings and health condition.
[0704] "Data integrity" means that the data received is in the correct format and is not missing or corrupted.
[0705] "Analysis" is the process of using a generative AI model to analyze a pet's facial expressions and movements to infer its feelings and health.
[0706] "Feedback information" is information about the pet's feelings and health condition that is generated by the server based on the analysis results and provided to the user.
[0707] The "emotion engine" is an algorithm that analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[0708] "Customized feedback" is feedback information that is tailored to the user's emotional state.
[0709] "What to do when your pet is unwell" refers to instructions and advice on how to respond appropriately when your pet becomes ill.
[0710] "Information about the nearest veterinary clinic" is information about the location and contact details of the nearest veterinary clinic where you can get your pet examined.
[0711] MODE FOR CARRYING OUT THE INVENTION
[0712] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[0713] Photo / video shooting and transmission phase
[0714] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos or videos are taken, the device converts the data format (compresses it, etc.) before sending it to the server via the application, efficiently uploading the data to the server.
[0715] A concrete example is a process in which a user takes a video of their cat on their smartphone, the device compresses the video, and sends it to a server using an HTTP POST request.
[0716] Data Analysis Phase
[0717] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[0718] For example, the generative AI model in the server uses deep learning frameworks such as TensorFlow to analyze a video of a cat and determine that the cat is relaxed.
[0719] User emotion recognition phase
[0720] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the situation of their pet. This data is obtained from the user's device.
[0721] As a specific example, the server collects facial expression data and tone of voice data from the user's device, inputs it into the emotion engine for analysis, and evaluates the user's emotional state by analyzing whether the user is relaxed while watching the video.
[0722] Analysis results feedback phase
[0723] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device, which then displays the received feedback on the application screen and notifies the user of the pet's status.
[0724] For example, a message might appear saying, "Your cat is relaxing. Come enjoy some relaxation time with him."
[0725] Phases of response when feeling unwell
[0726] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[0727] The server uses an emotion engine to customize this information according to the user's emotional state and generate countermeasure information. Finally, the server sends the countermeasure information to the user's device, which then displays it on the app screen.
[0728] For example, you might see feedback like, "Your pet may have indigestion. Keep your pet hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0729] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[0730] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0731] Step 1:
[0732] The user takes photos and videos of their pet using the device. The user launches the camera app on their smartphone and captures the pet's behavior. The photos and videos are then saved as digital data on the device.
[0733] Input: Pet photos and videos
[0734] Output: Data stored in the device
[0735] Step 2:
[0736] The device saves the captured data in the application. The device saves the captured photos and videos in a specific folder so that they can be sent to the server later.
[0737] Input: Photo and video data
[0738] Output: Data stored within the application
[0739] Step 3:
[0740] Before the device sends the captured data to the server, it performs format conversion, compressing and encoding the data to convert it into a format that can be transmitted efficiently.
[0741] Input: Data stored within the application
[0742] Output: Compressed and encoded data
[0743] Step 4:
[0744] The device sends the converted data to the server, using an HTTP POST request to upload the data to the server.
[0745] Input: Compressed and encoded data
[0746] Output: Data sent to the server
[0747] Step 5:
[0748] The server receives the data sent from the device, takes in the data, and stores it in storage.
[0749] Input: Data sent from the terminal
[0750] Output: Data stored in the server
[0751] Step 6:
[0752] The server checks the integrity of the data, ensuring that the data received is not corrupted and is in the correct format, and if there is a problem it issues a request to resend it.
[0753] Input: Data stored on the server
[0754] Output: Data integrity check result (normal or resend request)
[0755] Step 7:
[0756] The server inputs the data into the generative AI model, passing the data to be analyzed to the AI model and starting the analysis process.
[0757] Input: Normal data
[0758] Output: The data fed into the generative AI model
[0759] Step 8:
[0760] A generative AI model analyzes the data, analyzing your pet's facial expressions and movements to infer its mood and health.
[0761] Input: Data fed into a generative AI model
[0762] Output: Analysis results (predictions of pet's mood and health)
[0763] Step 9:
[0764] The server recognizes the user's emotional state based on the analysis results, using an emotion engine to analyze the user's facial expressions, tone of voice, etc.
[0765] Input: Analysis results, and user facial expression and tone of voice data
[0766] Output: User's emotional state analysis results
[0767] Step 10:
[0768] The server integrates the analysis results and the emotional state, and generates feedback information based on the pet's state and the user's emotional state.
[0769] Input: Analysis results, user emotional state analysis results
[0770] Output: Integrated feedback information
[0771] Step 11:
[0772] The server sends feedback information to the user's device, and information based on the analysis results is sent to the device and notified to the user.
[0773] Input: Integrated feedback information
[0774] Output: Feedback information sent to the terminal
[0775] Step 12:
[0776] The terminal displays the feedback information to the user. The sent feedback information is displayed on the application screen to notify the user.
[0777] Input: Feedback information sent to the device
[0778] Output: Feedback information displayed on the application screen
[0779] (Application example 2)
[0780] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0781] When pets are left alone at home, it is difficult for owners to monitor their pets' health and emotions in real time and ensure their safety. Furthermore, if a pet shows signs of anxiety or poor health while the owner is away for an extended period of time, there is a lack of means to take prompt and appropriate action. This increases the risk of overlooking abnormal behavior or health problems in pets, and also reduces the sense of security for pet owners.
[0782] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal that takes and analyzes photos and videos of the pet, means including a generative AI model that receives photo and video data transmitted from the terminal and analyzes the data, means for transmitting feedback on the pet's feelings and health status analyzed by the server to the terminal, means including an emotion engine that collects and analyzes user emotion data, and means for the server to generate and transmit customized feedback information based on the analysis results and the user's emotional state. This makes it possible to monitor the pet's health status and emotions in real time and to respond quickly if an abnormality occurs. Furthermore, providing feedback according to the user's emotional state can provide a sense of security to the owner.
[0783] A "terminal" is a device that takes photos and videos of pets and sends the data to a server.
[0784] The "generative AI model" is an artificial intelligence model that analyzes received pet photos and video data to predict the pet's feelings and health condition.
[0785] A "server" is a central computer system that receives and analyzes data sent from terminals.
[0786] The "emotion engine" is an analytical engine that analyzes the user's facial expressions and tone of voice to understand the user's emotional state.
[0787] "Feedback information" is information about the pet's status that is generated by the server based on the analysis results and the user's emotional state.
[0788] "Customized feedback information" is feedback information whose content is adjusted according to the user's emotional state.
[0789] A specific system for implementing this invention will be described. The entire system is composed of multiple means, including a "terminal," a "server," and an "emotion engine." The purpose of this system is to analyze photos and videos of pets and provide feedback to users in real time.
[0790] 1. Photo / video shooting and transmission phase
[0791] Users take photos and videos of their pets using a device such as a smartphone. The captured data is converted into an appropriate format (JPEG, MP4, etc.) on the device. The device then sends the data to a server. This transmission is performed using the HTTP or RTSP protocol.
[0792] 2. Data analysis phase
[0793] The server receives the photo and video data sent from the device. It checks the integrity of the received data and requests a resend if it is incomplete. Once the data is verified, it is input into a generative AI model. This generative AI model uses machine learning algorithms to analyze the pet's facial expressions and movements, and performs data calculations to infer the pet's mood and health. Specifically, libraries such as TensorFlow and PyTorch are used.
[0794] 3. User emotion recognition phase
[0795] The server also incorporates an emotion engine that analyzes facial images and voice data acquired from the user's device. This analysis uses facial expression recognition and voice analysis technologies (OpenCV, DeepFace, SpeechRecognition, etc.). Based on this data, the user's emotional state is understood.
[0796] 4. Analysis results feedback phase
[0797] The server generates customized feedback information based on the analysis results of the generative AI model and the emotion engine. For example, if the user is worried, detailed analysis results and reassuring advice are provided. The generated feedback information is sent from the server to the device, which then displays it on the application screen.
[0798] 5. Response phase when you feel unwell
[0799] If a pet shows signs of illness, the user uses the device to input the pet's symptoms, take additional video footage, and send it to the server. The server then inputs the received data back into the generative AI model to analyze the cause of the pet's illness and the appropriate course of action. The analysis results include a specific action plan (e.g., "keep the pet hydrated and rest for several hours," "avoid feeding certain foods," etc.). Information about the nearest veterinary clinic is also provided. Using an emotion engine, customized information is generated according to the user's emotional state and sent to the device.
[0800] Specific examples
[0801] The following is an example of a prompt that the user enters:
[0802] Example prompt sentence:
[0803] "It analyzes pet behavior to determine if the pet is relaxed or not, and if the user is feeling anxious, it provides reassuring advice about the pet's situation."
[0804] In this way, this system allows users to safely monitor their pet's condition in real time and take prompt and appropriate action when necessary.
[0805] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0806] Step 1:
[0807] A user takes photos and videos of their pet using a device such as a smartphone.
[0808] Specific behavior:
[0809] The user launches the camera app on their smartphone and takes photos or videos of their pet. The captured data is converted into a format (e.g., JPEG or MP4).
[0810] Input: Pet photos and videos
[0811] Output: Format-converted photo and video data
[0812] Step 2:
[0813] The device sends the captured data to the server.
[0814] Specific behavior:
[0815] The terminal compresses the data and sends it to the server using the HTTP or RTSP protocol.
[0816] Input: Format-converted photo and video data
[0817] Output: Data sent to the server
[0818] Step 3:
[0819] The server checks the integrity of the photo and video data received.
[0820] Specific behavior:
[0821] The server calculates a checksum for the data and compares it with the received data, and if any incomplete data is found, it requests a retransmission.
[0822] Input: Data sent to the server
[0823] Output: Data with integrity checked
[0824] Step 4:
[0825] The server analyzes the data using the generative AI model.
[0826] Specific behavior:
[0827] The server then inputs the verified data into a generative AI model to analyze the pet's facial expressions and movements, using TensorFlow or PyTorch, for example, to perform data calculations to infer the pet's emotions and health status.
[0828] Input: Integrity checked data
[0829] Output: Analysis of pet's mood and health condition
[0830] Step 5:
[0831] The server compares the analysis results with the user's emotion engine.
[0832] Specific behavior:
[0833] The server inputs facial images and voice data acquired from the user's device into the emotion engine and analyzes the user's emotional state. The analysis uses libraries such as OpenCV, DeepFace, and SpeechRecognition.
[0834] Input: User's facial image and voice data
[0835] Output: User's emotional state
[0836] Step 6:
[0837] The server generates customized feedback information based on the analysis results and the user's emotional state.
[0838] Specific behavior:
[0839] The server generates feedback information based on the analysis results of the generative AI model and the emotion engine, depending on the user's emotional state. For example, if the user is worried, the server provides reassuring advice such as, "Your pet seems a little anxious, but there is nothing serious going on."
[0840] Input: Analysis results of pet's feelings and health condition, user's emotional state
[0841] Output: Customized feedback information
[0842] Step 7:
[0843] The server transmits the generated feedback information to the terminal.
[0844] Specific behavior:
[0845] The server transmits the generated feedback information to the user's terminal using the HTTP protocol or the like.
[0846] Input:Customized feedback information
[0847] Output: Feedback information sent to the device
[0848] Step 8:
[0849] The feedback information received by the terminal is displayed on the application screen.
[0850] Specific behavior:
[0851] The device analyzes the feedback information received from the server and displays it on the application screen, allowing the user to check it and understand the status of their pet.
[0852] Input: Feedback information sent to the device
[0853] Output: Feedback information displayed on the screen
[0854] The above is the processing flow of the system of the present invention. Through the specific operations performed at each step, it is possible to monitor the health and emotions of pets in real time and take prompt action if necessary.
[0855] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0856] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0857] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0858] [Third embodiment]
[0859] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0860] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0861] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0862] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0863] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0864] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0865] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0866] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0867] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0868] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0869] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0870] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0871] This invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system mainly includes the following elements: a device that takes photos and videos of pets, a server that contains a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[0872] System configuration and operation
[0873] 1. Photo / video shooting and transmission phase
[0874] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[0875] 2. Data analysis phase
[0876] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[0877] 3. Analysis results feedback phase
[0878] Based on the analysis results of the generative AI model, the server generates feedback information such as "your pet is relaxed," "your pet is excited," or "your pet looks unwell." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen and notifies the user of the pet's status.
[0879] 4. Phase of responding when you are unwell
[0880] Furthermore, if a pet shows signs of illness, the user can use the device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model, which analyzes the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[0881] Specific examples
[0882] Typical feedback examples
[0883] 1. The user takes a video of their cat with their smartphone.
[0884] 2. The device sends the captured video to the server via the app.
[0885] 3. The server receives the video data and passes it to the generative AI model.
[0886] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0887] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0888] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0889] Examples of what to do when you are feeling unwell
[0890] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0891] 2. The device sends this data to the server.
[0892] 3. The server receives the data and inputs it into the generative AI model.
[0893] 4. The server's AI model analyzes that the pet may have indigestion.
[0894] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0895] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0896] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[0897] The processing flow will be explained below.
[0898] Step 1:
[0899] Users use their smartphone camera to take photos or videos of their pets.
[0900] Step 2:
[0901] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[0902] Step 3:
[0903] The terminal uploads the converted data to the server.
[0904] Step 4:
[0905] The server receives the photo and video data sent from the terminal.
[0906] Step 5:
[0907] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[0908] Step 6:
[0909] The server feeds the complete data into a generative AI model.
[0910] Step 7:
[0911] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[0912] Step 8:
[0913] The server generates feedback information based on the analysis results of the generative AI model.
[0914] Step 9:
[0915] The server transmits the generated feedback information to the user's terminal.
[0916] Step 10:
[0917] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[0918] Step 11:
[0919] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[0920] Step 12:
[0921] The device sends the input symptom information and captured video data to the server.
[0922] Step 13:
[0923] The server receives the input symptom information and video data.
[0924] Step 14:
[0925] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[0926] Step 15:
[0927] Based on the analysis results of the generated AI model, the server generates appropriate countermeasures and information on the nearest veterinary clinic.
[0928] Step 16:
[0929] The server sends the generated countermeasure information to the user's terminal.
[0930] Step 17:
[0931] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[0932] Example 1
[0933] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0934] There is a need to quickly and accurately grasp the mood and health status of pets and provide users with appropriate feedback and solutions. However, conventional systems lack analytical accuracy and data completeness, which often results in delays in providing information to users. Another problem is the lack of a means to quickly provide appropriate solutions when pets are unwell or information about nearby veterinary clinics.
[0935] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0936] In this invention, the server includes means for taking and transmitting photos and videos of the pet using a photographing device, means for checking the integrity of the photo and video data transmitted from the photographing device and requesting retransmission if there is a problem, means for receiving the data whose integrity has been confirmed and analyzing the pet's facial expressions and movements using a generative AI model, and means for generating feedback information based on the analysis results and transmitting it to the user's device. This makes it possible to quickly and accurately analyze the pet's mood and health condition and provide the user with appropriate feedback and advice.
[0937] "Photography device" refers to equipment that can take photos and videos of pets and transmit the data.
[0938] "Means" refers to a method, device, function, etc. used to achieve a specific purpose.
[0939] "Checking the integrity of photo and video data" refers to the process of verifying that the received data is not lost or corrupted and is accurately received from the sender.
[0940] A "retry request" refers to a request to the sender to resend data when the data is incomplete or corrupted.
[0941] A "generative AI model" is a type of artificial intelligence model that uses algorithms or networks to generate results for specific tasks or data analysis.
[0942] "Analyzing facial expressions and movements" refers to the process of analyzing the pet's facial expressions and body movements from the received image and video data, and inferring their meaning and state.
[0943] "Generating feedback information" refers to the process of creating specific information to provide to the user based on the analysis results, such as whether the pet is relaxed or not feeling well.
[0944] "User device" refers to a device that receives analysis results and feedback and notifies the user of them. Specifically, this applies to smartphones and tablets.
[0945] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. This system mainly includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[0946] System configuration and operation
[0947] Photo and video shooting
[0948] Users use devices such as smartphones and tablets to take photos and videos of their pets. This is done using the device's built-in camera function. For example, if a user wants to take a video of their cat, they launch the device's camera app and capture the desired scene.
[0949] Data submission and preprocessing
[0950] Once the photo or video data is captured, the device converts it into the appropriate format within the application and performs pre-processing such as compression. Once pre-processing is complete, the device sends the data to the server, typically using the HTTPS protocol to ensure data protection and security.
[0951] Data reception and integrity check
[0952] The server receives the photo and video data sent from the device. Upon receiving the data, it checks its integrity. For example, it calculates a hash value (such as MD5 or SHA-256) to ensure the data is not corrupted. If the data is incomplete or corrupted, the server issues a resend request and repeats this process until the correct data is sent.
[0953] Data analysis
[0954] Once complete data is obtained, the server analyzes the data using a generative AI model, often based on TensorFlow or PyTorch models. The generative AI model analyzes the pet's facial expressions, gestures, movements, and other features contained in the received data to infer the pet's mood and health. Specifically, it processes images using libraries such as OpenCV and utilizes deep learning-based object detection algorithms (such as YOLO and SSD).
[0955] Generate feedback
[0956] The server generates feedback information based on the analysis results of the generative AI model. For example, the analysis results may include "The cat is relaxed" or "The dog is excited." This information is sent to the user's device via a RESTful API.
[0957] Providing feedback information
[0958] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. For example, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the application.
[0959] What to do when you are unwell
[0960] In particular, if a pet's condition worsens, the user uses their device to enter the pet's symptoms into the application and send additional video data. The server receives this data and again inputs it into the generative AI model to analyze the cause of the illness and specific countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for several hours") and information about nearby veterinary clinics. The analysis results are then sent back to the user's device and displayed on the application screen.
[0961] Specific examples
[0962] Typical feedback examples
[0963] 1. The user takes a video of their cat with their smartphone.
[0964] 2. The device sends the captured video to the server via the app.
[0965] 3. The server receives the video data and passes it to the generative AI model.
[0966] 4. The server's generated AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[0967] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[0968] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[0969] Examples of what to do when you are feeling unwell
[0970] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[0971] 2. The device sends this data to the server.
[0972] 3. The server receives the data and inputs it into the generative AI model.
[0973] 4. The server's generative AI model analyzes that "your pet may have indigestion."
[0974] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[0975] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0976] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0977] Step 1:
[0978] Users take photos and videos of their pets using devices such as smartphones or tablets. Once the photos and videos are taken, the device converts the photo and video data into the appropriate format within the application and performs preprocessing such as compression. This preprocessing reduces the data size and makes transmission more efficient. The input is the pet's photo and video data, and the output is compressed, format-converted data. Specifically, the user takes a video of their cat, converts it to MP4 format, and then compresses it using H.264 encoding. The converted data is then ready to be transmitted.
[0979] Step 2:
[0980] Once preprocessing is complete, the device sends the data to the server. At this time, the HTTPS protocol is used to ensure data protection and security. The input is the format-converted compressed data, and the output is the data sent to the server. Specifically, the data is uploaded to the server by pressing the send button in the application.
[0981] Step 3:
[0982] The server receives photo and video data sent from the device. To check the integrity of the received data, it calculates a hash value (such as MD5 or SHA-256) to see if the data is corrupted. If the data is incomplete or corrupted, the server sends a resend request to the device. The input is the data sent from the device, and the output is data whose integrity has been confirmed or a resend request. Specifically, the server calculates the hash value of the received data, and if it does not match, it sends an HTTP request for resend to the device.
[0983] Step 4:
[0984] Once complete data is obtained, the server analyzes the data using a generative AI model. This model uses a TensorFlow model or a PyTorch model. The generative AI model uses the received data as input and analyzes the pet's facial expressions and movements. The output is the analysis result, which includes information such as "The pet is relaxed" or "The pet is excited." Specific operations include image processing for each frame and running face recognition using OpenCV and deep learning object detection algorithms (such as YOLO and SSD).
[0985] Step 5:
[0986] The server generates feedback information based on the analysis results of the generative AI model. Based on the analysis results, it generates feedback in text format, such as "The cat is relaxed" or "The dog is excited." The input is the analysis results of the generative AI model, and the output is the feedback information. This feedback information is again sent from the server to the user's device. Specifically, the feedback information is sent to the user's device via a RESTful API, and receipt is confirmed.
[0987] Step 6:
[0988] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. The input is the feedback information sent from the server, and the output is the notification displayed on the application screen. Specifically, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the app.
[0989] Step 7:
[0990] If a pet's condition worsens, the user uses the device to enter the pet's symptoms into the application and shoot and send additional video data. The input is text data and video data about the pet's symptoms, and the output is data sent to the server. Specifically, the user enters the symptoms in the text box, shoots additional video, and presses the send button from the application.
[0991] Step 8:
[0992] The server receives the data sent by the user and analyzes it again using the generative AI model. The input is text and video data about the symptoms, and the output is the analysis results regarding the cause of the illness and countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for two hours") and information about nearby veterinary clinics.
[0993] Step 9:
[0994] After the analysis results are generated, the server sends them to the user's device. The input is the analysis result for the poor health condition, and the output is information on countermeasures provided to the user. Specifically, data such as "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____" is sent via a RESTful API.
[0995] Step 10:
[0996] The terminal displays the countermeasure information received from the server on the application screen and notifies the user of specific countermeasures. The input is the countermeasure information sent from the server, and the output is a display of specific countermeasures provided to the user. Specifically, the message displayed is, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[0997] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[0998] (Application example 1)
[0999] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1000] Conventional machine monitoring systems in factories typically check the machine's condition visually or through regular inspections. However, these methods make it difficult to respond quickly when an abnormality occurs, making it difficult to prevent machine breakdowns and accidents. In extreme cases, this could lead to a serious accident or the shutdown of the production line, significantly impacting factory operations. To solve this problem, a system is needed that can constantly monitor the machine's condition and quickly detect and notify when an abnormality occurs.
[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1002] In this invention, the server includes a means including a generative AI model that analyzes photos and videos of pets, a means including a generative AI model that analyzes video data from cameras installed in the factory, and a means for transmitting the analyzed machine status to an operator terminal, thereby enabling constant monitoring of the machine status in the factory and immediate detection and notification of any abnormalities that occur.
[1003] A "terminal" is a device that takes photos and videos of pets and machines and sends the data to a server.
[1004] The "server" is a device that receives photo and video data sent from a device, analyzes the data using a generative AI model, and feeds back the results.
[1005] A "generative AI model" is an artificial intelligence model that analyzes received photo and video data and infers the state and feelings of the subject.
[1006] "Pets" refer to animals kept at home and are the subject of analysis in the present invention.
[1007] The term "machine" refers to equipment and devices installed in a factory, and is the subject of analysis in the present invention.
[1008] A "camera" is a device for taking still or video images, and is often built into a terminal.
[1009] "Analysis" is a process carried out to understand the state and feelings of the target object based on the received data.
[1010] "Feedback" refers to the exchange of information to notify the user of the analysis results.
[1011] The "worker terminal" is a terminal used in a factory, and is a device for receiving and displaying analysis results.
[1012] An "abnormality" refers to a state in which the machine is in a different state from its normal state, and early detection is required.
[1013] This invention is a system that analyzes the status of pets and machinery in a factory and immediately notifies you if an abnormality occurs. This system is composed of the following elements.
[1014] System configuration and operation
[1015] 1. Photo / video shooting and transmission phase
[1016] The device takes photos and videos of pets and factory machinery. In the case of pets, the target is animals kept at home. In the case of factory machinery, the target is equipment and devices. Once the photos are taken, the device converts the format of the data (compresses it, etc.) and efficiently sends it to the server. The device has a built-in camera, so the photos and transmission are done automatically. The hardware used includes smartphones and factory surveillance cameras.
[1017] 2. Data analysis phase
[1018] The server receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server. This model is built using deep learning frameworks such as Keras. The generative AI model analyzes the pet's facial expressions and behavior, as well as the operation of machinery in the factory, to infer the pet's feelings and health condition, and any abnormalities in the machinery.
[1019] 3. Analysis results feedback phase
[1020] Based on the analysis results of the generative AI model, the server generates feedback information about the pet's mood and health status, or about abnormal machine conditions. For example, the information may be, "The pet is relaxed," or "There is an abnormality in the machine." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen, informing the user of the pet's condition or the status of the machinery in the factory.
[1021] 4. Phase of responding when you are unwell
[1022] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. As with the previous procedure, the server inputs this data into a generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Avoid feeding certain foods," and "The nearest veterinary clinic is here." In the case of machinery in a factory, the system also suggests appropriate measures to take when an abnormality occurs.
[1023] Specific examples
[1024] For example, a video of a cat is taken using a smartphone and sent to a server. The server receives the video data and passes it to a generative AI model. The resulting analysis concludes that "the cat is relaxed." Furthermore, if a camera on a machine in a factory detects abnormal movement, the server immediately interprets this as "an abnormality has occurred in the machine" and notifies the worker's terminal. In this case, an example of a prompt sentence to be input to the generative AI model is as follows:
[1025] "Analyze the current video footage obtained from the robot's camera and detect if there are any abnormal conditions."
[1026] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1027] Step 1:
[1028] The device takes photos and videos of pets and factory machinery. Specifically, it uses the device's built-in camera to capture continuous or still image data. This data is compressed and pre-processed for efficiency. The input is the video of the pet or machinery, and the output is compressed video data.
[1029] Step 2:
[1030] The terminal transmits the captured video data to the server. Specifically, the video data is uploaded to the server via a network. The input is the preprocessed video data, and the output is the data transmitted to the server.
[1031] Step 3:
[1032] The server receives the video data sent from the terminal. Here, the data integrity is checked. If incomplete data is detected, a retransmission is requested. The input is the video data sent from the terminal, and the output is the complete video data.
[1033] Step 4:
[1034] The server inputs the received video data into a generative AI model for analysis. Specifically, it uses deep learning frameworks such as Keras to perform data preprocessing, feature extraction, and analysis. The input is the complete video data, and the output is the analysis results.
[1035] Step 5:
[1036] The server generates feedback information about the pet's mood and health, or about abnormal conditions in the machine, based on the analysis results of the generative AI model. Specifically, it converts the analysis results into text format and prepares the feedback in a format that is easy for the user to understand. The input is the analysis results, and the output is the feedback information.
[1037] Step 6:
[1038] The server sends the generated feedback information to the user's terminal. Specifically, the server uploads the feedback information to the terminal through the network. The input is the feedback information, and the output is the information sent to the terminal.
[1039] Step 7:
[1040] The device displays the received feedback information on the application screen and notifies the user of the status of the pet or machine. Specific operations include adjusting the layout and arranging the content for display on the screen. The input is the feedback information sent from the server, and the output is the feedback content displayed to the user.
[1041] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1042] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[1043] System configuration and operation
[1044] 1. Photo / video shooting and transmission phase
[1045] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[1046] 2. Data analysis phase
[1047] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[1048] 3. User Emotion Recognition Phase
[1049] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the pet's situation. This data is obtained from the user's device.
[1050] 4. Analysis results feedback phase
[1051] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device. The device then displays the received feedback on the application screen and notifies the user of the pet's status.
[1052] 5. Phase of responding when you are unwell
[1053] Furthermore, if a pet shows signs of illness, the user uses their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____." The server then uses an emotion engine to customize this information according to the user's emotional state and generate recommended action information. Finally, the server sends the recommended action information to the user's device, which displays it on the app screen.
[1054] Specific examples
[1055] Typical feedback examples
[1056] 1. The user takes a video of their cat with their smartphone.
[1057] 2. The device sends the captured video to the server via the app.
[1058] 3. The server receives the video data and passes it to the generative AI model.
[1059] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[1060] 5. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is relaxed.
[1061] 6. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[1062] 7. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[1063] Examples of what to do when you are feeling unwell
[1064] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[1065] 2. The device sends this data to the server.
[1066] 3. The server receives the data and inputs it into the generative AI model.
[1067] 4. The server's AI model analyzes that the pet may have indigestion.
[1068] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[1069] 6. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is worried.
[1070] 7. The server customizes countermeasure information based on the user's emotional state and sends it to the user's device.
[1071] 8. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[1072] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[1073] The processing flow will be explained below.
[1074] Step 1:
[1075] Users use their smartphone camera to take photos or videos of their pets.
[1076] Step 2:
[1077] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[1078] Step 3:
[1079] The terminal uploads the converted data to the server.
[1080] Step 4:
[1081] The server receives the photo and video data sent from the terminal.
[1082] Step 5:
[1083] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[1084] Step 6:
[1085] The server feeds the complete data into a generative AI model.
[1086] Step 7:
[1087] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[1088] Step 8:
[1089] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[1090] Step 9:
[1091] The server combines the analysis results of the generative AI model and the emotion engine to generate feedback information.
[1092] Step 10:
[1093] The server transmits the generated feedback information to the user's terminal.
[1094] Step 11:
[1095] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[1096] Step 12:
[1097] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[1098] Step 13:
[1099] The device sends the input symptom information and captured video data to the server.
[1100] Step 14:
[1101] The server receives the input symptom information and video data.
[1102] Step 15:
[1103] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[1104] Step 16:
[1105] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[1106] Step 17:
[1107] The server generates customized countermeasure information according to the user's emotional state based on the countermeasure information and the results of the emotion engine.
[1108] Step 18:
[1109] The server sends the generated countermeasure information to the user's terminal.
[1110] Step 19:
[1111] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[1112] Example 2
[1113] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1114] It is difficult to quickly and accurately grasp the health condition and mood of a pet, and there is a lack of technology to provide appropriate feedback according to the user's emotional state. Furthermore, when a pet becomes ill, there is a need for a means to quickly provide appropriate measures and alleviate the user's anxiety.
[1115] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1116] In this invention, the server includes a device for a user to take photos and videos of their pet, a means including a generative AI model that receives the photo and video data sent from the device and checks and analyzes the data for completeness, a means for transmitting feedback information on the pet's mood and health status analyzed by the server to the device, and a means for the server to analyze the user's emotional state and provide customized feedback based on that state. This makes it possible to quickly and accurately grasp the pet's health status and mood. Furthermore, appropriate feedback can be provided according to the user's emotional state, reducing the user's anxiety.
[1117] A "user" is someone who uses the system to take photos and videos of their pet and receive the analysis results.
[1118] A "terminal" is a photographic device used by a user, such as a smartphone or tablet.
[1119] "Photo and video data" refers to visual data of pets taken using a device.
[1120] A "server" is a computer system that receives and analyzes photo and video data sent from a terminal.
[1121] The "generative AI model" is an artificial intelligence model that resides on a server and analyzes the received data to predict the pet's feelings and health condition.
[1122] "Data integrity" means that the data received is in the correct format and is not missing or corrupted.
[1123] "Analysis" is the process of using a generative AI model to analyze a pet's facial expressions and movements to infer its feelings and health.
[1124] "Feedback information" is information about the pet's feelings and health condition that is generated by the server based on the analysis results and provided to the user.
[1125] The "emotion engine" is an algorithm that analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[1126] "Customized feedback" is feedback information that is tailored to the user's emotional state.
[1127] "What to do when your pet is unwell" refers to instructions and advice on how to respond appropriately when your pet becomes ill.
[1128] "Information about the nearest veterinary clinic" is information about the location and contact details of the nearest veterinary clinic where you can get your pet examined.
[1129] MODE FOR CARRYING OUT THE INVENTION
[1130] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[1131] Photo / video shooting and transmission phase
[1132] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos or videos are taken, the device converts the data format (compresses it, etc.) before sending it to the server via the application, efficiently uploading the data to the server.
[1133] A concrete example is a process in which a user takes a video of their cat on their smartphone, the device compresses the video, and sends it to a server using an HTTP POST request.
[1134] Data Analysis Phase
[1135] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[1136] For example, the generative AI model in the server uses deep learning frameworks such as TensorFlow to analyze a video of a cat and determine that the cat is relaxed.
[1137] User emotion recognition phase
[1138] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the situation of their pet. This data is obtained from the user's device.
[1139] As a specific example, the server collects facial expression data and tone of voice data from the user's device, inputs it into the emotion engine for analysis, and evaluates the user's emotional state by analyzing whether the user is relaxed while watching the video.
[1140] Analysis results feedback phase
[1141] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device, which then displays the received feedback on the application screen and notifies the user of the pet's status.
[1142] For example, a message might appear saying, "Your cat is relaxing. Come enjoy some relaxation time with him."
[1143] Phases of response when feeling unwell
[1144] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[1145] The server uses an emotion engine to customize this information according to the user's emotional state and generate countermeasure information. Finally, the server sends the countermeasure information to the user's device, which then displays it on the app screen.
[1146] For example, you might see feedback like, "Your pet may have indigestion. Keep your pet hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[1147] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[1148] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1149] Step 1:
[1150] The user takes photos and videos of their pet using the device. The user launches the camera app on their smartphone and captures the pet's behavior. The photos and videos are then saved as digital data on the device.
[1151] Input: Pet photos and videos
[1152] Output: Data stored in the device
[1153] Step 2:
[1154] The device saves the captured data in the application. The device saves the captured photos and videos in a specific folder so that they can be sent to the server later.
[1155] Input: Photo and video data
[1156] Output: Data stored within the application
[1157] Step 3:
[1158] Before the device sends the captured data to the server, it performs format conversion, compressing and encoding the data to convert it into a format that can be transmitted efficiently.
[1159] Input: Data stored within the application
[1160] Output: Compressed and encoded data
[1161] Step 4:
[1162] The device sends the converted data to the server, using an HTTP POST request to upload the data to the server.
[1163] Input: Compressed and encoded data
[1164] Output: Data sent to the server
[1165] Step 5:
[1166] The server receives the data sent from the device, takes in the data, and stores it in storage.
[1167] Input: Data sent from the terminal
[1168] Output: Data stored in the server
[1169] Step 6:
[1170] The server checks the integrity of the data, ensuring that the data received is not corrupted and is in the correct format, and if there is a problem it issues a request to resend it.
[1171] Input: Data stored on the server
[1172] Output: Data integrity check result (normal or resend request)
[1173] Step 7:
[1174] The server inputs the data into the generative AI model, passing the data to be analyzed to the AI model and starting the analysis process.
[1175] Input: Normal data
[1176] Output: The data fed into the generative AI model
[1177] Step 8:
[1178] A generative AI model analyzes the data, analyzing your pet's facial expressions and movements to infer its mood and health.
[1179] Input: Data fed into a generative AI model
[1180] Output: Analysis results (predictions of pet's mood and health)
[1181] Step 9:
[1182] The server recognizes the user's emotional state based on the analysis results, using an emotion engine to analyze the user's facial expressions, tone of voice, etc.
[1183] Input: Analysis results, and user facial expression and tone of voice data
[1184] Output: User's emotional state analysis results
[1185] Step 10:
[1186] The server integrates the analysis results and the emotional state, and generates feedback information based on the pet's state and the user's emotional state.
[1187] Input: Analysis results, user emotional state analysis results
[1188] Output: Integrated feedback information
[1189] Step 11:
[1190] The server sends feedback information to the user's device, and information based on the analysis results is sent to the device and notified to the user.
[1191] Input: Integrated feedback information
[1192] Output: Feedback information sent to the terminal
[1193] Step 12:
[1194] The terminal displays the feedback information to the user. The sent feedback information is displayed on the application screen to notify the user.
[1195] Input: Feedback information sent to the device
[1196] Output: Feedback information displayed on the application screen
[1197] (Application example 2)
[1198] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1199] When pets are left alone at home, it is difficult for owners to monitor their pets' health and emotions in real time and ensure their safety. Furthermore, if a pet shows signs of anxiety or poor health while the owner is away for an extended period of time, there is a lack of means to take prompt and appropriate action. This increases the risk of overlooking abnormal behavior or health problems in pets, and also reduces the sense of security for pet owners.
[1200] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal that takes and analyzes photos and videos of the pet, means including a generative AI model that receives photo and video data transmitted from the terminal and analyzes the data, means for transmitting feedback on the pet's feelings and health status analyzed by the server to the terminal, means including an emotion engine that collects and analyzes user emotion data, and means for the server to generate and transmit customized feedback information based on the analysis results and the user's emotional state. This makes it possible to monitor the pet's health status and emotions in real time and to respond quickly if an abnormality occurs. Furthermore, providing feedback according to the user's emotional state can provide a sense of security to the owner.
[1201] A "terminal" is a device that takes photos and videos of pets and sends the data to a server.
[1202] The "generative AI model" is an artificial intelligence model that analyzes received pet photos and video data to predict the pet's feelings and health condition.
[1203] A "server" is a central computer system that receives and analyzes data sent from terminals.
[1204] The "emotion engine" is an analytical engine that analyzes the user's facial expressions and tone of voice to understand the user's emotional state.
[1205] "Feedback information" is information about the pet's status that is generated by the server based on the analysis results and the user's emotional state.
[1206] "Customized feedback information" is feedback information whose content is adjusted according to the user's emotional state.
[1207] A specific system for implementing this invention will be described. The entire system is composed of multiple means, including a "terminal," a "server," and an "emotion engine." The purpose of this system is to analyze photos and videos of pets and provide feedback to users in real time.
[1208] 1. Photo / video shooting and transmission phase
[1209] Users take photos and videos of their pets using a device such as a smartphone. The captured data is converted into an appropriate format (JPEG, MP4, etc.) on the device. The device then sends the data to a server. This transmission is performed using the HTTP or RTSP protocol.
[1210] 2. Data analysis phase
[1211] The server receives the photo and video data sent from the device. It checks the integrity of the received data and requests a resend if it is incomplete. Once the data is verified, it is input into a generative AI model. This generative AI model uses machine learning algorithms to analyze the pet's facial expressions and movements, and performs data calculations to infer the pet's mood and health. Specifically, libraries such as TensorFlow and PyTorch are used.
[1212] 3. User emotion recognition phase
[1213] The server also incorporates an emotion engine that analyzes facial images and voice data acquired from the user's device. This analysis uses facial expression recognition and voice analysis technologies (OpenCV, DeepFace, SpeechRecognition, etc.). Based on this data, the user's emotional state is understood.
[1214] 4. Analysis results feedback phase
[1215] The server generates customized feedback information based on the analysis results of the generative AI model and the emotion engine. For example, if the user is worried, detailed analysis results and reassuring advice are provided. The generated feedback information is sent from the server to the device, which then displays it on the application screen.
[1216] 5. Response phase when you feel unwell
[1217] If a pet shows signs of illness, the user uses the device to input the pet's symptoms, take additional video footage, and send it to the server. The server then inputs the received data back into the generative AI model to analyze the cause of the pet's illness and the appropriate course of action. The analysis results include a specific action plan (e.g., "keep the pet hydrated and rest for several hours," "avoid feeding certain foods," etc.). Information about the nearest veterinary clinic is also provided. Using an emotion engine, customized information is generated according to the user's emotional state and sent to the device.
[1218] Specific examples
[1219] The following is an example of a prompt that the user enters:
[1220] Example prompt sentence:
[1221] "It analyzes pet behavior to determine if the pet is relaxed or not, and if the user is feeling anxious, it provides reassuring advice about the pet's situation."
[1222] In this way, this system allows users to safely monitor their pet's condition in real time and take prompt and appropriate action when necessary.
[1223] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1224] Step 1:
[1225] A user takes photos and videos of their pet using a device such as a smartphone.
[1226] Specific behavior:
[1227] The user launches the camera app on their smartphone and takes photos or videos of their pet. The captured data is converted into a format (e.g., JPEG or MP4).
[1228] Input: Pet photos and videos
[1229] Output: Format-converted photo and video data
[1230] Step 2:
[1231] The device sends the captured data to the server.
[1232] Specific behavior:
[1233] The terminal compresses the data and sends it to the server using the HTTP or RTSP protocol.
[1234] Input: Format-converted photo and video data
[1235] Output: Data sent to the server
[1236] Step 3:
[1237] The server checks the integrity of the photo and video data received.
[1238] Specific behavior:
[1239] The server calculates a checksum for the data and compares it with the received data, and if any incomplete data is found, it requests a retransmission.
[1240] Input: Data sent to the server
[1241] Output: Data with integrity checked
[1242] Step 4:
[1243] The server analyzes the data using the generative AI model.
[1244] Specific behavior:
[1245] The server then inputs the verified data into a generative AI model to analyze the pet's facial expressions and movements, using TensorFlow or PyTorch, for example, to perform data calculations to infer the pet's emotions and health status.
[1246] Input: Integrity checked data
[1247] Output: Analysis of pet's mood and health condition
[1248] Step 5:
[1249] The server compares the analysis results with the user's emotion engine.
[1250] Specific behavior:
[1251] The server inputs facial images and voice data acquired from the user's device into the emotion engine and analyzes the user's emotional state. The analysis uses libraries such as OpenCV, DeepFace, and SpeechRecognition.
[1252] Input: User's facial image and voice data
[1253] Output: User's emotional state
[1254] Step 6:
[1255] The server generates customized feedback information based on the analysis results and the user's emotional state.
[1256] Specific behavior:
[1257] The server generates feedback information based on the analysis results of the generative AI model and the emotion engine, depending on the user's emotional state. For example, if the user is worried, the server provides reassuring advice such as, "Your pet seems a little anxious, but there is nothing serious going on."
[1258] Input: Analysis results of pet's feelings and health condition, user's emotional state
[1259] Output: Customized feedback information
[1260] Step 7:
[1261] The server transmits the generated feedback information to the terminal.
[1262] Specific behavior:
[1263] The server transmits the generated feedback information to the user's terminal using the HTTP protocol or the like.
[1264] Input:Customized feedback information
[1265] Output: Feedback information sent to the device
[1266] Step 8:
[1267] The feedback information received by the terminal is displayed on the application screen.
[1268] Specific behavior:
[1269] The device analyzes the feedback information received from the server and displays it on the application screen, allowing the user to check it and understand the status of their pet.
[1270] Input: Feedback information sent to the device
[1271] Output: Feedback information displayed on the screen
[1272] The above is the processing flow of the system of the present invention. Through the specific operations performed at each step, it is possible to monitor the health and emotions of pets in real time and take prompt action if necessary.
[1273] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1274] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1275] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1276] [Fourth embodiment]
[1277] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1278] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1279] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1280] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1281] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1282] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1283] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1284] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1285] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1286] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1287] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1288] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1289] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1290] This invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system mainly includes the following elements: a device that takes photos and videos of pets, a server that contains a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[1291] System configuration and operation
[1292] 1. Photo / video shooting and transmission phase
[1293] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[1294] 2. Data analysis phase
[1295] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[1296] 3. Analysis results feedback phase
[1297] Based on the analysis results of the generative AI model, the server generates feedback information such as "your pet is relaxed," "your pet is excited," or "your pet looks unwell." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen and notifies the user of the pet's status.
[1298] 4. Phase of responding when you are unwell
[1299] Furthermore, if a pet shows signs of illness, the user can use the device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model, which analyzes the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[1300] Specific examples
[1301] Typical feedback examples
[1302] 1. The user takes a video of their cat with their smartphone.
[1303] 2. The device sends the captured video to the server via the app.
[1304] 3. The server receives the video data and passes it to the generative AI model.
[1305] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[1306] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[1307] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[1308] Examples of what to do when you are feeling unwell
[1309] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[1310] 2. The device sends this data to the server.
[1311] 3. The server receives the data and inputs it into the generative AI model.
[1312] 4. The server's AI model analyzes that the pet may have indigestion.
[1313] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[1314] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[1315] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[1316] The processing flow will be explained below.
[1317] Step 1:
[1318] Users use their smartphone camera to take photos or videos of their pets.
[1319] Step 2:
[1320] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[1321] Step 3:
[1322] The terminal uploads the converted data to the server.
[1323] Step 4:
[1324] The server receives the photo and video data sent from the terminal.
[1325] Step 5:
[1326] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[1327] Step 6:
[1328] The server feeds the complete data into a generative AI model.
[1329] Step 7:
[1330] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[1331] Step 8:
[1332] The server generates feedback information based on the analysis results of the generative AI model.
[1333] Step 9:
[1334] The server transmits the generated feedback information to the user's terminal.
[1335] Step 10:
[1336] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[1337] Step 11:
[1338] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[1339] Step 12:
[1340] The device sends the input symptom information and captured video data to the server.
[1341] Step 13:
[1342] The server receives the input symptom information and video data.
[1343] Step 14:
[1344] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[1345] Step 15:
[1346] Based on the analysis results of the generated AI model, the server generates appropriate countermeasures and information on the nearest veterinary clinic.
[1347] Step 16:
[1348] The server sends the generated countermeasure information to the user's terminal.
[1349] Step 17:
[1350] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[1351] Example 1
[1352] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1353] There is a need to quickly and accurately grasp the mood and health status of pets and provide users with appropriate feedback and solutions. However, conventional systems lack analytical accuracy and data completeness, which often results in delays in providing information to users. Another problem is the lack of a means to quickly provide appropriate solutions when pets are unwell or information about nearby veterinary clinics.
[1354] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1355] In this invention, the server includes means for taking and transmitting photos and videos of the pet using a photographing device, means for checking the integrity of the photo and video data transmitted from the photographing device and requesting retransmission if there is a problem, means for receiving the data whose integrity has been confirmed and analyzing the pet's facial expressions and movements using a generative AI model, and means for generating feedback information based on the analysis results and transmitting it to the user's device. This makes it possible to quickly and accurately analyze the pet's mood and health condition and provide the user with appropriate feedback and advice.
[1356] "Photography device" refers to equipment that can take photos and videos of pets and transmit the data.
[1357] "Means" refers to a method, device, function, etc. used to achieve a specific purpose.
[1358] "Checking the integrity of photo and video data" refers to the process of verifying that the received data is not lost or corrupted and is accurately received from the sender.
[1359] A "retry request" refers to a request to the sender to resend data when the data is incomplete or corrupted.
[1360] A "generative AI model" is a type of artificial intelligence model that uses algorithms or networks to generate results for specific tasks or data analysis.
[1361] "Analyzing facial expressions and movements" refers to the process of analyzing the pet's facial expressions and body movements from the received image and video data, and inferring their meaning and state.
[1362] "Generating feedback information" refers to the process of creating specific information to provide to the user based on the analysis results, such as whether the pet is relaxed or not feeling well.
[1363] "User device" refers to a device that receives analysis results and feedback and notifies the user of them. Specifically, this applies to smartphones and tablets.
[1364] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. This system mainly includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, and a means for providing the analysis results to users.
[1365] System configuration and operation
[1366] Photo and video shooting
[1367] Users use devices such as smartphones and tablets to take photos and videos of their pets. This is done using the device's built-in camera function. For example, if a user wants to take a video of their cat, they launch the device's camera app and capture the desired scene.
[1368] Data submission and preprocessing
[1369] Once the photo or video data is captured, the device converts it into the appropriate format within the application and performs pre-processing such as compression. Once pre-processing is complete, the device sends the data to the server, typically using the HTTPS protocol to ensure data protection and security.
[1370] Data reception and integrity check
[1371] The server receives the photo and video data sent from the device. Upon receiving the data, it checks its integrity. For example, it calculates a hash value (such as MD5 or SHA-256) to ensure the data is not corrupted. If the data is incomplete or corrupted, the server issues a resend request and repeats this process until the correct data is sent.
[1372] Data analysis
[1373] Once complete data is obtained, the server analyzes the data using a generative AI model, often based on TensorFlow or PyTorch models. The generative AI model analyzes the pet's facial expressions, gestures, movements, and other features contained in the received data to infer the pet's mood and health. Specifically, it processes images using libraries such as OpenCV and utilizes deep learning-based object detection algorithms (such as YOLO and SSD).
[1374] Generate feedback
[1375] The server generates feedback information based on the analysis results of the generative AI model. For example, the analysis results may include "The cat is relaxed" or "The dog is excited." This information is sent to the user's device via a RESTful API.
[1376] Providing feedback information
[1377] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. For example, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the application.
[1378] What to do when you are unwell
[1379] In particular, if a pet's condition worsens, the user uses their device to enter the pet's symptoms into the application and send additional video data. The server receives this data and again inputs it into the generative AI model to analyze the cause of the illness and specific countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for several hours") and information about nearby veterinary clinics. The analysis results are then sent back to the user's device and displayed on the application screen.
[1380] Specific examples
[1381] Typical feedback examples
[1382] 1. The user takes a video of their cat with their smartphone.
[1383] 2. The device sends the captured video to the server via the app.
[1384] 3. The server receives the video data and passes it to the generative AI model.
[1385] 4. The server's generated AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[1386] 5. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[1387] 6. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[1388] Examples of what to do when you are feeling unwell
[1389] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[1390] 2. The device sends this data to the server.
[1391] 3. The server receives the data and inputs it into the generative AI model.
[1392] 4. The server's generative AI model analyzes that "your pet may have indigestion."
[1393] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[1394] 6. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[1395] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1396] Step 1:
[1397] Users take photos and videos of their pets using devices such as smartphones or tablets. Once the photos and videos are taken, the device converts the photo and video data into the appropriate format within the application and performs preprocessing such as compression. This preprocessing reduces the data size and makes transmission more efficient. The input is the pet's photo and video data, and the output is compressed, format-converted data. Specifically, the user takes a video of their cat, converts it to MP4 format, and then compresses it using H.264 encoding. The converted data is then ready to be transmitted.
[1398] Step 2:
[1399] Once preprocessing is complete, the device sends the data to the server. At this time, the HTTPS protocol is used to ensure data protection and security. The input is the format-converted compressed data, and the output is the data sent to the server. Specifically, the data is uploaded to the server by pressing the send button in the application.
[1400] Step 3:
[1401] The server receives photo and video data sent from the device. To check the integrity of the received data, it calculates a hash value (such as MD5 or SHA-256) to see if the data is corrupted. If the data is incomplete or corrupted, the server sends a resend request to the device. The input is the data sent from the device, and the output is data whose integrity has been confirmed or a resend request. Specifically, the server calculates the hash value of the received data, and if it does not match, it sends an HTTP request for resend to the device.
[1402] Step 4:
[1403] Once complete data is obtained, the server analyzes the data using a generative AI model. This model uses a TensorFlow model or a PyTorch model. The generative AI model uses the received data as input and analyzes the pet's facial expressions and movements. The output is the analysis result, which includes information such as "The pet is relaxed" or "The pet is excited." Specific operations include image processing for each frame and running face recognition using OpenCV and deep learning object detection algorithms (such as YOLO and SSD).
[1404] Step 5:
[1405] The server generates feedback information based on the analysis results of the generative AI model. Based on the analysis results, it generates feedback in text format, such as "The cat is relaxed" or "The dog is excited." The input is the analysis results of the generative AI model, and the output is the feedback information. This feedback information is again sent from the server to the user's device. Specifically, the feedback information is sent to the user's device via a RESTful API, and receipt is confirmed.
[1406] Step 6:
[1407] The device displays the feedback information received from the server within the application and notifies the user of the pet's status. The input is the feedback information sent from the server, and the output is the notification displayed on the application screen. Specifically, messages such as "The cat is currently relaxed" or "The dog is excited" are displayed in the device's notification area or on the main screen of the app.
[1408] Step 7:
[1409] If a pet's condition worsens, the user uses the device to enter the pet's symptoms into the application and shoot and send additional video data. The input is text data and video data about the pet's symptoms, and the output is data sent to the server. Specifically, the user enters the symptoms in the text box, shoots additional video, and presses the send button from the application.
[1410] Step 8:
[1411] The server receives the data sent by the user and analyzes it again using the generative AI model. The input is text and video data about the symptoms, and the output is the analysis results regarding the cause of the illness and countermeasures. The generated analysis results include specific countermeasures (e.g., "Keep the pet hydrated and rest for two hours") and information about nearby veterinary clinics.
[1412] Step 9:
[1413] After the analysis results are generated, the server sends them to the user's device. The input is the analysis result for the poor health condition, and the output is information on countermeasures provided to the user. Specifically, data such as "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____" is sent via a RESTful API.
[1414] Step 10:
[1415] The terminal displays the countermeasure information received from the server on the application screen and notifies the user of specific countermeasures. The input is the countermeasure information sent from the server, and the output is a display of specific countermeasures provided to the user. Specifically, the message displayed is, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[1416] In this way, the system of the present invention can quickly and accurately analyze the pet's mood and health condition, and provide the user with appropriate feedback and advice.
[1417] (Application example 1)
[1418] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1419] Conventional machine monitoring systems in factories typically check the machine's condition visually or through regular inspections. However, these methods make it difficult to respond quickly when an abnormality occurs, making it difficult to prevent machine breakdowns and accidents. In extreme cases, this could lead to a serious accident or the shutdown of the production line, significantly impacting factory operations. To solve this problem, a system is needed that can constantly monitor the machine's condition and quickly detect and notify when an abnormality occurs.
[1420] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1421] In this invention, the server includes a means including a generative AI model that analyzes photos and videos of pets, a means including a generative AI model that analyzes video data from cameras installed in the factory, and a means for transmitting the analyzed machine status to an operator terminal, thereby enabling constant monitoring of the machine status in the factory and immediate detection and notification of any abnormalities that occur.
[1422] A "terminal" is a device that takes photos and videos of pets and machines and sends the data to a server.
[1423] The "server" is a device that receives photo and video data sent from a device, analyzes the data using a generative AI model, and feeds back the results.
[1424] A "generative AI model" is an artificial intelligence model that analyzes received photo and video data and infers the state and feelings of the subject.
[1425] "Pets" refer to animals kept at home and are the subject of analysis in the present invention.
[1426] The term "machine" refers to equipment and devices installed in a factory, and is the subject of analysis in the present invention.
[1427] A "camera" is a device for taking still or video images, and is often built into a terminal.
[1428] "Analysis" is a process carried out to understand the state and feelings of the target object based on the received data.
[1429] "Feedback" refers to the exchange of information to notify the user of the analysis results.
[1430] The "worker terminal" is a terminal used in a factory, and is a device for receiving and displaying analysis results.
[1431] An "abnormality" refers to a state in which the machine is in a different state from its normal state, and early detection is required.
[1432] This invention is a system that analyzes the status of pets and machinery in a factory and immediately notifies you if an abnormality occurs. This system is composed of the following elements.
[1433] System configuration and operation
[1434] 1. Photo / video shooting and transmission phase
[1435] The device takes photos and videos of pets and factory machinery. In the case of pets, the target is animals kept at home. In the case of factory machinery, the target is equipment and devices. Once the photos are taken, the device converts the format of the data (compresses it, etc.) and efficiently sends it to the server. The device has a built-in camera, so the photos and transmission are done automatically. The hardware used includes smartphones and factory surveillance cameras.
[1436] 2. Data analysis phase
[1437] The server receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server. This model is built using deep learning frameworks such as Keras. The generative AI model analyzes the pet's facial expressions and behavior, as well as the operation of machinery in the factory, to infer the pet's feelings and health condition, and any abnormalities in the machinery.
[1438] 3. Analysis results feedback phase
[1439] Based on the analysis results of the generative AI model, the server generates feedback information about the pet's mood and health status, or about abnormal machine conditions. For example, the information may be, "The pet is relaxed," or "There is an abnormality in the machine." This feedback information is then sent back from the server to the user's device. The device displays the received feedback on the application screen, informing the user of the pet's condition or the status of the machinery in the factory.
[1440] 4. Phase of responding when you are unwell
[1441] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. As with the previous procedure, the server inputs this data into a generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Avoid feeding certain foods," and "The nearest veterinary clinic is here." In the case of machinery in a factory, the system also suggests appropriate measures to take when an abnormality occurs.
[1442] Specific examples
[1443] For example, a video of a cat is taken using a smartphone and sent to a server. The server receives the video data and passes it to a generative AI model. The resulting analysis concludes that "the cat is relaxed." Furthermore, if a camera on a machine in a factory detects abnormal movement, the server immediately interprets this as "an abnormality has occurred in the machine" and notifies the worker's terminal. In this case, an example of a prompt sentence to be input to the generative AI model is as follows:
[1444] "Analyze the current video footage obtained from the robot's camera and detect if there are any abnormal conditions."
[1445] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1446] Step 1:
[1447] The device takes photos and videos of pets and factory machinery. Specifically, it uses the device's built-in camera to capture continuous or still image data. This data is compressed and pre-processed for efficiency. The input is the video of the pet or machinery, and the output is compressed video data.
[1448] Step 2:
[1449] The terminal transmits the captured video data to the server. Specifically, the video data is uploaded to the server via a network. The input is the preprocessed video data, and the output is the data transmitted to the server.
[1450] Step 3:
[1451] The server receives the video data sent from the terminal. Here, the data integrity is checked. If incomplete data is detected, a retransmission is requested. The input is the video data sent from the terminal, and the output is the complete video data.
[1452] Step 4:
[1453] The server inputs the received video data into a generative AI model for analysis. Specifically, it uses deep learning frameworks such as Keras to perform data preprocessing, feature extraction, and analysis. The input is the complete video data, and the output is the analysis results.
[1454] Step 5:
[1455] The server generates feedback information about the pet's mood and health, or about abnormal conditions in the machine, based on the analysis results of the generative AI model. Specifically, it converts the analysis results into text format and prepares the feedback in a format that is easy for the user to understand. The input is the analysis results, and the output is the feedback information.
[1456] Step 6:
[1457] The server sends the generated feedback information to the user's terminal. Specifically, the server uploads the feedback information to the terminal through the network. The input is the feedback information, and the output is the information sent to the terminal.
[1458] Step 7:
[1459] The device displays the received feedback information on the application screen and notifies the user of the status of the pet or machine. Specific operations include adjusting the layout and arranging the content for display on the screen. The input is the feedback information sent from the server, and the output is the feedback content displayed to the user.
[1460] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1461] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[1462] System configuration and operation
[1463] 1. Photo / video shooting and transmission phase
[1464] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos are taken, the device sends the photos or videos to the server via an application. Before sending, the device converts the data format (compresses it, etc.) to efficiently upload the data to the server.
[1465] 2. Data analysis phase
[1466] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[1467] 3. User Emotion Recognition Phase
[1468] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the pet's situation. This data is obtained from the user's device.
[1469] 4. Analysis results feedback phase
[1470] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device. The device then displays the received feedback on the application screen and notifies the user of the pet's status.
[1471] 5. Phase of responding when you are unwell
[1472] Furthermore, if a pet shows signs of illness, the user uses their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____." The server then uses an emotion engine to customize this information according to the user's emotional state and generate recommended action information. Finally, the server sends the recommended action information to the user's device, which displays it on the app screen.
[1473] Specific examples
[1474] Typical feedback examples
[1475] 1. The user takes a video of their cat with their smartphone.
[1476] 2. The device sends the captured video to the server via the app.
[1477] 3. The server receives the video data and passes it to the generative AI model.
[1478] 4. The server's AI model analyzes the cat's movements in the video and infers that the cat is relaxed.
[1479] 5. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is relaxed.
[1480] 6. The server generates feedback information that "the cat is relaxed" and sends it to the user's device.
[1481] 7. The device receives the feedback and displays "Your cat is currently relaxed" in the app.
[1482] Examples of what to do when you are feeling unwell
[1483] 1. If a user feels that their pet is not feeling well, they enter the symptoms into the app and take a video of their pet and send it.
[1484] 2. The device sends this data to the server.
[1485] 3. The server receives the data and inputs it into the generative AI model.
[1486] 4. The server's AI model analyzes that the pet may have indigestion.
[1487] 5. The server generates a solution to the problem, such as "Keep the pet hydrated and do not feed for two hours," along with information about the nearest veterinary clinic.
[1488] 6. The server's emotion engine analyzes the user's facial expressions and tone of voice to determine whether the user is worried.
[1489] 7. The server customizes countermeasure information based on the user's emotional state and sends it to the user's device.
[1490] 8. The device displays the countermeasure information received from the server on the app screen and tells the user, "Your pet may have indigestion. Keep it hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[1491] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[1492] The processing flow will be explained below.
[1493] Step 1:
[1494] Users use their smartphone camera to take photos or videos of their pets.
[1495] Step 2:
[1496] The device compresses and converts the format of the photos and videos taken before sending them to the server via the app.
[1497] Step 3:
[1498] The terminal uploads the converted data to the server.
[1499] Step 4:
[1500] The server receives the photo and video data sent from the terminal.
[1501] Step 5:
[1502] The server checks the integrity of the received data and issues a retransmission request if there is any incomplete data.
[1503] Step 6:
[1504] The server feeds the complete data into a generative AI model.
[1505] Step 7:
[1506] The AI model generated by the server analyzes photo and video data and predicts your pet's feelings and health based on its facial expressions, movements, and gestures.
[1507] Step 8:
[1508] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[1509] Step 9:
[1510] The server combines the analysis results of the generative AI model and the emotion engine to generate feedback information.
[1511] Step 10:
[1512] The server transmits the generated feedback information to the user's terminal.
[1513] Step 11:
[1514] The terminal displays the feedback information received from the server on the application screen to inform the user of the pet's status.
[1515] Step 12:
[1516] If a user feels that their pet is unwell, they input the symptoms and take a video of their pet.
[1517] Step 13:
[1518] The device sends the input symptom information and captured video data to the server.
[1519] Step 14:
[1520] The server receives the input symptom information and video data.
[1521] Step 15:
[1522] The server inputs the received data into a generative AI model to analyze the cause of the illness and appropriate countermeasures.
[1523] Step 16:
[1524] The server's emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[1525] Step 17:
[1526] The server generates customized countermeasure information according to the user's emotional state based on the countermeasure information and the results of the emotion engine.
[1527] Step 18:
[1528] The server sends the generated countermeasure information to the user's terminal.
[1529] Step 19:
[1530] The device displays the countermeasure information received from the server on the application screen and guides the user on how to respond promptly.
[1531] Example 2
[1532] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1533] It is difficult to quickly and accurately grasp the health condition and mood of a pet, and there is a lack of technology to provide appropriate feedback according to the user's emotional state. Furthermore, when a pet becomes ill, there is a need for a means to quickly provide appropriate measures and alleviate the user's anxiety.
[1534] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1535] In this invention, the server includes a device for a user to take photos and videos of their pet, a means including a generative AI model that receives the photo and video data sent from the device and checks and analyzes the data for completeness, a means for transmitting feedback information on the pet's mood and health status analyzed by the server to the device, and a means for the server to analyze the user's emotional state and provide customized feedback based on that state. This makes it possible to quickly and accurately grasp the pet's health status and mood. Furthermore, appropriate feedback can be provided according to the user's emotional state, reducing the user's anxiety.
[1536] A "user" is someone who uses the system to take photos and videos of their pet and receive the analysis results.
[1537] A "terminal" is a photographic device used by a user, such as a smartphone or tablet.
[1538] "Photo and video data" refers to visual data of pets taken using a device.
[1539] A "server" is a computer system that receives and analyzes photo and video data sent from a terminal.
[1540] The "generative AI model" is an artificial intelligence model that resides on a server and analyzes the received data to predict the pet's feelings and health condition.
[1541] "Data integrity" means that the data received is in the correct format and is not missing or corrupted.
[1542] "Analysis" is the process of using a generative AI model to analyze a pet's facial expressions and movements to infer its feelings and health.
[1543] "Feedback information" is information about the pet's feelings and health condition that is generated by the server based on the analysis results and provided to the user.
[1544] The "emotion engine" is an algorithm that analyzes the user's facial expressions and tone of voice to determine the user's emotional state.
[1545] "Customized feedback" is feedback information that is tailored to the user's emotional state.
[1546] "What to do when your pet is unwell" refers to instructions and advice on how to respond appropriately when your pet becomes ill.
[1547] "Information about the nearest veterinary clinic" is information about the location and contact details of the nearest veterinary clinic where you can get your pet examined.
[1548] MODE FOR CARRYING OUT THE INVENTION
[1549] The present invention is a system that analyzes photos and videos of pets and provides feedback on their moods and health to users. The system includes a device that takes photos and videos of pets, a server that includes a generative AI model that analyzes the data, a means for providing the analysis results to the user, and an emotion engine that recognizes the user's emotions.
[1550] Photo / video shooting and transmission phase
[1551] First, the user takes photos or videos of their pet using a device such as a smartphone. Once the photos or videos are taken, the device converts the data format (compresses it, etc.) before sending it to the server via the application, efficiently uploading the data to the server.
[1552] A concrete example is a process in which a user takes a video of their cat on their smartphone, the device compresses the video, and sends it to a server using an HTTP POST request.
[1553] Data Analysis Phase
[1554] The server then receives the photo and video data sent from the device. The server checks the integrity of the data and requests a resend if there is a problem. The received data is input into a generative AI model on the server, which analyzes the pet's facial expressions, gestures, and movements. Based on this analysis, the pet's mood and health condition are inferred.
[1555] For example, the generative AI model in the server uses deep learning frameworks such as TensorFlow to analyze a video of a cat and determine that the cat is relaxed.
[1556] User emotion recognition phase
[1557] The server is equipped with an emotion engine that recognizes the user's emotions. The server collects data such as the user's facial expressions and tone of voice to analyze how the user feels about the situation of their pet. This data is obtained from the user's device.
[1558] As a specific example, the server collects facial expression data and tone of voice data from the user's device, inputs it into the emotion engine for analysis, and evaluates the user's emotional state by analyzing whether the user is relaxed while watching the video.
[1559] Analysis results feedback phase
[1560] The server generates feedback information based on the user's emotional state based on the analysis results of the generative AI model and the emotion engine. For example, if the user is anxious, more detailed analysis results and reassuring advice are provided. The server then sends this customized feedback information to the user's device, which then displays the received feedback on the application screen and notifies the user of the pet's status.
[1561] For example, a message might appear saying, "Your cat is relaxing. Come enjoy some relaxation time with him."
[1562] Phases of response when feeling unwell
[1563] Furthermore, if a pet shows signs of illness, the user can use their device to input the pet's symptoms, take a video, and send it to the server via the application. The server then inputs this data back into the generative AI model to analyze the cause of the illness and the appropriate course of action. The analysis results include information such as "Keep the pet hydrated and rest for several hours," "Do not feed certain foods," and "The nearest veterinary clinic is ____."
[1564] The server uses an emotion engine to customize this information according to the user's emotional state and generate countermeasure information. Finally, the server sends the countermeasure information to the user's device, which then displays it on the app screen.
[1565] For example, you might see feedback like, "Your pet may have indigestion. Keep your pet hydrated and do not feed it for two hours. The nearest veterinary clinic is ____."
[1566] In this way, the system of the present invention quickly and accurately analyzes the mood and health status of pets and provides customized feedback based on the user's emotional state, allowing users to more easily manage their pet's health and understand its emotions.
[1567] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1568] Step 1:
[1569] The user takes photos and videos of their pet using the device. The user launches the camera app on their smartphone and captures the pet's behavior. The photos and videos are then saved as digital data on the device.
[1570] Input: Pet photos and videos
[1571] Output: Data stored in the device
[1572] Step 2:
[1573] The device saves the captured data in the application. The device saves the captured photos and videos in a specific folder so that they can be sent to the server later.
[1574] Input: Photo and video data
[1575] Output: Data stored within the application
[1576] Step 3:
[1577] Before the device sends the captured data to the server, it performs format conversion, compressing and encoding the data to convert it into a format that can be transmitted efficiently.
[1578] Input: Data stored within the application
[1579] Output: Compressed and encoded data
[1580] Step 4:
[1581] The device sends the converted data to the server, using an HTTP POST request to upload the data to the server.
[1582] Input: Compressed and encoded data
[1583] Output: Data sent to the server
[1584] Step 5:
[1585] The server receives the data sent from the device, takes in the data, and stores it in storage.
[1586] Input: Data sent from the terminal
[1587] Output: Data stored in the server
[1588] Step 6:
[1589] The server checks the integrity of the data, ensuring that the data received is not corrupted and is in the correct format, and if there is a problem it issues a request to resend it.
[1590] Input: Data stored on the server
[1591] Output: Data integrity check result (normal or resend request)
[1592] Step 7:
[1593] The server inputs the data into the generative AI model, passing the data to be analyzed to the AI model and starting the analysis process.
[1594] Input: Normal data
[1595] Output: The data fed into the generative AI model
[1596] Step 8:
[1597] A generative AI model analyzes the data, analyzing your pet's facial expressions and movements to infer its mood and health.
[1598] Input: Data fed into a generative AI model
[1599] Output: Analysis results (predictions of pet's mood and health)
[1600] Step 9:
[1601] The server recognizes the user's emotional state based on the analysis results, using an emotion engine to analyze the user's facial expressions, tone of voice, etc.
[1602] Input: Analysis results, and user facial expression and tone of voice data
[1603] Output: User's emotional state analysis results
[1604] Step 10:
[1605] The server integrates the analysis results and the emotional state, and generates feedback information based on the pet's state and the user's emotional state.
[1606] Input: Analysis results, user emotional state analysis results
[1607] Output: Integrated feedback information
[1608] Step 11:
[1609] The server sends feedback information to the user's device, and information based on the analysis results is sent to the device and notified to the user.
[1610] Input: Integrated feedback information
[1611] Output: Feedback information sent to the terminal
[1612] Step 12:
[1613] The terminal displays the feedback information to the user. The sent feedback information is displayed on the application screen to notify the user.
[1614] Input: Feedback information sent to the device
[1615] Output: Feedback information displayed on the application screen
[1616] (Application example 2)
[1617] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1618] When pets are left alone at home, it is difficult for owners to monitor their pets' health and emotions in real time and ensure their safety. Furthermore, if a pet shows signs of anxiety or poor health while the owner is away for an extended period of time, there is a lack of means to take prompt and appropriate action. This increases the risk of overlooking abnormal behavior or health problems in pets, and also reduces the sense of security for pet owners.
[1619] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a terminal that takes and analyzes photos and videos of the pet, means including a generative AI model that receives photo and video data transmitted from the terminal and analyzes the data, means for transmitting feedback on the pet's feelings and health status analyzed by the server to the terminal, means including an emotion engine that collects and analyzes user emotion data, and means for the server to generate and transmit customized feedback information based on the analysis results and the user's emotional state. This makes it possible to monitor the pet's health status and emotions in real time and to respond quickly if an abnormality occurs. Furthermore, providing feedback according to the user's emotional state can provide a sense of security to the owner.
[1620] A "terminal" is a device that takes photos and videos of pets and sends the data to a server.
[1621] The "generative AI model" is an artificial intelligence model that analyzes received pet photos and video data to predict the pet's feelings and health condition.
[1622] A "server" is a central computer system that receives and analyzes data sent from terminals.
[1623] The "emotion engine" is an analytical engine that analyzes the user's facial expressions and tone of voice to understand the user's emotional state.
[1624] "Feedback information" is information about the pet's status that is generated by the server based on the analysis results and the user's emotional state.
[1625] "Customized feedback information" is feedback information whose content is adjusted according to the user's emotional state.
[1626] A specific system for implementing this invention will be described. The entire system is composed of multiple means, including a "terminal," a "server," and an "emotion engine." The purpose of this system is to analyze photos and videos of pets and provide feedback to users in real time.
[1627] 1. Photo / video shooting and transmission phase
[1628] Users take photos and videos of their pets using a device such as a smartphone. The captured data is converted into an appropriate format (JPEG, MP4, etc.) on the device. The device then sends the data to a server. This transmission is performed using the HTTP or RTSP protocol.
[1629] 2. Data analysis phase
[1630] The server receives the photo and video data sent from the device. It checks the integrity of the received data and requests a resend if it is incomplete. Once the data is verified, it is input into a generative AI model. This generative AI model uses machine learning algorithms to analyze the pet's facial expressions and movements, and performs data calculations to infer the pet's mood and health. Specifically, libraries such as TensorFlow and PyTorch are used.
[1631] 3. User emotion recognition phase
[1632] The server also incorporates an emotion engine that analyzes facial images and voice data acquired from the user's device. This analysis uses facial expression recognition and voice analysis technologies (OpenCV, DeepFace, SpeechRecognition, etc.). Based on this data, the user's emotional state is understood.
[1633] 4. Analysis results feedback phase
[1634] The server generates customized feedback information based on the analysis results of the generative AI model and the emotion engine. For example, if the user is worried, detailed analysis results and reassuring advice are provided. The generated feedback information is sent from the server to the device, which then displays it on the application screen.
[1635] 5. Response phase when you feel unwell
[1636] If a pet shows signs of illness, the user uses the device to input the pet's symptoms, take additional video footage, and send it to the server. The server then inputs the received data back into the generative AI model to analyze the cause of the pet's illness and the appropriate course of action. The analysis results include a specific action plan (e.g., "keep the pet hydrated and rest for several hours," "avoid feeding certain foods," etc.). Information about the nearest veterinary clinic is also provided. Using an emotion engine, customized information is generated according to the user's emotional state and sent to the device.
[1637] Specific examples
[1638] The following is an example of a prompt that the user enters:
[1639] Example prompt sentence:
[1640] "It analyzes pet behavior to determine if the pet is relaxed or not, and if the user is feeling anxious, it provides reassuring advice about the pet's situation."
[1641] In this way, this system allows users to safely monitor their pet's condition in real time and take prompt and appropriate action when necessary.
[1642] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1643] Step 1:
[1644] A user takes photos and videos of their pet using a device such as a smartphone.
[1645] Specific behavior:
[1646] The user launches the camera app on their smartphone and takes photos or videos of their pet. The captured data is converted into a format (e.g., JPEG or MP4).
[1647] Input: Pet photos and videos
[1648] Output: Format-converted photo and video data
[1649] Step 2:
[1650] The device sends the captured data to the server.
[1651] Specific behavior:
[1652] The terminal compresses the data and sends it to the server using the HTTP or RTSP protocol.
[1653] Input: Format-converted photo and video data
[1654] Output: Data sent to the server
[1655] Step 3:
[1656] The server checks the integrity of the photo and video data received.
[1657] Specific behavior:
[1658] The server calculates a checksum for the data and compares it with the received data, and if any incomplete data is found, it requests a retransmission.
[1659] Input: Data sent to the server
[1660] Output: Data with integrity checked
[1661] Step 4:
[1662] The server analyzes the data using the generative AI model.
[1663] Specific behavior:
[1664] The server then inputs the verified data into a generative AI model to analyze the pet's facial expressions and movements, using TensorFlow or PyTorch, for example, to perform data calculations to infer the pet's emotions and health status.
[1665] Input: Integrity checked data
[1666] Output: Analysis of pet's mood and health condition
[1667] Step 5:
[1668] The server compares the analysis results with the user's emotion engine.
[1669] Specific behavior:
[1670] The server inputs facial images and voice data acquired from the user's device into the emotion engine and analyzes the user's emotional state. The analysis uses libraries such as OpenCV, DeepFace, and SpeechRecognition.
[1671] Input: User's facial image and voice data
[1672] Output: User's emotional state
[1673] Step 6:
[1674] The server generates customized feedback information based on the analysis results and the user's emotional state.
[1675] Specific behavior:
[1676] The server generates feedback information based on the analysis results of the generative AI model and the emotion engine, depending on the user's emotional state. For example, if the user is worried, the server provides reassuring advice such as, "Your pet seems a little anxious, but there is nothing serious going on."
[1677] Input: Analysis results of pet's feelings and health condition, user's emotional state
[1678] Output: Customized feedback information
[1679] Step 7:
[1680] The server transmits the generated feedback information to the terminal.
[1681] Specific behavior:
[1682] The server transmits the generated feedback information to the user's terminal using the HTTP protocol or the like.
[1683] Input:Customized feedback information
[1684] Output: Feedback information sent to the device
[1685] Step 8:
[1686] The feedback information received by the terminal is displayed on the application screen.
[1687] Specific behavior:
[1688] The device analyzes the feedback information received from the server and displays it on the application screen, allowing the user to check it and understand the status of their pet.
[1689] Input: Feedback information sent to the device
[1690] Output: Feedback information displayed on the screen
[1691] The above is the processing flow of the system of the present invention. Through the specific operations performed at each step, it is possible to monitor the health and emotions of pets in real time and take prompt action if necessary.
[1692] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1693] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1694] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1695] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1696] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1697] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1698] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1699] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1700] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1701] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1702] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1703] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1704] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1705] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1706] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1707] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1708] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1709] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1710] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1711] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1712] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1713] The following is further disclosed regarding the above embodiment.
[1714] (Claim 1)
[1715] A device that takes and analyzes photos and videos of pets,
[1716] a server including a generative AI model that receives photo and video data sent from the device and analyzes the data;
[1717] a means for transmitting to the terminal a feedback of the pet's feelings and health condition analyzed by the server;
[1718] A system including:
[1719] (Claim 2)
[1720] 2. The system according to claim 1, wherein the server provides appropriate measures to be taken when a pet becomes ill, and generates and transmits information about the nearest veterinary clinic.
[1721] (Claim 3)
[1722] 2. The system according to claim 1, wherein the terminal has means for inputting symptoms of the pet, taking a video, and transmitting the video to the server.
[1723] "Example 1"
[1724] (Claim 1)
[1725] A means for taking and transmitting photos and videos of the pet using a photographing device;
[1726] means for checking the integrity of the photographic and video data transmitted from the photographing device and requesting retransmission if there is a problem;
[1727] a means for receiving the integrity-verified data and analyzing the facial expressions and movements of the pet using a generative AI model;
[1728] means for generating feedback information based on the analysis results and transmitting the feedback information to a user device;
[1729] A system including:
[1730] (Claim 2)
[1731] A means for a user to input symptoms and send a video when the pet is unwell,
[1732] A means for the server to analyze the cause of the poor health and provide appropriate measures and information on the nearest veterinary hospital;
[1733] 10. The system of claim 1, comprising:
[1734] (Claim 3)
[1735] The system according to claim 1, further comprising means for presenting the analysis results to a user device in various data formats (text, notification, etc.).
[1736] "Application Example 1"
[1737] (Claim 1)
[1738] A device that takes and analyzes photos and videos of pets,
[1739] a server including a generative AI model that receives photo and video data sent from the device and analyzes the data;
[1740] a means for transmitting to the terminal a feedback of the pet's feelings and health condition analyzed by the server;
[1741] A terminal that captures images of equipment and machinery using cameras installed in the factory,
[1742] a server including a generative AI model that receives video data of the machine and analyzes its condition;
[1743] means for transmitting the machine status analyzed by the server to an operator terminal;
[1744] A system including:
[1745] (Claim 2)
[1746] 2. The system according to claim 1, wherein the server provides appropriate measures to be taken when a pet becomes ill, and generates and transmits information about the nearest veterinary clinic.
[1747] (Claim 3)
[1748] 2. The system according to claim 1, wherein the terminal has means for inputting symptoms of the pet, taking a video, and transmitting the video to the server.
[1749] "Example 2: Combining Emotion Engines"
[1750] (Claim 1)
[1751] a device for users to take photos and videos of their pets;
[1752] a server including a generative AI model that receives the photo and video data sent from the device, checks the integrity of the data, and analyzes it;
[1753] means for transmitting feedback information on the feelings and health status of the pet analyzed by the server to the terminal;
[1754] means for the server to also analyze the user's emotional state and provide customized feedback based on that state;
[1755] A system including:
[1756] (Claim 2)
[1757] 2. The system according to claim 1, wherein the server provides appropriate measures when a pet becomes unwell, and generates and transmits information about the nearest veterinary clinic.
[1758] (Claim 3)
[1759] 2. The system according to claim 1, wherein the terminal has a means for inputting symptoms of the pet, taking a video, and transmitting the video to the server.
[1760] "Application example 2 when combining emotion engines"
[1761] (Claim 1)
[1762] A device that takes and analyzes photos and videos of pets,
[1763] a server including a generative AI model that receives photo and video data sent from the device and analyzes the data;
[1764] a means for transmitting to the terminal a feedback of the pet's feelings and health condition analyzed by the server;
[1765] A means including an emotion engine for collecting and analyzing user emotion data;
[1766] means for generating and transmitting customized feedback information based on the analysis results and the user's emotional state;
[1767] A system including:
[1768] (Claim 2)
[1769] 2. The system according to claim 1, wherein the server provides appropriate measures to be taken when a pet becomes ill, and generates and transmits information about the nearest veterinary clinic.
[1770] (Claim 3)
[1771] 2. The system according to claim 1, wherein the terminal has means for inputting symptoms of the pet, taking a video, and transmitting the video to the server. [Explanation of symbols]
[1772] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A device that takes and analyzes photos and videos of pets, a server including a generative AI model that receives photo and video data sent from the device and analyzes the data; a means for transmitting to the terminal a feedback of the pet's feelings and health condition analyzed by the server; A system including:
2. 2. The system according to claim 1, wherein the server provides appropriate measures to be taken when a pet becomes ill, and generates and transmits information about the nearest veterinary clinic.
3. 2. The system according to claim 1, wherein the terminal has means for inputting symptoms of the pet, taking a video, and transmitting the video to the server.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A