Electronic device for performing data preprocessing for pet behavior learning and operating method thereof

The electronic device preprocesses various data types to create an adaptive AI model for companion animals, effectively predicting their behavior and emotions using a multi-modal language model, addressing the limitations of existing systems.

WO2025226114A1PCT designated stage Publication Date: 2025-10-30JEONG SOYOUNG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/095273
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-15
Filing Date
2025-04-22
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing systems fail to provide an adaptive AI model tailored to the temperament of each companion animal, making it difficult for pet owners to understand their pets' emotions, health, and future actions, especially when they are not constantly with the pet.

Method used

An electronic device preprocesses a behavioral data set including audio, olfactory, video, IMU, and biometric data to generate category-specific inference data, using a multi-modal artificial intelligence language model to predict the pet's behavior and emotions.

Benefits of technology

Enables accurate prediction of a companion animal's behavior, emotions, and future actions by generating high-quality training data and utilizing a learned AI model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025095273_30102025_PF_FP_ABST
    Figure KR2025095273_30102025_PF_FP_ABST
Patent Text Reader

Abstract

According to the present disclosure, provided are an electronic device and an operating method thereof, the electronic device comprising: a preprocessing unit for acquiring a behavior data set, preprocessing the behavior data set on the basis of a degree of influence of one category of data on another category in the behavior data set to generate category-specific inference data, and generating a training data set on the basis of the category-specific inference data; and a prediction unit for predicting, as an output, interpretation data indicating text interpreting the pet using a newly inputted behavior data set as an input, by using a multimodal artificial intelligence language model trained on the basis of the training data set.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for performing data preprocessing for pet behavior learning and its operating method

[0001] Embodiments of the present disclosure relate to an electronic device and a method of operating the same, and more particularly, to an electronic device and a method of operating the same for preprocessing data to learn the behavior of a companion animal using an artificial intelligence model.

[0002] The number of people keeping pets has been growing exponentially due to recent factors such as rising income levels and the increasing number of single-person households resulting from an aging population. Companion animals are animals that people love, keep close, and cherish, and include animals such as dogs, cats, birds, and goldfish. In particular, some companion animals, such as dogs and cats, are expanding their role in our increasingly individualistic modern society, sharing life with their owners and sharing emotional bonds.

[0003] Pets can convey messages to their owners through specific behaviors, such as barking or moving. Because pets cannot speak human language, direct communication with humans can be challenging. Consequently, even when their pets are sick, it can be difficult to recognize their symptoms, and an increasing number of users are frustrated by the inability to gauge their pets' emotions or health.

[0004] Although various technologies have been proposed to understand the behavior of companion animals, the existing systems for communicating with companion animals have the problem that they only unilaterally provide the pet owner with information on the pet's physical condition, activity level, or food intake. In addition, unless the pet owner is constantly with the pet, it is difficult to easily understand the pet's appearance, possible future actions, or emotional state.

[0005] However, these conventional electronic devices and their operating methods have the problem of failing to provide an adaptive AI model tailored to the temperament of each companion animal. Embodiments of the present disclosure address these and other issues, and provide an electronic device and its operating method that provide an adaptive AI model tailored to the temperament of each companion animal for AI services. However, these tasks are exemplary and are not intended to limit the scope of the present disclosure.

[0006] According to one aspect of the present disclosure, there is provided a preprocessing unit that obtains a behavioral data set including first audio data representing audio of a companion animal, second audio data representing the companion animal's voice, olfactory data representing a smell (olfactory) generated around the companion animal, video data representing a video in which the companion animal was filmed, IMU data representing an IMU (Inertial Measurement Unit) sensing result corresponding to a behavioral pattern of the companion animal, biometric data representing biometric information of the companion animal, first text data representing the surrounding environment of the companion animal in text, second text data representing a profile of the companion animal in text, and third text data representing experimental conditions of the companion animal in text, and preprocesses the behavioral data set based on an influence of one category of data on another category in the behavioral data set, thereby generating category-specific inference data, and generating a training data set based on the category-specific inference data; An electronic device is provided, including a prediction unit that uses a multi-modal artificial intelligence language model trained based on the training data set to predict interpretation data representing a text interpreted about the companion animal as an output using a newly input behavioral data set as an input.

[0007] According to the present embodiment, the preprocessing unit may receive the first audio data, the second audio data, the all-factory data, and the IMU data from a first sensing device attached to the companion animal, receive the video data from a second sensing device located around the companion animal, and receive the biometric data, the first text data, the second text data, and the third text data from an external device, or may store the biometric data, the first text data, the second text data, and the third text data in advance.

[0008] According to the present embodiment, the preprocessing unit may include a context parser that generates inference data for each of the all-factory data, the biometric data, the first text data, the second text data, and the third text data, analyzes the influence between the first audio data, the second audio data, the video data, and the IMU data, generates inference data for each of the first audio data, the second audio data, the video data, and the IMU data based on the influence, and synthesizes all generated inference data, thereby generating the training data set.

[0009] According to the present embodiment, the context parser generates spectrogram data by analyzing a spectrogram for frequency in the first audio data and the second audio data,

[0010] By recognizing an object in the video of the video data, object recognition data can be generated, by synthesizing a frequency analysis result and an object recognition result, synthetic data can be generated, by performing a signal processing operation on the IMU data, signal processed data can be generated, by performing a labeling operation on each of the first audio data and the second audio data and the synthetic data, a plurality of labeled data can be generated, by synthesizing the signal processed data and the plurality of labeled data, inference data corresponding to the IMU data can be generated, and by synthesizing all of the generated inference data, the training data set can be generated.

[0011] According to the present embodiment, the interpretation data may include text describing the appearance of the companion animal, including the posture of the companion animal, the behavior of the companion animal, and the expression of the companion animal, and text describing an expected circumstance for the companion animal.

[0012] According to the present embodiment, the interpretation data may further include text describing the current emotions of the companion animal and text describing the next action to be taken by the companion animal.

[0013] According to the present embodiment, the multimodal artificial intelligence language model may be a generative language model. The text of the interpretation data may be output in a descriptive format, describing the texts interpreting the companion animal in one or more complete sentences, based on the instruction text entered in the prompt.

[0014] According to this embodiment, the multimodal artificial intelligence language model may be a generative language model. The text of the interpretation data may be output in a question-and-answer format, matching the texts interpreting the companion animal with answers to the question text entered in the prompt.

[0015] According to another aspect of the present disclosure, there is provided a method of operating an electronic device, the method comprising: obtaining a behavioral data set including first audio data representing audio of a companion animal, second audio data representing the companion animal's voice, olfactory data representing a smell (olfactory) generated around the companion animal, video data representing a video in which the companion animal was filmed, IMU data representing an Inertial Measurement Unit (IMU) sensing result corresponding to a behavioral pattern of the companion animal, biometric data representing biometric information of the companion animal, first text data representing the surrounding environment of the companion animal in text, second text data representing a profile of the companion animal in text, and third text data representing experimental conditions of the companion animal in text; generating category-specific inference data by preprocessing the behavioral data set based on an influence of one category of data on another category in the behavioral data set; and generating a training data set based on the category-specific inference data.

[0016] Other aspects, features and advantages other than those described above will become apparent from the following detailed description, claims and drawings for carrying out the invention.

[0017] Additionally, these general and specific aspects may be implemented using any system, method, computer program, or combination of any system, method, or computer program.

[0018] An exemplary embodiment of the present disclosure has the effect of generating high-quality training data by preprocessing data to generate training data necessary for training an artificial intelligence model.

[0019] According to exemplary embodiments of the present disclosure, by using a learned artificial intelligence model, it is possible to predict the behavior and expression of a companion animal.

[0020] According to exemplary embodiments of the present disclosure, by using a learned artificial intelligence model, it is possible to predict the current emotions of a companion animal and the next action that the companion animal will take.

[0021] FIG. 1 is a block diagram illustrating a system according to an exemplary embodiment of the present disclosure.

[0022] FIG. 2 is a block diagram schematically illustrating a preprocessing unit according to an exemplary embodiment of the present disclosure.

[0023] FIG. 3 is a diagram illustrating a context parser according to an exemplary embodiment of the present disclosure.

[0024] FIG. 4 is a diagram illustrating a behavioral data set according to an exemplary embodiment of the present disclosure.

[0025] FIG. 5 is a block diagram schematically illustrating a prediction unit according to an exemplary embodiment of the present disclosure.

[0026] FIG. 6 is a diagram schematically illustrating an artificial intelligence model according to an exemplary embodiment of the present disclosure.

[0027] FIG. 7 is a flowchart illustrating a method of operating an electronic device according to an exemplary embodiment of the present disclosure.

[0028] The present disclosure is capable of various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present disclosure, as well as methods for achieving them, will become clearer with reference to the embodiments described in detail below, along with the drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various forms.

[0029] In the examples below, the terms first, second, etc. are not used in a limiting sense, but are used for the purpose of distinguishing one component from another.

[0030] In the examples below, singular expressions include plural expressions unless the context clearly indicates otherwise.

[0031] In the following examples, terms such as “include” or “have” mean that a feature or component described in the specification is present, and do not preclude the possibility that one or more other features or components may be added.

[0032] In the following examples, when a part such as a layer, region, or component is said to be on or above another part, it includes not only the case where it is directly above the other part, but also the case where another region, component, or the like is interposed in between.

[0033] For convenience of explanation, the sizes of components in the drawings may be exaggerated or reduced. For example, the sizes and thicknesses of each component shown in the drawings are arbitrarily indicated for convenience of explanation, and thus the present disclosure is not necessarily limited to the figures shown.

[0034] In some embodiments, where implementations are otherwise feasible, specific sequences of operations may be performed in a different order than described. For example, two steps described in succession may be performed substantially simultaneously, or in a reverse order from the described order.

[0035] In this specification, “A and / or B” refers to the case where it is A, or B, or both A and B. And, “at least one of A and B refers to the case where it is A, or B, or both A and B.

[0036] In the following examples, when it is said that layers, regions, components, etc. are connected, it includes cases where the layers, regions, components, etc. are directly connected, and / or cases where other layers, regions, components, etc. are interposed between the layers, regions, and components and are indirectly connected. For example, when it is said in this specification that layers, regions, components, etc. are electrically connected, it refers to cases where the layers, regions, components, etc. are directly electrically connected, and / or cases where other layers, regions, components, etc. are interposed between them and are indirectly electrically connected.

[0037] The x-axis, y-axis, and z-axis are not limited to the three axes in the Cartesian coordinate system, but can be interpreted in a broader sense that includes them. For example, the x-axis, y-axis, and z-axis may be orthogonal to each other, but they can also refer to different directions that are not orthogonal to each other.

[0038] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure is complete and to fully inform those skilled in the art of the scope of the present disclosure, and the present disclosure is defined solely by the scope of the claims.

[0039] The terminology used in this disclosure is for the purpose of describing embodiments only and is not intended to limit the present disclosure. In this disclosure, the singular may also include the plural unless specifically stated otherwise. The terms "comprises" and / or "comprising" as used herein do not exclude the presence or addition of one or more other components in addition to the mentioned components. Like reference numerals refer to like components throughout the disclosure, and "and / or" may include each and any combination of one or more of the mentioned components. Although "first", "second", etc. are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it should be understood that a first component mentioned below may also be a second component within the technical spirit of the present disclosure.

[0040] The word "exemplary" is used herein to mean "serving as an example or illustration." Any embodiment described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments.

[0041] Embodiments of the present disclosure may be described in terms of a function or a block that performs a function. A block, which may be referred to as a "unit" or a "module" in the present disclosure, may be physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memories, passive electronic components, active electronic components, optical components, hardwired circuits, etc., and may optionally be driven by firmware and software. Furthermore, the term "unit" as used in the disclosure refers to software, hardware elements such as FPGAs or ASICs, and the "unit" may perform certain roles. However, the "unit" is not limited to software or hardware. The "unit" may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, a "part" may include elements such as software elements, object-oriented software elements, class elements, and task elements, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided within the elements and "parts" may be combined into a smaller number of elements and "parts" or further separated into additional elements and "parts."

[0042] Embodiments of the present disclosure can be implemented using at least one software program running on at least one hardware device and capable of performing network management functions to control elements.

[0043] Spatially relative terms such as "below," "beneath," "lower," "above," and "upper" may be used to readily describe the relationship between one component and other components as depicted in the drawings. Spatially relative terms may be understood to encompass different orientations of components during use or operation in addition to the orientations depicted in the drawings. For example, if a component depicted in the drawings were flipped over, a component described as "below" or "beneath" another component may end up "above" the other component. Thus, the exemplary term "below" may encompass both the above and below orientations. Components may also be oriented in other directions, and thus spatially relative terms may be interpreted accordingly.

[0044] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure may be used with the meaning commonly understood by those skilled in the art to which this disclosure pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0045] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals and redundant descriptions thereof will be omitted.

[0046] FIG. 1 is a block diagram illustrating a system (100) according to an exemplary embodiment of the present disclosure.

[0047] Referring to FIG. 1, the system (100) may include a first sensing device (110), a second sensing device (120), an external device (130), and an electronic device (140).

[0048] The first sensing device (110) can sense physical variables associated with a companion animal and / or a companion. Companion animals may include dogs, cats, etc. Companion animals may refer to users who own companion animals. In an exemplary embodiment, the first sensing device (110) may be attached to the body of the companion animal or worn on a specific part of the body of the companion animal. For example, the first sensing device (110) may be a device in the form of a leash or collar worn around the neck of the companion animal, a harness worn around the front paws and chest of the companion animal, or a wearable device. In an exemplary embodiment, the first sensing device (110) may generate and output audio data, all-factory data, and IMU (Inertial Measurement Unit) sensing data. The audio data may include sounds produced by the companion animal or the companion animal. In an exemplary embodiment, the audio data may include first audio data and second audio data. The first audio data may be data representing audio of the companion animal. Secondary audio data may represent the companion's voice. Olfactory data may represent odors (olfactory) occurring around the companion. IMU data may represent IMU sensing results corresponding to the companion's behavioral patterns.

[0049] The second sensing device (120) can sense physical variables related to the companion animal and the companion animal's surroundings. In an exemplary embodiment, the second sensing device (120) can perform sensing operations from a third-person perspective. For example, a user other than the companion animal may operate the second sensing device (120), or the second sensing device (120) may be installed in the companion animal's surrounding space (e.g., on a wall) and perform sensing operations toward the companion animal. In an exemplary embodiment, the second sensing device (120) can photograph the companion animal and generate and output video data representing a video of the companion animal being photographed. In an exemplary embodiment, the second sensing device (120) can be implemented as a device that captures images, such as a camera.

[0050] The external device (130) can store various information about the companion animal. In an exemplary embodiment, the external device (130) can store biometric data, text data, etc., and provide the biometric data and text data to the electronic device (140). The biometric data may be data representing the biometric information of the companion animal. The text data may be data representing information about the companion animal in text format. In an exemplary embodiment, the text data may include first to third text data. The first text data may be data representing the companion animal's surroundings in text format. The second text data may be data representing the companion animal's profile in text format. The third text data may be data representing the companion animal's experimental conditions in text format.

[0051] Data provided from each of the first sensing device (110), the second sensing device (120), and the external device (130) may be provided to the electronic device (140). In an exemplary embodiment, a set of data provided from each of the first sensing device (110), the second sensing device (120), and the external device (130) may be referred to as a behavioral data set.

[0052] The electronic device (140) can generate a training data set based on a behavioral data set and train an artificial intelligence model using the training data set. Using the trained artificial intelligence model, the electronic device (140) can predict and output interpretation data about a companion animal expressed in human language using a behavioral data set provided from an external source (e.g., the first sensing device (110)) as input.

[0053] In an exemplary embodiment, the electronic device (140) may include a preprocessing unit (141) and a prediction unit (142).

[0054] The preprocessing unit (141) may preprocess a behavioral data set to generate a training data set used to train an artificial intelligence model, or the electronic device (140) may preprocess a behavioral data set to predict and output data to be predicted through a trained artificial intelligence model.

[0055] In an exemplary embodiment, the preprocessing unit (141) may obtain a behavioral data set, preprocess the behavioral data set based on the influence of one category of data on another category in the behavioral data set, thereby generating category-specific inference data, and generating training data based on the category-specific inference data. Here, the categories may include voice or audio, image (video), smell (or olfactory), behavioral patterns or IMU, companion's voice, environmental information, profile, etc.

[0056] In an exemplary embodiment, the preprocessing unit (141) may receive first audio data, second audio data, all-factory data, and IMU data from the first sensing device (110) attached to the companion animal. In addition, the preprocessing unit (141) may receive video data from the second sensing device (120) located around the companion animal. Meanwhile, the preprocessing unit (141) may receive biometric data, first text data, second text data, and third text data from the external device (130). Alternatively, the preprocessing unit (141) may store the biometric data, the first text data, the second text data, and the third text data in advance and utilize the stored data as metadata.

[0057] The prediction unit (142) can use an artificial intelligence model trained based on a training data set to predict interpretation data representing text interpreted about a companion animal as output, using a newly input behavioral data set as input. In an exemplary embodiment, the artificial intelligence model may be a multi-modal artificial intelligence language model. The multi-modal artificial intelligence language model according to an exemplary embodiment may be a generative language model. The generative language model may be, for example, a model designed based on a transformer, but is not limited thereto. In an exemplary embodiment, when a user inputs an input text (TXT1) through a prompt, the prediction unit (142) can output an output text (TXT2) for the input text (TXT1) using the trained multi-modal artificial intelligence language model.

[0058] The electronic device (140) can communicate with various terminals and execute programs of program data. In this specification, the 'device according to the present disclosure' includes all various devices that can perform computational processing and provide results to a user. For example, the device according to the present disclosure may be a computing device, including all of a computer, a server device, and a portable terminal, or may be in the form of any one of them. Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser. The server device is a server that communicates with external devices to process information, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server. A portable terminal is, for example, a wireless communication device that ensures portability and mobility, and may include all kinds of handheld-based wireless communication devices such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handy-phone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminals, smart phones, etc., and wearable devices such as watches, rings, bracelets, anklets, necklaces, glasses, contact lenses, or head-mounted devices (HMD).

[0059] In an exemplary embodiment, the electronic device (140) may be implemented as a device including a memory and a processor. Functions related to artificial intelligence according to the present disclosure are operated through the processor and memory.

[0060] The memory can store data for an algorithm for controlling the operation of components within the device or a program that reproduces the algorithm, and can be implemented with at least one processor that performs the aforementioned operation using the data stored in the memory. Here, the memory and the processor can be implemented as separate chips. Alternatively, the memory and the processor can be implemented as a single chip. The memory can store data that supports various functions of the device, a program for the operation of the processor, and can store input / output data. It can store a plurality of application programs (or applications) running on the device, data for the operation of the device, commands, and one or more instructions. At least some of these application programs can be downloaded from an external server via wireless communication.

[0061] A processor can execute one or more instructions stored in memory. The processor may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, AP, or DSP (Digital Signal Processor), a graphics-only processor such as a GPU or VPU (Vision Processing Unit), or an AI-only processor such as an NPU. One or more processors control the processing of input data according to predefined operating rules or AI models stored in memory. Alternatively, if one or more processors are AI-only processors, the AI-only processor may be designed with a hardware structure specialized for processing a specific AI model.

[0062] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is learned by a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0063] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.

[0064] According to an exemplary embodiment of the present disclosure, a processor can implement artificial intelligence. Artificial intelligence refers to a machine learning method based on an artificial neural network that imitates human neurons (biological neurons) to enable machines to learn. Artificial intelligence methodologies can be categorized into supervised learning, in which input data and output data are provided together as training data depending on the learning method, so that the solution (output data) to the problem (input data) is determined; unsupervised learning, in which only input data is provided without output data, so that the solution (output data) to the problem (input data) is not determined; and reinforcement learning, in which a reward (Reward) is provided from an external environment whenever an action (Action) is taken in the current state (State), and learning is performed in a direction to maximize this reward. In addition, artificial intelligence methodologies can be categorized according to the architecture of the learning model. The architectures of widely used deep learning technologies can be categorized into convolutional neural networks (CNNs), recurrent neural networks (RNNs), transformers, and generative adversarial networks (GANs).

[0065] The present device and system may include an artificial intelligence model. The artificial intelligence model may be a single artificial intelligence model or may be implemented as multiple artificial intelligence models. The artificial intelligence model may be composed of a neural network (or artificial neural network) and may include statistical learning algorithms that mimic biological neurons in machine learning and cognitive science. A neural network may refer to a model in general that has problem-solving capabilities by changing the binding strength of synapses through learning, formed by artificial neurons (nodes) that form a network by combining synapses. The neurons of the neural network may include a combination of weights or biases. The neural network may include one or more layers composed of one or more neurons or nodes. For example, the device may include an input layer, a hidden layer, and an output layer. The neural network constituting the device can infer a desired result (output) from an arbitrary input (input) by changing the weights of neurons through learning.

[0066] The processor can create a neural network, train (or learn) a neural network, perform a calculation based on received input data, generate an information signal based on the calculation result, or retrain the neural network. The models of the neural network can include various types of models such as CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, R-CNN (Region with Convolution Neural Network), RPN (Region Proposal Network), RNN (Recurrent Neural Network), S-DNN (Stacking-based deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), Deconvolution Network, DBN (Deep Belief Network), RBM (Restrcted Boltzman Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory) Network, Classification Network, etc., but are not limited thereto. The processor can include one or more processors for performing calculations according to the models of the neural network. For example, the neural network can be a deep neural network. It may include a deep neural network.

[0067] Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Network), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Depp Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), Generative Adversarial Network (GAN), Liquid State Machine (LSM), Extreme Learning Machine (ELM), It will be understood by those skilled in the art that any neural network may be included, including but not limited to ESN (Echo State Network), DRN (Deep Residual Network), DNC (Differentiable Neural Computer), NTM (Neural Turning Machine), CN (Capsule Network), KN (Kohonen Network), and AN (Attention Network).

[0068] According to an exemplary embodiment of the present disclosure, the processor may be configured to perform a process for generating a CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, Region with Convolution Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restrcted Boltzman Machine (RBM), Fully Convolutional Network, Long Short-Term Memory (LSTM) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA for natural language processing, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics for vision processing, Visual Understanding, Video Synthesis, ResNet for data intelligence, Anomaly Detection, Prediction, Time-Series Forecasting, Various artificial intelligence structures and algorithms can be used, including but not limited to Optimization, Recommendation, and Data Creation.

[0069] Although not shown, the electronic device (100) may further include a communication unit. The communication unit may perform communication with at least one user terminal, etc. In this case, in addition to a Wi-Fi module, the communication unit may include a wireless communication module that supports various wireless communication methods such as LTE (Long Term Evolution), 4G, 5G, and 6G.

[0070] According to the above-described embodiments, there is an effect of being able to generate high-quality training data by preprocessing data to generate training data necessary for training an artificial intelligence model.

[0071] In addition, according to the above-described embodiments, by using the learned artificial intelligence model, there is an effect of being able to predict the behavior and expression of a companion animal.

[0072] In addition, according to the above-described embodiments, by using the learned artificial intelligence model, there is an effect of being able to predict the current emotions of the companion animal and the next action the companion animal will take.

[0073] FIG. 2 is a block diagram schematically illustrating a preprocessing unit (200) according to an exemplary embodiment of the present disclosure.

[0074] Referring to FIG. 2, the preprocessing unit (200) may be provided with a behavioral data set (220). The behavioral data set (220) may include first data that may be generated at home (21), second data that may be generated at a laboratory (22), and third data provided from a device (23). In an exemplary embodiment, the first data provided from home (21) may include profiling of the companion animal through a questionnaire from the companion animal owner (or the parent of the companion animal). For example, the first data may be the second text data described above. The second data provided from the laboratory (22) may include experimental results from indoor and / or outdoor experiments. In this case, the experiment may be conducted under conditions that minimize stress on the companion animal by designing and executing the experimental routine with the assistance of an animal behavior expert or a veterinarian. The experiment may be conducted under specific, periodically set conditions, and data bundled at 10-second intervals based on each experiment may be provided as the behavioral data set (220). For example, 1,000 conditional data bundles can be collected. For example, the second data may include the first text data, third text data, and biometric data described above. The third data provided from the device (23) may be data continuously collected throughout the entire experimental period, and may include, for example, first audio data, second audio data, all-factory data, IMU data, and video data. For example, the video data of the third data may include video recorded for more than 3,670 hours. At least 500 hours of total data within three months may be collected as the third data. The IMU data and video data may be used to align in an image bind.

[0075] In an exemplary embodiment, the preprocessing unit (200) may include a context parser (220). The context parser (220) may generate inference data for each of the all-factory data, the biometric data, the first text data, the second text data, and the third text data. In addition, the context parser (220) may analyze the influence between the first audio data, the second audio data, the video data, and the IMU data, and generate inference data for each of the first audio data, the second audio data, the video data, and the IMU data based on the influence. In addition, the context parser (220) may generate a training data set (230) by synthesizing the generated plurality of inference data.

[0076] In an exemplary embodiment, the training data set (230) may be stored in a database (not shown). Alternatively, the training data set (230) may be provided to the prediction unit (141).

[0077] FIG. 3 is a diagram illustrating a context parser (300) according to an exemplary embodiment of the present disclosure.

[0078] Referring to FIG. 3, the categories of the behavioral data set (210) may include, for example, voice, video, behavioral pattern, olfactory, biometric information, surrounding environment, profile, and condition, and the behavioral data set (210) may include voice data (311), video data (312), behavioral pattern data (313), olfactory data (314), biometric information data (315), surrounding environment data (316), profile data (317), and condition data (318), which represent voice, video, behavioral pattern, olfactory, biometric information, surrounding environment, profile, and condition, respectively. In an exemplary embodiment, the voice data (311) may correspond to the first and second audio data described above with reference to FIG. 1. The video data (312) may correspond to the video data described above with reference to FIG. 1. The behavioral pattern data (313) may correspond to the IMU data described above with reference to FIG. 1. The olfactory data (314) may correspond to the all-factory data described above with reference to FIG. 1. The biometric information data (315) may correspond to the biometric data described above with reference to FIG. 1. Each physical variable represented by the surrounding environment data (316), the profile data (317), and the condition data (318) may be expressed as text, and for example, the surrounding environment data (316), the profile data (317), and the condition data (318) may correspond to the first to third text data, respectively.

[0079] The context parser (300) can generate inference data for each category of data in the behavioral data set (210). In an exemplary embodiment, the context parser (300) can generate inference data for each of the all-factory data, the biometric data, the first text data, the second text data, and the third text data. Referring to FIG. 3, for example, the context parser (300) can generate inference data for each of the olfactory data (314), the biometric information data (315), the surrounding environment data (316), the profile data (317), and the condition data (318).

[0080] The context parser (300) may generate inference data by considering the influence of a specific category among several categories on other categories. In an exemplary embodiment, the influence between first audio data, second audio data, video data, and IMU data may be analyzed, and inference data for each of the first audio data, second audio data, video data, and IMU data may be generated based on the influence. Referring to FIG. 3, for example, the context parser (300) may generate spectrogram data by analyzing a spectrogram for frequencies in audio data (311) (321). The context parser (300) may generate object recognition data by recognizing an object in an image of image data (312) (322). The context parser (300) may generate synthesized data by synthesizing the frequency analysis result and the object recognition result (331). The context parser (300) may generate signal-processed data by performing a signal processing operation on IMU data (323). The context parser (300) can generate a plurality of labeling data by performing a labeling operation on each of the first audio data, the second audio data, and the synthesized data (324). Meanwhile, the context parser (300) can generate inference data corresponding to the IMU data by synthesizing the signal-processed data and the plurality of labeling data (332).

[0081] The context parser (300) can synthesize all generated inference data (340) (333) and generate a training data set (230) corresponding to the synthesized data.

[0082] FIG. 4 is a diagram illustrating a behavioral data set according to an exemplary embodiment of the present disclosure.

[0083] Referring to FIG. 4, the IMU data (410) may be six-axis (or six-dimensional) time series data composed of acceleration and angular velocity along tree axes. The category of the IMU data (410) is IMU, and the type of the IMU data (410) is time series. The IMU may include an acceleration sensor and an angular velocity sensor, and in some cases, may further include a geomagnetic sensor. When the IMU includes an acceleration sensor and a geomagnetic sensor, the IMU data (410) may be nine-axis (or nine-dimensional) time series data composed of acceleration, angular velocity, and magnetic field along tree axes. The equipment used to acquire the IMU data (410) in the laboratory and field may be an IMU.

[0084] The first audio data (420) may include HFA (High Frequency Audio) data (421) and LFA (Low Frequency Audio) data (422). Each of the HFA data (421) and the LFA data (422) is of the audio type. The HFA data (421) may be data having a frequency range other than the frequency range of a human voice. For example, the HFA data (421) may have an audible frequency range (e.g., about 20 to 40 [kHz]) of a companion animal (e.g., a puppy or a dog). The category of the HFA data (421) is HFA. The equipment used to obtain the HFA data (421) in the laboratory may be a controlled speaker and microphone (Controlled Speaker & MIC). The equipment used to obtain the HFA data (421) in the field may be a wide-range MIC. LFA data (422) may be data having a frequency range other than the frequency range of a human voice but lower than the frequency range of the HFA data (421), or a frequency range of a human or pet voice. For example, the LFA data (422) may have a human audible frequency range (e.g., below about 20 [kHz]) or a frequency of a human or canine voice. The category of the LFA data (422) is LFA. The equipment used to acquire the LFA data (422) in the laboratory may be a controlled speaker and camcorder. The equipment used to acquire the HFA data (421) in the field may be a microphone (MIC).

[0085] In the audio domain, for example, a plant may produce a specific pattern of sound when it is dehydrated, or a tree may produce a specific pattern of sound when it is overly wet. Specifically, a stressed plant may produce an ultrasonic sound that can be recorded. The form of the ultrasonic sound may vary depending on the type of plant, the type of stress (drying, cutting), and the intensity of the stress (the number of days of drying). The form of the sound may be captured in a graph that can be defined by frequency, duration, and power (dB / Hz). The average maximum sound intensity recorded for dry plants may be approximately 61.6±0.1 [dBSPL] and 65.6±0.4 [dBSPL], and the average maximum frequency at 10 [cm] for tomatoes and tobacco, respectively, may be 49.6±0.4 [kHz] and 54.8±1.1 [kHz]. The average peak intensity of the sounds emitted by cut plants can be 65.6±0.2 [dBSPL] and 63.3±0.2 [dBSPL], and the average peak frequency at 10.0 [cm] for tomatoes and tobacco, respectively, can be 57.3±0.7 [kHz] and 57.8±0.7 [kHz]. For example, when certain organisms, especially species specialized in the ultrasonic range, such as dogs, walk in a park or mountain with many plants, dogs can hear the sounds of the aforementioned plants, so when a human does not hear any sounds, the dog may suddenly look into the distance and prick up its ears. In general, sharp ultrasonic emitting machines can be used for anti-barking purposes, but the soft or other high-frequency sounds emitted by nature, such as tomatoes, can be felt as everyday sounds for dogs, or can even be a factor that stimulates the enjoyment and curiosity of a walk.Based on the dynamic relationship between natural high-frequency sounds and the sensory organs and / or reactions of dogs, the present disclosure uses an artificial intelligence model to predict the actions, emotions, and future actions of companion animals by inputting physical characteristics such as hearing and smelling of the companion animal. In HFA, a spectrogram for frequency can be visualized and used for image learning. For HFA, two microphones, an all-band MIC and a stereo LFA MIC, can be installed on the companion animal's collar, and operations such as data cross-validation, data correction, missing value supplementation, and data augmentation are possible based on the data acquired by the two microphones.

[0086] The second audio data (430) may be data representing a corpus containing the voices of the pet's parents. The category of the second audio data (430) is PPV, and the type of the second audio data (430) is audio. The equipment used to obtain the second audio data (430) in the laboratory may be a controlled speaker and microphone (Controlled Speaker & MIC). The equipment used to obtain the second audio data (430) in the field may be a microphone (MIC).

[0087] The Olfactory data (440) may be a reference database of 32 odors at three concentration levels to address the noise issue of field gas sensor graphs. The Olfactory data (440) category is Olfactory (Smell & Gas), which includes odors and gases, and the Olfactory data (440) type is time series. The equipment used to acquire the Olfactory data (440) in the laboratory may be a Smell Simulator. The equipment used to acquire the Olfactory data (440) in the field may be a Gas Sensor.

[0088] In the Allfactory domain, for example, 32 odors (house odors, fabric softener, bread, food, diffuser, etc.) can be classified based on multiple concentration levels. For example, pets are sensitive to body odor, so their body odors can be classified based on five or more concentration levels. In the laboratory, the body odor of the pet's owner may be primarily used, while in the field, natural odors, textile odors, etc. may be applied. Classification criteria in the Allfactory domain may include, for example, familiar and unfamiliar odors, odors based on racial classification, odors of people you've met, odors of spaces you've been to, unique household odors or incense odors, odors of commonly eaten foods, or odors based on the gender, race, cleanliness, food intake, and health of the primary caregiver.

[0089] The first text data (450) may be data that expresses various environmental information around the companion animal, such as temperature, humidity, and location, in text format. The category of the first text data (450) is environmental information (Env (Environmental Info.)), and the type of the first text data (450) is text. The equipment used to obtain the first text data (450) in the laboratory is a measuring equipment for direct measurement, and in the field, the first text data (450) may be stored in the device as metadata.

[0090] The second text data (460) may be data that expresses detailed information about the companion animal, such as the breed, age, sex, and neutering status of the companion animal, in text format, or may be data that indicates dog temperament types. The category of the second text data (460) is profile, and the type of the second text data (460) is text. The second text data (460) may be obtained through text labels submitted by owners of companion animals in laboratories and fields (Owner-Submitted Text Labels) and / or through personal information labels from our custom survey (Personality Labels from Our Custom Survey).

[0091] Third text data (470) may be data that represents experimental conditions, such as who said what and where the dog goes for a walk, in text form. The category of third text data (470) is "Cond" and the type of third text data (470) is "text." Third text data (470) can be acquired in a laboratory through metadata.

[0092] Video data (480) may be data representing images and / or videos captured by a third-person perspective camera that captures the appearance of a target companion animal (e.g., a target puppy). The category of the video data (480) is video, and the type of the video data (480) is video (images). The video data (480) may be acquired in a laboratory using a camcorder.

[0093] In an exemplary embodiment, the behavioral data set may include lab data obtained in a laboratory and field data obtained in the field. Lab data may include basic collected data such as video, LFA, HFA, IMU, olfactory, pet parent's voice (PPV), dog's profile, experimental conditions, GPS, etc., and health data such as ECG, EEG, PPG, and IMUs mounted at different locations. Field data may include LFA, HFA, IMU, olfactory, GPS, etc. The collected data may be stored on a server directly or over a network. Since various modalities are collected simultaneously, the data is organized and stored in sync with respect to time.

[0094] The collected data requires processing to be used in AI model training. Voice data (e.g., first audio data (420)) must be divided into LFA and HFA, and sound annotation must be performed for the LFA. Video data (480) is used to annotate the dog's appearance. Since the AI ​​model of the present disclosure does not use video data (480) as input, the annotation data obtained in this manner is paired with the IMU and used as training data during the model training process. All-factory data requires labeling according to odor type and concentration. If one training data clip is 10 seconds long, a spectrogram of the voice data is created and stored in 10-second units. Analysis of the IMU data provides direct information about the dog's movements, so the results of feature extraction from the IMU data (410) are stored as training data. After processing each data, the processed data are grouped and stored in a database so that they can be used as training data.

[0095] In model training, data augmentation techniques can be used to dramatically increase the amount of training data. Specifically, slightly perturbing the MU or voice information can increase the diversity of inputs while maintaining a nearly identical output. The correct answer (e.g., predicted value or ground truth, abbreviated as GT) for the model's output can be derived from analysis of audio data, such as the dog's appearance, emotions, and behavior. Based on this, the model can be trained so that its output resembles the video.

[0096] To deploy the model, the trained model is uploaded to a server, and when new animal behavior data arrives, predicted results are obtained based on the model trained on the server. The model of this disclosure is trained by observing the predicted appearance based on a video as ground truth. However, there are areas where the model of this disclosure has difficulty learning on its own, such as surrounding information or health analysis of the dog that are not revealed in the image. In these cases, there may be differences from the actual correct answer, and these differences can be used to periodically update the already trained model. This update method is called SDILS. SDILS refers to a system that periodically updates the already trained model to narrow the gap between the model output and the actual expected result. The ground truth for the output of this disclosure uses the results extracted from the video data (480). However, even if the ground truth is obtained from the video using a high-performance Vision-Language LLM, errors may be present in the analysis of the dog's emotions or behavior when viewed by an animal behavior expert. And it is even more difficult to trust the inference of a high-performance LLM as the correct answer when it comes to information related to the surrounding environment that is not revealed in the video, or areas that require veterinary knowledge such as ECG, EEG, PPG, or gait / respiratory analysis. Therefore, by analyzing the true correct answer for such data, the model of this disclosure can enable analysis that requires more advanced knowledge beyond the level of simple visual analysis. The results analyzed by SDILS are implemented through an additional RoLA fine-tuning technique. The reason for using the RoLA technique is that it learns new parameters without touching the parameters of the previously learned model, so the general inference ability is maintained while further improving the overall performance. In addition, since RoLA merges the parameters into the model parameters during inference, it can have the advantage of maintaining the size of the inference model and the inference time.

[0097] FIG. 5 is a block diagram schematically illustrating a prediction unit (500) according to an exemplary embodiment of the present disclosure.

[0098] Referring to FIG. 5, a measuring device (51) can be worn on a companion animal such as a dog or puppy to sense various physical variables of the companion animal and output sensing data. In an exemplary embodiment, the measuring device (51) can include an IMU unit (51_1), an audio unit (51_2), a video unit (51_3), a GPS unit (51_4), an all-factory unit (51_5), and a biometric information unit (51_6). The IMU unit (51_1) can output IMU data, the audio unit (51_2) can output audio data, the video unit (51_3) can generate video data, the GPS unit (51_4) can generate GPS data, the all-factory unit (51_5) can generate all-factory data, and the biometric information unit (51_6) can output biometric data.

[0099] In an exemplary embodiment, the biometric information unit (51_6) may include sensors that sense an electrocardiogram (ECG) (or EKG, an abbreviation for the German word "Elektrokardiogramm"), an Electro Encephalo Graphy (EEG), and a Photoplethysmogram (PPG). An ECG interprets the electrical activity of the heart at a given time. An electrocardiogram is recorded by electrodes attached to the skin and equipment outside the body. An EEG is an electrical signal that can indirectly measure the electrical activity of nerve cells that make up the brain through electrodes on the scalp. In other words, brain waves indirectly capture electrical activity information occurring inside the brain through an electric field. A PPG is also referred to as a photoplethysmography, a photoplethysmography, a photoplethysmography, etc. A PPG is identified through a minute amount of blood flow that changes according to a pulse wave. When a PPG sensor shines light onto the skin, the amount of light absorbed varies depending on blood flow. By measuring how much light is absorbed, changes in blood flow can be identified. The instruction (52) may be information entered through a prompt.

[0100] The prediction unit (500) may include an encoder (510) and a transformer learning model (530). The encoder (510) may receive various data provided from a measurement device (51), encode the data to extract characteristics of the data, and provide the encoded data to the transformer learning model (530). In an exemplary embodiment, the encoder (510) may include a demeanor encoder (511) and a behavioral encoder (512). The demeanor encoder (511) may encode IMU data, audio data, and video data, and provide the encoded data to the transformer learning model (530). In an exemplary embodiment, a demagnetizer encoder (511) may encode IMU data, LFA data, and video data, align video features, IMU features, and LFA features in the encoded data, and bind each feature to an image to provide the bound data to a transformer learning model (530). A behavioral encoder (512) may encode audio data, video data, all-factory data, etc., and provide the encoded data to a transformer learning model (530). In an exemplary embodiment, the behavioral encoder (512) may encode first audio data (e.g., LFA data and HFA data) and second audio data as an audio encoder, and output text tokens for the encoded data based on a Q-Former. Text tokens and instructions (52) may be synthesized (520), and the synthesized data may be provided to a transformer learning model (530). The Transformer learning model (530) may be the aforementioned artificial intelligence model. In an exemplary embodiment, the Transformer learning model (530) may be a decoder-based artificial intelligence language model. The Transformer learning model (530) may take synthetic data and encoded data as input and predict interpretation data representing text interpreting a companion animal as output.A detailed description of the transformer learning model (530) is described below with reference to FIG. 6.

[0101] FIG. 6 is a diagram schematically illustrating an artificial intelligence model according to an exemplary embodiment of the present disclosure.

[0102] Referring to FIG. 6, the artificial intelligence model (600) may be referred to as PetAI as a multi-modal language model. The artificial intelligence model (600) may receive input data (611) from a device (60). The device (60) may be worn, coupled, or attached to a companion animal (e.g., a dog) and may be configured with a collar and an application. In an exemplary embodiment, the input data (611) may include first audio data, second audio data, all-factory data, video data, IMU data, biometric data, first text data, second text data, and third text data, such as the behavioral data set described above. The artificial intelligence model (600) may further receive instructions (612). In an exemplary embodiment, when a user inputs input data (611) and instructions (612) through a prompt, the input data (611) and instructions (612) are encoded and converted into a plurality of tokens, and the plurality of tokens can be input into an artificial intelligence model (600). The artificial intelligence model (600) can predict and output interpretation data.

[0103] In an exemplary embodiment, the interpretation data may include text describing the appearance (or appearance) of the pet, including the pet's posture, the pet's behavior, and the pet's expression, and text describing an expected circumstance for the pet.

[0104] In an exemplary embodiment, the interpretation data may further include text describing the pet's appearance and situation, as well as text describing the pet's current emotions and text describing the pet's next action.

[0105] In an exemplary embodiment, the text of the interpretation data may be output in a descriptive format (631) that describes the text of the interpretation of the companion animal in one or more complete sentences for the instruction text (621) entered into the prompt.

[0106] According to an exemplary embodiment, a prompt (621) for instruction text may be as follows: “Describe the dog's overall appearance, posture, and behavior in detail, from the general to the specific. Also, describe the situation the dog is currently in, what its expression and mood seem to be, and estimate its emotions. Lastly, predict what the dog's next action might be.”

[0107] In an exemplary embodiment, the training and application of an AI model may require detailed answers to maximize the capabilities of large-scale language models. In the descriptive format (631), objective and subjective questions may be included, such as emotion, appearance, posture, and feelings. An objective describing the current appearance, etc., corresponds to the output as "P(current)", and an objective describing the next appearance, etc., may be expressed as in Mathematical Expression 1 below.

[0108] [Mathematical Formula 1]

[0109] Future_Action=argmax_future_action_i{P(future_action_i l current_situation)}

[0110] Here, the current situation includes the current dog's information, appearance, behavior, and environmental information. Among the possible future_action_i, the one with the highest probability can be designated as the Future_Action. The predicted Future_Action is trained so that the GT results are the same as the learning model.

[0111] In an exemplary embodiment, by inputting instructions (612) and / or prompts (621) into images obtained from video data of a pet using a high-level Vision-Language LLM, such as generative AI (e.g., ChatGPT 4.0), ground truth can be obtained, which PetAI can then use to interpret the input data. In this case, techniques such as LoRA can be utilized for fine-tuning the learning model and applied to learning.

[0112] Output data can be divided into present-related and future-related data. The dog's future appearance and behavior are determined based on analysis of given input data from the present perspective.

[0113] The exemplary embodiment of the present disclosure is obtained in the form of a descriptive output, thereby allowing relatively free use of the output format of the LLM and maximizing the inherent capabilities of the LLM.

[0114] In an exemplary embodiment, the text of the interpretation data may be output in a question-and-answer format (632), which matches the texts interpreting the companion animal with answers to the question text entered in the prompt (622). The purpose of requiring answers to these multiple-choice questions is to provide a consistent and detailed description of the companion animal's appearance, facilitate easy evaluation, and analyze the model's output logits. The list of questions in the question-and-answer format (632) is a simple example, and in reality, there may be more questions. Since audio or IMU do not see the appearance, it is sometimes difficult to interpret the pose itself. To improve this difficulty, the prediction accuracy can be improved by presenting the text to the user in a question-and-answer format (632) as a constraint regulation.

[0115] While LLM's capabilities can be relied on to analyze pet behavior, numerically analyzing LLM results and utilizing them in services requires more specific and limited output data formats and data structures. For example, LLM can describe a dog's movements using a wide variety of expressions. However, this presents a challenge: categorizing the dog's movements is difficult. To overcome this, providing multiple answer options or alternatives allows for the responses to be categorized within a predefined classification system. For example, limiting a dog's behavior to six behavioral patterns—"stay still," "go forward," "go backward," "climb," "go down," and "roll"—limits the categories of dog movement, making analysis easier. In addition, if LLM is analyzed at the code level, the predicted probability for each movement of the dog can be calculated, so quantitative analysis and services such as "stay still (probability 10%), go forward (probability 70%), go backward (probability 2%), go up (probability 10%), and go down (probability 8%)" are possible, and probabilistic behavioral decisions applied with Bayesian probability distribution and Markov chain can also be achieved.

[0116] The exemplary embodiments of the present disclosure, like the Descriptive Output embodiments, allow for fine-tuning using ground truth. For example, the few-shot learning Multiple Choice Q&A method can obtain a desired output value by providing a desired question and answer and an example of an answer through few-shot learning using a small amount of learning and inference.

[0117] In the case of the ground truth regarding the mismatch between the output of the artificial intelligence model (600) and the sensing value, the result of the predicted appearance can be verified based on video data, and the situation can be verified based on data representing experimental conditions.

[0118] According to the above-described embodiments, there is an effect of being able to provide an adaptive artificial intelligence model tailored to the temperament of each companion animal with respect to the AI ​​service.

[0119] Additionally, according to the embodiments described above, there is an effect that multiple adaptive artificial intelligence models can be trained and distributed using the system of the present disclosure.

[0120] In addition, according to the above-described embodiments, there is an effect of being able to provide a multi-modal language model (MLLM) that can predict output as accurately as possible based on a given input.

[0121] In addition, according to the above-described embodiments, there is an effect of providing a PetAI model with the desired performance in the present disclosure by matching predictable problems and given inputs as accurately as possible, and by providing answers that reflect the probability distribution of actual data for stochastic problems.

[0122] FIG. 7 is a flowchart illustrating a method of operating an electronic device according to an exemplary embodiment of the present disclosure.

[0123] Referring to FIG. 7, a step (S100) of acquiring a behavioral data set is performed. The behavioral data set may include first audio data representing audio of a companion animal, second audio data representing the voice of the companion animal, olfactory data representing a smell (olfactory) generated around the companion animal, video data representing a video in which the companion animal was filmed, IMU data representing an IMU (Inertial Measurement Unit) sensing result corresponding to a behavioral pattern of the companion animal, biometric data representing biometric information of the companion animal, first text data representing the surrounding environment of the companion animal in text, second text data representing the profile of the companion animal in text, and third text data representing experimental conditions of the companion animal in text.

[0124] A step (S200) of generating category-specific inference data is performed by preprocessing the behavioral data set based on the influence of one category of data on another category in the behavioral data set.

[0125] A step (S300) of creating a training data set based on category-specific inference data is performed.

[0126] While the present disclosure has primarily focused on electronic devices, it is not limited thereto. For example, a computer program stored on a recording medium that executes the operating method of FIG. 10 in combination with such hardware is also within the scope of the present disclosure.

[0127] While this disclosure has been described with reference to the embodiments illustrated in the drawings, these are merely exemplary, and those skilled in the art will appreciate that various modifications and equivalent alternative embodiments are possible. Therefore, the true scope of technical protection of this disclosure should be determined by the technical spirit of the appended claims.

Claims

1. Obtaining a behavioral data set including first audio data representing audio of a companion animal, second audio data representing the companion animal's voice, third audio data of a non-audible frequency band generated around the companion animal, all-factory data representing a smell generated around the companion animal, video data representing a video in which the companion animal was filmed, IMU data representing an IMU sensing result corresponding to a behavioral pattern of the companion animal, biometric data representing biometric information of the companion animal, first text data representing the surrounding environment of the companion animal in text, second text data representing the profile of the companion animal in text, and third text data representing experimental conditions based on metadata designed from experimental conditions including a specific cycle and bundle amount of the companion animal, and performing frequency analysis on at least some of the first to third audio data to generate spectrogram data for each of the audible and non-audible bands, object recognition is performed on the video data, signal processing is performed on the IMU data, and based on the influence of one category of data on another category in the behavioral data set, A preprocessing unit that generates preprocessed data by performing labeling on at least one of the first and second audio data, generates category-specific inference data including context information for random inputs by category by learning based on the preprocessed data, and generates a training data set based on the category-specific inference data; and An electronic device comprising a prediction unit that predicts interpretation data representing text interpreted about the companion animal as an output using a multi-modal artificial intelligence language model trained based on the training data set, using a newly input behavioral data set as an input.

2. In paragraph 1, The above preprocessing unit, An electronic device characterized in that it receives the first audio data, the second audio data, the all-factory data, and the IMU data from a first sensing device attached to the companion animal, receives the video data from a second sensing device located around the companion animal, and receives the biometric data, the first text data, the second text data, and the third text data from an external device, or stores the biometric data, the first text data, the second text data, and the third text data in advance.

3. In paragraph 1, The above preprocessing unit, An electronic device characterized in that it includes a context parser that generates inference data for each of the above-mentioned all-factory data, the biometric data, the first text data, the second text data, and the third text data, analyzes the influence between the first audio data, the second audio data, the video data, and the IMU data, generates inference data for each of the first audio data, the second audio data, the video data, and the IMU data based on the influence, and synthesizes all generated inference data, thereby generating the training data set.

4. In paragraph 3, The above context parser is, An electronic device characterized in that the electronic device generates spectrogram data by analyzing a spectrogram for frequency in the first audio data and the second audio data, generates object recognition data by recognizing an object in the video of the video data, generates synthetic data by synthesizing a frequency analysis result and an object recognition result, generates signal-processed data by performing a signal processing operation on the IMU data, generates a plurality of labeling data by performing a labeling operation on each of the first audio data and the second audio data and the synthetic data, generates inference data corresponding to the IMU data by synthesizing the signal-processed data and the plurality of labeling data, and generates the training data set by synthesizing all of the generated inference data.

5. In paragraph 1, The above interpretation data is, An electronic device characterized by including text describing the appearance of the companion animal, including the posture of the companion animal, the behavior of the companion animal, and the facial expression of the companion animal, and text describing an expected situation for the companion animal.

6. In paragraph 5, The above interpretation data is, An electronic device further characterized by including text describing the current emotion of the pet, and text describing the next action to be taken by the pet.

7. In paragraph 1, The above multi-modal artificial intelligence language model is a generative language model, An electronic device characterized in that the text of the above interpretation data is output in a descriptive format that describes texts that interpret the companion animal in the form of one or more complete sentences for the instruction text entered in the prompt.

8. In paragraph 1, The above multi-modal artificial intelligence language model is a generative language model, An electronic device characterized in that the text of the above interpretation data is output in a question-answer format that matches the texts interpreted about the companion animal with answers to the question text entered in the prompt.

9. A step of obtaining a behavioral data set including first audio data representing audio of a companion animal, second audio data representing the companion animal's voice, third audio data of an audible frequency band occurring around the companion animal, all-factory data representing a smell occurring around the companion animal, video data representing a video in which the companion animal was filmed, IMU data representing an IMU sensing result corresponding to a behavioral pattern of the companion animal, biometric data representing biometric information of the companion animal, first text data representing the surrounding environment of the companion animal in text, second text data representing the profile of the companion animal in text, and third text data representing experimental conditions based on metadata designed from experimental conditions including a specific cycle and bundle amount of the companion animal in text; A step of generating spectrogram data for each of the audible and inaudible bands by performing frequency analysis on at least some of the first to third audio data; A step of generating preprocessed data by performing labeling on at least one of the first and second audio data based on the influence of one category of data on another category in the above behavioral data set; A step of generating category-specific inference data including context information for random inputs by category by learning based on the above preprocessing data; and A method of operating an electronic device, comprising the step of generating a training data set based on the above category-specific inference data.

10. A computer program stored on a recording medium that executes the method of Article 9 in combination with hardware.

Citation Information

Patent Citations

  • Polymer films and electronic devices

    KR1020200143295A

  • Language model for processing a multi-mode query input

    US20230350936A1

  • KR20220055519A

  • KR20230103666A