Information processing apparatus, information processing method, and information processing program

The information processing apparatus improves animal emotion estimation by integrating body movements and vocalizations with environmental factors, addressing limitations of visual-only systems and enhancing accuracy through weight coefficients and pre-trained models.

JP7713279B1Active Publication Date: 2025-07-25SHISHIMARO CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025073778
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-25
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing animal emotion estimation systems rely primarily on visual data, which can lead to estimation errors due to the limited information available, and do not effectively utilize vocalizations for accurate emotion assessment.

Method used

An information processing apparatus that estimates animal emotions by acquiring both body movements and vocalizations, incorporating environmental factors like time, location, and weather to calculate weight coefficients, and integrates these with pre-trained models for enhanced emotion estimation.

Benefits of technology

Enhances the accuracy of animal emotion estimation by leveraging both body movements and vocalizations, adjusting for environmental influences, and prioritizing movement-based estimation to improve overall precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713279000001_ABST
    Figure 0007713279000001_ABST
Patent Text Reader

Abstract

Provided is a technique capable of more appropriately estimating an animal's emotion based on the movements and sounds of the animal's body. **Solution**: The information processing apparatus of the present disclosure is an information processing apparatus that estimates an animal's emotion based on the animal's biological information. This information processing apparatus includes a control unit that executes: acquiring the movements and sounds of the animal's body, which are biological information; acquiring environmental information regarding the environment to which the animal belongs; and estimating the animal's emotion based on the movements and sounds of the animal's body. Then, the control unit calculates a weight coefficient based on the environmental information for the biological information, and estimates the animal's emotion based on the movements and sounds of the animal's body including the weight coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program for estimating an animal's emotion based on the animal's biological information.

Background Art

[0002] Conventionally, animals such as dogs and cats have been kept as pets. And such animals tend to express emotions according to their owners and the surrounding situations, so the owners often desire to understand those emotions.

[0003] However, since animals cannot speak and explain their own emotions, systems for estimating animal emotions have been proposed.

[0004] For example, Patent Document 1 discloses an emotion determination system that generates behavior information related to an animal's behavior from an image of the animal and determines the animal's emotion based on the behavior information.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] As conventionally known, it seems that an animal's emotion can be estimated from behavior information on what actions the animal is taking. On the other hand, animals may change their vocalizations according to their mood, vigilance, physical condition, and physiological phenomena. Therefore, it is considered that an animal's emotion can also be estimated by using the sound emitted by the animal.

[0007] Here, the technology described in Patent Document 1 determines the emotion of an animal based on an image of the animal, and is not a technology that uses the voice emitted by the animal for estimating the emotion of the animal. Then, in the estimation of the emotion of an animal based only on the image of the animal as in the technology described in Patent Document 1, there is a risk that an estimation error may easily occur due to the small amount of information used for the estimation.

[0008] An object of the present disclosure is to provide a technology capable of more appropriately estimating the emotion of an animal based on the body movements and voice of the animal.

Means for Solving the Problem

[0009] The information processing apparatus of the present disclosure is an information processing apparatus that estimates the emotion of an animal based on the biological information of the animal. This information processing apparatus includes a control unit that executes acquiring the body movements and voice of the animal, which are the biological information, acquiring environmental information regarding the environment to which the animal belongs, and estimating the emotion of the animal based on the body movements and voice of the animal. Then, the control unit calculates a weight coefficient based on the environmental information for the biological information, and estimates the emotion of the animal based on the body movements and voice of the animal including the weight coefficient.

[0010] And, in the above information processing apparatus, the environmental information includes information regarding time, place, and weather, the weight coefficient is a coefficient representing the influence of the environmental information on the acquisition of the voice of the animal, and the control unit may estimate the emotion of the animal by adding together the product of the weight coefficient and the emotion of the animal estimated based on the body movements of the animal and the emotion of the animal estimated based on the voice of the animal.

[0011] In this case, the weight coefficient multiplied by the emotion of the animal estimated based on the voice of the animal is increased when the time information included in the environmental information is during the day and late at night compared to when it is in the morning and evening, and is increased when the location information included in the environmental information is indoors compared to when it is outdoors, and can be increased when the weather information included in the environmental information is sunny and cloudy compared to when it is rainy. Further, in the weight coefficient, the weight coefficient multiplied by the emotion of the animal estimated based on the movement of the animal's body can be made larger than the weight coefficient multiplied by the emotion of the animal estimated based on the voice of the animal.

[0012] Further, the control unit may acquire information regarding the movement of the animal's body as video data, acquire the size of a patch extracted from the viewing angle in the video data, divide the video data into the patches so that a plurality of frames constituting the time series of the video data are included, extract feature parts for each frame for each of the patches divided from the video data, extract time-series feature parts based on the feature parts for each frame for each of the patches divided from the video data, and estimate the emotion of the animal based on the time-series feature parts for each of the patches divided from the video data.

[0013] Further, the control unit may acquire information regarding the voice of the animal as voice data, extract a plurality of characteristic sounds from the voice data, extract an immediate sound whose voice duration is less than a predetermined time from among the plurality of characteristic sounds, and estimate the emotion of the animal based on the characteristic sounds excluding the immediate sound among the plurality of characteristic sounds.

[0014] In addition, the present disclosure can be understood from the aspect of an information processing method by a computer. That is, the information processing method of the present disclosure is an information processing method for estimating the emotion of an animal based on the biological information of the animal, and the computer acquires the body movements and voices of the animal, which are the biological information, and acquires environmental information regarding the environment to which the animal belongs, and estimates the emotion of the animal based on the body movements and voices of the animal, and the computer calculates a weight coefficient based on the environmental information for the biological information, and estimates the emotion of the animal based on the body movements and voices of the animal including the weight coefficient.

[0015] In addition, the present disclosure can be understood from the aspect of an information processing program. That is, the information processing program of the present disclosure is an information processing program for estimating the emotion of an animal based on the biological information of the animal, and causes the computer to acquire the body movements and voices of the animal, which are the biological information, and acquire environmental information regarding the environment to which the animal belongs, and estimate the emotion of the animal based on the body movements and voices of the animal, and causes the computer to calculate a weight coefficient based on the environmental information for the biological information, and estimate the emotion of the animal based on the body movements and voices of the animal including the weight coefficient.

Advantages of the Invention

[0016] According to the present disclosure, the emotion of an animal can be estimated more appropriately based on the body movements and voices of the animal.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0018] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The configurations of the following embodiments are examples, and the present disclosure is not limited to the configurations of the embodiments.

[0019] <First Embodiment> The outline of the emotion estimation system in the first embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing the schematic configuration of the emotion estimation system in this embodiment. The emotion estimation system 100 according to this embodiment includes a network 200, a server 300, and a user terminal 400. Note that the emotion estimation system of the present disclosure is a system that estimates the emotion of an animal based on the biological information of the animal, and the emotion estimation of the animal is executed by the server 300. Also, the animal in this embodiment is, for example, a pet such as a dog or a cat.

[0020] The network 200 is, for example, an IP network. If the network 200 is an IP network, it may be wireless, wired, or a combination of wireless and wired. For example, if it is wireless communication, the user terminal 400 may access a wireless LAN access point (not shown) and communicate with the server 300 via a LAN or WAN. Further, the network 200 is not limited to these examples, and may be, for example, a public switched telephone network, an optical fiber line, an ADSL line, a satellite communication network, or the like.

[0021] The server 300 is connected to the user terminal 400 via the network 200. In FIG. 1, for the sake of simplicity of explanation, one server 300 and four user terminals 400 are shown, but it goes without saying that these are not limited thereto.

[0022] The server 300 may be any electronic device as long as it is a computer device having a processing ability for arithmetic processing and processing such as data acquisition, generation, and update. For example, it may be a personal computer, a server, a mainframe, or other electronic devices. That is, the server 300 can be configured as a computer having a processor such as a CPU or GPU, a main storage device such as a RAM or ROM, and an auxiliary storage device such as an EPROM, a hard disk drive, or a removable medium. The removable medium may be, for example, a USB memory or a disk recording medium such as a CD or DVD. The auxiliary storage device stores an operating system (OS), various programs, various tables, and the like.

[0023] Further, the server 300 may appropriately use SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service) provided by a cloud server without providing software, hardware, an OS, etc. dedicated to the emotion estimation system 100 according to the present embodiment.

[0024] The user terminal 400 may be an electronic device such as a mobile terminal owned by a user (who may be a pet owner or an operator providing animal emotion estimation services, etc.) that uses the emotion estimation system 100. For example, it may be a mobile terminal, a tablet terminal, a smartphone, a wearable terminal, a personal computer, or other terminal devices.

[0025] Next, based on FIG. 2, the components of the server 300 will be mainly described in detail. FIG. 2 is a diagram showing in more detail the components of the server 300 included in the emotion estimation system 100 in the first embodiment, and also showing the components of the user terminal 400 that communicates with the server 300.

[0026] The server 300 has a communication unit 301, a storage unit 302, and a control unit 303 as functional units. It loads the program stored in the auxiliary storage device into the working area of the main storage device and executes it. By controlling each functional unit through the execution of the program, each function that meets the predetermined purpose in each functional unit can be realized. However, part or all of the functions may also be realized by a hardware circuit such as an ASIC or an FPGA.

[0027] Here, the communication unit 301 is a communication interface for connecting the server 300 to the network 200. The communication unit 301 includes, for example, a network interface board or a wireless communication circuit for wireless communication. The server 300 is communicably connected to the user terminal 400 and other external devices via the communication unit 301.

[0028] The storage unit 302 is configured to include a main storage device and an auxiliary storage device. The main storage device is a memory in which programs executed by the control unit 303 and data used by the control programs are expanded. The auxiliary storage device is a device that stores programs executed in the control unit 303 and data used by the control programs. The storage unit 302 stores data transmitted from the user terminal 400 or the like, and the storage unit 302 stores biometric information (animal video data, audio data) and environmental information, which will be described later. Note that the server 300 can acquire data transmitted from the user terminal 400 or other external devices via the communication unit 301.

[0029] The control unit 303 is a functional unit that controls the operations performed by the server 300. The control unit 303 can be realized by an arithmetic processing device such as a CPU. The control unit 303 further includes four functional units: a first acquisition unit 3031, a second acquisition unit 3032, a calculation unit 3033, and an estimation unit 3034. Each functional unit may be realized by executing a stored program using the CPU.

[0030] The first acquisition unit 3031 acquires the body movements and sounds of an animal, which are biometric information of the animal. Here, the first acquisition unit 3031 acquires the body movements of the animal from video data and the sounds (such as cries) of the animal from audio data. The first acquisition unit 3031 acquires biometric information by acquiring this data transmitted from the user terminal 400 of the user who uses the emotion estimation system 100, and stores it in the storage unit 302.

[0031] Here, the user terminal 400 in this embodiment has a communication unit 401, an input / output unit 402, and a storage unit 403 as functional units. The communication unit 401 is a communication interface for connecting the user terminal 400 to the network 200, and is configured to include, for example, a network interface board or a wireless communication circuit for wireless communication. The input / output unit 402 is a functional unit for displaying information such as information transmitted from the outside via the communication unit 401, or for inputting the information when transmitting the information to the outside via the communication unit 401. The storage unit 403 is configured to include a main storage device and an auxiliary storage device, similar to the storage unit 302 of the server 300.

[0032] The input / output unit 402 further has a display unit 4021, an operation input unit 4022, and an image / audio input / output unit 4023. The display unit 4021 has a function of displaying various information, and is realized by, for example, an LCD (Liquid Crystal Display) display, an LED (Light Emitting Diode) display, an OLED (Organic Light Emitting Diode) display, etc. The operation input unit 4022 has a function of receiving operation input from the user, and is specifically realized by soft keys or hard keys such as a touch panel. The image / audio input / output unit 4023 has a function of receiving input of images such as still images and moving images, and is specifically realized by a camera using an image sensor such as Charged-Coupled Devices (CCD), Metal-oxide-semiconductor (MOS), or Complementary Metal-Oxide-Semiconductor (CMOS). Further, the image / audio input / output unit 4023 has a function of receiving input / output of audio, and is specifically realized by a microphone and a speaker.

[0033] Then, a user who uses the emotion estimation system 100 can transmit video data and audio data to the server 300 using the user terminal 400 configured as described above. Note that the user can transmit video data and audio data including the movements and sounds of the animal's body, which are input to the user terminal 400 using the image / audio input / output unit 4023, from the user terminal 400 to the server 300. Further, the server 300 may provide an interface for inputting and transmitting these data to the user's user terminal 400.

[0034] The second acquisition unit 3032 acquires environmental information. Here, the environmental information is information regarding the environment to which the animal belongs, and specifically includes information regarding time, location, and weather. The second acquisition unit 3032 acquires environmental information by acquiring information regarding time and location transmitted from the user terminal 400 of the user who uses the emotion estimation system 100, and by acquiring information regarding weather from an external device, and stores it in the storage unit 302.

[0035] The calculation unit 3033 calculates a weight coefficient based on the environmental information with respect to the biological information. Here, the weight coefficient is a coefficient representing the influence of the environmental information on the acquisition of the animal's voice. The details of the process executed by the calculation unit 3033 will be described based on FIG. 4 described later.

[0036] The estimation unit 3034 estimates the emotion of the animal based on the movements and sounds of the animal's body including the weight coefficient. The details of the process executed by the estimation unit 3034 will be described based on FIGS. 5 to 9 described later.

[0037] Note that the control unit 303 functions as the control unit according to the present disclosure by executing the processes of the first acquisition unit 3031, the second acquisition unit 3032, the calculation unit 3033, and the estimation unit 3034.

[0038] Next, the operation flow of the emotion estimation system 100 in the present embodiment will be described. FIG. 3 is a first diagram illustrating the operation flow of the emotion estimation system 100 in the present embodiment. In FIG. 3, the operation flow between the server 300, the user terminal 400, and the external device in the emotion estimation system 100 in the present embodiment, and the processes executed by the server 300, the user terminal 400, and the external device are described.

[0039] In the present embodiment, first, animal video data and audio data are input to the user terminal 400 of the user who uses the emotion estimation system 100 (S101). For example, the user can input the animal video data and audio data to the user terminal 400 by using the image / audio input / output unit 4023 of the user terminal 400 to capture the body movements and sounds of the pet animal. Then, the data input in this way is transmitted from the user terminal 400 to the server 300. Then, the server 300 acquires the biological information of the animal transmitted from the user terminal 400 (S102) and stores it in the storage unit 302.

[0040] Furthermore, time data and location data, which are environmental information in the present embodiment, are transmitted from the user terminal 400 to the server 300 (S103). Here, the time data and location data are data related to the time and location when the body movements and sounds of the animal are captured, and can be input to the user terminal 400 simultaneously as metadata for the above video data and audio data. Then, the server 300 acquires these environmental information transmitted from the user terminal 400 (S104) and stores it in the storage unit 302.

[0041] In addition, weather data and additional data, which are environmental information in the present embodiment, are transmitted from an external device to the server 300 (S105). Here, the weather data is data regarding the weather at the time when the movement or voice of the animal's body is photographed. For example, when the movement or voice of the animal's body is photographed using a predetermined application installed on the user terminal 400, information regarding the weather at the photographed location can be obtained by the application accessing a predetermined website. That is, the weather data is transmitted from the external device to the server 300 via the above application. The additional data is other additional data regarding the environment to which the animal belongs. For example, it is data regarding activities (events, construction, etc.) being carried out around the photographed location when the movement or voice of the animal's body is photographed. And such additional data can also be transmitted from the external device to the server 300 via the above application. Then, the server 300 acquires these environmental information transmitted from the external device (S106) and stores it in the storage unit 302.

[0042] Next, the server 300 calculates a weighting coefficient (S107). Here, the weighting coefficient is a coefficient representing the influence of environmental information on the acquisition of the animal's voice. As will be described later, it is multiplied by the emotion of the animal estimated based on the movement of the animal's body and the emotion of the animal estimated based on the voice of the animal, respectively.

[0043] In the weighting coefficient in the present embodiment, the weighting coefficient multiplied by the emotion of the animal estimated based on the voice of the animal is made larger when the time information included in the environmental information is daytime and midnight than when the time information is morning and night, made larger when the location information included in the environmental information is indoors than when the location information is outdoors, and made larger when the weather information included in the environmental information is sunny and cloudy than when the weather information is rainy.

[0044] Here, FIG. 4 is a diagram illustrating the weighting coefficient multiplied by the emotion of the animal estimated based on the voice of the animal in the present embodiment.

[0045] FIG. 4(a) is a diagram illustrating a weight coefficient representing the influence of time information on the acquisition of animal voices. The weight coefficients in the daytime (e.g., 12:00 to 18:00) and late at night (e.g., 24:00 to 6:00) divisions are made larger than the weight coefficients in the morning (e.g., 6:00 to 12:00) and evening (e.g., 18:00 to 24:00) divisions. This is because it has been found that the daytime and late-night time zones are more likely to provide a quieter environment for animals than the morning and evening time zones. That is, it has been found that the daytime and late-night time zones have less influence of ambient noise on the acquisition of animal voices than the morning and evening time zones. Thus, in time zones where the influence of ambient noise on the acquisition of animal voices is small, the weight multiplied by the animal emotion estimated based on the animal voice is increased.

[0046] Also, FIG. 4(b) is a diagram illustrating a weight coefficient representing the influence of location information on the acquisition of animal voices. The weight coefficient in the indoor division is made larger than the weight coefficient in the outdoor division. This is because it has been found that indoors is more likely to provide a quieter environment for animals than outdoors. That is, it has been found that indoors has less influence of ambient noise on the acquisition of animal voices than outdoors. Thus, in locations where the influence of ambient noise on the acquisition of animal voices is small, the weight multiplied by the animal emotion estimated based on the animal voice is increased.

[0047] Also, FIG. 4(c) is a diagram illustrating a weight coefficient representing the influence of weather information on the acquisition of animal voices. The weight coefficients in the sunny and cloudy divisions are made larger than the weight coefficients in the rainy, thunderous, and snowy divisions. This is because it has been found that sunny and cloudy weather is more likely to provide a quieter environment for animals than rainy, thunderous, and snowy weather. That is, it has been found that sunny and cloudy weather has less influence of ambient noise on the acquisition of animal voices than rainy, thunderous, and snowy weather. Thus, in weather where the influence of ambient noise on the acquisition of animal voices is small, the weight multiplied by the animal emotion estimated based on the animal voice is increased.

[0048] In addition, among the above weight coefficients, the weight coefficient multiplied by the emotion of the animal estimated based on the movement of the animal's body is made larger than the weight coefficient multiplied by the emotion of the animal estimated based on the animal's voice.

[0049] Specifically, when the environment where the animal belongs is outdoors in the morning time zone and the weather is rainy, the weight coefficient multiplied by the emotion of the animal estimated based on the movement of the animal's body is 0.9, and the weight coefficient multiplied by the emotion of the animal estimated based on the animal's voice is 0.1. On the other hand, when the weather is sunny indoors during the daytime, the weight coefficient multiplied by the emotion of the animal estimated based on the movement of the animal's body is 0.7, and the weight coefficient multiplied by the emotion of the animal estimated based on the animal's voice is 0.3.

[0050] Note that even in an environment where the influence of ambient noise on the acquisition of the animal's voice is small, in the estimation of the animal's emotion based on the movement of the animal's body and the animal's voice, the weight coefficient multiplied by the emotion of the animal estimated based on the movement of the animal's body is 0.7, and the weight coefficient multiplied by the emotion of the animal estimated based on the animal's voice is 0.3. In this way, the emotion of the animal estimated based on the movement of the animal's body is prioritized. Thereby, while the voice emitted from the animal is also used for the estimation of the animal's emotion, the emotion estimation based on the movement of the animal's body can be performed more accurately.

[0051] Then, returning to FIG. 3, next, the server 300 executes a process of estimating the emotion of the animal (S108). This will be described in detail with reference to FIG. 5.

[0052] FIG. 5 is a second diagram illustrating the operation flow of the emotion estimation system 100 in the present embodiment. In FIG. 5, the process executed by the server 300 in the emotion estimation system 100 in the present embodiment is described.

[0053] In the flow shown in FIG. 5, after the process of S107 shown in FIG. 3 above, a process of estimating the emotion of an animal based on voice data (S1081) and a process of estimating the emotion of an animal based on video data (S1082) are executed. Here, in these estimation processes, by inputting voice data and video data into the pre-trained model, the estimated emotion of the animal is output. And the pre-trained model is constructed by performing learning using data including the body movements and voices of animals.

[0054] Here, FIG. 6 is a diagram for explaining the discrimination result obtained from the input to the pre-trained model in the present embodiment and the neural network constituting the pre-trained model. In the present embodiment, as the pre-trained model, a neural network model generated by deep learning is used. The pre-trained model 30 for estimating the emotion of an animal based on voice data shown in FIG. 6(a) has an input layer 31 that receives an input of predetermined voice data, an intermediate layer (hidden layer) 32 that extracts a feature amount representing the emotion of the animal from the voice data input to the input layer 31, and an output layer 33 that outputs a discrimination result based on the feature amount. Also, the pre-trained model 30 for estimating the emotion of an animal based on video data shown in FIG. 6(b) has an input layer 31 that receives an input of predetermined video data, an intermediate layer (hidden layer) 32 that extracts a feature amount representing the emotion of the animal from the video data input to the input layer 31, and an output layer 33 that outputs a discrimination result based on the feature amount. In the example of FIG. 6, the pre-trained model 30 has one intermediate layer 32, the output of the input layer 31 is input to the intermediate layer 32, and the output of the intermediate layer 32 is input to the output layer 33. However, the number of intermediate layers 32 is not limited to one layer, and the pre-trained model 30 may have two or more intermediate layers 32.

[0055] Also, according to FIG. 6, each of the layers 31 to 33 includes one or more neurons. For example, the number of neurons in the input layer 31 can be set according to the input voice data and video data. Also, the number of neurons in the output layer 33 can be set according to the emotion of the animal that is the discrimination result.

[0056] Then, neurons in adjacent layers are appropriately connected, and weights (connection loads) are set for each connection based on the results of machine learning. In the example of FIG. 6, each neuron is connected to all neurons in the adjacent layer, but the connection of neurons is not limited to such an example and can be set as appropriate.

[0057] Such a pre-trained model 30 is constructed, for example, by performing supervised learning using teacher data that is a set of voice data and video data defined as data representing characteristic animal emotions, and voice and video labels representing animal emotions. Specifically, a set of feature quantities and labels is given to a neural network, and the weights of the connections between neurons are tuned so that the output of the neural network is the same as the label. In this way, a pre-trained model for learning the characteristics of teacher data and estimating results from inputs is inductively obtained. Further, such a pre-trained model 30 may be updated based on feedback from a user who uses the emotion estimation system 100. In this case, after the server 300 transmits the result of the animal emotion estimation process in S108 of FIG. 3 described below to the user terminal 400, it acquires feedback from the user. Here, the above feedback is, for example, information on the correctness of the estimation result in the process of S108, and when the server 300 acquires such feedback, it can re-learn the above pre-trained model 30 based on the feedback information.

[0058] Note that the pre-trained model 30 may be constructed by performing unsupervised learning. For unsupervised learning, for example, domain adaptation, which is a type of transfer learning, can be used. According to this, a pre-trained model can be obtained without preparing a large amount of labeled teacher data.

[0059] Furthermore, it has been newly found that the technology of Vision Transformer, which is applied to image recognition, can be applied to the extraction of feature portions from video data. This will be described with reference to FIG. 7.

[0060] FIG. 7 is a third diagram illustrating the operation flow of the emotion estimation system 100 in the present embodiment. In FIG. 7, the processing executed by the server 300 in the emotion estimation system 100 in the present embodiment will be described.

[0061] In the flow shown in FIG. 7, in the process of S1082 shown in FIG. 5 above, first, a process of acquiring a patch size is executed (S10821). Here, the above patch size represents the size of a patch extracted by dividing from the viewing angle in video data, and the server 300 may acquire an arbitrarily predetermined and stored patch size, or may acquire a patch size input by the user.

[0062] Next, the server 300 divides the video data into patches so that a plurality of frames constituting the time series of the video data are included (S10822).

[0063] Here, FIG. 8 is a diagram for explaining the extraction of a feature portion from video data in the present embodiment.

[0064] And FIG. 8(a) is a diagram for explaining the process of dividing video data into patches. As shown in FIG. 8(a), the video data includes a plurality of frames 3 (in the example shown in FIG. 8, frames 3a to 3j) constituting a time series. Also, in the example shown in FIG. 8(a), nine patches (P1 to P9) are extracted with respect to the viewing angle of the video data (exemplified based on frame 3j in FIG. 8(a)). And when dividing the video data into patches, the divided patches are processed so as to include a plurality of frames 3.

[0065] Then, returning to FIG. 7, for each patch segmented from the video data, the server 300 executes a process of extracting feature parts frame by frame (S10823), and a process of extracting time-series feature parts based on the feature parts for each frame for each patch segmented from the video data (S10824). Note that the concept of the process of extracting feature parts frame by frame is shown in FIG. 8(b), and the concept of the process of extracting time-series feature parts based on the feature parts for each frame is shown in FIG. 8(c). These are applied by applying the technology of the Vision Transformer.

[0066] Then, based on the time-series feature parts for each patch segmented from the video data, the server 300 estimates the emotion of the animal (S10825).

[0067] Then, returning to FIG. 3, in the animal emotion estimation process in S108, the server 300 adds the product of the weight coefficient calculated in the process of S107 to the emotion of the animal estimated based on the video data representing the body movements of the animal and the emotion of the animal estimated based on the audio data representing the voice of the animal, thereby estimating the emotion of the animal. Note that when the presence of a plurality of animals is recognized in the animal video data and audio data acquired in the process of S101, the server 300 identifies the individual of each animal from among the plurality of animals, and for each of the identified animals, the process shown in FIG. 5 above can be applied to estimate the emotion. In this case, the server 300 can identify the individual of each animal from among the plurality of animals, for example, by using the above-described pre-trained model.

[0068] Here, FIG. 9 is a diagram illustrating the emotion of the animal estimated based on the video data, the emotion of the animal estimated based on the audio data, and the emotion of the animal estimated by multiplying these by the weight coefficient and adding them together.

[0069] FIG. 9(a) is a diagram showing the emotions of an animal estimated based on video data in 10 levels according to the classification of joy, anger, sorrow, and pleasure. In this embodiment, the emotion of joy is estimated to be 8, the emotion of anger is estimated to be 1, the emotion of pleasure is estimated to be 6, and the emotion of sorrow is estimated to be 3. On the other hand, FIG. 9(b) is a diagram showing the emotions of an animal estimated based on audio data in 10 levels according to the classification of joy, anger, sorrow, and pleasure. In this embodiment, the emotion of joy is estimated to be 10, the emotion of anger is estimated to be 3, the emotion of pleasure is estimated to be 6, and the emotion of sorrow is estimated to be 5.

[0070] And FIG. 9(c) is a diagram showing the emotions of an animal estimated by multiplying these by weight coefficients and summing them up in 10 levels according to the classification of joy, anger, sorrow, and pleasure. In this embodiment, the emotion of joy is estimated to be 8.6, the emotion of anger is estimated to be 1.6, the emotion of pleasure is estimated to be 6, and the emotion of sorrow is estimated to be 3.6. At this time, the weight coefficient multiplied by the emotion of the animal estimated based on the video data representing the body movement of the animal is 0.7, and the weight coefficient multiplied by the emotion of the animal estimated based on the audio data representing the voice of the animal is 0.3.

[0071] Then, returning to FIG. 3, the server 300 transmits the information regarding the emotion of the animal estimated as described above to the user terminal 400 of the user who uses the emotion estimation system 100. Then, the user terminal 400 acquires the information transmitted from the server 300 (S109), and the user can grasp the emotion of the animal, which is the pet photographed in the process of S101. As described above, the server 300 may acquire feedback from the user to whom the information regarding the emotion of the animal has been provided in this way. Then, based on the feedback information, the accuracy of the estimation process by the above-mentioned pre-trained model can be further improved.

[0072] According to the emotion estimation system 100 described above, the emotion of the animal can be estimated more appropriately based on the body movement and voice of the animal.

[0073] <Second Embodiment> The second embodiment will be described with reference to FIG. 10. In this embodiment, when estimating the emotion of an animal based on voice data, the server 300 estimates the emotion of the animal based on the characteristic sounds other than the immediate sounds among the plurality of characteristic sounds included in the voice data.

[0074] FIG. 10 is a diagram illustrating the flow of operations of the emotion estimation system 100 in this embodiment. In FIG. 10, the processing executed by the server 300 in the emotion estimation system 100 in this embodiment will be described.

[0075] In the flow shown in FIG. 10, in the process of S1081 in FIG. 5 described in the description of the first embodiment, first, a process of extracting a plurality of characteristic sounds from the voice data is executed (S10811). Here, the above-mentioned characteristic sounds are, for example, sounds corresponding to those defined in advance as sounds representing characteristic animal emotions.

[0076] Next, the server 300 extracts immediate sounds from among the above plurality of characteristic sounds, the duration of which is less than a predetermined time (S10812). Here, the immediate sound can be defined as a sound representing a mood state determined from an emotional factor that appears immediately or for a short time only. And the above-mentioned predetermined time is several seconds (for example, 3 seconds).

[0077] Here, it has been found that the immediate mood state is likely to change in its nature, and the immediate sounds representing such a mood state are not suitable for appropriately expressing the continuous emotions of animals useful to the user.

[0078] Therefore, the server 300 estimates the emotion of the animal based on the characteristic sounds other than the immediate sounds among the plurality of characteristic sounds (S10813).

[0079] Thereby, while appropriately using the voice emitted from the animal for estimating the emotion of the animal, it is possible to more accurately perform the emotion estimation based on the body movements of the animal.

[0080] <Other Modification Examples> The above-described embodiments are merely examples, and the present disclosure can be implemented with appropriate modifications without departing from the gist thereof. For example, the processes and means described in the present disclosure can be freely combined and implemented as long as no technical contradiction occurs.

[0081] Also, the process described as being performed by one device may be shared and executed by a plurality of devices. For example, the first acquisition unit 3031 and the second acquisition unit 3032 may be formed in another arithmetic processing device. At this time, these arithmetic processing devices are preferably configured to be able to cooperate well. Further, the processes described as being performed by different devices may be executed by one device. In a computer system, it is possible to flexibly change how each function is realized by a hardware configuration (server configuration).

[0082] The present disclosure can also be realized by supplying a computer program that implements the functions described in the above embodiments to a computer, and causing one or more processors included in the computer to read and execute the program. Such a computer program may be provided to the computer by a non-transitory computer-readable storage medium connectable to the system bus of the computer, or may be provided to the computer via a network. The non-transitory computer-readable storage medium includes, for example, any type of disk such as a magnetic disk (floppy (registered trademark) disk, hard disk drive (HDD), etc.), an optical disk (CD-ROM, DVD disk, Blu-ray disk, etc.), a read-only memory (ROM), a random access memory (RAM), an EPROM, an EEPROM, a magnetic card, a flash memory, an optical card, and any type of medium suitable for storing electronic instructions.

Explanation of Reference Numerals

[0083] 100 ··· Emotion Estimation System 200 ··· Network 300 ··· Server 301 ··· Communication Unit 302 ··· Storage Unit 303 ··· Control Unit 400 ··· User Terminal

Claims

1. An information processing apparatus for estimating an animal's emotion based on the animal's biological information, comprising: acquiring the physical movements and voices of the animal, which are the biological information; acquiring environmental information regarding the environment to which the animal belongs; estimating the animal's emotion based on the physical movements and voices of the animal; and a control unit that executes the above; wherein the control unit: calculates a weighting coefficient based on the environmental information for the biological information, and estimates the animal's emotion based on the physical movements and voices of the animal including the weighting coefficient. An information processing apparatus.

2. The environmental information includes information regarding time, location, and weather. The weighting coefficient is a coefficient representing the influence of the environmental information on the acquisition of the animal's voice. The control unit: estimates the animal's emotion by adding the products of the weighting coefficient and the emotion of the animal estimated based on the physical movements of the animal and the emotion of the animal estimated based on the voice of the animal. The information processing apparatus according to claim 1.

3. The weighting coefficient multiplied by the emotion of the animal estimated based on the animal's voice is made larger when the time information included in the environmental information is daytime and midnight than when the time information is morning and night, larger when the location information included in the environmental information is indoors than when the location information is outdoors, and larger when the weather information included in the environmental information is sunny and cloudy than when the weather information is rainy. The information processing apparatus according to claim 2.

4. In the weighting coefficient, the weighting coefficient multiplied by the emotion of the animal estimated based on the physical movements of the animal is made larger than the weighting coefficient multiplied by the emotion of the animal estimated based on the voice of the animal. The information processing apparatus according to claim 3.

5. The control unit: acquires information regarding the physical movements of the animal as video data; acquires the size of a patch to be extracted from the field of view in the video data; divides the video data into the patches such that a plurality of frames constituting the time series of the video data are included; extracts feature parts for each frame for each of the patches divided from the video data; extracts time-series feature parts based on the feature parts for each frame for each of the patches divided from the video data. ​ ​ estimating the emotion of the animal based on the time-series feature parts of the respective patches divided from the video data; The information processing apparatus according to claim 2.

6. The control unit, acquires information regarding the voice of the animal as voice data, extracts a plurality of characteristic sounds from the voice data, extracts an immediate sound whose voice duration is less than a predetermined time from among the plurality of characteristic sounds, and estimates the emotion of the animal based on the characteristic sounds excluding the immediate sound among the plurality of characteristic sounds; The information processing apparatus according to claim 2.

7. The control unit, acquires information regarding the movement of the animal's body as input video data and information regarding the voice of the animal as input voice data, estimates the emotion based on the movement of the animal's body by inputting the input video data into a first pre-trained model constructed by performing learning using predetermined video data, and estimates the emotion based on the voice of the animal by inputting the input voice data into a second pre-trained model constructed by performing learning using predetermined voice data, provides the estimated emotion of the animal to the user, acquires feedback from the user to whom the emotion of the animal has been provided, and updates the first pre-trained model or the second pre-trained model based on the acquired feedback. The information processing apparatus according to claim 2.

8. The control unit, identifies the individual of each animal from among the plurality of animals based on the movement or voice of the animal, and estimates the emotion for each of the identified animals; The information processing apparatus according to claim 1.

9. An information processing method for estimating an animal's emotion based on the animal's biological information, comprising: a computer, acquires the movement and voice of the animal's body, which are the biological information, acquires environmental information regarding the environment to which the animal belongs, and estimates the emotion of the animal based on the movement and voice of the animal's body, wherein the computer, calculates a weight coefficient based on the environmental information for the biological information, and estimates the emotion of the animal based on the movement and voice of the animal's body including the weight coefficient. Information processing method.

10. An information processing program for estimating an animal's emotion based on the animal's biological information, comprising: causing a computer to, Obtaining the movements and voices of the animal's body, which are the biological information; Obtaining environmental information regarding the environment to which the animal belongs; Executing estimating the emotion of the animal based on the movements and voices of the animal's body; Causing the computer to calculate a weighting coefficient based on the environmental information for the biological information, and estimate the emotion of the animal based on the movements and voices of the animal's body including the weighting coefficient; An information processing program.

Citation Information

Patent Citations

  • Method and device for translating will of animals or the like

    JP1998003479A

  • Notification control system, notification control method, and program

    JP2019091233A

  • Information processor, and information processing method

    JP2020170916A

  • Emotion determination system and method of determining emotion

    JP2023094426A