Information processing device, information processing method, and information processing program

JP2025178452APending Publication Date: 2025-12-05FUJIFILM CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025165564
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-01
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing emotion tagging systems assign emotional information based on facial expressions captured in images, which are outdated by the time the user views the content, failing to reflect the user's current emotions during an event.

Method used

An emotion tagging system that detects and analyzes voice data from event participants to assign emotion tags indicating their real-time emotions, using a processor, voice detector, and emotion recognizer to calculate and attach emotion ranks to content.

Benefits of technology

Enables the quantification of event effectiveness by reflecting user emotions during the event, allowing for improved event planning based on these tags.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025178452000001_ABST
    Figure 2025178452000001_ABST
Patent Text Reader

Abstract

To provide an emotion tagging system, method, and program for attaching an emotion tag indicating a user's emotion at the time of conducting an event using a content to the content.SOLUTION: The emotion tagging method includes a step S16 of detecting audio data representing audio uttered by a person participating in an event during the event using a content by an audio detector, a step S18 of recognizing human emotions on the basis of the audio data by an emotion recognizer, a step of acquiring emotion information indicative of human emotions recognized during the event using the content by a processor, and steps S20 and S22 of attaching an emotion rank calculated from the acquired emotion information to the content as an emotion tag by the emotion recognizer.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an emotion tagging system, method, and program, and more particularly to a technology for assigning emotion tags relating to a user's emotions to content. [Background technology]

[0002] 2. Description of the Related Art Conventionally, a system has been known in which a large amount of content data, such as image files (eg, videos, photographs) and audio files, is stored in a storage device, and files are selectively read from the storage device and played back continuously like a slideshow.

[0003] Patent Document 1 proposes an information processing system that is used in such a system and reduces the burden on the user from registering tagged content data to searching and viewing it.

[0004] The information processing system described in Patent Document 1 generates one or more pieces of tag information for content data of a user from information about the user and the user's behavior, and associates the generated one or more pieces of tag information with the content data and registers them in a content database.

[0005] In addition, when searching and browsing, the information processing system selects tag information based on information detected from the content viewing environment including the user, and searches for one or more content data from the content database based on the selected tag information and plays them sequentially.

[0006] Incidentally, Patent Document 1 describes that when tag information is added to content data of an image file, the content data (photographed image) of the image file is analyzed, emotional information is estimated from the facial expressions of people included in the photographed image, and tag information including the estimated emotional information is added to the image file. Note that the emotional information is information estimated from facial expressions such as "happy" or "sad," for example. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] International Publication No. 2020 / 158536 Summary of the Invention [Problem to be solved by the invention]

[0008] When tag information is assigned to the content data of an image file, the information processing system described in Patent Document 1 analyzes the content data (captured image) of the image file, estimates emotional information from the facial expressions of people included in the captured image, and assigns tag information including the estimated emotional information to the image file.

[0009] However, emotional information estimated from the facial expression of a person included in a captured image is emotional information at the time the captured image was captured (in the past), and is not emotional information at the time the user who viewed the captured image viewed the content.

[0010] Furthermore, Patent Document 1 describes a method of acquiring an image of the content viewing environment using a fixed observation camera or the like when viewing content, recognizing the person in the image, and estimating the emotion from the person's facial expression. However, the emotion estimated here is used when, for example, the emotion changes to a specific emotion such as "fun" (a smiling face), by updating the search tag information to "fun" tag information, and re-searching the content data and switching to playback presentation.

[0011] The present invention has been made in consideration of the above circumstances, and aims to provide an emotion tagging system, method, and program that can assign emotion tags to content that indicate the user's emotions when an event using the content is held. [Means for solving the problem]

[0012] In order to achieve the above object, a first aspect of the invention is an emotion tagging system including a processor, a voice detector that detects voice data indicating voices uttered by people participating in an event using content while the event is being held, and an emotion recognizer that recognizes human emotions based on the voice data, wherein the processor acquires emotion information indicating the human emotions recognized by the emotion recognizer while the event using the content is being held, and assigns an emotion rank calculated from the acquired emotion information to the content as an emotion tag.

[0013] According to a first aspect of the present invention, during an event using content, audio data representing voices uttered by people participating in the event is detected, emotion information representing human emotions is acquired based on the detected audio data, and an emotion rank calculated from the acquired emotion information is assigned to the content as an emotion tag.

[0014] This makes it possible to assign emotion tags to content that indicate the user's emotions at the time of the event, and quantify the effect (value of the event) on the enjoyment of event participants by time period, allowing anyone to clearly determine which content was effective. Furthermore, emotion tags can be used when holding the next event, making it possible to hold a more effective event that reflects the emotion tags.

[0015] In the emotion tagging system according to the second aspect of the present invention, the emotion recognizer is preferably a recognizer that has undergone machine learning using a large amount of speech data as training data, including speech data of speech uttered when a person is happy and speech data of speech uttered when a person is not happy. This makes it possible to accurately recognize emotion information that indicates the emotions of users (people) who participated in the event, and that indicates the degree of joy.

[0016] In the emotion tagging system according to the third aspect of the present invention, it is preferable that the content is a plurality of images, and the event is a viewing event in which the plurality of images are sequentially played back by an image playback device and the played back plurality of images are viewed.

[0017] In the emotion tagging system according to the fourth aspect of the present invention, the plurality of images preferably includes photos or videos of people who attended the event, in order to attract the interest of the participants in the event and make it more enjoyable.

[0018] In the emotion tagging system according to the fifth aspect of the present invention, it is preferable that the processor acquires from the emotion recognizer a plurality of pieces of emotion information for a time period in which a plurality of images are played back, calculates an emotion rank corresponding to each image from a representative value of the plurality of pieces of emotion information, and assigns the calculated emotion rank to each image as an emotion tag.

[0019] In the emotion tagging system according to the sixth aspect of the present invention, when multiple people are participating in an event, the processor preferably identifies one or more dominant speakers during a time period when the multiple images are played back, based on audio data detected by the audio detector, and assigns speaker identification information indicating the identified one or more dominant speakers to each image.

[0020] In the emotion tagging system according to the seventh aspect of the present invention, the processor preferably causes at least one of the emotion rank and the speaker identification information to be displayed simultaneously while the image playback device is playing back the plurality of images.

[0021] In the emotion tagging system according to the eighth aspect of the present invention, it is preferable that the processor converts the audio data detected by the audio detector into text data, and assigns at least a portion of the text data as a comment tag to a corresponding image of the plurality of images.

[0022] A ninth aspect of the invention is an emotion tagging method including the steps of: a voice detector detecting, during the event using content, audio data indicating audio spoken by humans participating in the event; an emotion recognizer recognizing human emotions based on the audio data; a processor acquiring emotion information indicating the human emotions recognized during the event using the content; and the emotion recognizer calculating an emotion rank from the acquired emotion information and assigning it to the content as an emotion tag.

[0023] In the emotion tagging method according to the tenth aspect of the present invention, it is preferable that the content is a plurality of images, and the event is a viewing event in which the plurality of images are sequentially played back by an image playback device and the played back plurality of images are viewed.

[0024] In the emotion tagging method according to the eleventh aspect of the present invention, the plurality of images preferably include photographs or videos showing people who participated in the event.

[0025] In the emotion tagging method according to the twelfth aspect of the present invention, it is preferable that the processor obtains from the emotion recognizer a plurality of pieces of emotion information for a time period in which a plurality of images are played back, calculates an emotion rank corresponding to each image from a representative value of the plurality of pieces of emotion information, and assigns the calculated emotion rank to each image as an emotion tag.

[0026] In the emotion tagging method according to the thirteenth aspect of the present invention, when multiple people are participating in an event, it is preferable that the processor identifies one or more main speakers during a time period when the multiple images are played, based on audio data detected by the audio detector, and assigns speaker identification information indicating the identified one or more main speakers to each image.

[0027] In the emotion tagging method according to the fourteenth aspect of the present invention, it is preferable that the processor simultaneously displays at least one of the emotion rank and speaker information identified by the speaker identification information while the image playback device is playing back the multiple images.

[0028] In the emotion tagging method according to the fifteenth aspect of the present invention, it is preferable that the processor converts the voice data detected by the voice detector into text data, and assigns at least a portion of the text data as a comment tag to a corresponding image of the plurality of images.

[0029] A sixteenth aspect of the invention is an emotion tagging program that causes a computer to implement the following functions: during an event using content, acquire from a voice detector voice data indicating voices spoken by people participating in the event; recognize human emotions based on the voice data; acquire emotion information indicating the human emotions recognized during the event using the content; and assign an emotion rank calculated from the acquired emotion information to the content as an emotion tag. [Effects of the Invention]

[0030] According to the present invention, it is possible to assign to content an emotion tag that indicates the user's emotion at the time of an event using the content. This makes it possible to quantify the value of the event and clearly determine which content was effective. Furthermore, by utilizing the emotion tag when holding the next event, it becomes possible to hold a highly effective event that reflects the emotion tag. [Brief explanation of the drawings]

[0031] [Figure 1] FIG. 1 is a schematic diagram showing the configuration of an emotion tagging system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram illustrating an embodiment of an emotion tagging system according to the present invention. [Figure 3] FIG. 3 is a timing chart showing the first embodiment of the emotion tagging system according to the present invention. [Figure 4] FIG. 4 illustrates an embodiment in which emotion tags are assigned to content based on audio data representing sounds made by event participants during the event. [Figure 5]FIG. 5 is a timing chart showing a second embodiment of the emotion tagging system according to the present invention. [Figure 6] FIG. 6 is a diagram showing an embodiment in which emotion tags, speaker IDs, and comment tags are assigned to content based on audio data representing sounds uttered by event participants during the event. [Figure 7] FIG. 7 is a flowchart showing a first embodiment of the emotion tagging method according to the present invention. [Figure 8] FIG. 8 is a flowchart showing a second embodiment of the emotion tagging method according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the emotion tagging system, method and program according to the present invention will now be described with reference to the accompanying drawings.

[0033] [Configuration of emotion tagging system] FIG. 1 is a schematic diagram showing the configuration of an emotion tagging system according to an embodiment of the present invention.

[0034] 1 includes a tablet PC (personal computer) 10 and a voice detector 20. The voice detector 20 may be built into the tablet PC 10.

[0035] FIG. 1 also shows a large display 30 that displays a video (image) based on a video signal from the tablet PC 10.

[0036] FIG. 2 is a block diagram showing an embodiment of an emotion tagging system according to the present invention, and in particular a block diagram of a tablet PC 10. As shown in FIG.

[0037] The tablet PC 10 constituting the emotion tagging system 1 shown in FIG. 2 includes a processor 11, a memory 12, a display unit 14, an input / output interface 16, an operation unit 18, and the like.

[0038] The processor 11 is composed of a CPU (Central Processing Unit) and the like, and controls each unit of the tablet PC 10 in an integrated manner, and also functions as, for example, an emotion recognizer 13A, an emotion rank calculation unit 13B, and an emotion tag assignment unit 13C shown in FIG.

[0039] The memory 12 includes flash memory, ROM (Read-only Memory), RAM (Random Access Memory), a hard disk drive, etc. The flash memory, ROM, or hard disk drive is a non-volatile memory that stores an operating system, an emotion tagging program according to the present invention, content (in this example, a plurality of images), various programs including a trained model that causes the tablet PC 10 (processor 11) to function as the emotion recognizer 13A, etc. The RAM functions as a work area for processing by the processor 11. It also temporarily stores the emotion tagging program, etc., stored in the flash memory, etc. Note that the processor 11 may have a portion of the memory 12 (RAM) built in.

[0040] The tablet PC 10 of this example can hold an event using content stored in the memory 12. If the content is multiple images, the tablet PC 10 can hold an appreciation event in which the multiple images are played back sequentially using the large display 30, allowing event participants to view the multiple images that have been played back.

[0041] The display unit 14 not only displays a screen for operating the tablet PC 10, but also displays a list of images (thumbnail images) to be used in a viewing event when the viewing event is held, and is also used as part of a GUI (Graphical User Interface) when receiving input instructions for selecting and switching images (photographs) from the operation unit 18 by touching the touch panel provided on the screen of the display unit 14.

[0042] The input / output interface 16 includes a connection unit connectable to an external device, a communication unit connectable to a network, etc. As the connection unit connectable to an external device, a microphone input terminal, a USB (Universal Serial Bus), an HDMI (High-Definition Multimedia Interface) (HDMI is a registered trademark), etc. can be applied.

[0043] The processor 11 can acquire audio data detected by an audio detector (microphone) 20 via the input / output interface 16. The processor 11 can also output a video signal to a large display 30 via the input / output interface 16. When holding an entertainment event, a projector may be connected to the input / output interface 16 instead of the large display 30, and the video signal may be output to the projector.

[0044] The operation unit 18 includes a start button and a volume button (not shown), as well as a touch panel provided on the screen of the display unit 14, and the touch panel functions as part of a GUI that accepts various specifications from the user.

[0045] [Event example] Examples of viewing events using multiple images include photo slideshow viewing parties held at facilities such as nursing homes for the elderly, special nursing homes for the elderly, kindergartens, and schools.

[0046] In the case of a photo slideshow viewing event held at a facility, the event organizer may be a staff member of the facility or a party contracted by the facility to host the event, and the event participants may be users of the facility.

[0047] If the event is a photo slideshow viewing party, the event organizer will prepare multiple images as content. The "multiple images" can be multiple still images (photographs), a video created from multiple photos, or both photos and videos. The "photographs" can be not only photos taken with a regular camera, but also images such as pictures and text, including drawings or short sentences drawn by the person, or drawings or text written by the person (brush calligraphy, etc.). Images that are deeply emotional to the person can also produce a reminiscence effect.

[0048] The event organizer borrows photos, albums, etc. of the event participants, and if the photos, etc. are paper-based, scans the photos with a scanner to create electronic image files and store them in the memory 12, and if the photos, etc. are image files stored in a memory card or mobile terminal, the stored image files are stored in the memory 12.

[0049] It is preferable that the photos include images of people who participated in the event (event participants). Alternatively, it is also preferable that the photos or videos include photos or videos of the event participants' family and friends, alma mater, organization or region, favorite hobbies, their own work, favorite works, favorite talents, favorite regions, facilities or events, etc. This is to attract the interest of the event participants and make them enjoy the event more. It is also preferable that image files of the event participants are saved in a folder created for each event participant.

[0050] Photo slideshow software is installed in the flash memory or the like of memory 12, and by starting the photo slideshow software, tablet PC 10 functions as an image playback device for performing a photo slideshow using multiple photos stored in memory 12.

[0051] When holding a photo slideshow, the event organizer displays a list of thumbnail images to be used in the photo slideshow on the display unit 14, and operates the tablet PC 10 while watching the reactions of the event participants, switching among multiple images (photographs) at the right time to the next photo, returning to the previous photo, skipping the next photo, or lining up two photos, thereby holding a photo slideshow in a format that creates good reactions.

[0052] Viewing events using multiple photos are not limited to the photo slideshows mentioned above, but also include viewing parties where a video created from multiple photos is created by stitching together multiple images (photos) using publicly known photo movie creation software, and applying special effects to the switching and display of the photos (photo movie), and the video is played on a tablet PC 10 and a large display 30 that function as image playback devices, and the photo movie is viewed.

[0053] <Summary of the Invention> In the present invention, during an event using content, audio data indicating the voices uttered by event participants is detected by an audio detector 20, the emotions of the event participants at the time of the event are analyzed from the audio data, emotion tags indicating the chronological emotions of the event participants are assigned to the content used in the event, and the emotion tags are used the next time an event using the same content is held, thereby enabling an event with a greater effect on the event participants to be held.

[0054] In the case of an appreciation event where a photo slideshow is viewed, emotion information indicating the emotion of the event participants is acquired for each photo used in the photo slideshow, and an emotion rank calculated from the acquired emotion information is assigned to the content (each photo) as an emotion tag.

[0055] [First embodiment of emotion tagging system] FIG. 3 is a timing chart showing the first embodiment of the emotion tagging system according to the present invention, particularly showing the case where a photo slideshow viewing event is carried out.

[0056] 3-1 in FIG. 3 shows the time-series photos that are displayed in sequence during a photo slideshow viewing event. During this viewing event, multiple images (n photos P1 to P n ) are displayed in sequence. As described above, when holding a photo slideshow, the event organizer displays a list of thumbnail images to be used in the photo slideshow on the display unit 14, and operates the tablet PC 10 (touch panel) while watching the reactions of the event participants, and displays n photos P1 to P n Switch and display as appropriate.

[0057] 3-2 in FIG. 3 is a waveform diagram of audio data (analog data) indicating the audio generated by the event participants detected by the audio detector 20 during the photo slideshow viewing event.

[0058] 3-3 in Fig. 3 is a diagram showing the emotion rank of the event participants for each photo calculated from the voice data. The calculation method of the emotion rank will be described later.

[0059] The emotion rank corresponding to each photo shown in 3-3 of FIG. 3 is assigned to the corresponding photo as an emotion tag.

[0060] 3-4 in FIG. 3 shows the playback time of each photo in the viewing event. Photo P1 is played from time t0 to t1, where t0 is the start of the photo slide show. Similarly, photo P2 is played from time t1 to t2, photo P3 is played from time t2 to t3, and so on. n is the time t n-1 ~t n It will be played for a period of time.

[0061] <Adding emotion tags> FIG. 4 illustrates an embodiment in which emotion tags are assigned to content based on audio data representing sounds made by event participants during the event.

[0062] The processor 11 shown in FIG. 2 functions as an emotion recognizer 13A, an emotion rank calculation unit 13B, and an emotion tag assignment unit 13C as shown in FIG.

[0063] The emotion recognizer 13A can be realized, for example, by executing a trained model for emotion recognition stored in the memory 12, and recognizes the emotion of the person (event participant) who made the sound based on the audio data indicating the audio spoken by the event participant during the event from the audio detector 20, and outputs emotional information indicating the emotion of the event participant from time to time.

[0064] The trained model can be configured, for example, by a convolutional neural network (CNN), which is one of the training models.

[0065] CNN can be made into a trained CNN (trained model) by performing machine learning using a large number of training data datasets shown below.

[0066] From the subject's voice data acquired at a data rate of 16.6 kHz, only the voice data from the sections judged by the evaluator to be happy was extracted in 0.2-second increments and designated as Vah, and only the voice data from the sections judged to be unhappy was extracted in 0.2-second increments and designated as Nv.

[0067] The Vt and Vin audio data were each divided into 1-second intervals, and the energy (power spectrum) for each frequency from 100 Hz to 8000 Hz was calculated every 20 ms. A 2D energy pattern was then created, with the horizontal axis representing time (10 points for 0.2 seconds every 20 ms) and the vertical axis representing frequency (80 points up to 8000 Hz every 100 Hz). The multiple 2D Vo patterns and multiple 2D Nv patterns created in this way were used as training data.

[0068] Using the above dataset of training data, a CNN with a multi-layer structure is trained by machine learning, optimizing multiple weight parameters, etc., to create a trained CNN.

[0069] Furthermore, it is preferable that the teacher data dataset is created from a plurality of voice data of subjects in the age group of facility users.

[0070] When the emotion recognizer 13A receives voice data from the voice detector 20, it creates a two-dimensional pattern of energy every 20 ms in the same way as when creating teacher data, extracts feature amounts from this two-dimensional pattern, and outputs emotion information (inference result) indicating whether the subject is happy or not.

[0071] The emotion rank calculation unit 13B acquires the emotion information in the time period of displaying one photo, which is sequentially output from the emotion recognizer 13A, and calculates, for example, the confidence level indicating happiness as a five-level emotion rank (rank A to E). In this example, it is assumed that the confidence level indicating happiness increases in the order of A < B < C < D < E.

[0072] The emotion rank can be calculated from representative values (maximum value, average value, most frequent value, etc.) among a plurality of emotion information in the time period of displaying one photo.

[0073] In this example, the emotion rank is calculated based on the confidence level indicating happiness, but it is not limited to this. For example, an emotion rank corresponding to the degree of happiness may be calculated. Also, it is not limited to the emotion of joy, and emotion ranks for different types of emotions such as joy, anger, sorrow, and pleasure may be calculated.

[0074] The emotion tagging unit 13C assigns the emotion rank calculated by the emotion rank calculation unit 13B as an emotion tag to the photo (content) displayed in the time period when the emotion information was acquired. The assignment of the emotion tag to the photo can be performed by recording the emotion tag in the header of the photo image file, or by creating a text file or the like in which the emotion tag is described in relation to the file name of the image file or the playback time period of the content.

[0075] In this way, emotion tags that indicate the emotions of users (event participants) at the time of the event can be attached to the content, the value of the event can be quantified by time period, and anyone can clearly determine which content was effective. Furthermore, emotion tags can be used when holding the next event, making it possible to hold a more effective event that reflects the emotion tags.

[0076] [Second embodiment of emotion tagging system] FIG. 5 is a timing chart showing a second embodiment of the emotion tagging system according to the present invention.

[0077] 5-1 in Figure 5 shows a chronological sequence of photos displayed sequentially during a photo slideshow viewing event, 5-2 is a waveform diagram of audio data showing the audio generated by event participants detected by audio detector 20 during the photo slideshow viewing event, 5-3 is a diagram showing the emotional rank of event participants for each photo calculated from the audio data, and 5-6 is a diagram showing the playback time of each photo during the viewing event.

[0078] 5. Since 5-1 to 5-3 and 5-6 in FIG. 5 are common to 3-1 to 3-4 shown in FIG. 3, detailed description thereof will be omitted.

[0079] FIG. 5-4 is a diagram showing speaker identification information (speaker ID (identification)) indicating one or more main speakers identified from the speech data.

[0080] 5-5 in FIG. 5 shows text data D1 to D2 converted from voice data. n FIG.

[0081] In the second embodiment shown in FIG. 5, a speaker ID indicating a main speaker identified based on audio data detected during the time period when each photo is displayed at an appreciation event is assigned to the corresponding photo, and text data D1 to D2 converted based on the audio data detected during the time period when each photo is displayed are also provided. nThis embodiment differs from the first embodiment shown in FIG. 3 in that the text data (including at least a part of the text data converted for each display time period) is added as a comment tag for the corresponding photo.

[0082] <Assigning emotion tags, speaker IDs, and comment tags> FIG. 6 is a diagram showing an embodiment in which emotion tags, speaker IDs, and comment tags are assigned to content based on audio data representing sounds uttered by event participants during the event.

[0083] The processor 11 shown in FIG. 2 functions as an emotion recognizer 13A, an emotion rank calculation unit 13B, an emotion tag assignment unit 13C, a speaker recognizer 15A, a speaker ID assignment unit 15B, a speech-to-text converter 17A, and a comment tag assignment unit 17B, as shown in FIG. 6.

[0084] In FIG. 6, parts common to those in the embodiment shown in FIG. 4 are given the same reference numerals, and detailed description thereof will be omitted.

[0085] 6, speaker recognizer 15A recognizes individuals (speakers) from their voices, and information indicating the voice waveforms of each speaker who is an event participant (for example, "voiceprint") is registered in advance in speaker recognizer 15A in association with the speaker ID, speaker information identified by the speaker ID (speaker name), etc. The speaker ID shown in 5-4 in FIG. 5 is a three-digit number.

[0086] Based on the audio data detected by the audio detector 20 during the viewing event, the speaker recognizer 15A identifies one or more main speakers during each time period when multiple photos are played back sequentially, based on the degree of match of their "voiceprints," and outputs a speaker ID indicating the identified one or more main speakers.

[0087] The speaker ID assigning unit 15B assigns one or more speaker IDs calculated by the speaker recognizer 15A to the photograph (content) displayed during the time period in which the speaker ID was acquired.

[0088] The speech-to-text converter 17A converts audio data into text data based on audio data detected by the audio detector 20 during the viewing event. The processor 11 functions as the speech-to-text converter 17A by executing known speech-to-text conversion software.

[0089] The comment tag assigning unit 17B assigns at least a part of the text data converted by the speech-to-text converter 17A as a comment tag to the photo (content) displayed during the time period when the text data was acquired.

[0090] According to the second embodiment of the emotion tagging system, it is possible to assign a speaker ID and a comment tag to content in addition to an emotion tag.

[0091] Additionally, the processor 11 can display at least one of the emotion rank, speaker information, and text data simultaneously with the corresponding photo while the photo is being played back.

[0092] This allows participants with high and low emotional rankings to be identified, allowing feedback to be used to improve the event. For example, the event can be improved by displaying photos of events with high emotional rankings in the past for participants with low emotional rankings.

[0093] [First embodiment of emotion tagging method] FIG. 7 is a flowchart showing a first embodiment of the emotion tagging method according to the present invention.

[0094] The first embodiment of the emotion tagging method shown in FIG. 7 is an emotion tagging method that is used when holding an appreciation event in which a photo slideshow is appreciated.

[0095] The event organizer prepares a plurality of photos of the event participants as content and stores them in a playable manner in memory 12. When starting a photo slideshow viewing event, the emotion tagging program is executed and the photo slideshow software is started, and a photo slideshow using the plurality of photos stored in memory 12 is performed.

[0096] In this example, a plurality of photographs are used as content, but the content is not limited to photographs, and may be a video, or may be a mixture of photographs and video.

[0097] The event organizer displays a list of thumbnail images to be used in the photo slideshow on the display unit 14, and operates the touch panel of the tablet PC 10 to select photos to be displayed on the large display 30. In this way, multiple photos (n photos P1 to P n ) Selected photos P i is displayed (step S10), where i is a parameter that specifies the currently displayed photo, and can vary within the range of 1 to n.

[0098] The photo selection instructions were given based on the reactions of the event participants, and photos P1 to P n The selection instruction may be to switch between the images in order, or to return to the previous photo or skip to the next photo.

[0099] Processor 11 is a photo processor i It is determined whether or not an instruction to end the event has been input during the display of the event (step S12). If an instruction to end the event has been input (if "Yes"), the process ends, and if an instruction to end the event has not been input (if "No"), the process proceeds to step S14.

[0100] Meanwhile, the voice detector 20 detects voice data representing voices uttered by the event participants during the viewing event, and the emotion recognizer 13A of the processor 11 recognizes the voices uttered by the event participants during the viewing event. iWhile the event is being displayed, voice data is acquired from the voice detector 20 (step S16), and the emotions of the event participants (emotions indicating whether they are happy or not) are recognized (estimated) based on the acquired voice data, and emotion information indicating the recognized emotions is sequentially output (step S18).

[0101] In step S14, the currently displayed photo P i Different photos from P j If there is no input to switch (if "No"), the process proceeds to step S10, where the currently displayed photo P i If a switching instruction is input (if "Yes"), the process proceeds to step S20.

[0102] In step S20, the processor 11 (emotion rank calculation unit 13B) calculates the plurality of pieces of emotion information (photo P i The system obtains multiple pieces of emotional information during the display time period, and calculates an emotional rank indicating happiness from the representative value of the multiple pieces of emotional information.

[0103] The processor 11 (emotion tag assigning unit 13C) assigns the emotion rank calculated in step S20 to the emotion tag T i Content (Photo P i ) (step S22).

[0104] Next, the photo P that can be switched j The parameter j of the photograph is changed to the parameter i of the photograph currently being displayed (step S24), and the process returns to step S10.

[0105] Although not shown in the flowchart of FIG. 7, at the end of the viewing event (if "Yes" in step S12), the processes of steps S20 and S22 are also performed, and the last displayed photo P i For emotion tag T i After granting the bonus, the viewing event will end.

[0106] Furthermore, if multiple people are participating in a viewing event, the processor 11 may identify one or more main speakers during the time period when multiple photos are played based on the audio data detected by the audio detector 20, and assign a speaker ID indicating the identified one or more main speakers to each photo.

[0107] In the example shown in 5-4 of Figure 5, the main speakers identified during the time period (time points t0 to t1) when photo P1 was played were two speakers: one with speaker ID "005" and the other with speaker ID "002," who were also main speakers during different time periods within that time period (time points t0 to t1).

[0108] In addition, the processor 11 may convert the voice data detected by the voice detector 20 into text data, and assign at least a portion of the text data as a comment tag to a corresponding photo among the multiple photos.

[0109] Furthermore, the processor 11 may simultaneously display at least one of the emotion rank and speaker information identified by the speaker ID while a plurality of photos are being played back.

[0110] [Second embodiment of emotion tagging method] FIG. 8 is a flowchart showing a second embodiment of the emotion tagging method according to the present invention.

[0111] The second embodiment shown in FIG. 8 is a method for attaching emotion tags when a photo movie viewing event is held.

[0112] A photo movie is made by combining multiple photos (n photos P1 to P2) using photo movie creation software. n ) are connected together in sequence. Photo movie creation software creates videos according to user settings such as the display time of each photo, the transition between photos (fade in, fade out, scroll, etc.), and other user settings.

[0113] 8, the processor 11 initializes t=0 and i=0 when starting playback of the photo movie (step S50). Note that t is a parameter indicating the playback time of the photo movie, and i is a parameter specifying the photo currently being displayed, which varies within the range of 1 to n in this example.

[0114] The processor 11 plays back the photo movie. i are displayed on the screen via the large display 30 (step S52). When playback of the photo movie starts, photo P1 is displayed.

[0115] Photo P i While the picture P is displayed, the playback time t is measured (step S54). i It is then determined whether the reproduction (display) of the photo P has ended (step S56). i If the playback time is within the set playback time (Photo P i If the playback of the photo P is not to be ended, the process returns to step S52. i playback will be performed.

[0116] On the other hand, the voice detector 20 detects voice data indicating the voices uttered by the event participants during the photo movie viewing event. The emotion recognizer 13A of the processor 11 detects the voices uttered by the event participants during the photo movie viewing event. i While the event is being displayed, voice data is acquired from the voice detector 20 (step S58), the emotions of the event participants (emotions indicating whether they are happy or not) are recognized based on the acquired voice data, and emotion information indicating the recognized emotions is sequentially output (step S60).

[0117] Photo P i When the playback time reaches the set playback time (Photo P i When it is determined that the reproduction of the photograph P has ended, the process proceeds to step S62, where the processor 11 (emotion rank calculation unit 13B) calculates the plurality of pieces of emotion information (photo P) recognized in step S60 and output sequentially. i The system obtains multiple pieces of emotional information during the display time period, and calculates an emotional rank indicating happiness from the representative value of the multiple pieces of emotional information.

[0118] The processor 11 (emotion tag assigning unit 13C) assigns the emotion rank calculated in step S62 to the emotion tag T i Photo as P i (Step S64). Here, the emotion tag T i Photo P to be awarded i can be determined by the playback time t of the photo movie. i (See Figure 5 for example.) Emotion tags can be assigned to each photo by recording the emotion tag in the header of the video file of the photo movie, or by creating a text file in which the emotion tag is written in association with the file name of the video file or the playback time of the photo movie.

[0119] Next, the parameter i is incremented by 1 (step S66), and it is determined whether the incremented parameter i exceeds n (i>n) (step S68).

[0120] If the parameter i does not exceed n (i≦n), the process returns to step S52, and the processes from step S52 to step S68 are repeated.

[0121] If the parameter i is greater than n (i>n), all photos P1 to P n The playback of the photo movie ends, and the process ends.

[0122] [others] The event using the content is not limited to the event of viewing the photo slideshow or photo movie of this embodiment, but may also be an event such as music playback or game play. In this case, a time-series emotion rank obtained by emotion analysis of the user's voice while playing music or playing a game can be assigned as a time-series emotion tag to the music playback, game play, etc.

[0123] Furthermore, the devices that make up the emotion tagging system (in this example, tablet PCs) can be devices that also serve as devices for holding events using content, or they can be separate devices. If they are separate devices, it is preferable that the two devices operate in sync so that the content to which emotion tags are to be assigned can be identified.

[0124] The various processors of the tablet PC and the like that make up the emotion tagging system according to the present invention include a CPU (Central Processing Unit), which is a general-purpose processor that executes programs and functions as various processing units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), which is a processor whose circuit configuration can be changed after manufacture, and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing specific processing.

[0125] A processing unit constituting the emotion tagging system may be composed of one of the various processors described above, or may be composed of two or more processors of the same or different types. For example, a processing unit may be composed of multiple FPGAs or a combination of a CPU and an FPGA. Furthermore, multiple processing units may be composed of a single processor. Examples of multiple processing units composed of a single processor include, first, a configuration in which a single processor is composed of a combination of one or more CPUs and software, as typified by computers such as client and server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are composed of one or more of the various processors described above as a hardware structure. Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit composed of a combination of circuit elements such as semiconductor devices.

[0126] The present invention also includes an emotion tagging program that, when installed in a computer, causes the computer to function as the emotion tagging system of the present invention, and a non-volatile storage medium on which this emotion tagging program is recorded.

[0127] Furthermore, the present invention is not limited to the above-described embodiment, and it goes without saying that various modifications are possible without departing from the spirit of the present invention. [Explanation of symbols]

[0128] 1. Emotion tagging system 11 processors 12 Memory 13A Emotion Recognizer 13B Emotion Rank Calculation Unit 13C Emotion tagging section 14 Display section 15A Speaker Recognizer 15B Speaker ID assignment unit 16 Input / Output Interfaces 17A Speech-to-text converter 17B Comment tagging section 18 Control section 20. Audio Detector 30 Large Display S10~S24, S50~S68 steps

Claims

1. a processor; a voice detector that detects voice data indicating voices uttered by people who participate in an event using the content while the event is being held; an emotion recognizer that recognizes human emotions based on the voice data, The processor: acquiring emotion information indicating the human emotion recognized by the emotion recognizer during the implementation of the event using the content; assigning an emotion rank calculated from the acquired emotion information to the content as an emotion tag; Emotion tagging system.

2. The emotion recognizer is a recognizer that has undergone machine learning using a large amount of speech data as training data, including speech data of speech uttered when a person is happy and speech data of speech uttered when a person is not happy. The emotion tagging system of claim 1 .

3. the content is a plurality of images; the event is a viewing event in which the plurality of images are sequentially played back by an image playback device and the played back images are viewed; The emotion tagging system of claim 1 or 2.

4. the plurality of images includes photographs or videos of people attending the event; The emotion tagging system of claim 3 .

5. the processor acquires from the emotion recognizer a plurality of pieces of emotional information for a time period during which the plurality of images are played back, calculates an emotional rank corresponding to each image from a representative value of the plurality of pieces of emotional information, and assigns the calculated emotional rank to each image as the emotion tag; 5. The emotion tagging system of claim 3 or 4.

6. The processor: If multiple people are participating in the event, one or more dominant speakers during the time period in which the multiple images are played are identified based on the audio data detected by the audio detector, and speaker identification information indicating the identified one or more dominant speakers is assigned to each image. The emotion tagging system of claim 5 .

7. The processor: simultaneously displaying at least one of the emotion rank and the speaker identification information while the image playback device is playing back the plurality of images; The emotion tagging system of claim 6.

8. The processor: converting the voice data detected by the voice detector into text data, and assigning at least a part of the text data as a comment tag to a corresponding image of the plurality of images; An emotion tagging system according to any one of claims 3 to 7.

9. a step of detecting, by a voice detector, voice data indicating voices uttered by people participating in an event while the event using the content is being held; an emotion recognizer recognizing human emotions based on the audio data; a processor acquiring emotion information indicating the recognized human emotion during an event using the content; a step in which the emotion recognizer assigns an emotion rank calculated from the acquired emotion information to the content as an emotion tag; An emotion tagging method comprising:

10. the content is a plurality of images; the event is a viewing event in which the plurality of images are sequentially played back by an image playback device and the played back images are viewed; The emotion tagging method of claim 9.

11. the plurality of images includes photographs or videos of people attending the event; The emotion tagging method of claim 10.

12. the processor acquires from the emotion recognizer a plurality of pieces of emotional information for a time period during which the plurality of images are played back, calculates an emotional rank corresponding to each image from a representative value of the plurality of pieces of emotional information, and assigns the calculated emotional rank to each image as the emotion tag; The emotion tagging method according to claim 10 or 11.

13. When multiple people are participating in the event, the processor identifies one or more dominant speakers during a time period when the multiple images are played back based on the audio data detected by the audio detector, and assigns speaker identification information indicating the identified one or more dominant speakers to each image. The emotion tagging method of claim 12.

14. the processor simultaneously displays at least one of the emotion rank and speaker information identified by the speaker identification information while the image playback device is playing back the plurality of images. The emotion tagging method of claim 13.

15. The processor converts the audio data detected by the audio detector into text data, and assigns at least a portion of the text data to a corresponding image of the plurality of images as a comment tag. The emotion tagging method according to any one of claims 10 to 14.

16. a function of acquiring, from a voice detector, voice data indicating voices uttered by people who have participated in an event using the content during the event; a function of recognizing human emotions based on the voice data; a function of acquiring emotion information indicating the recognized human emotion during an event using the content; a function of assigning an emotion rank calculated from the acquired emotion information to the content as an emotion tag; An emotion tagging program that uses a computer to achieve this.

Citation Information

Patent Citations

  • Image recording device, and image recording method

    JP2010057003A

  • Emotion estimation device creation method, emotion estimation device creation device, emotion estimation method, emotion estimation device and program

    JP2017111760A

  • Anti-enterovirus agent

    JP2021130617A

  • Client device, control method, system and program

    WO2014192457A1

  • Information processing system, information processing method, and information processing device

    WO2020158536A1