Device for tagging digital data, tagging method, program, and recording medium.

The digital data tagging device addresses the challenge of homonyms and synonyms in voice data by selecting relevant tag candidates based on thresholds, enabling accurate and efficient tag assignment and search functionality.

JP7860082B2Active Publication Date: 2026-05-15FUJIFILM CORP
View PDF 11 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FUJIFILM CORP
Filing Date
2022-03-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Conventional tagging devices struggle with distinguishing homonyms and synonyms in voice data, leading to difficulties in accurate tag assignment and effective search functionality.

Method used

A digital data tagging device that utilizes a processor to pre-store tag candidates, select relevant tag candidates based on thresholds for pronunciation and semantic similarity, and display options for user selection, allowing for easy assignment of desired tags to digital data.

Benefits of technology

Enables users to assign tags to digital data using voice data effectively, regardless of homophones or synonyms with different expressions, improving search accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007860082000001
    Figure 0007860082000001
  • Figure 0007860082000002
    Figure 0007860082000002
  • Figure 0007860082000003
    Figure 0007860082000003
Patent Text Reader

Abstract

The present invention enables a user to assign a desired tag easily through speech, regardless of homonyms and synonyms with different expressions. In a digital data tagging device, tagging method, program, and recording medium according to the present invention, a digital data acquisition unit acquires digital data to be tagged and a speech data acquisition unit acquires speech data related to the digital data. A phrase extraction unit extracts a phrase from the speech data, a tag candidate determination unit determines, as a first tag candidate, one or more tag candidates having a degree of association with the phrase equal to or greater than a first threshold value from among a plurality of tag candidates stored in advance in a tag candidate storage unit, and a tag assignment unit assigns at least one of the phrase or a tag candidate group including the first tag candidate to the digital data as a tag.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a tagging device, a tagging method, a program, and a recording medium for attaching tags to digital data.

Background Art

[0002] Conventionally, a tagging device that extracts phrases from voice data and attaches the phrases extracted from the voice data as tags has been known (see Patent Documents 1 to 3).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in conventional tagging devices that use voice data, there has been a problem that it is difficult to distinguish homonyms. For example, when the Japanese voice contains a phrase "ぞう", it is difficult to determine whether this "ぞう" means "elephant" (象) or "statue" (像). Similarly, in English voice, it is difficult to determine whether the spoken phrase means "ant" (蟻) or "aunt" (叔母). Furthermore, with conventional tagging devices, there are many synonyms with different expressions, and if these synonyms are directly assigned as tags, it becomes difficult to search using the tags. For example, synonyms for the Japanese word "sanpo" (meaning walk) include "osanpo," "burabura," and "sansaku." Therefore, when searching using "sanpo," "sanpo" and "osanpo" are found, but "burabura" and "sansaku" are not. Similarly, in English, synonyms for "walk" include "stroll" and "ramble." Therefore, when searching using "walk," "walk" and "walking" are found, but "stroll" and "ramble" are not.

[0005] The object of the present invention is to provide a digital data tagging device, tagging method, program, and recording medium that enable users to easily assign desired tags to digital data using voice, regardless of whether the words are homophones or synonyms with different expressions. [Means for solving the problem]

[0006] To achieve the above objective, the present invention comprises a processor and a tag candidate memory that pre-stores a plurality of tag candidates. The processor is We obtain digital data to which tags are to be assigned, Acquire audio data related to digital data, Extract words and phrases from audio data, From among multiple tag candidates, one or more tag candidates whose relevance to the word is above the first threshold are selected as the first tag candidate. The present invention provides a digital data tagging device that assigns at least one tag from a group of tag candidates, including words and first tag candidates, to digital data.

[0007] Here, equipped with a display, the processor is Convert audio data into text data, and extract one or more words from the text data. Display the text corresponding to the text data on the screen. Based on the word or phrase selected by the user from among the one or more words or phrases contained in the text displayed on the screen, the first tag candidate is determined. Display the tag candidates on the screen. It is preferable to assign at least one tag selected by the user from a group of tag candidates displayed on the screen to the digital data.

[0008] Furthermore, it is preferable for the processor to include the first synonym among the synonyms of the word or phrase, whose pronunciation similarity to the word or phrase is equal to or greater than the first threshold, as the first tag candidate.

[0009] Furthermore, it is preferable for the processor to include a second synonym among the synonyms of the phrase, whose semantic similarity to the phrase is equal to or greater than the first threshold, as a first tag candidate.

[0010] Furthermore, it is preferable that the processor includes both a first synonym, which has a similarity in pronunciation to the phrase that is equal to or greater than the first threshold, and a second synonym, which has a similarity in meaning to the phrase that is equal to or greater than the first threshold, as first tag candidates.

[0011] Furthermore, it is preferable for the processor to determine the number of first and second synonyms to include in the first tag candidate such that the number of first synonyms is greater than the number of second synonyms.

[0012] Furthermore, it is preferable for the processor to include homophones of words as first tag candidates.

[0013] Furthermore, it is preferable that the processor prioritizes displaying words or tag candidates previously selected by the user from the group of tag candidates over words or tag candidates that the user has not previously selected.

[0014] Also, the processor preferably causes the words or tag candidates that have been selected more frequently in the past among the words or tag candidates previously selected by the user to be displayed preferentially over the words or tag candidates that have been selected less frequently in the past.

[0015] Also, the digital data is image data, and the processor recognizes the subject included in the image corresponding to the image data, represents the name of the subject corresponding to the word, and determines a word different from the word as a second tag candidate, and preferably causes the second tag candidate to be included in the tag candidate group and displayed on the display.

[0016] Also, the digital data is image data, and the processor recognizes at least one of the subject and the scene included in the image corresponding to the image data, When there are a certain number or more of tag candidates having a relevance degree with the word of a first threshold value or more among a plurality of tag candidates, it is preferable to determine only the tag candidates having a relevance degree with at least one of the subject and the scene of a second threshold value or more as first tag candidates from among the tag candidates of the certain number or more.

[0017] Also, the digital data is image data, and the processor recognizes at least one of the subject and the scene included in the image corresponding to the image data, determines a tag candidate having a relevance degree with at least one of the subject and the scene of a second threshold value or more and a similarity degree with the pronunciation of the word of a third threshold value or more as a third tag candidate from among a plurality of tag candidates, and preferably causes the third tag candidate to be included in the tag candidate group and displayed on the display.

[0018] Also, the digital data is image data, and a person tag representing the name of the subject included in the image corresponding to the image data has been assigned by a first user to the image data, and the processor recognizes the subject included in the image, From voice data including voice in which a second user different from the first user speaks the name of a subject with respect to an image, extract the name of the subject, Determine one or more tag candidates whose degree of relevance to the name of the subject is equal to or higher than a first threshold as first tag candidates, and when the first tag candidates are different from the person tag, determine the person tag as a fourth tag candidate, Preferably, the fourth tag candidate is included in the tag candidate group and displayed on the display.

[0019] Also, the digital data is image data, and the processor Obtain information on the shooting position of the image corresponding to the image data, Based on the information on the shooting position of the image, among a plurality of tag candidates, determine a tag candidate representing a place name that is located within a range of not more than a fourth threshold from the shooting position of the image and whose similarity to the pronunciation of the phrase is equal to or higher than a third threshold as a fifth tag candidate, Preferably, the fifth tag candidate is included in the tag candidate group and displayed on the display.

[0020] Also, the digital data is image data, and the processor Recognize the subject included in the image corresponding to the image data, Obtain information on the shooting position of the image, Extract the name of the subject from voice data including the name of the subject included in the image, Based on the information on the shooting position of the image, when the name of the subject is different from the actual name of the subject located within a range of not more than a fourth threshold from the shooting position of the image, determine the actual name of the subject as a sixth tag candidate, Preferably, the sixth tag candidate is included in the tag candidate group and displayed on the display.

[0021] Also, the processor When the user selects the sixth tag candidate from a group of tag candidates, including the sixth tag candidate displayed on the screen, for each of the multiple image data corresponding to multiple images taken within a specified period, the actual name corresponding to the subject contained in each of the multiple images is determined as the seventh tag candidate. It is preferable to assign a corresponding seventh tag candidate as a tag to each of the multiple image data.

[0022] Also, the processor, Extract place names from audio data that includes place names. When a place name has multiple locations, a tag candidate consisting of a combination of the place name and each of the multiple locations is selected as the eighth tag candidate. It is preferable to include the eighth tag candidate among the group of tag candidates and display them on the screen.

[0023] Also, the processor, Extract at least one onomatopoeic word or mimetic word corresponding to the ambient sounds contained in the audio data. We decided on at least one of onomatopoeia and mimetic words as a candidate for the 9th tag. It is preferable to include the 9th tag candidate among the group of tag candidates and display them on the screen.

[0024] Furthermore, it is equipped with an audio data memory for storing audio data. The processor preferably stores audio data, which contains information relating to digital data, in the audio data memory.

[0025] Furthermore, digital data is video data, The processor preferably extracts words and phrases from the audio data contained in the video data.

[0026] Furthermore, the present invention includes the step of a digital data acquisition unit acquiring digital data to which tags are to be assigned, The audio data acquisition unit acquires audio data related to digital data, The word extraction unit performs the step of extracting words from the audio data, The tag candidate determination unit selects one or more tag candidates from among multiple tag candidates pre-stored in the tag candidate memory unit, the first tag candidate, whose degree of relevance to the word is equal to or greater than a first threshold, and The present invention provides a tagging method that includes the step of a tagging unit assigning at least one tag from a group of tag candidates, including a word or phrase and a first tag candidate, as a tag to digital data.

[0027] Furthermore, the present invention provides a program for causing a computer to perform each of the above-described steps of the tagging method.

[0028] Furthermore, the present invention provides a computer-readable recording medium on which a program for causing a computer to perform each step of the above-described tagging method is recorded. [Effects of the Invention]

[0029] In this invention, a phrase is extracted from audio data, and from a plurality of pre-stored tag candidates, a tag candidate with a high degree of relevance to the phrase is selected as the first tag candidate. At least one of the tag candidate group, including the phrase and the first tag candidate, is then assigned as a tag to the digital data. Therefore, according to this invention, regardless of homophones or synonyms with different expressions, a user can use audio data to assign a desired tag to digital data. [Brief explanation of the drawing]

[0030] [Figure 1] This is a block diagram of one embodiment showing the configuration of the tagging device of the present invention. [Figure 2] This is a flowchart illustrating one embodiment of the operation of a tagging device. [Figure 3] This is a conceptual diagram of one embodiment representing the tagging operation screen. [Figure 4] This is a conceptual diagram of one embodiment showing a state in which text corresponding to audio data is displayed. [Figure 5] This is a conceptual diagram of one embodiment representing a state in which words or phrases are selected from the text. [Figure 6] This is a conceptual diagram of one embodiment showing a state where the list of tags has been updated. [Figure 7] This is a conceptual diagram of one embodiment showing a state where a group of tag candidates is displayed. [Figure 8] This is a conceptual diagram of another embodiment showing a state where the list of tags has been updated. [Modes for carrying out the invention]

[0031] The digital data tagging apparatus, tagging method, program, and recording medium of the present invention will be described in detail below based on preferred embodiments shown in the attached drawings.

[0032] Figure 1 is a block diagram of one embodiment showing the configuration of the tagging device of the present invention. The tagging device 10 shown in Figure 1 is a device that assigns tags to digital data that are related to words and phrases contained in audio data, and comprises a digital data acquisition unit 12, an audio data acquisition unit 14, an audio data storage unit 16, a word and phrase extraction unit 18, a tag candidate storage unit 20, a tag candidate determination unit 22, a tag assignment unit 24, an image analysis unit 26, a location information acquisition unit 30, a display unit 32, a display control unit 34, and an instruction acquisition unit 36.

[0033] The digital data acquisition unit 12 is connected to the image analysis unit 26 and the location information acquisition unit 30, and the audio data acquisition unit 14 is connected to the phrase extraction unit 18. The tag candidate determination unit 22 is connected to the phrase extraction unit 18, the image analysis unit 26, the location information acquisition unit 30, the instruction acquisition unit 36, and the tag candidate storage unit 20. The tag assignment unit 24 is connected to the digital data acquisition unit 12, the tag candidate determination unit 22, and the instruction acquisition unit 36. The audio data storage unit 16 is connected to the audio data acquisition unit 14 and the tag assignment unit 24. The display unit 32 is connected to the display control unit 34, and the display control unit 34 is connected to the phrase extraction unit 18 and the tag candidate determination unit 22.

[0034] The digital data acquisition unit 12 acquires digital data to which tags will be assigned. Digital data can be anything that can be tagged, and is not particularly limited, but includes image data, video data, and text data, etc. The method for acquiring digital data is not particularly limited. The digital data acquisition unit 12 can acquire, for example, image data of images currently being taken by a smartphone camera or digital camera, or image data selected by the user from image data stored in a storage unit for previously taken image data (not shown). The same applies to video data and text data.

[0035] The audio data acquisition unit 14 acquires audio data related to the digital data acquired by the digital data acquisition unit 12. Audio data includes, but is not limited to, audio of a user speaking or conversing about digital data, ambient sounds at the time of the user's speech or conversation, etc. The audio data acquisition unit 14 can acquire one or more audio data from one digital data. One audio data may contain the voices of one or more users, and the two or more audio data may contain the voices of different users or the voices of the same user. The method for acquiring audio data is not particularly limited. For example, the audio data acquisition unit 14 can acquire audio by recording the voice of a user speaking or conversing with digital data using the voice recorder function of a smartphone or digital camera. Alternatively, it may acquire audio data selected by the user from audio data that has been recorded in the past and stored in the audio data storage unit 16.

[0036] The audio data storage unit (audio data memory) 16 stores the audio data acquired by the audio data acquisition unit 14. The audio data storage unit 16, for example, under the control of the tagging unit 24, associates digital data with audio data related to this digital data and stores audio data that has information about the association with the digital data.

[0037] The word extraction unit 18 extracts words from the audio data acquired by the audio data acquisition unit 14. The word extraction unit 18 can also extract words from the audio data stored in the audio data storage unit 16. The words extracted by the word extraction unit 18 (hereinafter also referred to as extracted words) can be assigned as tags to digital data, and may be words consisting of one or more characters (strings), or phrases such as "That was fun." The method for extracting words and phrases is not particularly limited, but the word and phrase extraction unit 18 can, for example, convert speech data into text data using speech recognition and extract one or more words and phrases from this text data.

[0038] The tag candidate storage unit (tag candidate memory) 20 is a database that pre-stores multiple tag candidates that are candidates for tags to be assigned to digital data. The words stored as tag candidates are not particularly limited, but for example, for a given word, synonyms and homonyms can be stored as tag candidates in association with that word. The tag candidate memory unit 20 stores synonyms such as "Furo" (in katakana), "Furo" (in kanji), "Ofuro" (in hiragana), a bath emoji, "pool," and "sento" (public bath) in association with "Ofuro" (in katakana), in the case of a Japanese environment. The tag candidate memory unit 20 also stores homophones such as "Zou" (in katakana), meaning "statue," in association with "Zou" (in katakana), meaning "elephant."

[0039] The tag candidate determination unit 22 selects one or more tag candidates from among the multiple tag candidates stored in the tag candidate storage unit 20 that have a degree of relevance to the extracted phrase of a first threshold or higher, including homophones and synonyms with different expressions, as the first tag candidate. In other words, it selects a tag candidate that has a higher degree of relevance to the extracted phrase than other tag candidates. The tag candidate determination unit 22 can determine as the first tag candidate not only the tag candidates associated with the extracted phrase, but also tag candidates whose degree of relevance to the extracted phrase is equal to or greater than the first threshold, from among the multiple tag candidates stored in the tag candidate storage unit 20. Furthermore, the tag candidate determination unit 22 can determine as the first tag candidate not only the tag candidates stored in the tag candidate storage unit 20, but also any phrase whose degree of relevance to the extracted phrase is equal to or greater than the first threshold. The specific method for determining potential tags will be explained later.

[0040] The tagging unit 24 assigns at least one tag from a group of tag candidates, including the extracted phrase and the first tag candidate determined by the tag candidate determination unit 22, as a tag to the digital data. The assigned tag is associated with the digital data and stored. The tag can be stored anywhere; if the digital data has an Exif (Exchangeable image file format) header area, that header area may be used as the tag storage location, or a dedicated storage area provided for tags in the tagging device 10 may be used as the tag storage location.

[0041] The image analysis unit 26 recognizes at least one of the subject and the scene contained in the image corresponding to the image data. The method for extracting a subject or scene from an image is not particularly limited, and various conventionally known methods can be used.

[0042] The location information acquisition unit 30 acquires information about the shooting location of the image corresponding to the image data when the digital data is image data. The method for obtaining information about the shooting location is not particularly limited. For example, images taken with a smartphone camera or digital camera are assigned Exif-format header information (image information). This header information includes information such as the date and time the image was taken and the shooting location. Therefore, the location information acquisition unit 30 can, for example, obtain information about the shooting location from the image header information.

[0043] The display control unit 34 controls the display by the display unit 32. That is, the display unit (display) 32 displays various types of information under the control of the display control unit 34. The display control unit 34 displays on the display unit 32 the operation screen for assigning tags to digital data, text corresponding to text data, a group of tag candidates, a list of tags assigned to the digital data, etc. The specific method for displaying tag suggestions will be explained later.

[0044] The instruction acquisition unit 36 ​​acquires various instructions input by the user. Instructions input by the user include, for example, an instruction to select an extraction phrase from one or more extraction phrases contained in the text displayed on the display unit 32 in order to display tag candidates, an instruction to select an extraction phrase or a first tag candidate from the group of tag candidates displayed on the display unit 32, and so on.

[0045] Next, the operation of the tagging device 10 will be explained with reference to the flowchart shown in Figure 2. In the following explanation, as an example, we will assume a case where tags are added to image data using an application for the tagging device 10 that runs on a smartphone.

[0046] When a user performs tagging, the display control unit 34 displays the tagging operation screen on the display unit 32, that is, on the smartphone's display screen.

[0047] On the tagging operation screen, the user first selects the image data to be tagged from the user's image data stored on their smartphone. For example, the user can select the image data to be tagged by tapping (pressing) a desired image from a list of images corresponding to the image data displayed on the smartphone screen.

[0048] In response, the digital data acquisition unit 12 acquires this image data (step S1), and the display control unit 34 displays the image corresponding to this image data on the tagging operation screen, as shown in Figure 3.

[0049] At the top of the tagging operation screen shown in Figure 3, an image (photo) 40 corresponding to the image data to be tagged is displayed, and below this image 40, the image's capture date and time information 42, "March 10, 2018, 8:56 PM," is displayed. In the center of the tagging operation screen, a list of tags 44, "2018" and "March," which were automatically assigned to the image data based on the image's capture date and time information 42, is displayed. At the bottom of the tagging operation screen, a text display area 46 for displaying text corresponding to the text data converted from the audio data is displayed, and within this text display area 46, an "OK" button 48 and an "End" button 50 are displayed. In the lower left of the tagging operation screen, an audio input button 52 is displayed.

[0050] Next, the user, while looking at the image 40 displayed on the tagging operation screen, presses the voice input button 52 to use the voice recorder function of their smartphone to record a voice message for the image 40, for example, saying "When he played in a bath" in Japanese (which means "When he played in a bath" in English).

[0051] In response, the voice data acquisition unit 14 acquires voice data of the voice spoken by the user (step S2). Next, the phrase extraction unit 18 converts this audio data into text data, for example. The phrase extraction unit 18 converts the audio data "When I played in the bath" into text data corresponding to the Japanese text "When I played in the bath". Next, the word extraction unit 18 extracts one or more words from the text data (step S3). For example, the word extraction unit 18 extracts three words from the text "When I played in the bath" corresponding to the text data: "bath", "play", and "when". Next, the display control unit 34 displays this text in the text display area 46 (step S4). The display control unit 34 displays these three words in the text 54, for example, as shown in Figure 4, enclosed in a border. This allows the user to know that the three words enclosed in this border are words that can be added to the image data as tags.

[0052] Next, the user selects a word or phrase to be assigned as a tag to the image data from among one or more words or phrases contained in the text 54 displayed in the text display area 46 (step S5). For example, the user selects "bath" from among "bath," "play," and "time."

[0053] In response, the display control unit 34 highlights the word or phrase selected by the user, as shown in Figure 5. The display control unit 34 highlights "bath" by, for example, changing its display color to a different color from the text's display color. If the text's display color is black, the display control unit 34 changes the display color of "bath" to yellow. From this state, if the user selects "play" or "time," the display color of "bath" returns to black, and the respective selected texts are changed to yellow. Also, pressing an area outside the selectable area returns to the state of step S4. Note that in Figure 5, instead of changing the display color of "bath," it is shown with a thick line. This allows the user to know that "bath" has been selected.

[0054] Next, the user can choose to press the "OK" button 48 on the tagging operation screen, press the selected word "bath" again, or press the "Finish" button 50 (step S6).

[0055] If the user presses the "OK" button 48 (selection 1 in step S6), the tagging unit 24 adds the selected phrase as a tag to the image data (step S7). Next, the display control unit 34 displays the phrase selected by the user in the tag list 44. That is, as shown in Figure 6, the display control unit 34 adds "bath" to the tag list 44 on the tagging operation screen and displays it. The display control unit 34 also changes the display color of the text 54 "bath" back to black. After that, the process returns to step S4. If you want to add another phrase as a tag, simply select the other phrase and press the "OK" button 48.

[0056] If the user presses the currently selected word "bath" again (selection 2 in step S6), the system enters tag candidate display mode. The tag candidate determination unit 22 then selects one or more tag candidates from among the multiple tag candidates stored in the tag candidate storage unit 20, based on the word selected by the user from among the one or more words contained in the text displayed in the text display area 46, and determines one or more tag candidates whose degree of relevance to the word is at or above the first threshold as the first tag candidate (step S8). For example, the tag candidate determination unit 22 selects the following tag candidates from among the multiple tag candidates stored in the tag candidate storage unit 20 as the first tag candidate, because their degree of relevance to "bath" is at or above the first threshold: "Furo" in katakana, "Bath" in kanji, and "Ofuro" in hiragana. Subsequently, the display control unit 34 displays a tag candidate group including the word and the first tag candidates (step S9). That is, as shown in FIG. 7, the display control unit 34 displays a window screen (popup screen) 56 of a tag candidate group including the katakana "FRO", the kanji "風呂", and the hiragana "おふろ" as the first tag candidates in addition to the extracted word "お風呂" in a form of a balloon from the extracted word "お風呂" so that it can be understood that they are the first tag candidates for the extracted word "お風呂", and overlays and displays it on the tagging operation screen.

[0057] In the example of FIG. 7, the window screen of the tag candidate group is displayed as one window including all of the extracted word "お風呂", the katakana "FRO", the kanji "風呂", and the hiragana "おふろ", but it is not limited thereto, and four independent windows each including one of these four words may be displayed. Also, the window screen of the tag candidate group may be displayed so as not to overlap with the text 54, the "OK" button 48, the "End" button 50, etc., or may be overlaid and displayed on the text 54, the "OK" button 48, the "End" button 50, etc.

[0058] Subsequently, the user selects at least one of the word and the first tag candidates as a tag from the tag candidate group displayed in the window screen 56 (step S10). For example, the user selects the kanji "風呂" from the katakana "FRO", the kanji "風呂", and the hiragana "おふろ".

[0059] In response to this, the tagging unit 24 assigns at least one selected by the user from the tag candidate group displayed in the window screen 56 to the image data as a tag (step S11). That is, the tagging unit 24 assigns the kanji "風呂" to the image data as a tag. Next, the display control unit 34 displays the phrase selected by the user in the tag list 44. That is, as shown in Figure 8, the display control unit 34 adds "bath" to the tag list 44 and displays it on the tagging operation screen. The display control unit 34 also changes the display color of the text 54 "bath" back to black on the tagging operation screen and hides the display of the tag candidate group window screen 56. After that, the process returns to step S4. If the user wants to add another phrase, for example, a first tag candidate related to "play", the user selects "play" and then selects "play" again. In response, the first tag candidates related to "play" are determined and displayed, and the user simply selects one of the displayed first tag candidates related to "play".

[0060] If the user presses the "Finish" button 50 (selection 3 in step S6), a message box will appear, for example, "Confirming tagging. The text currently displayed in the text area will be discarded. Is this OK?". If the user presses the "Do not finish" button that is simultaneously displayed in the message box, the user returns to the state before pressing the "Finish" button 50. On the other hand, if the user presses the "Finish" button that is simultaneously displayed in the message box, the tagging process ends (step S12), and the display control unit 34 removes the text display from the tagging operation screen. The "Finish" button 50 can also be pressed at any step other than step S6. This allows the user to return to the tagging operation screen shown in Figure 3. If no tag candidates are found, the tagging flow using the acquired audio data will be terminated, and the audio data will be acquired again to perform the tagging flow again.

[0061] The tagging device 10 uses voice data to assign tags, making it easy to assign tags to digital data, and even multiple tags can be easily assigned. Furthermore, since the tagging device 10 can use voice data of users speaking or conversing in spoken language, it can assign emotional tags such as "Much fun." Furthermore, the tagging device 10 extracts words from the audio data, selects a tag candidate with a high degree of relevance to the extracted words from a group of pre-stored tag candidates, and assigns at least one of the tag candidate group, including the words and the first tag candidate, as a tag to the digital data. Therefore, the tagging device 10 allows users to assign desired tags to digital data using audio, regardless of whether the words are homophones or synonyms with different expressions.

[0062] Next, we will explain the method for determining and displaying tag candidates, using specific examples.

[0063] For example, synonyms with a high degree of similarity in pronunciation to the extracted phrase may be used as first tag candidates. That is, the tag candidate determination unit 22 may include first synonyms among the synonyms of the extracted phrase whose degree of similarity in pronunciation to the extracted phrase is equal to or greater than a first threshold as first tag candidates. For example, if the word "bath" is extracted from the audio data, the tag candidate determination unit 22 may include, for example, the katakana "Furo", the kanji "Furo", and the hiragana "Ofuro" as first tag candidates, among the synonyms of "bath" that have a high degree of similarity in pronunciation to "bath".

[0064] Furthermore, synonyms with a high degree of semantic similarity to the extracted phrase may be used as first tag candidates. That is, the tag candidate determination unit 22 may include second synonyms among the synonyms of the extracted phrase whose semantic similarity to the extracted phrase is equal to or greater than a first threshold as first tag candidates. Similarly, if the tag candidate determination unit 22 extracts the phrase "bath" from the audio data, it can include "bathroom," "bath," "Bath," and the bathtub emoji as first tag candidates, which are synonyms for "bath" and have a high degree of semantic similarity to "bath."

[0065] Furthermore, both the aforementioned first synonym and second synonym may be used as first tag candidates. That is, the tag candidate determination unit 22 may include both the first synonym, which has a similarity in pronunciation to the extracted phrase of at least one threshold, and the second synonym, which has a similarity in meaning to the extracted phrase of at least one threshold, as first tag candidates. Similarly, if the tag candidate determination unit 22 extracts the phrase "bath" from the audio data, it can include the following as first tag candidates: "Furo" in katakana, "Furo" in kanji, "Ofuro" in hiragana, "Bathroom", "Bath", "Bath", and a bathtub emoji.

[0066] Furthermore, when the tag candidate determination unit 22 uses both the first synonym and the second synonym as first tag candidates, it is desirable to determine the number of first synonyms and second synonyms to include as first tag candidates such that the number of first synonyms with high similarity in pronunciation is greater than the number of second synonyms with high similarity in meaning. Similarly, if the tag candidate determination unit 22 extracts the phrase "bath" from the audio data, it may include, for example, the extracted phrase "bath," the first synonyms "Furo" in katakana and "Furo" in kanji, and the synonym "bathroom" as first tag candidates.

[0067] Furthermore, the tag candidate determination unit 22 may use a tag candidate that is a homophone of the extracted phrase as the first tag candidate. For example, in Japanese, it is known that there are two words "kaki": the fruit "kaki" and the seafood "oyster." Therefore, the tag candidate memory unit 20 can be pre-programmed to store two tag candidates, "kaki" and "oyster," in relation to the utterance "kaki." If the tag candidate determination unit 22 extracts the word "persimmon" from audio data containing the speech "kaki, oishii!" (which in English would be "kaki is delicious!"), it can include "oyster," a homophone of "kaki," as the first tag candidate. Similarly, in the case of English speech, if the audio data can be interpreted as either "The hare is beautiful" or "The hair is beautiful," both "hare" and "hair" can be included as the first tag candidates. Furthermore, the tag candidate determination unit 22 may simultaneously use three first tag candidates: a first synonym, a second synonym, and a homophone.

[0068] Extracted keywords or tag candidates previously selected by the user are more likely to be the user's preferred keywords or tag candidates than extracted keywords or tag candidates that have not been previously selected. Accordingly, the display control unit 34 may, from among the tag candidate group, display extracted terms or tag candidates previously selected by the user for the extracted term, with priority given to those selected by the user for the same extracted term, over extracted terms or tag candidates that have not been previously selected by the user. Furthermore, the display control unit 34 may, among the extracted terms or tag candidates previously selected by the user for the extracted term, display extracted terms or tag candidates that have been selected more frequently in the past for the same extracted term, with priority given to those selected less frequently in the past. This prioritizes the display of extraction terms or tag candidates that are likely to be the user's preference, thereby improving the user's convenience when selecting extraction terms or tag candidates from a group of tag candidates.

[0069] If the digital data is image data, you may use a phrase that represents the name of the subject contained in the image corresponding to the image data as a tag candidate. In this case, the image analysis unit 26 recognizes the subject contained in the image corresponding to the image data. Next, the tag candidate determination unit 22 determines a second tag candidate that represents the name of the subject corresponding to the extracted phrase, and is different from the extracted phrase. The display control unit 34 then displays the second tag candidate, along with the tag candidate group, on the display unit 32.

[0070] For example, suppose an image of a baby playing in a plastic pool is captured, and audio data of the mother saying "It was much fun like a bath." is obtained, and the word "bath" is extracted from this audio data. In this case, when the mother presses "bathtub" twice in a row, the image analysis unit 26 enters tag candidate display mode and recognizes that the subject in the image is a "plastic pool." The tag candidate determination unit 22 determines that the extracted phrases "bath" and "vinyl pool" are different, and therefore selects "vinyl pool" as the second tag candidate. Then, the display control unit 34 displays "inflatable pool" in addition to "bathtub" among the tag candidate group. This allows the system to use the correct subject's name as a tag suggestion even if the user has misidentified the subject in the image, or if they are using a figurative expression and the subject being referred to is different from the correct subject.

[0071] While it is acceptable to display the second tag candidate alongside the first tag candidate, it is preferable to associate "vinyl pool" with "bathtub" since "vinyl pool" is the correct term for "bathtub." For example, if multiple first tag candidates are displayed vertically, the second tag candidate "vinyl pool" should be displayed horizontally alongside the first tag candidate "bathtub."

[0072] Furthermore, if the digital data is image data, the number of first tag candidates may be limited based on at least one of the subjects and scenes included in the image corresponding to the image data. In this case, the image analysis unit 26 recognizes at least one of the subject and the scene included in the image corresponding to the image data. Then, if there is a predetermined number of tag candidates among the multiple tag candidates stored in the tag candidate storage unit 20 whose degree of relevance to the extracted phrase is at or above the first threshold, the tag candidate determination unit 22 determines from this predetermined number of tag candidates only those whose degree of relevance to at least one of the subject and / or scene is at or above the second threshold as the first tag candidate.

[0073] For example, if the tag candidate determination unit 22 has 10 tag candidates that are highly related to "bath," it will select only 5 of these 10 tag candidates that are highly related to the "baby" shown in the image as the first tag candidates. This allows the number of tag candidates to be limited, even when there are many tag candidates with a high degree of relevance to the extracted phrase, preventing the display of a large number of primary tag candidates exceeding a predetermined limit.

[0074] If the digital data is image data, then based on at least one of the subjects and scenes contained in the image data, words with a high similarity to the pronunciation of the extracted words may be used as tag candidates. In this case, the image analysis unit 26 recognizes at least one of the subject and scene included in the image corresponding to the image data, The tag candidate determination unit 22 selects a tag candidate as the third tag candidate from among the multiple tag candidates stored in the tag candidate storage unit 20, such that the degree of relevance to at least one of the subject and the scene is at or above the second threshold, and the degree of similarity to the pronunciation of the extracted phrase is at or above the third threshold. The display control unit 34 then displays the third tag candidate, including it in the group of tag candidates, on the display unit 32.

[0075] For example, suppose an image showing a large red lantern at Kaminarimon is captured, and the user says "Now in Akasaka!" (Now in Akasaka!), and the word "Akasaka" is extracted from this audio data. In this case, the image analysis unit 26 recognizes that the subject included in the image is the "red lantern of Kaminarimon," a famous landmark in Asakusa. Next, the tag candidate determination unit 22 selects "Asakusa" as the second tag candidate because it has a high degree of relevance to "Kaminarimon's red lantern" and a high degree of similarity in pronunciation to "Akasaka". Then, the display control unit 34 displays "Asakusa" in addition to "Akasaka" among the tag candidate group. This allows users to select the desired tag candidate that best matches their intention from among "Akasaka" and "Asakusa," even if they mistakenly say "Akasaka" instead of "Asakusa" or if speech recognition misinterprets "Asakusa" as "Akasaka." The same applies to English. For example, suppose an image of the Reunion Tower in Dallas is captured, and the user says "Now in Dulles!", and the word "Dulles" is extracted from this audio data. In this case, the image analysis unit 26 recognizes that the subject in the image is the "Reunion Tower," a famous landmark in Dallas. Next, the tag candidate determination unit 22 selects "Dallas" as the second tag candidate because it has a high degree of relevance to "Reunion Tower" and a high degree of similarity in pronunciation to "Dulles". Then, the display control unit 34 displays "Dallas" in addition to "Dulles" among the tag candidate group. This allows users to select the desired tag candidate that best matches their intent from both "Dulles" and "Dallas," even if they mistakenly pronounce "Dallas" as "Dulles" or if speech recognition misinterprets "Dallas" as "Dulles."

[0076] If the digital data is image data, and a person tag representing the name of the subject contained in the image has already been assigned to the image corresponding to the image data by the first user, then the name of the subject, which may be referred to differently by different speakers, may be used as a tag candidate. In this case, the image analysis unit 26 recognizes the subject contained in the image. Next, the word extraction unit 18 extracts the name of the subject from the audio data, which includes audio in which a second user, different from the first user, speaks the name of the subject in relation to this image. Next, the tag candidate determination unit 22 determines one or more tag candidates whose degree of relevance to the subject's name is equal to or greater than the first threshold as the first tag candidate, and if the first tag candidate is different from the person tag assigned to the image, it determines this person tag as the fourth tag candidate. The display control unit 34 then displays the fourth tag candidate, including it in the group of tag candidates, on the display unit 32.

[0077] For example, suppose a user typically tags images of their mother with the person tag "mother". On the other hand, suppose that in response to an image of the user's mother, the word "grandma" is extracted from audio data of the user's child saying, "Grandma, come to play again!" In this case, the image analysis unit 26 recognizes that the subject in the image is "Mother" because the image data has been tagged with the person "Mother". Next, the tag candidate determination unit 22 determines the phrase "grandmother" as the first tag candidate, and since this "grandmother" is different from "mother," it determines "mother" as the fourth tag candidate. Then, the display control unit 34 displays "mother" in addition to "grandmother" among the tag candidate group. In some countries, such as Japan, it is customary to address people not by their first names, but by their relationship within the family. Therefore, the same person may be called "Okaasan" (mother) by her daughter and "Obaachan" (grandmother) by her grandchild. In other words, the same person may be referred to by different words. However, according to this embodiment, even if the speaker refers to the subject differently, the user can select their desired tag candidate from "Obaachan" and "Okaasan".

[0078] If the digital data is image data, place names with a high degree of similarity to the pronunciation of the extracted phrase may be used as tag candidates based on information about the location where the image data was taken. In this case, the location information acquisition unit 30 acquires information about the shooting location of the image corresponding to the image data. Next, the tag candidate determination unit 22, based on the information of the image's shooting location, selects a tag candidate representing a place name from among the multiple tag candidates stored in the tag candidate storage unit 20, which is located within a range of 4th threshold or less from the image's shooting location and has a similarity to the pronunciation of the extracted phrase of 3rd threshold or more, as the 5th tag candidate. The display control unit 34 then displays the fifth tag candidate, including it in the group of tag candidates, on the display unit 32.

[0079] For example, suppose the word "Akasaka" is extracted from audio data containing the utterance "Akasaka," but information about the image's shooting location reveals that "Asakusa," not "Akasaka," was actually located near the image's shooting location. In this case, the tag candidate determination unit 22 determines "Asakusa" as the fifth tag candidate because it is close to the image's shooting location and has a high degree of similarity in pronunciation to "Akasaka". Then, the display control unit 34 displays "Asakusa" in addition to "Akasaka" among the tag candidate group. This allows users to select their desired tag from among "Akasaka" and "Asakusa," even if they mistakenly say "Akasaka" instead of "Asakusa" or if speech recognition misinterprets "Asakusa" as "Akasaka." The same applies to English. For example, suppose the word "Dulles" is extracted from audio data containing the utterance "Dulles," but based on the image's location information, there was actually "Dallas" in the vicinity of the image's location, not "Dulles." In this case, the tag candidate determination unit 22 determines "Dallas" as the fifth tag candidate because it is close to the image's shooting location and has a high degree of similarity in pronunciation to "Dulles". Then, the display control unit 34 displays "Dallas" in addition to "Dulles" among the tag candidate group. This allows the user to select their preferred tag from both "Dulles" and "Dallas," even if they mistakenly pronounce "Dallas" as "Dulles" or if speech recognition incorrectly identifies "Dallas" as "Dulles."

[0080] If the digital data is image data, the names of the subjects contained in the corresponding images may be used as tag candidates. In this case, the image analysis unit 26 recognizes the subject contained in the image corresponding to the image data, and the location information acquisition unit 30 acquires information about the shooting location of this image. Next, the word extraction unit 18 extracts the name of the subject from the audio data that includes the name of the subject in the image. The tag candidate determination unit 22, based on the information of the image's shooting location, determines the actual name of a subject as the sixth tag candidate if the name of the subject differs from the actual name of a subject located within a range of the fourth threshold or less from the image's shooting location. The display control unit 34 then displays the sixth tag candidate in the tag candidate group on the display unit 32.

[0081] For example, suppose an image of a theme park attraction is taken, and audio data containing the utterance "Now at “Star Travel!”" is used to extract the phrase "Star Travel." However, based on the location information of the image, it is determined that the attraction is actually "Space Fantasy," not "Star Travel." In this case, the tag candidate determination unit 22 determines that "Space Fantasy" is the fifth tag candidate because it is different from "Start Label" and "Space Fantasy" is located near the image's shooting location. The display control unit 34 then displays "Space Fantasy" in addition to "Start Label" among the tag candidate group. This allows users to select their desired tag from both "Star Travel" and "Space Fantasy," even if they mistakenly say "Star Travel" instead of "Space Fantasy."

[0082] Furthermore, if there are multiple images, the actual names of the subjects contained in each image may be automatically assigned as tags, similar to the method described above. In other words, when the user selects the sixth tag candidate from the group of tag candidates, including the sixth tag candidate displayed on the display unit 32, for one image data, the tag candidate determination unit 22 determines the actual name corresponding to the subject contained in each of the multiple image data corresponding to multiple images taken within a specified period as the seventh tag candidate. The tagging unit 24 then assigns a seventh tag candidate, corresponding to each of the multiple image data, as a tag to each of the multiple image data.

[0083] If the extracted phrase is a place name and there are multiple locations for that place name, the place name including the location may be used as a tag candidate. In other words, the word extraction unit 18 extracts place names from audio data that includes place names. The tag candidate determination unit 22 determines, as the eighth tag candidate, multiple tag candidates consisting of combinations of the place name and each of the multiple locations, when there are multiple locations for the place name. The display control unit 34 then displays the eighth tag candidate, including it in the group of tag candidates, on the display unit 32.

[0084] For example, if "Otemachi" is extracted from audio data containing the sound "Otemachi", the tag candidate determination unit 22 will determine "Otemachi (Tokyo)" and "Otemachi (Ehime)" as the eighth tag candidate. The display control unit 34 then displays "Otemachi (Tokyo)" and "Otemachi (Ehime)" in addition to "Otemachi" among the tag candidate group. This allows users to select their desired tag information from either "Otemachi" in Tokyo or "Otemachi" in Ehime.

[0085] For example, for a user residing in Tokyo, the display "Otemachi (Tokyo)" may be redundant. Conversely, if the user's residency in Tokyo is registered in advance, the display may simply show "Otemachi" instead of "Otemachi (Tokyo)". Furthermore, if it is desired to distinguish locations within the tagging device 10, "Otemachi (Tokyo)" and "Otemachi (Ehime)" may be stored separately. Alternatively, both "Otemachi (Tokyo)" and "Otemachi (Ehime)" may be displayed, and when one of them is selected as a tag by the user, the location display may be removed, and only "Otemachi" may be assigned as a tag to the image data.

[0086] In addition to the audio contained in the audio data, onomatopoeia corresponding to ambient sounds, such as onomatopoeia and mimetic words, may also be used as tag candidates. In this case, the word extraction unit 18 extracts at least one of onomatopoeic words and mimetic words that correspond to the ambient sounds contained in the audio data. Next, the tag candidate determination unit 22 determines at least one of the onomatopoeic words and mimetic words as the ninth tag candidate. The display control unit 34 then displays the 9th tag candidate, including it in the group of tag candidates, on the display unit 32.

[0087] For example, suppose the phrase "zaa-zaa," which is an onomatopoeic representation of rain sounds, is extracted from audio data that includes the sound of rain. In this case, the tag candidate determination unit 22 determines "Zaa-zaa" as the ninth tag candidate. In addition to "Zaa-zaa," the tag candidate determination unit 22 may also use "Ame" (rain) as a tag candidate. Then, the display control unit 34 displays "Zzzzz" among the tag candidate group. This allows users to easily tag image data with onomatopoeic words corresponding to ambient sounds.

[0088] If, for example, a user's voice uttering something in response to an image is captured, this voice data could represent one of the memories associated with the moment the image was taken. This applies not only to images but to all digital data. Accordingly, the tagging unit 24 may associate the digital data with the audio data relating to the digital data and store the audio data having association information with the digital data in the audio data storage unit 16. This allows users, for example, to play and listen to audio data associated with the image data corresponding to an image when they view that image.

[0089] Video data often includes audio data. Accordingly, if the digital data is video data, the audio data acquisition unit 14 may acquire audio data from the video data, and the phrase extraction unit 18 may extract phrases from the audio data acquired from the video data. In this case, the user can use extracted keywords automatically obtained from the audio data contained in the video data to tag the image data.

[0090] In the apparatus of the present invention, the hardware configuration of the processing unit (Processing Unit) that performs various processes such as the digital data acquisition unit 12, the audio data acquisition unit 14, the phrase extraction unit 18, the tag candidate determination unit 22, the tag assignment unit 24, the image analysis unit 26, the location information acquisition unit 30, the display control unit 34, and the instruction acquisition unit 36 ​​may be dedicated hardware or may be various processors or computers that execute programs. Furthermore, the audio data storage unit 16 and the tag candidate storage unit 20 may be configured with memory such as semiconductor memory, HDD (Hard Disk Drive), or SSD (Solid State Drive).

[0091] Various types of processors include CPUs (Central Processing Units), which are general-purpose processors that execute software (programs) and function as various processing units; Programmable Logic Devices (PLDs), such as FPGAs (Field Programmable Gate Arrays), which are processors whose circuit configuration can be changed after manufacturing; and Dedicated Electrical Circuits, such as ASICs (Application Specific Integrated Circuits), which are processors with circuit configurations specifically designed for performing particular processing.

[0092] A single processing unit may be composed of one of these various processors, or it may be composed of a combination of two or more processors of the same or different types, such as a combination of multiple FPGAs, or a combination of an FPGA and a CPU. Furthermore, multiple processing units may be composed of one of these various processors, or two or more of the multiple processing units may be combined and composed of a single processor.

[0093] For example, as exemplified by servers and client computers, one processor is composed of a combination of one or more CPUs and software, and this processor functions as multiple processing units. Alternatively, as exemplified by System on Chip (SoC), a processor is used that realizes the functions of the entire system, including multiple processing units, on a single IC (Integrated Circuit) chip.

[0094] Furthermore, the hardware configuration of these various processors is, more specifically, an electrical circuit (Circuitry) made up of circuit elements such as semiconductor devices.

[0095] Furthermore, the method of the present invention can be implemented, for example, by a program that causes a computer to execute each of its steps. A computer-readable recording medium on which this program is recorded can also be provided.

[0096] Although the present invention has been described in detail above, the present invention is not limited to the embodiments described above, and various improvements and modifications may be made without departing from the spirit of the present invention. [Explanation of Symbols]

[0097] 10 Tagging device 12 Digital Data Acquisition Unit 14. Audio data acquisition unit 16. Audio data storage unit (audio data memory) 18. Word / phrase extraction section 20 Tag candidate memory (Tag candidate memory) 22 Tag Candidate Selection Section 24. Tagging section 26 Image Analysis Department 30 Location information acquisition unit 32 Display Unit 34 Display Control Unit 36 Instruction acquisition part 40 images 42 Information on the date and time of shooting List of 44 tags 46 Text display area 48. OK button 50 "End" button 52 Voice input button 54 Text 56. Window screen (popup screen)

Claims

1. It comprises a processor, a display, and a tag candidate memory that pre-stores multiple tag candidates. The aforementioned processor, We obtain digital data to which tags are to be assigned, Acquire audio data related to the aforementioned digital data, Convert the aforementioned audio data into text data, and extract one or more words from the text data. The text corresponding to the aforementioned text data is displayed on the display. From among the multiple tag candidates, one or more tag candidates whose degree of relevance to the phrase selected by the user from among the one or more phrases contained in the text displayed on the display is equal to or greater than a first threshold are determined as the first tag candidate. A device for tagging digital data, which assigns at least one of a group of tag candidates, including the phrase and the first tag candidate, as the tag to the digital data.

2. The processor is The group of tag candidates is displayed on the display, A digital data tagging device according to claim 1, which assigns to the digital data at least one selected by the user from the group of tag candidates displayed on the display as the tag.

3. The digital data tagging device according to claim 2, wherein the processor includes a first synonym among the synonyms of the phrase whose similarity in pronunciation to the phrase is equal to or greater than the first threshold as a first tag candidate.

4. The digital data tagging device according to claim 2 or 3, wherein the processor includes a second synonym among the synonyms of the phrase whose semantic similarity to the phrase is equal to or greater than the first threshold as a candidate for the first tag.

5. The digital data tagging device according to claim 2, wherein the processor includes both a first synonym among the synonyms of the phrase whose pronunciation similarity to the phrase is equal to or greater than the first threshold, and a second synonym whose semantic similarity to the phrase is equal to or greater than the first threshold, as the first tag candidate.

6. The digital data tagging device according to claim 5, wherein the processor determines the number of first synonyms and second synonyms to be included in the first tag candidate such that the number of first synonyms is greater than the number of second synonyms.

7. The digital data tagging device according to any one of claims 2 to 6, wherein the processor includes homophones of the phrase as first tag candidates.

8. The digital data tagging device according to any one of claims 2 to 7, wherein the processor displays, from the group of tag candidates, words or tag candidates previously selected by the user in the past, with priority given to words or tag candidates not previously selected by the user.

9. The digital data tagging device according to claim 8, wherein the processor displays, with priority given to, the words or tag candidates that have been selected more frequently in the past than the words or tag candidates that have been selected less frequently in the past.

10. The aforementioned digital data is image data, and the processor is Recognizing the subject contained in the image corresponding to the aforementioned image data, A second tag candidate is determined to represent the name of the subject corresponding to the aforementioned phrase, and to be different from the aforementioned phrase. A digital data tagging device according to any one of claims 2 to 9, wherein the second tag candidate is included in the group of tag candidates and displayed on the display.

11. The aforementioned digital data is image data, and the processor is Recognizing at least one of the subject and scene included in the image corresponding to the aforementioned image data, A digital data tagging device according to any one of claims 2 to 9, wherein, among the plurality of tag candidates, if there is a predetermined number or more tag candidates whose degree of relevance to the phrase is equal to or greater than the first threshold, only the tag candidates whose degree of relevance to at least one of the subject and the scene is equal to or greater than the second threshold are selected from among the predetermined number or more tag candidates as the first tag candidate.

12. The aforementioned digital data is image data, and the processor is Recognizing at least one of the subject and scene included in the image corresponding to the aforementioned image data, From among the multiple tag candidates, a tag candidate is selected as the third tag candidate if its relevance to at least one of the subject and the scene is at or above the second threshold, and its similarity to the pronunciation of the word is at or above the third threshold. A digital data tagging device according to any one of claims 2 to 11, wherein the third tag candidate is included in the group of tag candidates and displayed on the display.

13. The aforementioned digital data is image data, and a person tag representing the name of the subject included in the image corresponding to the image data has been assigned to the image data by a first user, and the processor is Recognizing the subject included in the aforementioned image, With respect to the aforementioned image, the name of the subject is extracted from audio data that includes audio in which a second user, different from the first user, speaks the name of the subject. One or more tag candidates whose degree of relevance to the subject's name is equal to or greater than the first threshold are selected as the first tag candidate, and if the first tag candidate and the person tag are different, the person tag is selected as the fourth tag candidate. A digital data tagging device according to any one of claims 2 to 12, wherein the fourth tag candidate is included in the group of tag candidates and displayed on the display.

14. The aforementioned digital data is image data, and the processor is The information of the shooting location of the image corresponding to the aforementioned image data is obtained, Based on the information of the image's shooting location, a fifth tag candidate is selected from among the multiple tag candidates that is located within a range of the image's shooting location and less than or equal to the fourth threshold, and whose similarity to the pronunciation of the phrase is greater than or equal to the third threshold, representing a place name. A digital data tagging device according to any one of claims 2 to 13, wherein the fifth tag candidate is included in the group of tag candidates and displayed on the display.

15. The aforementioned digital data is image data, and the processor is Recognizing the subject contained in the image corresponding to the aforementioned image data, The information of the shooting location of the aforementioned image is obtained, The name of the subject is extracted from the audio data containing the name of the subject in the aforementioned image. Based on the information of the image's shooting location, if the name of the subject differs from the actual name of the subject located within a range of the fourth threshold from the image's shooting location, the actual name of the subject is determined as a sixth tag candidate. A digital data tagging device according to any one of claims 2 to 14, wherein the sixth tag candidate is included in the group of tag candidates and displayed on the display.

16. The aforementioned processor, When the user selects the sixth tag candidate from the group of tag candidates, including the sixth tag candidate displayed on the display, for each of the multiple image data corresponding to multiple images taken within a specified period, the actual name corresponding to the subject contained in each of the multiple images is determined as the seventh tag candidate. A digital data tagging device according to claim 15, wherein for each of the plurality of image data, the seventh tag candidate corresponding to each of the plurality of image data is assigned as the tag.

17. The aforementioned processor, Extract the place name from the audio data containing the place name, If there are multiple locations for the aforementioned place name, a tag candidate consisting of a combination of the aforementioned place name and each of the multiple locations is determined as the eighth tag candidate. A digital data tagging device according to any one of claims 2 to 16, wherein the group of tag candidates includes the eighth tag candidate and displays it on the display.

18. The aforementioned processor, From the audio data, at least one of onomatopoeic words and mimetic words corresponding to the ambient sounds contained in the audio data is extracted. At least one of the aforementioned onomatopoeia and the aforementioned sound-like word is selected as a candidate for the ninth tag. A digital data tagging device according to any one of claims 2 to 17, wherein the 9th tag candidate is included in the group of tag candidates and displayed on the display.

19. The system includes an audio data memory for storing the aforementioned audio data, The digital data tagging device according to any one of claims 1 to 18, wherein the processor stores the audio data having association information with the digital data in the audio data memory.

20. The aforementioned digital data is video data, The digital data tagging device according to any one of claims 1 to 19, wherein the processor extracts the phrase from the audio data contained in the video data.

21. The digital data acquisition unit acquires digital data to which tags are to be assigned, The steps include: the audio data acquisition unit acquires audio data relating to the digital data; The word extraction unit performs the steps of converting the audio data into text data and extracting one or more words from the text data, The display control unit causes the text corresponding to the text data to be displayed on the display, The tag candidate determination unit determines, from among a plurality of tag candidates pre-stored in the tag candidate storage unit, one or more tag candidates whose degree of relevance to the phrase selected by the user from among the one or more phrases contained in the text displayed on the display is at or above a first threshold, as the first tag candidate. A method for tagging digital data, comprising the step of a tagging unit assigning at least one of a group of tag candidates, including the word and the first tag candidate, to the digital data as the tag.

22. A program for causing a computer to perform each step of the method for tagging digital data described in claim 21.

23. A computer-readable recording medium on which a program for causing a computer to perform each step of the method for tagging digital data according to claim 21 is recorded.