Method for identifying encountered words, program for identifying encountered words, and device for identifying encountered words

JP7925211B1Active Publication Date: 2026-09-25藤坂 勇斗
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2026028741
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2026-02-25
Publication Date
2026-09-25
Estimated Expiration
2046-02-25

AI Technical Summary

Benefits of technology

【0007】 インターネット上で利用者が遭遇した単語を正確に特定することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007925211000001_ABST
    Figure 0007925211000001_ABST
Patent Text Reader

Abstract

To accurately record words that users encounter on the internet. [Solution] The method for identifying words encountered by a user includes the steps of: obtaining candidate words that may be presented to the user from content obtained from multiple types of information sources using at least one of subtitle acquisition, character recognition, speech recognition, and document structure analysis; obtaining individual confidence scores for each of the obtained candidates for each of the multiple methods; obtaining an integrated confidence score for each of the obtained candidates by combining the multiple individual confidence scores; identifying a candidate that has been presented to the user from among the candidates using the integrated confidence score; obtaining the user's level of interest in the identified candidate based on multiple indicators; and identifying a word encountered by the user from among the identified candidates using the level of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an encountered word identification method, an encountered word identification program, and an encountered word identification apparatus. [Background Art]

[0002] In recent years, when learning words such as English words, wordbook applications that operate on terminals such as personal computers and smartphones, instead of paper wordbooks, have been widely used. Some wordbook applications are configured to allow users to create their own wordbook data in which words to be learned are registered.

[0003] For example, Patent Document 1 discloses a technique that enables generation of a wordbook suitable for a user with a simple configuration. [Prior Art Documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Unexamined Patent Publication No. 2020-160382 [Summary of the Invention] [Problem to be Solved by the Invention]

[0005] Here, when creating wordbook data in which words to be learned are registered, it is preferable in terms of both learning motivation and the large number of opportunities to utilize learning results to create wordbook data that allows users to preferentially learn words that are often used in fields that interest the user and that the user often encounters. Accordingly, recording words that a user encounters when accessing various contents on the Internet in daily life is considered useful for creating suitable wordbook data. However, since contents on the Internet are diverse, it has been difficult to accurately identify words encountered by a user. [Means for Solving the Problem]

[0006] In one aspect of this disclosure, an encounter word identification method for identifying words encountered by a user includes the steps of: an information processing device obtaining candidate words that may be presented to the user from content obtained from multiple types of information sources using at least one of subtitle acquisition, character recognition, speech recognition, and document structure analysis; obtaining individual confidence scores for each of the obtained candidates for each of the multiple methods; obtaining an integrated confidence score for each of the obtained candidates by combining the multiple individual confidence scores; identifying a candidate presented to the user from among the candidates using the integrated confidence score; obtaining the user's level of interest in the identified candidate based on multiple indicators; and identifying a word encountered by the user from among the identified candidates using the level of interest. [Effects of the Invention]

[0007] It can accurately identify words that users encounter on the internet. [Brief explanation of the drawing]

[0008] The drawings illustrate specific embodiments of the present invention relating to this disclosure, including not only essential components of the invention but also optional and preferred embodiments. [Figure 1] Figure 1 shows an example of the configuration of the encounter information management system according to this embodiment. [Figure 2] Figure 2 shows an example of a user terminal configuration. [Figure 3] Figure 3 shows examples of programs and data stored in the memory unit. [Figure 4] Figure 4 shows an example of a candidate buffer. [Figure 5] Figure 5 shows an example of an encounter information cache. [Figure 6] Figure 6 shows an example of word learning information. [Figure 7] Figure 7 shows an example of the operation of the processing unit when the encounter information update program is executed. [Figure 8]FIG. 8 is a diagram illustrating an example of an operation of a processing unit when an encounter information providing program is executed. [Figure 9] FIG. 9 is a diagram illustrating an example of a screen displayed by a function provided by the encounter information providing program. [Figure 10] FIG. 10 is a diagram illustrating another example of a screen displayed by a function provided by the encounter information providing program. [Figure 11] FIG. 11 is a diagram illustrating another example of a screen displayed by a function provided by the encounter information providing program. [Figure 12] FIG. 12 is a diagram illustrating another example of a screen displayed by a function provided by the encounter information providing program. [Figure 13] FIG. 13 is a diagram illustrating an example of a configuration of an encounter information management apparatus. [Figure 14] FIG. 14 is a diagram illustrating an example of a program and data stored in a storage unit. [Figure 15] FIG. 15 is a diagram illustrating an example of an encounter information DB. [Figure 16] FIG. 16 is a diagram illustrating an example of an encounter history. [Figure 17] FIG. 17 is a diagram illustrating an example of an operation of a processing unit when an encounter information management program is executed. [Figure 18] FIG. 18 is a diagram illustrating an example of a configuration of an administrator terminal. [Figure 19] FIG. 19 is a diagram illustrating an example of a program and data stored in a storage unit. [Figure 20] FIG. 20 is a diagram illustrating an example of a screen displayed by a function provided by a management work program. [Figure 21] FIG. 21 is a diagram illustrating another example of a screen displayed by a function provided by the management work program. MODE FOR CARRYING OUT THE INVENTION

[0009] Modes for implementing the present disclosure (embodiments) will be described in detail with reference to the drawings. The present disclosure is not limited by the contents described in the following embodiments. The components described below include those that can be easily conceived by those skilled in the art and those that are substantially identical. Furthermore, the components described below can be combined as appropriate.

[0010] <1.Overall Configuration Example> Figure 1 is a diagram showing a configuration example of an encounter information management system 1 according to the present embodiment. The encounter information management system 1 is a system that manages the number of times a user encounters each word in a language the user is learning, and includes user terminals 10a to 10c, an encounter information management device 20, an administrator terminal 30, and an alternative content management device 40. The number of devices shown in Figure 1 is an example, and the number of devices included in the encounter information management system 1 is not limited to the example shown in Figure 1. For example, although Figure 1 includes three user terminals, the number of user terminals included in the encounter information management system 1 may be two or less, or four or more.

[0011] The user terminals 10a to 10c are terminals used by a user who is learning a language. The user terminals 10a to 10c present words in the language the user is learning to the user in forms such as video, audio, and text, and determine whether the user has encountered the presented words. Then, the user terminals 10a to 10c transmit information about words determined to have been encountered by the user to the encounter information management device 20, and accumulate the number of encounters for each word. In addition, the user terminals 10a to 10c request the encounter information management device 20 to transmit information on the number of encounters for each word in accordance with the user's operation, and present the responded information to the user.

[0012] In this specification, "encounter" shall mean that a user views and listens to a word with sufficient interest. Therefore, words presented when the user is not looking at the screen or listening to audio do not correspond to words encountered by the user.

[0013] The encounter information management device 20 manages the number of times each word has been encountered by the user. Specifically, the encounter information management device 20 accumulates the number of encounters for each word based on the information transmitted from user terminals 10a to 10c. Furthermore, the encounter information management device 20 transmits information regarding the number of encounters for each word (hereinafter referred to as "encounter information") to the requesting party in response to requests from user terminals 10a to 10c and the administrator terminal 30.

[0014] The administrator terminal 30 is a terminal used by the administrator. The administrator is, for example, the homeroom teacher of the user's class, the user's parent, or the operator of the encounter information management system 1. The administrator terminal 30 requests the encounter information management device 20 to transmit information regarding the number of encounters for each word, in response to the administrator's actions, and presents the received information to the administrator. While the information presented to users by user terminals 10a to 10c is basically the information of that user, the information presented to the administrator by the administrator terminal 30 is the information of one or more users under the administrator's management.

[0015] The alternative content management device 40 manages content that shows the context, such as video, audio, and text, when a user encounters a word. The encounter information management system 1 stores reference information to the content, such as video, audio, and DOM text, at the time of the encounter, associated with the encountered word, so that the context when the user encountered the word can be reproduced. However, even if reference information to the original content is stored, the user may not be able to access the original content afterward due to restrictions such as copyright protection, usage period, and usage fees. Therefore, the encounter information management system 1 is configured to generate alternative content that serves as a substitute for such restricted content, store it in the alternative content management device 40, and store reference information to the alternative content associated with the encountered word.

[0016] The user terminals 10a to 10c, the encounter information management device 20, the administrator terminal 30, and the alternative content management device 40 are interconnected via network 2, enabling them to communicate with each other. Network 2 may be a wired network, a wireless network, or a combination of both. The wired and wireless networks may be implemented using any known technology.

[0017] The user terminals 10a-10c, the administrator terminal 30, and the alternative content management device 40 are further connected via network 3 to communicate with video distribution sites 50a-50c, audio distribution sites 60a-60b, and DOM (Document Object Model) text distribution sites 70a-70b. Network 3 may be a wired network, a wireless network, or a combination thereof. The wired and wireless networks may be implemented using any known technology. Note that the number of distribution sites shown in Figure 1 is an example, and the number of distribution sites connected to the encounter information management system 1 is not limited to the example shown in Figure 1.

[0018] In the example shown in Figure 1, networks 2 and 3 are depicted as separate networks, but they may be the same network. Hereafter, video distribution sites 50a-50c, audio distribution sites 60a-60b, and DOM text distribution sites 70a-70b will be collectively referred to as "sites."

[0019] Video streaming sites 50a to 50c are websites that distribute videos. These include, for example, program distribution sites that distribute movies and dramas, and video sharing sites that distribute videos uploaded by users. Videos distributed by video streaming sites 50a to 50c may include not only video data and audio data, but also subtitle data corresponding to the audio data.

[0020] Audio distribution sites 60a to 60b are websites that distribute audio. These include, for example, podcast distribution sites, internet radio distribution sites, and audiobook distribution sites. The audio distributed by audio distribution sites 60a to 60b may include not only audio data but also corresponding subtitle data.

[0021] DOM text distribution sites 70a-70b are sites that distribute DOM text. DOM text distribution sites 70a-70b include, for example, sites that distribute web pages. DOM text includes various types of text that statically constitute web pages, such as headings, article body text, captions, button labels, and link text, as well as text such as titles, descriptions, and comments (excluding video, audio, and subtitles) within content played by web browser extensions or built-in players, and text on dynamically rendered SPAs (Single Page Applications) and infinitely scrolling pages.

[0022] <2. User Terminal> (2-1. Example of user terminal configuration) Next, we will explain the configuration of user terminals 10a to 10c. Since user terminals 10a to 10c have similar configurations, we will use user terminal 10a as an example to explain the configuration below.

[0023] Figure 2 shows an example of the configuration of a user terminal 10a. In the example shown in Figure 2, the user terminal 10a includes a display unit 110, an audio output unit 120, an operation unit 130, a communication unit 140, a processing unit 150, and a storage unit 160.

[0024] The display unit 110 displays various information such as characters, graphics, still images, and videos. The display unit 110 can be configured using any display material such as liquid crystal or OLED (Organic Light Emitting Diode). The audio output unit 120 outputs sound. The audio output unit 120 includes, for example, one or more interfaces for wired connection of audio output devices such as speakers, earphones, or headphones, and one or more interfaces for wireless connection.

[0025] The control unit 130 accepts user input. The control unit 130 includes one or more input devices such as a keyboard, mouse, and microphone. The communication unit 140 includes a communication interface for communicating with networks 2 and 3. The communication unit 140 supports standards corresponding to networks 2 and 3.

[0026] The processing unit 150 controls the user terminal 10a as a whole. The processing unit 150 includes, for example, a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), and NPU (Neural Processing Unit), and memory such as DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). It loads computer programs stored in the storage unit 160, etc., into the memory and executes them using the processor to perform various processes.

[0027] The storage unit 160 stores various programs and data. The storage unit 160 is implemented by a storage device built into or locally connected to the user terminal 10a, such as a hard disk drive or solid-state drive. The storage unit 160 may also be implemented by connecting an external storage device, such as a storage server or cloud storage, to the user terminal 10a via a network.

[0028] Figure 3 shows an example of programs and data stored in the memory unit 160. In the example shown in Figure 3, the memory unit 160 stores an encounter information update program (encounter word identification program) 161, an encounter information provision program 162, a playback program 163, a word learning program 164, multiple application programs 165, a candidate buffer 166, an encounter information cache 167, word learning information 168, and setting information 169. Note that some or all of the programs and data stored in the memory unit 160 in the example shown in Figure 3 may be stored in the memory of the processing unit 150.

[0029] The encounter information update program 161 determines which words the user has encountered and provides the processing unit 150 with a function to send information about the words the user has encountered to the encounter information management device 20, thereby accumulating the number of encounters for each word. Details of the functions provided by the encounter information update program 161 will be described later.

[0030] The encounter information provision program 162 transmits a request for the number of encounters for each word to the encounter information management device 20 and provides the processing unit 150 with a function to present the responded information to the user. The functions provided by the encounter information provision program 162 also include functions to display encounter information and output encounter information in a predetermined format in cooperation with functions provided by other computer programs such as the playback program 163, the word learning program 164, and the application program 165. The cooperation between the functions provided by the encounter information provision program 162 and the functions provided by other computer programs is realized, for example, by one program using an API (Application Programming Interface) provided by the other. Details of the functions provided by the encounter information provision program 162 will be described later.

[0031] The playback program 163 provides the processing unit 150 with the function to play content provided by the video distribution sites 50a-50c, the audio distribution sites 60a-60b, and the DOM text distribution sites 70a-70b. The playback program 163 includes, for example, a media player program and a web browser program. The web browser program may be a standalone computer program or may be incorporated as a module into another program. The playback program 163 may be provided separately for each of the video distribution sites 50a-50c, the audio distribution sites 60a-60b, and the DOM text distribution sites 70a-70b.

[0032] The word learning program 164 provides the processing unit 150 with functions for the user to learn words. The word learning program 164 includes, for example, a word list app. The functions provided by the word learning program 164 include, for example, a function to present the user with questions about words, determine whether the user's answers are correct or not, update information indicating the user's progress in learning words according to the determination result, and a function to generate questions about words to present to the user.

[0033] Each of the multiple application programs 165 provides a predetermined function to the processing unit 150. The functions provided by the multiple application programs 165 include, for example, functions that display encounter information or output encounter information in a predetermined format, in cooperation with the functions provided by the encounter information provision program 162. The multiple application programs 165 include, for example, an AI agent program, an e-book reader program, and a dictionary program.

[0034] The candidate buffer 166 temporarily holds one or more words that the content being played back by the playback program 163 may currently be presenting to the user. The candidate buffer 166 is used by the encounter information update program 161.

[0035] An example of candidate buffer 166 is shown in Figure 4. In the example shown in Figure 4, candidate buffer 166 has items such as word, source type, source ID, title, reference information, and encounter date and time. The word item is set to the word that the content being played may currently be presenting to the user. The source type item is set to the type of site from which the content being played is distributed (e.g., "video", "audio", "DOM text"). The source ID item is set to the identifier of the site from which the content being played is distributed. The title item is set to the title of the content being played.

[0036] The reference information field contains references to the portion of the content where the word was presented, or to alternative content for that portion, so that the context in which the user encountered the word can be reconstructed. The reference information is expressed, for example, using a URL (Uniform Resource Locator), or a combination of a URL and at least one of the location and time of the relevant portion. The encounter date and time field contains the date and time when the user may have encountered the word.

[0037] The encounter information cache 167 temporarily holds encounter information obtained from the encounter information management device 20 in order to conceal the latency when obtaining encounter information from the encounter information management device 20 from the user. The encounter information cache 167 is used by the functions provided by the encounter information provision program 162.

[0038] An example of the encounter information cache 167 is shown in Figure 5. In the example shown in Figure 5, the encounter information cache 167 has items such as word, source type, source ID, title, reference information, number of encounters, and last encounter date and time. The word item is set to the word encountered by the user. The source type item is set to the type of site from which the content presenting the word originated (e.g., "video", "audio", "DOM text"). The source ID item is set to the identifier of the site from which the content presenting the word originated. The title item is set to the title of the content presenting the word.

[0039] The Reference Information field contains reference information to the portion of the content where the word is presented, or to alternative content for that portion. Reference information is expressed, for example, using a URL (Uniform Resource Locator), or a combination of a URL and at least one of the location and time of the relevant portion. The Encounter Count field contains the cumulative number of times the user has encountered the word. The Last Encounter Date and Time field contains the date and time the user last encountered the word.

[0040] Word learning information 168 holds information about the words the user is learning. Word learning information 168 is used by the functions provided by the word learning program 164.

[0041] An example of word learning information 168 is shown in Figure 6. In the example shown in Figure 6, word learning information 168 has items such as word, translation, pronunciation, definition, and learning information. The word item is set to a word that the user may learn. The translation item is set to the translation of the word. The pronunciation item is set to the pronunciation of the word. The definition item is set to the dictionary definition or explanation of the word.

[0042] The learning information item has sub-items such as selection flag, proficiency level, number of times asked, and number of correct answers. The selection flag item is set to a value indicating whether or not the word has been selected as a learning target. The proficiency level item is set to a value indicating the degree to which the user understands the word. The number of times asked is set to the number of times the user has been asked questions about the word. The number of correct answers is set to the number of times the user has answered questions about the word correctly.

[0043] The configuration information 169 holds various information necessary for the operation of the user terminal 10a. The information held by the configuration information 169 includes, for example, the user's user ID, the network address for each site accessed, definitions that specify how to obtain various information and how to set reference information, and parameters such as thresholds used by various computer programs.

[0044] (2-2. Example of behavior when running the encounter information update program) Figure 7 shows an example of the operation of the processing unit 150 when the encounter information update program 161 is executed. The process shown in the example in Figure 7 starts when playback of content obtained from any site is started by a function provided by the playback program 163.

[0045] First, the processing unit 150 executes the following processes in parallel: subtitle acquisition (step S101), character recognition (step S102), speech recognition (step S103), DOM analysis (step S104), reference information acquisition (step S105), and interest level determination (step S106). These processes may be executed asynchronously with the processes in steps S107 to S118, in a form that allows the processing results to be provided to the processes in steps S107 to S118 at any time.

[0046] The subtitle acquisition process is the process of acquiring zero or more words contained in the subtitles currently displayed to the user as the content is played. The subtitles are superimposed on the video or displayed in a predetermined area of ​​the screen provided by the playback program 163, using a function provided by the playback program 163. Acquisition of words contained in the subtitles currently displayed to the user may be achieved, for example, by using an API provided by the playback program 163, or by extracting data for the currently playing portion from subtitle data included in the content being played or from subtitle data acquired separately from the content, based on a time code or the like.

[0047] In the subtitle acquisition process, in addition to the words contained in the subtitles currently presented to the user, the display position and / or output time of each word may also be acquired. The display position may be acquired, for example, using an API provided by the playback program 163, or by calculation using parameters set in the configuration information 169. The configuration information 169 includes parameters set for each playback program 163 and / or site, such as the position and size at which the subtitles are displayed, the type and size of the font used to display the subtitles, and the hyphenation rules to be adopted. The output time may be a relative time from the start of content playback, or it may be the time on a clock.

[0048] Character recognition processing is the process of obtaining zero or more words contained in the screen currently displayed to the user as the content is played, by applying character recognition to the screen in question. The screen currently displayed to the user as the content is played refers to the frames that make up the video if the content is a video, or, if the content is DOM text, it refers to the combination of shapes, images, characters, etc., displayed in the content display area of ​​the web browser excluding the header and footer. Obtaining the screen currently displayed to the user may be achieved, for example, using the API provided by the playback program 163, or by using the screen capture function.

[0049] The character recognition process may target the entire screen where the content is played, or it may target only the area of ​​the screen where subtitles are displayed. The position and size of the subtitle display area on the screen are set in the configuration information 169 for each site, for example.

[0050] In character recognition processing, in addition to the words currently displayed on the screen to the user, the display position and / or output time of each word may be obtained. The display position of a word can be obtained during the process of applying character recognition to the screen. The output time may be a relative time since the content playback started, or it may be the time on a clock.

[0051] The speech recognition process involves applying speech recognition to the audio currently presented to the user as the content is played, thereby obtaining zero or more words contained in that audio. Obtaining the audio currently presented to the user may be achieved, for example, using an API provided by the playback program 163, or using an audio capture function.

[0052] In speech recognition processing, in addition to the words contained in the audio currently presented to the user, the output time of each word may be obtained. The output time may be relative to the start of content playback, or it may be the time on a clock.

[0053] DOM parsing is the process of obtaining zero or more words contained in the DOM text currently presented to the user during content playback. This acquisition of words in the DOM text can be achieved, for example, using the API provided by playback program 163.

[0054] In DOM parsing, in addition to obtaining the words contained in the DOM text, the display position of each word may also be obtained. Obtaining the display position can be achieved, for example, by obtaining a position identifier such as the DOM path, section ID, or offset.

[0055] Subtitle acquisition, character recognition, speech recognition, and DOM parsing processes do not necessarily need to be performed all at once; one or more may be selectively executed depending on the source site and type of content. Which processes to execute may be predetermined in configuration information 169 for each site and / or content type, or it may be determined by experimentally performing the processes on the acquired content to determine whether they are feasible.

[0056] If a video streaming site distributes videos in a format that allows for subtitle acquisition, the processing unit 150 may perform only the subtitle acquisition process. This is because words obtained from subtitles are considered to be more accurate than words obtained by character recognition or speech recognition. However, if the video streaming site 50a distributes videos in a format that allows for subtitle acquisition, but the subtitles omit part of the audio, the processing unit 150 may perform both the subtitle acquisition process and the speech recognition process, and then acquire words by comprehensively judging the results of these processes.

[0057] If a video streaming site distributes videos in a format that makes it impossible to obtain subtitles, the processing unit 150 may perform character recognition processing and speech recognition processing. Performing character recognition processing and speech recognition processing is done because it is thought that words can be obtained more accurately by comprehensively judging the recognition results of the two processes.

[0058] If the audio is delivered in a format that allows for the acquisition of subtitles, the processing unit 150 may perform only the subtitle acquisition process. This is because the words obtained from subtitles are considered to be more accurate than the words obtained by speech recognition. However, if the audio is delivered in a format that allows for the acquisition of subtitles, but the subtitles omit parts of the audio, the processing unit 150 may perform both the subtitle acquisition process and the speech recognition process, and then acquire the words by comprehensively judging the results of these processes.

[0059] The reference information retrieval process retrieves reference information for the currently playing portion of the content. This reference information is used to recreate the context of the content when the user encountered the word. Reference information can be represented, for example, using a URL, or a combination of the URL and the temporal relative position of the relevant portion. The reference information may also include positional information that identifies the exact location of the word within the content.

[0060] In the reference information acquisition process, if the processing unit 150 holds reference information to the above-mentioned part of the content, but the user is unable to access that part later due to restrictions such as copyright protection, usage period, or usage fees, it will have the alternative content for that part stored in the alternative content management device 40, and instead of the reference information to the above-mentioned part, it will acquire reference information to the alternative content stored in the alternative content management device 40.

[0061] Whether or not the user can access the above portion retrospectively may be determined by referring to metadata associated with the content, or by referring to parameters held in the configuration information 169 regarding the site from which the content including the above portion is distributed. The generation of alternative content may be performed by the processing unit 150, or the processing unit 150 may provide information indicating the above portion to the alternative content management device 40 and have the alternative content management device 40 perform the operation.

[0062] Alternative content should be something that helps users recall the above-mentioned portion retrospectively. Alternative content may be, for example, a thumbnail image, a summary of the context containing the above-mentioned portion, a section identifier, or at least one of a hash.

[0063] The interest level determination process determines the level of user interest in the content being played. In the interest level determination process, the processing unit 150 acquires multiple indicators and uses them comprehensively to determine the level of user interest. The indicators used by the processing unit 150 in the interest level determination process include, for example, a visibility indicator, a foreground indicator, an operation indicator, a media session indicator, a volume indicator, a path indicator, a continuous playback time indicator, and an unlock operation indicator.

[0064] The visibility index is an indicator of how easy the content playback screen is to see. A higher visibility index indicates a higher level of user interest in the content. The visibility index is highest when the content playback screen is displayed in full screen within the display area of ​​the display unit 110, and decreases as the content playback screen becomes smaller or as a larger portion of the content playback screen is located outside the display area of ​​the display unit 110. In addition, if the display position of a word is obtained, the visibility index decreases the shorter the period during which the word's display position is within the display area of ​​the display unit 110.

[0065] The foreground index indicates whether the content playback screen is displayed in the foreground of the display area of ​​the display unit 110. When the content playback screen is displayed in the foreground of the display area of ​​the display unit 110, it is determined that the user has a high level of interest in the content.

[0066] The operation metric is an indicator that shows how short the time has been since the user last performed an operation on the operation unit 130. The shorter the time elapsed since the user last performed an operation on the operation unit 130, the higher the user's interest in the content is judged to be.

[0067] The media session metric indicates the shortest elapsed time since the last detected operation on command buttons such as play start, pause, resume, fast forward, and rewind. Command buttons can be on-screen buttons or physical buttons. The shorter the elapsed time since the last detected operation on a command button, the higher the user's interest in the content is considered to be.

[0068] The volume index is an indicator of the volume level of the content being played. A higher volume index indicates a higher level of user interest in the content. The volume index may be set to increase proportionally to the volume level of the content, and then remain constant once the volume is sufficiently loud for the user to hear.

[0069] The path metric is an indicator that shows how easily users are attracted to content based on the output path of the audio. A higher path metric indicates a higher level of user interest in the content. For example, if the audio is output from earphones or headphones, the path metric will be high, while if the audio is output from speakers, the path metric will be low.

[0070] The continuous playback time index is an indicator that shows the length of time a content can be played continuously. A higher continuous playback time index indicates a higher level of user interest in the content. The continuous playback time index may be set to maintain a constant value until it reaches a predetermined time corresponding to the time when users' concentration begins to decline, and then decrease as the continuous playback time increases thereafter.

[0071] The unlock activity metric indicates how quickly the lock screen has been unlocked since the last detection. A shorter time elapsed since the last unlock activity suggests a higher level of user interest in the content.

[0072] The processing unit 150, for example, treats a value obtained by aggregating the values ​​obtained by multiplying each of these indicators by a predetermined weighting coefficient (e.g., a weighted average) as a value indicating the level of user interest in the content being played. The weighting coefficient for each indicator may differ, for example, depending on the content distribution site or the type of content.

[0073] The metrics described above are examples, and the number of metrics used to determine the level of user interest in the content being played may be greater or less than those shown in these examples. Furthermore, metrics may be selectively used depending on the type of content being played. For example, if the content being played is audio, the level of user interest in the content being played may be determined using media session metrics, volume metrics, path metrics, continuous playback time metrics, and unlock operation metrics.

[0074] Next, the processing unit 150 determines whether a new section of the content has been played (step S107). A new section is, for example, a section in which a new subtitle is displayed, a section in which the displayed content has changed by a predetermined amount or more, a section in which the volume has changed from a silent state, or a section in which a predetermined time interval has elapsed since the previous section. Whether or not a new subtitle is displayed can be determined, for example, based on a change in the acquisition result of the subtitle acquisition process or the character recognition process.

[0075] If it is determined that a new section of the content has not been played (step S107: No), the processing unit 150 returns to steps S101 to S106. If it is determined that a new section of the content has been played (step S107: Yes), the processing unit 150 proceeds to step S108.

[0076] In step S108, the processing unit 150 obtains word candidates, which are words that may be presented to the user. Specifically, the processing unit 150 obtains words as word candidates that have been obtained by the executed process among the subtitle acquisition process, character recognition process, speech recognition process, and DOM analysis process.

[0077] Next, the processing unit 150 determines the confidence level of each word candidate (step S109). The processing unit 150 determines multiple individual confidence levels for each word candidate and comprehensively evaluates the determined individual confidence levels to determine the combined confidence level of each word candidate.

[0078] The processing unit 150 may determine an individual confidence level for each word candidate based on the type of processing that obtained the word candidate. For example, the individual confidence level of a word candidate obtained by subtitle acquisition processing or DOM analysis processing may be determined to be high, while the individual confidence level of a word candidate obtained by character recognition processing or speech recognition processing, which may result in misrecognition, may be determined to be low.

[0079] If the same word candidate is detected in multiple processes, the processing unit 150 may determine that the individual confidence level of that word candidate is high. In this case, if the order and display position of the word candidates are consistent among the results obtained from multiple processes, the processing unit 150 may determine that the individual confidence level of that word candidate is even higher.

[0080] The processing unit 150 may determine the individual confidence level of word candidates based on the morphological normalization distance or the language model likelihood.

[0081] The processing unit 150 may use methods such as weighted consensus, voting, or probability estimation to comprehensively evaluate the individual confidence levels and determine the combined confidence level of each word candidate.

[0082] Next, the processing unit 150 identifies the word candidates presented by the content by discarding word candidates whose integrated confidence level is lower than a predetermined threshold (step S110). In this way, by using integrated confidence level, the word candidates presented by the content can be identified with high accuracy.

[0083] Next, the processing unit 150 determines whether the user's level of interest, obtained through the interest level determination process, is higher than the first threshold (step S111). The first threshold is a threshold used to determine whether the user's level of interest is sufficiently high.

[0084] If the user's level of interest is higher than the first threshold, in other words, if the user's level of interest is sufficiently high (step S111: Yes), the processing unit 150 determines whether or not data exists in the candidate buffer 166 (step S112). The data present in the candidate buffer 166 is data for which the final determination of whether or not the user's level of interest is sufficiently high was temporarily postponed. If data exists in the candidate buffer 166 (step S112: Yes), the processing unit 150 sends the data present in the candidate buffer 166 to the encounter information management device 20 as an encounter information update request, requesting that the cumulative value of the word encounter count be updated (step S113). The encounter information update request includes the user's user ID in addition to the items contained in the candidate buffer 166. If no data exists in the candidate buffer 166 (step S112: No), step S113 is not executed.

[0085] Next, the processing unit 150 sends the word candidates and related information, such as the reference information obtained in the reference information acquisition process, to the encounter information management device 20 as an encounter information update request, requesting that the cumulative value of the word encounter count be updated (step S114). After that, the processing unit 150 proceeds to step S117.

[0086] On the other hand, if the user's level of interest is not higher than the first threshold, in other words, if the user's level of interest is not sufficiently high (step S111: No), the processing unit 150 determines whether the user's level of interest obtained by the level of interest determination process is higher than the second threshold (step S115). The second threshold is a threshold used to determine whether to temporarily suspend the final determination of whether the user's level of interest is sufficiently high and to continue observation, and it is set lower than the first threshold.

[0087] If the user's level of interest is higher than the second threshold, in other words, if the final determination of whether the user's level of interest is sufficiently high should be temporarily suspended and observation should be continued (step S115: Yes), the processing unit 150 stores the word candidates and related information such as the reference information obtained in the reference information acquisition process in the candidate buffer 166 (step S116). After that, the processing unit 150 proceeds to step S117. By storing the word candidates in the candidate buffer 166, if a level of interest higher than the first threshold is obtained in the next level of interest determination process, the word candidates will be sent to the encounter information management device 20 as an encounter information update request.

[0088] If the user's level of interest is not higher than the second threshold, in other words, if the user's level of interest is not sufficiently high (step S115: No), the word candidates are discarded without being stored in the candidate buffer 166, and the processing unit 150 proceeds to step S117.

[0089] In step S117, the processing unit 150 removes data from the candidate buffer 166 that has been elapsed for a predetermined time or longer. The predetermined time is short enough so that the user does not fail to determine if the content is of sufficient interest to the user, for example, one to a few seconds.

[0090] Next, the processing unit 150 determines whether or not the content playback has finished (step S118). If it determines that the content playback has not finished (step S118: No), the processing unit 150 returns to steps S101 to S106. If it determines that the content playback has finished (step S118: Yes), the processing unit 150 terminates the process.

[0091] Thus, the encounter information management system 1 can comprehensively record users' encounters with words contained in content from multiple types of sites, such as video streaming sites, audio streaming sites, and DOM text streaming sites. For this reason, the encounter information management system 1 can record users' encounters with words not only when they access a predetermined specific site, but also when they access a site that suits their preferences.

[0092] Furthermore, the encounter information management system 1 comprehensively evaluates the results of multiple processes, including subtitle acquisition, character recognition, speech recognition, and DOM analysis, to determine the word presented to the user. Therefore, the encounter information management system 1 can identify the word presented to the user with high accuracy.

[0093] Furthermore, the encounter information management system 1 uses multiple indicators to determine whether the user was interested in the presented words. Therefore, the encounter information management system 1 can accurately exclude words that were presented to the user but that the user was not interested in from its records.

[0094] Furthermore, the determination of the level of user interest in a word based on multiple indicators may be implemented using layers within neural networks or machine learning models. The various thresholds and weight coefficients used to evaluate the multiple indicators may be adjusted using rule-based, statistical, machine learning, Bayesian, or a combination thereof.

[0095] Furthermore, even if the content being played is a video, depending on the user's settings, audio-only playback may be treated as an encounter (passive listening). In this case, the level of interest of the user in the word may be determined without using display-related metrics such as visibility and foreground.

[0096] (2-3. Example of operation when the encounter information provision program is executed) Figure 8 shows an example of the operation of the processing unit 150 when the encounter information provision program 162 is executed. The process shown in the example in Figure 8 is started, for example, when one of the computer programs that cooperates with the encounter information provision program 162 is started, or when the encounter information provision program 162 is started on its own.

[0097] First, the processing unit 150 determines whether a display request has been made (step S201). A display request is made when a specific function provided by a computer program that cooperates with the encounter information provision program 162 is called, or when a specific operation by the user on the screen provided by the encounter information provision program 162 is detected. The display request includes information that identifies the word to be displayed.

[0098] If it is determined that a display request has been made (Step S201: Yes), the processing unit 150 attempts to retrieve the data of the word to be displayed from the encounter information cache 167 (Step S202). If the data of the word to be displayed is successfully retrieved from the encounter information cache 167 (Step S203: Yes), the processing unit 150 displays the retrieved data on the display unit 110 (Step S204) and proceeds to Step S211.

[0099] On the other hand, if data for the word to be displayed cannot be obtained from the encounter information cache 167 (step S203: No), the processing unit 150 requests data for the word to be displayed from the encounter information management device 20 (step S205) and displays a provisional value on the display unit 110 (step S206). The provisional value is, for example, a graphic or character indicating that data is being acquired.

[0100] Next, the processing unit 150 determines whether or not it has received a response from the encounter information management device 20 (step S207). If it has not received a response (step S207: No), it repeats this determination. If it has received a response from the encounter information management device 20 (step S207: Yes), the processing unit 150 updates the displayed provisional value with the received data (step S208) and stores the received data in the encounter information cache 167 (step S209). After that, the processing unit 150 proceeds to step S211.

[0101] If it is determined in step S201 that no display request has been made (step S201: No), the processing unit 150 organizes the encounter information cache 167 (step S210). For example, the processing unit 150 uses logic such as LRU (Least Recently Used) to delete old and unused data from the encounter information cache 167 to free up space. After that, the processing unit 150 proceeds to step S211.

[0102] In step S211, the processing unit 150 determines whether or not to terminate the process. If it determines not to terminate the process (step S211: No), the processing unit 150 returns to step S201. If it determines to terminate the process (step S211: Yes), the processing unit 150 terminates the process.

[0103] Figure 9 shows an example of a screen displayed by a function provided by the encounter information provision program 162. In the example shown in Figure 9, content distributed from a video streaming site is played by a function provided by the playback program 163, and subtitles are displayed at the bottom of the screen. When the user selects "organize" from the subtitles and performs a predetermined operation, a display request to display "organize" is sent to the processing unit 150. As a result, the number of times the user has encountered the word "organize" is displayed as a pop-up screen for each site, as shown in the example in Figure 8. In this example, in addition to the number of word encounters, word information from word learning information 168 is also displayed on the same pop-up screen.

[0104] A card creation button is located in the lower left corner of the pop-up screen. When this button is pressed, the processing unit 150 updates the selection flag of the word learning information 168 so that "organize" becomes the target of learning, using a function provided by the encounter information provision program 162.

[0105] The functions provided by the encounter information provision program 162 can also be used in cooperation with functions provided by other computer programs, such as an AI agent program. For example, when an AI agent is resident on the user terminal 10 due to a function provided by the AI ​​agent program, if the user requests the AI ​​agent to perform processing related to encounter information, the AI ​​agent will acquire data from the encounter information management device 20 via the functions provided by the encounter information provision program 162, process the data, and display or output the results in a predetermined format.

[0106] Figure 10 shows another example of a screen displayed by the functions provided by the encounter information provision program 162. In the example shown in Figure 10, the screen is displayed solely by the functions provided by the encounter information provision program 162. In this example, in addition to the information obtained from the encounter information management device 20, the word information in the word learning information 168 is also displayed on the same screen, and by combining the information from both, the user can understand their learning progress from various perspectives.

[0107] Figures 11 and 12 show other examples of screens displayed by the functions provided by the encounter information provision program 162. In the example shown in Figure 11, the screen is displayed solely by the functions provided by the encounter information provision program 162. In this example, in addition to the information obtained from the encounter information management device 20, the word information in the word learning information 168 is also displayed on the same screen.

[0108] At the top of the screen, the title, the total number of words (vocabulary) the user has encountered, the total number of words the user is currently learning, and the total number of words the user has retained through learning are displayed.

[0109] Below that, there is an input field for entering search criteria for words, a legend field showing the types of sources for words the user has encountered, a filter setting field for narrowing down the words displayed, and a sorting field for specifying the order in which the words should be displayed.

[0110] In the filter section, users can specify whether or not to display words they have encountered for each source type: video, audio, web (DOM text), and SRS (word learning). The filter section also allows users to specify that words encountered across two or more sources should be displayed. In this embodiment, since the number of word encounters is recorded across multiple sources, such filtering options can be provided to the user. The filter section also allows for further filtering for more detailed results. Detailed filtering includes, for example, filtering by encounter date and time, and filtering by encounter count.

[0111] Below that, there is an area for displaying information for each word. This information includes the word and its meaning, the source structure, the number of encounters, the retention level, and the date and time of the last encounter. The source structure allows for a compact representation of how many times the user has encountered the word in each source. In the example shown in Figure 11, the source structure shows the number of encounters in each source using the ratio of bar length and numerical values, but other representations may be used if a compact representation is possible. Furthermore, by displaying the number of encounters and the retention level for each word, it becomes easier to identify words that have been encountered frequently but have not been retained, which is useful in supporting word learning.

[0112] When any row in this area is selected by the user, the details screen shown in Figure 12 is displayed. The details screen displays detailed information about the word corresponding to the selected row. In the example shown in Figure 12, the detailed information about the word includes pronunciation, meaning, total number of encounters, number of sources, retention level, last encounter date and time, encounter status by source, and encounter timeline by source. The encounter timeline by source shows chronologically when and in which source the user encountered the word. The display of intervals between encounters visualizes the time until the next encounter, helping to assess learning continuity and the risk of forgetting. In addition, each encounter is associated with a UI component that allows for the recreation of the context of the encounter using reference information.

[0113] <3. Encounter information management device> (3-1. Example of an encounter information management device configuration) Next, the configuration of the encounter information management device 20 will be described.

[0114] Figure 13 shows an example of the configuration of the encounter information management device 20. In the example shown in Figure 13, the encounter information management device 20 includes a communication unit 210, a processing unit 220, and a storage unit 230. The communication unit 210 supports standards corresponding to network 2.

[0115] The processing unit 220 controls the encounter information management device 20 as a whole. The processing unit 220 includes, for example, a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), and NPU (Neural Processing Unit), and memory such as DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). It loads computer programs stored in the storage unit 230, etc., into the memory and executes them using the processor to perform various processes.

[0116] The storage unit 230 stores various programs and data. The storage unit 230 is implemented by a storage device built into or locally connected to the encounter information management device 20, such as a hard disk drive or solid-state drive. The storage unit 230 may also be implemented by connecting an external storage device, such as a storage server or cloud storage, to the encounter information management device 20 via a network.

[0117] Figure 14 shows an example of programs and data stored in the storage unit 230. In the example shown in Figure 14, the storage unit 230 stores an encounter information management program 231, a statistics program 232, an encounter information DB (Data Base) 233, an encounter history 234, and setting information 235. Note that some or all of the programs and data stored in the storage unit 230 in the example shown in Figure 13 may be stored in the memory of the processing unit 220.

[0118] The encounter information management program 231 provides the processing unit 220 with the functions to update the encounter information DB 233 in response to encounter information update requests from user terminals 10a to 10c, and to acquire data from the encounter information DB 233 and encounter history 234 in response to requests from user terminals 10a to 10c and the administrator terminal 30, and to respond to the requester with the acquired results. Details of the functions provided by the encounter information management program 231 will be described later.

[0119] The statistical program 232 works in cooperation with the functions provided by the encounter information management program 231 to provide the processing unit 220 with the function to statistically process the data stored in the encounter information DB 233 and the encounter history 234 in response to requests from user terminals 10a to 10c and the administrator terminal 30.

[0120] The encounter information DB233 stores encounter information for multiple users. An example of the encounter information DB233 is shown in Figure 15. In the example shown in Figure 15, the encounter information DB233 has items such as user ID, word, source type, source ID, title, reference information, number of encounters, and last encounter date and time.

[0121] The User ID field contains an identifier to identify the user who encountered the word. The Word field contains the word the user encountered. The Source Type field contains the type of site from which the content presenting the word originated (e.g., "video," "audio," "DOM text"). The Source ID field contains the identifier of the site from which the content presenting the word originated. The Title field contains the title of the content presenting the word.

[0122] The Reference Information field contains reference information to the portion of the content where the word is presented, or to alternative content for that portion. Reference information is expressed, for example, using a URL (Uniform Resource Locator), or a combination of a URL and at least one of the location and time of the relevant portion. The Encounter Count field contains the cumulative number of times the user has encountered the word. The Last Encounter Date and Time field contains the date and time the user last encountered the word.

[0123] Furthermore, the number of encounters may be calculated dynamically using the encounter history 234, rather than being statically stored in the encounter information DB 233. In this case, considering the possibility of user forgetting, encounters with a longer elapsed time may be evaluated as a value closer to 0, and encounters with a shorter elapsed time may be evaluated as a value closer to 1.

[0124] The encounter history 234 stores encounter information update requests sent from user terminals 10a to 10c in chronological order. An example of the encounter history 234 is shown in Figure 16. In the example shown in Figure 16, the encounter history 234 has items such as user ID, word, source type, source ID, title, reference information, and last encounter date and time. In this example, the items of the encounter history 234 are the same as the items of the encounter information DB 233, except that there is no item for the number of encounters, so redundant explanations are omitted.

[0125] The configuration information 235 holds various types of information necessary for the operation of the encounter information management device 20. The information held by the configuration information 235 includes, for example, parameters used by various computer programs.

[0126] (3-2. Example of operation when the encounter information provision program is executed) Figure 17 shows an example of the operation of the processing unit 220 when the encounter information management program 231 is executed. The process shown in the example in Figure 17 is started, for example, when an encounter information update request is received from user terminals 10a to 10c.

[0127] First, the processing unit 220 obtains the user ID and the word encountered by the user corresponding to the user ID from the encounter information update request (step S301). Then, the processing unit 220 determines whether an encounter information update request with the same user ID and word as the obtained user ID and word has been received within a predetermined period (step S302). This determination may be achieved, for example, by maintaining a data structure such as a queue or list that holds the combination of user ID and word of received encounter information update requests for a predetermined period and referring to this data structure, or by setting a timer for each combination of user ID and word of an encounter information update request and performing debounce processing using this timer.

[0128] If the processing unit 220 determines that an encounter information update request with the same user ID and words as the acquired user ID and words has been received within a predetermined period (step S303: Yes), the processing unit 220 discards the encounter information update request and terminates processing.

[0129] On the other hand, if it is determined that no encounter information update request with the same user ID and word as the acquired user ID and word has been received within a predetermined period (step S303: No), the processing unit 220 reflects the contents of the encounter information update request in the encounter information DB 233 (step S304). Specifically, the processing unit 220 determines whether there is data in the encounter information DB 233 with the same user ID and word as the user ID and word acquired from the encounter information update request. If it is determined that such data does not exist, the processing unit 220 adds data to the encounter information DB 233 with 1 set in the encounter count field and the corresponding values ​​from the encounter information update request set in the other fields. If it is determined that such data exists, the processing unit 220 adds 1 to the value of the encounter count field of the data and overwrites the reference information and the last encounter date and time values ​​of the data with the corresponding values ​​from the encounter information update request.

[0130] Next, the processing unit 220 adds the contents of the encounter information update request to the encounter history 234 (step S305).

[0131] Thus, the encounter information management system 1 treats a user's encounter with the same word multiple times within a predetermined period as if it were only encountered once. Even if a user encounters the same word multiple times in a short period, they will not gain learning benefits for each encounter. By handling the number of encounters in this manner, the number of times a user encounters a word can be maintained at a significant value.

[0132] <4. Administrator Terminal> (4-1. Example of Administrator Terminal Configuration) Next, we will describe the configuration of the administrator terminal 30.

[0133] Figure 18 shows an example of the configuration of the administrator terminal 30. In the example shown in Figure 18, the administrator terminal 30 includes a display unit 310, an audio output unit 320, an operation unit 330, a communication unit 340, a processing unit 350, and a storage unit 360. The display unit 310, audio output unit 320, operation unit 330, communication unit 340, processing unit 350, and storage unit 360 are equivalent to the display unit 110, audio output unit 120, operation unit 130, communication unit 140, processing unit 150, and storage unit 160 in the example in Figure 2, respectively, so redundant explanations are omitted.

[0134] Figure 19 shows an example of programs and data stored in the storage unit 360. In the example shown in Figure 19, the storage unit 360 stores a management business program 361 and setting information 362. In the example shown in Figure 19, some or all of the programs and data stored in the storage unit 360 may be stored in the memory of the processing unit 350.

[0135] The management program 361 provides the processing unit 350 with various functions useful for administrators to manage users under their control. The information provided by the management program 361 includes, for example, a function to analyze the word encounter status of one or more users under the administrator's control, and a function to generate data necessary for the administrator's work based on the information stored in the encounter information management device 20. Encounter information may be aggregated and presented in association with classes or rosters within the organizational tenant. The organizational tenant may have Class and Enrollment, and learning logs may be aggregated and presented on a class basis. The dashboard for administrators (teachers) can present viewing progress, encounter heatmaps, unlearned concepts, review status, submissions (summaries / quizzes), etc., within a predetermined delay, and the retention period, sharing scope, and attachment permission may be controlled based on organizational policy.

[0136] Configuration information 362 holds various information necessary for the operation of the administrator terminal 30. The information held by configuration information 362 includes, for example, the administrator's user ID, the user IDs of users under the administrator's management, the network addresses of each site to be accessed, and parameters such as thresholds used by various computer programs.

[0137] (4-2. Examples of operations when a management program is executed) Figure 20 shows an example of a screen displayed by the functions provided by the management program 361. In the example shown in Figure 20, the user "Yuto" under the administrator's management has a history of viewing content, and it also shows how many of the words the user encountered while viewing each piece of content have been extracted as words to be learned, and how many have not yet been extracted as words to be learned.

[0138] Figure 21 shows another example of a screen displayed by the functions provided by the management program 361. In the example shown in Figure 21, a process is performed to extract words that the user encountered while viewing content within a predetermined period, and that have been encountered at least once a predetermined number of times. By selecting words to be learned from these extracted words using logic such as i+1, words that are frequently used in fields in which the user has a strong interest become the learning targets, allowing the user to approach learning with high motivation and improving the practicality of the learning results.

[0139] <5. Variation> The embodiments described above may be modified without departing from the spirit of this disclosure. For example, some or all of the processing that the processing unit is supposed to implement by executing a computer program in the embodiments described above may be implemented by wired logic.

[0140] Furthermore, the encounter information management system 1 may have a function to resolve inconsistencies that arise when a user uses multiple user terminals simultaneously. Such inconsistencies may occur, for example, when a user uses the same content on multiple user terminals, when data such as word learning information 168 is edited differently on user terminals, or when settings are different on user terminals. The encounter information management system 1 resolves inconsistencies using methods such as causal order priority, latest first, merge, CRDT (Conflict-free Replicated Data Type), and OT (Operational Transformation) according to a pre-set policy.

[0141] Furthermore, the various functions in the above-described embodiment may be appropriately distributed or consolidated. For example, some or all of the functions of the encounter information management device 20 may be distributed across multiple physical or virtual servers. Alternatively, some or all of the functions of the encounter information management device 20 may be placed on one of the user terminals 10a to 10c and the administrator terminal 30. Alternatively, some or all of the functions of the encounter information management device 20 may be placed on each of the user terminals 10a to 10c and the administrator terminal 30. [Explanation of Symbols]

[0142] 1. Encounter Information Management System 2,3 Network 10a, 10b, 10c User terminals 20 Encounter information management device 30 Administrator terminals 40 Alternative Content Management Devices 50a, 50b, 50c Video streaming sites 60a, 60b Audio distribution site 70a, 70b DOM text distribution site 110 Display section 120 Audio output section 130 Operation section 140 Communications Department 150 Processing Unit 160 Storage section 161 Encounter Information Update 162 Encounter Information Provision Program 163 Regeneration Program 164 Word Learning Program 165 Application Programs 166 candidate buffers 167 Encounter Information Cache 168 Vocabulary Learning Information 169 Configuration Information 210 Communications Department 220 processing units 230 Storage section 231 Encounter Information Management Program 232 Statistical Program 233 Encounter information DB 234 Encounter History 235 Configuration Information 310 Display section 320 Audio Output Section 330 Operation section 340 Communications Department 350 Processing Unit 360 storage section 361 Management Business Program 362 Configuration Information

Claims

1. A method for identifying words encountered by a user, Information processing device, A process of obtaining candidate words that may be presented to the user from content obtained from multiple types of information sources, using at least one of the following: subtitle acquisition processing, character recognition processing, speech recognition processing, and DOM (Document Object Model) analysis processing, A step of determining the reliability of each of the acquired candidates based on the type of processing in which the candidate was acquired, wherein if the same candidate is detected in multiple processing steps, the reliability of the candidate is determined to be higher than if it were detected in a single processing step. Using the aforementioned confidence level, the process involves identifying the candidate presented to the user from among the candidates, A step of obtaining the user's level of interest in the identified candidate based on multiple indicators, Using the aforementioned level of interest, the process involves identifying the word encountered by the user from among the identified candidates. A method for identifying encountered words that include this word.

2. The method for identifying an encountered word according to claim 1, further comprising the step of storing reference information indicating the context of the content when the word identified as having been encountered by the user was presented to the user, in association with the word identified as having been encountered by the user.

3. The method for identifying an encountered word according to claim 2, wherein the reference information is at least one of a summary, a section identifier, or a hash indicating the context of the content when the word identified as having been encountered by the user was presented to the user.

4. The method for identifying encountered words according to claim 1, wherein the indicator includes at least one of a visibility indicator, a foreground indicator, an operation indicator, a media session indicator, a volume indicator, a path indicator, a continuous playback time indicator, and an unlock operation indicator.

5. A word identification program that identifies words encountered by the user, In the processing unit of the information processing device, A process of obtaining candidate words that may be presented to the user from content obtained from multiple types of information sources, using at least one of the following: subtitle acquisition processing, character recognition processing, speech recognition processing, and DOM (Document Object Model) analysis processing, A step of determining the reliability of each of the acquired candidates based on the type of processing in which the candidate was acquired, wherein if the same candidate is detected in multiple processing steps, the reliability of the candidate is determined to be higher than if it were detected in a single processing step. Using the aforementioned confidence level, the process involves identifying the candidate presented to the user from among the candidates, A step of obtaining the user's level of interest in the identified candidate based on multiple indicators, Using the aforementioned level of interest, the process involves identifying the word encountered by the user from among the identified candidates. A program that identifies encountered words to execute.

6. A word identification device that identifies words encountered by the user, A process to obtain candidate words that may be presented to the user from content obtained from multiple types of information sources, using at least one of the following: subtitle acquisition process, character recognition process, speech recognition process, and DOM (Document Object Model) analysis process. A process for determining the confidence level of each of the acquired candidates based on the type of process in which the candidate was acquired, wherein if the same candidate is detected in multiple processes, the confidence level of the candidate is determined to be higher than if it were detected in a single process. Using the aforementioned confidence level, a process is performed to identify the candidate presented to the user from among the candidates, A process to obtain the user's level of interest in the identified candidate based on multiple indicators, Using the aforementioned level of interest, the process involves identifying the word encountered by the user from among the identified candidates. An encounter word identification device equipped with a processing unit that performs the following actions.

Citation Information

Patent Citations

  • Foreign language learning system and foreign language learning method based on communication network

    JP2008209922A

  • Method and system for providing word book learning content, and program

    JP2016115335A

  • Document retrieval learning system

    JP2018195222A

  • Learning support system, learning support method, and learning support program

    JP2020160382A

  • Information processing device, learning support method and program

    JP2023060860A