Question creation device, question creation system, and question creation method

The question creation device and system address the challenge of creating questions for video-related information by extracting and analyzing video elements to generate candidate questions, enhancing viewer interaction and information retrieval.

JP7845690B2Active Publication Date: 2026-04-14PARONYM INC
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Viewers struggle to create appropriate questions for obtaining item-related information from videos due to lack of search keywords, especially when interacting with chatbots.

Method used

A question creation device and system that extracts video elements through user actions, analyzes them, and generates candidate questions using text elements, allowing viewers to easily obtain information without interrupting video viewing.

Benefits of technology

Enables easy creation of questions to obtain information related to video objects, facilitating seamless interaction and information retrieval without disrupting the video experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007845690000001
    Figure 0007845690000001
  • Figure 0007845690000002
    Figure 0007845690000002
  • Figure 0007845690000003
    Figure 0007845690000003
Patent Text Reader

Abstract

To easily generate a question for obtaining information related to an object displayed in a moving image.SOLUTION: A question generation device includes: an element output part which extracts a moving image element from the moving image on which an action is performed when the action is performed on the displayed moving image, and outputs it to an analyzer; a text acquisition part which acquires a text-converted element created according to the moving image element by the analyzer; and a question generation part which generates a question candidate related to the moving image element from the text-converted element.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a question creation device, a question creation system, and a question creation method.

Background Art

[0002] In recent years, videos have been viewed on terminals on the user side (for example, tablet-type terminals, smartphones, personal computers, etc.). For example, Patent Documents 1 to 3 describe viewing broadcast programs such as dramas on terminals on the viewer side.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, when a viewer of a video is interested in an object (hereinafter sometimes referred to as an "item") shown in the video, the viewer may try to collect information related to the item (hereinafter sometimes referred to as "item-related information"). For example, when a viewer is watching a drama and wants to buy the bag carried by the protagonist of the drama, the viewer may try to search for the brand name and product name of the bag online.

[0005] Furthermore, viewers may attempt to gather item-related information using, for example, online chatbots. In recent years, many chatbots utilize software built on Large Language Models (LLMs). Viewers can obtain item-related information through text and voice interactions with the chatbot.

[0006] However, even if viewers try to obtain item-related information through text or voice interaction, it will be difficult to find the desired information if they do not know the search keywords for the item. In other words, because viewers do not know the search keywords for the item, they cannot easily come up with appropriate questions to ask the chatbot.

[0007] One example of the object of the present invention is to facilitate the creation of questions for obtaining information related to objects displayed in a video. Other objects of the present invention will become apparent from the description herein. [Means for solving the problem]

[0008] One aspect of the present invention is a question creation device comprising: an element output unit that extracts video elements from a video on which an action has been performed when an action has been performed on a displayed video and outputs them to an analysis device; a text acquisition unit that acquires text elements created by the analysis device based on the video elements; and a question creation unit that creates candidate questions relating to the video elements from the text elements.

[0009] One aspect of the present invention is a question creation system comprising: an element output unit that extracts video elements from a video on which an action has been performed when an action has been performed on a displayed video and outputs them to an analysis device; a text acquisition unit that acquires text elements created by the analysis device based on the video elements; a question creation unit that creates candidate questions relating to the video elements from the text elements; and a display unit that displays the candidate questions.

[0010] One aspect of the present invention is a question creation method comprising: extracting video elements from a video on which an action has been performed when an action has been performed on a displayed video, and outputting them to an analysis device; obtaining text elements created by the analysis device based on the video elements; and creating candidate questions relating to the video elements from the text elements. [Effects of the Invention]

[0011] According to the above embodiment of the present invention, questions can be easily created to obtain information related to an object displayed in a video. [Brief explanation of the drawing]

[0012] [Figure 1] Figure 1 is an explanatory diagram showing an example of an action performed on a video in this embodiment. [Figure 2] Figure 2 shows a question creation system 100 having the question creation device 10 of this embodiment. [Figure 3] Figure 3 is a flowchart of the question creation process performed by the question creation device 10 of this embodiment. [Figure 4] Figure 4 shows an example of a question template. [Figure 5] Figure 5 shows how to input text elements into a question template when there are multiple text elements. [Figure 6] Figure 6 shows an example of the screen in the first example of the question creation process. [Figure 7] Figure 7 shows an example of the screen used when acquiring video elements and creating question candidates in the first example of the question creation process. [Figure 8] Figure 8 shows an example of a screen displaying candidate questions in the first example of the question creation process. [Figure 9] Figure 9 shows an example of the screen in the second example of the question creation process. [Figure 10]FIG. 10 is a diagram showing an example of a screen when acquiring video elements in a second example of question creation processing. [Figure 11] FIG. 11 is a diagram showing an example of a screen when acquiring textified elements and creating question candidates in a second example of question creation processing.

Embodiments for Carrying Out the Invention

[0013] From the descriptions in the specification and drawings to be described later, at least the following matters will become clear.

[0014] When an action is executed on the displayed video, an element output unit that extracts video elements in the video on which the action was executed and outputs them to an analysis device, a text acquisition unit that acquires textified elements created based on the video elements by the analysis device, and a question creation unit that creates question candidates regarding the video elements from the textified elements. It becomes clear that there is a question creation device provided. According to such a question creation device, it is possible to easily create questions for obtaining information related to the object shown in the video.

[0015] It is desirable that the video elements have at least one or more of image information, audio information, and character information regarding the video. Thereby, it is possible to easily create questions for obtaining information related to the object shown in the video.

[0016] The video elements have the image information including type information regarding the video elements, the text acquisition unit acquires the type information linked to the textified elements, and the question creation unit creates the question candidates using a question template corresponding to the type information. It is desirable. Thereby, by using the type information possessed by the image information, it is possible to create highly accurate question candidates based on the type information.

[0017] Preferably, the video element consists of at least one of image information, audio information, and text information that does not include type information related to the video element, the text acquisition unit acquires the text elements associated with the type information, and the question creation unit creates the candidate questions using a question template corresponding to the type information. This allows the use of a question template corresponding to type information even if the video element does not have type information.

[0018] The text acquisition unit acquires a first text element associated with the first type information and a second text element associated with the second type information, and the question template has a first input element related to the first type information and a second input element related to the second type information, and it is desirable that the question creation unit creates a question candidate in which the first text element is input to the first input element and the second text element is input to the second input element. This can prevent the question creation unit from creating a question candidate that is expressed in an unnatural way as a question.

[0019] It is desirable to include a question output unit that outputs the aforementioned question candidates to a display device. This allows viewers to visually view the question candidates via the display device.

[0020] The question generation unit preferably generates a plurality of candidate questions, and the question output unit preferably outputs the plurality of candidate questions to the display device in a manner that allows the viewer of the video to select one. This allows the viewer to select a candidate question that suits their intention from among the multiple candidate questions.

[0021] It is desirable that the question output unit adds priority information to the multiple candidate questions and outputs them to the display device. This allows viewers to quickly select the question that best suits their intent from among the multiple candidate questions.

[0022] It is desirable that the question output unit outputs the candidate questions confirmed by the viewer of the video to the answer creation device. This allows for the automatic creation of answers to the questions confirmed by the viewer.

[0023] The aforementioned answer generation device should preferably be a large-scale language model. This allows for the automatic generation of answers to questions confirmed by the audience.

[0024] The aforementioned actions should preferably be performed by clicking, touching, voice input, or keyboard input, or a combination of any of these operations. This makes it easy to create questions to obtain information related to the subject shown in the video.

[0025] A question creation system is revealed that comprises: an element output unit that extracts video elements from the video on which an action was performed when an action was performed on the displayed video and outputs them to an analysis device; a text acquisition unit that acquires text elements created by the analysis device based on the video elements; a question creation unit that creates candidate questions about the video elements from the text elements; and a display unit that displays the candidate questions. With such a question creation system, it is possible to easily create questions to obtain information related to the object shown in the video.

[0026] A question creation method is revealed that comprises: extracting video elements from the video on which an action was performed when an action was performed on the displayed video, outputting them to an analysis device; obtaining text elements created by the analysis device based on the video elements; and creating candidate questions related to the video elements from the text elements. With such a question creation method, it is possible to easily create questions to obtain information related to the objects shown in the video.

[0027] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings. The same or equivalent components, members, etc. shown in each drawing are denoted by the same reference numerals, and redundant explanations will be omitted as appropriate.

[0028] ===Execution=== <Video Summary> Before describing the question creation device, question creation system, and question creation method of this embodiment, we will first describe an example of a video used in this embodiment and the actions performed on the video, with reference to Figure 1.

[0029] Figure 1 is an explanatory diagram showing an example of an action performed on a video according to this embodiment. Figure 1A shows an example of a screen in a video according to this embodiment, Figure 1B shows an example of an operation to stock item area information in a video (described later), and Figure 1C shows an example of a screen in a state where item area information has been stored by stocking.

[0030] As shown in Figure 1A, the video is being played on the screen (in this case, the touch panel 6) of the user terminal 5 (for example, the viewer's smartphone or tablet device). "Video" refers to moving images delivered as data from the distribution entity (in this case, the video distribution server 1, which will be described later), such as variety shows, dramas, news, sports, music, and anime. The video distribution format is not limited to streaming, as will be described later; it may also be download format, progressive download format, or live format.

[0031] Figure 1A shows a touch panel 6 displaying an image of a person wearing sportswear (specifically, a jacket) on their upper body. In this embodiment, since item areas are pre-set in the frames (still images) included in the video, the video contains information about the items for which the item areas are set (hereinafter sometimes referred to as "item area information").

[0032] Here, "item area" refers to the area on the screen corresponding to an item displayed in the video. In this embodiment, an item area is set for an item displayed in the video, and by guiding the viewer to the item area information, item area information in the video can be provided to the viewer in a simple manner. For example, when a viewer becomes interested in an item in the video, they can simply select the item area to display the item area information associated with that item area. In this embodiment, an item area is pre-set in the area of ​​the jacket on the screen. In Figure 1B, the item area is represented by a rectangular dashed frame 41.

[0033] However, the frame 41 representing the item area is not displayed on the actual touch panel 6. This prevents the frame 41 representing the item area from interfering with video viewing. In other words, Figure 1B shows the frame 41 indicating the item area for the sake of explanation. However, the frame 41 representing the item area may be displayed on the actual touch panel 6 during video playback, rather than being hidden.

[0034] When a video is played, frames (still images) are displayed sequentially, so the area occupied by the jacket on the video screen (on the frame) changes moment by moment. However, in this embodiment, the item area is set to change moment by moment in accordance with the video, and the frame 41 representing the item area also changes moment by moment.

[0035] As shown in Figure 1B, when user terminal 5 detects that an item area pre-set in the video has been tapped, it stores the item area information associated with that item area as stock information. Here, "tap operation" refers to the operation of touching and releasing the screen with a finger, and is included in "touch operation." "Touch operation" is a general term for operations performed by touching the screen with a finger. In addition to the "tap operation" described above, "touch operation" includes various operations such as "swipe," "double tap," "flick," "scroll," "drag," "long press," and "pinch in or pinch out." User terminal 5 may also store the item area information associated with other touch operations other than tap operations as stock information.

[0036] As shown in Figure 1C, when predetermined item area information is stored as stock information on the user terminal 5, an item image 44 (for example, a thumbnail image) corresponding to the stock information (stocked item area information) is displayed on the stock information display unit 43 on the touch panel 6. In other words, when a viewer becomes interested in an album cover in a video, they can tap the cover on the touch panel 6 to stock the item area information of that cover. The viewer can also confirm on the stock information display unit 43 that the item area information of that cover has been stored as stock information.

[0037] Although not shown in Figures 1B and 1C, when a viewer stores item area information, the item image (e.g., a thumbnail image) associated with that item area may be displayed. For example, when the user terminal 5 detects that an item area pre-set in the video has been tapped, it displays the item image associated with that item area. Even if the frame 41 indicating the item area (the dashed rectangle in Figure 1B) is not displayed, the viewer can recognize that if they are interested in the album cover in the video, tapping the album cover on the touch panel 6 will display the item image of the album cover, allowing them to obtain related information about that album cover. The user terminal 5 may also display the item image associated with an item area if it detects a touch operation other than a tap, such as a double tap. The viewer may store information for multiple item areas.

[0038] When the user terminal 5 detects that an area of ​​an item image 44 displayed on the stock information display unit 43 has been tapped, it performs processing according to the event information associated with that item image 44 (item area information) (for example, displaying the jacket purchase page (payment page)). In addition, the user terminal 5 may also perform processing according to the event information associated with the item image 44 (item area information) if it detects any touch operation other than a tap.

[0039] Here, the item area information for the jacket is associated with the address of the jacket's purchase page (payment page). When a viewer taps the jacket's item image 44, the jacket's purchase page (payment page) is displayed on the touch panel 6. The display screen of the jacket's purchase page (payment page) is both the item area information for the jacket and the jacket's item image 44. The web page can be displayed as part of a multi-screen setup along with the currently playing video, or it can be displayed independently.

[0040] In the following explanation, actions taken by viewers in response to the video they are watching may be referred to as "actions." Actions are not limited to the touch operations (tap and swipe operations) described above. Actions taken by viewers in response to the video they are watching may include, for example, voice input, or, if the user device 5 is a personal computer, mouse clicks or keyboard input. Actions are performed by any of the following operations: click, touch, voice input, and keyboard input, or a combination of any of these operations.

[0041] Furthermore, in the example described above, the terminal (device) on which the action is performed is the same terminal (device) on which the video is displayed. In other words, on user terminal 5 on which the video is displayed, the storage of stock information via tap operation and the display of the item purchase page (payment page) via tap operation are performed.

[0042] However, the terminal (device) on which the action is performed may be different from the terminal (device) on which the video is displayed. The action may be performed on a terminal other than the user terminal 5 on which the video is displayed, for example, on a different smartphone.

[0043] Furthermore, the frames (still images) included in the video do not need to have item areas set. In other words, the video does not need to contain item area information.

[0044] <Question Creation System 100> Next, with reference to Figure 2, the configuration of the question creation system 100 having the question creation device 10 of this embodiment will be described.

[0045] Figure 2 shows a question creation system 100 having the question creation device 10 of this embodiment.

[0046] The question creation system 100 is a system for creating questions to obtain item-related information. This item-related information includes not only information directly related to the item (e.g., the item's name, size, color, price, etc., and the item area information shown in Figure 1 above), but also indirect information about the item (e.g., recommended information related to the item).

[0047] In this embodiment, when a viewer becomes interested in an object (item) displayed in a video, they can easily create a question to obtain item-related information simply by performing an action on the displayed video. For example, if a viewer were to input text on their user terminal when creating a question in a chatbot, it would interfere with watching the video. In contrast, according to this embodiment, a question to obtain item-related information can be automatically created simply by performing an action on the displayed video, thus preventing interference with video viewing. In the question creation system 100, the question creation process is executed by multiple devices via the communication network 9 described later.

[0048] As shown in Figure 2, the question creation system 100 includes a video distribution server 1, a metadata distribution server 3, a user terminal 5, a question creation device 10, an analysis device 20, and an answer creation device 30. The video distribution server 1, the metadata distribution server 3, the user terminal 5, the question creation device 10, the analysis device 20, and the answer creation device 30 are interconnected and can communicate with each other via a communication network 9. Here, the communication network 9 can be, for example, the internet, a telephone network, a wireless communication network, a LAN, a VAN, etc., and in this case, the internet is assumed.

[0049] The video distribution server 1 is a server for distributing videos. In this embodiment, the video distribution server 1 distributes video data to the user terminal 5 in streaming format. However, the video data may be distributed in download format or progressive download format. In the case of streaming format distribution, the video data will be temporarily stored on the user terminal 5, and in the case of download format distribution, the downloaded video data will be stored and retained on the user terminal 5. Furthermore, the video distribution server 1 may also distribute videos in live format.

[0050] Metadata distribution server 3 is a server for distributing metadata, including the item area information described above. In this embodiment, a portion of the metadata is distributed in a preload format before video playback, while another portion of the metadata is distributed in a progressive download format. However, the method of distributing the metadata is not limited to this; for example, it may be in download format or streaming format.

[0051] In this embodiment, metadata and video data are described separately for the sake of explanation, but metadata may be stored in the video data (video file). When the video distribution server 1 distributes video data with metadata stored in the video data, the question creation system 100 does not need to have a metadata distribution server 3.

[0052] The metadata distributed by metadata distribution server 3 may be created by metadata distribution server 3, or it may be created by a metadata creation terminal (not shown).

[0053] User terminal 5 is an information terminal (video playback device) capable of playing videos. Here, user terminal 5 is a smartphone. However, user terminal 5 is not limited to a smartphone; for example, it could be a tablet-type mobile device or a personal computer. User terminal 5 is equipped with hardware such as a CPU, memory, storage device, communication module, and touch panel 6 (corresponding to the display unit 7A and input unit 7B), which are not shown. A video playback program is installed on user terminal 5, and the various operations described above are realized when user terminal 5 executes the video playback program. The video playback program can be downloaded to user terminal 5 from a program distribution server, which is not shown.

[0054] The user terminal 5 comprises a display unit 7A, an input unit 7B, a control unit 8A, and a communication unit 8B.

[0055] The display unit 7A is a function for displaying various screens. In this embodiment, the display unit 7A is realized by the display of the touch panel 6 and a controller that controls the display of that display. The input unit 7B is a function for inputting and detecting instructions from the user. In this embodiment, the input unit 7B is realized by the touch sensor of the touch panel 6. When the user terminal 5 is a smartphone or a tablet-type mobile terminal, the display unit 7A and the input unit 7B are mainly realized by the touch panel 6, but the display unit 7A and the input unit 7B may be made up of separate components. For example, when the user terminal 5 is a personal computer, the display unit 7A will be made up of, for example, a liquid crystal display, and the input unit 7B will be made up of a mouse or keyboard.

[0056] The control unit 8A has the function of controlling the user terminal 5. The control unit 8A has functions for processing video data and playing (displaying) the video, as well as functions for processing metadata. The control unit 8A also has a browser function that retrieves information from a web page and displays the web page. In this embodiment, the control unit 8A is realized by a CPU (not shown), a memory and storage device that store the video playback program, etc.

[0057] The communication unit 8B is responsible for connecting to the communication network 9. The communication unit 8B receives video data from the video distribution server 1, receives metadata from the metadata distribution server 3, and requests data from the video distribution server 1 and the metadata distribution server 3.

[0058] Although not shown in the diagram, the user terminal 5 may also have, in addition to the above-described configuration, a video data storage unit having the function of storing video data, a metadata storage unit having the function of storing metadata, and a stock information storage unit that stores stocked item area information in association with video data.

[0059] The question creation device 10 is a device for creating questions to obtain item-related information. The question creation device 10 performs the question creation process described later. In this embodiment, the question creation device 10 can perform various processes, including the question creation process, through the cooperation of various hardware and software (not shown) of the question creation device 10.

[0060] The question creation device 10 is, for example, a computer such as a server, and is composed of an arithmetic unit (CPU, etc.), memory, storage device, communication device, etc. The storage device stores various programs and data, including the question creation program. The arithmetic unit reads the question creation program stored in the storage device into memory and executes it, thereby realizing various processes, including the question creation process, i.e., the functions of each part described later (element output unit 11, text acquisition unit 12, question creation unit 13, and question output unit 14). Details of the functions of the element output unit 11, text acquisition unit 12, question creation unit 13, and question output unit 14 will be described later.

[0061] The question creation device 10 in this embodiment may have multiple computers. Various processes, including the question creation process, may be executed through the cooperation of these multiple computers via a network. In the question creation system 100 shown in Figure 2, the question creation device 10 is connected to one user terminal 5 via a communication network 9, but the question creation device 10 may be connected to a large number of user terminals 5 via a communication network 9.

[0062] The analysis device 20 is a device that analyzes video elements in a video and creates text elements.

[0063] Here, a video element is an element that can be extracted as data from a video, and a video element has at least one of the following related to the video: image information, audio information, and text information. Image information is, for example, image data of people or objects in the video, and is not limited to image data that includes item area information (item areas are set), but also includes image data that does not include item area information (item areas are not set). Audio information is audio data such as sounds, dialogue, and background music in the video. Text information is not limited to text information within the video such as subtitles and captions, but also includes text information outside the video such as viewer comments.

[0064] Furthermore, a text element is textual information that represents the video element mentioned above. For example, if the video element is image data of a jacket as shown in Figure 1 above, the textual information would be "clothing," "sportswear," "jacket," etc. If the color of the jacket is red, the textual information would be "red," etc. Note that there is no one-to-one correspondence between video elements and text elements; a single video element may be associated with multiple "text elements."

[0065] In this embodiment, the analysis device 20 can perform various processes, including analysis processing, through the cooperation of various hardware and software (not shown) of the analysis device 20.

[0066] The analysis device 20 is, for example, a computer such as a server, and is composed of a processing unit (CPU, etc.), memory, storage device, communication device, etc. The storage device stores various programs, including the analysis program, and various data. Various processes, including the analysis process, are realized when the processing unit reads the analysis program stored in the storage device into memory and executes it. The analysis device 20 in this embodiment may have multiple computers. Various processes, including the analysis process, may be executed through the cooperation of these multiple computers via a network.

[0067] The analysis device 20 specifically includes large-scale language models, text search engines, image search / image generation engines, sound source search / sound source generation engines, video search / video generation engines, recommendation engines, etc. The analysis device 20 is not limited to devices using AI (Artificial Intelligence); any device that analyzes video elements in a video and creates text elements may also be used, even if it does not use AI.

[0068] In the analysis device 20, for example, the text search engine and the image search engine may be configured on separate computers. Furthermore, the computer constituting the text search engine may be composed of multiple computers.

[0069] In this embodiment, the analysis device 20 was described as a separate device from the question creation device 10. That is, the analysis device 20 was described as an external device to the question creation device 10. However, the question creation device 10 may also have the functions of the analysis device 20. Furthermore, the question creation system 100 does not necessarily have to include the analysis device 20.

[0070] The answer generation device 30 is a device that generates answers to candidate questions related to video elements created from the text elements described above. In this embodiment, the cooperation of various hardware and software (not shown) of the answer generation device 30 enables the execution of various processes, including the answer generation process, in the answer generation device 30.

[0071] The answer creation device 30 is, for example, a computer such as a server, and is composed of an arithmetic unit (CPU, etc.), memory, storage device, communication device, etc. Various programs, including the answer creation program, and various data are stored in the storage device. Various processes, including the answer creation process, are realized when the arithmetic unit reads the answer creation program stored in the storage device into memory and executes it. The answer creation device 30 in this embodiment may have multiple computers. Various processes, including the answer creation process, may be executed through the cooperation of these multiple computers via a network.

[0072] The answer generation device 30 specifically includes a large-scale language model, a text search engine, an image search / image generation engine, an audio search / audio generation engine, a video search / video generation engine, a recommendation engine, etc. In this embodiment, specifically, the answer generation device 30 is a generative AI built on a large-scale language model such as CHATGPT®. The answer generation device 30 is not limited to a device using AI; it may be a device that generates answers to candidate questions, even if it does not use AI.

[0073] In this embodiment, the answer creation device 30 was described as a separate device from the question creation device 10. That is, the answer creation device 30 was described as an external device to the question creation device 10. However, the question creation device 10 may have the functions of the answer creation device 30. Also, the question creation system 100 does not have to include the answer creation device 30.

[0074] <Question generation device 10> As shown in Figure 2, the question creation device 10 includes an element output unit 11, a text acquisition unit 12, a question creation unit 13, a question output unit 14, and a recording unit 16.

[0075] The element output unit 11 is the part that extracts video elements (for example, image information of the jacket) from the video when an action is performed on the displayed video and outputs them to the analysis device 20. As described above, the analysis device 20 analyzes the video elements (for example, image information of the jacket) output from the element output unit 11 and creates text elements (for example, the text information "jacket").

[0076] Furthermore, the analysis device 20 may create multiple text elements (for example, text information such as "jacket," "clothes," and "sportswear") for a single video element (for example, image information of a jacket). The analysis device 20 creates text elements (for example, text information such as "jacket") based on the video element (for example, image information of a jacket) output from the element output unit 11 and outputs them to the text acquisition unit 12.

[0077] The text acquisition unit 12 is the part that acquires at least one text element created based on the video elements by the analysis device 20.

[0078] The question creation unit 13 is responsible for creating candidate questions about video elements (for example, "How much does the jacket cost?") from the text elements acquired by the text acquisition unit 12 (for example, the text information "jacket"). The question creation unit 13 may also create multiple candidate questions (for example, "How much does the jacket cost?", "Where is the jacket sold?") from a single text element (for example, the text information "jacket").

[0079] The question output unit 14 is the part that outputs the question candidates created by the question creation unit 13 to a display device (for example, the display unit 7A of the user terminal 5). This allows viewers to visually view the question candidates via the display device.

[0080] The recording unit 16 is the part where the video elements, text elements, and question candidate data described above are recorded. However, the recording unit 16 may also record data other than the video elements, text elements, and question candidate data. Furthermore, the question creation device 10 does not have to have a recording unit 16, and the user terminal 5 may have a recording unit 16. Specifically, the user terminal 5 may have a recording unit 16 as a memory area. This allows specific video elements, text elements, and question candidate data to be recorded on the user terminal 5 (i.e., the terminal owned by the viewer) while protecting the viewer's privacy. Moreover, both the question creation device 10 and the user terminal 5 may have a recording unit 16.

[0081] <Overview of the question creation process> Next, with reference to Figure 3, we will describe a question creation process that is an example of the question creation method of this embodiment.

[0082] Figure 3 is a flowchart of the question creation process performed by the question creation device 10 of this embodiment.

[0083] First, when an action is performed on the displayed video, the element output unit 11 extracts video elements (for example, image information of the jacket) from the video in which the action was performed and outputs them to the analysis device 20 (S001). Here, the action is not limited to one; multiple actions may be performed, and the element output unit 11 may extract at least one video element between multiple actions. For example, the element output unit 11 may extract video elements between multiple touch operations performed by the viewer within a predetermined time. This allows the element output unit 11 to extract not only video elements at the time the action was performed, but also video elements over a certain time range (a range within a predetermined time). Furthermore, the element output unit 11 may extract video elements between long press operations performed by the viewer within a predetermined time.

[0084] The analysis device 20 creates text elements (for example, the word "jacket") based on video elements (for example, image information of a jacket) output from the element output unit 11, and outputs them to the text acquisition unit 12.

[0085] Next, the text acquisition unit 12 acquires text elements (for example, the word "jacket") from the analysis device 20 (S002). Next, the question creation unit 13 creates a candidate question (for example, the word "How much does the jacket cost?") (S003).

[0086] In this embodiment, the question creation unit 13 creates candidate questions using a question template corresponding to the type information. Here, "type information" refers to text information relating to the attributes of a video element, and is data that can be linked to the text element of the video element. For example, if the video element is image data of a jacket, the text element of the video element is the text information "jacket," and the type information is the text information "proper name / product name."

[0087] If a video element is image information that includes item area information (i.e., an image of a video as shown in Figure 1), then type information may be included in the item area information of the video element. In other words, type information may be included as the item area information of the video. In this case, the text acquisition unit 12 acquires the type information (for example, text information such as "proper name / product name") by associating it with a text element (for example, text information such as "jacket"). By utilizing the type information contained in the image information, it is possible to create highly accurate question candidates based on the type information.

[0088] If a video element consists of at least one of the following: image information, audio information, or text information, which does not include item area information, the analysis device 20 may analyze the video element and create a text element (for example, the text information "jacket") associated with type information (for example, the text information "proper name / product name"). The text acquisition unit 12 acquires the text element associated with the type information. This allows the use of a question template corresponding to type information even if the video element does not have type information.

[0089] Question templates are pre-recorded in the recording unit 16, and the question creation unit 13 reads the question template data from the recording unit 16 to create candidate questions.

[0090] Figure 4 shows an example of a question template.

[0091] As shown in Figure 4, question templates are provided, such as "I want to know more about ○○○○○," "How much does ○○○○○ cost?", "Where is ○○○○○ sold?", and "I want to see different colors of ○○○○○." Here, the "○○○○○" part of the question template is where text elements (for example, the text information "jacket") are entered (hereinafter sometimes referred to as "input elements"). Furthermore, the four question templates mentioned above correspond to category information, which is "proper noun / product name."

[0092] In the case of a text element (for example, the text information "jacket") associated with type information (for example, the text information "proper name / product name"), the question creation unit 13 uses the question template corresponding to the type information "proper name / product name" from among multiple question templates. By using a question template corresponding to the type information, rather than the question creation unit 13 arbitrarily selecting and using a question template (unrelated to the type information), it is possible to create more accurate question candidates that better align with the viewer's intentions.

[0093] Figure 5 shows how to input text elements into a question template when there are multiple text elements.

[0094] The question template is not limited to having only one input element; as shown in Figure 5A, it may have multiple input elements (in this case, a first input element and a second input element). Furthermore, each input element may be linked to a predetermined type of information. That is, the first input element (○○○) may be linked to a first type of information (for example, "person's name"), and the second input element (△△△) may be linked to a second type of information (for example, "proper name / product name").

[0095] Here, if the text elements entered into the first input element and the text elements entered into the second input element are arbitrary (unrelated to type information), then, for example, a candidate question like the one shown in Figure 5B may be created. In other words, a candidate question with an unnatural expression may be created, such as "I want to know more about Mr. ABC that the jacket was wearing."

[0096] In this embodiment, since each input element is associated with predetermined type information, the question creation unit 13 creates a candidate question in which the first text element (Mr.ABC) is input to the first input element (○○○) and the second text element (Jacket) is input to the second input element (△△△), as shown in Figure 5C. This makes it possible to create a candidate question that is expressed in a natural way, such as "I would like to know more about the jacket that Mr.ABC was wearing." In other words, it is possible to suppress the creation of a candidate question that is expressed in an unnatural way, as shown in Figure 5B.

[0097] Next, the question output unit 14 outputs the question candidates to the display device (S003). This allows viewers to visually view the question candidates via the display device. Furthermore, in this embodiment, since the question candidates are displayed during video playback (in parallel with video playback), viewers can check the question candidates without interrupting their viewing of the video.

[0098] When the question creation unit 13 creates multiple question candidates, the question output unit 14 may output the multiple question candidates to a display device in a manner that allows the video viewer to select one. This allows the viewer to select the question candidate that best suits their intentions from among the multiple candidate questions.

[0099] Furthermore, the question output unit 14 may also output multiple question candidates to a display device with priority information added to them. For example, multiple question candidates may be displayed from top to bottom of the screen in order of priority. Alternatively, multiple question candidates may be displayed in order of priority through screen transitions. In addition, each of the multiple question candidates may be displayed with information on relevance (relevance in %) or priority ranking (1st, 2nd, 3rd, etc.). This allows viewers to quickly select the question candidate that best suits their intent from among multiple question candidates.

[0100] Finally, the question output unit 14 outputs the question candidate selected by the viewer (hereinafter sometimes simply referred to as "question") to the answer creation device 30 (S004). The answer creation device 30 creates an answer to the question and outputs it to the question creation device 10. The question creation device 10 outputs the answer to the question to a display device (for example, the display unit 7A of the user terminal 5). This allows the viewer to visually see the answer to the question via the display device.

[0101] The question generation device 10 may also record the answers to the questions obtained from the answer generation device 30 in the recording unit 16. This allows for further analysis of the answers to the questions and their use in the videos to be distributed. Furthermore, it is possible to extract viewer trends from the analysis results and create recommendations based on those trends.

[0102] <Specific example of the question creation process> Next, we will explain a specific example of the question creation process, referring to Figures 6 to 11. • Example 1 Figure 6 shows an example of the screen in the first example of the question creation process. Figure 7 shows an example of the screen when acquiring video elements and creating question candidates in the first example of the question creation process. Figure 8 shows an example of the screen where question candidates are displayed in the first example of the question creation process.

[0103] The first example of the question generation process shown in Figures 6 to 8 is an example in which, when a viewer becomes interested in an item with an item area set (i.e., a jacket), a question is generated to obtain item-related information (information about the jacket) about that item.

[0104] As shown in Figure 6, a question button 50 is located in the lower right corner of the touch panel 6 of the user terminal 5. This allows viewers to recognize the existence of a function that generates question candidates. As shown in Figure 7, when the user terminal 5 detects that an item area pre-set in the video has been tapped, the question creation device 10 uses its element output unit 11, text acquisition unit 12, and question creation unit 13 (each function) to create question candidates.

[0105] In this embodiment, simply performing an action (in this case, a tap) on the displayed video can automatically generate questions to obtain item-related information about the jacket, thus preventing interruptions to video viewing. As mentioned above, when the user terminal 5 detects that an item area pre-set in the video has been tapped, it stores the item area information associated with that item area as stock information. Therefore, a single action (in this case, a tap) on the displayed video can simultaneously generate questions to obtain item-related information and store it as stock information.

[0106] In Figure 7, although a list of potential questions has been created, the specific content of the questions is not displayed on the touch panel 6. However, the viewer can recognize that a list of potential questions has been created by the change in color of the question button 50 located in the lower right corner of the touch panel 6.

[0107] As shown in Figure 8, when the user terminal 5 detects that the question button 50 has been tapped, the question output unit 14 of the question creation device 10 displays the candidate question 52. Since the viewer can view the candidate question 52 at their own intended timing (the timing of tapping the question button 50), it is possible to prevent the unexpected display of the candidate question 52 from interrupting the viewing of the video. In Figure 8, the question output unit 14 outputs multiple (in this case, three) candidate questions to the touch panel 6 in a manner that can be selected by the viewer. This allows the viewer to select the candidate question that best suits their intentions from among the multiple candidate questions.

[0108] • Second example Figure 9 shows an example of the screen in the second example of the question creation process. Figure 10 shows an example of the screen when acquiring video elements in the second example of the question creation process. Figure 11 shows an example of the screen when acquiring text elements and creating question candidates in the second example of the question creation process.

[0109] The second example of the question creation process shown in Figures 9 to 11 is an example in which, when a viewer becomes interested in an item other than the item for which an item area is set (i.e., a person), a question is created to obtain item-related information (information about the jacket) for that item.

[0110] As shown in Figure 9, when the user terminal 5 detects that a tap operation has been performed on an area other than the item areas pre-set in the video, the element output unit 11 of the question creation device 10 outputs image information of a predetermined size area including the tapped area to the analysis device 20. As shown in Figure 10, when a viewer taps the image information of the predetermined size area, the item image within that area may be displayed. Viewers can recognize that even if the area is outside the item areas pre-set in the video, if they are interested in a person in the video, tapping the person on the touch panel 6 will display the person's item image, allowing them to obtain item-related information about that person.

[0111] The analysis device 20 analyzes the image information and creates text elements. As shown in Figure 10, the text elements ("ABC") are displayed on the touch panel 6, allowing the user to recognize that the text elements were created based on the image information of the tapped area. Also, as shown in Figure 11, the question creation unit 13 of the question creation device 10 creates a candidate question 52, and the question output unit 14 displays the candidate question 52.

[0112] Other examples Examples of the question creation process are not limited to the first and second examples described above. For example, in the first example described above, the question button 50 for displaying the question candidate 52 does not need to be placed. In this case, when a viewer becomes interested in an item with an item area set (i.e., a jacket), for example, by long-pressing the item, the question candidate 52 shown in Figure 11 in the second example may be displayed. Furthermore, in the second example described above, the question button 50 may be placed, and the question candidate 52 may be displayed when the viewer taps the question button 50.

[0113] ===Other=== The embodiments described above are provided to facilitate understanding of the present invention and are not intended to limit its interpretation. The present invention can be modified and improved without departing from its spirit, and it goes without saying that the present invention includes equivalents thereof. [Explanation of Symbols]

[0114] 1. Video streaming server 3. Metadata distribution server 5 User terminals 6 Touch panel 7A Display section 7B Input Section 8A Control Unit 8B Communications Department 9. Communication Network 10 Question generation device 11 Element Output Section 12 Text acquisition section 13 Question Creation Department 14 Question Output Section 16 Records Section 20 Analyzer 30. Answer generation device 41 slots 43. Stock Information Display Section 44 item details 50 Question Buttons 51 Text Elements 52 possible questions 100 Question Creation System

Claims

1. An element output unit that, when an action is performed on a displayed video, extracts video elements from the video on which the action was performed and outputs them to an analysis device, A text acquisition unit that acquires text elements created based on the video elements by the analysis device, A question creation unit that creates candidate questions related to the video elements from the text elements, A question generation device equipped with the following features.

2. The aforementioned video element has at least one of the following related to the video: image information, audio information, and text information. The question generation device according to claim 1.

3. The aforementioned video element has the image information which includes type information relating to the video element, The text acquisition unit acquires the type information and associates it with the text element. The question creation unit creates the candidate questions using a question template corresponding to the type information. The question generation device according to claim 2.

4. The aforementioned video element is at least one of image information, audio information, and text information that does not include type information related to the video element. The text acquisition unit acquires the text elements associated with the type information, The question creation unit creates the candidate questions using a question template corresponding to the type information. The question generation device according to claim 2.

5. The text acquisition unit acquires a first text element associated with the first type information and a second text element associated with the second type information. The aforementioned question template includes a first input element related to the first type information and a second input element related to the second type information. The question creation unit creates a candidate question in which the first text element is input to the first input element and the second text element is input to the second input element. The question generation device according to claim 3 or 4.

6. The system includes a question output unit that outputs the aforementioned candidate questions to a display device. The question generation device according to claim 1.

7. The question generation unit generates a plurality of candidate questions, The question output unit outputs the plurality of question candidates to the display device in a manner that can be selected by the viewer of the video. The question generation device according to claim 6.

8. The question output unit adds priority information to the plurality of candidate questions and outputs them to the display device. The question generation device according to claim 7.

9. The question output unit outputs the candidate questions determined by the viewer of the video to the answer creation device. A question generation device according to any one of claims 6 to 8.

10. The aforementioned response generation device is a large-scale language model, The question generation device according to claim 9.

11. The aforementioned action is performed by clicking, touching, voice input, or keyboard input, or a combination of any of these operations. The question generation device according to claim 1.

12. An element output unit that, when an action is performed on a displayed video, extracts video elements from the video on which the action was performed and outputs them to an analysis device, A text acquisition unit that acquires text elements created based on the video elements by the analysis device, A question creation unit that creates candidate questions related to the video elements from the text elements, A display unit that displays the aforementioned candidate questions, A question creation system equipped with the following features.

13. When an action is performed on the displayed video, the video elements of the video on which the action was performed are extracted and output to the analysis device. The analysis device obtains text elements created based on the video elements, To generate candidate questions related to the video elements from the aforementioned text elements, A question creation method that includes the following features.

Citation Information

Patent Citations

  • System and method that enables presentation of personal movie and creation of personal movie collection

    JP1997027936A

  • System for selling and buying article by using broadcast program

    JP2004023425A

  • Device, method, and program for retrieving information

    JP2007018068A

  • Goods information acquisition system

    JP2007306399A

  • Information processing apparatus, keyword registration method, and program

    JP2011180729A