Segment retrieval device and program
The segment search device addresses the challenge of searching complex content by dividing it into segments and creating abstract or Q&A-based search sentences, allowing for more flexible and effective retrieval of information.
Patent Information
- Application Number
- PCT/JP2024/044532
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-26
AI Technical Summary
Existing search technologies struggle to effectively search contents like videos, audio, and texts, especially when they contain technical terms or colloquial language, making it difficult for non-experts to find relevant information.
A segment search device that divides content into segments based on arguments and topics, creates search sentences such as abstracts and Q&A collections for each segment, and uses these search sentences to facilitate flexible searching by collating user-input search terms.
Enables more flexible and effective searching of content segments, even when the content is in colloquial language or contains technical terms, by creating abstract or Q&A-based search sentences that are easier to search through.
Smart Images

Figure JP2024044532_26062025_PF_FP_ABST
Abstract
Description
Segment search device and program
[0001] The present invention relates to a search device and a computer program for dividing content such as video, audio, and text into a plurality of segments and for searching each of the segments.
[0002] The applicant of the present application has previously proposed a system capable of appropriately extracting candidate search terms for each page of a presentation document (Patent Document 1). The system described in Patent Document 1 extracts terms (document terms) and related topic words from each page of the document and stores them in memory. Topical words related to the document terms are then extracted from this memory, and the extracted topical words are used as candidate search terms for each page of the document. This system is said to effectively provide search terms related to each page of the presentation document, enabling effective searching of each page.
[0003] Japanese Patent Application Laid-Open No. 2020-074144
[0004] In the system of Patent Document 1, the terms and topic words actually appearing on each page of a presentation document are used as search terms to be entered when searching for each page. Therefore, if a person has a thorough understanding of the terms and topic words in the document, they can certainly use those terms and topic words as search terms to effectively search for pages within the document. However, if the terms and topic words in the document are, for example, technical terms or neologisms, ordinary people who are not experts or document creators may not fully understand those terms and topics. In such cases, the problem of ordinary people still having difficulty searching for each page within the document remains.
[0005] Furthermore, it is also possible to search not only presentation materials but also text data transcribed from video or audio of lectures (with one speaker) or conversations (with multiple speakers). In such cases, the text data often contains characters written in colloquial language, which may result in a limited number of searchable keywords. Such colloquial text data presents a problem in that it is unsuitable for searches using search terms. In particular, colloquial text data may not contain correct terms or topical words, making it difficult to use the system of Patent Document 1 as a search target.
[0006] Therefore, a main object of the present invention is to provide a technology that enables each segment that constitutes content such as video, audio, and text to be searched more flexibly.
[0007] The inventor of the present invention has diligently studied means for solving the problems of the conventional inventions described above, and has discovered that by dividing content such as videos and text into multiple segments based on issues or topics, creating search statements such as summaries and question and answer collections for each segment, and by matching a search term entered by a user with the search statement created in advance to search for a segment corresponding to the search term, it becomes possible to more flexibly target various types of content for search. Based on this discovery, the inventor has conceived the idea that the problems of the conventional inventions can be solved, and has completed the present invention.
[0008] A first aspect of the present invention relates to a segment search device. The segment search device includes a segment division unit, a search query creation unit, a search term input unit, a corresponding segment search unit, and a corresponding segment output unit. The segment division unit divides content, such as video, audio, or text, into segments. For example, if the content contains multiple issues or topics, the segment division unit may divide the content into segments for each of those issues or topics. The search query creation unit creates search queries for each segment. Segments may include, for example, text converted from video or audio, or raw text. The search query creation unit may create search queries from these texts. Note that search queries contain multiple words, and preferably multiple words grouped together to form a series of sentences. Furthermore, multiple sentences may also be grouped together. The created search queries are stored in a memory or the like in association with the segments from which they were created. The search term input unit accepts search terms input from a user. The search terms may be a single word, multiple words, or a series of sentences grouped together. The corresponding segment search unit compares the search term with the search text to find a corresponding segment that corresponds to the search term from among multiple segments. The corresponding segment search unit may also search the text contained in the segment itself in addition to the search text. The corresponding segment output unit outputs the corresponding segment, access information for the corresponding segment, or a search text related to the corresponding segment. The access information for the corresponding segment may be a path specifying the location of a file or folder within a computer, or a URL specifying the location of a file or folder on the Internet.
[0009] As configured above, in the present invention, rather than using the text contained in each segment constituting content as a search target, search sentences are created from the text contained in each segment, and these search sentences are used as search targets. For example, if the text contained in a segment is written in colloquial language, the search sentences can be written in a more formal style that is more suitable for searching. Also, for example, if the text contained in a segment is technical terminology, the search sentences can be replaced with words that are more generally understandable. In this way, by creating search sentences for segments, each segment can be searched more flexibly. Furthermore, a dataset of text contained in a segment and search sentences created from it has the secondary effect of being easily usable as training data for machine learning (especially generative AI).
[0010] In the segment search device according to the present invention, it is preferable that the content is a video (including audio), and the search query is a summary of the segment or a collection of questions and answers about the segment. Such a summary summarizing the contents of the video or a collection of questions and answers (Q&A) about the contents of the video are related to the purpose of the user's search or concerns, so by creating such a summary or a collection of questions and answers as search queries, when a user inputs a search term, it is possible to more appropriately find a segment corresponding to the search term.
[0011] In the segment search device according to the present invention, the search term may be an input term to the generative AI. The corresponding segment output unit may output the corresponding segment, access information for the corresponding segment, or a search phrase related to the corresponding segment as a response from the generative AI. By incorporating generative AI into the segment search device in this way, users can more easily find segments within content.
[0012] A second aspect of the present invention is a program for causing a computer to function as the segment search device according to the first aspect. As described above, the segment search device includes a segment division unit, a search term input unit, a corresponding segment search unit, and a corresponding segment output unit. This program may be pre-installed on the computer or may be downloaded via the Internet.
[0013] A third aspect of the present invention is a non-transitory information recording medium storing the program according to the second aspect. Examples of the non-transitory information recording medium are a CD-ROM and a flash memory.
[0014] According to the present invention, each segment that constitutes content can be searched more flexibly.
[0015] Fig. 1 is an overall diagram showing an example of a system including a segment search device. Fig. 2 is a block diagram showing an example of the functional configuration of a segment search device. Fig. 3 is a schematic diagram showing an example of a process for creating a search query for each segment. Fig. 4 is a schematic diagram showing an example of segment information. Fig. 5 is a schematic diagram showing an example of output by the segment search device.
[0016] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The present invention is not limited to the embodiments described below, and includes appropriate modifications of the embodiments described below within the scope obvious to those skilled in the art.
[0017] FIG. 1 shows an example of a system 100 including a segment search device 10 and its peripheral devices. In the example shown in FIG. 1, the system 100 includes, in addition to the segment search device 10, peripheral devices such as an imaging device 20, a sound collection device 30, an input device 40, a display device 50, and a sound emission device 60, and these peripheral devices are connected to the segment search device 10. In the example shown in FIG. 1, it is assumed that a conversation between two people is photographed and recorded using the imaging device 20 and the sound collection device 30. Video data (including audio data) obtained by the imaging device 20 and the sound collection device 30 is input to the segment search device 10, and processing such as analysis and editing is performed by this segment search device 10. Therefore, in this example, video data including audio and moving images becomes the "content" to be processed by the segment search device 10.
[0018] Furthermore, a user can use this segment search device 10 to search for a desired portion of the content. For example, if the content is video data of a conversation between two people, a user can input any search term to search for one or more scenes of the conversation corresponding to the search term from the video data. In this case, the user simply inputs the search term into an input device 40 connected to the segment search device 10. Furthermore, the search results by the segment search device 10 are output from a display device 50 and a sound output device 60. Note that the content listed here is an example, and the content may be video data (including audio) of one person talking alone, or audio data that does not include moving images. Alternatively, the content may be text data.
[0019] The segment search device 10 is basically realized by a computer. Generally, a computer has an input unit, an output unit, a control unit, a calculation unit, and a memory unit, and each element is connected by a bus or the like to exchange information. The various information is digital information, and the computer can process the digital information to perform various calculations, and the memory unit can store the digital information. For example, the memory unit may store a control program or various information. When predetermined information is input from the input unit, the control unit reads the control program stored in the memory unit. The control unit then reads the information stored in the memory unit as appropriate and transmits it to the calculation unit. The control unit also transmits the input information as appropriate to the calculation unit. The calculation unit performs calculations using the received various information and stores the results in the memory unit. The control unit reads the calculation results stored in the memory unit and outputs them from the output unit. In this way, various processes and steps are performed. Each unit or means executes these various processes. The computer may have a processor, and the processor may realize various functions and steps. The computer may be standalone. A computer may have some of its functions distributed between a server and a terminal. In this case, it is preferable that the server and the terminal can exchange information via a network such as the Internet or an intranet. The computer may include a processor and a memory coupled to the processor. The memory may store instructions that, when executed by the processor, cause the computer to perform various processes or function as various elements. The computer may build a trained model by performing machine learning using various training data, and obtain desired results by inputting various information into this trained model. The obtained results may also be input into the computer and fed back to improve the accuracy of the trained model. In this case, the computer may perform various analyses using a trained model created by machine learning or deep learning of AI (artificial intelligence).
[0020] The imaging device 20 is a device for acquiring image data of still images or videos. A known digital camera can be used as the imaging device 20. The imaging device 20 may be an external device attached to the segment search device 10 as shown in FIG. 1 , or may be built into the computer constituting the segment search device 10. The imaging device 20 includes, for example, a lens, a mechanical shutter, a shutter driver, a photoelectric conversion element such as a CCD image sensor unit or a CMOS image sensor unit, a digital signal processor (DSP) that reads the charge amount from the photoelectric conversion element and generates image data, and an IC memory. The imaging device 20 converts the acquired image data into a digital signal and sends it to the segment search device 10. The segment search device 10 at least temporarily stores the image data received from the imaging device 20 in the storage unit 12.
[0021] The sound collection device 30 is a device for acquiring voice data. The sound collection device 30 can use a known microphone, such as a dynamic microphone, a condenser microphone, or a MEMS (Micro-Electrical-Mechanical Systems) microphone. The microphone may be a directional microphone or an omnidirectional (omnidirectional) microphone. The sound collection device 30 may be an external device to the segment search device 10 as shown in FIG. 1 , or may be built into the computer that constitutes the segment search device 10. The sound collection device 30 converts sound into an electrical signal, amplifies the electrical signal using an amplifier circuit, converts it into a digital signal using an A / D conversion circuit, and outputs it to the segment search device 10. The segment search device 10 at least temporarily stores the voice data received from the sound collection device 30 in the memory unit 12.
[0022] The input device 40 is a device for transmitting information input by a user to the segment search device 10. As the input device 40, a known input device such as a mouse, keyboard, touch panel, or stylus pen may be used. The input device 40 may be physically separable from the segment search device 10, in which case the input device 40 is connected to the segment search device 10 via a short-range wireless communication standard such as Bluetooth (registered trademark). Also, a touch panel display may be configured by placing a touch panel in front of a display.
[0023] The display device 50 is a device for displaying predetermined images (still images and videos). The display device 50 is configured with a known display device such as a liquid crystal display or an organic EL display. The display device 50 may be a device external to the segment search device 10 as shown in FIG. 1, or may be built into the computer that constitutes the segment search device 10.
[0024] The sound emitting device 60 is a device for outputting sound. The sound emitting device 60 converts an electrical signal into physical vibrations (i.e., sound) and outputs the converted sound. An example of the sound emitting device 60 is a general speaker that transmits sound to the wearer by air vibrations. The sound emitting device 60 may be a device external to the segment search device 10 as shown in FIG. 1 , or may be built into the computer that constitutes the segment search device 10. The sound emitting device 60 may also be a bone conduction speaker that transmits sound to the wearer by vibrating the wearer's bones.
[0025] FIG. 2 is a block diagram showing the functional configuration of the system 100. In particular, FIG. 2 shows the functional blocks of the segment search device 10. The segment search device 10 mainly includes a processing unit 11, a storage unit 12, and a communication unit 13. The processing unit 11 is composed of, for example, a processor and a memory. Examples of the processor include a known CPU or GPU. The processor performs predetermined arithmetic processing and image processing according to programs and data stored in the memory, and executes various control processes while writing the results of the processing to a working space in the memory. The memory is composed of, for example, a volatile memory such as RAM (Random Access Memory) and is used for the arithmetic processing by the processor. The storage unit 12 is a storage element (element) for storing data mainly used for the arithmetic processing by the processing unit 11. The storage unit 12 is composed of a non-volatile memory such as ROM (Read Only Memory) or flash memory, or a hard disk drive (HDD). The storage unit 12 may also store a computer program for causing the processing unit 11 to execute predetermined processes. Furthermore, as will be described in detail later, the storage unit 12 stores content, segment information, and trained models as information used in the arithmetic processing in the processing unit 11 or information obtained as a result of the arithmetic processing in the processing unit 11. The communication unit 13 is an element that enables the segment search device 10 to send and receive data to and from an external computer (e.g., a web server) via the Internet. The communication unit 13 may be any unit that can send and receive data via wired or wireless communication. When wireless communication is performed, a communication module that complies with a known wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) can be used as the communication unit 13.
[0026] 2, the processing unit 11 has functional blocks such as a text conversion unit 11a, a segment division unit 11b, a search query creation unit 11c, a search term input unit 11d, a corresponding segment search unit 11e, and a corresponding segment output unit 11f. These functional blocks are functions obtained by the processing unit 11 executing a predetermined program, and are actually realized mainly by cooperation between a processor and a main memory. Details of these functional blocks will be described with reference to the examples shown in FIGS. 3 to 5 in addition to FIG. 2.
[0027] As shown in Fig. 3, content such as video (including audio), audio, and text is processed by the segment search device 10. For example, if the content to be processed is video or audio, the segment search device 10 can acquire the video data or audio data using the imaging device 20 and / or sound collection device 30 described above. In addition, the segment search device 10 can also receive (download) content such as video, audio, and text via the Internet using the communication unit 13. Data related to the content is stored in the memory unit 12 of the segment search device 10.
[0028] Next, if the content to be processed is video or audio, the text conversion unit 11a of the segment search device 10 reads the video data or audio data stored in the storage unit 12 and performs speech recognition processing on the video data or audio data. The speech recognition processing involves, for example, preprocessing the audio data included in the video data or the audio data itself, such as noise removal and signal amplification, and then extracting features from the audio data. Examples of the features include well-known Mel spectrum and Mel Frequency Cepstrum Coefficients (MFCC). This allows the frequency components and temporal patterns of the audio data to be extracted. The extracted features are then supplied to a well-known speech recognition engine or speech recognition model (a machine learning trained model) to convert the audio data into text data. The speech recognition model can analyze the features and generate the most likely text data. The text conversion unit 11a preferably also performs mistranslation correction processing in conjunction with the speech recognition processing. The mistranslation correction processing can generate text data with context-based errors corrected, for example, by introducing a language model that takes context into the aforementioned speech recognition model. This allows the text conversion unit 11a to convert the voice data into text data with high accuracy. The voice recognition and erroneous conversion correction are not limited to the above-mentioned methods, and any known method can be used as appropriate.
[0029] If the content to be processed is a sentence, the document contains text data, and therefore the function of the text conversion unit 11a does not need to be executed. However, the text conversion unit 11a may correct typos and omissions by correcting mistranslations in the text data contained in the sentence.
[0030] Next, the segmentation unit 11b of the segment search device 10 divides the text data generated by the text conversion unit 11a into multiple segments (also called clusters). Each of the multiple segments contains a different topic or issue. Each segment basically consists of a single topic or issue, forming a group of documents containing similar topics or issues. The segmentation unit 11b may segment the text data by topic or issue using known natural language processing. For example, the segmentation unit 11b may use an unsupervised data classification method such as k-means to group text data with similar content into a single segment. The segmentation unit 11b may also create segments for each topic using a topic model such as LDA (Latent Dirichlet Allocation), which divides the entire text into topics and groups sentences related to each topic. The segmentation unit 11b may also use a method such as TF-IDF (Term Frequency-Inverse Document Frequency) to extract important words and phrases from the text data and segment the text based on the importance of those words and phrases. Furthermore, the segmentation unit 11b can use a pre-trained language model, such as a recurrent neural network (RNN), a transformer-based model, or a generative pre-trained transformer (GPT), to segment text data based on topics or points of discussion, taking into account contextual information of the text data. These methods may be combined. Furthermore, a user may be able to input information regarding whether the segmentation is correct or incorrect to the trained model. This allows the accuracy of the model to be improved by repeating the cluster division process. Figure 3 shows a conceptual diagram of segmented conversation between two speakers, speaker A and speaker B. In this example, the conversation between the two speakers is divided into smaller points of discussion based on keywords and key phrases each time the content of the conversation changes, and the segments are visualized.
[0031] Next, the search sentence creation unit 11c creates a search sentence for each segment. The search sentence is created based on the character strings of the text data that make up the segment, and is, for example, a converted, abstracted, summarized, or supplemented version of the character strings included in the text data. Examples of search sentences are summaries and question and answer collections.
[0032] A summary is a concise summary that extracts key information from the text data that constitutes each segment. There are two types of summaries: extractive and generative. Extractive summaries can be created by extracting important words and phrases from a text using techniques such as TF-ID and then extracting sentences or passages that contain them. Extractive summaries can also be created by using algorithms such as PageRank to evaluate the importance of sentences in a text and then ranking and extracting the most important sentences. Generative summaries can be generated by using neural network models such as Seq2Seq with an encoder-decoder structure to understand the meaning of sentences and generate new sentences. Generative summaries can also be generated using a Transformer architecture with a self-attention mechanism to identify important parts of a sentence. Automatic evaluation metrics (e.g., ROUGE, BLEU) can also be used to evaluate the quality of summaries.
[0033] The Q&A collection pairs anticipated questions and answers regarding the text data constituting each segment. The segment search device 10 is connected to the Internet, allowing it to access various web pages and crawl the content of those web pages. The search query creation unit 11c obtains information from various web pages based on keywords contained in each segment and stores the information appropriately in the memory unit 12. The search query creation unit 11c may then automatically create a Q&A collection related to the keywords contained in each segment. For example, if the keyword is "Tablet A," the search query creation unit 11c obtains various information from the package insert for "Tablet A" and stores it in the memory unit 12. The search query creation unit 11c then creates a question such as, "Are there any contraindications for using Tablet A?" and a response such as, "Yes. Tablet A cannot be prescribed to patients who have been prescribed a drug for disease B." The search query creation unit 11c can also use natural language processing technology to create questions and answers from the text data contained in the segments. For example, the query creation unit 11c can generate questions and answers from the text data of each segment using a pre-trained language model such as BERT (Bidirectional Encoder Representations from Transformers) or GPT. This method uses the text data as input to the language model, and can output questions of a specific format and their optimal answers. It is also possible to manually create a large number of question-answer pairs and use them as training data to train a machine learning model (e.g., a Seq2Seq or Transformer-based model) to automatically generate answers to new questions.
[0034] The search queries (e.g., summaries and question and answer collections) created by the search query creation unit 11c are stored in the storage unit 12 in association with each segment, along with the text data contained in the segment and information (metadata) about the segment. In this specification, information associated with each segment is referred to as segment information. Figure 4 shows an example of segment information. As shown in Figure 4, the segment information includes, for example, the segment name, the title of the topic or issue contained in the segment, a timeline (time information) for each segment in the content, a summary, an explanatory text, and a question and answer collection. The segment name is a file name unique to the segment. For example, if content is divided into eight segments, segment names such as Issues 1 to 8 are automatically assigned in chronological order. The title succinctly expresses the topic or issue contained in the segment and may be created manually. Alternatively, when the search query creation unit 11c creates a summary using the above-described method, a title that more succinctly expresses the summary may be automatically generated. The timeline is information regarding the time of each segment in the content when the content to be processed is time-series data such as video or audio. Furthermore, when the content to be processed is a text, instead of a timeline, information that can identify the appearance position of each segment in the text, such as the page number, paragraph number, or line number in the text, may be used. The summary text is created by the search text creation unit 11c using the method described above. The explanatory text is created by the text conversion unit 11a using voice recognition processing. When the text data is in a conversational format, the text data may be organized for each speaker (A or B). The question and answer collection is created by the search text creation unit 11c using the method described above.
[0035] Segment information including the above information is stored as a database in the storage unit 12 and is used when searching for segments. As shown in FIG. 3 , the segment search device 10 performs preprocessing before executing a search by converting each piece of content into text, dividing the text data into multiple segments, and creating search statements (such as summaries and question-and-answer collections) for each segment. This allows content that does not actually contain text, such as video or audio, to be searched by text. Furthermore, because content is divided into multiple segments, even more detailed segments of the content can be searched. Furthermore, each content segment is associated not only with the text that constitutes it, but also with summaries and question-and-answer collections created based on that text. Therefore, even if the characters entered by the user during a search do not match the text, the characters may match the characters in the summaries or question-and-answer collections, allowing for more flexible search of each segment. In particular, when the content is a conversation between two or more people (colloquial language), searching by text is usually difficult, but by creating a summary of the conversation or a collection of questions and answers (written language), it becomes easier to search by text for the desired segment of the conversation.
[0036] Next, an example of a method for searching for segments within content will be described with reference to FIG. 5 . In particular, FIG. 5 illustrates an example in which generative AI is used to search for segments. Generative AI generally refers to technology in the fields of machine learning and artificial intelligence that has the ability to generate new data and information. Such generative AI is a trained model that learns from given data and patterns and can generate new data and content based on the learned content. Such models are trained on large amounts of data, understand the patterns and characteristics of the data, and learn parameters for generating new data. A typical example of generative AI is an architecture called GAN (Generative Adversarial Network), which learns by having two networks, a generator and a discriminator, compete with each other to generate data similar to real data. Note that this generative AI (trained model) may be stored in the memory unit 12 of the segment search device 10, or may be stored on a web server so that the segment search device 10 can access the generative AI via the Internet as needed.
[0037] As shown in FIG. 5, the search term input unit 11d of the segment search device 10 provides a UI (user interface) for inputting search terms (questions) to the generation AI. Specifically, the search term input unit 11d displays the UI on the display device 50 and accepts input of search terms from the user via the UI. The search term may be a single word, multiple words, a single sentence, or a sentence consisting of multiple sentences. In the example shown in FIG. 5, the search term functions as a prompt for the generation AI.
[0038] Next, the corresponding segment search unit 11e of the segment search device 10 searches for a corresponding segment corresponding to the search term entered by the user from among the segments of the content created in advance. In the example shown in FIG. 5, segment information including a search sentence created in advance (see FIG. 4) is input as training data to the generation AI, and a secondary model is created by retraining the generation AI. Therefore, by inputting a search term (prompt) to the generation AI, the generation AI can obtain an answer that takes into account the segment information created in advance. The corresponding segment search unit 11e can use this generation AI to search for a corresponding segment corresponding to the search term.
[0039] Next, the corresponding segment output unit 11f of the segment search device 10 outputs the information extracted by the corresponding segment search unit 11e. As shown in FIG. 5, the corresponding segment output unit 11f may output, for example, the answer of the generation AI by displaying it on the display device 50. Specifically, the generation AI generates an answer to the search term. Furthermore, if the search term includes a question about searching for segments within content, the display device 50 displays the corresponding segments extracted by the search, access information to the corresponding segments, and search phrases related to the corresponding segments. For example, if the search term is "Please search for issues related to Tablet A in video content α," the response of the generation AI may include segment information for segments containing issues related to Tablet A among the multiple segments constituting video content α, as well as a portion of the video containing issues related to Tablet A and its time period. Furthermore, access information for the video of issues related to Tablet A may be displayed in the form of a path specifying the location of a file or folder within a computer or a URL specifying the location of a file or folder on the Internet. This allows users to easily search for desired segments from any content. Furthermore, when a user instructs playback of video content, the corresponding segment output unit 11f displays an image of the video content on the display device 50 and outputs the sound of the video content from the sound output device 60. Furthermore, the corresponding segment output unit 11f can transmit (upload) the segment information searched for by the corresponding segment search unit 11e to a web server via the communication unit 13, or can transmit it individually to another computer.
[0040] 5, the corresponding segment search unit 11e and the corresponding segment output unit 11f each use generation AI to search for and output segments. However, the corresponding segment search unit 11e may simply compare the search term input by the user via the search term input unit 11d with the explanatory text or search text (summary, Q&A collection) in the segment information stored in the memory unit 12, and extract segment information containing terms that match the search term. In this case, the corresponding segment output unit 11f may simply provide the segment information extracted by the corresponding segment search unit 11e to the user by, for example, displaying it on the display device 50.
[0041] In the above description of the present invention, the embodiments of the present invention have been described with reference to the drawings in order to express the contents of the present invention. However, the present invention is not limited to the above embodiments, and includes modifications and improvements that are obvious to those skilled in the art based on the matters described in the present specification.
[0042] The present invention relates to a segment search device and program, and can be used in information-related industries.
[0043] DESCRIPTION OF SYMBOLS 10...Segment search device 11...Processing unit 11a...Text conversion unit 11b...Segment division unit 11c...Search sentence creation unit 11d...Search term input unit 11e...Corresponding segment search unit 11f...Corresponding segment output unit 12...Storage unit 13...Communication unit 20...Imaging device 30...Sound collection device 40...Input device 50...Display device 60...Sound emission device 100...System
Claims
1. A segment search device comprising: a segment division unit which divides video, audio or text content into segments; a search sentence creation unit which creates a search sentence related to each of the segments; a search term input unit to which search terms are input; a corresponding segment search unit which compares the search terms with the search sentences to find corresponding segments among the segments which correspond to the search terms; and a corresponding segment output unit which outputs the corresponding segments, access information for the corresponding segments, or search sentences related to the corresponding segments, wherein the search sentences include a question and answer collection which pairs expected questions and answers regarding text data constituting the segments.
2. A segment search device according to claim 1, wherein the content is a video, and the search text further includes a summary of the segment.
3. A segment search device as described in claim 1, wherein the search term is an input term of a generative AI, and the corresponding segment output unit outputs the corresponding segment, access information for the corresponding segment, or a search sentence related to the corresponding segment as a response from the generative AI.
4. A program for causing a computer to function as a segment search device having: a segment division unit which divides content, such as video or text, into segments; a search sentence creation unit which creates search sentences related to each of the segments; a search term input unit to which search terms are input; a corresponding segment search unit which matches the search terms with the search sentences to find corresponding segments among the segments which correspond to the search terms; and a corresponding segment output unit which outputs the corresponding segments, access information for the corresponding segments, or search sentences related to the corresponding segments, wherein the search sentences include a collection of questions and answers paired with expected questions regarding text data constituting the segments and their answers.
5. A non-transitory information recording medium storing the program according to claim 4.
Citation Information
Patent Citations
Moving picture retrieving information generating apparatus, summary information generating apparatus for moving picture, recording medium and summary information storing format for moving picture
JP2004104836A
Hint information description method
JP2008086030A
Video service providing method and service server using the same
JP2019212308A
System for providing knowledge sharing service based on multimedia platform
KR102434880B1