Segment search device and program
By segmenting content and creating search sentences for each segment, the system addresses the challenge of searching colloquial and technical terms, enabling more flexible and accurate content retrieval.
Patent Information
- Application Number
- JP2023216002
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2043-12-21
AI Technical Summary
Existing systems struggle to effectively search content in colloquial language or technical terms, especially in videos, voices, and texts, as ordinary people may not understand technical terms or coined words, making it difficult to use existing search technologies.
The system divides content into segments, creates search sentences such as abstracts and Q&A collections for each segment, and uses these to facilitate flexible searching by collating user input terms with pre-created search texts.
This approach allows for more flexible and accurate searching of content by using search sentences that are easier to understand, even when the original content is in colloquial language or technical terms, and also serves as teacher data for machine learning.
Smart Images

Figure 2025099374000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a search device that can divide content such as videos, voices, and texts into a plurality of segments and make each segment searchable, and a computer program related thereto.
Background Art
[0002] The applicant of the present application has previously proposed a system that can appropriately extract candidates for search terms related to each page of presentation materials (Patent Document 1). In the system described in this Patent Document 1, terms (terms in the material) and related topic words are extracted from each page of the material and stored in a memory, and topic words related to the terms in the material are extracted from this memory, and the topic words extracted here are used as candidates for search terms for each page in the material. Thus, according to the system of Patent Document 1, it is said that search terms related to each page of presentation materials can be effectively provided and each page can be effectively searched.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, in the system of Patent Document 1, the terms and topic words actually described on each page of the presentation material are used as search terms to be input when searching each page. Therefore, for those who fully understand the terms and topic words in the material, it is indeed possible to effectively search for the pages in the material by using those terms and topic words as search terms. However, when these terms and topic words in the material are, for example, technical terms or coined words, ordinary people who are not experts or material creators may not be able to fully understand these terms and topics. In such a case, the problem remains that it is still difficult for ordinary people to search for each page in the material.
[0005] Further, consider searching text data obtained by transcribing videos and voices of lectures (one speaker) or conversations (multiple speakers), not limited to presentation materials. In this case, the character strings included in the text data are often mainly in colloquial language, and there may be few search keywords to be targeted. Such text data in colloquial language has the problem of being unsuitable for search using search terms. In particular, since the text data in colloquial language may not include correct terms or topic words, it is also difficult to use the system of Patent Document 1 for searching.
[0006] Therefore, the main object of the present invention is to provide a technology that can make each segment constituting contents such as videos, voices, and articles more flexible as a search target.
Means for Solving the Problems
[0007] As a result of earnestly studying means for solving the problems of the above-described conventional invention, the inventor of the present invention divided contents such as videos and articles into a plurality of segments according to arguments and topics, created search texts such as abstracts and Q&A collections for each segment, and when a search term is input by a user, collated this search term with the search texts created in advance, and obtained the finding that various contents can be made more flexible search targets by searching for the segment corresponding to this search term. Then, based on the above finding, the inventor conceived that the problems of the conventional invention can be solved, and completed the present invention.
[0008] A first aspect of the present invention relates to a segment search device. The segment search device includes a segment division unit, a search sentence creation unit, a search term input unit, a corresponding segment search unit, and a corresponding segment output unit. The segment division unit divides content, which is video, audio, or text, into segments. For example, when the content includes a plurality of arguments or topics, the segment division unit may delimit segments for each of those arguments or topics. The search sentence creation unit creates a search sentence for each segment regarding the segment. The segment includes, for example, text converted from video or audio, or the text of the article itself, and the search sentence creation unit may create a search sentence from these texts. Note that the search sentence includes a plurality of words, and in particular, it is preferable that a plurality of words form a series of sentences, and furthermore, it may be a text formed by a plurality of sentences. The search sentence created here is stored in a memory or the like in association with the segment on which it is based. The search term input unit receives an input of a search term from a user. The search term may be a single word, a plurality of words, or a series of sentences formed by a plurality of words. The corresponding segment search unit collates the search term and the search sentence, and finds a corresponding segment, which is a segment corresponding to the search term, from among a plurality of segments. Note that the corresponding segment search unit may use, in addition to the search sentence, the text itself included in the segment as a search target. The corresponding segment output unit outputs the corresponding segment, access information of the corresponding segment, or a search sentence regarding the corresponding segment. Note that the access information of the corresponding segment may be a path specifying the location of a file or folder inside a computer, or a URL specifying the location of a file or folder on the Internet.
[0009] As described above, in the present invention, instead of directly using the text included in each segment that constitutes the content as the search target, a search sentence is created from the text included in each segment, and this search sentence is used as the search target. For example, when the text included in the segment is in colloquial language, the search sentence may be in classical Chinese, which is more suitable for searching. Also, for example, when the text included in the segment is a technical term, it may be replaced with a more commonly understood word in the search sentence. By creating a search sentence for each segment in this way, each segment can be searched more flexibly. There is also a secondary effect that the dataset of the text included in the segment and the search sentence created therefrom is easy to use as teacher data for machine learning (especially generative AI).
[0010] In the segment search device according to the present invention, the content is a video (including audio), and the search sentence is preferably an abstract sentence of the segment or a Q&A set of the segment. Since such an abstract sentence summarizing the content of the video or a Q&A set (Q&A) regarding the content of the video is related to the purpose and questions of the user's search, by creating these abstract sentences or Q&A sets as search sentences, when a search term is input by the user, the segment corresponding to the search term can be found more appropriately.
[0011] In the segment search device according to the present invention, the search term may be an input term for generative AI. Also, the corresponding segment output unit may output, as an answer of the generative AI, the corresponding segment, access information of the corresponding segment, or a search sentence regarding the corresponding segment. By introducing generative AI into the segment search device in this way, the user can discover the segments in the content in a simpler way.
[0012] A second aspect of the present invention is a program for causing a computer to function as the segment search device according to the first aspect described above. As described above, the segment search device includes a segment division unit, a search term input unit, a corresponding segment search unit, and a corresponding segment output unit. Note that this program may be pre-installed in the computer or may be downloaded via the Internet.
[0013] A third aspect of the present invention is a non-transitory information recording medium storing the program according to the second aspect. Examples of the non-transitory information recording medium are a CD-ROM and a flash memory.
Advantages of the Invention
[0014] According to the present invention, each segment constituting the content can be made a search target more flexibly.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Modes for Carrying Out the Invention
[0016] Hereinafter, modes for carrying out the present invention will be described with reference to the drawings. The present invention is not limited to the modes described below, and includes those appropriately modified by those skilled in the art within an obvious range from the following modes.
[0017] Figure 1 shows an example of a system 100 including a segment search device 10 and its peripheral devices. In the example shown in Figure 1, this system 100 includes peripheral devices such as an imaging device 20, a sound collection device 30, an input device 40, a display device 50, and a sound playback device 60, in addition to the segment search device 10, and these peripheral devices are connected to the segment search device 10. In the example shown in Figure 1, it is assumed that a conversation between two people is photographed and recorded using the imaging device 20 and the sound collection device 30. The video data (including audio data) obtained by the imaging device 20 and the sound collection device 30 is input to the segment search device 10, and processing such as analysis and editing is performed by this segment search device 10. Therefore, in this example, the video data including audio and moving images becomes the "content" to be processed by the segment search device 10.
[0018] Also, the user can use this segment search device 10 to search for a desired part from this content. For example, when the content is video data of a conversation between two people, the user can input an arbitrary search term to find one or more scenes of the conversation corresponding to the search term from the video data. At this time, the user may input the search term to the input device 40 connected to the segment search device 10. Also, the search results by the segment search device 10 are output from the display device 50 and the sound playback device 60. Note that the content mentioned here is just an example, and the content may be video data of a monologue of one person (including sound), or may be audio data not including moving images. In addition, the content may be text data.
[0019] The segment search device 10 is basically implemented by a computer. Generally, a computer has an input unit, an output unit, a control unit, an arithmetic unit, and a storage unit, and each element is connected by a bus or the like so that information can be exchanged. Various types of information are digital information, and the computer can process digital information to perform various operations, and the storage unit can store digital information. For example, a control program may be stored in the storage unit, or various types of information may be stored. When predetermined information is input from the input unit, the control unit reads out the control program stored in the storage unit. Then, the control unit appropriately reads out the information stored in the storage unit and transmits it to the arithmetic unit. Also, the control unit transmits the appropriately input information to the arithmetic unit. The arithmetic unit performs arithmetic processing using the received various types of information and stores it in the storage unit. The control unit reads out the arithmetic result stored in the storage unit and outputs it from the output unit. In this way, various processes and each step are executed. The ones that execute these various processes are each unit and each means. The computer may have a processor, and the processor may realize various functions and various steps. The computer may be a stand-alone one. Part of the functions of the computer may be distributed between a server and a terminal. In that case, it is preferable that the server and the terminal can exchange information through a network such as the Internet or an intranet. The computer may include a processor and a memory connected to the processor. And, if the memory stores instructions, and when the instructions are executed by the processor, the computer may be caused to perform various steps or the computer may function as various elements. The computer may construct a learned model by performing machine learning by giving various types of training data, and input various types of information into this learned model to obtain a desired result. Also, the obtained result may be input into the computer and fed back to improve the accuracy of the learned model. In this case, the computer may execute various analyses and examinations using a learning model created by machine learning and deep learning of AI (artificial intelligence).
[0020] The imaging device 20 is a device for acquiring image data of still images or moving images. As the imaging device 20, a known digital camera can be used. The imaging device 20 may be an externally attached device to the segment search device 10 as shown in FIG. 1, or may be built into the computer constituting the segment search device 10. The imaging device 20 includes, for example, a photoelectric conversion element such as a lens, a mechanical shutter, a shutter driver, a CCD image sensor unit or a CMOS image sensor unit, a digital signal processor (DSP) that reads out the charge amount from the photoelectric conversion element and generates image data, and an IC memory. The imaging device 20 converts the acquired image data into a digital signal and sends it to the segment search device 10. The segment search device 10 stores the image data received from the imaging device 20 at least temporarily in the storage unit 12.
[0021] The sound collection device 30 is a device for acquiring voice data. As the sound collection device 30, a known microphone such as a dynamic microphone, a condenser microphone, or a MEMS (Micro-Electrical-Mechanical Systems) microphone can be used. The microphone may be a directional microphone or an omnidirectional (non-directional) microphone. Also, the sound collection device 30 may be an externally attached device to the segment search device 10 as shown in FIG. 1, or may be built into the computer constituting the segment search device 10. The sound collection device 30 converts sound into an electrical signal, amplifies the electrical signal by an amplifier circuit, and then converts it into a digital signal by an A / D conversion circuit and outputs it to the segment search device 10. The segment search device 10 stores the voice data received from the sound collection device 30 at least temporarily in the storage unit 12.
[0022] The input device 40 is a device for transmitting information input by the user to the segment search device 10. As the input device 40, a known input device such as a mouse, a keyboard, a touch panel, or a stylus pen may be used. Note that the input device 40 may be physically separable from the segment search device 10. In that case, the input device 40 is connected to the segment search device 10 by a short-range wireless communication standard such as Bluetooth (registered trademark). Also, a touch panel display may be configured by arranging a touch panel on the front of the display.
[0023] The display device 50 is a device for displaying a predetermined image (still image and moving image). The display device 50 is configured by a known display device such as a liquid crystal display or an organic EL display. The display device 50 may be an externally attached device to the segment search device 10 as shown in FIG. 1, or may be built into the computer constituting the segment search device 10.
[0024] The sound output device 60 is a device for outputting sound. The sound output device 60 converts an electrical signal into physical vibration (i.e., sound) and outputs it. An example of the sound output device 60 is a general speaker that transmits sound to the wearer by air vibration. The sound output device 60 may be an externally attached device to the segment search device 10 as shown in FIG. 1, or may be built into the computer constituting the segment search device 10. Note that the sound output device 60 may be a bone conduction speaker that transmits sound to the wearer by vibrating the wearer's bone.
[0025] FIG. 2 is a block diagram showing the functional configuration of the system 100. In particular, FIG. 2 shows the functional blocks of the segment search device 10. The segment search device 10 mainly includes a processing unit 11, a storage unit 12, and a communication unit 13. The processing unit 11 is composed of, for example, a processor and a memory. Examples of the processor are well-known CPUs and GPUs. The processor performs predetermined arithmetic processing and image processing according to programs and data stored in the memory, and executes various control processes while writing the results of the processing to the working space of the memory. The memory is composed of, for example, a volatile memory such as RAM (Random Access Memory), and is used for the arithmetic processing by the above-described processor. The storage unit 12 is an element (storage) mainly for storing data used in the arithmetic processing by the processing unit 11. The storage unit 12 is composed of a non-volatile memory such as ROM (Read Only Memory) or flash memory, or an HDD (hard disk drive). Further, a computer program for causing the processing unit 11 to execute a predetermined process may be stored in the storage unit 12. Further, although details will be described later, the storage unit 12 stores contents, segment information, and a learned model as information used in the arithmetic processing by the processing unit 11 or information obtained as a result of the arithmetic processing by the processing unit 11. The communication unit 13 is an element for the segment search device 10 to transmit and receive data to and from an external computer (for example, a web server) via the Internet. The communication unit 13 may be any device that can transmit and receive data by wire or wirelessly. When performing wireless communication, as the communication unit 13, a communication module conforming to a well-known wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) can be adopted.
[0026] Further, as shown in FIG. 2, the processing unit 11 includes functional blocks such as a text conversion unit 11a, a segment division unit 11b, a search sentence creation unit 11c, a search term input unit 11d, a corresponding segment search unit 11e, and a corresponding segment output unit 11f. These functional blocks are functions obtained by the processing unit 11 executing a predetermined program, and are actually mainly realized by the cooperation of a processor and a main memory. For details of these functional blocks, in addition to FIG. 2, reference will be made to the examples described in FIGS. 3 to 5 for explanation.
[0027] As shown in FIG. 3, contents such as videos (including audio), audio, and texts are the processing targets of the segment search device 10. For example, when the content to be processed is a video or audio, the segment search device 10 can acquire these video data or audio data by using the imaging device 20 and / or the sound collection device 30 described above. In addition, the segment search device 10 can also receive (download) contents such as videos, audio, and texts via the Internet by the communication unit 13. The data related to the content is stored in the storage unit 12 of the segment search device 10.
[0028] Next, when the content to be processed is video or audio, the text conversion unit 11a of the segment search device 10 reads the video data or audio data stored in the storage unit 12 and performs speech recognition processing. The speech recognition processing, for example, for the audio data included in the video data or the audio data itself, after performing preprocessing such as noise removal and signal amplification, extracts feature amounts from the audio data. As the feature amounts, for example, known mel spectrum or mel frequency cepstral coefficients (MFCC) may be used. Thereby, the frequency components and temporal patterns of the audio data are extracted. Then, the extracted feature amounts are supplied to a known speech recognition engine or speech recognition model (a learned model of machine learning) to convert the audio data into text data. Note that the speech recognition model can analyze the feature amounts and generate the most likely text data. Also, the text conversion unit 11a preferably performs error conversion correction processing together with the speech recognition processing. The error conversion correction processing can generate text data in which errors based on the context are corrected, for example, by introducing a language model that takes the context into account into the above-described speech recognition model. Thereby, the text conversion unit 11a can accurately convert the audio data into text data. In addition, for speech recognition and error conversion correction, it is not limited to the above-described method, and known methods can be appropriately adopted.
[0029] Note that when the content to be processed is a document, since the document already contains text data, it is not necessary to execute the function of the text conversion unit 11a. However, the text conversion unit 11a may perform error conversion correction on the text data included in the document to correct typographical errors.
[0030] Next, the segment division unit 11b of the segment search device 10 divides the text data generated by the text conversion unit 11a into a plurality of segments (also referred to as clusters). The plurality of segments each contain different topics or arguments. Each segment is basically composed of one topic or argument and forms a group for documents containing topics or arguments in the same direction. The segment division unit 11b may segment the text data for each topic or argument by known natural language processing. For example, the segment division unit 11b can use an unsupervised data classification method such as the k-means method to group text data with similar content into one segment. Also, the segment division unit 11b may create segments for each topic by using a topic model such as LDA (Latent Dirichlet Allocation) that divides the entire text into topics and groups the sentences related to each topic. Further, the segment division unit 11b may use a method such as TF-IDF (Term Frequency-Inverse Document Frequency) to extract important words and phrases from the text data and segment them based on the importance of those words and phrases. Additionally, the segment division unit 11b can use a pre-trained language model such as RNN (Recurrent Neural Networks), a Transformer-based model, or GPT (Generative Pre-trained Transformer) to segment the text data based on topics and arguments considering the context information of the text data. Note that these methods may be combined. Furthermore, it may be possible for the user to input information regarding whether the segmenting of the pre-trained model is correct or incorrect. Thereby, the accuracy of the model can be improved by repeating the cluster division process. FIG. 3 shows a conceptual diagram of segmenting the conversation between speaker A and speaker B. In this example, for the conversation between the two speakers, based on keywords and key phrases, the argument is finely divided each time the content being spoken changes, and it is visualized as segments.
[0031] Next, the search sentence creation unit 11c creates a search sentence for each segment. The search sentence is a sentence created based on the character string of the text data constituting the segment, and is, for example, a conversion, abstraction, summary, or supplementation of the character strings included in the text data. Examples of search sentences are summary sentences and Q&A sets.
[0032] The summary sentence extracts the main information from the text data constituting each segment and summarizes it in a concise form. There are, for example, extractive summary sentences and generative summary sentences. The extractive summary sentence can be created, for example, by extracting important words and phrases in the text using a method such as TF-ID and extracting sentences or passages containing them. Also, the extractive summary sentence may be created by evaluating the importance of sentences in the text using an algorithm such as PageRank, ranking the important sentences, and extracting them. On the other hand, the generative summary sentence can be generated as a new sentence by understanding the meaning of the sentence using a neural network model such as Seq2Seq with an Encoder-Decoder structure. Also, the generative summary sentence may be generated using the Transformer architecture with a self-attention mechanism to identify important parts of the sentence and using them. Note that it is also possible to use automatic evaluation metrics (such as ROUGE and BLEU) to evaluate the quality of the summary sentence.
[0033] The Q&A set is a pair of a question sentence expected for the text data constituting each segment and an answer sentence for the question. The segment search device 10 is connected to the Internet, can access various web pages, and can crawl the contents of various web pages. The search sentence creation unit 11c obtains information from various web pages based on the keywords included in each segment and appropriately stores it in the storage unit 12. Moreover, the search sentence creation unit 11c may automatically create a Q&A set related to the keywords included in each segment. For example, when the keyword is "Tablet A", the search sentence creation unit 11c obtains various information from the attached document regarding "Tablet A" and stores it in the storage unit 12. Then, the search sentence creation unit 11c creates a question sentence "Is there any contraindication in using Tablet A?" and an answer sentence "Yes. Tablet A cannot be prescribed to patients who are prescribed a therapeutic drug for Disease B." Also, the search sentence creation unit 11c can use natural language processing techniques to create a question and an answer for it from the text data included in the segment. For example, the search sentence creation unit 11c can use pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers) and GPT to generate questions and answers from the text data of each segment. In this method, this text data can be used as input to the language model, and a question in a specific format and an optimal answer for it can be output. Also, it is possible to manually create a large number of question-and-answer pairs, use them as teacher data to train a machine learning model (for example, a Seq2Seq or Transformer-based model), and automatically generate an answer for a new question.
[0034] The search sentences (e.g., summary sentences and Q&A sets) created by the search sentence creation unit 11c are stored in the storage unit 12 in association with each segment together with the text data included in the segment and the information (metadata) regarding the segment. In the present specification, the information thus associated with each segment is referred to as segment information. FIG. 4 shows an example of segment information. As shown in FIG. 4, the segment information includes, for example, a segment name, a title of a topic or argument included in the segment, a timeline (time information) of each segment in the content, a summary sentence, an explanatory sentence, and a Q&A set. The segment name is a file name unique to the segment. For example, when the content is divided into eight segments, segment names such as Argument 1 to 8 are automatically assigned in chronological order. The title is a concise representation of the topic or argument included in the segment, and may be created manually, or when the search sentence creation unit 11c creates a summary sentence by the method described above, a title that more concisely represents the summary sentence may be automatically generated. The timeline is information regarding the time of each segment in the content when the content to be processed is time-series data such as a video or audio. When the content to be processed is a text, instead of the timeline, information that can identify the appearance position of each segment in the text, such as the page number, paragraph number, and line number in the text, may be used. The summary sentence is created by the search sentence creation unit 11c by the method described above. The explanatory sentence is created by the text conversion unit 11a by voice recognition processing. When the text data is in a conversation format, the text data may be grouped for each speaker (A or B). The Q&A set is created by the search sentence creation unit 11c by the method described above.
[0035] The segment information including the above information is stored in the form of a database in the storage unit 12 and is used when performing a search for segments. As shown in FIG. 3, as preprocessing performed before executing a search, the segment search device 10 converts each content into text, divides this text data into a plurality of segments, and creates a search sentence (abstract sentence, Q&A set, etc.) for each segment. Thereby, even if the content does not actually include text data as in the case where the content is video or audio, such content can be made a target of character-based search. Further, since the content is divided into a plurality of segments, segments obtained by further finely classifying the content can be made search targets. Furthermore, for each segment of the content, not only the text constituting the segment but also an abstract sentence or Q&A set created based on the text is associated therewith. Therefore, even when the characters input by the user at the time of search do not match the text, since there is a possibility that the characters match the characters in the abstract sentence or Q&A set, each segment can be searched more flexibly. In particular, in the case where the content is a conversation (spoken language) between two or more parties, it is usually difficult to perform a character-based search. However, by creating an abstract sentence or Q&A set (classical language) of this conversation, it becomes easier to search for a desired segment in the conversation by characters.
[0036] Next, with reference to FIG. 5, an example of a method for searching for segments in content will be described. In particular, FIG. 5 shows an example of using generative AI for segment search. Generative AI generally refers to technologies that have the ability to generate new data and information in the fields of machine learning and artificial intelligence. Such generative AI is a trained model that has learned from given data and patterns, and can generate new data and content based on what it has learned. Such a model is trained based on a large amount of data, understands the patterns and features of that data, and learns the parameters for generating new data. A typical example of generative AI is an architecture called GAN (Generative Adversarial Network), in which two networks, a generator and a discriminator, compete to learn and generate data similar to real data. Note that this generative AI (trained model) may be stored in the storage unit 12 of the segment search device 10, or may be stored on a web server and configured such that the segment search device 10 can access the generative AI via the Internet as needed.
[0037] As shown in FIG. 5, the search term input unit 11d of the segment search device 10 provides a UI (user interface) for inputting search terms (questions) for the generative AI. Specifically, the search term input unit 11d displays the UI on the display device 50 and accepts input of search terms from the user via the UI. The search term may be a single word, multiple words, a single sentence, or a passage consisting of multiple sentences. In the example shown in FIG. 5, this search term functions as a prompt for the generative AI.
[0038] Next, the corresponding segment search unit 11e of the segment search device 10 searches for a corresponding segment that corresponds to the search term input by the user from among the segments of the content created in advance. In the example shown in FIG. 5, a secondary model is created by re-educating the generation AI by inputting, as teacher data, the segment information (see FIG. 4) including the search sentence created in advance into the generation AI. Therefore, by inputting a search term (prompt) into this generation AI, an answer considering the segment information created in advance can be obtained from this generation AI. The corresponding segment search unit 11e may search for a corresponding segment corresponding to the search term by using such a generation AI.
[0039] Next, the corresponding segment output unit 11f of the segment search device 10 outputs the information extracted by the corresponding segment search unit 11e. As shown in FIG. 5, the corresponding segment output unit 11f may output, for example, by causing the display device 50 to display the answer of the generation AI. Specifically, according to the generation AI, an answer to the search term is generated. When the search term includes a question regarding the search for segments in the content, the corresponding segments extracted by the search, the access information to the corresponding segments, and the search sentence regarding the corresponding segments are displayed on the display device 50. For example, when the search term is "Please search for the argument regarding tablet A from video content α.", as the answer of the generation AI, the segment information of the segment including the argument regarding tablet A among the plurality of segments constituting the video content α, a part of the video including the argument regarding tablet A, and the time zone thereof are displayed on the display device 50. Also, the access information to the video regarding the argument regarding tablet A may be displayed in a format such as a path specifying the location of a file or folder inside the computer or a URL specifying the location of a file or folder on the Internet. In this way, the user can easily search for a desired segment from arbitrary content. Further, when the reproduction of the video content is instructed by the user, the corresponding segment output unit 11f displays the image of the video content on the display device 50 and outputs the audio of the video content from the sound output device 60. Also, the corresponding segment output unit 11f can transmit (upload) the segment information searched by the corresponding segment search unit 11e to the Web server via the communication unit 13 or individually transmit it to other computers.
[0040] In the example shown in FIG. 5 above, the corresponding segment search unit 11e and the corresponding segment output unit 11f are each configured to search for and output segments using a generation AI. However, the corresponding segment search unit 11e may simply collate the search terms input by the user via the search term input unit 11d with the explanatory texts and search sentences (abstracts, Q&A collections) in the segment information stored in the storage unit 12, and extract segment information containing terms that match the search terms. In this case, the corresponding segment output unit 11f may simply provide the segment information extracted by the corresponding segment search unit 11e to the user, such as by displaying it on the display device 50.
[0041] As described above, in this specification, in order to describe the content of the present invention, embodiments of the present invention have been described with reference to the drawings. However, the present invention is not limited to the above embodiments, and includes obvious modifications and improvements that can be made by those skilled in the art based on the matters described in this specification.
Industrial Applicability
[0042] The present invention relates to a segment search device and a program, and can be used in the information-related industry.
Explanation of Signs
[0043] 10…Segment search device 11…Processing unit 11a…Text conversion unit 11b…Segment division unit 11c…Search sentence creation unit 11d…Search term input unit 11e…Corresponding segment search unit 11f…Corresponding segment output unit 12…Storage unit 13…Communication unit 20…Imaging device 30…Sound collection device 40…Input device 50…Display device 60…Sound playback device 100…System
Claims
1. A segment division unit that divides content, which is video, audio, or text, into segments; A search sentence creation unit that creates a search sentence for each of the segments regarding the segment; A search term input unit where search terms are input; A corresponding segment search unit that collates the search terms and the search sentences, and finds a corresponding segment that is a segment corresponding to the search terms among the segments; A corresponding segment output unit that outputs the corresponding segment, access information of the corresponding segment, or a search sentence regarding the corresponding segment; A segment search device.
2. The segment search device according to Claim 1, wherein the content is video, and the search sentence is an abstract sentence of the segment or a Q&A set of the segment. A segment search device.
3. The segment search device according to Claim 1, wherein the search terms are input terms for a generative AI, and the corresponding segment output unit outputs the corresponding segment, access information of the corresponding segment, or a search sentence regarding the corresponding segment as an answer of the generative AI. A segment search device.
4. A program for causing a computer to function as a segment division unit that divides content, which is video or text, into segments; a search sentence creation unit that creates a search sentence for each of the segments regarding the segment; a search term input unit where search terms are input; a corresponding segment search unit that collates the search terms and the search sentences, and finds a corresponding segment that is a segment corresponding to the search terms among the segments; a corresponding segment output unit that outputs the corresponding segment, access information of the corresponding segment, or a search sentence regarding the corresponding segment; a segment search device.
5. A non-transitory information recording medium storing the program according to Claim 4.
Citation Information
Patent Citations
Moving picture retrieving information generating apparatus, summary information generating apparatus for moving picture, recording medium and summary information storing format for moving picture
JP2004104836A
Hint information description method
JP2008086030A
Video service providing method and service server using the same
JP2019212308A
System for providing knowledge sharing service based on multimedia platform
KR102434880B1
Search document information storing apparatus
JP2020074144A