Search device, search method, and program
Patent Information
- Application Number
- JP2024567413
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-06-18
- Publication Date
- 2025-08-26
AI Technical Summary
Existing technologies for creating digest videos struggle to consistently extract relevant scenes that meet user-specific needs, as they rely on machine learning models that produce the same scenes every time, failing to adapt to different content types and focus areas, and users face confusion when specifying attribute values for precise searches.
A search device and method that allow users to select a search query template based on content type, scene type, person, or organization, which includes predefined attribute values, enabling the selection of appropriate search queries for extracting specific scenes from target video images.
Enables users to set appropriate search queries effectively, ensuring that extracted scenes match their needs for each digest video session by providing a structured approach to specifying attribute values, thus improving the relevance and precision of scene extraction.
Abstract
Description
Search device, search method, and recording medium
[0001] The present invention relates to a search device, a search method, and a program.
[0002] Techniques related to the present invention are disclosed in Patent Documents 1 to 6.
[0003] The technology disclosed in Patent Document 1 extracts audience scenes and important scenes from moving images, and generates a digest video including these scenes.
[0004] The technology disclosed in Patent Document 2 identifies content or a position within the content that matches a specified search query.
[0005] The technology disclosed in Patent Document 3 extracts an image of an object (performer) from a moving image based on the performer's facial features, clothing, etc., and displays the extracted image of the object at a different magnification from the other images.
[0006] The technology disclosed in Patent Document 4 extracts sport play sections from a video based on the posture of a referee in the video.
[0007] The technology disclosed in Patent Document 5 divides a video into a plurality of scenes and searches for scenes that match a search query specified by a user.
[0008] The technology disclosed in Patent Document 6 targets videos of sports, fashion shows, etc., and searches for videos that match search data indicating attributes of subjects specified by the user.
[0009] International Publication No. 2021 / 240678 International Publication No. 2007 / 043679 Japanese Patent Application Laid-Open No. 2021-15417 International Publication No. 2016 / 067553 Japanese Patent Application Laid-Open No. 2009-171624 Japanese Patent Application Laid-Open No. 2001-230994
[0010] There is a demand for a technology that can create a digest video of content (moving images). The technology disclosed in Patent Literature 1 extracts important scenes using a learning model generated by machine learning. In this technology, the extracted scenes depend on the learning model. However, the scenes to be extracted are not necessarily the same each time.
[0011] For example, the scenes to be extracted may differ depending on the type of content, and the characteristics of the scenes to be extracted will differ when creating a digest video of a sporting event and when creating a digest video of a singer's concert.
[0012] Furthermore, the scenes to be extracted may differ depending on what is to be focused on in the digest video. For example, a digest video may be created focusing on a specific person, a specific organization, or a specific scene (such as a scoring scene in a sports game).
[0013] In the case of the technology disclosed in Patent Document 1, which extracts important scenes using a learning model generated by machine learning, similar scenes are extracted each time, making it difficult to extract scenes that meet the needs of each session.
[0014] The problem disclosed in Patent Document 1 can be solved by utilizing the technology disclosed in Patent Document 2, which extracts scenes that match a search query set by a user. However, this technology has the following problems.
[0015] In a search, if a wide variety of attribute values (values indicating scene characteristics) can be specified in a search query, it becomes possible to search for a desired scene with high accuracy. However, if there are many options for specifiable attribute values, a user may be confused about which attribute value to specify and may not be able to set an appropriate search query. None of Patent Documents 1 to 6 discloses such a problem or a solution to the problem.
[0016] One example of the objective of the present invention is to provide a search device, a search method, and a program that solve the problem of enabling users to set appropriate search queries in technology for creating digest videos of content, in consideration of the above-mentioned problems.
[0017] According to one aspect of the present invention, there is provided a search device having: a selection means for selecting a search query template corresponding to a classification specified by a user input from a group of search query templates each having at least one attribute value specified therein, the search query template being classified by at least one of content type, scene type, person, and organization; a setting means for setting the selected search query template as a search query; and a search means for searching for scenes matching the set search query from target videos.
[0018] According to one aspect of the present invention, a search method is provided in which one or more computers select a search query template corresponding to a classification specified by a user input from a group of search query templates, the search query template having at least one attribute value specified therein, classified by at least one of content type, scene type, person, and organization; set the selected search query template as a search query; and search for scenes matching the set search query from target videos.
[0019] According to one aspect of the present invention, a program is provided that causes a computer to function as: a selection means for selecting a search query template corresponding to a classification specified by user input from a group of search query templates that include search query templates, each having at least one specified attribute value, classified by at least one of content type, scene type, person, and organization; a setting means for setting the selected search query template as a search query; and a search means for searching for scenes that match the set search query from target videos.
[0020] According to one aspect of the present invention, a search device, a search method, and a program are realized that solve the problem of enabling a user to set an appropriate search query in a technology for creating a digest video of content.
[0021] The above-mentioned objects and other objects, features and advantages will become more apparent from the following description of the preferred embodiments and the accompanying drawings.
[0022] It is a figure which shows an example of a functional block diagram of the search device. It is a figure which shows an example of a search query template group. It is a figure which shows an example of a search screen. It is a figure which shows an example of a hardware configuration of the search device. It is a flowchart which shows an example of a processing flow of the search device.
[0023] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0024] 1 is a functional block diagram showing an overview of a search device 10 according to a first embodiment. The search device 10 includes a selection unit 11, a setting unit 12, and a search unit 13.
[0025] The selection unit 11 selects a search query template from a group of search query templates that corresponds to a classification specified by a user input. The search query template group includes search query templates, each of which specifies at least one attribute value, for a classification defined by at least one of content type, scene type, person, and organization. The setting unit 12 sets the selected search query template as a search query. The search unit 13 searches for scenes that match the set search query from target videos (content). The searched scenes are presented to the user as candidates for scenes to be included in a digest video of the content.
[0026] In this way, the search device 10 searches the target video for scenes that match the search query set by the user. The searched scenes become candidates for scenes to be used in the digest video of the target video. With this search device 10, the user can obtain search results that meet the needs of each episode by appropriately setting a search query that meets the needs of each episode when creating a digest video for each episode.
[0027] The search device 10 also stores search query templates for categories defined by at least one of content type, scene type, person, and organization. The search device 10 then sets, as a search query, a search query template corresponding to a category specified by user input. With this search device 10, a user can set an appropriate search query by inputting a "category" that corresponds to the type of content or the focus of the current digest video. In other words, the user does not need to manually specify the attribute values to be included in the search query. This search device 10 solves the problem of enabling a user to set an appropriate search query.
[0028] Second Embodiment Overview A search device 10 according to a second embodiment is a specific implementation of the search device 10 according to the first embodiment.
[0029] The search device 10 stores in advance a group of search query templates as shown in Fig. 2. The search query template group includes search query templates, each of which specifies at least one attribute value, for each category defined by at least one of content type, scene type, person, and organization. In the information shown in Fig. 2, the definition of each category and the search query template are registered in association with category identification information that distinguishes the categories from one another.
[0030] The search device 10 accepts a user input specifying at least one of content type, scene type, person, and organization on a search screen such as that shown in FIG. 3 . The user input specifies a category. The search device 10 then selects a search query template corresponding to the category specified by the user input from among a group of search query templates such as that shown in FIG. 2 . Next, the search device 10 displays the selected search query template in a search query field, as shown in FIG. 3 . Thereafter, in response to a search execution instruction input by the user, the search device 10 executes a search based on the search query displayed in the search query field. The configuration of the search device 10 will be described in detail below.
[0031] "Hardware Configuration" An example of the hardware configuration of the search device 10 will be described. Each functional unit of the search device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various variations in the realization method and device. Software includes programs that are pre-loaded when the device is shipped, and programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet.
[0032] FIG. 4 is a block diagram illustrating an example of the hardware configuration of the search device 10. As shown in FIG. 4, the search device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The search device 10 does not necessarily have to have the peripheral circuit 4A. Note that the search device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration.
[0033] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to mutually transmit and receive data. The processor 1A is, for example, a processing unit such as a CPU or a graphics processing unit (GPU). The memory 2A is, for example, a random access memory (RAM) or a read-only memory (ROM). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes an interface for connecting to a communication network such as the Internet. Examples of input devices include a keyboard, mouse, microphone, physical buttons, touch panel, etc. Examples of output devices include a display, speaker, printer, mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0034] "Functional Configuration" Next, the functional configuration of the search device 10 of this embodiment will be described in detail. Fig. 1 shows an example of a functional block diagram of the search device 10. As shown in the figure, the search device 10 has a selection unit 11, a setting unit 12, and a search unit 13.
[0035] The selection unit 11 selects a search query template corresponding to a category designated by a user input from among a group of search query templates.
[0036] A "search query template" is a template of a search query that is generated in advance and stored in the search device 10. In the search query template, an attribute value of at least one attribute is specified.
[0037] "Attributes" are attributes of objects that appear in the scene to be searched. The objects are people, objects, etc. The scene to be searched is a candidate for a scene to be included in the digest video.
[0038] Examples of person attributes include, but are not limited to, gender, age, nationality, physique, height, weight, hairstyle, clothing characteristics, possessions characteristics, posture, movement, facial expression, and voice (content of speech, voice volume, voiceprint, etc.). These attributes may be further subdivided. For example, clothing characteristics may be subdivided into hat characteristics, eyeglass characteristics, jacket characteristics, shoe characteristics, and a number attached to the clothing.
[0039] Examples of object attributes include, but are not limited to, the type of object (e.g., soccer goal, soccer ball, baseball, etc.) and the external characteristics of the object (e.g., color, shape, size, etc.).
[0040] The "attribute value" indicates the specific content of the attribute as described above. For example, the attribute value of "attribute: gender" is "male" or "female." A plurality of selectable attribute values may be set in advance for each attribute.
[0041] A "search query template group" is a collection of search query templates as described above. The search device 10 stores the search query template group in advance. FIG. 2 shows an example of the search query template group stored in the search device 10. As shown in FIG. 2, the search query template group includes search query templates for classifications defined by at least one of content type, scene type, person, and organization.
[0042] In the example shown in FIG. 2, category identification information, category definitions, and search query templates are registered in association with one another.
[0043] "Classification identification information" is information that distinguishes multiple classifications from one another.
[0044] "Classification definition" indicates the definition of each classification. For example, the classification identified by "classification identification information: 000001" shown in FIG. 2 corresponds to "Type of content: baseball." The classification identified by "classification identification information: 000093" shown in FIG. 2 corresponds to "Person: Tokyo Taro." The classification identified by "classification identification information: 001118" shown in FIG. 2 corresponds to "Type of scene: scoring scene (soccer)."
[0045] In Fig. 2, one classification is defined by one content (value). Although not shown, a classification that combines multiple contents (values) may be defined, such as a classification corresponding to "Person: Tokyo Taro" and "Scene type: Good fielding (baseball)."
[0046] The search query template group is generated in advance by a system administrator and stored in the search device 10. The system administrator can update the content of the search query template group at any time. Alternatively, each user may be able to edit the search query template group. For example, each user may be able to register a new search query template, change the content of a registered search query template, or delete a registered search query template.
[0047] In the "user input specifying a classification," at least one of the type of content, the type of scene, a person, and an organization is specified.
[0048] For example, the selection unit 11 can accept a user input specifying at least one of a content type, a scene type, a person, and an organization on a search screen such as that shown in Fig. 3. The search screen shown in Fig. 3 displays a UI (user interface) component that allows the user to specify at least one of a content type, a scene type, a person, and an organization.
[0049] The search screen shown in Fig. 3 allows one content type to be specified. In other words, for example, "content type: baseball" and "content type: soccer" cannot be specified simultaneously on the search screen shown in Fig. 3. However, the search screen may be configured so that multiple content types can be specified simultaneously.
[0050] The search screen shown in Fig. 3 allows multiple people, groups, and scenes to be simultaneously specified. For example, the search screen shown in Fig. 3 allows "People: Taro Tokyo" and "People: Ichiro Tanaka" to be simultaneously specified. However, it may be configured so that only one of these can be specified.
[0051] In addition, the search screen shown in Fig. 3 allows multiple content types, people, organizations, and scene types to be specified at the same time. That is, on the search screen shown in Fig. 3, for example, "People: Tokyo Taro" and "Scene Type: Good Defense" can be specified at the same time. However, it may be configured so that only one of them can be specified.
[0052] For example, when a user input specifying a category is received from the search screen shown in Fig. 3, the selection unit 11 refers to the category definition column shown in Fig. 2 and identifies one or more categories that include the specified content (value). Then, the selection unit 11 selects a search query template corresponding to the identified one or more categories.
[0053] For example, when "Content type: Baseball," "Person: Tokyo Taro," and "Scene type: Good fielding" are specified as in the search screen shown in FIG. 3, the selection unit 11 can specify a classification that includes each of them. In the case of FIG. 2, the selection unit 11 specifies "classification identification information: 000001" corresponding to "Content type: Baseball." The selection unit 11 also specifies "classification identification information: 000093" corresponding to "Person: Tokyo Taro." The selection unit 11 also specifies "classification identification information: 001117" corresponding to "Content type: Baseball" and "Scene type: Good fielding."
[0054] The user input may be realized via an input device included in the search device 10. Alternatively, the search device 10 may be a server. The user input may be performed via a client terminal. In this case, data indicating the content of the user input is transmitted from the client terminal to the search device 10.
[0055] 1 , the setting unit 12 sets the search query template selected in response to a user input as a search query. The setting unit 12 displays the set search query in the search query field, as shown in Fig. 3 . The setting unit 12 may also accept a user input to modify the search query template displayed in the search query field.
[0056] When the selection unit 11 selects one search query template corresponding to one classification, the setting unit 12 sets the one search query template as a search query.
[0057] When the selection unit 11 selects multiple search query templates corresponding to multiple classifications, the setting unit 12 generates a condition by connecting the multiple search query templates with any logical operator such as an AND condition or an OR condition, and sets it as a search query.
[0058] The search unit 13 searches for scenes matching the set search query from among the target video. The search unit 13 searches for scenes in which an object having an attribute value specified in the search query appears. A scene may be composed of one frame image or multiple frame images. The search unit 13 may output, as a search result, a frame image in which an object having an attribute value specified in the search query appears. The search unit 13 may also output, as a search result, a video composed of a frame image in which an object having an attribute value specified in the search query appears and a predetermined number of consecutive frame images (design matters) before and / or after the frame image. The target video is a video that is the target of processing to create a digest video. The user specifies the target video. The search by the search unit 13 is realized using any technology.
[0059] Next, a specific example of a search query template will be described.
[0060] - Search query templates prepared for classifications defined by content type - The content type is defined by at least one of the following: type of sports event, type of television content, and type of Internet content.
[0061] The types of sports include, but are not limited to, soccer, basketball, volleyball, baseball, swimming, etc.
[0062] The types of television content include, but are not limited to, dramas, fashion shows, comedy shows, and the like.
[0063] Internet content includes, but is not limited to, comedy, outdoor activities, gambling, etc.
[0064] The search query template group can include a search query template for each type of content to create a digest of that type of content. The scenes to be included in the digest video are determined to some extent for each type of content. For example, goal scenes in soccer, scoring scenes in baseball, and emotional scenes in dramas (scenes in which actors are crying, laughing, angry, etc.). Therefore, for each type of content, a search query template specifying attribute values specific to typical scenes to be included in the digest video of each content can be generated and included in the search query template group.
[0065] A search query template for soccer specifies attribute values such as age, gender, clothing (top and bottom), hairstyle, face, voice, shoes, national flag, commentary (volume and content of remarks), referee's posture, referee's voice (volume and content of remarks), uniform number, socks, ball, posture (kicking posture, throw-in posture, etc.), captain's bandana, etc.
[0066] A search query template for basketball specifies attribute values such as age, gender, clothing (top and bottom), hairstyle, face, voice, shoes, national flag, commentary (volume and content of remarks), referee's posture, referee's voice (volume and content of remarks), uniform number, posture (shooting posture, blocking posture, etc.), and ball.
[0067] A search query template for volleyball specifies attribute values such as age, gender, clothing (top and bottom), hairstyle, face, voice, shoes, national flag, commentary (volume and content), referee's posture, referee's voice (volume and content), uniform number, voice (volume and content), posture (spiking posture, receiving posture, serving posture, etc.), player formation, and ball.
[0068] A search query template for baseball specifies attribute values such as age, gender, clothing (top and bottom), hairstyle, face, voice, shoes, national flag, commentary (volume and content of remarks), umpire's posture, umpire's voice (volume and content of remarks), uniform number, socks, ball, bat, base, posture (pitching posture, swing posture, etc.).
[0069] A search query template for swimming specifies attribute values such as age, gender, costume (swimsuit, cap, goggles, etc.), hairstyle, face, voice, national flag, commentary (volume and content of remarks), referee's posture, referee's voice (volume and content of remarks), lane number, and swimming style.
[0070] In a search query template corresponding to a drama, attribute values such as age, sex, clothing (top and bottom), hairstyle, facial expression, voice, glasses, hat, car, etc. are specified.
[0071] In a search query template corresponding to a fashion show, attribute values such as age, sex, outfit (top and bottom), hairstyle, facial expression, voice, glasses, hat, and signature pose are specified.
[0072] In a search query template for comedy, attribute values such as age, gender, costume (top and bottom), hairstyle, face, voice (person), voice (audience), glasses, hat, facial expression, and posture are specified.
[0073] The search query template may be able to specify the content of the audio of the content as an attribute value. For example, in a search query template corresponding to soccer, a phrase such as "goal" may be specified as an attribute value. In addition, in a search query template corresponding to baseball, a phrase such as "home run" may be specified as an attribute value. When such an attribute value specifying the content of the audio of the content is included in the search query, the search unit 13 analyzes the audio of the content (target video) and searches for scenes in which the phrase is spoken.
[0074] --Search query templates prepared for classifications defined by scene type-- Scene types are defined for each content. For example, in baseball, they include, but are not limited to, scoring scenes, good fielding scenes, and base stealing scenes. In dramas, they include, but are not limited to, scenes in which characters laugh, cry, or shout.
[0075] The search query template group may include a search query template for each type of scene, specifying attribute values specific to that type of scene, such as the posture of a person (such as a player, referee, or other character), facial expression, possessions of the person, and voice volume.
[0076] The search query template may be able to specify the content of the audio of the content as an attribute value. For example, in a search query template corresponding to a soccer goal scene, a phrase such as "goal" may be specified as an attribute value. In addition, in a search query template corresponding to a baseball goal scene, a phrase such as "home run" may be specified as an attribute value. When such an attribute value specifying the content of the audio of the content is included in the search query, the search unit 13 analyzes the audio of the content (target video) and searches for scenes in which the phrase is spoken.
[0077] - Search query templates prepared for classifications defined by people - The search query template group can include a search query template for each person that specifies attribute values specific to that person. The search query template specifies attribute values such as the person's gender, age, nationality, physique, height, weight, hairstyle, clothing characteristics, possessions characteristics, posture, movement, and voice.
[0078] The search query template may specify the content of the audio in the content as an attribute value. For example, a search query template corresponding to a specific person may specify the name or nickname of the person as an attribute value. When such an attribute value specifying the content of the audio is included in the search query, the search unit 13 analyzes the audio in the content (target video) and searches for scenes in which the phrase is spoken.
[0079] - Search query templates prepared for classifications defined by organizations - Organizations are, for example, sports teams. The search query template group can include a search query template for each organization that specifies attribute values specific to that organization. The search query template specifies attribute values such as the characteristics of the organization's uniform, the organization's logo or mark, and the organization's team colors.
[0080] The search query template may specify the content of the audio of the content as an attribute value. For example, a search query template corresponding to a specific organization may specify the name or nickname of the organization as an attribute value. When such an attribute value specifying the content of the audio of the content is included in the search query, the search unit 13 analyzes the audio of the content (target video) and searches for scenes in which the phrase is spoken.
[0081] Next, an example of the processing flow of the search device 10 will be described with reference to the flowchart of FIG.
[0082] The search device 10 accepts a user input specifying a category (S10) on a search screen such as that shown in Fig. 3. The user input specifying a category specifies at least one of a type of content, a type of scene, a person, and an organization.
[0083] Next, the search device 10 selects a search query template corresponding to the classification specified in S10 from the search query template group (S11). The search query template group includes search query templates, each of which specifies at least one attribute value, for a classification defined by at least one of content type, scene type, person, and organization.
[0084] Next, the search device 10 sets the search query template selected in S11 as a search query (S12). The search device 10 displays the set search query in a search query field, as shown in Fig. 3. The search device 10 may accept an input to modify the search query displayed in the search query field, and modify the search query displayed in the search query field in accordance with the input.
[0085] Thereafter, the search device 10 performs a search based on the set search query (S13). For example, in response to a user's input of a search start instruction, the search device 10 performs a search based on the search query that is set at that time. In this search, the search device 10 searches for scenes that match the set search query from among the target videos.
[0086] The search device 10 can output search results. The search results may display a list of images of the searched scenes. The search results may be output via a display provided in the search device 10. Alternatively, the search device 10 may be a server. The search device 10 may then transmit the search results to a client terminal. In this case, the search results are displayed on the display of the client terminal.
[0087] <Effects> According to the search device 10 of this embodiment, the user can set an appropriate examination query by inputting a designation of a predetermined category.
[0088] For example, the user may input a content type of the target video as an input for specifying a predetermined classification, and the search device 10 selects a search query template for creating a digest of the specified content in response to the input and sets the selected search query as the search query.
[0089] Furthermore, the user can input an input specifying an object to be focused on in the digest video as an input specifying a predetermined classification. The object to be focused on can be, for example, a person, an organization, a predetermined scene (such as a sports scoring scene), etc. In response to the input, the search device 10 selects a search query template in which attribute values specific to the specified person are specified and sets it as a search query. In response to the input, the search device 10 selects a search query template in which attribute values specific to the specified organization are specified and sets it as a search query. In response to the input, the search device 10 selects a search query template in which attribute values specific to a specified scene are specified and sets it as a search query.
[0090] In this way, according to the search device 10 of this embodiment, the user can set an appropriate examination query by performing a relatively simple input such as specifying a predetermined category.
[0091] <Third Embodiment> In a third embodiment, a search query template group includes a search query template for a zoomed-in scene and a search query template for a zoomed-out scene for each type of content, each type of scene, each person, or each organization.
[0092] Attribute values (feature amounts of people and objects) that can be detected from an image may differ between a zoomed-in scene and a zoomed-out scene. For example, in a zoomed-in scene, attribute values related to relatively detailed features such as a person's facial expression, hairstyle, and uniform number can be detected. On the other hand, in a zoomed-out scene, it is difficult to detect attribute values related to such relatively detailed features. However, even in a zoomed-out scene, there are detectable attribute values such as a person's hair color, clothing color, and posture. In consideration of this point, the third embodiment provides separate search query templates suitable for zoomed-in scenes and zoomed-out scenes.
[0093] In the third embodiment, the selection unit 11 accepts a user input specifying at least one of a content type, a scene type, a person, and an organization, as well as a user input specifying a zoomed-in scene or a zoomed-out scene, and selects a search query template corresponding to the category specified by the user input.
[0094] Other configurations of the search device 10 of the third embodiment are similar to those of the search device 10 of the first and second embodiments.
[0095] The search device 10 of the third embodiment achieves the same effects as the search devices 10 of the first and second embodiments. Furthermore, the search device 10 of the third embodiment allows the user to appropriately set a search query for searching for a zoomed-in scene and a search query for searching for a zoomed-out scene.
[0096] <Modification> The search query template may be able to specify changes in the volume of content as attribute values. For example, a user may wish to include exciting scenes in a digest video. Exciting scenes tend to have louder volumes. Therefore, the search query template may specify, for example, "the volume has increased by more than a threshold value compared to the reference volume" as an attribute value. When such an attribute value specifying changes in the volume of content is included in the search query, the search unit 13 analyzes the volume of the content (target video image) and searches for scenes where the volume has increased by more than a threshold value compared to the reference volume. The reference volume may be, for example, a statistical value (average, median, mode, minimum, etc.) of the volume of the entire content. Alternatively, the reference volume may be, for example, a statistical value (average, median, mode, minimum, etc.) of the volume of multiple sample points selected from the content according to an arbitrary rule.
[0097] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations may be adopted. The configurations of the above-described embodiments may be combined with each other, or some of the configurations may be replaced with other configurations. Furthermore, various modifications may be made to the configurations of the above-described embodiments without departing from the spirit of the invention. Furthermore, the configurations and processes disclosed in the above-described embodiments and modified examples may be combined with each other.
[0098] In addition, in the flowcharts used in the above explanation, multiple steps (processes) are described in order. However, the order of execution of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the figures can be changed to the extent that the content is not affected. Furthermore, each of the above-mentioned embodiments can be combined to the extent that the content is not contradictory.
[0099] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes: 1. A search device comprising: a selection means for selecting a search query template corresponding to a classification specified by a user input from a search query template group including search query templates, each search query template having at least one attribute value specified therein, the search query template being classified by at least one of content type, scene type, person, and organization; a setting means for setting the selected search query template as a search query; and a search means for searching for scenes matching the set search query from target videos. 2. The search device according to 1, wherein the user input specifying the classification specifies at least one of the content type, scene type, person, and organization. 3. The search device according to 2, wherein the content type is defined by at least one of sporting event type, television content type, and Internet content type, and the search query template group includes, for each content type, a search query template for creating a digest of the content of each type. 4. The search device according to 2 or 3, wherein the search query template group includes, for each scene type, a search query template specifying an attribute value specific to the scene of each type. 5. The search device according to any one of 2 to 4, wherein the search query template group comprises, for each person, the search query template specifying the attribute value specific to each person. 6. The search device according to any one of 2 to 5, wherein the search query template group comprises, for each organization, the search query template specifying the attribute value specific to each organization. 7. The search device according to any one of 2 to 6, wherein the search query template group comprises, for each type of content, each type of scene, each person, or each organization, the search query template for a zoomed-in scene and the search query template for a zoomed-out scene, and the user input specifies the zoomed-in scene or the zoomed-out scene.8. The search device according to any one of 1 to 7, wherein the setting means displays the search query template selected in response to the user input in a search query field, and accepts input to modify the search query template displayed in the search query field. 9. A search method, in which one or more computers select the search query template corresponding to a classification specified by a user input from a search query template group including search query templates, each including at least one attribute value specified, classified by a classification defined by at least one of content type, scene type, person, and organization, set the selected search query template as a search query, and search for scenes matching the set search query from target videos. 10. A program that causes a computer to function as: a selection means that selects the search query template corresponding to the classification specified by a user input from a search query template group including search query templates, each including at least one attribute value specified, classified by a classification defined by at least one of content type, scene type, person, and organization; a setting means that sets the selected search query template as a search query; and a search means that searches for scenes matching the set search query from target videos.
[0100] This application claims priority based on Japanese Patent Application No. 2022-210275, filed December 27, 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0101] 10 Search device 11 Selection unit 12 Setting unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus
Claims
1. a selection means for selecting a search query template corresponding to a classification specified by a user input from a group of search query templates each having at least one attribute value specified therein, the search query template being classified by at least one of content type, scene type, person, and organization; a setting means for setting the selected search query template as a search query; a search means for searching for scenes matching the set search query from among target videos; A search device having the above configuration.
2. The search device according to claim 1 , wherein the user input specifying the classification specifies at least one of the type of content, the type of scene, a person, and an organization.
3. the content type is defined by at least one of a sporting event type, a television content type, and an internet content type; The search device according to claim 2 , wherein the group of search query templates includes, for each type of content, a search query template for creating a digest of the content of each type.
4. The search device according to claim 2 , wherein the group of search query templates includes, for each type of scene, a search query template that specifies the attribute value specific to each type of scene.
5. The search device according to claim 2 , wherein the group of search query templates includes, for each person, the search query template that specifies the attribute value specific to each person.
6. The search device according to claim 2 , wherein the group of search query templates includes, for each organization, a search query template that specifies the attribute value specific to each organization.
7. the search query template group includes, for each type of content, each type of scene, each person, or each organization, the search query template for a zoomed-in scene and the search query template for a zoomed-out scene; 4. The search device according to claim 2, wherein the user input specifies a zoomed-in scene or a zoomed-out scene.
8. The setting means displaying the search query template selected in response to the user input in a search query field; The search device according to claim 1 , wherein an input for modifying the search query template displayed in the search query field is accepted.
9. One or more computers selecting a search query template corresponding to a classification specified by a user input from a group of search query templates each having at least one attribute value specified therein, the search query template being classified by at least one of content type, scene type, person, and organization; setting the selected search query template as a search query; A search method for searching for scenes that match the set search query from among target videos.
10. Computer, a selection means for selecting a search query template corresponding to a classification specified by a user input from a group of search query templates each having at least one attribute value specified therein, the search query template being classified by at least one of content type, scene type, person, and organization; a setting means for setting the selected search query template as a search query; a search means for searching for scenes matching the set search query from among target videos; A program that functions as a