Video search device, video search method, and program

The video search device and method improve search accuracy by generating and updating explanatory information based on user feedback, addressing the issues of insufficient and inaccurate video database information.

JP7798114B2Active Publication Date: 2026-01-14NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023563366
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2026-01-14
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

Existing video search systems struggle with inaccurate search results due to insufficient video information and low information accuracy in databases, leading to poor search accuracy.

Method used

A video search device and method that generates explanatory information for each video, performs searches using this information, accepts user feedback on search results, and updates the information based on the feedback to improve accuracy.

Benefits of technology

Enhances search accuracy by generating and updating explanatory information, allowing for precise video retrieval even when information is insufficient or inaccurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798114000001
    Figure 0007798114000001
  • Figure 0007798114000002
    Figure 0007798114000002
  • Figure 0007798114000003
    Figure 0007798114000003
Patent Text Reader

Abstract

In order to solve the problem of improving video search accuracy even in a case where the accuracy or amount of information concerning video is not sufficient, a video search device (1) comprises: a generation unit (11) that generates explanation information for each video image stored in a video storage device; an acquisition unit (12) that acquires a search query; a search unit (13) that searches the video storage device for a video image by using the search query and explanation information; an output unit (14) that outputs a search result by the search unit (13); an input unit (15) that receives input of a determination result of a user with respect to the search result; and an update unit (16) that updates the explanation information on the basis of the determination result and the search query.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for searching for video. [Background technology]

[0002] Patent Document 1 describes a video search system that searches a video database based on input search criteria. This video search system allows a user to select and classify videos similar to a target video from a set of videos obtained by the search, and extracts video information related to the classified videos from the video database. This video search system also determines feature quantities related to the target video using the extracted video information and classification information, and re-searches the video database using the determined feature quantities. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-331009 Summary of the Invention [Problem to be solved by the invention]

[0004] In the video search system described in Patent Document 1, if the video database does not store sufficient video information, the video information related to the classified video cannot be extracted sufficiently. Furthermore, if the accuracy of the video information stored in the video database is insufficient, the accuracy of the video information extracted related to the classified video will also be insufficient. Therefore, it is not possible to accurately determine the feature amounts related to the target video, and there is a possibility that the search accuracy cannot be improved.

[0005] One aspect of the present invention has been made in consideration of the above-mentioned problems, and one example of its purpose is to provide a technology that improves the accuracy of video search even when the amount or accuracy of information about the video is insufficient. [Means for solving the problem]

[0006] A video search device according to one aspect of the present invention comprises a generation means for generating explanatory information for each video stored in a video storage device, an acquisition means for acquiring a search query, a search means for searching for video from the video storage device using the search query and the explanatory information, an output means for outputting the search results by the search means, an input means for accepting input of a user's judgment result on the search results, and an update means for updating the explanatory information based on the judgment result and the search query.

[0007] A video search system according to one aspect of the present invention comprises a generation means for generating explanatory information for each video stored in a video storage device, an acquisition means for acquiring a search query, a search means for searching for video from the video storage device using the search query and the explanatory information, an output means for outputting the search results by the search means, an input means for accepting input of a user's judgment result on the search results, and an update means for updating the explanatory information based on the judgment result and the search query.

[0008] A video search method according to one aspect of the present invention generates explanatory information for each video stored in a video storage device, obtains a search query, searches for video from the video storage device using the search query and the explanatory information, outputs search results, accepts input of a user's judgment result on the search results, and updates the explanatory information based on the judgment result and the search query.

[0009] A program according to one aspect of the present invention is a program for causing a computer to function as a video search device, and causes the computer to function as a generation means for generating explanatory information for each video stored in a video storage device, an acquisition means for acquiring a search query, a search means for searching for video from the video storage device using the search query and the explanatory information, an output means for outputting the search results by the search means, an input means for accepting input of a user's judgment result on the search results, and an update means for updating the explanatory information based on the judgment result and the search query. [Effects of the Invention]

[0010] According to one aspect of the present invention, it is possible to improve the accuracy of video search even when the amount or accuracy of information related to the video is insufficient. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing the configuration of a video search device according to a first exemplary embodiment of the present invention. [Figure 2] 1 is a flowchart showing the flow of a video search method according to a first exemplary embodiment of the present invention. [Figure 3] 1 is a block diagram showing the configuration of a video search system according to a first exemplary embodiment of the present invention. [Figure 4] FIG. 10 is a block diagram showing the configuration of a video search system according to a second exemplary embodiment of the present invention. [Figure 5] 10 is a schematic diagram illustrating details of a moving image and sensor information according to the second exemplary embodiment of the present invention. FIG. [Figure 6] 10 is a flowchart showing the flow of a video search method according to a second exemplary embodiment of the present invention. [Figure 7] FIG. 10 is a diagram showing an example of explanatory information according to the second exemplary embodiment of the present invention. [Figure 8] FIG. 10 is a schematic diagram showing a specific example of a video search method according to a second exemplary embodiment of the present invention. [Figure 9] FIG. 10 is a schematic diagram showing another specific example of a video search method according to exemplary embodiment 2 of the present invention. [Figure 10] FIG. 10 is a schematic diagram showing yet another specific example of a video search method according to exemplary embodiment 2 of the present invention. [Figure 11] FIG. 10 is a block diagram showing the configuration of a video search system according to a third exemplary embodiment of the present invention. [Figure 12] FIG. 10 is a flow chart showing the flow of a video search method according to a third exemplary embodiment of the present invention. [Figure 13]FIG. 1 is a diagram illustrating an example of a hardware configuration of a video search device according to each exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] Exemplary Embodiment 1 A first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form of the exemplary embodiments described below.

[0013] <Configuration of video search device 1> The configuration of a video search device 1 according to this exemplary embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the video search device 1.

[0014] As shown in FIG. 1, the video search device 1 includes a generation unit 11, an acquisition unit 12, a search unit 13, an output unit 14, an input unit 15, and an update unit 16. The generation unit 11 is an example of a configuration that realizes the generation means described in the claims. The acquisition unit 12 is an example of a configuration that realizes the acquisition means described in the claims. The search unit 13 is an example of a configuration that realizes the search means described in the claims. The output unit 14 is an example of a configuration that realizes the output means described in the claims. The input unit 15 is an example of a configuration that realizes the input means described in the claims. The update unit 16 is an example of a configuration that realizes the update means described in the claims.

[0015] The generation unit 11 generates description information for each video stored in the video storage device. The acquisition unit 12 acquires a search query. The search unit 13 searches for video from the video storage device using the search query and description information. The output unit 14 outputs the search results by the search unit 13. The input unit 15 accepts input of the user's judgment result on the search results. The update unit 16 updates the description information based on the judgment result and the search query. The "description information," "search query," and "judgment result" will be explained in detail in the flow of the video search method S1 described below.

[0016] <Video search method S1 flow> The video search device 1 executes a video search method S1 according to this exemplary embodiment. The flow of the video search method S1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the video search method S1. As shown in Fig. 2, the video search method S1 includes steps S11 to S16.

[0017] (Step S11) In step S11, the generation unit 11 generates explanation information for each video stored in the video storage device.

[0018] Here, the video storage device is a device that stores multiple videos to be searched. The video retrieval device 1 is communicably connected to the video storage device via a network, for example. The videos to be searched may be still images or videos. In the case of videos, the units to be searched may be video segments obtained by dividing the videos along the time axis. The video storage device may be provided in the video retrieval device 1 as a video storage unit.

[0019] The explanatory information is information that explains each video to be searched. The explanatory information may be, for example, a key-value pair or a natural language sentence. However, the expression format of the explanatory information is not limited to this. For example, the generation unit 11 analyzes each video and generates explanatory information based on the analysis results. For example, the generation unit 11 may acquire explanatory text entered by a user for each video and generate explanatory information based on the acquired explanatory text. In this case, the explanatory text entered by the user is acquired via an input device or a network. The generation unit 11 stores the generated explanatory information in memory in association with the video. Since the generation unit 11 generates explanatory information for each of multiple videos, the generation unit 11 generates multiple pieces of explanatory information.

[0020] (Step S12) In step S12, the acquisition unit 12 acquires a search query.

[0021] The search query includes information for identifying a desired video. Specifically, the search query is a query for searching for explanatory information. The search query may be, for example, a set of a key and a value, or may be a natural language sentence. However, the expression format of the search query is not limited to this.

[0022] In this step, the acquisition unit 12 may acquire the search query entered by the user via an input device or a network, or may acquire the search query by reading it from a memory. The acquisition unit 12 may also acquire a search query generated by another device or another functional block (not shown).

[0023] (Step S13) In step S13, the search unit 13 searches for a video from the video storage device using the search query and the description information.

[0024] For example, the search unit 13 extracts explanatory information that at least partially matches the search query from among the multiple pieces of explanatory information generated by the generation unit 11. The search unit 13 also sets the video linked to the extracted explanatory information as a search result. The number of videos obtained as a search result by the search unit 13 may be one or more. The number of videos obtained as a search result is multiple when the search unit 13 extracts multiple pieces of explanatory information that at least partially match the search query. In this case, the search unit 13 sets the video linked to each of the extracted multiple pieces of explanatory information as a search result.

[0025] (Step S14) In step S14, the output unit 14 outputs the search results by the search unit 13. The search results include one or more videos. Here, the output unit 14 may output the search results by the search unit 13 by transmitting them to the user's terminal device. In this case, the terminal device displays the received search results on a display connected to the terminal device. The output unit 14 may also display the search results by the search unit 13 on a display connected to the video search device 1. By outputting the search results in this manner, the output unit 14 can present the search results to the user.

[0026] (Step S15) In step S15, the input unit 15 receives an input of the user's judgment result on the search results.

[0027] The determination result is the result of the user's determination as to whether each video included in the search result is the desired video. As a specific example, the input unit 15 displays a user interface component that allows the user to select "appropriate (the desired video)" or "inappropriate (not the desired video)" near each video displayed as a search result. The user interface component may be displayed on a display connected to the video retrieval device 1 or on the user's terminal device. For example, when the search result is displayed on the user's terminal device, the input unit 15 transmits information indicating the user interface component to the terminal device, thereby displaying the user interface component near each video. The input unit 15 also accepts an input of the determination result of the video in response to a user's selection operation on the user interface component. For example, the user's selection operation may be performed using an input device connected to the video retrieval device 1, or may be performed on the user's terminal device. When the user interface component is displayed on the user's terminal device, the terminal device accepts the user's selection operation on the user interface component and transmits information indicating the selection operation to the video retrieval device 1. The input unit 15 receives information indicating a selection operation from the terminal device, thereby accepting the input of the determination result. However, the method of accepting the input of the determination result is not limited to this specific example.

[0028] The determination result is not limited to "whether the video is the target video" but may also indicate "the degree of match with the target video." In this case, the input unit 15 may display a user interface component that allows the selection of three or more options, or any numerical value included in a predetermined range (for example, from 1 to 100).

[0029] (Step S16) In step S16, the update unit 16 updates the description information based on the determination result and the search query. For example, the update unit 16 updates the part of the description information that partially matches the search query but does not match the search query, according to the determination result. For example, if the determination result for the video related to the description information is "appropriate," the update unit 16 updates the part of the description information that does not match the search query so that it matches the search query.

[0030] <Advantages of this exemplary embodiment> As described above, the video search device 1 and video search method S1 according to this exemplary embodiment are configured to generate description information for each video stored in a video storage device, obtain a search query, use the search query and description information to search for one or more videos from the video storage device, output search results, accept input of a user's judgment result on the search results, and update the description information based on the judgment result and search query.

[0031] According to this configuration, the generation unit 11 generates explanatory information about the video and performs a search using the generated explanatory information, so that searches can be performed with high accuracy even when the amount or accuracy of information previously associated with the video is insufficient. Furthermore, according to this configuration, the explanatory information can be updated with high accuracy based on user feedback on search results. As a result, searches can be performed using the updated explanatory information, improving search accuracy. Thus, according to this configuration, a technology can be provided that improves the accuracy of video searches even when the amount or accuracy of information about the video is insufficient.

[0032] Other Aspects of the Exemplary Embodiment Another aspect of this exemplary embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of a video retrieval system 10 according to another aspect. As shown in Fig. 3, the video retrieval system 10 includes a generation unit 11, an acquisition unit 12, a search unit 13, an output unit 14, an input unit 15, and an update unit 16. The video retrieval system 10 includes a plurality of physically different devices, and one or more of these units are distributed across the plurality of devices. The configuration and operation of each unit are as described above in detail.

[0033] Exemplary Embodiment 2 A second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are given the same reference numerals, and their description will be omitted as appropriate.

[0034] <Configuration of video search system 20> The configuration of the video search system 20 according to this exemplary embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing the configuration of the video search system 20.

[0035] 4, the video retrieval system 20 includes a video retrieval device 2 and a video storage device 9. The video retrieval device 2 includes a control unit 210, a storage unit 220, an input / output unit 230, and a communication unit 240.

[0036] (Video Memory Device 9) The video storage device 9 stores one or more moving images and one or more types of sensor information. The moving images and sensor information will be described with reference to Fig. 5. Fig. 5 is a schematic diagram for explaining the details of the moving images and sensor information.

[0037] Videos are captured by a camera mounted on a moving object. Examples of moving objects and camera devices include automobiles and dashcams. However, the moving objects and camera devices are not limited to these. As shown in FIG. 5, a moving object ID is associated with each video. The moving object ID identifies the moving object equipped with the camera that captured the video. Furthermore, time information at which each frame was captured is associated with each frame constituting the video. Furthermore, a video is composed of multiple video segments divided along a time axis. A video segment includes multiple frames. The time length of each video segment is, for example, 10 to 20 seconds, but is not limited to this. The video segments constituting a video are an example of "video" as defined in the claims, and are the units to be searched.

[0038] Sensor information is information acquired by a sensor mounted on a mobile object. Examples of sensors include a vehicle speed sensor, a steering angle sensor, an engine rotation speed sensor, and a positioning sensor. The "time-series data of vehicle speed" shown in FIG. 5 is an example of sensor information acquired by a vehicle speed sensor. Furthermore, the "time-series data of location information" is an example of sensor information acquired by a positioning sensor. However, the types of sensors and sensor information are not limited to these. Furthermore, a mobile object ID is linked to the sensor information. The mobile object ID identifies the mobile object equipped with the sensor that acquired the sensor information. Furthermore, the sensor information is linked to time information at which the sensor information was acquired.

[0039] Furthermore, as shown in Fig. 5, each video segment is linked to sensor information. Video segments and sensor information can be linked using the mobile object ID and time information associated with each. For example, a certain video segment is linked to a video segment that has the same mobile object ID and is the time-series data of sensor information acquired from the start to the end of filming of the video segment.

[0040] (Storage unit 220) The storage unit 220 stores the generative model, the explanation information, and the search query.

[0041] A generative model is a model generated to receive at least a video as input and output explanatory information. Generative models include machine learning models and rule-based models.

[0042] The machine learning model is, for example, a model generated using training data so as to receive at least a video segment as input and output explanatory information. Examples of the machine learning model include, but are not limited to, a support vector machine, a decision tree, a random forest, and a neural network model. The machine learning model may be generated by a generation unit 21 (described later) or an external device. Note that the input to the machine learning model may include sensor information linked to the video segment in addition to or instead of the video segment itself.

[0043] The rule-based model may include, for example, one or more rules. Each rule may include a condition related to the sensor information and explanatory information to be used when the condition is met. Each rule may include a condition related to information obtained by analyzing the video segments in addition to or instead of the condition related to the sensor information. The information obtained by analyzing the video segments may be, for example, but is not limited to, the type and color of the subject.

[0044] The description information is generated and stored by a generating unit 21, which will be described later. The search query is acquired and stored by an acquiring unit 22, which will be described later. The description information and the search query will be described in detail later.

[0045] (Input / output section 230) The input / output unit 230 controls input and output to and from the video search device 2. The input / output unit 230 includes, for example, a keyboard, a mouse, a touchpad, a display, and the like.

[0046] (Communication unit 240) The communication unit 240 connects to a network and controls communication with the video storage device 9. The network to be connected may be, for example, a wireless local area network (LAN), a wired LAN, the Internet, a mobile data communication network, or a combination of these.

[0047] (control unit 210) The control unit 210 controls the storage unit 220, the input / output unit 230, and the communication unit 240 to control the overall operation of the video search device 2. The control unit 210 includes a generation unit 21, an acquisition unit 22, a search unit 23, an output unit 24, an input unit 25, and an update unit 26. The acquisition unit 22, the output unit 24, and the input unit 25 are configured in the same manner as the acquisition unit 12, the output unit 14, and the input unit 15 in exemplary embodiment 1, and therefore detailed description thereof will not be repeated.

[0048] The generation unit 21 generates explanatory information using a generative model. The generation unit 21 also generates explanatory information using video segments and sensor information. The search unit 23 searches the video storage device 9 for video segments whose explanatory information at least partially matches the search query. The update unit 26 updates the parts of the explanatory information related to the searched video segments that do not match the search query according to the determination result. Details of "searching for partially matching video segments" and "updating the non-matching parts" will be explained in the flow of the video search method S2 described below.

[0049] <Video search method S2 flow> The video search device 2 configured as above executes a video search method S2 according to this exemplary embodiment. The flow of the video search method S2 will be described with reference to Fig. 6. Fig. 6 is a flow diagram showing the flow of the video search method S2. As shown in Fig. 6, the video search method S2 includes steps S21 to S26.

[0050] (Step S21) In step S21, the generation unit 21 uses the video segments and sensor information to generate explanatory information for each video segment using a generative model. Specifically, the generation unit 21 inputs the video segments to a machine learning model. The generation unit 21 also inputs the sensor information linked to the video segments to a rule-based model. The generation unit 21 then stores the explanatory information output from the machine learning model and the rule-based model in the storage unit 220, linking them to the video segments.

[0051] Here, a specific example of the explanatory information generated in step S21 will be described with reference to FIG. 7. FIG. 7 is a diagram illustrating a specific example of the explanatory information. In this specific example, the explanatory information is expressed as a set of a key and a value. Note that the explanatory information may include a key whose value is null. In the example of FIG. 7, for example, the value of the key "status" included in the road information is null. Hereinafter, the set of key "x" and value "y" will also be referred to as the value "y" of key "x", the value "y" held by key "x", etc.

[0052] Examples of types of keys that may be included in the explanatory information include (i) "own vehicle information," (ii) "traffic participant information (individual)," (iii) traffic participant information (collective)," (iv) "own vehicle / other vehicle relative information," (v) "road information," (vi) "event information," and (vii) "meta information."

[0053] (i) "Own vehicle information" includes keys related to the own vehicle itself, such as "vehicle type," "lane type," and "action." Note that "own vehicle" refers to a moving object equipped with an imaging device that captured the video including the video segment. The "vehicle type" key indicates the attributes of the own vehicle, and in this example, its value is "standard vehicle." The "lane type" key indicates one of the driving states of the own vehicle during video segment capture, and in this example, its value is "passing lane." Other examples of keys that indicate the driving state of the own vehicle include the "position," "speed," and "acceleration" keys (not shown). The "action" key indicates one of the actions of the own vehicle during video segment capture, and in this example, its value is "brake operation." Other examples of values ​​that the "action" key can take include the values ​​"steering (right or left turn)," "merging / merging / lane change," and "overtaking / passing," (not shown).

[0054] (ii) "Traffic participant information (single)" includes keys such as "driver" and "type" related to each traffic participant during the video segment capture. A traffic participant is a person, object, or vehicle participating in traffic, inside or outside the vehicle. The value of the key "driver" is "female" in this example. The key "type" indicates the type of traffic participant other than the driver, and in this example, its value is "motorcycle." Examples of other values ​​that the key "type" can take include "other vehicle," "motorcycle," "bicycle," "pedestrian," and "animal."

[0055] (iii) Transportation participant information (gathering) The "traffic participant information (set)" includes keys such as "centroid" and "range" that relate to multiple traffic participants during the capture of a video segment. The key "centroid" indicates the centroid of the positions of multiple traffic participants, and in this example, its value is null. The key "range" indicates the range that includes multiple traffic participants, and in this example, its value is null.

[0056] (iv) "Relative information between own vehicle and other vehicles" The "own vehicle / other vehicle relative information" includes keys such as "relative distance" and "relative motion" that indicate the relationship between the own vehicle and other vehicles during video segment capture. The "relative distance" key indicates the relative distance between the own vehicle and other vehicles, and in this example, its value is null. The "relative motion" key indicates the relative motion between the own vehicle and other vehicles, and in this example, its value is "approach." Other examples of keys that indicate the relationship between the own vehicle and other vehicles include keys such as "relative speed" and "relative acceleration," not shown.

[0057] (v) “Road information” "Road Information" includes keys such as "Shape," "Area," and "Status" that relate to the road on which the vehicle traveled during the capture of the video segment. The "Shape" key indicates the shape of the road, and in this example, its value is "Branch." Other examples of values ​​that the "Shape" key can take include "Lane Increase / Decrease," "Merge," and "Intersection." The "Area" key indicates the area in which the road is located, and in this example, its value is "Tunnel." Other examples of values ​​that the "Area" key can take include "No Lane Change," "Zebra Zone," "Safety Zone," "Parking Lot," "Expressway," "Urban Area," and "Place Name." The "Status" key indicates the condition of the road, and in this example, its value is null. Other examples of values ​​that the "Status" key can take include weather indicators such as "Rainfall" and "Snowfall," as well as "Paved."

[0058] (vi) "Event Information" "Event Information" includes keys such as "Near Miss" and "Traffic Jam" that relate to events that occurred during the filming of a video segment. The key "Near Miss" indicates whether a so-called near miss occurred, and in this example, its value is "Yes." The key "Traffic Jam" indicates whether a traffic jam occurred, and in this example, its value is "Yes." Other examples of keys that may be included in "Event Information" include "Accident," "Construction," "Visibility," "Visibility (Fog, Backlight, Heavy Rain)," and "Hit-and-Run Accident."

[0059] (vii) "Meta Information" etc. "Meta information" includes keys such as "motion blur" and "likely to appear in a commercial" that indicate meta information for the video segment. These keys indicate the characteristics of the video segment as an image, regardless of the traffic conditions depicted in the video segment. In this example, the value of the key "motion blur" is "none." In addition, the value of the key "likely to appear in a commercial" is null.

[0060] Although FIG. 7 shows an example in which one key has one value, one key may have multiple values. In other words, the explanatory information may include a combination of one key and multiple values. For example, in FIG. 7, the key "action" (hereinafter also referred to as "own vehicle action") included in the category "own vehicle information" may have multiple values, such as "brake operation" and "left turn." Furthermore, a value corresponding to one key may be expressed as a range value. For example, the value of the key "speed" (hereinafter also referred to as "vehicle speed") not shown included in the category "own vehicle information" may be "10 to 15 km / h." Here, "X to Y" represents a range from X to Y, and "km / h" represents kilometers per hour.

[0061] (Step S22) In step S22, the acquisition unit 22 acquires a search query. The operation of this step is substantially the same as the operation of step S12 described in the first exemplary embodiment. However, the search query acquired in this step includes one or more queries. When the description information is expressed as a set of keys and values ​​shown in FIG. 7, each query included in the search query is expressed as a set of keys and values. In other words, the search query includes multiple sets of keys and values. Hereinafter, "keys and values ​​representing each query included in the search query" will also be referred to as "keys and values ​​specified in the search query (or query)", etc.

[0062] (Step S23) In step S23, the search unit 23 searches the video storage device 9 for video segments whose explanatory information at least partially matches the search query. For example, when the search query includes multiple queries, the search unit 23 extracts explanatory information that satisfies at least some of the queries from the storage unit 220. The search unit 23 also sets video segments linked to the extracted explanatory information as search results. For example, assume that the search query includes a first query and a second query. The first query is represented by a pair of a first key and a first value, and the second query is represented by a pair of a second key and a second value. In this case, the search unit 23 extracts, from the explanatory information stored in the storage unit 220, (i) explanatory information that matches at least the first query (including the pair of the first key and the first value) and (ii) explanatory information that matches at least the second query (including the pair of the second key and the second value). The explanatory information (i) includes information that matches the second query and information that does not match the second query. The explanatory information that matches the first query but does not match the second query does not completely match the search query, but only partially matches it. The explanatory information (ii) includes information that matches the first query and information that does not match the first query. The explanatory information that matches the second query but does not match the first query does not completely match the search query, but only partially matches it. Note that when the explanatory information includes a key that is not specified in the search query (a key other than the first key and the second key), the search unit 23 extracts the key, assuming that any value is acceptable for the key.

[0063] Here, specific examples will be used to explain how to determine whether or not the explanatory information matches each query included in a search query. The first specific example relates to a query that specifies a key having only one value (for example, "car model"). For example, such a query is expressed as a set of the key "car model" and the value "standard car." In this case, if the key "car model" has the value "standard car" in the explanatory information, the explanatory information matches the query. On the other hand, if the key "car model" has the value "light car" in the explanatory information, the explanatory information does not match the query.

[0064] A second specific example relates to a query that specifies a key (for example, "subject vehicle action") that may have multiple values. For example, such a query is expressed as a set of the key "subject vehicle action" and the value "brake action." In this case, if the key "subject vehicle action" in the explanatory information has multiple values, "brake action" and "left turn," the explanatory information matches the query. On the other hand, if the key "subject vehicle action" in the explanatory information has multiple values, "accelerate" and "left turn," the explanatory information does not match the query. In other words, if the key specified in the query in the explanatory information has at least the value specified in the query, the explanatory information matches the query. Note that a query may also be expressed as a set of one key and multiple values. In this case, if the key specified in the query in the explanatory information has at least all the values ​​specified in the query, the explanatory information may match the query, and may not match in other cases. Alternatively, if the key specified in the query in the explanatory information has at least one of the multiple values ​​specified in the query, the explanatory information may match the query. In this case, if the key specified in the query in the description information does not have any of the multiple values ​​specified in the query, the description information may not match the query.

[0065] A third specific example relates to a query that specifies a key whose value is expressed as a range (for example, "vehicle speed"). For example, such a query is expressed as a set of the key "vehicle speed" and the value "10-30 km / h." In this case, if the key "vehicle speed" has the value "10-15 km / h" in the explanatory information, the explanatory information matches the query. Also, if the key "vehicle speed" has the value "40-50 km / h" in the explanatory information, the explanatory information does not match the query. In other words, if the range value indicated by the value of the key specified in the query (hereinafter also referred to as the range value of the explanatory information) is included in the range value specified in the query, the explanatory information matches the query. Also, if there is no overlap between the range value of the explanatory information and the range value specified in the query, the explanatory information does not match the query. Note that the range value of the explanatory information may include both overlapping and non-overlapping parts with respect to the range value specified in the query. For example, if the range value of the description information is "0 to 15 km / h" and the range value specified in the query is "10 to 40 km / h", such description information may or may not be considered to match.

[0066] The determination of whether the description information matches each query included in the search query is not limited to the specific example described above. Furthermore, the matching conditions used in such a determination may be optionally specified by the user.

[0067] (Step S24) In step S24, the output unit 24 outputs the search results from the search unit 23. The operation of this step is almost the same as the operation of step S14 described in exemplary embodiment 1. However, the difference is that the search results are output in units of video segments.

[0068] (Step S25) In step S25, the input unit 25 accepts input of the user's judgment result for the search results. The operation of this step is almost the same as the operation of step S15 described in exemplary embodiment 1. However, it differs in that the unit for accepting the input of the judgment result is a video segment.

[0069] (Step S26) In step S26, the update unit 26 updates the parts of the description information related to the searched video segments that do not match the search query in accordance with the determination result. A specific example of the update process in this step will be described with reference to Figs. 8 to 10.

[0070] (Example 1) Fig. 8 is a schematic diagram illustrating a specific example 1 of the video search method S2. As shown in Fig. 8, in this specific example, the search query acquired in step S22 includes "the value of the first key 'shape' is 'confluence'" and "the value of the second key 'state' is 'snowfall'".

[0071] In the description information extracted in step S23, the value of the first key "state" is "confluence", but the value of the second key "state" is null. Therefore, this description information satisfies the search query for the first key but does not satisfy the search query for the second key, and therefore partially matches the search query.

[0072] In step S24, the video segment associated with this explanatory information is displayed on the display. In step S25, the result of the determination received indicates "appropriate."

[0073] In this case, in step S26, the update unit 26 updates the value of the second key "state" in the description information that does not match the search query to "snowfall" so that it matches the search query.

[0074] In this way, when the update unit 26 obtains a judgment result indicating that the video segment is appropriate, it updates the value of the key in the description information that does not match the search query so that it matches the search query.

[0075] (Example 2) 9 is a schematic diagram illustrating Example 2 of the video search method S2. As shown in FIG. 9, the search query acquired in step S22 of this example is the same as that of Example 1.

[0076] The description information extracted in step S23 has a value of "confluence" for the first key "state" but does not include the second key. Therefore, this description information satisfies the first query but not the second query, and therefore partially matches the search query.

[0077] In step S24, the video segment associated with such explanatory information is displayed on the display. Also, the result of the determination received in step S25 indicates "appropriate."

[0078] In this case, in step S26, the update unit 26 adds a second key "state" to the explanation information, and updates its value to "snowfall" so that it matches the search query.

[0079] In this way, when the update unit 26 obtains a judgment result indicating that the video segment is appropriate, it adds a new key to the description information that is not included in the search query and updates its value to match the search query.

[0080] (Example 3) 10 is a schematic diagram illustrating Example 3 of the video search method S2. As shown in FIG. 10, the search query acquired in step S22 of this example is the same as in Examples 1 and 2.

[0081] In the description information extracted in step S23, the value of the first key "state" is "confluence", but the value of the second key "state" is null. Therefore, this description information satisfies the search query for the first key but does not satisfy the search query for the second key, and therefore partially matches the search query.

[0082] In step S24, the video segment associated with such explanatory information is displayed on the display. Also, the determination result received in step S25 indicates "inappropriate."

[0083] In this case, in step S26, the update unit 26 updates the value of the second key "state" in the description information that does not match the search query to "not snowfall" so as to negate the search query.

[0084] In this way, when a determination result indicating that the video segment is inappropriate is obtained, the update unit 26 updates the value of the key in the description information that does not match the search query to negate the search query. In this case, when extracting description information that satisfies at least a part of the search query from the storage unit 220, the search unit 23 does not extract description information that includes information that negates the search query.

[0085] (If it matches the search query exactly) In addition, in step S26, if the description information completely matches the search query and the judgment result is ``inappropriate,'' the update unit 26 may update at least a portion of the description information that matches the search query so that it no longer matches.

[0086] <Advantages of this exemplary embodiment> As described above, the video storage device 9 referenced by the video retrieval device 2 and video retrieval method S2 according to this exemplary embodiment stores video images captured by a camera mounted on a mobile object and sensor information acquired by a sensor mounted on the mobile object. The sensor information is linked to video segments obtained by dividing the video images along the time axis. Furthermore, the video retrieval device 2 and video retrieval method S2 employ a configuration similar to that of the exemplary embodiment, in which explanatory information is generated using a generative model generated to input video segments and sensor information and output explanatory information.

[0087] According to this configuration, since the explanatory information is generated using a generative model, the explanatory information can be generated with high accuracy. Furthermore, since the explanatory information is generated using sensor information in addition to video segments, the explanatory information can be generated with high accuracy. Therefore, this exemplary embodiment can more accurately search for video segments using accurately generated explanatory information even when there is no or insufficient information previously associated with a video.

[0088] Furthermore, according to the video search device 2 and the video search method S2, in addition to the same configuration as the exemplary embodiment, a configuration is adopted in which a video whose explanatory information partially matches the search query is searched from the video storage device 9, and the part of the explanatory information related to the searched video that does not match the search query is updated according to the determination result.

[0089] According to this configuration, it is possible to accurately update the portion of the description information related to the searched video that does not match the search query.

[0090] Other Aspects of the Exemplary Embodiment Other aspects 1 to 8 that are modifications of this exemplary embodiment will be described below.

[0091] (Aspect 1) In the first mode, priority is given to searching for a desired video segment. In the first mode, the output unit 24 and step S24 are modified as follows.

[0092] In step S24, if the search results include a plurality of video segments, the output unit 24 outputs the search results in descending order of search accuracy by the search unit 23.

[0093] Here, specific examples of high search accuracy will be described. As a first specific example, high search accuracy may mean that the reliability of the portion of the explanatory information that matches the search query is high. As such reliability, the reliability output from the machine learning model together with the explanatory information can be used. For example, the generation unit 21 associates the explanatory information and reliability output from the machine learning model with the video segments and stores them in the storage unit 220. In this case, the output unit 24 outputs the video segments in descending order of reliability associated with the portion of the explanatory information that matches the search query.

[0094] As a second specific example, high search precision may mean that the description information matches the search query in many parts. For example, if the search query includes three queries, the search precision is higher in the following order: description information that matches all three queries, description information that matches two queries but not one, and description information that matches one query but not two.

[0095] As a third specific example, high search accuracy may mean that the weight of a matching query is high. In this case, it is assumed that weights are assigned to multiple queries included in the search query. The weights may be specified by the user. Furthermore, the weights may be specified in advance or may be specified together with the search query. For example, suppose a search query includes two queries, one specifying the key "subject vehicle operation" and the other specifying the key "vehicle speed," and the key "subject vehicle operation" has a higher weight than the key "vehicle speed." In this case, the search accuracy is higher in the order of explanatory information that at least matches the key "subject vehicle operation," and explanatory information that does not match the key "subject vehicle operation" but matches the key "vehicle speed."

[0096] The "output order" may be realized, for example, by the arrangement order on the display, or by a chronological order. For example, the output unit 24 displays the multiple video segments included in the search results on the display in a predetermined direction (e.g., from top to bottom) in descending order of search accuracy. The output unit 24 also displays a predetermined number of video segments on the display in descending order of search accuracy, and upon receiving a determination result for those segments, repeats this process of displaying a predetermined number of video segments with the next highest search accuracy on the display. However, the methods for realizing the "output order" are not limited to these.

[0097] According to the configuration of aspect 1, search results are output in descending order of search accuracy, and the video segments are presented to the user in the order in which they are output. This allows the user to recognize the video segments in descending order of search accuracy, and provides the advantage of making it easier for the user to find the desired video segment.

[0098] (Aspect 2) In the second aspect, priority is given to improving the accuracy of the explanatory information. In the second aspect, the output unit 24 and step S24 are modified as follows.

[0099] In step S24, if the search results include a plurality of video segments, the output unit 24 outputs the search results in descending order of search accuracy by the search unit 23.

[0100] Here, specific examples of low search accuracy will be described. As a first specific example, for example, the degree to which description information matches the search query may be low. For example, if the search query includes three queries, the search accuracy can be said to be lower when only one matches, when only two matches, and when all three match. In this case, the output unit 24 outputs video segments in descending order of the degree to which description information matches the search query.

[0101] As a second specific example, low search precision may mean that the description information has few matches with the search query. For example, if the search query includes three queries, the search precision decreases in descending order from description information that matches one query but not two, description information that matches two queries but not one, and description information that matches all three queries.

[0102] As a third specific example, low search accuracy may mean that the weight of the matching query is small. The weight is as explained in the third specific example of high search accuracy. For example, suppose a search query includes two queries, one specifying the key "subject vehicle operation" and the other specifying the key "vehicle speed", and the key "vehicle speed" has a smaller weight than the key "subject vehicle operation". In this case, the search accuracy decreases in order from description information that at least matches the key "vehicle speed" to description information that does not match the key "vehicle speed" but matches the key "subject vehicle operation".

[0103] As a fourth specific example, low search accuracy may mean that the description information contains a large number of null values. In this case, the output unit 24 outputs video segments in descending order of the number of null values ​​in the description information.

[0104] Note that a specific example of the order in which the information is output to the user is the same as in the first embodiment, and therefore a detailed description thereof will be omitted.

[0105] Here, in step S25, the user may not input the judgment results for all of the video segments included in the search results, but may input the judgment results for some of the video segments that are output earlier. This tendency is thought to be particularly strong when the search results include a large number of video segments.

[0106] Therefore, according to the configuration of aspect 2, search results are output in descending order of search accuracy, and the video segments are presented to the user in the order in which they are output. As a result, the user recognizes the video segments in descending order of search accuracy, and it is expected that the earlier the segments are recognized, the more likely the user will input a judgment result. As a result, it is possible to accept more judgment results for video segments with lower search accuracy, and it is possible to update the description information more accurately.

[0107] (Aspect 3) Aspect 3 is an aspect in which it is possible to switch between Aspect 1 and Aspect 2 as modes. In Aspect 3, the video retrieval device 2 is modified so as to receive an input from the user as to which mode to select. The video retrieval device 2 operates as Aspect 1 or Aspect 2 according to the mode selected by the user.

[0108] According to the configuration of aspect 3, the user can enjoy the advantage of being able to switch between prioritizing searching for a desired video segment and prioritizing improving the accuracy of the explanatory information depending on the situation.

[0109] (Aspect 4) In the fourth aspect, the search results are classified. In the fourth aspect, the output unit 24 and step S24, and the input unit 25 and step S25 are modified as follows.

[0110] In step S24, if the search results include multiple video segments, the output unit 24 classifies and outputs the search results. For example, the output unit 24 may classify the multiple video segments according to description information. For example, the multiple video segments included in the search results may be classified according to the value of the key "area." In this case, the key used for classification may be a key included in the search query or a key not included. Alternatively, the output unit 24 may classify the multiple video segments according to the video characteristics of the video segments (e.g., type of subject, color, etc.). The output unit 24 may also classify the multiple video segments using a classification model. In this case, the classification model is generated using machine learning to input the video segments and output their classification. The classification model may be stored in the storage unit 220 of the video retrieval device 2 or in an external device. If stored in an external device, the video retrieval device 2 uses the classification model by communicating with the external device. The classification model may be generated by a functional block (not shown) of the video retrieval device 2 or by another device.

[0111] Note that one method of "classifying and outputting" is to divide the display area of ​​the display into multiple areas and associate the areas with the classifications. Another method is to generate a different screen for each classification and switch between the screens. Note that the "classifying and outputting" method is not limited to these.

[0112] In step S25, the input unit 25 accepts input of the determination result for each classification. For example, if the video segments are classified into a plurality of areas and displayed, the input unit 25 may display a user interface component for accepting the determination result for each area and accept input operations for each user interface component. However, the method of "accepting input of the determination result for each classification" is not limited to this.

[0113] According to the configuration of the fourth aspect, the user does not need to input the judgment result for each video segment included in the search result individually, but can input the judgment results for each category all at once. Therefore, it is possible to receive the judgment results for a larger number of video segments, and to update the description information more accurately.

[0114] (Aspect 5) Aspect 5 is an aspect in which a plurality of determination results are used. In aspect 5, the input unit 25 and step S25, and the update unit 26 and step S26 are modified as follows.

[0115] In step S25, the input unit 25 accepts input of multiple judgment results for the search results. For example, the video retrieval device 2 may repeat steps S24 to S25 to re-output the video segment for which the judgment result has been accepted and accept the judgment result again. In this case, one user inputs multiple judgment results. Also, for example, the video retrieval device 2 may output the search results to multiple terminals in step S24 and accept input of judgment results from the multiple terminals in step S25. In this case, multiple users each input a judgment result.

[0116] In step S26, the update unit 26 updates the explanation information using the multiple determination results. For example, the update unit 26 may use the most common determination result among the multiple determination results. As a specific example, if three of five determination results indicate "appropriate" and two indicate "inappropriate," the update unit 26 updates the explanation information by adopting the "appropriate" determination result that is most common. The update unit 26 may also weight each of the multiple determination results. For example, when steps S24 to S25 are repeated to receive input of multiple determination results, the weight may be increased as the input of the determination result is received most recently.

[0117] For example, when multiple determination results are received from a single user, the user may be unsure whether the output video segment is the desired one and may change the determination result each time it is input. Also, when determination results are received from multiple users, the determination result of one user may differ from the determination result of another user. According to the configuration of aspect 5, since multiple determination results are used, the description information can be updated more accurately than when a single determination result is used.

[0118] (Aspect 6) In the sixth aspect, the same interpolation is applied to similar video segments. In the sixth aspect, the update unit 26 and step S26 are modified as follows.

[0119] As described above, time information and location information are linked to each video segment of the video stored in the video storage device 9. This linking is possible by comparing the timestamp attached to each frame of the video with the time-series data of the location information included in the sensor information.

[0120] In step S26, the update unit 26 extracts other video segments from the videos stored in the video storage device 9, the other video segments having similar time information and / or position information to the video segment whose description information is to be updated. The update unit 26 also updates the description information related to the other extracted videos. More specifically, the update unit 26 updates the description information related to the other extracted videos in the same manner as the description information to be updated.

[0121] Here, as described above, the video segment for which the description information is to be updated is, for example, a video segment whose description information at least partially matches the search query.

[0122] For example, in the specific example of updating explanatory information shown in FIG. 8, the value of the key "status" is updated from a null value to "snowfall." In this specific example, in this aspect, the update unit 26 extracts other video segments whose temporal distance and spatial distance to the video segment linked to the explanatory information are within a threshold value. The extracted other video segments are, for example, video segments captured by other moving objects that were traveling around the moving object at the time the video segment was captured. The update unit 26 then updates the value of the key "status" to "snowfall" for the explanatory information linked to the extracted other video segments.

[0123] Note that each video segment may be associated with either time information or position information, not necessarily both.

[0124] Furthermore, in step S26, the update unit 26 may extract other video segments that have similar time information and location information as well as similar driving directions. For example, even when driving on the same road at a similar time, the explanatory information to be added to the video may differ depending on whether the driving direction is uphill or downhill. By adding the driving direction condition, other video segments whose explanatory information is to be updated in a similar manner can be extracted with greater accuracy.

[0125] Specifically, in step S26, the update unit 26 identifies the traveling direction of the moving object when the video segment for which the description information is to be updated was captured. For example, the update unit 26 can identify the traveling direction by using time-series data of location information linked to the video segment. The other video segments to be extracted are, for example, video segments captured by other moving objects traveling in the same direction (uphill or downhill) on the same road as the moving object when the video segment was captured.

[0126] According to the configuration of the sixth aspect, the description information for other video segments for which the user's judgment results have not been accepted can be updated in accordance with the user's judgment results for a certain video segment. This allows the description information for more video segments to be updated with higher accuracy.

[0127] (Aspect 7) The seventh aspect is an aspect that takes into consideration the dependency between the explanation information. In the seventh aspect, the update unit 26 and step S26 are modified as follows.

[0128] In this embodiment, the explanatory information includes first explanatory information and second explanatory information. The first explanatory information and the second explanatory information have a dependency relationship. Information relating to such a dependency relationship is stored in the storage unit 220. For example, in the explanatory information described with reference to FIG. 7, an example of the first explanatory information is the key "area." Furthermore, an example of the second explanatory information is the key "status." For example, if the value of the key "area" is "tunnel," the value of the key "status" cannot be "rainfall" or "snowfall." In other words, there is a dependency relationship between the key "area" and the key "status."

[0129] In step S26, the update unit 26 updates the explanatory information using the dependency between the first explanatory information and the second explanatory information.

[0130] For example, in the specific example of updating explanatory information shown in Figure 8, the value of the key "status" was updated from a null value to "snowfall". In this specific example, if the value of the key "area" was "tunnel", the update unit 26 would not update the value of the key "status" to "snowfall", taking into account the dependency between the key "area" and the key "status".

[0131] According to the configuration of the seventh aspect, the explanatory information is updated in consideration of the dependency between the first explanatory information and the second explanatory information, so that the explanatory information can be updated with higher accuracy.

[0132] (Aspect 8) In the eighth aspect, the types of description information to be updated are limited. In the fourth aspect, the update unit 26 and step S26 are modified as follows.

[0133] In this aspect, the explanatory information includes third explanatory information and fourth explanatory information. The generation unit 21 generates the third explanatory information using a rule-based model. The generation unit 21 generates the fourth explanatory information based on a machine learning model or user input. The rule-based model and the machine learning model are stored in the storage unit 220, and details thereof are as described above. The generation unit 21 may also acquire explanatory text entered by a user for each video, and generate the fourth explanatory information based on the acquired explanatory text. Details of generating explanatory information based on explanatory text entered by a user are as described in the first exemplary embodiment. The storage unit 220 stores information indicating whether the explanatory information is the third explanatory information or the fourth explanatory information, depending on the type of explanatory information (e.g., key).

[0134] In step S26, the update unit 26 does not update the third explanation information, but updates the fourth explanation information.

[0135] Here, since the third explanation information is derived based on a rule-based model, it is likely to be highly objective and clearly defined. Therefore, it can be said that the third explanation information is information with relatively high accuracy. Since the fourth explanation information is derived based on a machine learning model or user input, it may be difficult to clearly define or may have low objectivity. Therefore, it can be said that the fourth explanation information is information with room for improvement in accuracy through feedback of the determination result.

[0136] According to the configuration of the eighth aspect, the highly accurate third explanation information is not updated, but the fourth explanation information, which has room for improvement in accuracy, is updated, so that the explanation information can be updated with higher accuracy.

[0137] Exemplary Embodiment 3 A third exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the second exemplary embodiment are given the same reference numerals, and their description will be omitted as appropriate.

[0138] <Configuration of video search system 30> The configuration of the video search system 30 according to this exemplary embodiment will be described with reference to Fig. 11. Fig. 11 is a block diagram showing the configuration of the video search system 30.

[0139] 11, the video retrieval system 30 includes a video retrieval device 3 and a video storage device 9. The video retrieval device 3 includes a control unit 310, a storage unit 320, an input / output unit 330, and a communication unit 340. The video storage device 9 is as described in exemplary embodiment 2. Furthermore, the storage unit 320, the input / output unit 330, and the communication unit 340 are similar to the storage unit 220, the input / output unit 230, and the communication unit 240 described in exemplary embodiment 2, and therefore detailed description thereof will not be repeated.

[0140] 11, the control unit 310 includes a generation unit 31, an acquisition unit 32, a search unit 33, an output unit 34, an input unit 35, an update unit 36, and a model update unit 37. Here, the configuration of the model update unit 37 will be described. The other functional blocks are configured in the same manner as in the second exemplary embodiment, and therefore detailed description thereof will not be repeated.

[0141] The model update unit 37 updates the generative model using the explanation information updated by the update unit 36. Details of updating the generative model will be explained later in the flow of the video search method S3.

[0142] <Video search method S3 flow> The video retrieval device 3 configured as above executes a video retrieval method S3 according to this exemplary embodiment. The flow of the video retrieval method S3 will be described with reference to Fig. 12. Fig. 12 is a flow diagram showing the flow of the video retrieval method S3. As shown in Fig. 12, the video retrieval method S3 includes steps S31 to S37. The operations of steps S31 to S36 are the same as the operations of steps S21 to S26 described as exemplary embodiment 2. Here, the operation of step S37 will be described.

[0143] (Step S37) In step S37, the model update unit 37 updates the generative model using the explanation information updated in step S36.

[0144] For example, the model update unit 37 performs additional learning on the machine learning model included in the generative model using the updated explanatory information as training data. As a specific example, a case will be described in which the value of the key "status" is updated from a null value to "snowfall" as described with reference to FIG. 8. In this case, the model update unit 37 performs additional learning on the machine learning model so that it outputs a set of the key "status" and the value "snowfall" when a corresponding video segment is input.

[0145] <Advantages of this exemplary embodiment> The video retrieval device 3 and the video retrieval method S3 according to this exemplary embodiment employ a configuration in which the explanation information updated by the update unit 36 ​​is used to update the generative model.

[0146] According to this configuration, the generative model is updated to output explanatory information that matches the user's judgment results, so that searches can be performed using explanatory information generated using the updated generative model, thereby improving search accuracy.

[0147] [Modification] Each of the exemplary embodiments 2 and 3 can be modified as follows.

[0148] In each exemplary embodiment, the video storage device 9 may store still images, and the still images may be searched. In this case, the still images are an example of the video described in the claims. Alternatively, the video storage device 9 may store moving images, and the moving images may be searched in file units rather than in video segment units. In this case, the moving image files are an example of the video described in the claims.

[0149] In each exemplary embodiment, the generative model may include only one of a machine learning model and a rule-based model, but is not limited to both.

[0150] In each exemplary embodiment, the generators 21 and 31 may generate explanatory information using various information that can be linked to the video segments in addition to the video segments and sensor information. Examples of such various information include, but are not limited to, weather information observed near the moving object when the video segments were captured.

[0151] In each exemplary embodiment, one or both of the descriptive information and the search query may be in natural language.

[0152] In each exemplary embodiment, each functional block of the video search devices 2 and 3 may be included in a physically single device, or may be included in a distributed manner across multiple physically different devices.

[0153] [Software implementation example] Some or all of the functions of the video search devices 1, 2, and 3 may be realized by hardware such as an integrated circuit (IC chip), or by software.

[0154] In the latter case, the video retrieval devices 1, 2, and 3 are realized, for example, by a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in FIG. 13. The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for operating the computer C as the video retrieval devices 1, 2, and 3. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing each function of the video retrieval devices 1, 2, and 3.

[0155] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0156] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, mouse, display, and printer.

[0157] Furthermore, the program P can be recorded on a non-transitory tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0158] [Appendix 1] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiments are also included in the technical scope of the present invention.

[0159] [Appendix 2] Some or all of the above-described embodiments can also be described as follows: However, the present invention is not limited to the following described aspects.

[0160] (Appendix 1) a generating means for generating explanatory information for each video stored in the video storage device; An acquisition means for acquiring a search query; a search means for searching for a video from the video storage device using the search query and the description information; an output means for outputting a search result obtained by the search means; an input means for receiving an input of a user's judgment result on the search results; an update means for updating the description information based on the determination result and the search query; A video search device comprising:

[0161] According to the above configuration, even if the amount or accuracy of information regarding the video is insufficient, the generated explanatory information can be updated accurately, and the accuracy of searches using the updated explanatory information can be improved.

[0162] (Appendix 2) the generation means generates the explanatory information using a generative model generated to input at least a video and output explanatory information; 2. A video search device according to claim 1.

[0163] According to the above configuration, by using a generative model, even when there is no or insufficient information about a video, it is possible to generate explanatory information about the video with high accuracy.

[0164] (Appendix 3) further comprising a model update means for updating the generative model using the explanation information updated by the update means; 3. A video search device according to claim 2.

[0165] According to the above configuration, by using the updated generative model, it is possible to generate explanation information with even greater accuracy.

[0166] (Appendix 4) the searching means searches the video storage device for videos whose description information at least partially matches the search query; the updating means updates a portion of description information relating to the searched video that does not match the search query in accordance with the determination result. 4. A video search device according to any one of appendices 1 to 3.

[0167] According to the above configuration, it is possible to accurately update the part of the description information related to the searched video that does not match the search query.

[0168] (Appendix 5) When the search result includes a plurality of videos, the output means outputs the search result in descending order of search accuracy by the search means. 5. A video search device according to any one of appendices 1 to 4.

[0169] According to the above configuration, the user can enjoy the advantage of being able to easily find a desired video segment.

[0170] (Appendix 6) When the search result includes a plurality of videos, the output means outputs the search result in descending order of search accuracy by the search means. 5. A video search device according to any one of appendices 1 to 4.

[0171] According to the above configuration, the user inputs the determination results in order of lowest search accuracy, which allows the description information for videos with low search accuracy to be updated with higher accuracy.

[0172] (Appendix 7) When the search results include a plurality of videos, the output means classifies and outputs the search results; the input means accepts input of the determination result for each of the classifications; 7. A video search device according to any one of appendices 1 to 6.

[0173] According to the above configuration, the user does not need to input the judgment result for each video included in the search results individually, but can input the judgment results for each category all at once, making it easier to input the judgment results. This increases the likelihood that judgment results will be input for more videos, and allows the description information to be updated more accurately.

[0174] (Appendix 8) the input means accepts input of a plurality of determination results for the search results, the update means updates the explanation information using a plurality of the determination results. 8. A video search device according to any one of appendices 1 to 7.

[0175] According to the above configuration, a more reliable determination result can be obtained compared to when a single determination result is used, thereby enabling the explanation information to be updated with higher accuracy.

[0176] (Appendix 9) Each video stored in the video storage device is associated with one or both of time information and position information, the updating means further updates the description information for other videos stored in the video storage device that have similar time information and / or position information to the video for which the description information is to be updated. 9. A video search device according to any one of appendices 1 to 8.

[0177] According to the above configuration, it is possible to accurately update the explanatory information even for a video for which input of a determination result has not been accepted.

[0178] (Appendix 10) the explanatory information includes first explanatory information and second explanatory information, the update means updates the explanatory information using a dependency relationship between the first explanatory information and the second explanatory information. 10. A video search device according to any one of appendices 1 to 9.

[0179] According to the above configuration, the first explanation information and the second explanation information having a dependency relationship can be updated with higher accuracy.

[0180] (Appendix 11) Each image stored in the image storage device is The video is taken by a camera mounted on a moving object. Each image is linked to sensor information acquired by a sensor mounted on the moving object, the generating means generates the explanatory information using the video and the sensor information. 11. A video search device according to any one of appendices 1 to 10.

[0181] According to the above configuration, even if the amount or accuracy of information regarding the video captured by the imaging device mounted on the moving body is insufficient, the accuracy of searching for the video can be improved.

[0182] (Appendix 12) The update means identifying a traveling direction of the moving object at the time of capturing the video for which the description information is to be updated; further updating the explanatory information regarding other images, among the images stored in the image storage device, whose traveling direction is similar to that of the image for which the explanatory information is to be updated; 12. The video search device according to claim 11.

[0183] According to the above configuration, by taking into consideration the traveling direction of the moving object, it is possible to accurately update the explanation information even for video for which input of the determination result has not been accepted. It is possible.

[0184] (Appendix 13) the explanatory information includes third explanatory information and fourth explanatory information; the generating means generates the third explanation information using a rule-based model, and generates the fourth explanation information based on a machine learning model or user input; the update means does not update the third description information, but updates the fourth description information. 13. A video search device according to claim 11 or 12.

[0185] The third explanation information is generated by a rule-based model, and therefore is likely to be highly objective and clearly defined. In contrast, the fourth explanation information is generated based on a machine learning model or user input, and therefore may be difficult to define or have low objectivity. According to the above configuration, the third explanation information, which is highly objective and clearly defined, is adopted from the generation unit, and the fourth explanation information, which is less objective or difficult to define, is updated, thereby enabling the explanation information to be updated with high accuracy.

[0186] (Appendix 14) a generating means for generating explanatory information for each video stored in the video storage device; An acquisition means for acquiring a search query; a search means for searching for a video from the video storage device using the search query and the description information; an output means for outputting a search result obtained by the search means; an input means for receiving an input of a user's judgment result on the search results; an update means for updating the description information based on the determination result and the search query; A video search system comprising:

[0187] According to the above configuration, the same effects as those in Supplementary Note 1 are achieved.

[0188] (Appendix 15) generating explanatory information for each video stored in the video storage device; Get the search query, retrieving a video from the video storage device using the search query and the description information; Output the search results, Accepting an input of a user's judgment result regarding the search results; The video search method further comprises updating the description information based on the determination result and the search query.

[0189] According to the above configuration, the same effects as those in Supplementary Note 1 are achieved.

[0190] (Appendix 16) A program for causing a computer to function as a video search device, the program comprising: a generating means for generating explanatory information for each video stored in the video storage device; An acquisition means for acquiring a search query; a search means for searching for a video from the video storage device using the search query and the description information; an output means for outputting a search result obtained by the search means; an input means for receiving an input of a user's judgment result on the search results; an update means for updating the description information based on the determination result and the search query; A program that functions as a

[0191] According to the above configuration, the same effects as those in Supplementary Note 1 are achieved.

[0192] [Appendix 3] Some or all of the above-described embodiments can also be expressed as follows.

[0193] at least one processor; The processor: a generation process for generating description information for each video stored in the video storage device; a retrieval process for retrieving a search query; a search process for searching for a video from the video storage device using the search query and the description information; an output process for outputting search results from the search process; an input process for receiving an input of a user's judgment result on the search results; an update process for updating the description information based on the determination result and the search query; A video search device that executes the above.

[0194] The video search device may further include a memory that stores a program for causing the processor to execute the generating process, the acquiring process, the searching process, the output process, the input process, and the updating process. The program may be recorded on a computer-readable, non-transitory, tangible recording medium. [Explanation of symbols]

[0195] 10, 20, 30 Video Search System 1, 2, 3 Video search device 9 Video storage device 11, 21, 31 generation section 12, 22, 32 Acquisition part 13, 23, 33 Search Section 14, 24, 34 output section 15, 25, 35 Input section 16, 26, 36 update part 37 Model Update Section 210, 310 control section 220, 320 storage section 230, 330 input / output section 240, 340 Communications Department C1 processor C2 Memory S1, S2, S3 video search method

Claims

1. a generating means for generating explanatory information for each video using the video stored in the video storage device and the sensor information associated with the video; An acquisition means for acquiring a search query; a search means for searching for a video from the video storage device using the search query and the description information; an output means for outputting a search result obtained by the search means; an input means for receiving an input of a user's judgment result on the search results; an update means for updating the description information based on the determination result and the search query; A video search device comprising:

2. the generation means generates the explanatory information using a generative model generated to input at least a video and output explanatory information; The video search device according to claim 1 .

3. further comprising a model update means for updating the generative model using the explanation information updated by the update means; The video search device according to claim 2 .

4. When the search results include a plurality of videos, the output means classifies and outputs the search results; the input means accepts input of the determination result for each of the classifications; The video search device according to any one of claims 1 to 3.

5. the input means accepts input of the judgment results for each of the plurality of search results, or accepts input of judgment results of a plurality of users for the search results; the update means updates the explanation information using a plurality of the determination results. The video search device according to any one of claims 1 to 4.

6. Each video stored in the video storage device is associated with one or both of time information and position information, the updating means further updates the description information for other videos stored in the video storage device that have similar time information and / or position information to the video for which the description information is to be updated. The video search device according to any one of claims 1 to 5.

7. the explanatory information includes first explanatory information and second explanatory information, the update means updates the explanatory information using a dependency relationship between the first explanatory information and the second explanatory information. The video search device according to any one of claims 1 to 6.

8. Each image stored in the image storage device is The video is taken by a camera mounted on a moving object. Each image is linked to sensor information acquired by a sensor mounted on the moving object, the generating means generates the explanatory information using the video and the sensor information; The update means identifying a traveling direction of the moving object at the time of capturing the video for which the description information is to be updated; further updating the explanatory information regarding other images, among the images stored in the image storage device, whose traveling direction is similar to that of the image for which the explanatory information is to be updated; The video search device according to claim 1 .

9. A computer comprising: generating explanatory information for each video using the video stored in the video storage device and the sensor information associated with the video; Get the search query, retrieving a video from the video storage device using the search query and the description information; Output the search results, Accepting an input of a user's judgment result regarding the search results; The video search method further comprises updating the description information based on the determination result and the search query.

10. A program for causing a computer to function as a video search device, the program comprising: a generating means for generating explanatory information for each video using the video stored in the video storage device and the sensor information associated with the video; An acquisition means for acquiring a search query; a search means for searching for a video from the video storage device using the search query and the description information; an output means for outputting a search result obtained by the search means; an input means for receiving an input of a user's judgment result on the search results; an update means for updating the description information based on the determination result and the search query; A program that functions as a

11. A generation means for generating explanatory information for each video stored in a video storage device using a generative model generated to input at least a video and output explanatory information; An acquisition means for acquiring a search query; a search means for searching for a video from the video storage device using the search query and the description information; an output means for outputting a search result obtained by the search means; an input means for receiving an input of a user's judgment result on the search results; an update means for updating the description information based on the determination result and the search query; a model update means for updating the generative model using the explanation information updated by the update means; A video search device comprising:

12. When the search results include a plurality of videos, the output means classifies and outputs the search results; the input means accepts input of the determination result for each of the classifications; The video search device according to claim 11.

13. The input means receives input of the judgment results for each of the plurality of search results, or receives input of judgment results of a plurality of users for the search results, the update means updates the explanation information using a plurality of the determination results.

13. The video search device according to claim 11 or 12.

14. Each image stored in the image storage device is associated with one or both of time information and position information, the updating means further updates the description information for other videos stored in the video storage device that have similar time information and / or position information to the video for which the description information is to be updated. The video search device according to any one of claims 11 to 13.

15. The explanatory information includes first explanatory information and second explanatory information, the update means updates the explanatory information using a dependency relationship between the first explanatory information and the second explanatory information. The video search device according to any one of claims 11 to 14.

16. Each image stored in the image storage device is The video is taken by a camera mounted on a moving object. Each image is linked to sensor information acquired by a sensor mounted on the moving object, the generating means generates the explanatory information using the video and the sensor information; The update means identifying a traveling direction of the moving object at the time of capturing the video for which the description information is to be updated; further updating the explanatory information regarding other images, among the images stored in the image storage device, whose traveling direction is similar to that of the image for which the explanatory information is to be updated; The video search device according to claim 11.

Citation Information

Patent Citations

  • Method and device for video retrieval and recording medium recording video retrieval program

    JP2000331009A

  • Information processor and information processing method

    JP2012208656A

  • Rule generation device, rule generation method, and rule generation program

    JP2020077343A

  • Generation device, generation system, and generation method

    JP2020201434A