Information processing system, information processing method, and program
The information processing system addresses the limitation of traditional video search by deriving attributes and behaviors from text input, allowing for comprehensive video search across multiple types and contexts.
Patent Information
- Application Number
- PCT/JP2025/013328
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-02
- Filing Date
- 2025-04-01
- Publication Date
- 2025-10-09
AI Technical Summary
Existing video search technologies limit the types of videos that can be searched by requiring users to click attribute check boxes, restricting the variety of search criteria that can be applied.
An information processing system that receives text input, derives attribute information and behavioral features from the text, and acquires videos based on these attributes and features, using deep learning-based models to enable comprehensive video search.
Enables users to search for a wide variety of videos by considering object, third-party, and environmental attributes and behaviors, improving search accuracy and efficiency.
Smart Images

Figure JP2025013328_09102025_PF_FP_ABST
Abstract
Description
Information processing system, information processing method and program
[0001] The present disclosure relates to an information processing system, an information processing method, and a program.
[0002] As computer capabilities improve, technology for users to search for desired videos is also improving.
[0003] For example, Patent Literature 1 discloses an image retrieval device that, when a user generates a search condition by clicking a checkbox of an attribute, searches for images containing similar poses according to the search condition. The image retrieval device searches an image database for an image by using image features as a key.
[0004] Japanese Patent Application Laid-Open No. 2019-091138
[0005] In video searches, it is desirable for users to be able to search for as many different types of videos as possible. However, the image search device disclosed in Patent Document 1 generates search criteria by the user clicking attribute check boxes. This may limit the types of videos that can be searched.
[0006] One of the objectives to be achieved by the embodiments of the present disclosure is to provide an information processing system, an information processing method, and a program that enable a user to search for various types of videos. It should be noted that this objective is only one of the objectives to be achieved by the embodiments disclosed herein. Other objectives or problems and novel features will become apparent from the description of this specification or the accompanying drawings.
[0007] An information processing system according to one embodiment includes: a receiving means for receiving text input; an attribute derivation means for deriving attribute information based on the text; a feature derivation means for deriving behavioral features of an object from which video is to be acquired based on the text; and an acquisition means for acquiring video based on the derived attribute information and the behavioral features of the object.
[0008] In one aspect, the information processing method is performed by a computer to receive text input, derive attribute information based on the text, derive behavioral characteristics of an object from which video is to be acquired based on the text, and acquire video based on the derived attribute information and the behavioral characteristics of the object.
[0009] In one embodiment, the program causes a computer to receive input of text, derive attribute information based on the text, derive behavioral characteristics of an object from which video is to be acquired based on the text, and acquire video based on the derived attribute information and the behavioral characteristics of the object.
[0010] The present disclosure can provide an information processing system, an information processing method, and a program that allow a user to search for many types of video.
[0011] FIG. 1 is a block diagram showing an example of an information processing system according to the present disclosure. FIG. 2 is a flowchart showing an example of a representative process of the information processing system. FIG. 3 is a block diagram showing an example of a video analysis system according to the present disclosure. FIG. 4 is an example of a screen displayed on a display unit. FIG. 5A is a flowchart showing an example of a representative process of the video analysis system. FIG. 5B is a flowchart showing an example of a representative process of the video analysis system. FIG. 6 is another example of a screen displayed on the display unit. FIG. 7 is another example of a screen displayed on the display unit. FIG. 8 is a block diagram showing another example of a video analysis system according to the present disclosure. FIG. 9 is another example of a screen displayed on the display unit. FIG. 10 is another example of a screen displayed on the display unit. FIG. 11A is a flowchart showing another example of a representative process of the video analysis system. FIG. 11B is a flowchart showing another example of a representative process of the video analysis system. FIG. 12 is a block diagram showing an example of the hardware configuration of an information processing device on which processing of the system or device is executed.
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the following descriptions and drawings in the embodiments have been omitted or simplified as appropriate for clarity of explanation. Furthermore, in this disclosure, unless otherwise specified, when multiple items are defined as "at least one of multiple items," the definition may mean any one item, or any multiple items including all items.
[0013] Each drawing referenced in the embodiments is merely an example for describing one or more embodiments. Each drawing is not related to only one particular embodiment, but may also be related to one or more other embodiments. As will be understood by those skilled in the art, various features or steps described with reference to any one drawing can be combined with features or steps shown in one or more other drawings to create, for example, an embodiment not explicitly shown or described. Not all features or steps shown in any one drawing are necessarily required to describe an exemplary embodiment, and some features or steps may be omitted. The order of steps described in any drawing may be changed as appropriate.
[0014] [Definition] In this disclosure, "video" refers to a single unit of still image (hereinafter also referred to simply as image) or video. Here, a single unit of video may be video data of a partial time domain among the video data to be confirmed, or may be one of multiple video data to be confirmed. Furthermore, source data refers to the image or video data to be confirmed from which video can be extracted.
[0015] In the present disclosure, an "object" may be a living thing such as a person or an animal, or may be a non-living thing. Non-living things may include, for example, structures such as buildings, roads, and walls, products for sale, vehicles, and machines. Objects may include manned vehicles or machines such as automobiles and ships, as well as robots that can operate unmanned, such as drones and unmanned vehicles.
[0016] In the present disclosure, "attribute information" includes, for example, at least one of the following pieces of information: (1) Information indicating the properties of the object itself indicated as a specific target by the input text; (2) Information indicating the properties of a second object itself, which is an object different from the first object indicated as a specific target by the input text; and (3) Information indicating the surrounding environment of the object indicated by the input text. The object indicated as a specific target by the text can also be said to be an object from which video is acquired. Hereinafter, (1) will also be referred to as object attribute information, (2) as third-party attribute information, and (3) as environment attribute information.
[0017] The object attribute information (1) indicates at least one of information that identifies the object, or information about the object such as the shape, color, and size of the object. The information that identifies the object may include the type of object (e.g., person, animal, vehicle, etc.). The object attribute information may further include information about detailed characteristics of the object (e.g., gender, age, clothing, and occupation of a person, or type of animal or vehicle, etc.).
[0018] The third-party attribute information (2) indicates at least one of information that identifies the second object or information about the second object, such as its shape, color, size, etc. A detailed description of this information is omitted here because it is similar to the object attribute information (1).
[0019] In the third-party attribute information (2), at least one of the first object and the second object performs an action related to the other. For example, in (2), a situation is assumed in which the first object performs some action toward the second object. Specific examples of actions include the first object looking at the second object, the first object moving toward the second object, and the first object picking up the second object. Here, the first object performing the action may be a person or an animal. Alternatively, the first object may be an object that has a visual sensor or a movable structure, such as a manned vehicle, a machine, or an unmanned robot. The second object may be a living thing such as a person or an animal, or may be any non-living object.
[0020] The environmental attribute information (3) is information indicating the environment that the object is in. For example, the environmental attribute information may include at least any of information indicating the presence or absence of a predetermined natural environment such as a forest or an ocean, information indicating the presence or absence of predetermined man-made objects such as buildings or roads, information indicating the weather such as sunny, cloudy, or rainy, information indicating the time of day such as daytime, morning, evening, or night, and information indicating brightness.
[0021] In the present disclosure, a "behavioral feature" is information indicating an object's behavior or a change in behavior. For example, the object's behavior may include looking, moving (e.g., a person walking or running, or a vehicle driving), turning the body, talking, falling, and extending or retracting a part of the object (e.g., a limb or a part of a robot). A behavior may be a single behavior such as looking or moving, or may include multiple behaviors such as "moving while looking at a visual target." A change in behavior may indicate a change in one or more behaviors themselves, or may indicate a change in the behavioral aspect, even if the one or more behaviors themselves do not change. An example of a change in the behavioral aspect is a change in the moving speed of an object while it is moving (increasing or decreasing the moving speed).
[0022] 1 is a block diagram showing an example of an information processing system according to the present disclosure. The information processing system 10 includes a receiving unit 12, an attribute derivation unit 14, a feature derivation unit 16, and an acquisition unit 18. Each unit of the information processing system 10 will be described below.
[0023] The reception unit 12 is an interface that receives input of text (i.e., character information). The reception unit 12 may receive information input directly from a user via a touch panel, a keyboard, a mouse, or the like. Alternatively, the reception unit 12 may be an input / output interface that receives information from a network external to the information processing system 10.
[0024] The attribute derivation unit 14 derives first attribute information based on the text input from the reception unit 12. The first attribute information includes at least one of the above-mentioned (1) object attribute information, (2) third-party attribute information, and (3) environment attribute information.
[0025] The feature derivation unit 16 derives behavioral features of a first object (hereinafter simply referred to as “first behavioral features”), which is the target of video capture, based on the text input from the reception unit 12. The behavioral features are as described above. The first object may be, for example, a person, an animal, a manned vehicle, a machine, or an unmanned robot.
[0026] The attribute derivation unit 14 can derive the first attribute information from text by using, for example, a deep learning-based LLM (Large Language Models) model. However, the attribute derivation unit 14 may derive the first attribute information from text by using language processing such as other trained models or algorithms. Similarly, the feature derivation unit 16 can derive the first behavioral feature by using any language processing such as LLM.
[0027] The acquisition unit 18 acquires a first video based on the first attribute information derived by the attribute derivation unit 14 and the first behavioral feature derived by the feature derivation unit 16. The acquisition unit 18 acquires a first video in which the first attribute information and the first behavioral feature appear.
[0028] As a first example, assume that the first attribute information is object attribute information of "female" and the first behavioral characteristic is "running." In this case, the acquisition unit 18 searches the source data to extract and acquire a video in which a "running female" appears.
[0029] As a second example, assume that the first attribute information includes the object attribute information "person" and the third-party attribute information "female," and the first behavioral characteristic is "increasing walking speed while looking at the visual target." In this case, the acquisition unit 18 searches the source data to extract and acquire a video in which "a person increases walking speed while looking at a female."
[0030] As a third example, assume that the first attribute information includes object attribute information "person" and environment attribute information "rainy day," and the first behavioral feature is "falling." In this case, the acquisition unit 18 searches the source data to extract and acquire a video in which a "person falling on a rainy day" is captured.
[0031] As a fourth example, assume that the first attribute information includes object attribute information "person," third-party attribute information "female," and environmental attribute information "rainy day," and the first behavioral characteristic is "increasing walking speed while looking at a visual target." In this case, the acquisition unit 18 searches the source data to extract and acquire a video in which "a person increases walking speed while looking at a female on a rainy day" appears.
[0032] As a fifth example, assume that the first attribute information includes object attribute information "motorcycle," third-party attribute information "elderly person," and environmental attribute information "night," and the first behavioral characteristic is "increasing speed while looking at a visual target." In this case, the acquisition unit 18 searches the source data to extract and acquire a video showing "a motorbike increasing speed while looking at an elderly person at night."
[0033] For example, the acquisition unit 18 may extract feature amounts from each video of the source data. If the source data is long video data, the acquisition unit 18 may divide the source data into predetermined time regions and extract feature amounts from each divided video. The acquisition unit 18 identifies, as the first video, a video whose extracted feature amounts include the first attribute information and the first behavioral feature.
[0034] The acquisition unit 18 may extract features by using a model trained using a deep learning technique. Specific examples of such a technique include a CNN (Convolutional Neural Network) and a variation of a CNN such as Residual Neural Networks (ResNet). The trained model includes layers such as a feature extraction layer and a fully connected layer.
[0035] As another example, the acquisition unit 18 may extract information about the object itself and information about the surrounding area of the object included in the source data. The acquisition unit 18 may perform this extraction process by using a model trained using a V&L (Vision and Language) method. The acquisition unit 18 identifies, as the first video, a video in which the extracted information about the object itself and information about the surrounding area of the object include first attribute information and first behavioral characteristics.
[0036] The above-described model is trained by inputting training data including, for example, multiple sets of sample videos and information indicating attribute information or behavioral features that serve as correct labels. After this training, source data information is input to the model. The model analyzes the source data videos to extract attribute information or behavioral features that appear in the videos. Using the results extracted by the model, the acquisition unit 18 identifies, as the first video, a video that appears with the same attribute information and behavioral features as the first attribute information and first behavioral features derived from the text received by the reception unit 12.
[0037] However, the acquisition unit 18 may identify the first video by using any algorithm instead of the trained model described above.
[0038] Furthermore, the information processing system 10 may perform analysis of source data in advance using CNN, V&L, or the like to assign tags indicating at least one of attribute information and behavioral features to information in the video. The acquisition unit 18 can acquire video in which the first attribute information appears by comparing the first attribute information derived by the attribute derivation unit 14 with the tags attached to the video. The acquisition unit 18 may perform similar processing on the first behavioral features derived by the feature derivation unit 16 instead of the first attribute information derived by the attribute derivation unit 14. In this way, the acquisition unit 18 identifies video in which the first attribute information and the first behavioral features appear as the first video.
[0039] 2 is a flowchart showing an example of a typical process of the information processing system 10. This flowchart explains the process of the information processing system 10. Note that the details of each process are as described above, and therefore will not be explained as appropriate.
[0040] First, the receiving unit 12 receives input of text (step S12). Then, the attribute derivation unit 14 derives first attribute information based on the text acquired in step S12 (step S14). Furthermore, the feature derivation unit 16 derives first behavioral features based on the text acquired in step S12 (step S16). Note that the process of step S14 and the process of step S16 may be executed first, or both processes may be executed in parallel. The acquiring unit 18 acquires a first video based on the first attribute information derived in step S14 and the first behavioral features derived in step S16 (step S18).
[0041] [Explanation of Effects] As described above, the information processing system 10 derives attribute information and behavioral features based on input text and acquires video based on the derived information. Here, the user can freely input text information. Therefore, the user can detect many types of video, not just video of patterns preset in the information processing system 10. Furthermore, the information processing system 10 can detect objects by taking into consideration attributes and behaviors in a composite manner.
[0042] Furthermore, the information processing system 10 can detect video using not only (1) object attribute information, but also (2) third-party attribute information and (3) environmental attribute information. In other words, the information processing system 10 can detect video using not only information on the detection target, but also environmental information and information on a third party that is thought to be related to the detection target as a trigger. Therefore, a user can detect many types of video by inputting environmental information and third-party information in text. A detailed example of this will be described in embodiment 2.
[0043] The information processing system 10 may be configured as a single computer device, or may be configured as a distributed system having multiple computer devices. In a distributed system, the processing performed by the information processing system 10 can be shared and executed by multiple computer devices. In other words, the reception unit 12 through the acquisition unit 18 may be distributed and installed on two or more computer devices. The information processing system 10 is realized by distributing and installing each unit of the information processing system 10 on two or more computer devices, and enabling the two or more computer devices to communicate with each other.
[0044] Some or all of the components of the information processing system 10 may be provided on a cloud server built on a cloud, or on other types of virtualized servers created using virtualization technology, etc. Functions other than those provided on servers such as cloud servers or virtualized servers are placed on edges. For example, in a system that acquires video footage captured on-site via a network and monitors the footage, edges are devices placed on or near the site, and are also devices that are close to the terminals in terms of the network hierarchy.
[0045] In each of the following embodiments, a specific example of the information processing system 10 described in embodiment 1 will be disclosed. However, the specific example of the information processing system 10 described in embodiment 1 is not limited to the one shown below. Furthermore, the configurations and processes described below are examples and are not limited to these.
[0046] Embodiment 2 [Configuration Description] Fig. 3 is a block diagram showing an example of a video analysis system according to the present disclosure. The video analysis system 20 includes a video acquisition unit 202, an analysis unit 204, a storage unit 206, an input unit 208, an attribute search unit 210, a behavior search unit 212, and a display unit 214. The video analysis system 20 functions as a center server that analyzes video captured by a surveillance camera CA installed at the site. Although only one surveillance camera CA is shown in Fig. 3, multiple surveillance cameras CA may be installed. The video analysis system 20 stores and analyzes video data captured by one or more surveillance cameras CA. Note that, hereinafter, a person is illustrated as an example of a target to be identified by the video analysis system 20, but the following processing can also be applied to other living things, robots, etc.
[0047] The video acquisition unit 202 is an interface that acquires video data captured by the surveillance camera CA via the network NW. The video acquisition unit 202 outputs the video data to the analysis unit 204. The surveillance camera CA continues to output the video data while capturing images. Therefore, the video acquisition unit 202 continues to output the video data to the analysis unit 204 while the surveillance camera CA continues capturing images.
[0048] The analysis unit 204 performs an analysis of the output video data using the trained model to identify attribute information of objects appearing in the video (e.g., attribute information of people) and assigns tags of the identified attribute information to the video. The specific analysis method is as described in the first embodiment. The analysis unit 204 stores the video data to which the attribute information tags have been assigned in the storage unit 206.
[0049] The storage unit 206 stores video data tagged with attribute information, and also stores trained models used by the analysis unit 204, the attribute search unit 210, and the behavior search unit 212 for searches.
[0050] Note that the video acquisition unit 202 to the storage unit 206 may perform the above-described processing on video data captured by not only one surveillance camera CA but also multiple surveillance cameras CA, and the video data may be stored in the storage unit 206. Furthermore, the video data stored in the storage unit 206 may be associated with identification information (e.g., location information) of the surveillance camera CA that captured the video data. Furthermore, the video data may be associated with information indicating the capture range of the surveillance camera CA.
[0051] The input unit 208 is an interface that accepts input of search text from an operator. The input unit 208 corresponds to the accepting unit 12 according to the first embodiment. The text input by the input unit 208 is displayed on the display unit 214. The text includes first attribute information and first behavioral characteristics. The input unit 208 outputs the text information to the attribute search unit 210.
[0052] The attribute search unit 210 derives first attribute information contained in the text by applying the trained LLM model to the text output from the input unit 208. The attribute search unit 210 then references the video data stored in the storage unit 206 and compares the tag of the assigned attribute information with the first attribute information. The attribute search unit 210 identifies, from the video data, videos of a predetermined time domain that are tagged with the same information as the first attribute information. In this manner, the attribute search unit 210 searches for and extracts one or more videos from the video data. The attribute search unit 210 corresponds to a part of the attribute derivation unit 14 and the acquisition unit 18 according to the first embodiment. The attribute search unit 210 outputs the extracted videos to the behavior search unit 212.
[0053] The behavior search unit 212 derives first behavioral features contained in the text by applying the trained LLM model to the text output from the input unit 208. As described above, the first behavioral features are behavioral features of the object indicated by the text. The behavior search unit 212 then analyzes the video output by the attribute search unit 210 and extracts behavioral features appearing in the video. The behavior search unit 212 searches for and extracts videos that contain the same features as the first behavioral features from the analyzed videos. The behavior search unit 212 uses a trained V&L model for this analysis. The attribute search unit 210 corresponds to part of the feature derivation unit 16 and the acquisition unit 18 according to embodiment 1. The behavior search unit 212 outputs the extracted first video to the display unit 214.
[0054] The display unit 214 is an interface (for example, a display or a touch panel) that displays the first video extracted by the behavior search unit 212. The operator can check the video corresponding to the input text by visually checking the display unit 214. The display unit 214 may also display identification information of the surveillance camera CA that captured the video and information indicating the capture range.
[0055] 4 shows an example of a screen displayed on the display unit 214. A specific example of the processing executed by the input unit 208 to the display unit 214 will be described below with reference to FIG.
[0056] In this example, the operator uses the input unit 208 to input the text "a person who saw a police officer and turned back" to search for footage in the video data that shows a person who saw a police officer and turned back. "A person who saw a police officer and turned back" refers to a person who recognized that a police officer was in the direction in which they were moving and changed the direction in which they were moving (i.e., moved in a direction to avoid approaching the police officer). This text is displayed as a search query in the text field IT on the display unit 214, allowing the operator to confirm whether the input was correct.
[0057] The attribute search unit 210 derives the object attribute information "person" and the third-party attribute information "police officer" as first attribute information included in the text. Then, the attribute search unit 210 searches for video to which "person" and "police officer" are added as attribute information tags by referencing the video data stored in the storage unit 206. Here, the attribute search unit 210 searches for video data from the surveillance camera CA captured in "X Station Front," "Y Intersection," and "Z Park." Location information indicating the locations where the surveillance camera CA captured the video is stored in association with the video data in the storage unit 206. In this way, the attribute search unit 210 searches for and extracts video containing people and police officers from the video data. The attribute search unit 210 outputs the extracted video to the behavior search unit 212.
[0058] The behavior search unit 212 derives "looks at a visual target and turns back" as a first behavior feature included in the text. Using the trained V&L model, the behavior search unit 212 searches for and extracts video showing the behavior of looking at some visual target and turning back from the video output by the attribute search unit 210. The behavior search unit 212 outputs the extracted first video to the display unit 214.
[0059] The display unit 214 displays the extracted first videos V1, V2, and V3. Video V1 shows person P1, upon seeing police officer C1, changing his direction of movement to avoid approaching the police officer. Similarly, videos V2 and V3 show people P2 and P3, upon seeing police officers C2 and C3, respectively, changing their direction of movement to avoid approaching the police officers. The display unit 214 also displays "X Station Front," "Y Intersection," and "Z Park" as location information for the surveillance camera CA that captured each of videos V1, V2, and V3. As described above, this location information is information stored in the storage unit 206 in association with the video data.
[0060] By checking the images V1 to V3 shown in Figure 4, the operator can determine that there is a suspicious person avoiding the police in each of the locations "in front of Station X," "intersection Y," and "park Z." This suspicious person may be a criminal engaged in drug trafficking or robbery, for example.
[0061] In the above, the text "The person who saw the police officer turned back" was used as an example, but the content of the text is not limited to this. Further examples will be given below.
[0062] (i) Assume that the operator inputs the text "A person who increases his walking speed when he sees a woman." The attribute search unit 210 derives the object attribute information "person" and the third-party attribute information "woman" as first attribute information contained in this text. Then, the attribute search unit 210 searches for videos to which "person" and "woman" are added as attribute information tags by referring to the video data stored in the storage unit 206. The attribute search unit 210 outputs the videos extracted as videos showing a person and a woman to the behavior search unit 212.
[0063] The behavior search unit 212 derives "increasing walking speed while looking at the visual target" as a first behavioral feature included in the text. The behavior search unit 212 searches for and extracts, from the video output by the attribute search unit 210, video in which the behavior of increasing walking speed while looking at the visual target is exhibited. The behavior search unit 212 outputs the extracted first video to the display unit 214.
[0064] Through the above processing, the display unit 214 displays the first video image showing the "person who increases his walking speed after seeing the woman." By checking this video image, the operator can determine that there is a suspicious person who may commit a crime against the woman (e.g., theft, stalking, sexual crime, etc.) at each point of the surveillance camera CA that captured the video image.
[0065] (ii) Assume that the operator inputs the text "A person who looks at the store clerk and puts the products on display into clothes." The attribute search unit 210 derives the object attribute information "person" and the third-party attribute information "store clerk" as first attribute information contained in this text. Then, the attribute search unit 210 searches for video to which "person" and "store clerk" are assigned as attribute information tags by referring to the video data stored in the storage unit 206. The attribute search unit 210 outputs the video extracted as video showing a person and a store clerk to the behavior search unit 212.
[0066] The behavior search unit 212 derives "looking at the visual target and putting a displayed product into clothes" as a first behavioral feature included in the text. The behavior search unit 212 searches for and extracts, from the video output by the attribute search unit 210, video in which the behavior of looking at the visual target and putting a displayed product into clothes is displayed. The behavior search unit 212 outputs the extracted first video to the display unit 214.
[0067] Through the above process, the display unit 214 displays the first video image showing "a person looking at the store clerk and putting a displayed item into his / her clothes." By checking this video image, the operator can determine whether there is a suspicious person who may be attempting to shoplift an item at each point of the surveillance camera CA that captured the video image.
[0068] (iii) Assume that the operator inputs the text "A person who increases his walking speed when he sees a child." The attribute search unit 210 derives the object attribute information "person" and the third-party attribute information "child" as first attribute information contained in this text. Then, the attribute search unit 210 searches for videos to which "person" and "child" are added as attribute information tags by referring to the video data stored in the storage unit 206. The attribute search unit 210 outputs the videos extracted as videos showing people and store clerks to the behavior search unit 212.
[0069] The behavior search unit 212 derives "increasing walking speed while looking at the visual target" as a first behavioral feature included in the text. The behavior search unit 212 searches for and extracts, from the video output by the attribute search unit 210, video in which the behavior of increasing walking speed while looking at the visual target is exhibited. The behavior search unit 212 outputs the extracted first video to the display unit 214.
[0070] Through the above process, the display unit 214 displays the first video image showing the "person seeing the child and increasing their walking speed." By checking this video image, the operator can determine whether there is a suspicious person who may be attempting to abduct a child at each point of the surveillance camera CA that captured the video image.
[0071] In the cases described above, the operator can prevent crime or resolve it quickly by promptly contacting the police and other relevant parties as necessary.
[0072] 5A and 5B are flowcharts showing an example of a typical process of the video analysis system 20. This flowchart explains the process of the video analysis system 20. Note that the details of each process are as described above, and therefore will not be explained further.
[0073] 5A shows the process that the video analysis system 20 performs on video data before text input. First, the video acquisition unit 202 acquires video data including video from the surveillance camera CA (step S22). The analysis unit 204 analyzes the video data and assigns tags of attribute information to the video (step S24). The analysis unit 204 stores the video data with the tag of attribute information in the storage unit 206 (step S26).
[0074] 5B shows the process by which the video analysis system 20 searches for videos after text input. First, the input unit 208 accepts search text input from an operator (step S32). The attribute search unit 210 derives first attribute information contained in the text and uses this attribute information to search for and extract videos containing the attribute information from the video data. In this way, the attribute search unit 210 narrows down the videos to be searched by the behavior search unit 212 (step S34).
[0075] The behavior search unit 212 derives a first behavioral feature included in the text. Using this behavioral feature, the behavior search unit 212 searches for and extracts videos in which the behavioral feature appears from the videos extracted by the behavior search unit 212 (step S36). The display unit 214 displays the videos extracted by the behavior search unit 212 (step S38).
[0076] [Explanation of Effect] Currently, surveillance cameras and live cameras are installed in various locations both outdoors and indoors (for example, in buildings and factories). Enabling operators to freely search the footage captured by such cameras is important for achieving goals such as preventing crimes and accidents and improving work efficiency.
[0077] In video search, it is desirable for operators to be able to search for as many different types of video as possible. However, video search has had issues such as only being able to detect predetermined behaviors of a target person and not being able to easily perform detection from a composite perspective using behaviors and person attributes. Furthermore, even if video search focusing on the behavior of a target person is possible, there is also the issue that video search using third-party information related to the target person is not possible.
[0078] The video analysis system 20 can solve the above-mentioned problems. Specifically, the video analysis system 20 allows an operator to input text, thereby enabling detection of various behaviors of a target person and detection from a composite perspective using the behavior and attributes of the person.
[0079] Furthermore, by inputting text, attribute information of a second object (a third party) on which the first object (the person being searched for) performs an action may be set as the first attribute information. This allows the video analysis system 20 to perform a video search using information about the third party related to the person being searched for, enabling the operator to search for various types of video.
[0080] Furthermore, attribute information of an object other than a person may be set as the first attribute information by text input. This allows the video analysis system 20 to search for videos that show objects other than people as targets, enabling the operator to search for various types of video.
[0081] Alternatively, the attribute search unit 210 may identify the video based on the first attribute information, and then the behavior search unit 212 may identify the first video based on the first behavioral feature. In this way, the video analysis system 20 can achieve effects such as shortening the time required to finally identify the video and improving the accuracy of identification by first narrowing down the videos to be identified using the first attribute information.
[0082] Note that the input unit 208 may further receive, as a query, the order of the process of identifying videos, so that the attribute search unit 210 first identifies videos and then the behavior search unit 212 identifies videos. Fig. 6 is another example of a screen displayed on the display unit 214. Fig. 6 shows a screen displayed on the display unit 214 when an operator inputs text (i.e., before videos are identified and displayed). The display unit 214 displays a text field IT that displays the input text, as well as a query that specifies the order of searches based on the first attribute information and first behavioral feature indicated by the text.
[0083] The operator uses the input unit 208 to input first attribute information to be searched first and first behavioral characteristics to be searched next. The input first attribute information is displayed in a text field I1 in FIG. 6 , and the input first behavioral characteristics are displayed in a text field I2 in FIG. 6 . In FIG. 6 , the first attribute information "police officer" is displayed in the text field I1, and the first behavioral characteristic "looked at (the visual target) and turned back" is displayed in the text field I2. In this way, the text of the first attribute information and the first behavioral characteristics are displayed as search queries in the text fields I1 and I2 on the display unit 214, respectively, allowing the operator to confirm whether the input was correct.
[0084] After the operator inputs text into text fields I1 and I2, he or she selects search execution instruction I3 using input unit 208, thereby confirming a query to video analysis system 20. In response to the input query, attribute search unit 210 narrows down the videos based on the first attribute information, and then behavior search unit 212 identifies the first video based on the first behavioral feature. In this way, by being able to accept the order of the process of identifying videos as a query, video analysis system 20 can more reliably achieve effects such as time reduction and improved identification accuracy.
[0085] After the operator inputs text into the text field IT, the attribute search unit 210 may analyze the input text and derive first attribute information contained in the text before inputting the text into the text field I1. The attribute search unit 210 displays the derived first attribute information in the text field I1 as a query candidate. The first attribute information displayed in the text field I1 may be set to a state that allows the operator to change it by operating the input unit 208. The operator refers to the first attribute information displayed in the text field I1 and checks whether there are any errors in the first attribute information. If there are any errors, the operator changes the first attribute information to the correct information by operating the input unit 208.
[0086] Similarly, after the operator inputs text into the text field IT, but before inputting text into the text field I1, the behavior search unit 212 may analyze the input text and derive first behavioral features included in the text. The behavior search unit 212 displays the derived first behavioral features in the text field I2 as query candidates. Note that the first attribute information displayed in the text field I2 may be set to a state where it can be changed by the operator operating the input unit 208. The operator refers to the first behavioral feature displayed in the text field I2 and checks whether there are any errors in the first behavioral feature. If there are any errors, the operator operates the input unit 208 to change the first behavioral feature to the correct information.
[0087] Displaying and changing text in the text fields described above may be performed in either text field I1 or I2, or in both text fields I1 and I2. After the operator confirms that the text displayed in text fields I1 and I2 is correct, they use the input unit 208 to select the search execution instruction I3. This confirms the query to the video analysis system 20. The attribute search unit 210 and the behavior search unit 212 perform the above-described processes in response to the input query. In this way, by displaying editable query candidates on the display unit 214 and then having the input unit 208 accept the official query, the video analysis system 20 can minimize the effort required for the user to input information. Furthermore, if an error occurs in the information derivation process for at least one of the attribute search unit 210 and the behavior search unit 212, the user can correct the error, thereby improving the accuracy of video identification in the video analysis system 20.
[0088] In the following third and fourth embodiments, variations of the video analysis system 20 described in the second embodiment will be disclosed. However, the variations of the video analysis system 20 are not limited to those described below. Furthermore, the explanation of the points already explained in the second embodiment will be omitted as appropriate, and only processing specific to each embodiment will be explained.
[0089] Embodiment 3 [Explanation of Processing] Fig. 7 is another example of a screen displayed on the display unit 214. Fig. 7 shows a screen displayed on the display unit 214 when an operator inputs text (i.e., before an image is specified and displayed). In addition to the text field IT, the display unit 214 displays a text field I4 that presents multiple related information candidates generated based on the input text.
[0090] When text is input into the text field IT, the attribute search unit 210 analyzes the input text to detect a second object that appears in the text. In the example shown in Fig. 7, the text "a person who increases walking speed when looking at a woman" is input, so the attribute search unit 210 detects "woman" as the second object.
[0091] The attribute search unit 210 generates a question as to which of the following information, (A) object attribute, (B) third-party attribute, or (C) environmental attribute, the woman corresponds to, as a candidate for relationship information indicating the relationship between the first object, the image of which is being acquired, and the detected second object. The attribute search unit 210 displays the generated question in a text field I4.
[0092] The attribute search unit 210 may present two or four or more options as candidates for related information. For example, the attribute search unit 210 may determine that "woman" in the text indicates a person, and generate a question as a candidate for related information asking whether the woman is (A) an object attribute or (B) a third-party attribute. The attribute search unit 210 displays the generated question in the text field I4.
[0093] The operator operates the input unit 208 to select one of (A) to (C) displayed in the text field I4. In this case, the operator selects (B) third-party attribute. The attribute search unit 210 determines that "female" is the third-party attribute information based on the candidate information received from the input unit 208. In other words, the attribute search unit 210 determines the first attribute information. Thereafter, the attribute search unit 210 executes the processing shown in the second embodiment. The processing executed by the behavior search unit 212 and the display unit 214 is also the same as in the second embodiment, and therefore description thereof will be omitted.
[0094] Note that, before presenting multiple candidates as related information, the attribute search unit 210 may present one candidate as attribute information for "female" in the text identified by the attribute search unit 210 through analysis. For example, if the attribute search unit 210 determines that the attribute of "female" in the text is (B) a third-party attribute, it creates a query as to whether the determination result is correct. Then, the attribute search unit 210 displays the generated query, such as "Is female a third-party attribute?" in the text field I4.
[0095] The operator operates the input unit 208 to input agreement or disagreement with the question. If the operator inputs agreement with the question, the attribute search unit 210 determines that "female" is a third-party attribute based on the information from the input unit 208. Thereafter, the attribute search unit 210, the behavior search unit 212, and the display unit 214 execute the processing described in the second embodiment. On the other hand, if the operator inputs disagreement with the question, the attribute search unit 210 presents multiple candidates as related information in the text field I4 based on the information from the input unit 208. For example, the content presented by the attribute search unit 210 in the text field I4 may be the content shown in FIG. 7 . Alternatively, in response to the operator's disagreement with the question, the attribute search unit 210 may delete (B) from candidates (A) to (C). The attribute search unit 210 generates a question asking whether "female" is (A) an object attribute or (C) an environment attribute, and displays the generated question in the text field I4. The operator operates the input unit 208 to select one of the displayed candidates. The processing of each part of the video analysis system 20 after the selection is as described above, and therefore a description thereof will be omitted.
[0096] As another example, when "running woman" is input as text, the behavior search unit 212 detects "woman" as the second object. Then, the display unit 214 displays the same content as shown in text field I4 in FIG. 7 . However, in this case, the operator selects (A) an object attribute. Based on the candidate information received from the input unit 208, the attribute search unit 210 determines that "woman" is the object attribute information. The attribute search unit 210, behavior search unit 212, and display unit 214 then execute the above-described processing.
[0097] [Explanation of Effect] As described above, the attribute search unit 210 displays one or more candidates of relationship information indicating the relationship between a first object, which is the target of video acquisition, and a second object detected from text on the display unit 214, and allows the operator to select a candidate. This allows the attribute search unit 210 to determine the first attribute information based on the information of the selected candidate. Therefore, the video analysis system 20 can improve the accuracy of video identification.
[0098] In addition, instead of the attribute search unit 210 presenting candidates for related information in the text field I4, the operator may determine the first attribute information by operating the input unit 208 to input the related information.
[0099] [Configuration Description] Figure 8 is a block diagram showing another example of a video analysis system according to the present disclosure. In addition to the components of the video analysis system 20 shown in Figure 3, the video analysis system 30 further includes a video search unit 216. Each unit of the video analysis system 30 is controlled by a hardware controller (not shown). Below, explanations of points already explained in the second embodiment will be omitted as appropriate, and the components and processing unique to the third embodiment will be particularly explained.
[0100] 9 is another example of a screen displayed on the display unit 214. This figure shows a screen displayed after the video search process described in the second embodiment is performed on sample video data based on the text entered in the text field IT. The searched text is "People who increase their walking speed when they see a woman."
[0101] In Fig. 9, a plurality of sample videos S1 to S3 are displayed as the first video displayed in step S38 of Fig. 5B. Here, the user operates the input unit 208 to select one or more videos from among the sample videos S1 to S3 as videos that reflect the situation described in the text. However, the user does not have to select any of the sample videos S1 to S3. Note that the sample videos S1 to S3 may be images or videos.
[0102] In this example, sample video S1 shows person P4 looking at woman W1 and increasing his walking speed. However, sample video S2 shows woman W2 and person P5, but person P5 is stationary. Furthermore, sample video S3 shows person P6 looking at woman W3 and increasing his walking speed. Therefore, the user selects sample videos S1 and S3 as videos that reflect the situation described in the text.
[0103] The video analysis system 30 receives information about the selected sample videos S1 and S3. These selected sample videos are associated with the text "Man increases walking speed when looking at a woman" and stored in the storage unit 206. The sample videos may be further associated with at least one piece of information, such as identification information of the surveillance camera CA that captured the sample video, the date and time of capture, and a flag indicating importance, and stored in the storage unit 206. The display unit 214 can then present the selected sample videos S1 and S3 as a prompt.
[0104] For example, when new text is input, the video analysis system 30 determines whether the new text means "a person who increases their walking speed when looking at a woman" based on the above analysis by the attribute search unit 210 and the behavior search unit 212. "A person who increases their walking speed when looking at a woman" is the text that was input when the sample video was previously selected. If the new text means "a person who increases their walking speed when looking at a woman," the display unit 214 presents the sample videos S1 and S3 associated with the text "a person who increases their walking speed when looking at a woman" as a prompt.
[0105] FIG. 10 is another example of a screen displayed on the display unit 214. Because the above process determines that the new text means "a person who increases their walking speed while looking at a woman," the display unit 214 presents the selected sample videos S1 and S3 as a prompt. The operator can select at least one of the presented videos using the input unit 208. However, the operator may not select any of the presented videos. The operator inputs the selected sample video as a query to the video analysis system 30 by selecting a search execution instruction I3 using the input unit 208. Note that the video data that is the subject of the video search in FIG. 9 may be different from or the same as the video data that is the subject of the sample video search in FIG. 8.
[0106] When the operator selects a sample video, the video search unit 216 uses the selected sample video to search for and identify a video similar to the sample video from among the video data to be searched. The video search unit 216 can use any method to search for a similar video. The video search unit 216 displays the identified video on the display unit 214.
[0107] 11A and 11B are flowcharts showing another example of typical processing by the video analysis system 30. This flowchart explains the processing by the video analysis system 30. Note that details of each process are as described above, and therefore will not be explained as appropriate. Furthermore, the processing that the video analysis system 30 performs on video data before text input is as described in FIG. 5A.
[0108] 11B shows the process by which the video analysis system 30 searches for videos after text input. First, the input unit 208 accepts search text input from an operator (step S42). The attribute search unit 210 derives first attribute information contained in the text and uses this attribute information to search for and extract sample videos containing the attribute information from the sample video data. In this way, the attribute search unit 210 narrows down the sample videos to be searched by the behavior search unit 212 (step S44).
[0109] The behavior search unit 212 derives a first behavioral feature contained in the text. Using this behavioral feature, the behavior search unit 212 searches for and extracts a video in which the behavioral feature appears among the sample videos extracted by the behavior search unit 212 (step S46). The display unit 214 displays the sample videos extracted by the behavior search unit 212 (step S48). When the operator selects a sample video, the selected sample video and the text obtained in step S42 are stored in the storage unit 206 in association with each other (step S50).
[0110] 11B illustrates the process by which the video analysis system 30 searches for a video after the sample video has been stored. First, the input unit 208 accepts search text input from the operator (step S52). The attribute search unit 210 and the behavior search unit 212 perform analysis to determine that the newly entered text has the same meaning as the text accepted in step S42. As a result, the display unit 214 displays the sample video stored in step S50 as a prompt (step S54).
[0111] The video search unit 216 uses the sample video selected by the operator to search for and identify videos similar to the sample video from the video data to be searched (step S56). The video search unit 216 displays the identified videos on the display unit 214 (step S58).
[0112] [Explanation of Effect] As described above, the video analysis system 30 can perform a video search using a sample video selected by an operator. In other words, the video analysis system 30 can use a video that clearly expresses information appearing in text as instruction information. Therefore, compared to using text as search information, it is possible to reduce search time and improve search accuracy.
[0113] In addition, in the video search processing of step S56 of FIG. 11B , in addition to the processing of the video search unit 216, at least one of the processing of the attribute search unit 210 or the behavior search unit 212 based on the text input in step S52 may be further executed. That is, in step S56 of FIG. 11B , at least one of the processing of step S34 or the processing of step S36 in FIG. 5B may be further executed to narrow down the identified videos. Either the attribute search processing performed by the attribute search unit 210 or the processing of the video search unit 216 may be executed first. Similarly, either the attribute search processing performed by the behavior search unit 212 or the processing of the video search unit 216 may be executed first. This further increases the likelihood that the identified and finally displayed video will satisfy the situation indicated in the text, thereby improving the accuracy of the video search.
[0114] 11B illustrates an example in which the display unit 214 displays a sample video as a prompt (step S54) after receiving input of search text (step S52). However, even if the operator does not input search text, the display unit 214 can display the sample video stored in step S50 as a prompt.
[0115] For example, when the operator operates the input unit 208 to input an instruction to display a sample video as a prompt, the display unit 214 may display the sample video as a prompt. At this time, if the operator does not specify any conditions for the sample video to be displayed, all or part of the sample video stored in the storage unit 206 is displayed on the display unit 214. On the other hand, if conditions for the sample video to be displayed are specified via the input unit 208, the video analysis system 30 displays all or part of the sample video stored in the storage unit 206 that satisfies the conditions on the display unit 214. The conditions are, for example, at least one of information such as identification information of the surveillance camera CA that captured the sample video, the date and time of capture, and a flag indicating importance. The video analysis system 30 references the storage unit 206 to display the sample video that satisfies the conditions on the display unit 214.
[0116] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0117] Like the information processing system 10, the video analysis systems 20 and 30 may be configured as a single computer device, or may be configured as a distributed system having multiple computer devices.
[0118] As described in the second embodiment, the system according to the present disclosure can be applied as a surveillance system for preventing or quickly resolving crimes, and can also be applied to the following uses, for example. However, the areas to which the system can be applied are not limited to the following: - An image acquisition system for improving the efficiency of employees at factory production sites or offices - An image acquisition system for determining which products or advertisements are most likely to attract customers' attention in commercial facilities
[0119] In the above-described embodiments, the present disclosure has been described as a hardware configuration, but the present disclosure is not limited to this. The present disclosure can also be realized by causing a processor in a computer to execute a computer program to perform the processing of the devices constituting the information processing system 10 or the video analysis systems 20 and 30 described in the above-described embodiments.
[0120] 12 is a block diagram showing an example of the hardware configuration of an information processing device (i.e., a computer) on which the processing of the system or device shown in each embodiment is executed. Referring to FIG. 12, an information processing device 90 includes a signal processing circuit 91, a processor 92, and a memory 93.
[0121] The signal processing circuit 91 is a circuit for processing signals in accordance with the control of the processor 92. The signal processing circuit 91 may include a communication circuit for receiving signals from a transmitting device.
[0122] The processor 92 is connected to the memory 93, and performs the processing of the system described in the above embodiment by reading and executing a computer program from the memory 93. As an example of the processor 92, one of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an FPGA (Field-Programmable Gate Array), a DSP (Demand-Side Platform), and an ASIC (Application Specific Integrated Circuit) may be used, or a plurality of these may be used in parallel.
[0123] The memory 93 may be a volatile memory, a nonvolatile memory, or a combination thereof. The memory 93 is not limited to one, and may be provided in multiple units. The volatile memory may be, for example, a random access memory (RAM) such as a dynamic random access memory (DRAM) or a static random access memory (SRAM). The nonvolatile memory may be, for example, a read only memory (ROM) such as a programmable random only memory (PROM) or an erasable programmable read only memory (EPROM), a flash memory, or a solid state drive (SSD).
[0124] The memory 93 is used to store one or more instructions. Here, the one or more instructions are stored as programs in the memory 93. The processor 92 can perform the processes described in the above embodiments by reading and executing these programs from the memory 93.
[0125] The memory 93 may include a memory provided outside the processor 92, as well as a memory built into the processor 92. The memory 93 may also include a storage device located away from the processors constituting the processor 92. In this case, the processor 92 can access the memory 93 via an I / O (Input / Output) interface.
[0126] As described above, one or more processors included in each system in the above-described embodiments execute one or more programs including instructions for causing a computer to execute the algorithms described using the drawings. Execution of the programs enables the information processing described in each embodiment to be realized.
[0127] The program includes instructions or software code that, when loaded into a computer, causes the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disk (DVD), Blu-ray disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagated signals. The transitory computer-readable medium or communication medium may provide the program to the computer via a wired communication path, such as an electric wire or optical fiber, or via a wireless communication path.
[0128] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes. (Supplementary Note 1) An information processing system comprising: a receiving means for receiving input of text; an attribute derivation means for deriving first attribute information based on the text; a feature derivation means for deriving, based on the text, a behavioral feature of a first object from which a video is to be acquired; and an acquisition means for acquiring a first video based on the derived first attribute information and the behavioral feature of the first object. (Supplementary Note 2) The information processing system according to Supplementary Note 1 further comprises a presentation means for presenting, based on the text, one or more candidates for relationship information indicating a relationship between a second object appearing in the text and the first object, wherein the receiving means accepts information of a candidate selected from one or more candidates, and the attribute derivation means determines the first attribute information based on the information of the selected candidate. (Supplementary Note 3) The information processing system according to Supplementary Note 1 or 2, further comprising a presentation means for presenting a plurality of the acquired first videos, wherein after the reception means receives information of one or more videos selected from the plurality of first videos, the presentation means presents the selected one or more first videos as a prompt, the reception means receives a selection of one or more videos from the one or more first videos presented as the prompt, and the acquisition means acquires a second video based on the received one or more videos. (Supplementary Note 4) The information processing system according to any one of Supplementary Notes 1 to 3, wherein the acquisition means identifies a video from video source data based on the first attribute information, and then identifies the first video from the identified video based on a behavioral feature of the first object. (Supplementary Note 5) The information processing system described in Supplementary Note 4, wherein the receiving means further receives, as a query, the order in which the acquisition means will identify images, and the acquisition means, in response to the query, identifies images from the source data based on the first attribute information, and then identifies the first image from the identified images based on behavioral characteristics of the first object.(Supplementary Note 6) The information processing system according to Supplementary Note 5, further comprising a presentation means for presenting at least one of the first attribute information derived by the attribute derivation means or the behavioral feature of the first object derived by the feature derivation means in a changeable state as a candidate for the query, wherein the accepting means accepts the query after the presentation means has presented the candidate query. (Supplementary Note 7) The information processing system according to any one of Supplements 1 to 6, wherein the first attribute information is attribute information of a second object on which the first object will perform an action. (Supplementary Note 8) The information processing system according to any one of Supplements 1 to 7, wherein the first attribute information is attribute information of an object. (Supplementary Note 9) An information processing method executed by a computer, comprising: accepting input of text; deriving attribute information based on the text; deriving behavioral features of an object from which video is to be acquired based on the text; and acquiring video based on the derived attribute information and the behavioral feature of the object. (Supplementary Note 10) A program that causes a computer to execute the following: accept text input; derive attribute information based on the text; derive behavioral characteristics of an object from which video is to be acquired based on the text; and acquire video based on the derived attribute information and behavioral characteristics of the object.
[0129] Some or all of the elements (e.g., configurations and functions) described in Supplementary Notes 2 to 8 that are dependent on Supplementary Note 1 may also be dependent on Supplementary Notes 9 and 10 in the same dependency relationship as Supplementary Notes 2 to 8. Some or all of the elements described in any Supplementary Note may be applied to various hardware, software, recording means for recording software, systems, and methods.
[0130] This application claims priority based on Japanese Patent Application No. 2024-059578, filed April 2, 2024, the disclosure of which is incorporated herein by reference in its entirety.
[0131] REFERENCE SIGNS LIST 10 Information processing system 12 Reception unit 14 Attribute derivation unit 16 Feature derivation unit 18 Acquisition unit 20, 30 Video analysis system 202 Video acquisition unit 204 Analysis unit 206 Storage unit 208 Input unit 210 Attribute search unit 212 Behavior search unit 214 Display unit 216 Video search unit
Claims
1. An information processing system comprising: a receiving means for receiving text input; an attribute derivation means for deriving first attribute information based on the text; a feature derivation means for deriving behavioral features of a first object that is the subject of video acquisition based on the text; and an acquisition means for acquiring a first video based on the derived first attribute information and behavioral features of the first object.
2. The information processing system of claim 1, further comprising a presentation means for presenting, based on the text, one or more candidates for relationship information indicating the relationship between a second object appearing in the text and the first object, wherein the reception means receives information on a candidate selected from one or more of the candidates, and the attribute derivation means determines the first attribute information based on the information on the selected candidate.
3. An information processing system as described in claim 1 or 2, further comprising a presentation means for presenting a plurality of the acquired first images, wherein after the reception means receives information on one or more images selected from the plurality of first images, the presentation means presents the selected one or more first images as a prompt, the reception means receives a selection of one or more images from the one or more first images presented as the prompt, and the acquisition means acquires a second image based on the received one or more images.
4. An information processing system according to any one of claims 1 to 3, wherein the acquisition means identifies an image from the source data of the image based on the first attribute information, and then identifies the first image from the identified image based on the behavioral characteristics of the first object.
5. The information processing system of claim 4, wherein the receiving means further receives as a query the order in which the acquisition means will identify images, and the acquisition means, in response to the query, identifies images from the source data based on the first attribute information, and then identifies the first image from the identified images based on the behavioral characteristics of the first object.
6. The information processing system of claim 5, further comprising a presentation means for presenting at least one of the first attribute information derived by the attribute derivation means or the behavioral characteristics of the first object derived by the characteristic derivation means in a changeable state as a candidate for the query, and wherein the acceptance means accepts the query after the presentation means has presented the candidate for the query.
7. An information processing system according to any one of claims 1 to 6, wherein the first attribute information is attribute information of a second object on which the first object performs an action.
8. An information processing system according to any one of claims 1 to 7, wherein the first attribute information is attribute information of an object.
9. An information processing method executed by a computer, which comprises: accepting text input; deriving attribute information based on the text; deriving behavioral characteristics of an object from which video is to be acquired based on the text; and acquiring video based on the derived attribute information and behavioral characteristics of the object.
10. A program that causes a computer to perform the following steps: accept text input; derive attribute information based on the text; derive behavioral characteristics of an object from which video is to be acquired based on the text; and acquire video based on the derived attribute information and behavioral characteristics of the object.
Citation Information
Patent Citations
Information processing apparatus, information processing method, and information processing program
JP2022180941A
Information processing device, information processing method, and information processing program
JP2022180942A
Video search system, video search method, and computer program
WO2021145030A1