Processing device, processing method, and program
Patent Information
- Application Number
- JP2025506540
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies face challenges in accurately selecting a frame image that can effectively extract appearance features of a person to be tracked, leading to difficulties in generating a search query that can accurately search for the person across multiple frame images.
A processing device and method that accepts user input to specify a person in a target frame image, detects the person in surrounding frames, extracts appearance features related to multiple items, and generates a search query by integrating these features from the target and surrounding frames, allowing for the selection of relevant features based on reliability, brightness, and other criteria to create a comprehensive search query.
This approach enables the generation of a search query that can accurately identify a person to be tracked, even when a single frame image cannot capture all appearance features, providing a more flexible and accurate search capability.
Abstract
Description
Processing device, processing method, and recording medium
[0001] The present invention relates to a processing device, a processing method, and a program.
[0002] Techniques related to the present invention are disclosed in Patent Documents 1 and 2.
[0003] The technology disclosed in Patent Literature 1 searches for an object specified by a user in a video from different time periods within the video or from different videos. When the technology receives a user input specifying an object in a frame image, it selects one query image from a series of frame images before and after the object and performs a similar image search. More specifically, the technology extracts a frame image in which a person is facing a specific direction from the series of frame images before and after the object and uses that frame image as the query image.
[0004] The technology disclosed in Patent Document 2 discloses a technology for tracking a target person within an image.
[0005] JP 2015-114685 A JP 2007-068008 A
[0006] By performing an image search using a search query that accurately describes the person to be tracked, the person to be tracked can be accurately detected from the image. The search query is composed of features (appearance features) related to multiple items that can be extracted from the person's appearance. The multiple items include, but are not limited to, gender, age, hairstyle, body type, clothing color, clothing design, color and design of belongings, etc.
[0007] In the technology disclosed in Patent Document 1, one frame image selected from a plurality of frame images serves as a query image. When using this technology, a plurality of appearance features related to a person to be tracked are extracted from the single frame image. However, it is not easy to select a single frame image from which the appearance features of all items can be accurately extracted.
[0008] The technology disclosed in Patent Document 2 is not a technology for generating search queries.
[0009] In view of the above-described problems, an example of an object of the present invention is to provide a processing device, a processing method, and a program that generate a search query that can accurately search for a person to be tracked.
[0010] According to one aspect of the present invention, there is provided a processing device having: a receiving means for receiving a user input specifying a person to be tracked within a target frame image, which is one of a plurality of frame images in a time series; a detecting means for detecting the person to be tracked within a plurality of surrounding frame images before and / or after the target frame image; an extracting means for extracting appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of surrounding frame images; and a generating means for integrating the appearance features extracted from the target frame image and each of the plurality of surrounding frame images for each of the items to generate a search query.
[0011] According to one aspect of the present invention, there is provided a processing method in which one or more computers receive user input specifying a person to be tracked within a target frame image, which is one of a plurality of frame images in a time series; detect the person to be tracked within a plurality of surrounding frame images before and / or after the target frame image; extract appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of surrounding frame images; and integrate the appearance features extracted from the target frame image and each of the plurality of surrounding frame images for each of the items to generate a search query.
[0012] According to one aspect of the present invention, there is provided a program that causes a computer to function as: a receiving means that receives user input specifying a person to be tracked within a target frame image, which is one of a plurality of frame images in a time series; a detecting means that detects the person to be tracked within a plurality of surrounding frame images before and / or after the target frame image; an extracting means that extracts appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of surrounding frame images; and a generating means that integrates the appearance features extracted from the target frame image and each of the plurality of surrounding frame images for each of the items to generate a search query.
[0013] According to one aspect of the present invention, a processing device, a processing method, and a program are provided that generate a search query that can accurately search for a person to be tracked.
[0014] The above-mentioned objects and other objects, features and advantages will become more apparent from the following description of the preferred embodiments and the accompanying drawings.
[0015] FIG. 1 is a diagram showing an example of a functional block diagram of a processing device. FIG. 2 is a diagram for explaining processing of the processing device. FIG. 3 is another diagram for explaining processing of the processing device. FIG. 4 is a diagram showing an example of a hardware configuration of a processing device. FIG. 5 is another diagram for explaining processing of the processing device. FIG. 6 is a flowchart showing an example of a flow of processing of the processing device. FIG. 7 is another diagram for explaining processing of the processing device. FIG. 8 is a diagram showing an example of a functional block diagram of a processing device.
[0016] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0017] 1 is a functional block diagram showing an overview of a processing device 10 according to a first embodiment. The processing device 10 includes a receiving unit 11, a detecting unit 12, an extracting unit 13, and a generating unit 14.
[0018] The receiving unit 11 receives user input specifying a person to be tracked in a target frame image, which is one of a plurality of frame images in a time series. The detecting unit 12 detects the person to be tracked in a plurality of peripheral frame images before and / or after the target frame image. The extracting unit 13 extracts appearance features related to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images. The generating unit 14 generates a search query by integrating the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each item.
[0019] In this way, the processing device 10 extracts appearance features relating to multiple items of the person to be tracked from each of multiple frame images consisting of the target frame image and the surrounding frame images before and after it.The processing device 10 then generates a search query by integrating the appearance features extracted from the multiple frame images "for each item."The processing device 10 of this embodiment generates search queries using such a unique method, making it possible to generate search queries that can accurately search for the person to be tracked.
[0020] Second Embodiment Overview A processing apparatus 10 according to a second embodiment is a specific embodiment of the processing apparatus 10 according to the first embodiment.
[0021] 2, the processing device 10 receives a user input specifying a person P to be tracked in a target frame image, which is one of a plurality of frame images in time series. The processing device 10 then detects the person P to be tracked in a plurality of peripheral frame images before and / or after the target frame image.
[0022] Next, the processing device 10 extracts appearance features F of the person P to be tracked from the target frame image and each of the multiple peripheral frame images. As shown in Fig. 3, the appearance features F include information on multiple items. The multiple items include, but are not limited to, gender, age, hairstyle, body type, clothing color, clothing design, the color and design of belongings, etc.
[0023] Then, as shown in FIGS. 2 and 3, the processing device 10 generates a search query by integrating the appearance features F extracted from the target frame image and each of the multiple peripheral frame images for each item.
[0024] The configuration of the processing device 10 will be described in detail below.
[0025] "Hardware Configuration" An example of the hardware configuration of the processing device 10 will be described. Each functional unit of the processing device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various variations in the realization method and device. Software includes programs that are pre-loaded in the device before shipping, and programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet.
[0026] FIG. 4 is a block diagram illustrating an example of the hardware configuration of a processing device 10. As shown in FIG. 4, the processing device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The processing device 10 does not necessarily have to have the peripheral circuit 4A. Note that the processing device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices may have the above hardware configuration.
[0027] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to mutually transmit and receive data. The processor 1A is, for example, a processing unit such as a CPU or a graphics processing unit (GPU). The memory 2A is, for example, a random access memory (RAM) or a read-only memory (ROM). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes an interface for connecting to a communication network such as the Internet. Examples of input devices include a keyboard, mouse, microphone, physical buttons, touch panel, etc. Examples of output devices include a display, speaker, printer, mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0028] "Functional Configuration" Next, the functional configuration of the processing device 10 of this embodiment will be described in detail. Fig. 1 shows an example of a functional block diagram of the processing device 10 of this embodiment. As shown in the figure, the processing device 10 of this embodiment has a receiving unit 11, a detecting unit 12, an extracting unit 13, and a generating unit 14.
[0029] The reception unit 11 receives a user input specifying a person to be tracked within a target frame image, which is one of a plurality of frame images in time series. The user performs an "input to specify one of the plurality of frame images as the target frame image" and an "input to specify a person to be tracked within the target frame image."
[0030] The "input to designate one of a plurality of frame images as the target frame image" can be realized using any technology. For example, a user plays a video and inputs an input to pause the playback at a scene in which a person to be tracked appears. The reception unit 11 identifies the frame image displayed on the display at the time of this pause as the target frame image.
[0031] Alternatively, the input may be received from the user by other means. For example, the user may input the amount of time that has elapsed since the start of the video. The receiving unit 11 may then identify one frame image identified by this amount of time as the target frame image.
[0032] The "input specifying a person to be tracked within a target frame image" can be realized using any technology. In one example, the processing device 10 executes a person detection process on the target frame image. Then, as shown in FIG. 5 , the reception unit 11 displays a rectangular area W including the detected person superimposed on the target frame image, and receives a user input specifying the rectangular area W. Although one rectangular area W is shown in FIG. 5 , multiple rectangular areas W may be shown. The reception unit 11 identifies a person present within the specified rectangular area W as a person to be tracked.
[0033] Note that the input from the user may be received by other means. For example, the user may input to designate an area within the target frame image that includes the person to be tracked. The input to designate a partial area within the image can be realized using any technology. The processing device 10 executes a person detection process on the image within the designated area. Then, the receiving unit 11 identifies the person detected within the designated area as the person to be tracked.
[0034] 1 , the detection unit 12 detects the person to be tracked in multiple peripheral frame images before and / or after the target frame image. "Before and / or after the target frame image" means before and / or after the target frame image in chronological order within the multiple frame images arranged in chronological order.
[0035] A "peripheral frame image" is any of the following frame images: - Frame images from the frame image that occurs a first predetermined time before the target frame image to the frame image that occurs a second predetermined time after the target frame image, excluding the target frame image - Frame images from the frame image that occurs a first predetermined time before the target frame image to the frame image that immediately precedes the target frame image - Frame images from the frame image that occurs immediately after the target frame image to the frame image that occurs a second predetermined time after the target frame image
[0036] The first predetermined time and the second predetermined time may be the same or different. The first predetermined time and the second predetermined time may be predetermined fixed values. Furthermore, the first predetermined time and the second predetermined time may be changeable by the user. As will be described in detail in the following embodiment, the detection unit 12 may have a function to determine optimal first predetermined time and second predetermined time for each moving image to be processed.
[0037] The detection of the tracking target person in each of the plurality of peripheral frame images can be realized using any technology. For example, the detection unit 12 may detect the tracking target person in each of the plurality of peripheral frame images by tracking the tracking target person specified in the target frame image in the moving image using an object tracking technology that tracks an object in the moving image.
[0038] The extraction unit 13 extracts appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images.
[0039] The "multiple items" refer to characteristics of a person that can be extracted from their appearance, i.e., characteristics of a person that can be extracted by image analysis. The multiple items include, but are not limited to, gender, age, hairstyle, body type, clothing color, clothing design, color and design of belongings, etc.
[0040] The appearance characteristics that the item "gender" can take are male and female.
[0041] The appearance feature that can be taken by the item "age" may be the age itself or an age range such as teens or twenties.
[0042] The appearance feature that the item "hairstyle" can take is a hairstyle classification such as shaved head, semi-long hair, etc.
[0043] The appearance characteristics that can be taken by the item "body type" are body type classifications such as thin type, chubby type, and the like.
[0044] The appearance feature that the item "clothing color" can take may be a single color such as red or black, or a combination of multiple colors such as red and black.
[0045] Appearance features that can be taken by the item "clothing design" are clothing classifications such as miniskirts, pants, T-shirts, and coats. Note that appearance features that can be taken by the item "clothing design" may be further subdivided into these clothing classifications. For example, coats can be subdivided into down coats and duffle coats.
[0046] The appearance feature that can be taken by the item "color of possession" may be a single color such as red or black, or a combination of multiple colors such as red and black.
[0047] The appearance feature that can be taken by the item "design of belongings" is the classification of belongings, such as a bag or an umbrella. Note that the appearance feature that can be taken by the item "design of belongings" may be a further classification of belongings. For example, a bag can be further classified into a backpack, a business bag, etc. If a person has multiple belongings, multiple features can be taken.
[0048] The extraction unit 13 can achieve this extraction using any technology that estimates or identifies various appearance features of a person appearing in an image through image analysis. For example, an estimation model (classifier) generated by machine learning can be used, but this is not limiting. When an estimation model is used, the reliability (sometimes referred to as certainty) of each appearance feature extracted from the image can be obtained.
[0049] The generation unit 14 generates a search query by integrating, for each item, the appearance features extracted from each of a plurality of frame images, including the target frame image and a plurality of peripheral frame images. Hereinafter, "a plurality of frame images, including the target frame image and a plurality of peripheral frame images" may be simply referred to as "a plurality of frame images."
[0050] The "integration" performed by the generation unit 14 involves selecting, for each item, appearance features to be included in the search query from among appearance features extracted from multiple frame images. The generation unit 14 may select one appearance feature or multiple appearance features corresponding to one item. Furthermore, the generation unit 14 may select no appearance features corresponding to one item. Furthermore, the generation unit 14 may select different appearance features for each item. That is, the generation unit 14 may select one appearance feature corresponding to one item, multiple appearance features corresponding to another item, and no appearance features corresponding to yet another item. The generation unit 14 generates a search query including the appearance features selected for each item in this manner.
[0051] Here, we will explain the process of selecting, for each item, appearance features to be included in a search query from among the appearance features extracted from multiple frame images. The generation unit 14 selects, for each item, at least one of the following appearance features 1 to 10 from among the appearance features extracted from multiple frame images.
[0052] Note that the appearance feature selected may differ for each item. For example, the generation unit 14 may select appearance feature 1 for one item and appearance feature 5 for another item. Alternatively, the generation unit 14 may select appearance feature 1 for one item and appearance feature 2 and appearance feature 3 for another item. It is determined in advance which of appearance features 1 to 10 will be selected for each item, and the generation unit 14 selects an appearance feature for each item according to that rule.
[0053] (Appearance feature 1) Appearance feature extracted from the largest number of frame images (Appearance feature 2) Appearance feature extracted from a predetermined percentage or more of frame images (Appearance feature 3) Appearance feature extracted from a predetermined number or more of frame images (Appearance feature 4) Appearance feature with the highest reliability of the extraction result (Appearance feature 5) Appearance feature with a reliability of the extraction result equal to or greater than a threshold (Appearance feature 6) Appearance feature extracted from a frame image in which the brightness of a rectangular area containing the person to be tracked in the frame image is the brightest (Appearance feature 7) Appearance feature extracted from a frame image in which the brightness of a rectangular area containing the person to be tracked in the frame image is equal to or greater than a threshold (Appearance feature 8) Appearance feature extracted from a frame image in which the size of a rectangular area containing the person to be tracked in the frame image is the largest (Appearance feature 9) Appearance feature extracted from a frame image in which the size of a rectangular area containing the person to be tracked in the frame image is equal to or greater than a threshold (Appearance feature 10) Appearance feature extracted from a frame image in which the person to be tracked in the frame image does not overlap with other people or objects in the frame image
[0054] Appearance feature 1 is the appearance feature extracted from the largest number of frame images. The generation unit 14 selects the appearance feature extracted from the largest number of frame images as appearance feature 1 from among the appearance features extracted from a plurality of frame images.
[0055] The process of selecting appearance feature 1 may be performed using only appearance features extracted from multiple frame images, whose extraction results output from the estimation model have a reliability equal to or greater than a threshold. That is, the generation unit 14 may select the most common appearance feature from the collection of appearance features whose reliability equals or exceeds a threshold as appearance feature 1. The threshold is a predetermined arbitrary value.
[0056] Appearance feature 2 is an appearance feature extracted from a predetermined proportion or more of frame images. The generation unit 14 selects, as appearance feature 2, an appearance feature that accounts for a predetermined proportion or more of the appearance features extracted from multiple frame images. The predetermined proportion is an arbitrary value that is determined in advance. If there is no appearance feature that accounts for a predetermined proportion or more, the generation unit 14 does not select any appearance feature as appearance feature 2.
[0057] The process of selecting appearance feature 2 may be performed using only appearance features extracted from multiple frame images whose extraction results output from the estimation model have a reliability equal to or greater than a threshold. That is, the generation unit 14 may select, as appearance feature 2, appearance features that account for a predetermined percentage or more of the collection of appearance features whose reliability equals or exceeds a threshold. The threshold is a predetermined arbitrary value.
[0058] Appearance feature 3 is an appearance feature extracted from a predetermined number or more of frame images. The generation unit 14 selects, from among appearance features extracted from a plurality of frame images, appearance features extracted from a predetermined number or more of frame images as appearance feature 3. The predetermined number is an arbitrary value that is determined in advance. If there are not appearance features equal to or greater than the predetermined number, the generation unit 14 does not select any appearance feature as appearance feature 3.
[0059] The process of selecting appearance feature 3 may be performed using only appearance features extracted from multiple frame images, whose extraction results output from the estimation model have a reliability equal to or greater than a threshold. That is, the generation unit 14 may select, from a collection of appearance features whose reliability is equal to or greater than a threshold, appearance features extracted from a predetermined number of frame images as appearance feature 3. The threshold is a predetermined arbitrary value.
[0060] Appearance feature 4 is the appearance feature with the highest reliability in the extraction result output from the estimation model described above. The generation unit 14 selects the appearance feature with the highest reliability as appearance feature 4.
[0061] Appearance feature 5 is an appearance feature for which the reliability of the extraction result output from the above-mentioned estimation model is equal to or greater than a threshold. The generation unit 14 selects an appearance feature for which the reliability is equal to or greater than the threshold as appearance feature 5. If there is no appearance feature for which the reliability is equal to or greater than the threshold, the generation unit 14 does not select any appearance feature as appearance feature 5. The threshold is an arbitrary value that is determined in advance.
[0062] Appearance feature 6 is an appearance feature extracted from a frame image in which the brightness of a rectangular area including the tracking target person within the frame image is the brightest. By performing a person detection process on the frame images, a rectangular area W including the person is detected, as shown in FIG. 5 . The generation unit 14 calculates the brightness of the rectangular area including the tracking target person detected in this manner for each frame image. The generation unit 14 then selects, as appearance feature 6, the appearance feature extracted from the frame image in which the calculated brightness of the rectangular area is the brightest among the multiple frame images. Indicators that can be used to indicate the brightness of a rectangular area include brightness, luminance, and luminosity. For example, the brightness of the rectangular area can be determined by the statistical value of these indices for the pixels included in the rectangular area. Examples of statistical values include, but are not limited to, the mean, median, mode, maximum, and minimum values.
[0063] The process of selecting the appearance feature 6 may be performed using only those appearance features extracted from multiple frame images whose extraction results output from the estimation model have a reliability equal to or greater than a threshold. That is, the generation unit 14 may select, as the appearance feature 6, the appearance feature extracted from the frame image in which the brightness of the rectangular area is the brightest among the appearance features whose reliability equals or exceeds a threshold. The threshold is a predetermined arbitrary value.
[0064] Appearance feature 7 is an appearance feature extracted from a frame image in which the brightness of a rectangular region including the person to be tracked in the frame image is equal to or greater than a threshold. The generation unit 14 selects the appearance feature extracted from the frame image in which the brightness is equal to or greater than the threshold as appearance feature 7. If there is no frame image in which the brightness is equal to or greater than the threshold, the generation unit 14 does not select any appearance feature as appearance feature 7. The threshold is an arbitrary value that is determined in advance.
[0065] The process of selecting the appearance feature 7 may be performed using only those appearance features extracted from multiple frame images whose extraction results output from the estimation model have a reliability equal to or greater than a threshold. That is, the generation unit 14 may select, as the appearance feature 7, an appearance feature extracted from a frame image whose brightness is equal to or greater than a threshold among the appearance features whose reliability is equal to or greater than a threshold. The threshold is a predetermined arbitrary value.
[0066] Appearance feature 8 is an appearance feature extracted from the frame image in which the rectangular area containing the person to be tracked is the largest. The generation unit 14 selects the appearance feature extracted from the frame image with the largest size as appearance feature 8. The size of the rectangular area can be indicated, for example, by the number of pixels contained in the rectangular area.
[0067] The process of selecting the appearance feature 8 may be performed using only those appearance features extracted from multiple frame images whose extraction results output from the estimation model have a reliability equal to or greater than a threshold. That is, the generation unit 14 may select, from among the appearance features whose reliability is equal to or greater than a threshold, the appearance feature extracted from the frame image with the largest size as the appearance feature 8. The threshold is a predetermined arbitrary value.
[0068] The appearance feature 9 is an appearance feature extracted from a frame image in which the size of a rectangular area containing the person to be tracked in the frame image is equal to or greater than a threshold. The generation unit 14 selects the appearance feature extracted from the frame image in which the size is equal to or greater than the threshold as the appearance feature 9. If there is no frame image in which the size is equal to or greater than the threshold, the generation unit 14 does not select any appearance feature as the appearance feature 9. The threshold is an arbitrary value that is determined in advance.
[0069] The process of selecting the appearance feature 9 may be performed using only those appearance features extracted from a plurality of frame images whose extraction results output from the estimation model have a reliability equal to or greater than a threshold. That is, the generation unit 14 may select, as the appearance feature 9, an appearance feature extracted from a frame image whose magnitude is equal to or greater than a threshold among the appearance features whose reliability is equal to or greater than a threshold. The threshold is a predetermined arbitrary value.
[0070] The appearance feature 10 is an appearance feature extracted from a frame image in which the tracking target person does not overlap with other people or objects within the frame image. The generation unit 14 determines, for each frame image, whether the tracking target person overlaps with other people or objects within the frame image. The generation unit 14 then identifies, as the appearance feature 10, the appearance feature extracted from the frame image in which the tracking target person does not overlap with other people or objects within the frame image.
[0071] The process of determining whether the tracked person overlaps with another person or object in the frame image can be realized using any technology. For example, the generation unit 14 may perform an object detection process on the frame image to identify a rectangular area including a person or object. Then, based on whether the rectangular area of the identified person or object overlaps with the rectangular area including the tracked person, it may be determined whether the tracked person overlaps with another person or object in the frame image. For example, if the rectangular area of the identified person or object overlaps with the rectangular area including the tracked person, it can be determined that the tracked person overlaps with another person or object in the frame image.
[0072] Next, a process for generating a search query including the appearance features selected for each item will be described. The generation unit 14 generates a search query by connecting the appearance features selected for each item with predetermined logical operators in accordance with predetermined rules.
[0073] For example, the generation unit 14 generates a condition for each item (item-specific condition) based on the result of integrating the appearance features for each item. Then, the generation unit 14 generates a search query by connecting multiple item-specific conditions with a predetermined logical operator (e.g., AND condition).
[0074] The item-specific condition may be a condition in which one appearance feature is set for one item, such as "gender: male." The generation unit 14 may select multiple appearance features corresponding to one item. In this case, the item-specific condition may be a condition in which multiple appearance features are connected by an OR condition. The generation unit 14 may also select no appearance features corresponding to one item. In this case, the generation unit 14 can generate a search query that does not include an item-specific condition for that item.
[0075] Next, an example of the processing flow of the processing device 10 will be described with reference to the flowchart of FIG.
[0076] First, the processing device 10 accepts a user input specifying a person to be tracked in a target frame image, which is one of a plurality of frame images in a time series (S10). Next, the processing device 10 detects the person to be tracked in a plurality of peripheral frame images before and / or after the target frame image (S11). Next, the processing device 10 extracts appearance features related to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images (S12). The processing device 10 then integrates the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each item to generate a search query (S13).
[0077] "Effects" The processing device 10 of this embodiment generates a search query for detecting a tracking target person based on a target frame image specified by the user and multiple surrounding frame images before and / or after it. Specifically, the processing device 10 extracts appearance features of the tracking target person from each of multiple frame images including the target frame image and multiple surrounding frame images. The processing device 10 then integrates the appearance features extracted from the multiple frame images for each item to generate a search query.
[0078] The processing device 10 that generates a search query based on such multiple frame images can generate a search query that can accurately search for a tracked person, even if there is not a single frame image that can accurately extract appearance features for all items. Furthermore, the processing device 10 that integrates by item can more flexibly integrate appearance features extracted from multiple frame images. As a result, a search query that can more accurately search for a tracked person can be generated.
[0079] The processing device 10 also selects at least one of the above-described appearance features 1 to 10 from the appearance features extracted from the multiple frame images, and generates a search query including the selected appearance feature. The processing device 10 can then use different integration methods for each item. Such a processing device 10 can generate a search query that can more accurately search for a tracking target person.
[0080] The processing device 10 of this embodiment differs from the processing device 10 of the second embodiment in the content of the process of selecting appearance features to be included in a search query from appearance features extracted from a plurality of frame images. This will be described in detail below.
[0081] The generation unit 14 calculates an evaluation value for each of a plurality of appearance features extracted from the target frame image and a plurality of peripheral frame images for each item. The generation unit 14 then selects appearance features whose evaluation values satisfy a predetermined condition, and generates a search query including the selected appearance features. The predetermined condition is, but is not limited to, an evaluation value equal to or greater than a threshold. The threshold is any predetermined value. In this way, the generation unit 14 differs from the second embodiment in the content of the process of selecting appearance features to be included in the search query. The method of generating a search query including the selected appearance features is the same as that of the second embodiment.
[0082] The process of selecting appearance features to be included in a search query, more specifically the process of calculating an evaluation value, will be described in detail below.
[0083] The generation unit 14 calculates an evaluation value for each appearance feature of each item based on at least one of the following: - The number of extracted frame images - The reliability of the extraction result output from the estimation model described in the second embodiment - The brightness of the rectangular area in the frame image that includes the person to be tracked - The size of the rectangular area in the frame image that includes the person to be tracked - Whether the person to be tracked overlaps with another person or object in the frame image - From which of the target frame image or multiple surrounding frame images the person to be tracked was extracted - The orientation of the person to be tracked - The chronological order of the extracted frame images within the target frame image and multiple surrounding frame images
[0084] The generation unit 14 calculates the evaluation value based on a predetermined evaluation value calculation model. The evaluation value calculation model may be a function, a table that associates the above content with the evaluation value, or other models.
[0085] The evaluation value calculation model may be configured to calculate a higher evaluation value for appearance features extracted from a larger number of frame images. By evaluating reliable appearance features extracted from a larger number of frame images more highly, a search query that accurately represents the person to be tracked can be generated.
[0086] The evaluation value calculation model may be configured to calculate a higher evaluation value for an appearance feature that has a higher reliability in the extraction result. By evaluating an appearance feature that has a higher reliability in the extraction result more highly, a search query that accurately represents the person to be tracked can be generated.
[0087] The evaluation value calculation model may also be configured to calculate a higher evaluation value for appearance features extracted from a frame image in which the rectangular area containing the tracked person is brighter. It is considered that the brighter the rectangular area, the better the appearance features of the tracked person are represented. By evaluating appearance features extracted from brighter rectangular areas more highly, a search query that accurately represents the tracked person can be generated.
[0088] The evaluation value calculation model may also be configured to calculate a higher evaluation value for appearance features extracted from frame images with larger rectangular areas that include the person to be tracked. The larger the rectangular area, the better the appearance features of the person to be tracked are considered to be represented. By evaluating appearance features extracted from larger rectangular areas more highly, a search query that accurately represents the person to be tracked can be generated.
[0089] Furthermore, the evaluation value calculation model may be configured to calculate a higher evaluation value for appearance features extracted from frame images in which the tracked person is not overlapped with other people or objects than for appearance features extracted from frame images in which the tracked person is overlapped with other people or objects. When the tracked person is not overlapped with other people or objects, the appearance features of the tracked person can be extracted more accurately than when the tracked person is overlapped with other people or objects. By evaluating the appearance features extracted from frame images in which the tracked person is not overlapped with other people or objects more highly, a search query that accurately represents the tracked person can be generated.
[0090] Furthermore, the evaluation value calculation model may be configured to calculate a higher evaluation value for appearance features extracted from a target frame image than for appearance features extracted from surrounding frame images. By evaluating appearance features extracted from a target frame image specified by the user higher than appearance features extracted from surrounding frame images not specified by the user, a search query that is more in line with the user's intentions can be generated.
[0091] Furthermore, the evaluation value calculation model may be configured to calculate a higher evaluation value for appearance features extracted from frame images in which the tracked person is facing a predetermined direction than for appearance features extracted from frame images in which the tracked person is not facing the predetermined direction. The predetermined direction is, for example, but is not limited to, facing forward (facing the camera). When the tracked person is facing the predetermined direction, the appearance features of the tracked person can be extracted with greater accuracy. By evaluating the appearance features extracted from frame images in which the tracked person is facing the predetermined direction more highly, a search query that accurately represents the tracked person can be generated.
[0092] Here, a process for calculating an evaluation value based on "the chronological order of the extracted frame images within the target frame image and the plurality of peripheral frame images" will be described.
[0093] First, in this example, the generation unit 14 generates at least one of a first search query and a second search query. The first search query is a search query used in a process of searching for a tracking target person in a frame image that is chronologically later than the target frame image. The second search query is a search query used in a process of searching for a tracking target person in a frame image that is chronologically earlier than the target frame image.
[0094] Then, in generating the first search query, the generation unit 14 calculates a higher evaluation value for appearance features extracted from a frame image that is later in chronological order within a plurality of frame images including the target frame image and a plurality of surrounding frame images (first calculation process).
[0095] In addition, when generating the second search query, the generation unit 14 calculates a higher evaluation value for appearance features extracted from a frame image that is earlier in chronological order within a plurality of frame images including the target frame image and a plurality of surrounding frame images (second calculation process).
[0096] The generation unit 14 can execute at least one of a first calculation process and a second calculation process.
[0097] The appearance features of a person to be tracked may change over time. For example, the appearance features of a person to be tracked may change due to disguise, leaving behind belongings, transferring belongings, etc. Taking this into consideration, in the process of searching for the person to be tracked in frame images that are later in the chronological order than the target frame image, appearance features extracted from newer frame images are useful. Also, in the process of searching for the person to be tracked in frame images that are earlier in the chronological order than the target frame image, appearance features extracted from older frame images are useful. By calculating the evaluation value taking these points into consideration, a search query that can accurately detect the person to be tracked is generated.
[0098] The other configurations of the processing apparatus 10 are the same as those of the first and second embodiments.
[0099] The processing device 10 of this embodiment achieves the same effects as those of the first and second embodiments. Furthermore, the processing device 10 of this embodiment calculates an evaluation value for each of a plurality of appearance features extracted from the target frame image and a plurality of peripheral frame images using the characteristic technique described above, and selects appearance features to be included in a search query based on the evaluation values. This processing device 10 can generate a search query that can accurately search for a tracking target person.
[0100] Fourth Embodiment The processing device 10 of this embodiment integrates the appearance features of a plurality of items using a plurality of different methods, which will be described in detail below.
[0101] The generation unit 14 integrates the appearance features of multiple items using multiple different methods. The generation unit 14 integrates the appearance features of a first item using a first method and integrates the appearance features of a second item using a second method. The first method and the second method are different. The generation unit 14 may classify the multiple items into two categories, first items and second items, and integrate the appearance features of each item using a method corresponding to each category. Alternatively, the generation unit 14 may classify the multiple items into three or more categories and integrate the appearance features of each item using a method corresponding to each category.
[0102] The classification contents of a plurality of items and the integration methods corresponding to each classification are determined in advance and registered in the processing device 10. Based on this information, the generation unit 14 identifies the classification to which each item belongs and also identifies the integration method corresponding to each classification.
[0103] The integration method corresponding to each classification may be, for example, any of the integration methods described in the second and third embodiments. That is, the integration method corresponding to a certain classification may be any of the integration methods described in the second and third embodiments, and the integration method corresponding to another classification may be any other of the integration methods described in the second and third embodiments.
[0104] Here, another specific example of the integration method corresponding to each classification will be described. In this example, a plurality of items are classified into two categories, a first item and a second item.
[0105] The first items are items such as gender, age, hairstyle, and body type, which are physically impossible for a single person to have multiple appearance characteristics. In a first method corresponding to the first items, the generation unit 14 selects, for each item, at least one from multiple appearance characteristics extracted from the target frame image and multiple peripheral frame images. The generation unit 14 then generates an item-specific condition that includes at least one selected feature, and generates a search query that includes the item-specific condition. The first method is the integration method described in the second and third embodiments.
[0106] The second items are items for which it is physically possible for one person to have multiple appearance characteristics, such as clothing color, clothing design, and the color and design of belongings. In a second method corresponding to the second items, the generation unit 14 generates, for each item, an item-specific condition that includes all of the multiple appearance characteristics extracted from the target frame image and multiple peripheral frame images. For example, the generation unit 14 generates an item-specific condition that connects all of the extracted multiple appearance characteristics with an OR condition. Then, the generation unit 14 generates a search query that includes the item-specific condition.
[0107] The other configurations of the processing apparatus 10 are the same as those of the first to third embodiments.
[0108] The processing device 10 of this embodiment achieves the same effects as those of the first to third embodiments. Furthermore, the processing device 10 of this embodiment can generate a search query by integrating the appearance features of each item using a method suited to the characteristics of each item. The processing device 10 can generate a search query that can accurately search for a tracking target person.
[0109] Fifth Embodiment The processing device 10 of this embodiment determines the frame images to be included in the peripheral frame images based on the analysis results of the moving image, as will be described in detail below.
[0110] When a frame image that is earlier in time series than the target frame image is included in the peripheral frame images, the detection unit 12 detects the target frame image M as shown in FIG. 0 The frame image M immediately before -1 The detection unit 12 determines whether each frame image satisfies a predetermined stop condition while tracing back in time series from the target frame image M 0 The frame image M immediately before -1 Arrow A from 1 The detection unit 12 then determines whether each frame image satisfies a predetermined stop condition while tracing back in the direction of the frame image M. -1 " to "the frame image that is determined to satisfy the stop condition for the Nth time" are determined to be included in the peripheral frame images, where N is an arbitrary integer of 1 or greater.
[0111] Furthermore, when a frame image that is chronologically later than the target frame image is included in the peripheral frame images, the detection unit 12 determines whether the target frame image M 0 Frame image M immediately after 1 In other words, the detection unit 12 determines whether each frame image satisfies a predetermined stop condition in chronological order from the target frame image M 0 Frame image M immediately after 1 Arrow A from 2 Then, the detection unit 12 determines whether each frame image satisfies a predetermined stop condition in the direction of the frame image M. 1" to "the frame image that is determined to satisfy the stop condition for the Nth time" are determined to be included in the peripheral frame images, where N is an arbitrary integer of 1 or greater.
[0112] The detection unit 12 determines frame images to be included in the plurality of peripheral frame images based on at least one of the following characteristics of frame images other than the target frame image. The stop condition is defined by at least one of the following characteristics: Brightness of the rectangular area in the frame image that includes the person to be tracked Size of the rectangular area in the frame image that includes the person to be tracked Whether the person to be tracked overlaps with other people or objects in the frame image Orientation of the person to be tracked in the frame image
[0113] The stop condition may be that the brightness of the rectangular area is equal to or greater than a threshold value, which may be a predetermined arbitrary value.
[0114] Alternatively, the stop condition may be that the size of the rectangular area is equal to or greater than a threshold value, the threshold value being a predetermined arbitrary value.
[0115] Alternatively, the stop condition may be that the person to be tracked does not overlap with other people or objects in the frame image.
[0116] Alternatively, the stop condition may be that the tracking target person faces a predetermined direction in the frame image, such as, but not limited to, facing forward (toward the camera).
[0117] Alternatively, the stop condition may be that, in the determinations made in the frame images immediately before or after the target frame image, it is determined that there are a plurality of frame images in which the tracking target person is oriented in each of a plurality of predetermined orientations within the frame images, such as, but not limited to, facing forward, facing backward, facing right, facing left, etc.
[0118] Alternatively, the stop condition may be a condition in which two or more of the above-mentioned stop conditions are connected by an arbitrary logical operator.
[0119] The other configurations of the processing apparatus 10 are the same as those of the first to fourth embodiments.
[0120] The processing device 10 of this embodiment achieves the same effects as those of the first to fourth embodiments. Furthermore, the processing device 10 of this embodiment can determine frame images to be included in the peripheral frame images based on the analysis results of the moving image. This processing device 10 can generate a search query that can accurately search for a tracking target person.
[0121] The processing device 10 of this embodiment determines whether to perform the above-described "integration of appearance features extracted from a plurality of frame images" for each item, and generates a search query using the determined method. This will be described in detail below.
[0122] For an item in which the reliability of the appearance feature extracted from the target frame specified by the user is equal to or greater than a threshold, the generation unit 14 generates a search query that includes the appearance feature in the item-by-item condition. In other words, for such an item, the generation unit 14 does not "integrate appearance features extracted from multiple frame images."
[0123] On the other hand, for items in which the reliability of the appearance features extracted from the target frame specified by the user is less than the threshold, the generation unit 14 performs "integration of appearance features extracted from multiple frame images." The generation unit 14 can perform this integration using the method described in the above embodiment.
[0124] The reliability is the reliability of the extraction result output from the estimation model described in the second embodiment. The threshold is an arbitrary value that is determined in advance.
[0125] The other configurations of the processing apparatus 10 are the same as those of the first to fifth embodiments.
[0126] The processing device 10 of this embodiment achieves the same effects as the first to fifth embodiments. Furthermore, for items for which highly reliable appearance features have been extracted from a target frame designated by the user, the processing device 10 generates a search query using those appearance features. For items for which highly reliable appearance features have not been extracted from a target frame designated by the user, the processing device 10 generates a search query by integrating appearance features extracted from the target frame image and multiple peripheral frame images. This processing device 10 can generate a search query that is more in line with the user's intentions and that can accurately search for a tracking target person.
[0127] Seventh Embodiment The processing device 10 of this embodiment searches for a person to be tracked using a search query generated by the generation unit 14. This will be described in detail below.
[0128] 8 shows an example of a functional block diagram of the processing device 10 of this embodiment. As shown in the figure, the processing device 10 of this embodiment has a receiving unit 11, a detecting unit 12, an extracting unit 13, a generating unit 14, and a searching unit 15.
[0129] The search unit 15 searches for the tracking target person within the video using the search query generated by the generation unit 14. That is, the search unit 15 searches for a person having the appearance characteristics indicated by the search query within the video. The video to be searched may be a video including the target frame image and surrounding frame images. Alternatively, the video to be searched may be a video not including the target frame image and surrounding frame images.
[0130] The search unit 15 searches within the video for a person who satisfies the conditions of the appearance characteristics indicated in the search query generated by the generation unit 14. The search can be realized using any technology.
[0131] The other configurations of the processing apparatus 10 are the same as those of the first to sixth embodiments.
[0132] The processing device 10 of this embodiment achieves the same effects as those of the first to sixth embodiments. Furthermore, the processing device 10 of this embodiment searches for a tracking target person using a search query generated by the characteristic method described in the above embodiments. With this processing device 10, the tracking target person can be searched for with high accuracy.
[0133] <Modification> This modification is applicable to the first to seventh embodiments. In this modification, the generation unit 14 regenerates a search query using the search results by the search unit 15 of the seventh embodiment.
[0134] First, the generation unit 14 generates a search query using the method described in the first to sixth embodiments. Then, the search unit 15 searches for a person to be tracked using the search query.
[0135] Thereafter, the generation unit 14 regenerates a search query using the search results by the search unit 15. Specifically, the generation unit 14 acquires a frame image to be used for regenerating the search query from among the multiple frame images included in the search results by the search unit 15.
[0136] For example, the generation unit 14 may display on a display a plurality of frame images included in the search results by the search unit 15, and accept a user input specifying a frame image to be used to regenerate a search query from among the frame images. In this case, the generation unit 14 acquires the frame image specified by the user input as the frame image to be used to regenerate the search query. The user specifies a frame image including a person to be tracked from among the plurality of frame images displayed on the display.
[0137] Additionally, the generation unit 14 may calculate the similarity in appearance between a person included in multiple frame images included in the search results and the person to be tracked specified by the user input received by the reception unit 11. The generation unit 14 may then acquire frame images for which the similarity is equal to or greater than a threshold as frame images to be used for regenerating a search query. The threshold is an arbitrary value determined in advance. The similarity is realized using any technology that calculates the similarity in the appearance of a person. For example, a technology that calculates the similarity of a person's face may be used.
[0138] After frame images to be used for regenerating a search query are acquired from the plurality of frame images included in the search results by the search unit 15, the extraction unit 13 extracts appearance features relating to a plurality of items of the person to be tracked from each of the acquired frame images. The generation unit 14 then integrates the appearance features extracted from the target frame image, the plurality of peripheral frame images, and each of the acquired frame images for each item to generate a search query. The integration method is the same as in the first to seventh embodiments.
[0139] The processing device 10 may perform a search again using the regenerated search query, and may regenerate a search query based on the search results. The processing device 10 may then repeat this loop multiple times.
[0140] According to this modification, a search query can be generated based on the appearance of the tracked person in a larger number of frame images, thereby enabling the generation of a search query that can accurately search for the tracked person.
[0141] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations may be adopted. The configurations of the above-described embodiments may be combined with each other, or some of the configurations may be replaced with other configurations. Furthermore, various modifications may be made to the configurations of the above-described embodiments without departing from the spirit of the invention. Furthermore, the configurations and processes disclosed in the above-described embodiments and modified examples may be combined with each other.
[0142] In addition, in the flowcharts used in the above explanation, multiple steps (processes) are described in order. However, the order of execution of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the figures can be changed to the extent that the content is not affected. Furthermore, each of the above-mentioned embodiments can be combined to the extent that the content is not contradictory.
[0143] Some or all of the above embodiments can be described as in the following supplementary notes, but are not limited to the following: 1. A processing device having: a receiving means for receiving a user input specifying a person to be tracked in a target frame image that is one of a plurality of frame images in time series; a detecting means for detecting the person to be tracked in a plurality of peripheral frame images before and / or after the target frame image; an extracting means for extracting appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images; and a generating means for generating a search query by integrating the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each of the items. 2. 2. The processing device according to claim 1, wherein the generation means selects, for each of the items, from the plurality of appearance features extracted from the target frame image and the plurality of peripheral frame images, at least one of the following: the appearance feature extracted from the largest number of frame images; the appearance feature extracted from a predetermined percentage or more of frame images; the appearance feature extracted from a predetermined number or more of frame images; the appearance feature having the highest reliability of the extraction result; the appearance feature having a reliability of the extraction result equal to or greater than a threshold; the appearance feature extracted from the frame image in which the brightness of a rectangular area containing the person to be tracked within the frame image is the brightest; the appearance feature extracted from the frame image in which the brightness of a rectangular area containing the person to be tracked within the frame image is equal to or greater than a threshold; the appearance feature extracted from the frame image in which the rectangular area containing the person to be tracked within the frame image is the largest in size; the appearance feature extracted from the frame image in which the size of a rectangular area containing the person to be tracked within the frame image is equal to or greater than a threshold; and the appearance feature extracted from the frame image in which the person to be tracked does not overlap with other people or objects within the frame image; and generates the search query including the selected appearance feature.3. The processing device according to 1, wherein the generation means calculates an evaluation value for each of the plurality of appearance features extracted from the target frame image and the plurality of peripheral frame images for each of the items, generates the search query including the appearance features whose evaluation values satisfy a predetermined condition, and calculates the evaluation value based on at least one of the number of extracted frame images, the reliability of the extraction result, the brightness of a rectangular area in a frame image including the person to be tracked, the size of a rectangular area in a frame image including the person to be tracked, whether the person to be tracked overlaps with another person or object in the frame image, from which of the target frame image and the plurality of peripheral frame images the person to be tracked was extracted, and the chronological order of the extracted frame images within the target frame image and the plurality of peripheral frame images. The processing device according to 3, wherein the generation means executes at least one of: in generating the search query to search for the tracking target person in frame images that are chronologically later than the target frame image, a process of calculating the evaluation value so that the appearance feature extracted from a frame image that is chronologically later among the target frame image and the plurality of peripheral frame images is higher; and in generating the search query to search for the tracking target person in frame images that are chronologically earlier than the target frame image, a process of calculating the evaluation value so that the appearance feature extracted from a frame image that is chronologically earlier among the target frame image and the plurality of peripheral frame images is higher. 5. The processing device according to 4, wherein the generation means generates at least one of the search query used in the process to search for the tracking target person in frame images that are chronologically later than the target frame image, and the search query used in the process to search for the tracking target person in frame images that are chronologically earlier than the target frame image. 6. The processing device according to any one of 1 to 5, wherein the generation means integrates the appearance features of a first item using a first method and integrates the appearance features of a second item using a second method, and the first method and the second method are different.7. The processing device according to 6, wherein the generation means, in the first method, selects at least one from the plurality of appearance features extracted from the target frame image and the plurality of peripheral frame images, and generates the search query including at least one selected feature, and in the second method, generates the search query including the plurality of appearance features extracted from the target frame image and the plurality of peripheral frame images. 8. The processing device according to any of 1 to 7, wherein the detection means determines frame images to be included in the plurality of peripheral frame images based on at least one of the following in frame images other than the target frame image: brightness of a rectangular area including the person to be tracked in a frame image; size of a rectangular area including the person to be tracked in a frame image; whether the person to be tracked overlaps with another person or object in the frame image; and orientation of the person to be tracked in the frame image. 9. 10. A processing method in which one or more computers accept user input specifying a person to be tracked in a target frame image that is one of a plurality of frame images in a time series, detect the person to be tracked in a plurality of peripheral frame images before and / or after the target frame image, extract appearance features related to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images, and generate a search query by integrating the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each of the items. 10. A program that causes a computer to function as: accepting means for accepting user input specifying the person to be tracked in a target frame image that is one of a plurality of frame images in a time series, detecting means for detecting the person to be tracked in a plurality of peripheral frame images before and / or after the target frame image, extracting means for extracting appearance features related to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images, and generating means for integrating the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each of the items to generate a search query.
[0144] This application claims priority based on Japanese Patent Application No. 2023-041792, filed March 16, 2023, the disclosure of which is incorporated herein in its entirety.
[0145] 10 Processing device 11 Reception unit 12 Detection unit 13 Extraction unit 14 Generation unit 15 Search unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus
Claims
1. a receiving means for receiving a user input specifying a person to be tracked within a target frame image that is one of a plurality of frame images in time series; a detection means for detecting the person to be tracked in a plurality of peripheral frame images before and / or after the target frame image; extraction means for extracting appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images; a generation means for generating a search query by integrating the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each of the items; A processing device having:
2. The generating means For each item, from among the plurality of appearance features extracted from the target frame image and the plurality of peripheral frame images, the appearance features extracted from the most frame images; the appearance features extracted from a predetermined percentage or more of the frame images; the appearance features extracted from a predetermined number or more of frame images; The appearance feature having the highest reliability of the extraction result; The appearance feature whose reliability of the extraction result is equal to or greater than a threshold value; the appearance feature extracted from the frame image in which the brightness of a rectangular region including the person to be tracked in the frame image is the brightest; the appearance feature extracted from the frame image in which the brightness of a rectangular region including the person to be tracked in the frame image is equal to or greater than a threshold; the appearance features extracted from the frame image having the largest rectangular region including the person to be tracked; the appearance features extracted from the frame images in which the size of a rectangular area including the person to be tracked in the frame images is equal to or greater than a threshold; and the appearance features extracted from a frame image in which the person to be tracked does not overlap with other people or objects in the frame image; and generating the search query including the selected appearance feature.
3. The generating means calculating an evaluation value for each of the plurality of appearance features extracted from the target frame image and the plurality of peripheral frame images for each of the items, and generating the search query including the appearance features whose evaluation values satisfy a predetermined condition; 2. The processing device according to claim 1, wherein the evaluation value is calculated based on at least one of the number of extracted frame images, the reliability of the extraction result, the brightness of the rectangular area in the frame image including the person to be tracked, the size of the rectangular area in the frame image including the person to be tracked, whether the person to be tracked overlaps with another person or object in the frame image, whether the person to be tracked is extracted from the target frame image or the plurality of surrounding frame images, and the chronological order of the extracted frame images within the target frame image and the plurality of surrounding frame images.
4. The generating means a process of calculating, in the generation of the search query for searching for the tracking target person in a frame image that is chronologically later than the target frame image, a higher evaluation value for the appearance feature extracted from a frame image that is chronologically later among the target frame image and the plurality of peripheral frame images; and a process of calculating, in the generation of the search query for searching for the tracking target person in a frame image that is earlier in time series than the target frame image, a higher evaluation value for the appearance feature extracted from a frame image that is earlier in time series among the target frame image and the plurality of peripheral frame images; 4. The processing device according to claim 3, wherein the processing device executes at least one of the above.
5. The generating means The processing device according to claim 4, wherein the processing device generates at least one of the search query used in a process of searching for the person to be tracked in a frame image that is chronologically later than the target frame image, and the search query used in a process of searching for the person to be tracked in a frame image that is chronologically earlier than the target frame image.
6. the generating means integrates the appearance features of a first item using a first technique and integrates the appearance features of a second item using a second technique; The processing device according to claim 1 , wherein the first technique and the second technique are different.
7. The generating means In the first method, at least one of the plurality of appearance features extracted from the target frame image and the plurality of peripheral frame images is selected, and the search query is generated including the selected at least one appearance feature; The processing device according to claim 6 , wherein the second method generates the search query including a plurality of the appearance features extracted from the target frame image and the plurality of peripheral frame images.
8. The detection means In a frame image other than the target frame image, the brightness of a rectangular area including the person to be tracked in the frame image; the size of a rectangular area including the person to be tracked in the frame image; Whether the person to be tracked overlaps with another person or object in the frame image; and the orientation of the person to be tracked in the frame image; The processing device according to claim 1 , wherein frame images to be included in the plurality of peripheral frame images are determined based on at least one of the following:
9. One or more computers Accepting a user input specifying a person to be tracked within a target frame image that is one of a plurality of frame images in time series; Detecting the person to be tracked in a plurality of surrounding frame images before and / or after the target frame image; extracting appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images; A processing method for generating a search query by integrating the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each item.
10. Computer, a receiving means for receiving a user input specifying a person to be tracked within a target frame image, which is one of a plurality of frame images in time series; a detection means for detecting the person to be tracked in a plurality of surrounding frame images before and / or after the target frame image; an extraction means for extracting appearance features relating to a plurality of items of the person to be tracked from the target frame image and each of the plurality of peripheral frame images; a generation means for generating a search query by integrating the appearance features extracted from the target frame image and each of the plurality of peripheral frame images for each of the items; A program that functions as a