Image processing device, image processing method, and program
The image processing apparatus enhances image search accuracy by incorporating posture and appearance information to correct search queries, ensuring precise image retrieval.
Patent Information
- Application Number
- JP2022131660
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-08-22
AI Technical Summary
Existing image search technologies suffer from low accuracy in identifying and retrieving images of a desired person due to insufficient utilization of posture and appearance information, leading to inaccurate search results.
An image processing apparatus and method that incorporates posture and appearance information to correct and refine search queries, enabling precise image searches by using corrected search queries to identify and retrieve target images.
Improves the accuracy of image search by utilizing posture and appearance information to correct search queries, allowing for high-precision retrieval of desired images.
Smart Images

Figure 0007910391000001 
Figure 0007910391000002 
Figure 0007910391000003
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.
Background Art
[0002] Technologies related to the present invention are disclosed in Patent Documents 1 to 3.
[0003] Patent Document 1 discloses performing image search using posture information and other information about a person. And it is disclosed that the other information is color information of a person or a wearable (which may be by part of the person), face information, gender, age group, body type, position in the image, etc.
[0004] Patent Document 2 discloses a technique for tracking a person in a video based on the color, pattern, shape, height, aspect ratio, etc. of the person.
[0005] Patent Document 3 discloses a technique for determining whether a person is a single mover or a group mover based on the face image, physique, age, gender, clothing, etc. of the person.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0007] By performing image search using a wide variety of information (gender, age, body type, posture, clothing, position in the image, etc.), it becomes possible to search for an image containing a desired person with high accuracy.
[0008] By the way, one possible method for inputting a search query is to input an image. In this case, by analyzing the input image, a wide variety of information that the image represents is identified, and this identification result is used to search for images.
[0009] In the case of this technology, if the accuracy of identifying the various types of information realized through the analysis of the input image is low, the accuracy of image search using the identification results will also be low. Patent documents 1 to 3 do not disclose this problem or any means of solving it.
[0010] One example of the object of the present invention is to provide an image processing apparatus, an image processing method, and a program that solve the problem of improving the accuracy of the process of searching for an image containing a desired person from among multiple images, in view of the problems described above. [Means for solving the problem]
[0011] According to one aspect of the present invention, A means for obtaining a search query that includes posture information showing a person's posture and appearance information showing a person's appearance, Correction means for correcting the other using one of the posture information and the appearance information, A search means that searches for a target image from among multiple reference images using the corrected search query, An image processing device having the following is provided.
[0012] According to one aspect of the present invention, One or more computers, A search query is obtained that includes posture information showing a person's stance and appearance information showing a person's appearance. Using one of the aforementioned posture information and the aforementioned appearance information, the other is corrected. An image processing method is provided that uses the corrected search query to search for a target image from among multiple reference images.
[0013] According to one aspect of the present invention, Computers, An acquisition means for acquiring a search query including pose information indicating a person's pose and appearance information indicating the person's appearance, A correction means for correcting one of the pose information and the appearance information using the other, A search means for searching for a target image from a plurality of reference images using the corrected search query, A program is provided that functions as such.
Effect of the Invention
[0014] According to one aspect of the present invention, an image processing apparatus, an image processing method, and a program are realized that solve the problem of improving the accuracy of processing for searching for an image including a desired person from a plurality of images.
Brief Description of the Drawings
[0015] The above-described object, as well as other objects, features, and advantages, will become more apparent from the following public embodiments and the accompanying drawings below.
[0016] [Figure 1] It is a diagram showing an example of a functional block diagram of an image processing apparatus. [Figure 2] It is a diagram showing an example of a hardware configuration of an image processing apparatus. [Figure 3] It is a diagram for explaining an example of keypoints. [Figure 4] It is a diagram schematically showing an example of information processed by an image processing apparatus. [Figure 5] It is a flowchart showing an example of the processing flow of an image processing apparatus. [Figure 6] It is a diagram showing an example of a screen output by an image processing apparatus. [Figure 7] It is a diagram showing another example of a screen output by an image processing apparatus. [Figure 8] It is a flowchart showing another example of the processing flow of an image processing apparatus. [Figure 9] It is a flowchart showing another example of the processing flow of an image processing apparatus. [Figure 10] This figure shows another example of a functional block diagram of an image processing device. [Figure 11] Here is another example flowchart of the processing flow of an image processing device. [Modes for carrying out the invention]
[0017] Embodiments of the present invention will be described below with reference to the drawings. In all drawings, similar components are denoted by the same reference numerals, and their descriptions are omitted as appropriate.
[0018] <First Embodiment> Figure 10 is a functional block diagram showing an overview of the image processing apparatus 10 according to the first embodiment. The image processing apparatus 10 includes an acquisition unit 11, a search unit 12, and a correction unit 13.
[0019] The acquisition unit 11 acquires a search query that includes posture information indicating a person's posture and appearance information indicating a person's appearance. The correction unit 13 corrects the other using either the posture information or the appearance information. The search unit 12 searches for the target image from among multiple reference images using the corrected search query.
[0020] The image processing device 10 with this configuration solves the problem of improving the accuracy of the process of searching for an image containing a desired person from among multiple images.
[0021] <Second Embodiment> "overview" By performing image searches using a wide variety of information, it becomes possible to search for images containing a desired person with high accuracy. The image processing device 10 of this embodiment searches for images containing a desired person with high accuracy by performing image searches using characteristic information that has not been used in the prior art. Specifically, the image processing device 10 performs image searches based on appearance information and posture information that indicate at least one of the following: "whether or not a predetermined type of attachment is being worn," "whether or not a predetermined type of attachment is being worn on a predetermined part of the body," "whether or not an attachment with a predetermined pattern is being worn," and "whether or not an attachment with a predetermined pattern is being worn on a predetermined part of the body." This will be explained in detail below.
[0022] "Hardware configuration" Next, an example of the hardware configuration of the image processing device 10 will be described. Each functional unit of the image processing device 10 is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, a program loaded into memory, a storage unit such as a hard disk that stores that program (which can store programs that are pre-installed at the time of shipment, as well as programs downloaded from recording media such as CDs (Compact Discs) or from servers on the Internet), and a network connection interface. It will be understood by those skilled in the art that there are various modifications to the implementation method and the device.
[0023] Figure 2 is a block diagram illustrating the hardware configuration of the image processing device 10. As shown in Figure 2, the image processing device 10 includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. Peripheral circuitry 4A includes various modules. The image processing device 10 does not necessarily have peripheral circuitry 4A. The image processing device 10 may also be composed of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.
[0024] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuits 4A, and input / output interface 3A to send and receive data to and from each other. Processor 1A is a processing unit such as a CPU or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). Input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. Input devices include, for example, keyboards, mice, microphones, physical buttons, touch panels, etc. Output devices include, for example, displays, speakers, printers, mailers, etc. Processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0025] "Functional Configuration" Next, the functional configuration of the image processing apparatus 10 of this embodiment will be described in detail. Figure 1 shows an example of a functional block diagram of the image processing apparatus 10. As shown in the figure, the image processing apparatus 10 has an acquisition unit 11 and a search unit 12.
[0026] The acquisition unit 11 acquires a search query that includes posture information indicating the posture of a person's body and appearance information indicating the person's appearance. The acquisition unit 11 may acquire an image as the search query, or it may acquire text or numerical data. When an image is acquired, the posture and appearance of the person contained in the image become the posture information and appearance information. When text or numerical data is acquired, the posture and appearance of the person are indicated by the text or numerical data. The acquisition unit 11 may also analyze the image acquired as the search query and generate a search query for text or numerical data.
[0027] "Posture information" indicates the posture of a person's body. A person's body posture can be classified, for example, as standing, sitting, or lying down. In addition, a person's body posture can be classified as standing with the right arm raised, standing with the left arm raised, etc. There are various ways to classify a person's body posture. Posture information may indicate any one of these posture classifications.
[0028] In addition, posture information may also include information about key points of the person's body. Examples of key points of the person's body include, for example, the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, left foot A82, etc. Key point detection is achieved using well-known techniques such as Openpose. Examples of information about key points include, but are not limited to, information indicating the relative positional relationship between multiple key points.
[0029] "Appearance information" refers to the person's physical appearance.
[0030] External information is, • Whether or not a specified type of attachment is being worn. • Whether or not a specified type of attachment is attached to a specified part of the body. • Whether or not an attachment with the specified pattern is attached, and • Whether or not an attachment with a specific pattern is attached to a specific part of the body. Show at least one of these. A combination of two or more of these may be used as appearance information.
[0031] "Wearings" are items worn by people. Examples of wearings include, but are not limited to, eyeglasses, sunglasses, hats, masks, watches, headphones, scarves, gloves, coats, shirts, trousers, skirts, shoes, sandals, and slippers. Furthermore, the types of wearings exemplified here may be further subdivided. For example, coats may be subdivided according to their design, such as trench coats, duffel coats, and pea coats.
[0032] "Specific body parts" are body parts that can be identified based on key points of the body detected using well-known techniques such as Openpose as described above. For example, as shown in Figure 3, examples include the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, left foot A82, etc. Other examples of specific body parts include body parts defined by grouping multiple key points, such as "head," "upper body," "lower body," "right half," "left half," "upper right half," "upper left half," "lower right half," and "lower left half."
[0033] Furthermore, the appearance information may include other information in addition to the information described above. This other information can include facial information, gender, age group, body type, position within the image, clothing color, and any other known technology.
[0034] Here, we will explain how to obtain a search query that includes the posture information and appearance information described above. The acquisition unit 11 can obtain the search query using, for example, one of the following first to third acquisition examples.
[0035] -First acquisition example- In this example, the acquisition unit 11 acquires a still image as a search query. The acquisition unit 11 then analyzes the still image to detect people in the image, as well as their posture and appearance. Person detection, posture detection, and appearance detection can be achieved using any conventional technology. Thus, in this example, the acquisition unit 11 analyzes the still image acquired as a search query and generates search queries for text and numerical data.
[0036] -Second acquisition example- In this example, the acquisition unit 11 acquires video footage as a search query. The acquisition unit 11 selects a representative frame image from the acquired video footage. The acquisition unit 11 then analyzes the representative frame image to detect people in the image, as well as their posture and appearance. Person detection, posture detection, and appearance detection can be achieved using any conventional technology. Furthermore, any method can be used to select the representative frame image. Thus, in this example, the acquisition unit 11 analyzes the acquired video footage as a search query and generates search queries for text and numerical data.
[0037] -Third acquisition example- In this example, the acquisition unit 11 acquires search queries that specify a person's posture and appearance using text or numerical data. For example, the acquisition unit 11 may accept user input by selecting one or more options from predetermined choices (posture options, appearance options) using UI (user interface) components such as a dropdown list. Alternatively, the acquisition unit 11 may accept user input by specifying posture and appearance in free text using UI components such as a text box. If free text specification is adopted, the acquisition unit 11 may use a pre-prepared word conversion dictionary to change the user's description to content suitable for processing by the search unit 12.
[0038] The search unit 12 searches for the target image from among multiple reference images using the posture information and appearance information included in the search query. The search unit 12 compares the search query with the multiple reference images stored in the memory unit and searches for the target image based on the comparison result. The multiple reference images may be multiple frame images included in a video, or multiple still images.
[0039] The "target image" is an image that includes the target person. The target person is a person who is in the posture indicated by the posture information in the search query and has the appearance indicated by the appearance information in the search query. For example, if the posture information in the search query is "standing posture" and the appearance information in the search query is "a male in his 30s wearing a watch on his left wrist and red pants," then the target person would be "a male in his 30s wearing a watch on his left wrist, wearing red pants, and in a standing posture." An image containing such a target person would then be the target image.
[0040] The matching method and the criteria for determining a match are design matters, and any known technology can be used. For example, image analysis may be performed on each of the multiple reference images in advance, and the posture and appearance of the person contained in each reference image may be identified. Then, as shown in Figure 4, information linking the posture information and appearance information of the person contained in each of the multiple reference images (reference image identification information) may be stored in the memory unit. The search unit 12 may then compare the posture information and appearance information contained in the search query with the posture information and appearance information of the person contained in each reference image stored in the memory unit.
[0041] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 5.
[0042] First, the image processing device 10 obtains a search query including posture information and appearance information (S10). In this embodiment, the appearance information indicates at least one of the following: "whether or not a predetermined type of attachment is being worn," "whether or not a predetermined type of attachment is being worn on a predetermined part of the body," "whether or not an attachment with a predetermined pattern is being worn," and "whether or not an attachment with a predetermined pattern is being worn on a predetermined part of the body." In addition to at least one of the above, the appearance information may also indicate other information.
[0043] Subsequently, the image processing device 10 searches for the target image from among multiple reference images using the posture information and appearance information included in the search query (S11). The target image is an image that includes the target person. The target person is in the posture indicated by the posture information of the search query and has the appearance indicated by the appearance information of the search query. The image processing device 10 then outputs the retrieved target image as a search result.
[0044] "Effects and Effects" According to the image processing device 10 of this embodiment, a desired target image is searched using both appearance information and posture information that indicates at least one of the following: "whether or not a predetermined type of attachment is being worn," "whether or not a predetermined type of attachment is being worn on a predetermined part of the body," "whether or not an attachment with a predetermined pattern is being worn," and "whether or not an attachment with a predetermined pattern is being worn on a predetermined part of the body." By performing an image search using both characteristic appearance information and posture information that have not been used in conventional image searches, it becomes possible to search for a desired target image with high accuracy.
[0045] Furthermore, according to the image processing device 10 of this embodiment, image searches can be performed using appearance information indicating "whether or not a predetermined type of attachment is attached to a predetermined part of the body" or "whether or not an attachment with a predetermined pattern is attached to a predetermined part of the body." In other words, by focusing on a predetermined part of a person, image searches can be performed based on whether or not a predetermined type of attachment is attached to that part, or whether or not an attachment with a predetermined pattern is attached.
[0046] By the way, Reference 1 discloses image retrieval based on color information of different body parts. However, color information is affected by the environment at the time of shooting (outdoors, indoors, brightness of light, weather, etc.). Therefore, the technology disclosed in Reference 1 may result in low accuracy in image retrieval.
[0047] As in this embodiment, the accuracy of image searches is improved by determining, for each predetermined part, whether or not a predetermined type of attachment is being worn or whether or not an attachment with a predetermined pattern is being worn, which are not significantly affected by the environment during shooting.
[0048] Furthermore, the "predetermined body part" in this embodiment can be defined as identifying the left or right side of the body, such as "right hand, left hand, right arm, left arm, right leg, or left leg." By detecting key points on the body as described above, it becomes possible to identify the left or right side of the body. By identifying the left and right side of the body in this way and performing an image search while considering attachments to each part, it becomes possible to search for the desired target image with high accuracy.
[0049] <Third Embodiment> "overview" By performing image searches using a wide variety of information, such as posture information and appearance information as described in the second embodiment, it becomes possible to search for desired target images with high accuracy. Incidentally, one example of a method for inputting a search query is to input an image. In this case, if an image in which some of the above-mentioned wide variety of information does not match the desired content is used as the search query, there is a problem that the accuracy of searching for images containing the desired person will be low. This problem can be solved by using an image in which all of the above-mentioned wide variety of information matches the desired content as the search query. However, finding an image in which all of the above-mentioned wide variety of information matches the desired content requires a lot of effort.
[0050] Therefore, the image processing device 10 of this embodiment has a function to perform image searches using a portion of the pose information and appearance information included in the acquired search query that is specified by the user. With such an image processing device 10, the user can input images in which some of the diverse information is of the desired content but other parts are not as search queries, and instruct the device to perform an image search using only the information that is of the desired content, thereby enabling high-precision retrieval of the desired target image. As a result, images in which some of the diverse information is of the desired content but other parts are not can also be used as search queries. This will be explained in detail below.
[0051] "Hardware configuration" An example of the hardware configuration of the image processing apparatus 10 in this embodiment is the same as that described in the second embodiment.
[0052] "Functional Configuration" Next, the functional configuration of the image processing apparatus 10 of this embodiment will be described in detail. Figure 1 shows an example of a functional block diagram of the image processing apparatus 10. As shown in the figure, the image processing apparatus 10 has an acquisition unit 11 and a search unit 12.
[0053] The configuration of the acquisition unit 11 is the same as that described in the second embodiment.
[0054] However, the definition of "appearance information" in this embodiment differs somewhat from that in the second embodiment. As described above, in the second embodiment, "appearance information" refers to at least one of the following: "whether or not a predetermined type of attachment is being worn," "whether or not a predetermined type of attachment is being worn on a predetermined part of the body," "whether or not an attachment with a predetermined pattern is being worn," and "whether or not an attachment with a predetermined pattern is being worn on a predetermined part of the body." In addition, it may also include other information such as facial information, gender, age group, body type, position in the image, and clothing color.
[0055] The appearance information in this embodiment is any information that indicates the appearance of a person, and is not limited to indicating at least one of the following, as in the second embodiment: "whether or not a predetermined type of garment is being worn," "whether or not a predetermined type of garment is being worn on a predetermined part of the body," "whether or not a garment with a predetermined pattern is being worn," and "whether or not a garment with a predetermined pattern is being worn on a predetermined part of the body." In other words, the appearance information in this embodiment may indicate at least one of the following: "whether or not a predetermined type of garment is being worn," "whether or not a predetermined type of garment is being worn on a predetermined part of the body," "whether or not a garment with a predetermined pattern is being worn," and "whether or not a garment with a predetermined pattern is being worn on a predetermined part of the body," or in addition to or instead of these, other information such as facial information, gender, age group, body type, position in the image, and clothing color may be indicated.
[0056] The search unit 12 uses partial posture information, which indicates the posture of a part of a person's body, from the posture information included in the search query, and partial appearance information, which is a part of the appearance information included in the search query, to search for the target image from among multiple reference images.
[0057] The search unit 12 can retrieve partial posture information and partial appearance information from the search query, for example, by processing one of the first to third retrieval examples described below.
[0058] -Example of extraction (first example)- First, the acquisition unit 11 analyzes the image input as a search query to detect a person in the image, and then analyzes the entire area of the image in which the person is pictured to detect the posture and appearance of the detected person.
[0059] Then, as shown in Figure 6, the search unit 12 presents the posture and appearance of the person detected in the above process to the user (e.g., on the screen) and accepts user input to specify the information to be used for image search. User input can be provided via any input device such as a touch panel, physical buttons, microphone, keyboard, or mouse.
[0060] For example, the user inputs a selection of key points (indicated by black circles in the figure) of the posture information shown in Figure 6, specifying which key points to use for image search. The search unit 12 then generates partial posture information indicated by the specified key points.
[0061] Furthermore, the user inputs key points to be used for image search from among the various appearance information listed as shown in Figure 6. The search unit 12 then generates partial appearance information that shows the specified portion of the appearance information.
[0062] -Second example of extraction- The acquisition unit 11 displays the image entered as a search query on the screen. Then, as shown in Figure 7, the acquisition unit 11 accepts user input to specify a portion of the image. In Figure 7, the portion enclosed by frame W is specified. The user specifies a portion of the image so as to include the area where the desired information is displayed and exclude the area where the unwanted information is displayed. Input to specify a portion of the image can be implemented using any known technology.
[0063] The acquisition unit 11 analyzes the image of a region specified by the user within the image entered as a search query to detect the posture and appearance of a person. The search unit 12 then acquires partial posture information, which shows the posture of the person detected by analyzing a portion of the image, and partial appearance information, which shows the appearance of the person, from the acquisition unit 11.
[0064] -Third extraction example- The acquisition unit 11 acquires the image entered as a search query and accepts user input specifying a part of the body using a method different from the "method of specifying a part of the image." For example, the acquisition unit 11 presents the user with selectable names indicating parts of the body, such as "upper body," "lower body," "right half," "left half," "upper right half," "upper left half," "lower right half," and "lower left half," and accepts user input to select one of them. This presentation and acceptance of user input can be implemented using well-known UI (user interface) components such as dropdown lists. Alternatively, the acquisition unit 11 may accept user input specifying a part of the body in free text using UI components such as text boxes. If free text specification is adopted, the acquisition unit 11 may use a pre-prepared word conversion dictionary to change the user's description to content suitable for processing by the acquisition unit 11.
[0065] The acquisition unit 11 then uses the posture information (information on detected key points of the body) obtained by analyzing the image to identify the region within the image where the body part specified by the user is located. For example, if "upper body" is specified by the user, the acquisition unit 11 identifies the region within the image that includes the locations where key points such as head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, and left hand A52 are detected, but does not include the locations where other key points are detected. Information indicating the correspondence between body parts and key points contained within those parts may be stored in the image processing device 10 in advance.
[0066] Next, the acquisition unit 11 analyzes the identified portion of the image input as a search query to detect the person's posture and appearance. Then, the search unit 12 acquires from the acquisition unit 11 partial posture information, which shows the person's posture, and partial appearance information, which shows the person's appearance, which were detected by analyzing the portion of the image.
[0067] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 8.
[0068] When the image processing device 10 receives a search query that includes posture information and appearance information (S20), it obtains partial posture information, which shows the posture of a part of a person's body from the posture information included in the search query, and partial appearance information, which is a part of the appearance information included in the search query, based on the obtained search query (S21). Then, the image processing device 10 uses the obtained partial posture information and partial appearance information to search for the target image from among multiple reference images (S22). Next, the image processing device 10 outputs the search results.
[0069] Next, we will explain an example of the process in S21 using the flowchart in Figure 9. Here, we will explain an example of the process flow for the third extraction example described above.
[0070] When the image processing device 10 receives user input specifying a part of a person's body (S30), it identifies the region in the image obtained as a search query where the part of the body specified in the input of S30 is located, based on the posture information (S31). The image processing device 10 then analyzes the image of the region identified in S31 within the image obtained as a search query and obtains partial posture information indicating the posture of the part of the body, and partial appearance information indicating the appearance of the human body (S32).
[0071] Other configurations of the image processing apparatus 10 in this embodiment are the same as those in the first and second embodiments.
[0072] "Effects and Effects" The image processing apparatus 10 of this embodiment achieves the same effects and advantages as those of the first and second embodiments.
[0073] Furthermore, the image processing device 10 of this embodiment can perform image searches using a portion of the pose information and appearance information contained in an image acquired as a search query, as specified by the user. With such an image processing device 10, the user can input an image as a search query in which some of the diverse information is of the desired content but other parts are not, and instruct the device to perform an image search using only the information that is of the desired content, thereby enabling high-precision retrieval of the desired target image. As a result, images in which some of the diverse information is of the desired content but other parts are not can also be used as search queries.
[0074] <Fourth Embodiment> "overview" As described in the second and third embodiments, one possible method for inputting a search query is to input an image. In this case, by analyzing the input image, a wide variety of information that the image represents is identified, and the image is then searched using the results of that identification.
[0075] By the way, when identifying diverse information through image analysis, there is a possibility that the identified information may contain errors. And if image searches are performed using such erroneous information, the accuracy of the image search will decrease.
[0076] Therefore, the image processing device 10 of this embodiment corrects the pose information and appearance information identified by the analysis of the image input as a search query, and then searches for the target image from among multiple reference images using the corrected search query. In this way, the image processing device 10 of this embodiment, which appropriately corrects the search query based on the relationship between pose information and appearance information, improves the accuracy of the search query, and as a result, the accuracy of the image search is also improved. A detailed explanation follows below.
[0077] "Hardware configuration" An example of the hardware configuration of the image processing apparatus 10 in this embodiment is the same as that described in the second embodiment.
[0078] "Functional Configuration" Next, the functional configuration of the image processing apparatus 10 of this embodiment will be described in detail. Figure 10 shows an example of a functional block diagram of the image processing apparatus 10. As shown in the figure, the image processing apparatus 10 has an acquisition unit 11, a search unit 12, and a correction unit 13.
[0079] The configuration of the acquisition unit 11 is the same as in any of the first to third embodiments. The acquisition unit 11 acquires an image as a search query, analyzes the image, and generates a search query (text or numerical data) that includes posture information and appearance information.
[0080] The correction unit 13 corrects the orientation information and appearance information contained in the search query (text or numerical data) acquired by the acquisition unit 11 using one of them. That is, the correction unit 13 can correct the appearance information using the orientation information contained in the search query acquired by the acquisition unit 11. Also, the correction unit 13 can correct the orientation information using the appearance information contained in the search query acquired by the acquisition unit 11. An example of the correction process is described below.
[0081] -Example of processing to correct appearance information using posture information- The posture information shows the detection results of multiple key points on the person's body. The correction unit 13 then corrects the appearance information based on the key point detection results.
[0082] For example, if a predetermined key point is not detected, the correction unit 13 can delete information of a type that is pre-associated with that key point from the appearance information. In this example, a key point on the body is pre-associated with a predetermined type of appearance information. For example, "Key point: Head A1" may be associated with appearance information related to items worn on the face or head, such as information about hats, glasses, sunglasses, and masks. In this case, if "Key point: Head A1" is not detected, the correction unit 13 deletes information related to items worn on the face or head, such as information about hats, glasses, sunglasses, and masks, from the appearance information.
[0083] As in this example, by linking key points for each body part and associating the appearance information of attachments worn on each part with that key point, if a key point for a body part is not detected, the information about the attachment worn on that part can be removed from the appearance information.
[0084] As another example, if a predetermined percentage (design-related) or more of the key points included in a part of a person's body are not detected, the correction unit 13 can delete information of a type pre-associated with that part of the person's body from the appearance information. In this example, a predetermined type of appearance information is pre-associated with a part of the body (upper body, lower body, right half, left half, upper right half, upper left half, lower right half, lower left half, etc.). For example, appearance information related to items worn on the upper body, such as information about hats, information about glasses, information about sunglasses, information about masks, information about gloves, information about scarves, information about jackets, etc., may be associated with "upper body". In this case, if a predetermined percentage or more of the key points included in the upper body are not detected, the correction unit 13 deletes information related to items worn on the upper body, such as information about hats, information about glasses, information about sunglasses, information about masks, information about gloves, information about scarves, information about jackets, etc., from the appearance information.
[0085] As in this example, by linking information about the attachments worn on each part of the body to specific body parts, if a predetermined percentage or more of the multiple key points in a particular part of the body are not detected, information about the attachments worn on that part can be removed from the visual information.
[0086] -Example of processing to correct posture information using visual information- In this example, key points on the body are pre-associated with a predetermined type of appearance information. For example, "Key point: Head A1" may be associated with appearance information related to items worn on the face or head, such as information about hats, glasses, sunglasses, or masks.
[0087] Then, the correction unit 13 deletes information relating to a part of a person's body from the posture information if the reliability of the information relating to a part of a person's body in the appearance information satisfies predetermined conditions.
[0088] If there are multiple types of external information relating to a part of a person's body, the correction unit 13 deletes information relating to that part of the person's body from the posture information if at least one of the reliability levels satisfies a predetermined condition, if all of the reliability levels satisfies a predetermined condition, or if a predetermined percentage or more of the reliability levels satisfies a predetermined condition.
[0089] For example, if the confidence level of at least one piece of appearance information related to an item worn on the face or head, such as information about a hat, glasses, sunglasses, or a mask, is below a certain threshold (under predetermined conditions), the correction unit 13 can delete the information related to "Key Point: Head A1" from the posture information.
[0090] As another example, if the reliability of all appearance information related to items worn on the face or head, such as information about hats, glasses, sunglasses, and masks, is below a standard value (under predetermined conditions), the correction unit 13 can delete the information related to "Key Point: Head A1" from the posture information.
[0091] As another example, if the confidence level of a certain percentage or more of the appearance information related to items worn on the face or head, such as information about hats, glasses, sunglasses, and masks, is below a certain threshold (a predetermined condition), the correction unit 13 can delete the information related to "Key Point: Head A1" from the posture information.
[0092] Confidence level is a value that indicates the degree of confidence in the results of image analysis, and any well-known technique can be used to calculate it.
[0093] The search unit 12 searches for the target image from among multiple reference images using the corrected search query.
[0094] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 11.
[0095] First, the image processing device 10 obtains a search query that includes posture information indicating the person's posture and appearance information indicating the person's appearance (S40). Specifically, the image processing device 10 obtains an image as a search query, analyzes the image, and generates a search query for text and numerical data.
[0096] Next, the image processing device 10 corrects the orientation information and appearance information included in the text or numerical data search query using one of them (S41). Then, the image processing device 10 searches for the target image from among multiple reference images using the corrected search query (text or numerical data search query) (S42). Subsequently, the image processing device 10 outputs the search results.
[0097] The other configurations of the image processing apparatus 10 in this embodiment are the same as those of the first to third embodiments.
[0098] "Effects and Effects" The image processing apparatus 10 of this embodiment achieves the same effects and advantages as those of the first to third embodiments.
[0099] Furthermore, according to the image processing device 10 of this embodiment, it is possible to correct the other using either the posture information or the appearance information obtained by analyzing the image acquired as a search query, and then perform an image search using the corrected search query. When generating posture information or appearance information through image analysis, there is a possibility that the generated posture information or appearance information may contain errors. If an image search is performed using such erroneous information, the accuracy of the image search will be low.
[0100] Therefore, the image processing device 10 of this embodiment corrects the other using either the posture information or the appearance information identified by the analysis of the image input as a search query, and then searches for the target image from among multiple reference images using the corrected search query. In this way, the image processing device 10 of this embodiment, which appropriately corrects the search query based on the relationship between posture information and appearance information, improves the accuracy of the search query, and as a result, the accuracy of the image search is also improved.
[0101] <Variation> The following describes modifications applicable to the first to fourth embodiments. These modifications also achieve the same effects and advantages as the first to fourth embodiments.
[0102] -Experimental Variation 1- In this example, the image processing device 10 searches for a target image from among multiple reference images by comparing the reference image appearance information and reference image pose information obtained by analyzing the reference image with the appearance information and pose information included in the search query. The reference image appearance information indicates the appearance of the person included in the reference image. The reference image pose information indicates the pose of the person included in the reference image.
[0103] Then, the image processing device 10 corrects the reference image appearance information and reference image orientation information obtained by analyzing the reference image using the method described in the fourth embodiment.
[0104] In other words, the correction unit 13 corrects the other using either the reference image pose information, which indicates the posture of the person, and the reference image appearance information, which indicates the appearance of the person, which are generated based on the reference image. The correction method is as described in the fourth embodiment.
[0105] For example, if a predetermined key point is not detected, the correction unit 13 may delete information of a type previously associated with that key point from the reference image appearance information.
[0106] In addition, if a predetermined percentage or more of the multiple keypoints contained in a part of a person's body are not detected, the correction unit 13 may delete information of a type that is pre-associated with that part of the person's body from the reference image appearance information.
[0107] In addition, the correction unit 13 may delete information relating to a part of a person's body from the reference image posture information if the reliability of the information relating to a part of a person's body in the reference image appearance information satisfies predetermined conditions.
[0108] -Variation 2- In this example, the image processing device 10 searches for the target image from among multiple reference images by comparing the reference image appearance information and reference image orientation information obtained by analyzing the reference image with the appearance information and orientation information included in the search query.
[0109] Furthermore, the search unit 12 does not select as the target image any reference image in which multiple keypoints of the same type are detected from a single person region among multiple reference images.
[0110] A "person region" is a region of a predetermined shape (e.g., a rectangular region) where a person is detected. Techniques for detecting person regions using regions of a predetermined shape are widely known.
[0111] "Multiple identical keypoints were detected within a single person domain" means, for example, that multiple instances of "Head A1" were detected within a single person domain.
[0112] Since the accuracy of the reference image appearance information and reference image orientation information obtained by analyzing such reference images is low, the search unit 12 does not use such reference images as target images. The search unit 12 may exclude such reference images from the search query matching targets in advance. Alternatively, the search unit 12 may match the reference image with the search query, but exclude it from the target images regardless of the result.
[0113] -Variation 3- If the appearance information of the search query indicates that a predetermined type of garment is being worn, and the predetermined type of garment is not detected in the person included in the reference image, the search unit 12 determines whether or not the person is wearing the predetermined type of garment based on information indicating the posture of the person's body included in the reference image.
[0114] Specifically, the search unit 12 determines that if the posture of the person in the reference image is facing a first direction, the person in the reference image is not wearing a predetermined type of attachment. On the other hand, if the posture of the person in the reference image is facing a second direction, the search unit 12 determines that it is unclear whether or not the person in the reference image is wearing a predetermined type of attachment.
[0115] Here, we will explain the process in detail using the example of a case where the search query indicates that the person is "wearing glasses." In this example, if the person in the reference image is not found to be wearing the specified type of eyewear (glasses), there are two possible reasons: (1) the person is not wearing glasses in the first place and therefore not detected, and (2) the person is facing away and therefore not detected. In case (1), it is acceptable to conclude that the search query "wearing glasses" is not satisfied, but in case (2), it may not be desirable to conclude that the search query "wearing glasses" is not satisfied.
[0116] Therefore, the search unit 12 uses posture information to identify whether it is case (1) or case (2). Then, in case (1), the search unit 12 determines that the person included in the reference image is not wearing the predetermined type of eyewear (glasses), and in case (2), it determines that it is unclear whether the person included in the reference image is wearing the predetermined type of eyewear (glasses).
[0117] The search unit 12 identifies whether the person is facing a first direction or a second direction based on the posture information, and determines whether it is case (1) or case (2) based on the identification result. Specifically, if the person is facing the first direction, it is determined to be case (1), and if the person is facing the second direction, it is determined to be case (2).
[0118] The first and second directions differ depending on the type and position of the attached object.
[0119] For items worn on the face, such as glasses, sunglasses, and masks, the first direction is forward (face facing the camera), and the second direction is backward (face not facing the camera). Whether other directions, such as the side, are included in the first or second direction is a design consideration.
[0120] Furthermore, in the case of a wristwatch worn on the left hand, the first direction is forward, backward, and right (facing right towards the camera), and the second direction is left (facing left towards the camera).
[0121] How to handle reference images that are deemed unreliable in terms of whether or not they match based on at least one of the visual information items included in the search query (i.e., whether or not they are wearing a specific type of accessory) is a design consideration.
[0122] For example, reference images that are determined to match the search query, excluding items that are deemed unknown, may be used as target images. In other words, such reference images may also be displayed on the search results screen as reference images that match the search query.
[0123] In addition, reference images that are determined to match the search query, other than those deemed unknown, may be extracted as candidate target images. These reference images may then be displayed separately from the target image on the search results screen as candidate target images.
[0124] -Variation 4- The search unit 12 narrows down the reference images to match the search query based on the content of the search query.
[0125] Specifically, in this modified version, each of the multiple reference images is analyzed in advance, and as shown in Figure 4, information is stored in the memory unit that links the identification information (reference image identification information) of each of the multiple reference images with the posture information and appearance information of the person contained in each reference image. In addition, key points of the body are pre-linked with the type of appearance information related to the attachments worn on those parts of the body.
[0126] The search unit 12 then analyzes the search query and identifies which body part's appearance information is included. It then extracts reference images containing the identified type of appearance information from the reference images. Finally, the search unit 12 compares the extracted reference images with the search query.
[0127] The embodiments of the present invention have been described above with reference to the drawings, but these are illustrative examples of the present invention, and various other configurations can be adopted. The configurations of the embodiments described above may be combined with each other, or some configurations may be replaced with other configurations. Furthermore, the configurations of the embodiments described above may be modified in various ways without departing from the spirit of the invention. In addition, the configurations and processes disclosed in each of the embodiments and modifications described above may be combined with each other.
[0128] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content. Also, the above embodiments can be combined to the extent that their contents do not conflict.
[0129] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. An acquisition means for acquiring a search query that includes posture information showing a person's posture and appearance information showing a person's appearance, Correction means for correcting the other using one of the posture information and the appearance information, A search means that searches for a target image from among multiple reference images using the corrected search query, An image processing device having 2. The posture information indicates the detection results of multiple key points on the person's body. The image processing apparatus according to claim 1, wherein the correction means corrects the appearance information based on the detection result of the key point. 3. The correction means is, The image processing apparatus according to claim 2, which, if a predetermined key point is not detected, deletes the type of information associated with that key point from the appearance information. 4. The correction means is, The image processing apparatus according to claim 2, which, if a predetermined proportion or more of the multiple key points contained in a part of a person's body are not detected, deletes information of a type associated with that part of a person's body from the appearance information. 5. The correction means is, An image processing apparatus according to any one of 1 to 4, which corrects the posture information based on the reliability of the aforementioned appearance information. 6. The correction means is, The image processing apparatus according to 5, which deletes information relating to a part of a person's body from the posture information if the reliability of the information relating to a part of a person's body in the aforementioned appearance information satisfies predetermined conditions. 7. The correction means is, Using either the reference image pose information, which shows the posture of a person, or the reference image appearance information, which shows the appearance of a person, generated based on the aforementioned reference image, the other is corrected. An image processing apparatus according to any one of 1 to 6, wherein if the reliability of the information relating to a part of a person's body in the reference image appearance information satisfies predetermined conditions, the information relating to that part of a person's body is deleted from the reference image posture information. 8. The search means is, An image processing apparatus according to any one of 1 to 7, wherein the appearance information of the search query indicates that a predetermined type of attachment is being worn, and the predetermined type of attachment is not detected on the person included in the reference image, and determines whether or not the person is wearing the predetermined type of attachment based on information indicating the posture of the person's body included in the reference image. 9. One or more computers, A search query is obtained that includes posture information showing a person's stance and appearance information showing a person's appearance. Using one of the aforementioned posture information and the aforementioned appearance information, the other is corrected. An image processing method that searches for a target image from among multiple reference images using the corrected search query. 10. Computers, A means for obtaining a search query that includes posture information showing a person's posture and appearance information showing a person's appearance. Correction means for correcting the other using one of the posture information and the appearance information, A search means that searches for a target image from among multiple reference images using the corrected search query. A program that makes it function as such. [Explanation of Symbols]
[0130] 10 Image Processing Device 11 Acquisition Department 12 Search Section 13 Correction section 1A Processor 2A Memory 3A input / output I / F 4A Peripheral Circuits 5A Bus
Claims
1. A means for obtaining a search query that includes posture information showing a person's posture and appearance information showing a person's appearance, Correction means for performing correction by deleting specific information from the other using one of the posture information and the appearance information, A search means that searches for a target image from among multiple reference images using the corrected search query, An image processing device having
2. The aforementioned posture information shows the detection results of multiple key points on the person's body. The image processing apparatus according to claim 1, wherein the correction means performs the correction on the appearance information based on the detection result of the key point.
3. The correction means is The image processing apparatus according to claim 2, wherein if a predetermined key point is not detected, the type of information associated with that key point is deleted from the appearance information.
4. The correction means is The image processing apparatus according to claim 2, wherein if a predetermined proportion or more of the multiple key points contained in a part of a person's body are not detected, the type of information associated with that part of a person's body is deleted from the appearance information.
5. The correction means is The image processing apparatus according to claim 1, which performs the correction to the posture information based on the reliability of the aforementioned appearance information.
6. The correction means is The image processing apparatus according to claim 5, wherein if the reliability of the information relating to a part of a person's body in the aforementioned appearance information satisfies predetermined conditions, the information relating to that part of a person's body is deleted from the posture information.
7. The correction means is Using either the reference image pose information, which shows the posture of a person, or the reference image appearance information, which shows the appearance of a person, generated based on the aforementioned reference image, a correction is performed to remove specific information from the other. The image processing apparatus according to any one of claims 1 to 6, wherein if the reliability of the information relating to a part of a person's body in the reference image appearance information satisfies a predetermined condition, the information relating to that part of a person's body is deleted from the reference image posture information.
8. The search means is, The image processing apparatus according to claim 1, wherein the appearance information of the search query indicates that a predetermined type of attachment is being worn, and the predetermined type of attachment is not detected on the person included in the reference image, and the apparatus determines, based on information indicating the posture of the person's body included in the reference image, that the person is not wearing the predetermined type of attachment, or that it is unclear whether or not they are wearing it.
9. One or more computers, A search query is obtained that includes posture information showing a person's stance and appearance information showing a person's appearance. Using either the posture information or the appearance information, a correction is performed to delete specific information from the other. An image processing method that searches for a target image from among multiple reference images using the corrected search query.
10. Computers, A means for obtaining a search query that includes posture information showing a person's posture and appearance information showing a person's appearance. Correction means that performs correction by deleting specific information from the other using one of the posture information and the appearance information, A search means that searches for a target image from among multiple reference images using the corrected search query. A program that makes it function as such.
Citation Information
Patent Citations
Image processing method and device, clothes identification method and device, equipment and storage medium
CN113869435A
Imaging device, printer, image processor and program
JP2005086516A
Apparatus for detecting lone person and person in group
JP2006092396A
Image processing apparatus
JP2018045287A
Image retrieving apparatus, image retrieving method, and setting screen used therefor
JP2019091138A