Image analysis device, image analysis method, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-17
AI Technical Summary
Existing image analysis techniques are not versatile enough to correctly detect objects or determine similarity between images that have been subjected to image processing or distortion, limiting their effectiveness.
An image analysis device and method that acquires image characteristic information to correct images and object information, allowing for detection and identity determination using appropriate correction methods based on the characteristics of the images, including distortion and image processing types.
Enables accurate detection and identity determination of objects in images that have undergone various forms of processing or distortion, enhancing the versatility and accuracy of image analysis.
Abstract
Description
Image analysis device, image analysis method, and recording medium
[0001] The present disclosure relates to an image analysis device, an image analysis method, and a program.
[0002] A technology related to the present disclosure is disclosed in Patent Document 1. The technology disclosed in Patent Document 1 processes a registered image read from a registered image database to generate a match image in which wide-angle distortion is simulated. Then, the technology disclosed in Patent Document 1 performs matching based on the image to be processed and the match image in which wide-angle distortion is simulated.
[0003] International Publication No. 2020 / 189266
[0004] If the image to be analyzed has been processed or distorted, it will be impossible to correctly detect the target object or determine the similarity of the images.
[0005] The technology disclosed in Patent Document 1 is effective for analyzing images that contain wide-angle distortion. However, this technology is not effective for analyzing images that have been subjected to image processing. As such, the technology disclosed in Patent Document 1 lacks versatility.
[0006] In view of the above-mentioned problems, an example of an object of the present disclosure is to provide an image analysis device, an image analysis method, and a program that are capable of performing image analysis with excellent versatility.
[0007] According to the present disclosure, there is provided an image analysis device having: an acquisition means for acquiring image characteristic information indicating the characteristics of an image to be analyzed; a correction means for correcting at least one of the image to be analyzed, information on an object extracted from the image to be analyzed, a reference image, and information on the reference object in a manner according to the characteristics; and an analysis means for performing at least one of a process for detecting the reference object from the image to be analyzed and a process for determining the identity of the image to be analyzed and the reference image using the data after the correction.
[0008] Furthermore, according to the present disclosure, there is provided an image analysis method in which one or more computers acquire image characteristic information indicating the characteristics of an image to be analyzed, correct at least one of the image to be analyzed, information on an object extracted from the image to be analyzed, a reference image, and information on the reference object using a method according to the characteristics, and use the corrected data to perform at least one of a process of detecting the reference object from within the image to be analyzed and a process of determining the identity of the image to be analyzed and the reference image.
[0009] Furthermore, according to the present disclosure, there is provided a program that causes a computer to function as: an acquisition means that acquires image characteristic information indicating the characteristics of an image to be analyzed; a correction means that corrects at least one of the image to be analyzed, information on an object extracted from the image to be analyzed, a reference image, and information on the reference object in a manner according to the characteristics; and an analysis means that uses the corrected data to execute at least one of a process of detecting the reference object from within the image to be analyzed and a process of determining the identity of the image to be analyzed and the reference image.
[0010] According to one aspect of the present disclosure, an image analysis device, an image analysis method, and a program capable of performing highly versatile image analysis are realized.
[0011] FIG. 1 is a diagram showing an example of a functional block diagram of an image analysis device according to the present disclosure. FIG. 2 is a flowchart showing an example of a processing flow of an image analysis device according to the present disclosure. FIG. 3 is a diagram showing an example of a hardware configuration of an image analysis device according to the present disclosure. FIG. 4 is a diagram showing an example of image processing applied to an image to be analyzed. FIG. 5 is a diagram showing an example of processing performed by an image analysis device according to the present disclosure. FIG. 6 is a diagram showing another example of processing performed by an image analysis device according to the present disclosure. FIG. 7 is a diagram showing another example of processing performed by an image analysis device according to the present disclosure. FIG. 8 is a diagram showing another example of processing performed by an image analysis device according to the present disclosure. FIG. 9 is a diagram showing another example of processing performed by an image analysis device according to the present disclosure.
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In this disclosure, the drawings relate to one or more embodiments. In all drawings, similar components are designated by similar reference numerals, and descriptions thereof will be omitted as appropriate.
[0013] First Embodiment Fig. 1 is a functional block diagram showing an overview of an image analysis device 10. Fig. 2 is a flowchart showing an example of the flow of processing executed by the image analysis device 10.
[0014] 1, the image analysis device 10 includes an acquisition unit 11, a correction unit 12, and an analysis unit 13. These functional units execute the process shown in FIG.
[0015] First, the acquisition unit 11 acquires image characteristic information indicating the characteristics of the image to be analyzed (S10).
[0016] Next, the correction unit 12 corrects at least one of the "image to be analyzed," "information on the object (e.g., feature amounts) extracted from the image to be analyzed," "reference image," and "information on the reference object (e.g., feature amounts)" in a manner according to the characteristics of the image to be analyzed indicated by the image characteristic information acquired in S10 (S11).
[0017] Then, the analysis unit 13 uses the corrected data obtained in S11 to perform at least one of the following processes: "a process of detecting a reference object from the image to be analyzed" and "a process of determining the identity of the image to be analyzed and the reference image" (S12).
[0018] In this way, the image analysis device 10 can acquire image characteristic information indicating the characteristics of the image to be analyzed, and perform the correction in S11 using a method that corresponds to the characteristics of the image to be analyzed indicated by the acquired image characteristic information.
[0019] For example, if the image characteristic information indicates that distortion has occurred in the image to be analyzed, the image analysis device 10 can correct S11 using a method corresponding to that. Furthermore, the image analysis device 10 can correct S11 using a method corresponding to the type of distortion indicated in the image characteristic information.
[0020] Additionally, if the image characteristic information indicates that the image to be analyzed has been subjected to image processing, the image analyzing device 10 can perform the correction in S11 in a manner corresponding to that. Also, the image analyzing device 10 can perform the correction in S11 in a manner corresponding to the content of the image processing indicated in the image characteristic information.
[0021] In this way, the image analysis device 10 that acquires image characteristic information can identify the characteristics of the image to be analyzed based on the acquired image characteristic information. Then, the image analysis device 10 can perform the correction in S11 using a method that corresponds to the identified characteristics of the image to be analyzed. Furthermore, the image analysis device 10 can use the corrected data to execute a "process for detecting a reference object from within the image to be analyzed" and a "process for determining the identity of the image to be analyzed and the reference image." In this way, the image analysis device 10 can correct data using a highly versatile method and perform detection processes, identity determination processes, and the like using the corrected data.
[0022] Second Embodiment Overview An image analysis device 10 according to a second embodiment is a specific implementation of the configuration of the image analysis device 10 according to the first embodiment. This will be described in detail below.
[0023] "Hardware Configuration" First, an example of the hardware configuration of the image analysis device 10 will be described. Each functional unit of the image analysis device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are various variations in the realization method and device. Software includes programs that are pre-loaded in the device before shipping, and programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet.
[0024] FIG. 3 is a block diagram illustrating an example of the hardware configuration of an image analysis device 10. As shown in FIG. 3, the image analysis device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The image analysis device 10 does not necessarily have to have the peripheral circuit 4A. Note that the image analysis device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration.
[0025] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to mutually transmit and receive data. The processor 1A is, for example, a processing unit such as a CPU or a graphics processing unit (GPU). The memory 2A is, for example, a random access memory (RAM) or a read-only memory (ROM). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes an interface for connecting to a communication network such as the Internet. Examples of input devices include a keyboard, mouse, microphone, physical buttons, touch panel, etc. Examples of output devices include a display, speaker, printer, mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0026] "Functional Configuration" Next, the functional configuration of the image analysis device 10 will be described in detail. FIG. 1 shows an example of a functional block diagram of the image analysis device 10. As shown in the figure, the image analysis device 10 has an acquisition unit 11, a correction unit 12, and an analysis unit 13. As described above, the image analysis device 10 may be configured as multiple physically and / or logically separated devices. In this case, the acquisition unit 11, the correction unit 12, and the analysis unit 13 can be configured as multiple devices separated in any manner. For example, the image analysis device 10 may be configured as three physically and / or logically separated devices, each of which includes one of the acquisition unit 11, the correction unit 12, and the analysis unit 13.
[0027] The acquisition unit 11 acquires image characteristic information indicating the characteristics of the image to be analyzed.
[0028] The "image to be analyzed" is an image that is the subject of image analysis. Image analysis is performed by the analysis unit 13. As will be described later, the analysis unit 13 performs at least one of a process of detecting a reference object from the image to be analyzed and a process of determining the identity (similarity) between the image to be analyzed and the reference image. The concept of an image includes still images and moving images. The image to be analyzed may be a still image or a moving image.
[0029] The image to be analyzed may be an image that has been subjected to image processing or an image that has been distorted.
[0030] For example, when image content (videos, still images) is uploaded to a server on the Internet and made public without the permission of the rights holder, the uploaded image content may be edited to avoid detection of the act. For example, such image content may be the image to be analyzed. In other words, an image that may have been uploaded to a server on the Internet and made public without the permission of the rights holder may be the image to be analyzed.
[0031] In addition, in recent years, cameras with wide angles of view have been used as surveillance cameras, in-vehicle cameras, etc. Images generated by such cameras may be distorted. For example, an image generated by such a camera with a wide angle of view may be the image to be analyzed.
[0032] The examples of the analysis target images shown here are merely examples, and the present invention is not limited to these.
[0033] The "characteristics of the image to be analyzed indicated by the image characteristic information" are characteristics that may affect the correction by the correction unit 12. The characteristics include at least one of the following characteristics: - Presence or absence of image processing - Contents of image processing - Presence or absence of distortion - Conditions of the camera that generated the image
[0034] "Presence or absence of image processing" indicates whether image processing has been applied to the image to be analyzed.
[0035] The "contents of image processing" refers to the content of image processing applied to the image to be analyzed. Image processing here includes all widely known techniques.
[0036] The image processing may include at least one of the following: distorting the image to be analyzed, modifying the content display area, changing it to a pictorial image, blurring, and making the object in the image to be analyzed unrecognizable. Note that the image processing described here is merely an example and is not limited to these.
[0037] Processing to Distort the Image to be Analyzed and Correcting the Content Display Area Here, an example of the above-mentioned "processing to distort the image to be analyzed" and "correction of the content display area" will be described using FIG. 4 . Such image processing is performed, for example, when uploading and publishing image content (videos and still images) to an Internet server without the permission of the rights holder. In the example of FIG. 4 , in the image P before processing, the content is displayed across the entire image P. That is, the entire image P is the content display area. Note that a person M is depicted in the content. In the image P after processing, the content is displayed in a partial region P' of the image P. That is, the partial region P' of the image P is the content display area. The shape of the partial region P' displaying the content is distorted. This is achieved, for example, by rotating the content image three-dimensionally using a three-dimensional transformation tool. In the example of FIG. 4 , the content displayed in the partial region P' is tilted so that the left edge L is located at the back and the right edge R is located at the front. Note that the tilting method is not limited to this.
[0038] When such image processing has been performed, the image characteristic information may further indicate the position and shape of a partial region P' within the image to be analyzed.
[0039] o Changing to a Painterly Image In recent years, technologies for changing images generated by a camera into a painterly image have become widespread. Changing to a painterly image is a process that uses these technologies to change the image being analyzed into a painterly image. Painterly images can include a variety of images, such as oil painting-style images and sketch-style images. When such image processing has been performed, the image characteristic information may further indicate which of multiple types of painterly images the image has been processed into.
[0040] Blur processing Blur processing is a process in which the data of each pixel is modified, for example, using the data of surrounding pixels. The degree of blurring can vary. For example, in some cases, the image is blurred to the extent that the object in the image can be identified, while in other cases, the image is blurred to the extent that the object in the image cannot be identified.
[0041] Processing to make objects in the image to be analyzed unrecognizable Processing to make objects in the image to be analyzed unrecognizable is processing to make the objects in the image to be analyzed unrecognizable to humans and / or computers. This type of image processing is performed, for example, when image content (videos, still images) is uploaded to a server on the Internet without the permission of the rights holder and made public, with the aim of enjoying it mainly through audio.
[0042] The target object may be a person or other object. Examples of the processing include, but are not limited to, blurring at a predetermined level or higher, mosaic processing, superimposing another image on at least a portion of the image to be analyzed, and painting at least a portion of the image to be analyzed in a predetermined color.
[0043] Note that the image processing exemplified here may be performed on the entire image to be analyzed, or on a partial region. Furthermore, different types of image processing may be performed on multiple regions of the image to be analyzed. The image characteristic information may further indicate these details. That is, the image characteristic information may indicate the content of the image processing performed on each partial region within the image to be analyzed.
[0044] "Presence or absence of distortion" indicates whether distortion has occurred in the image to be analyzed.
[0045] The "conditions of the camera that generated the image" include at least one of the characteristics of the camera that generated the image to be analyzed and the shooting conditions. The conditions of the camera that generated the image may indicate, for example, the type of lens of the camera that generated the image to be analyzed. The type of lens depends on the width of the angle of view, and examples include a fisheye lens, an ultra-wide-angle lens, a wide-angle lens, and a standard lens. It is known that images generated using lenses with wide angles of view, such as a fisheye lens, an ultra-wide-angle lens, and a wide-angle lens, tend to be distorted. In addition, the conditions of the camera that generated the image may indicate lens characteristic values such as focal length, F-number, and horizontal angle.
[0046] Additionally, the conditions of the camera that generated the image may indicate the height of the camera from the ground when the image was taken. Also, the conditions of the camera that generated the image may indicate various camera settings when the image was taken, such as the F-number, shutter speed, ISO sensitivity, etc. Also, the conditions of the camera that generated the image may indicate the time, weather, season, shooting environment (outdoors, indoors, etc.) when the image was taken, etc.
[0047] If the image to be analyzed is a moving image, the acquisition unit 11 may acquire one piece of image characteristic information for a predetermined length of moving image (a collection of multiple frame images). Alternatively, the acquisition unit 11 may acquire image characteristic information for each frame image.
[0048] Next, an example of a process in which the acquisition unit 11 acquires the image characteristic information described above will be described. In one example, an operator visually checks the image to be analyzed, collects various information related to the image to be analyzed and the camera, and identifies the characteristics. The operator then inputs the identified characteristics into the image analysis device 10 as image characteristic information. The image analysis device 10 can accept this input via any input device, such as a keyboard, a mouse, a touch panel, physical buttons, or a microphone.
[0049] Alternatively, the acquisition unit 11 may acquire image characteristic information using an estimation model previously generated by machine learning. The estimation model is generated by machine learning based on learning data linking multiple teacher images with image characteristic information (correct labels) of each teacher image. The multiple teacher images may include images that have been subjected to image processing and images that have not been subjected to image processing. The multiple teacher images may also include multiple images that have been subjected to different types of image processing. The multiple teacher images may also include images that have been distorted and images that have not been distorted. The multiple teacher images may also include images generated under different camera conditions.
[0050] Alternatively, the acquisition unit 11 may acquire image characteristic information by using both an input by an operator and an estimation model. For example, the acquisition unit 11 may acquire some types of information included in the image characteristic information (e.g., information obtained by image analysis) by using an estimation model, and acquire other types of information (e.g., information that is difficult to obtain by image analysis) by input by an operator.
[0051] Returning to Figure 1, the correction unit 12 corrects at least one of the image to be analyzed, the information on the object extracted from the image to be analyzed, the reference image, and the information on the reference object in a manner according to the characteristics indicated in the image characteristic information acquired by the acquisition unit 11.
[0052] "Objects" are people and / or other objects.
[0053] "Object information" is data indicating the external characteristics of an object. The object information may be an external image of the object, or may be feature quantities extracted from such an external image. If the object is a person, the object information may include a facial image, a full-body image, facial features, physique features, gender, age group, etc. Furthermore, if the object is another object, the object information may include the type of object, the shape of the object, etc.
[0054] A "reference image" is an image that is referenced to analyze the image to be analyzed. The reference image may be an image that is registered in a database in advance. Alternatively, the reference image may be an image that the user inputs and / or specifies in the image analysis device 10 each time an analysis is performed, or an image that the image analysis device 10 acquires from an external device. The image acquired each time an analysis is performed may be temporarily stored in a volatile memory or may be stored in a non-volatile memory. The reference image may be a still image or a moving image. In one example of image analysis performed by the analysis unit 13 described below, the identity (similarity) between the image to be analyzed and the reference image is determined. The image analysis device 10 may store the reference image, or an external device configured to be able to communicate with the image analysis device 10 may store the reference image.
[0055] The "reference object" is a person and / or other object to be detected from the image to be analyzed. In one example of image analysis performed by the analysis unit 13 described below, a process of detecting the reference object from the image to be analyzed is performed.
[0056] "Reference object information" refers to information referenced for analyzing the analysis target image. The reference object information may be registered in a database in advance. Alternatively, the reference object information may be information input and / or specified by the user in the image analysis device 10 each time an analysis is performed, or information acquired by the image analysis device 10 from an external device. The reference object information may also be information obtained by analyzing the reference image described above each time an analysis is performed. The information acquired each time an analysis is performed may be temporarily stored in a volatile memory or may be stored in a non-volatile memory. The reference object information is data indicating the external characteristics of the reference object. The reference object information may be an external image of the reference object, or may be feature quantities extracted from such an external image. Note that instead of extracting the information from an external image, an operator may identify the reference object information by any means and register it in a database or input it into the image analysis device 10. When the reference object is a person, the reference object information may include a face image, a whole-body image, facial features, physique features, data showing a three-dimensional model of facial and body features, gender, age group, etc. Furthermore, when the reference object is another object, the information on the reference object may include data showing the type of object, the shape of the object, the shape characteristics as a three-dimensional model, etc. Furthermore, the information on the reference object may include voiceprint information of the reference object.
[0057] Next, the correction performed by the correction unit 12 will be described in detail.
[0058] The correction unit 12 can execute at least one of the following processes: - A process of correcting the image to be analyzed to an image before predetermined image processing is applied - A process of correcting the image to be analyzed to an image without distortion - A process of correcting information about an object extracted from the image to be analyzed to information about the object extracted in the image before predetermined image processing is applied - A process of correcting information about an object extracted from the image to be analyzed to content extracted from an image without distortion - A process of correcting the reference image to an image to which image processing similar to that applied to the image to be analyzed - A process of correcting the reference image to an image as if taken with the lens of the camera that generated the image to be analyzed - A process of correcting information about the reference object to information extracted from an image to which image processing applied to the image to be analyzed - A process of correcting information about the reference object to information as if taken with the lens of the camera that generated the image to be analyzed - A process of generating data including audio from the image to be analyzed but not including an image as corrected data - A process of generating data including audio from the reference image but not including an image as corrected data
[0059] This will be explained in detail below.
[0060] "Process of correcting the analysis target image to the image before predetermined image processing" When the image characteristic information indicates that the analysis target image has been subjected to image processing, the correction unit 12 can perform the correction. The "predetermined image processing" is the image processing that has been performed on the analysis target image indicated by the image characteristic information.
[0061] First, correction will be described for a case where the "processing for distorting the analysis target image" and the "correction of the content display area" as shown in FIG. 4 have been performed on the analysis target image.
[0062] In this case, the correction unit 12 performs correction to obtain the image P displayed on the left side of FIG. 4 from the image P displayed on the right side of FIG. 4. This correction can be achieved using any technique. For example, the correction unit 12 identifies a partial region P' on the right side of FIG. 4 based on image characteristic information. Then, the correction unit 12 obtains the image P displayed on the left side of FIG. 4 by three-dimensionally rotating the image of the content displayed in the partial region P' using a three-dimensional conversion tool. The correction unit 12 may estimate the degree of tilt in the depth direction of the content displayed in the partial region P' based on the ratio between the length of the side of the left edge L of the content displayed in the partial region P' and the length of the side of the right edge R of the content. Then, the correction unit 12 may three-dimensionally rotate the image of the content displayed in the partial region P' to eliminate the estimated tilt in the depth direction. The correction unit 12 may also identify multiple candidates for the degree of tilt in the depth direction of the content displayed in the partial region P' based on the ratio. Then, the correction unit 12 may perform correction based on each of the multiple candidates and generate multiple corrected images.
[0063] Although the correction method has been described here as an example in which both the left and right edges of a content image are misaligned in the depth direction, the same correction method can be used when both the top and bottom edges of a content image are misaligned in the depth direction.Furthermore, the same correction method can be used when both the left and right edges and the top and bottom edges of a content image are misaligned in the depth direction.
[0064] Next, correction when the image to be analyzed has been changed to a pictorial image will be described.
[0065] In this case, the correction unit 12 performs correction to change the pictorial image into an image (a live-action image) that realistically depicts people and other objects. In recent years, technologies for changing illustrated images into live-action images have become widespread. The correction unit 12 can use, for example, these technologies to perform correction to change the pictorial image into an image (a live-action image) that realistically depicts people and other objects. With this correction, even if it is not possible to reproduce the actual image of a person, it is possible to obtain an image that is closer to the actual image of the person than the pictorial image.
[0066] "Process of correcting the image to be analyzed into an image without distortion" When the image characteristic information indicates that the image to be analyzed has distortion, the correction unit 12 can perform the correction.
[0067] In this case, the correction unit 12 performs correction to change a distorted image generated by a fisheye lens, an ultra-wide-angle lens, a wide-angle lens, or the like into an image with reduced distortion. The correction unit 12 can perform this correction using well-known techniques.
[0068] "Process of correcting information about an object extracted from an image to be analyzed to information extracted from an image before a predetermined image processing is applied" and "Process of correcting information about an object extracted from an image to be analyzed to information extracted from an image without distortion" The correction unit 12 can perform the correction when the image characteristic information indicates that image processing has been applied to the image to be analyzed. Furthermore, the correction unit 12 can perform the correction when the image characteristic information indicates that distortion has occurred in the image to be analyzed. "Predetermined image processing" is image processing applied to the image to be analyzed that is indicated by the image characteristic information.
[0069] The correction unit 12 may correct the information about the object using an estimation model previously generated by machine learning. The estimation model is generated by machine learning based on training data that pairs information about the object extracted from an image before image processing with information about the object extracted from an image after image processing. Machine learning learns the relationship between the information about the two paired objects. The estimation model receives information about the object extracted from the image after image processing as input and outputs information about the object extracted from the image before image processing. Multiple estimation models may be prepared for each type of image processing. The correction unit 12 may perform the correction using an estimation model corresponding to the type of image processing applied to the image to be analyzed. For example, in the image processing shown in FIG. 4, the shape of the partial region P' may be classified into multiple groups, and an estimation model may be prepared for each classification. The information about the extracted object may also change depending on the area in which the content is displayed.
[0070] The estimation model may also be generated by machine learning based on training data that pairs information about an object extracted from a distorted image with information about an object extracted from an undistorted image. Machine learning learns the relationship between the information about the two paired objects. The estimation model receives information about the object extracted from the distorted image as input and information about the object extracted from the undistorted image as output. Multiple estimation models may be prepared depending on the type of lens, various camera settings at the time of capture, the height of the camera from the ground at the time of capture, and / or the position of the object in the image to be analyzed. The items exemplified here may affect the distortion of the object. The correction unit 12 may perform the correction using an estimation model that corresponds to the details of these items related to the image to be analyzed.
[0071] "Process of correcting the reference image into an image that has undergone image processing similar to that applied to the image to be analyzed" The correction unit 12 can perform this correction when the image characteristic information indicates that the image to be analyzed has undergone image processing.
[0072] In this correction, the correction unit 12 performs image processing on the reference image, the image processing performed on the analysis target image indicated by the image characteristic information. The correction unit 12 can perform this image processing using any widely known technology.
[0073] "Process of correcting the reference image to an image that would appear if it were taken with the lens of the camera that generated the image to be analyzed" The correction unit 12 can perform this correction if the image characteristic information indicates that distortion has occurred in the image to be analyzed.
[0074] In this correction, the correction unit 12 corrects the reference image to an image captured under the conditions (lens type, etc.) of the camera that generated the analysis target image indicated by the image characteristic information. That is, the correction unit 12 adds distortion to the reference image. The correction unit 12 can use any technique to perform this correction. For example, the correction unit 12 may perform a correction process to remove distortion from the analysis target image that has distortion. Then, the correction unit 12 may identify the type of distortion that has occurred in the analysis target image based on the distorted analysis target image and the analysis target image after the distortion has been removed, and apply the identified distortion to the reference image. In addition, the correction unit 12 can use any technique, such as the technique disclosed in Patent Document 1 or the technique disclosed in Japanese Patent Laid-Open No. 2011-148369.
[0075] "Process of correcting information about a reference object to information extracted from an image that has been subjected to image processing similar to that applied to the image to be analyzed" and "Process of correcting information about a reference object to information as if the image were photographed with the lens of the camera that generated the image to be analyzed." The correction unit 12 can perform this correction when the image characteristic information indicates that image processing has been applied to the image to be analyzed. Furthermore, the correction unit 12 can perform this correction when the image characteristic information indicates that distortion has occurred in the image to be analyzed.
[0076] As described above, the information on the reference object is data indicating the external characteristics of the reference object. The correction unit 12 may correct the information on the reference object using an estimation model previously generated by machine learning. The estimation model is generated by machine learning based on learning data paired with information on the reference object extracted from an image before image processing and information on the reference object extracted from an image after image processing. The machine learning learns the relationship between the information on the two paired reference objects. The estimation model receives the information on the reference object extracted from the image before image processing and outputs the information on the reference object extracted from the image after image processing. Multiple estimation models may be prepared for each type of image processing. The correction unit 12 may perform the correction using an estimation model corresponding to the type of image processing applied to the image to be analyzed. For example, in the image processing shown in FIG. 4, the shape of the partial region P' may be classified into multiple groups, and an estimation model may be prepared for each classification. The information on the extracted object may also change depending on the area in which the content is displayed.
[0077] The estimation model may also be generated by machine learning based on training data paired with information on a reference object extracted from a distorted image and information on a reference object extracted from an undistorted image. Machine learning learns the relationship between the information on the two paired reference objects. The estimation model receives the information on the reference object extracted from the undistorted image as input and the information on the reference object extracted from the distorted image as output. Multiple estimation models may be prepared depending on the type of lens, various camera settings at the time of capture, the height of the camera from the ground at the time of capture, and / or the position of the object in the image to be analyzed. The items exemplified here may affect the distortion of the object. The correction unit 12 may perform the correction using an estimation model corresponding to the details of these items related to the image to be analyzed.
[0078] "Process for generating data containing the audio of the image to be analyzed but not the image as corrected data" and "Process for generating data containing the audio of the reference image but not the image as corrected data" The correction unit 12 can perform this correction when the image characteristic information indicates that image processing has been applied to make the objects appearing in the image to be analyzed unidentifiable.
[0079] In this correction, the correction unit 12 can extract audio data from the analysis target image, which is a moving image, and use the extracted audio data as corrected data (audio file).In addition, in this correction, the correction unit 12 can extract audio data from the reference image, which is a moving image, and use the extracted audio data as corrected data (audio file).
[0080] When an image to be analyzed has been processed to make the object in the image unrecognizable, it is difficult to detect the reference object based on the image or to determine its identity with the reference image. In this case, the image analysis device 10 can perform such detection and determination using audio data.
[0081] Returning to Figure 1, the analysis unit 13 uses the data corrected by the correction unit 12 to perform at least one of the following processes: detecting a reference object from the image to be analyzed; and determining the identity of the image to be analyzed and the reference image.
[0082] For example, the analysis unit 13 can compare the "analysis target image before predetermined image processing" with the "reference image without image processing," as shown in Fig. 5. The "analysis target image before predetermined image processing" is obtained by correcting the "analysis target image that has been processed."
[0083] 6, the analysis unit 13 can compare the "analysis target image that has undergone image processing" with the "reference image that has undergone image processing similar to that performed on the analysis target image." The "reference image that has undergone image processing similar to that performed on the analysis target image" is obtained by correcting the "reference image that has not undergone image processing."
[0084] 7, the analysis unit 13 can compare the "feature amounts of the object extracted in the analysis target image before the predetermined image processing is performed" with the "feature amounts of the reference object." The "feature amounts of the object extracted in the analysis target image before the predetermined image processing is performed" can be obtained by correcting the "feature amounts of the object extracted from the analysis target image that has been image processed."
[0085] 8, the analysis unit 13 can compare the "feature amount of the object extracted from the processed image to be analyzed" with the "feature amount of the reference object extracted from the image to be analyzed that has been subjected to image processing." The "feature amount of the reference object extracted from the image to be analyzed that has been subjected to image processing" can be obtained by correcting the "feature amount of the reference object."
[0086] Furthermore, as shown in Figure 9, the analysis unit 13 can compare the "distorted analysis target image" obtained by correcting the "distorted analysis target image" with the "undistorted reference image."
[0087] Furthermore, as shown in Figure 10, the analysis unit 13 can compare the "image to be analyzed that has distortion" with the "reference image that has distortion" obtained by correcting the "reference image that does not have distortion."
[0088] Furthermore, as shown in FIG. 11, the analysis unit 13 can compare the "feature values of the object extracted from the image to be analyzed that is not distorted" obtained by correcting the "feature values of the object extracted from the image to be analyzed that is distorted" with the "feature values of the reference object."
[0089] Furthermore, as shown in FIG. 12, the analysis unit 13 can compare the "feature values of the object extracted from the image to be analyzed in which distortion has occurred" with the "feature values of the reference object extracted from the image in which distortion has occurred" obtained by correcting the "feature values of the reference object."
[0090] Furthermore, as shown in Figure 13, the analysis unit 13 can compare the "audio" obtained by correcting the "analysis target image, which has been subjected to image processing to make the object unidentifiable," with the "audio" obtained by correcting the "reference image."
[0091] Furthermore, as shown in FIG. 14, the analysis unit 13 can compare the "voice" obtained by correcting the "image to be analyzed, which has been subjected to image processing to make the object unidentifiable," with the "voiceprint of the reference object."
[0092] The analysis unit 13 can perform the above matching using any widely known technology. The above matching realizes a process of detecting a reference object from the analysis target image and a process of determining the identity (similarity) between the analysis target image and the reference image. In the example of FIG. 13 , the analysis unit 13 may determine the identity (similarity) between the analysis target image and the reference image by, for example, performing transcription using voice analysis and then determining the identity (similarity) of the sentences. Furthermore, in the example of FIG. 14 , the analysis unit 13 may detect the reference object from the analysis target image by matching a voiceprint detected from a voice obtained by correcting the analysis target image with the voiceprint of the reference object.
[0093] The analysis unit 13 can output the matching result via any output device such as a display, a projection device, or a speaker.
[0094] Next, an example of the processing flow of the image analysis device 10 will be described with reference to the flowchart of FIG.
[0095] First, the image analysis device 10 acquires image characteristic information indicating the characteristics of the image to be analyzed (S10). Next, the image analysis device 10 corrects at least one of the "image to be analyzed," "feature amounts of an object extracted from the image to be analyzed," "reference image," and "feature amounts of a reference object" using a method according to the characteristics of the image to be analyzed indicated by the image characteristic information acquired in S10 (S11). Then, using the corrected data obtained in S11, the image analysis device 10 executes at least one of a "process for detecting a reference object in the image to be analyzed" and a "process for determining the identity of the image to be analyzed and the reference image" (S12).
[0096] "Effects" The image analysis device 10 of the second embodiment has an acquisition unit 11 that acquires image characteristic information that indicates the characteristics of the image to be analyzed. The image analysis device 10 can identify the characteristics of the image to be analyzed based on this image characteristic information. Even if the image to be analyzed is an image that has been processed or an image that is distorted, the image analysis device 10 can identify the characteristics of the image to be analyzed based on the image characteristic information.
[0097] The image analysis device 10 can then correct at least one of the analysis target image, the information on the object extracted from the analysis target image, the reference image, and the information on the reference object using a method that corresponds to the characteristics of the identified analysis target image. In this way, the image analysis device 10 can perform correction for each analysis target image using a method that corresponds to the characteristics of each analysis target image.
[0098] The image analysis device 10 can then use the corrected data to perform at least one of the following: detecting a reference object from within the image to be analyzed; and determining whether the image to be analyzed and the reference image are identical. Because the image analysis device 10 uses the corrected data, it can perform the above detection and determination processes with high accuracy even if the image to be analyzed has been processed or distorted.
[0099] The processing target of the image analysis device 10, which can perform such versatile image analysis, is not limited to predetermined images such as "only distorted images" or "only images that have been subjected to predetermined image processing." The image analysis device 10 can analyze any image.
[0100] Furthermore, when the image to be analyzed has been subjected to processing that makes the object appearing in the image to be analyzed unidentifiable, the image analysis device 10 can perform the above detection processing and determination processing using audio. When processing that makes the object unidentifiable has been performed, performing the above detection processing and determination processing based on image data (such as the values of each pixel) may result in erroneous results. In such cases, performing processing using audio can improve the accuracy of the processing compared to when image data is used.
[0101] <Modifications> Here, a modification of the configuration of the image analysis device 10 of the first and second embodiments will be described. When the image to be analyzed is a moving image, the acquisition unit 11 may acquire one piece of image characteristic information for a predetermined length of the moving image (a collection of multiple frame images). Then, the correction unit 12 may perform correction on all of the multiple frame images included in the predetermined length of the moving image in accordance with the characteristics indicated by that one piece of image characteristic information.
[0102] Alternatively, if the image to be analyzed is a moving image, the acquisition unit 11 may acquire one piece of image characteristic information for each frame image, and the correction unit 12 may perform correction on each frame image according to the characteristics indicated in the image characteristic information acquired corresponding to each frame image.
[0103] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0104] In addition, in the flowcharts used in the above explanation, multiple steps (processes) are described in order. However, the order of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed as long as it does not cause any problems in terms of the content.
[0105] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes. Note that some or all of Supplements 2 to 8, which are dependent on the image analysis device of Supplement 1, may also be dependent on the image analysis method of Supplement 9, the program of Supplement 10, and the recording medium of Supplement 11 in the same dependent relationship as Supplementary notes 2 to 8. Furthermore, within the scope of each of the above-mentioned embodiments, some or all of the configurations described as supplementary notes can be realized in various hardware, software, various recording means for recording software, or systems. 1. An image analysis device having: an acquisition means for acquiring image characteristic information indicating the characteristics of an image to be analyzed; a correction means for correcting at least one of the image to be analyzed, information on an object extracted from the image to be analyzed, a reference image, and information on the reference object, using a method according to the characteristics; and an analysis means for performing at least one of a process for detecting the reference object from the image to be analyzed and a process for determining the identity of the image to be analyzed and the reference image, using the data after the correction. 2. 2. The image analysis device according to claim 1, wherein the image characteristic information indicates the content of image processing applied to the image to be analyzed. 3. The image analysis device according to claim 1 or 2, wherein the image characteristic information indicates the type of lens of a camera that generated the image to be analyzed. 4. The image analysis device according to claim 2, wherein the correction means executes at least one of the following: a process of correcting the image to be analyzed to an image before the image processing; a process of correcting information of an object extracted from the image to be analyzed to information extracted from the image before the image processing; a process of correcting the reference image to an image after the image processing; and a process of correcting information of the reference object to information extracted from the image after the image processing. 5. The image analysis device according to claim 4, wherein the content of the image processing includes at least one of a process of distorting the image to be analyzed, a correction of a content display area, a change to a painterly image, and a blurring process.6. The image analysis device according to 2, wherein, when the image processing is a process to make an object appearing in the image to be analyzed unidentifiable, the correction means executes at least one of the following processes: a process to generate data including audio of the image to be analyzed but no image as data after the correction; and a process to generate data including audio of the reference image but no image as data after the correction. 7. The image analysis device according to 6, wherein, when the image processing is a process to make an object appearing in the image to be analyzed unidentifiable, the analysis means executes at least one of a process to detect the reference object from within the image to be analyzed based on audio, and a process to determine the identity of the image to be analyzed and the reference image. 8. 9. The image analysis device according to 3, wherein the correction means executes at least one of the following processes: correcting the image to be analyzed into an image without distortion; correcting information about the object extracted from the image to be analyzed into content that would be extracted from an image without distortion; correcting the reference image into an image as if it were taken with the lens of a camera that generated the image to be analyzed; and correcting information about the reference object into information as if it were taken with the lens of a camera that generated the image to be analyzed. 9. An image analysis method in which one or more computers acquire image characteristic information that indicates characteristics of the image to be analyzed, correct at least one of the image to be analyzed, information about the object extracted from the image to be analyzed, the reference image, and information about the reference object using a method according to the characteristics, and execute at least one of the processes of detecting the reference object from the image to be analyzed using the corrected data and determining the identity of the image to be analyzed and the reference image. A program that causes a computer to function as: an acquisition means that acquires image characteristic information that indicates the characteristics of an image to be analyzed; a correction means that corrects at least one of the image to be analyzed, information on an object extracted from the image to be analyzed, a reference image, and information on the reference object using a method according to the characteristics; and an analysis means that uses the corrected data to execute at least one of a process of detecting the reference object from within the image to be analyzed and a process of determining the identity of the image to be analyzed and the reference image.11. A recording medium that records a program that causes a computer to function as: an acquisition means that acquires image characteristic information that indicates the characteristics of an image to be analyzed; a correction means that corrects at least one of the image to be analyzed, information on an object extracted from the image to be analyzed, a reference image, and information on the reference object using a method according to the characteristics; and an analysis means that uses the corrected data to execute at least one of a process of detecting the reference object from within the image to be analyzed and a process of determining the identity of the image to be analyzed and the reference image.
[0106] This application claims priority based on Japanese Patent Application No. 2023-118280, filed on July 20, 2023, the disclosure of which is incorporated herein by reference in its entirety.
[0107] 10 Image analysis device 11 Acquisition unit 12 Correction unit 13 Analysis unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus
Claims
1. An acquisition means for acquiring image characteristic information that shows the characteristics of the image to be analyzed, Correction means for correcting at least one of the following in a manner corresponding to the characteristics: the image to be analyzed, the information of the object extracted from the image to be analyzed, the reference image, and the information of the reference object. An analysis means that performs at least one of the following processes using the corrected data: a process to detect the reference object from the image to be analyzed, and a process to determine the identity between the image to be analyzed and the reference image. An image analysis device having the following features.
2. The image analysis apparatus according to claim 1, wherein the image characteristic information indicates the content of image processing applied to the image to be analyzed.
3. The image analysis apparatus according to claim 1 or 2, wherein the image characteristic information indicates the type of lens of the camera that generated the image to be analyzed.
4. The correction means is A process of correcting the image to be analyzed to the image before the image processing was applied. A process of correcting the information of the object extracted from the image to be analyzed to the information extracted from the image before the image processing was applied. The process of modifying the aforementioned reference image into the image after the aforementioned image processing has been performed, and A process of modifying the information of the reference object to information extracted from the image that has undergone image processing. The image analysis device according to claim 2, which performs at least one of the following.
5. The image analysis apparatus according to claim 4, wherein the image processing includes at least one of the following: a process to distort the image to be analyzed, a process to modify the content display area, a process to change to a painterly image, and a blurring process.
6. If the image processing described above is a process that makes the objects in the image to be analyzed unidentifiable, The correction means is A process for generating data that includes the audio of the image to be analyzed but does not include the image, as the data after the correction, and A process of generating data that includes the audio of the aforementioned reference image but does not include the image, as the data after the correction. The image analysis device according to claim 2, which performs at least one of the following.
7. If the image processing described above is a process that makes the objects in the image to be analyzed unidentifiable, The image analysis apparatus according to claim 6, wherein the analysis means performs at least one of the following processes based on sound: detecting the reference object from the image to be analyzed, and determining the identity between the image to be analyzed and the reference image.
8. The correction means is A process to correct the aforementioned image to an image without distortion. A process to correct the information of the object extracted from the aforementioned image to match the content extracted from an image without distortion. The process of correcting the aforementioned reference image to an image taken with the lens of the camera that generated the aforementioned image to be analyzed, and A process to correct the information of the reference object to information obtained when it is captured by the lens of the camera that generated the image to be analyzed. The image analysis device according to claim 3, which performs at least one of the following:
9. One or more computers, We obtain image characteristic information that shows the characteristics of the image to be analyzed, Correcting at least one of the following in accordance with the characteristics: the image to be analyzed, the information of the object extracted from the image to be analyzed, the reference image, and the information of the reference object. An image analysis method that performs at least one of the following processes using the corrected data: detecting the reference object from the image to be analyzed, and determining the identity between the image to be analyzed and the reference image.
10. Computers, means for acquiring image characteristic information that shows the characteristics of the image to be analyzed, Correction means for correcting at least one of the following in a manner corresponding to the characteristics: the image to be analyzed, the information of an object extracted from the image to be analyzed, a reference image, and the information of a reference object. Analysis means that performs at least one of the following processes using the corrected data: detecting the reference object from the image to be analyzed, and determining the identity between the image to be analyzed and the reference image. A program that makes it function as such.