Image processing program, image processing device, and image processing method

The image processing program and device address the privacy concerns in constant authentication systems by using a machine-learned model to add noise to specific areas of images, ensuring accurate person re-identification while protecting individual privacy.

WO2025094226A1PCT designated stage expired Publication Date: 2025-05-08FUJITSU LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/039045
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Conventional human feature extraction in constant authentication systems does not adequately protect the privacy of individuals, as images used for identification are often viewed by third parties, compromising the photographer's privacy.

Method used

An image processing program and device that utilize a machine-learned image processing model to process the person area of an image, adding noise to specific areas to minimize feature extraction changes, thereby protecting privacy while maintaining accurate person re-identification.

Benefits of technology

The solution effectively protects the privacy of individuals by adding noise to specific areas of the image, ensuring that the accuracy of person re-identification is not compromised, thus addressing the privacy concerns in constant authentication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023039045_08052025_PF_FP_ABST
    Figure JP2023039045_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention protects the privacy of a photographed person in an image to be subjected to personal feature extraction. A control unit (11) obtains, from a person region (a1) of an input image, a first feature amount (V1) calculated by using a prescribed feature amount calculation (c1) and obtains, from a processed person region (a2) obtained by processing a portion of the person region of the input image, a second feature amount (V1a) calculated by using the prescribed feature amount calculation (c1). In addition, the control unit (11) determines the image processing region of the person region of the input image and the image processing characteristics of the image processing region such that the amount of change between the first feature amount (V1) and the second feature amount (V1a) is less than or equal to a prescribed value. Furthermore, the control unit (11) generates, for the input image, an image processing model (m0) that is machine-trained so as to generate an output image obtained by being processed on the basis of the determined image processing region and the image processing characteristics, and generates a processed image (g1a) obtained by using the image processing model (m0) to process a portion of the person region of a photographed image (g1).
Need to check novelty before this filing date? Find Prior Art

Description

Image processing program, image processing device, and image processing method

[0001] The present invention relates to an image processing program, an image processing device, and an image processing method.

[0002] In recent years, continuous authentication technology has been advancing, which extracts a person's features from images taken by surveillance cameras installed in the area and estimates their location in real time.By introducing such continuous authentication systems in places such as airports and shopping centers, it becomes possible to provide personalized services based on each customer's purchasing behavior, etc.

[0003] Related technologies include, for example, a technology that determines the degree of abstraction for each of multiple partial regions of an object to be confirmed contained in a captured image and generates an image in which the partial regions are abstracted. Another technology has been proposed that generates an obfuscated image by flattening the gradation of a privacy protection image. Another technology has been proposed that determines a person's face in a first image as a bounding box and changes the contrast value of the bounding box to generate a second image. Yet another technology has been proposed that displays a person's face in a surveillance image as a colored filler.

[0004] Furthermore, a technique has been proposed that extracts a masking target portion from an input image, masks the target portion, and then removes the masking based on a specific procedure. Furthermore, a technique has been proposed that performs anomaly detection using a mosaic of surveillance camera footage. Furthermore, a technique has been proposed that uses a multi-pinhole camera as a surveillance camera and performs object recognition using blurred images.

[0005] Japanese Patent Publication No. 2021-149747 Japanese Patent Publication No. 2022-96519 U.S. Patent No. 11,604,938 U.S. Patent Application Publication No. 2019 / 0377958 International Publication No. 2018 / 225775 International Publication No. 2021 / 210313 International Publication No. 2021 / 176899

[0006] However, in conventional continuous authentication, personal feature extraction uses images that can identify the person receiving personalized services. This means that the images can be viewed by third parties providing the services, which poses a problem of not protecting the privacy of the person being photographed.

[0007] In one aspect, the present invention aims to provide an image processing program, an image processing device, and an image processing method that are capable of protecting the privacy of a person photographed in an image from which personal features are extracted.

[0008] To solve the above problem, an image processing program is provided, which causes a computer to: determine a first feature amount calculated from a person area of ​​an input image by a predetermined feature amount calculation; determine a second feature amount calculated from an edited person area obtained by editing a part of the person area of ​​the input image by the predetermined feature amount calculation; determine an image editing area in the person area of ​​the input image and image editing characteristics for the image editing area so that a change amount between the first feature amount and the second feature amount is equal to or less than a predetermined value; generate an image editing model trained by machine learning to generate an output image edited on the input image based on the determined image editing area and image editing characteristics; and edit the part of the person area of ​​the captured image using the image editing model.

[0009] In order to solve the above problems, an image processing device that executes the same processing as the image processing program is provided. Furthermore, in order to solve the above problems, an image processing method that causes a computer to execute the same processing as the image processing program is provided.

[0010] According to one aspect, it is possible to protect the privacy of a person photographed in an image from which human feature extraction is performed. The above and other objects, features, and advantages of the present invention will become apparent from the following description taken in conjunction with the accompanying drawings illustrating preferred embodiments of the present invention by way of example.

[0011] 1 is a diagram for explaining an example of an image processing device. FIG. 1 is a diagram illustrating an example of the configuration of an image processing system. FIG. 2 is a diagram illustrating an example of the hardware configuration of an image processing device. FIG. 3 is a diagram illustrating an example of functional blocks of an image processing device and a continuous authentication device. FIG. 4 is a flowchart illustrating an example of a training operation of a noise image generation model. FIG. 5 is a flowchart illustrating an example of an image processing operation. FIG. 6 is a flowchart illustrating an example of a first learning process of a trained noise image generation model. FIG. 7 is a flowchart illustrating an example of a second learning process of a trained noise image generation model. FIG. 8 is a flowchart illustrating an example of a third learning process of a trained noise image generation model. FIG. 9 is a flowchart illustrating an example of a fourth learning process of a trained noise image generation model. FIG. 10 is a flowchart illustrating an example of a fifth learning process of a trained noise image generation model. FIG.

[0012] The present embodiment will be described below with reference to the drawings. FIG. 1 is a diagram illustrating an example of an image processing device. The image processing device 10 includes a control unit 11 and a photographing unit 13. The image processing device 10 is applied, for example, to a system in which a feature extraction model is used to extract features of a person from a photographed image and perform continuous authentication. The function of the control unit 11 is realized, for example, by a processor (not shown) included in the image processing device 10 executing a predetermined program. The photographing unit 13 is, for example, a surveillance camera.

[0013] [Step S1] The control unit 11 calculates a first feature amount V1 from a person area a1 of the input image using a predetermined feature amount calculation c1. The control unit 11 also calculates a second feature amount V1a from a processed person area a2, which is a processed portion of the person area of ​​the input image, using the predetermined feature amount calculation c1. The control unit 11 then determines an image processing area in the person area of ​​the input image and image processing characteristics for the image processing area so that the amount of change between the first feature amount V1 and the second feature amount V1a is equal to or less than a predetermined value.

[0014] The image processing region corresponds to, for example, a non-interest region that is not focused on when a feature extraction model that performs a predetermined feature calculation c1 extracts features of a person. Furthermore, the image processing characteristics in the image processing region correspond to, for example, noise that conceals the individuality of a person.

[0015] [Step S2] The control unit 11 generates an image processing model m0 that has been machine-learned to generate an output image processed based on the determined image processing area and image processing characteristics for the input image.

[0016] [Step S3] The control unit 11 generates a processed image g1a by processing a part of the person area of ​​the photographed image g1 using the image processing model m0. In the example of Fig. 1, the image processing area is a person's face, and image processing has been performed to add noise to the person's face as an image processing characteristic.

[0017] In this way, the image processing device 10 processes a part of the person area of ​​the captured image using a machine-learned model that generates a processed image based on the image processing area where the person feature amount does not change and the image processing characteristics in the image processing area, thereby making it possible to protect the privacy of the photographed person in the image from which person feature extraction is performed.

[0018] 2 is a diagram showing an example of the configuration of an image processing system sy1. The image processing system sy1 includes image processing devices 10-1 to 10-4, a continuous authentication device 20, and monitoring cameras 13-1 to 13-4.

[0019] The continuous authentication device 20 is connected to image processing devices 10-1, ..., 10-4 via a network NW. Furthermore, surveillance cameras 13-1, ..., 13-4 are connected to the image processing devices 10-1, ..., 10-4, respectively, and are placed within area 4. Note that, in the example of Fig. 2, area 4 is configured to include a plurality of image processing devices with surveillance cameras, but it may also be configured to include a single image processing device with a surveillance camera.

[0020] The surveillance cameras 13-1, ..., 13-4 capture images of subjects in the area 4 where continuous authentication is performed. The image processing device 10-1 processes the frame images captured by the surveillance camera 13-1 to generate processed images as described above in Fig. 1 and transmits the processed images to the continuous authentication device 20 via the network NW. The image processing device 10-2 processes the frame images captured by the surveillance camera 13-2 to generate processed images as described above in Fig. 1 and transmits the processed images to the continuous authentication device 20 via the network NW.

[0021] Similarly, the image processing device 10-3 processes the frame image captured by the surveillance camera 13-3 to generate a processed image as described above in Fig. 1 and transmits it via the network NW to the continuous authentication device 20. The image processing device 10-4 processes the frame image captured by the surveillance camera 13-4 to generate a processed image as described above in Fig. 1 and transmits it via the network NW to the continuous authentication device 20.

[0022] The continuous authentication device 20 performs continuous authentication by performing ID identification (person re-identification: Re-ID) of the same person across multiple surveillance cameras based on feature quantities extracted from the received processed image as needed, such as the person's clothing and stature. At this time, frame images may be saved.

[0023] The processed image used by the continuous authentication device 20 when performing continuous authentication based on the feature extraction model is a noise image in which noise that conceals individuality is added to the original frame image so that the effect on the calculation results of the feature amounts by the feature extraction model is minimal.

[0024] For example, a noise image is an image in which the noise is added to an image area that the feature extraction model does not focus on for human feature extraction. This allows the continuous authentication device 20 to protect the privacy of people in images transmitted to the continuous authentication device 20 without degrading the performance of continuous authentication.

[0025] The image processing devices 10-1, ..., 10-4 and the continuous authentication device 20 can be realized, for example, by a computer. Note that, although the example in Fig. 2 shows a configuration in which the image processing devices are connected to a surveillance camera, the functions of the image processing devices may be implemented within the surveillance camera.

[0026] 3 is a diagram showing an example of the hardware configuration of an image processing device. The image processing device 10 is entirely controlled by a processor 101 having the functions of a control unit 11. A memory 102 and multiple peripheral devices are connected to the processor 101 via a bus 110. The processor 101 may be a multiprocessor. The processor 101 is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a DSP (Digital Signal Processor). At least some of the functions realized by the processor 101 executing a program may be realized by an electronic circuit such as an ASIC (Application Specific Integrated Circuit) or a PLD (Programmable Logic Device).

[0027] The memory 102 is used as a main storage device of the image processing device 10. The memory 102 temporarily stores at least a portion of the OS (Operating System) programs and application programs to be executed by the processor 101. The memory 102 also stores various data used in processing by the processor 101. The memory 102 may be a volatile semiconductor storage device such as a RAM (Random Access Memory).

[0028] The peripheral devices connected to the bus 110 include a storage device 103, a GPU (Graphics Processing Unit) 104, an input interface 105, an optical drive device 106, a device connection interface 107, and a network interface 108.

[0029] The storage device 103 electrically or magnetically writes and reads data to and from a built-in recording medium. The storage device 103 is used as an auxiliary storage device for the image processing device 10. The storage device 103 stores the OS program, application programs, and various data. Note that the storage device 103 may be, for example, a hard disk drive (HDD) or a solid state drive (SSD).

[0030] The GPU 104 is an arithmetic unit that performs image processing and is also called a graphics controller. The GPU 104 is connected to a display 201. The GPU 104 displays an image on the screen of the display 201 in accordance with an instruction from the processor 101.

[0031] A keyboard 202 and a mouse 203 are connected to the input interface 105. The input interface 105 transmits signals sent from the keyboard 202 and the mouse 203 to the processor 101. The mouse 203 is an example of a pointing device, and other pointing devices can also be used. Examples of other pointing devices include a touch panel, a tablet, a touch pad, and a trackball.

[0032] The optical drive device 106 uses a laser beam or the like to read data recorded on an optical disc 204 or write data to the optical disc 204. The optical disc 204 is a portable recording medium on which data is recorded so that it can be read by reflected light. Examples of the optical disc 204 include a DVD (Digital Versatile Disc), a DVD-RAM, a CD-ROM (Compact Disc Read Only Memory), and a CD-R (Recordable) / RW (Re-Writable).

[0033] The device connection interface 107 is a communication interface for connecting peripheral devices to the image processing device 10. A surveillance camera 13a is connected to the device connection interface 107. In addition, for example, a memory device 205 or a memory reader / writer 206 can be connected to the device connection interface 107. The memory device 205 is a recording medium equipped with a function for communicating with the device connection interface 107. The memory reader / writer 206 is a device for writing data to the memory card 207 or reading data from the memory card 207. The memory card 207 is a card-type recording medium.

[0034] The network interface 108 is connected to the network NW. The network interface 108 transmits and receives data to and from the constant authentication device 20 and other computers or communication devices via the network NW. The network interface 108 is, for example, a wired communication interface connected by a cable to a wired communication device such as a switch or a router. The network interface 108 may also be a wireless communication interface connected by radio waves to a wireless communication device such as an access point.

[0035] The image processing device 10 can realize the processing functions of the present invention using the hardware described above. The image processing device 10 realizes the processing functions of the present invention by executing a program recorded on, for example, a computer-readable recording medium. The program describing the processing content to be executed by the image processing device 10 can be recorded on various recording media.

[0036] For example, the program to be executed by the image processing device 10 can be stored in the storage device 103. The processor 101 loads at least a part of the program in the storage device 103 into the memory 102 and executes the program. The program to be executed by the image processing device 10 can also be recorded on a portable recording medium such as an optical disk 204, a memory device 205, or a memory card 207. The program stored on the portable recording medium becomes executable after being installed on the storage device 103 under the control of the processor 101, for example. The processor 101 can also read and execute the program directly from the portable recording medium. The continuous authentication device 20 shown in FIG. 2 also includes a computer and can be realized by similar hardware as shown in FIG. 3.

[0037] 4 is a diagram showing an example of functional blocks of the image processing device and the continuous authentication device. In the following description, it is assumed that a noise image to which noise is added is generated as the processed image.

[0038] The image processing device 10 includes a control unit 11 and a storage unit 12. The control unit 11 includes a noise image generation model training unit 11a, an image processing unit 11b, and an output unit 11c. The storage unit 12 includes a teacher data DB (database) 12a, a machine learning model DB 12b, a trained noise image generation model DB 12c, and a frame image DB 12d.

[0039] The noise image generation model training unit 11a inputs teacher data into the machine learning model to train the noise image generation model by machine learning, generating a trained noise image generation model (parameters of the trained noise image generation model). The image processing unit 11b inputs frame images acquired via a surveillance camera into the trained noise image generation model and processes them to generate noise images in which individuality is concealed but which can be authenticated at all times. The output unit 11c outputs the generated noise images to the continuous authentication device 20.

[0040] The teacher data DB12a stores, for example, training frame images, training frame image person area coordinate data, and non-important color data that is not emphasized in Re-ID processing (or important color data that is emphasized in Re-ID processing).

[0041] The machine learning model DB 12b stores data of trained machine learning models used to train the noise image generation model. The machine learning model DB 12b stores at least the Re-ID feature extraction models (trained Re-ID feature extraction models) used in the Re-ID processing performed by the continuous authentication device 20 connected to the output unit 11c.

[0042] Furthermore, the machine learning model DB 12b stores, for example, a face detection model that detects faces from people, and Gradient-weighted-Class Activation Mapping (Grad-CAM). Note that Grad-CAM is a model that focuses on features extracted by the convolutional layer of a Convolutional Neural Network (CNN) and visualizes which parts of an image the machine learning is focusing on.

[0043] The trained noise image generation model DB 12c stores the trained noise image generation model (parameters of the trained noise image generation model) generated by the noise image generation model training unit 11a. The frame image DB 12d stores frame images captured by a surveillance camera.

[0044] The continuous authentication device 20 includes a Re-ID feature extraction model DB 21, a person feature extraction unit 22, and a continuous authentication processing unit 23. The Re-ID feature extraction model DB 21 stores a Re-ID feature extraction model. The person feature extraction unit 22 inputs a noise image output from the image processing device 10 (a noise image in which noise that conceals individuality is added to a non-interest area of ​​the Re-ID feature extraction model) into the Re-ID feature extraction model to extract a person area, and extracts person features from the area.

[0045] The continuous authentication processing unit 23 performs continuous authentication based on the Re-ID from the extracted personal features. The continuous authentication device 20 performs continuous authentication based on an image in which noise is added to non-interest areas of the Re-ID, so the privacy of the photographed person is protected and continuous authentication with high accuracy can be performed.

[0046] 5 is a flowchart showing an example of the training operation of the noise image generation model. [Step S11] The control unit 11 reads the trained Re-ID feature extraction model from the machine learning model DB 12b.

[0047] [Step S12] The control unit 11 reads training data from the training data DB 12a. [Step S13] The control unit 11 performs machine learning using the training data to train a noise image generation model. In this machine learning, Re-ID features obtained by inputting the original training data (images including people) into a trained Re-ID feature extraction model are compared with Re-ID features obtained by inputting image data obtained by adding noise to the training data into the trained Re-ID feature extraction model.

[0048] [Step S14] The control unit 11 generates parameters of the trained noise image generation model and stores the generated parameters in the trained noise image generation model DB 12 c. If the noise image generation model is a model formed by a neural network, the parameters include weighting coefficients between nodes included in the neural network.

[0049] 6 is a flowchart showing an example of the image processing operation. [Step S21] The control unit 11 reads the parameters of the trained noise image generation model from the trained noise image generation model DB 12c.

[0050] [Step S22] The control unit 11 reads a frame image from the frame image DB 12d. [Step S23] The control unit 11 inputs the frame image into a trained noise image generation model to generate a noise image in which noise is added to the frame image. This noise image is an image in which noise is added to a person area in the frame image without reducing the accuracy of continuous authentication and concealing individuality. Note that the trained noise image generation model may, for example, be a model that adds noise similar to the noise added to the person area to at least a portion of a background area (e.g., a randomly selected area).

[0051] [Step S24] The control unit 11 outputs the noise image to the continuous authentication device 20. Next, the learning procedure for training the noise image generation model will be described in detail with reference to Figures 7 to 12. In the first to fourth learning processes described below, different areas of the person area are targeted for noise addition.

[0052] The trained parameters are parameters that change the position of the noise addition region in the noise addition target region and the noise type (such as the color or shape of the noise) as an image processing characteristic. However, for privacy protection, it is desirable to train the noise addition region so that it is as large as possible. Furthermore, multiple noise addition regions may be set, and different types of noise may be added to each noise addition region. Conversely, the noise addition region may be set over the entire person region, and one type of noise may be uniformly added.

[0053] 7 is a flowchart showing an example of a first learning process of a trained noise image generation model, which changes pixel values ​​in a person region.

[0054] [Step S31] The control unit 11 reads a training frame image from the teacher data DB 12a. [Step S32] The control unit 11 reads training frame image person area coordinate data from the teacher data DB 12a.

[0055] [Step S33] The control unit 11 reads the Re-ID feature extraction model (trained Re-ID feature extraction model) from the machine learning model DB 12b. [Step S34] The control unit 11 initializes the parameters of the noise image generation model. In the first learning process, the entire person region becomes the noise-added region. Therefore, parameters that change the position of the noise-added region in the person region and the type of noise become the training target (the target of update in step S40). In step S34, the position of the noise-added region in the person region and the type of noise are initially set using, for example, random values.

[0056] [Step S35] The control unit 11 generates an original person region image from the training frame image and the training frame image person region coordinate data. Then, the control unit 11 inputs the original person region image to the Re-ID feature extraction model and extracts first Re-ID features from the original person region image.

[0057] [Step S36] The control unit 11 uses the noise image generation model with the current parameters set to modify and process the pixel values ​​of the original person area image (i.e., add noise) to generate a processed person area image.

[0058] [Step S37] The control unit 11 inputs the processed person area image generated in step S36 to the Re-ID feature extraction model, and extracts second Re-ID features from the processed person area image.

[0059] [Step S38] The control unit 11 calculates the similarity between the first Re-ID feature extracted from the original person area image and the second Re-ID feature extracted from the processed person area image.

[0060] [Step S39] If the calculated similarity is equal to or greater than the threshold, the control unit 11 determines that the similarity has converged and proceeds to processing in step S41. If the calculated similarity is less than the threshold, the control unit 11 determines that the similarity has not converged and proceeds to processing in step S40.

[0061] In addition, when the similarity between the first Re-ID feature and the second Re-ID feature converges, the processed person area image from which the second Re-ID feature has been extracted is an image in which the person's individuality is concealed and privacy is protected, but the accuracy of Re-ID processing for that person is not reduced.

[0062] [Step S40] The control unit 11 updates the parameters of the noise image generation model. The process returns to step S36. [Step S41] The control unit 11 outputs the current parameters of the noise image generation model as parameters of the trained noise image generation model. This generates a trained noise image generation model that generates a noise image with noise added to a person area when an image is input. The output parameters are stored in the trained noise image generation model DB 12c.

[0063] In the first learning process described above, an image processing area is searched for from the entire person area under the condition that the amount of change in the Re-ID feature amount is equal to or less than a certain magnitude corresponding to the threshold value, and noise is determined as the image processing characteristic of the image processing area. The searched image processing area is estimated to be an area of ​​the person area that receives little attention in the Re-ID processing. Through this learning, a trained noise image generation model is generated that adds noise to the person area in the input image. This trained noise image generation model makes it possible to generate an image that allows for highly accurate extraction of person features using Re-ID processing while protecting the privacy of the person being photographed.

[0064] 8 is a flowchart showing an example of a second learning process of the trained noise image generation model, which changes pixel values ​​in a face region.

[0065] [Step S51] The control unit 11 reads a training frame image from the teacher data DB 12a. [Step S52] The control unit 11 reads training frame image person area coordinate data from the teacher data DB 12a.

[0066] [Step S53] The control unit 11 reads the Re-ID feature extraction model (trained Re-ID feature extraction model) from the machine learning model DB 12b. [Step S54] The control unit 11 reads the face detection model from the machine learning model DB 12b.

[0067] [Step S55] The control unit 11 initializes the parameters of the noise image generation model. In the second learning process, the face area detected from the person area becomes the noise addition target area. Therefore, the updated parameters are parameters that change the position of the noise addition area in the face area and the noise type.

[0068] [Step S56] The control unit 11 generates an original person region image from the training frame image and the training frame image person region coordinate data. Then, the control unit 11 inputs the original person region image to the Re-ID feature extraction model and extracts first Re-ID features from the original person region image.

[0069] [Step S57] The control unit 11 performs face detection on the original person area image using the face detection model and acquires the face area coordinates. [Step S58] The control unit 11 generates an original face area image from the face area coordinates, and processes the pixel values ​​of the original face area image (i.e., adds noise) using the noise image generation model with the current parameters set, to generate a processed face area image.

[0070] [Step S59] The control unit 11 inputs the processed person area image, in which the image of the face area of ​​the original person area image has been replaced with the processed face area image generated in step S58, into the Re-ID feature extraction model, and extracts second Re-ID features from the processed person area image.

[0071] [Step S60] The control unit 11 calculates the similarity between the first Re-ID feature extracted from the original person area image and the second Re-ID feature extracted from the processed person area image.

[0072] [Step S61] If the calculated similarity is greater than or equal to the threshold, the control unit 11 determines that the similarity has converged and proceeds to processing in step S63. If the calculated similarity is less than the threshold, the control unit 11 determines that the similarity has not converged and proceeds to processing in step S62.

[0073] [Step S62] The control unit 11 updates the parameters of the noise image generation model. The process returns to step S58. [Step S63] The control unit 11 outputs the current parameters of the noise image generation model as parameters of the trained noise image generation model. This generates a trained noise image generation model that, when an image is input, generates a noise image in which noise is added to the face region of a person region. The output parameters are stored in the trained noise image generation model DB 12c.

[0074] In the second learning process described above, under the condition that the amount of change in the Re-ID feature amount is equal to or less than a certain magnitude corresponding to the threshold value, a face area in a person area is searched for as an image processing area, and noise is determined as an image processing characteristic in the image processing area. The searched face area is estimated to be an area that receives little attention in the Re-ID processing. Through this learning, a trained noise image generation model is generated that adds noise to the face area in the person area on the input image. This trained noise image generation model makes it possible to generate an image that allows for highly accurate extraction of person features using the Re-ID processing while protecting the privacy of the person being photographed.

[0075] A technique for adding noise to a facial region by changing the pixel values ​​of the facial region is disclosed, for example, in "Dietlmeier, J. Antony, K. McGuinness and N. O'Connor, "How important are faces for person re-identification?" in 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 2021, pp. 6912-6919." This technique claims that blurring or blacking out a face has little effect on the accuracy of continuous authentication.

[0076] Furthermore, by adding noise not only to the facial area of ​​a person but also to the background area where there is no face, it is possible to not only conceal the individual but also make the presence of the person difficult to discern, thereby protecting the privacy of the entire image.

[0077] 9 is a flowchart showing an example of a third learning process of a trained noise image generation model, which changes pixel values ​​of non-important colors.

[0078] [Step S71] The control unit 11 reads a training frame image from the teacher data DB 12a. [Step S72] The control unit 11 reads training frame image person area coordinate data from the teacher data DB 12a.

[0079] [Step S73] The control unit 11 reads the Re-ID feature extraction model (trained Re-ID feature extraction model) from the machine learning model DB 12b. [Step S74] The control unit 11 reads non-important color data from the training data DB 12a. Note that the important color data is color data that is important for extracting features in the Re-ID processing, and the non-important color data is color data that is not important for extracting features in the Re-ID processing.

[0080] The teacher data DB 12a stores in advance important color data, which is RGB information representing important colors converted into data, and non-important color data, which is RGB information representing non-important colors converted into data. A method for determining whether a color is important or non-important is, for example, to create a histogram of color information from a data set of human images and determine colors with low frequency of appearance as important colors and other colors as non-important colors.

[0081] [Step S75] The control unit 11 initializes the parameters of the noise image generation model. In the third learning process, areas of non-important colors within the person area are designated as noise-added areas. Therefore, parameters that change the position and type of noise in areas of non-important colors are the training targets (the parameters that are updated in step S81). In step S75, the position and type of noise in areas of non-important colors are initially set using, for example, random values.

[0082] [Step S76] The control unit 11 generates an original person region image from the training frame image and the training frame image person region coordinate data. Then, the control unit 11 inputs the original person region image to the Re-ID feature extraction model and extracts first Re-ID features from the original person region image.

[0083] [Step S77] The control unit 11 generates an edited person area image by changing pixel values ​​of non-important color areas (predetermined color areas) of the original person area image using parameters of the noise image generation model. When changing pixel values ​​of non-important colors, it is preferable to change the pixel values ​​to colors that are less likely to reveal individual characteristics, such as black, white, or colors close to skin.

[0084] [Step S78] The control unit 11 inputs the processed person area image generated in step S77 to the Re-ID feature extraction model, and extracts second Re-ID features from the processed person area image.

[0085] [Step S79] The control unit 11 calculates the similarity between the first Re-ID feature extracted from the original person area image and the second Re-ID feature extracted from the processed person area image.

[0086] [Step S80] If the calculated similarity is greater than or equal to the threshold, the control unit 11 determines that the similarity has converged and proceeds to processing in step S82. If the calculated similarity is less than the threshold, the control unit 11 determines that the similarity has not converged and proceeds to processing in step S81.

[0087] [Step S81] The control unit 11 updates the parameters of the noise image generation model. The process returns to step S77. [Step S82] The control unit 11 outputs the current parameters of the noise image generation model as parameters of the trained noise image generation model. This generates a trained noise image generation model that, when an image is input, generates a noise image in which noise is added to non-important color regions of a person area. The output parameters are stored in the trained noise image generation model DB 12c.

[0088] The Re-ID feature extraction model places emphasis on appearance features that contain a large amount of information, such as colorful colors and logos on clothing, and that are more representative of individuality, as disclosed, for example, in "Liu, C., Gong, S., Loy, CC, Lin, X. (2012). Person Re-identification: What Features Are Important?. In: Fusiello, A., Murino, V., Cucchiara, R. (eds) Computer Vision - ECCV 2012. Workshops and Demonstrations. ECCV 2012. Lecture Notes in Computer Science, vol. 7583. Springer, Berlin, Heidelberg. https: / / doi.org / 10.1007 / 978-3-642-33863-2_39." Therefore, the image processing device 10 generates a noise image by adding noise obtained by changing pixel values ​​in non-important color regions that the Re-ID feature extraction model does not emphasize to non-interest regions of the Re-ID feature extraction model. This makes it possible to generate an image that allows for the extraction of personal features by Re-ID processing while protecting the privacy of the person being photographed.

[0089] In the third learning process described above, an image processing area is searched for from a non-important color area of ​​a person area under the condition that the amount of change in the Re-ID feature amount is equal to or less than a certain magnitude corresponding to the threshold value, and noise is determined as the image processing characteristic of the image processing area. The searched image processing area is estimated to be a non-important color area that receives little attention in Re-ID processing. Through this learning, a trained noise image generation model is generated that adds noise to the non-important color area of ​​a person area on the input image. This trained noise image generation model makes it possible to generate an image that allows for highly accurate human feature extraction using Re-ID processing while protecting the privacy of the person being photographed.

[0090] 10 is a flowchart showing an example of a fourth learning process of the trained noise image generation model, which changes pixel values ​​of non-interest regions of the Re-ID feature extraction model.

[0091] [Step S91] The control unit 11 reads a training frame image from the teacher data DB 12a. [Step S92] The control unit 11 reads training frame image person area coordinate data from the teacher data DB 12a.

[0092] [Step S93] The control unit 11 reads the Re-ID feature extraction model (trained Re-ID feature extraction model) from the machine learning model DB 12b. [Step S94] The control unit 11 reads the Grad-CAM from the machine learning model DB 12b.

[0093] [Step S95] The control unit 11 initializes the parameters of the noise image generation model. In the fourth learning process, the image of a non-interest area in the person area becomes the noise-added area. Therefore, parameters that change the position and noise type of the non-interest area in the person area become the training target (the parameters that are updated in step S102). In step S95, the position and noise type of the noise-added area in the non-interest area are initially set using, for example, random values.

[0094] [Step S96] The control unit 11 generates an original person region image from the training frame image and the training frame image person region coordinate data. Then, the control unit 11 inputs the original person region image to the Re-ID feature extraction model and extracts first Re-ID features from the original person region image.

[0095] [Step S97] The control unit 11 inputs the original person area image into the Re-ID feature extraction model, and applies Grad-CAM to acquire area-of-interest information for the original person area image.

[0096] [Step S98] Based on the information on the region of interest, the control unit 11 detects a region of non-interest in the original person region image, and generates a processed person region image by changing the pixel values ​​of the region of non-interest in the original person region image using parameters of the noise image generation model.

[0097] [Step S99] The control unit 11 inputs the processed person area image generated in step S98 to the Re-ID feature extraction model, and extracts second Re-ID feature amounts for the processed person area image.

[0098] [Step S100] The control unit 11 calculates the similarity between the first Re-ID feature extracted from the original person area image and the second Re-ID feature extracted from the processed person area image.

[0099] [Step S101] If the calculated similarity is equal to or greater than the threshold, the control unit 11 determines that the similarity has converged and proceeds to processing in step S103. If the calculated similarity is less than the threshold, the control unit 11 determines that the similarity has not converged and proceeds to processing in step S102.

[0100] [Step S102] The control unit 11 updates the parameters of the noise image generation model. The process returns to step S98. [Step S103] The control unit 11 outputs the current parameters of the noise image generation model as parameters of the trained noise image generation model. This generates a trained noise image generation model that, when an image is input, generates a noise image in which noise is added to non-interest areas of the person area. The output parameters are stored in the trained noise image generation model DB 12c.

[0101] In the fourth learning process described above, an image processing area is searched for from a non-interest area in a person area under the condition that the amount of change in the Re-ID feature amount is equal to or less than a certain magnitude corresponding to the threshold value, and noise is determined as the image processing characteristic in the image processing area. The non-interest area among the searched image processing areas is an area that receives a low level of attention in the Re-ID processing. Through this learning, a trained noise image generation model is generated that adds noise to the non-interest area in a person area on the input image. This trained noise image generation model makes it possible to generate an image that allows for highly accurate extraction of person features using Re-ID processing while protecting the privacy of the person being photographed.

[0102] 11 and 12 are flowcharts showing an example of a fifth learning process of a trained noise image generation model, which changes pixel values ​​in a person region without reducing the accuracy of person detection.

[0103] [Step S111] The control unit 11 reads a training frame image from the teacher data DB 12a. [Step S112] The control unit 11 reads person area coordinate data of the training frame image from the teacher data DB 12a.

[0104] [Step S113] The control unit 11 reads the Re-ID feature extraction model (trained Re-ID feature extraction model) from the machine learning model DB 12b. [Step S114] The control unit 11 reads the person detection model from the machine learning model DB 12b.

[0105] [Step S115] The control unit 11 initializes the parameters of the noise image generation model. In the fifth learning process, as in the first learning process, the entire person region becomes the noise-added region. Therefore, parameters that change the position of the noise-added region in the person region and the type of noise become the training target (the parameters that are updated in step S125). In step S115, the position of the noise-added region in the person region and the type of noise are initially set using, for example, random values.

[0106] [Step S116] The control unit 11 generates an original person region image from the training frame image and the training frame image person region coordinate data. Then, the control unit 11 inputs the original person region image to the Re-ID feature extraction model and extracts first Re-ID features from the original person region image.

[0107] [Step S117] The control unit 11 generates a processed person area image by changing the pixel values ​​of the original person area image using the parameters of the noise image generation model. [Step S118] The control unit 11 inputs the original person area image to the person detection model to perform person detection and obtain first person area coordinates.

[0108] [Step S119] The control unit 11 inputs the processed person area image into the person detection model to perform person detection and obtains second person area coordinates. [Step S120] The control unit 11 detects a first person area from the original person area image based on the first person area coordinates, and detects a second person area from the processed person area image based on the second person area coordinates.

[0109] [Step S121] The control unit 11 calculates the Intersection over Union (IoU) between the first person area detected from the original person area image and the second person area detected from the processed person area image, and obtains an IoU value that indicates the degree of overlap between the first person area and the second person area.

[0110] [Step S122] The control unit 11 inputs the processed person area image generated in step S117 to the Re-ID feature extraction model, and extracts second Re-ID features from the processed person area image.

[0111] [Step S123] The control unit 11 calculates the similarity between the first Re-ID feature extracted from the original person area image and the second Re-ID feature extracted from the processed person area image.

[0112] [Step S124] The control unit 11 determines whether the similarity has converged to a threshold or more and the IoU value has converged to a threshold or more. If convergence has occurred, the process proceeds to step S126. If convergence has not occurred, the process proceeds to step S125.

[0113] [Step S125] The control unit 11 updates the parameters of the noise image generation model. The process returns to step S117. [Step S126] The control unit 11 outputs the current parameters of the noise image generation model as parameters of the trained noise image generation model. This generates a trained noise image generation model that generates a noise image with noise added to the person area when an image is input. The output parameters are stored in the trained noise image generation model DB 12c.

[0114] In the fifth learning process described above, an image processing area is searched for from the entire person area under the condition that the amount of change in the Re-ID feature amount is equal to or less than a certain magnitude corresponding to the threshold value, and noise is determined as the image processing characteristic of the image processing area. The searched image processing area is estimated to be an area of ​​the person area that receives little attention in the Re-ID processing. Through this learning, a trained noise image generation model is generated that adds noise to the person area in the input image. This trained noise image generation model makes it possible to generate images that enable highly accurate extraction of person features using Re-ID processing while protecting the privacy of the person being photographed without reducing person detection accuracy.

[0115] The control unit 11 may perform training by combining two or more of the first to fifth learning processes to generate a noise image generation model that selects an optimal process from the image processing processes used in the two or more learning processes to generate a noise image. Alternatively, the control unit 11 may use all of the image processing processes used in the two or more learning processes to generate a noise image that has little effect on feature calculation.

[0116] The image processing device of the present invention described above can be realized by a computer. In this case, a program describing the processing contents of the functions that the image processing device should have is provided. By executing the program on a computer, the processing functions are realized on the computer.

[0117] The program describing the processing contents can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include magnetic storage units, optical disks, magneto-optical recording media, and semiconductor memories. Examples of magnetic storage units include hard disk drives (HDDs), flexible disks (FDs), and magnetic tapes. Examples of optical disks include CD-ROMs / RWs. Examples of magneto-optical recording media include MOs (Magneto Optical disks).

[0118] When distributing a program, for example, the program may be recorded on a portable recording medium such as a CD-ROM and sold. Alternatively, the program may be stored in the memory of a server computer and transferred from the server computer to other computers via a network.

[0119] A computer that executes a program stores, for example, a program recorded on a portable recording medium or a program transferred from a server computer in its own storage unit. The computer then reads the program from its own storage unit and executes processing in accordance with the program. Note that the computer can also read the program directly from a portable recording medium and execute processing in accordance with the program.

[0120] The computer can also execute the processing according to the received program each time the program is transferred from a server computer connected via a network. At least a part of the processing functions can also be realized by electronic circuits such as DSPs, ASICs, and PLDs.

[0121] The foregoing merely illustrates the principles of the present invention. Further, since numerous modifications and changes will be apparent to those skilled in the art, the present invention is not limited to the exact construction and application shown and described above, and all corresponding modifications and equivalents are deemed to be within the scope of the present invention as defined by the appended claims and their equivalents.

[0122] REFERENCE SIGNS LIST 10 Image processing device 11 Control unit 13 Photography unit a1 Person area a2 Processed person area c1 Feature amount calculation V1 First feature amount V1a Second feature amount m0 ​​Image processing model g1 Photographed image g1a Processed image

Claims

1. An image processing program that causes a computer to execute the following processes: determine a first feature amount calculated from a person area of ​​an input image by a predetermined feature amount calculation; determine a second feature amount calculated from a processed person area obtained by processing a part of the person area of ​​the input image by the predetermined feature amount calculation; determine an image processing area in the person area of ​​the input image and image processing characteristics in the image processing area so that the change amount between the first feature amount and the second feature amount is equal to or less than a predetermined value; generate an image processing model that has been machine-learned to generate an output image that has been processed for the input image based on the determined image processing area and image processing characteristics; and process a part of the person area of ​​a captured image using the image processing model.

2. The image processing program according to claim 1, wherein the image processing characteristics indicate noise characteristics, and the processing generates an image processing model that adds the noise based on the image processing characteristics to a part of a person area of ​​the photographed image.

3. The image processing program according to claim 2, wherein said processing generates said image processing model for adding said noise to a face area of ​​a person area of ​​said photographed image.

4. The image processing program according to claim 2, wherein said processing generates said image processing model for adding said noise to a predetermined color area of ​​a person area of ​​said photographed image.

5. The image processing program of claim 2, wherein the processing determines a non-interest area excluding the interest area of ​​the person area in the captured image as the image processing area when the interest area is detected by a model that visualizes the interest area of ​​a feature extraction model that performs the specified feature calculation, and generates the image processing model that adds the noise to the non-interest area.

6. The image processing program of claim 1, wherein the processing determines the image processing area and the image processing characteristics in the person area of ​​the input image when a degree of overlapping coincidence between the person area of ​​the input image and the processed person area obtained by processing a portion of the person area of ​​the input image is equal to or greater than a predetermined value.

7. The image processing program of claim 1, wherein the processing comprises: extracting person features using a feature extraction model that performs the predetermined feature calculations, and outputting an image in which a portion of the person area in the captured image has been processed using the image processing model to a device that performs continuous authentication.

8. An image processing device having: a photographing unit; and a control unit that determines a first feature amount calculated from a person area of ​​an input image by a predetermined feature amount calculation, determines a second feature amount calculated from a processed person area obtained by processing a part of the person area of ​​the input image by the predetermined feature amount calculation, determines an image processing area in the person area of ​​the input image and image processing characteristics in the image processing area so that a change amount between the first feature amount and the second feature amount is not more than a predetermined value, generates an image processing model that has been machine-learned to generate an output image that has been processed for the input image based on the determined image processing area and image processing characteristics, and processes a part of the person area of ​​the captured image obtained from the photographing unit using the image processing model.

9. An image processing method in which a computer determines a first feature amount calculated from a person area of ​​an input image by a predetermined feature amount calculation, determines a second feature amount calculated from a processed person area obtained by processing a part of the person area of ​​the input image by the predetermined feature amount calculation, determines an image processing area in the person area of ​​the input image and image processing characteristics in the image processing area so that a change amount between the first feature amount and the second feature amount is not more than a predetermined value, generates an image processing model that has been machine-learned to generate an output image that has been processed for the input image based on the determined image processing area and image processing characteristics, and processes a part of the person area of ​​a captured image using the image processing model.

Citation Information

Patent Citations

  • Information processing unit, information processing method and program

    JP2021149747A

  • Learning method, image conversion device and program

    JP2022096519A

  • Systems for obscuring identifying information in images

    US11604938B1

  • Camera for monitoring a monitored area and monitoring device, and method for monitoring a monitored area

    US20190377958A1

  • Image masking device and image masking method

    WO2018225775A1