Computer program, image processing method, and image processing device

The described computer program and image processing method address the challenge of protecting individual privacy in real-life images by replacing facial images with realistic generated ones, maintaining the image scene and ensuring natural-looking replacements.

JP2025085486AActive Publication Date: 2025-06-05HIPERDYNE CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023199394
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-06-05
Estimated Expiration
2043-11-24

AI Technical Summary

Technical Problem

Existing technologies do not effectively protect the privacy of individuals in real-life images while maintaining the scene of the image.

Method used

A computer program and image processing method that acquires a real-life image, identifies attributes of individuals, and replaces their facial images with realistic generated facial images based on those attributes.

Benefits of technology

This approach effectively protects the privacy of individuals in real-life images while preserving the scene, ensuring that the replaced facial images are realistic and natural-looking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085486000001_ABST
    Figure 2025085486000001_ABST
Patent Text Reader

Abstract

To provide a computer program, an image processing method, and an image processing device capable of protecting a privacy of a person included in a photographed image while keeping the scene of the photographed image.SOLUTION: The method comprises acquiring a photographed image including a person, specifying attributes of the person included in the acquired photographed image, and replacing a face image of the person with a photorealistic generated face image generated according to the specified attributes.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a computer program, an image processing method, and an image processing device. [Background technology]

[0002] Patent Document 1 discloses a technology that acquires a face image, performs face recognition, and outputs an age-hidden image in which the age is changed while maintaining the features of the person so as to ensure the accuracy of face authentication. It is possible to ensure the accuracy of face recognition and protect the user's age privacy information, thereby improving the reliability of face recognition technology.

[0003] Patent Document 2 discloses a technology that recognizes the face of a person included in an input image, draws the face using lines including feature points used for emotion estimation, and replaces the face image of the person included in the input image with the drawn face image. By replacing the face image in this way, privacy can be protected and emotion estimation can be performed. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 6651038 [Patent Document 2] JP 2020-086921 A Summary of the Invention [Problem to be solved by the invention]

[0005] Patent Documents 1 and 2 do not disclose a technique for protecting the privacy of people in a real-life image while maintaining the scene of the real-life image.

[0006] An object of the present disclosure is to provide a computer program, an image processing method, and an image processing device that can protect the privacy of people included in a real-life image while maintaining the scenery of the real-life image. [Means for solving the problem]

[0007] A computer program according to one aspect of the present disclosure causes a computer to execute a process of acquiring a live-action image including a person, identifying attributes of the person included in the acquired live-action image, and replacing a facial image of the person with a realistic generated facial image generated in accordance with the identified attributes.

[0008] An image processing method according to one aspect of the present disclosure acquires a real-life image including a person, identifies attributes of the person included in the acquired real-life image, and replaces a facial image of the person with a realistic generated facial image generated in accordance with the identified attributes.

[0009] An image processing device according to one aspect of the present disclosure is an image processing device including a processing unit that acquires a real-life image including a person and performs image processing, wherein the processing unit identifies attributes of the person included in the acquired real-life image and replaces the face image of the person with a realistic generated face image generated in accordance with the identified attributes. Effect of the Invention

[0010] According to the present disclosure, it is possible to protect the privacy of people included in a real-life image while maintaining the scenery of the real-life image. [Brief description of the drawings]

[0011] [Figure 1] 1 is a block diagram showing an example of the configuration of an image processing device according to a first embodiment. [Diagram 2] FIG. 2 is a data flow diagram of the image processing method according to the first embodiment. [Diagram 3] 4 is a flowchart showing the processing steps of an image processing method according to the first embodiment. [Figure 4] FIG. 1 is a conceptual diagram showing a face image replacement method for privacy protection. [Diagram 5] FIG. 13 is a conceptual diagram illustrating a trace table. [Figure 6]1A and 1B are schematic diagrams showing moving images before and after image processing related to privacy protection. [Figure 7] 1 is a conceptual diagram showing a method for tracking a person to be replaced in a video and replacing a face image. [Figure 8] 10 is a flowchart showing the processing steps of an image processing method according to a second embodiment. [Figure 9] FIG. 13 is a conceptual diagram showing an image processing method for privacy protection using both a replacement process with a generated face image and a masking process. [Figure 10] 11 is a flowchart showing the processing steps of an image processing method according to a third embodiment. [Figure 11] 13 is a flowchart showing the processing steps of an image processing method according to a fourth embodiment. [Figure 12] 13A and 13B are schematic diagrams showing frame images of a moving image in which a face image has been replaced by the image processing method according to the fourth embodiment. [Figure 13] FIG. 13 is a conceptual diagram showing a tracking table when a generated face image is used in duplicate. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] An image processing device 1 according to an embodiment of the present disclosure will be described below with reference to the drawings. Note that the present disclosure is not limited to these examples, but is intended to include all modifications within the scope and meaning equivalent to the claims as set forth in the claims. In addition, at least some of the embodiments described below may be combined in any combination.

[0013] 1 is a block diagram showing an example of the configuration of an image processing device 1 according to embodiment 1. The image processing device 1 is a computer, and includes a processing unit 11, a display unit 12, an operation unit 13, a communication unit 14, and a storage unit 15. Each unit is connected via a bus.

[0014] The image processing device 1 according to this embodiment realizes privacy protection by identifying and tracking one or more persons included in a live-action video and replacing the facial image of the person with a realistic generated facial image P that is generated based on the attribute information of the person. In particular, the image processing device 1 according to the first embodiment does not perform image processing for privacy protection on the facial image of a specific person who does not require privacy protection (hereinafter referred to as a person not to be replaced), and replaces the facial images of other people (hereinafter referred to as a person to be replaced) with a generated facial image P (see FIG. 4). For example, the person not to be replaced is a celebrity who appears in the video, and the person to be replaced is a general audience member.

[0015] More specifically, the image processing device 1 estimates attribute information (age, sex, race, etc.) of each person appearing in the video, tracks the same person based on the results, and replaces the person's face with a different facial image. These replacement facial images are generated by AI according to the person's attributes, and although they look human, they are different from the person's original image. This allows the video to be used without worrying about portrait rights, while maintaining a more natural appearance than if the image were replaced with a character or mosaic, and protecting privacy.

[0016] The image processing device 1 may be a standalone computer or a server device connected to a network. The image processing device 1 may be a computer in an on-premise environment or a computer such as a server in a cloud environment. The image processing device 1 may be configured to perform distributed processing using multiple computers, may be realized by multiple virtual machines provided in one server, or may be realized using a cloud server.

[0017] The processing unit 11 is a processor having an arithmetic circuit such as a CPU (Central Processing Unit), a multi-core CPU, a GPU (Graphics Processing Unit), a GPGPU (General-purpose computing on graphics processing units), a TPU (Tensor Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), an NPU (Neural Processing Unit), etc., an internal storage device such as a ROM (Read Only Memory) and a RAM (Random Access Memory), an I / O terminal, a clock unit, etc. The processing unit 11 performs the image processing method according to the first embodiment by executing a computer program (program product) 15a stored in a storage unit 15 described later.

[0018] From another viewpoint, the processing unit 11 executes the computer program 15a to operate as functional units such as a human body / face detection unit 11a, an attribute extraction unit 11b, a person identification unit 11c, a replacement target person selection unit 11d, a face image generation unit 11e, and a face image replacement unit 11f (see FIG. 2). Note that each functional unit of the image processing device 1 may be realized by software, or a part or all of it may be realized by hardware.

[0019] The display unit 12 is, for example, a display device such as a liquid crystal panel or an organic EL (Electro Luminescence) display.

[0020] The operation unit 13 is an input device such as a hardware keyboard, a mouse, a touch panel, etc. A user can operate the image processing device 1 using the operation unit 13 to execute processes such as importing and playing back video data.

[0021] The communication unit 14 includes a communication circuit for performing communication processing with an external device. The communication unit 14 of the present embodiment 1 acquires video data by accessing a video camera, an external server storage, etc. (not shown). A video is composed of a plurality of frame images in a time series. The frame images constituting a live-action video are an example of live-action images that are the subject of image processing according to the present embodiment 1.

[0022] The storage unit 15 includes, for example, a main storage unit and an auxiliary storage unit. The main storage unit is a temporary storage area such as a static random access memory (SRAM), a dynamic random access memory (DRAM), or a flash memory, and temporarily stores data required for the processing unit 11 to execute arithmetic processing. The auxiliary storage unit is a storage device such as a hard disk or an electrically erasable programmable ROM (EEPROM). The storage unit 15 stores a computer program 15a executed by the processing unit 11, a face recognition trained model 15b, a face image generation trained model 15c, and a face image replacement trained model 15d. Details of the processing based on the computer program 15a will be described later.

[0023] The face recognition trained model 15b includes, for example, a model trained by deep learning such as a convolutional neural network (CNN) or a vision transformer. It may also be configured as a combination of these. The face recognition trained model 15b has an input layer to which frame images of a video are input, an intermediate layer to extract features of the frame images, and an output layer to output an inference result related to a detected object.

[0024] The inference result includes information such as the center coordinate position and horizontal and vertical sizes of a bounding box surrounding the face image, image features of the face (landmarks, key points, etc.), and "attributes." The "attributes" include at least the facial attributes of the person, such as the age, sex, race, face direction, whether or not the person is wearing glasses, whether or not the person is wearing a mask, and emotions. Furthermore, the "attributes" may include environmental attributes that affect the scenery of the face image and the frame image. The environmental attributes include, for example, information such as the lighting conditions and sunlight conditions for the person's face. The lighting conditions are information indicating the presence or absence of a lighting source that illuminates the face, the location of the lighting, and the color of the light source. The sunlight conditions are information indicating the shadows that appear on the face (the effects of shade and sunlight). The environmental attributes may also include information indicating the degree of wetness of the face (the effect of weather), the degree of sweating (the effect of temperature), etc.

[0025] In addition, the face recognition trained model 15b may be configured to infer a human body in addition to the face of a person included in a frame image of a video. In this case, the inference result includes information such as a bounding box (center coordinate position and vertical and horizontal sizes) surrounding an image of a human body, image features of the human body, etc.

[0026] The face image generation trained model 15c and the face image replacement trained model 15d include models such as generative adversarial networks (GANs), diffusion, and autoencoders that have been trained using deep learning. They may also be configured in combination. Of course, equivalent functions may be achieved using methods other than deep learning models.

[0027] The face image generation trained model 15c is a trained model that outputs an image having the features expressed in a prompt by inputting the features of an image to be generated and / or inputting a prompt expressing the features in text. The face image generation trained model 15c used in the present embodiment 1 has a function of outputting a realistic generated face image P according to the attributes by directly inputting the features of an image to be generated and / or giving a prompt input expressing the features in text to the face image generation trained model 15c. Specifically, when age, sex, and race are specified in the prompt, the face image generation trained model 15c has a function of outputting a generated face image P according to the specified age, sex, and race.

[0028] Furthermore, the face image generation trained model 15c may be configured to have a function of outputting a generated face image P according to a specified face direction or facial expression when a face direction or facial expression is specified. Of course, the face image generation trained model 15c may be configured to have a function of outputting a generated face image P according to a specified face direction and facial expression when a face direction and facial expression are specified. When the presence or absence of glasses, the presence or absence of a mask, and environmental attributes are specified, the face image P with glasses, the generated face image P with a mask, the generated face image P reflecting the lighting conditions according to the environmental attributes, etc. may be configured to have a function of outputting a generated face image P with glasses, the generated face image P with a mask, the generated face image P reflecting the lighting conditions according to the environmental attributes, etc., according to the specification.

[0029] The face image replacement learned model 15d may be configured to partially adjust the image features (landmarks, key points, etc.) of the face in the face image generated by the face image generation learned model 15c and regenerate a new face image. The face image replacement learned model 15d may be configured to input information on the image features (landmarks, key points, etc.) of the face of the person to be replaced and information on the image features (landmarks, key points, etc.) of the face of the generated face image to the face replacement learned model, and adjust the position and orientation of the image features (landmarks, key points, etc.) of the face in the input generated face image according to the position and orientation of the image features (landmarks, key points, etc.) of the face in the face image of the person to be replaced, thereby regenerating a generated face image P according to the facial orientation or expression of the person to be replaced. Needless to say, the face image replacement trained model 15d may be configured to regenerate a generated face image P according to the facial orientation and facial expression of the person to be replaced.

[0030] In this embodiment 1, an example has been described in which the memory unit 15 of the image processing device 1 stores the face recognition trained model 15b, the face image generation trained model 15c, and the face image replacement trained model 15d, but it may also be configured to utilize an external API (Application Programming Interface) that recognizes and generates face images.

[0031] The storage unit 15 also stores a non-replacement target person feature for identifying a specific person who does not require privacy protection. For example, a user may select a face image of a specific person who does not require privacy protection, and the processing unit 11 may extract a feature for identifying the person from the selected face image, and store the extracted feature as a non-replacement target person feature.

[0032] The auxiliary storage unit may be an external storage device connected to the image processing device 1. The computer program 15a, the face recognition learned model 15b, and the face image generation learned model 15c may be written to the storage unit 15 during the manufacturing stage of the image processing device 1, or the image processing device 1 may acquire through communication what is distributed by the external image processing device 1 and store it in the storage unit 15. The computer program 15a, the face recognition learned model 15b, and the face image generation learned model 15c may be readably recorded on a recording medium 1a such as a magnetic disk, an optical disk, or a semiconductor memory, or may be read from the recording medium 1a by a reading device and stored in the storage unit 15.

[0033] FIG. 2 is a data flow diagram of the image processing method according to the first embodiment, FIG. 3 is a flowchart showing the processing steps of the image processing method according to the first embodiment, and FIG. 4 is a conceptual diagram showing a face image replacement method for privacy protection.

[0034] The processing unit 11 of the image processing device 1 acquires an actual image (moving image or still image) (step S111). For example, the processing unit 11 acquires moving image data. The processing unit 11 may acquire the moving image data from the outside via the communication unit 14, or may acquire the moving image data by reading out the moving image data stored in the storage unit 15. Here, it is assumed that the moving image includes a plurality of people, and the plurality of people includes one person not to be replaced.

[0035] Next, the processing unit 11 functioning as the human body / face detection unit 11a and the attribute extraction unit 11b detects the regions of the human bodies and faces of multiple people included in the frame images of the video using the face recognition trained model 15b, and extracts the attributes of the people (step S112). The processing unit 11 inputs the frame images to the face recognition trained model 15b, thereby obtaining attribute information indicating the positions and sizes of the human body and face images of each of the multiple people included in the frame images, and the features of the human bodies and faces. Note that, although an example has been described in which the regions of the human body and face images of people included in the frame images are detected and attributes are extracted using a machine-learned trained model, such detection and attribute extraction processing may be performed using rule-based processing such as pattern matching.

[0036] Then, the processing unit 11 functioning as the person identification unit 11c identifies and tracks each of the multiple persons (step S113). Here, the processing unit 11 tracks each of the multiple identified persons and records data required to continuously replace the face image of each identified person with the same generated face image P in a tracking table. The tracking table is stored in the storage unit 15. Note that the tracking table is an example of a method of storing data required for tracking and managing identified persons, and is not limited to a table format. Needless to say, the process of step S113 is executed when the real image is a moving image. When the real image is a still image, the process of step S114 is executed after step S112. The same applies to the other embodiments.

[0037] 5 is a conceptual diagram showing a tracking table. The tracking table stores, for example, a person number (No.) for identifying each of a plurality of face images detected from a frame image, tracking information for tracking the face image, information for indicating whether the face image is a replacement image, and information indicating a generated face image P that replaces the face image, in association with each other.

[0038] The tracking information includes, for example, the whole body feature amount of the person, the facial feature amount, the whole body movement amount, the facial movement amount, etc. The whole body feature amount for tracking the face image includes, for example, the whole body image feature amount (landmarks, key points, etc.). The facial feature amount for tracking the face image includes, for example, the facial image feature amount (landmarks, key points, etc.). The tracking information may include information indicating the position and range of the face image.

[0039] In the case of video input, offline processing may be performed by playing the video in the reverse direction and collecting tracking information of people in the same way to complement the tracking information from forward playback (normal playback). This complementation makes the accuracy of tracking information after frame-in / fade-in / scene change in forward playback equivalent to the accuracy of tracking information before frame-out / fade-out / scene change in reverse playback.

[0040] The information indicating whether or not an image is a replacement image is information identified in step S114 described below, and the information indicating the generated facial image P is information for identifying the generated facial image P generated in step S115 described below, and continuously linking the generated facial image P with the facial image of a specific person.

[0041] Returning to FIG. 3, following the process of step S113, the processing unit 11 functioning as the replacement target person selection unit 11d identifies the face image of the replacement target person whose face image is to be replaced and the face image of the non-replacement target person whose face image is not to be replaced based on the feature amount of the face image of each of the detected multiple people and the non-replacement target person feature amount stored in the storage unit 15, and selects the replacement target person (step S114). Here, the processing unit 11 stores the identification result in the tracking table. A person having the same or similar feature amount as the non-replacement target person feature amount is identified as a non-replacement target person. The other people who are not identified as non-replacement targets are identified as replacement target people. In the example shown in FIG. 4, six people are included in the frame image, the central person is a non-replacement target person, and the other five people are replacement target people.

[0042] It should be noted that if a person not to be replaced is not designated or set, the face images of all people included in the frame image are specified as face images to be replaced.

[0043] Next, the processing unit 11 functioning as the face image generating unit 11e generates face images based on the attribute information of the plurality of persons to be replaced included in the frame image (step S115). Specifically, as shown in FIG. 4, the processing unit 11 generates an input (direct input and / or prompt input) including the attributes of the person to be replaced, and inputs the generated prompt to the face image generation trained model 15c, thereby generating a realistic face image for protecting the privacy of the person to be replaced. The prompt is, for example, text such as "Male in his twenties, Japanese, whole face, realistic image, 1 person". In order to associate the face image of the person being identified and tracked with the generated face image P, the processing unit 11 stores information indicating the generated face image P in association with the person number in the tracking table.

[0044] Next, the processing unit 11 may modify the face image generated by the face image generation trained model 15c by the method exemplified below, and regenerate the generated face image (step S116). For example, (1) convert the direction or expression of the generated face image P according to the face direction or expression of each person to be replaced. If a generated face image P facing forward is generated in step S115, processing may be performed to convert the direction of the generated face image P. In this case, the generated face images P facing forward, facing right, and facing left may be generated in advance in step S115, and the generated face image P to be used for replacing the face image may be selected according to the face direction of the person to be replaced. The same applies to facial expressions: if an expressionless generated facial image P is generated in step S115, processing may be performed to change facial expression-related parts such as the eyes and mouth of the generated facial image P. In this case, a configuration may be adopted in which generated facial images P with a plurality of types of facial expressions are generated in advance in step S115, and a generated facial image P to be used for replacing a facial image is selected according to the facial expression of the person to be replaced. Needless to say, the processing unit 11 may be configured to convert the direction and expression of the generated face image P in accordance with both the direction and expression of the face of each person to be replaced. (2) By inputting information on the image features (landmarks, key points, etc.) of the face in the face image of the person to be replaced and information on the image features (landmarks, key points, etc.) of the face in the generated face image into a face replacement trained model, a face image is regenerated according to the facial orientation or expression, or both the facial orientation and expression, of the person to be replaced. At that time, the position and orientation of the image features (landmarks, key points, etc.) of the face of the person to be replaced are adjusted according to the position and orientation of the image features (landmarks, key points, etc.) of the face of the person to be replaced.

[0045] Then, the processing unit 11 functioning as the facial image replacement unit 11f replaces each of the facial images of the multiple replacement target persons included in the frame image with the generated facial image P generated in steps S115 and S116 (step S117), and outputs the frame image in which the facial image has been replaced with the generated facial image P (step S118). For example, the processing unit 11 displays the frame image in which the facial image has been replaced on the display unit 12. The processing unit 11 may also transmit the frame image in which the facial image has been replaced to the outside via the communication unit 14. When the image processing device 1 is a video distribution server, the processing unit 11 sequentially performs privacy protection image processing on the multiple frame images constituting the video stored in the storage unit 15, and transmits the frame image to the outside.

[0046] In addition, since the face image of the person included in the frame image is identified and tracked by face recognition technology, even if the face image is occluded or framed out by another object, it is possible to estimate which part of the face image is occluded or framed out. For example, if the face image moves to the left and the center position of the face image is near the left edge, shortening the horizontal dimension of the face image, the processing unit 11 can recognize that the left side of the face image is framed out by the shortened dimension. Even if the face image is occluded or framed out, the processing unit 11 can generate a generated face image P based on the attribute information of the person. The processing unit 11 can simply cut out the generated face image P to match the missing part of the face image before replacement due to the occlusion or frame-out, and replace it with the face image.

[0047] After replacing the face image of the person to be replaced, a process may be performed in which a certain width of the edge side of the frame (four sides, top, bottom, left and right) is cut off or covered with black or the like.

[0048] FIG. 6 is a schematic diagram showing a moving image before and after image processing related to privacy protection. The upper diagram of FIG. 6 shows one frame image constituting the original moving image data. The frame image includes six people. Of the six people, the person in the center is a person not to be replaced, and the other people are people to be replaced. In this case, as shown in the lower diagram of FIG. 6, the facial images of the other people except for the person in the center are replaced with generated facial images P1, P2, P3, and P4 created by taking into account the attributes of the person, such as sex, age, and race. It is possible to generate a frame image in which the facial image of the person to be replaced, which requires privacy protection, is replaced, and the facial images of the people not to be replaced, which do not require privacy protection, are left as the original images.

[0049] We have explained the privacy-protecting image processing in which the facial image of one frame image is replaced with the generated facial image P by the processing of steps S111 to S118. However, by repeatedly executing the above processing and performing similar image processing on each frame image that constitutes a video, the facial image of the person to be replaced appearing in the video can always be replaced with the generated facial image P.

[0050] Fig. 7 is a conceptual diagram showing a method for tracking a person to be replaced in a video and replacing a face image. The people included in the video are continuously identified and tracked, and each person is associated with a generated face image P, so that even if the position of the person to be replaced in the video changes as shown in Fig. 7, the face image of each person is continuously replaced with the same generated face image P. In other words, it is possible to prevent the occurrence of an unnatural state in which the same person is replaced with a different generated face image P while moving.

[0051] The upper and lower diagrams of Fig. 7 show different frame images, but since the person included in each frame image is identified and tracked, it can be seen that they are continually replaced with the same generated face image P. Specifically, the person who is replaced with generated face image P2 in the upper diagram of Fig. 7 moves from the upper right to the lower left as shown in the lower diagram, but the face image of that person is replaced by the same generated face image P2.

[0052] According to the image processing device 1 of the first embodiment configured as described above, the facial image of a person included in a frame image of a video can be replaced with a generated facial image P according to the attributes of the person, thereby protecting the privacy of the person while maintaining the scene of the live-action video.

[0053] Also, a moving image can be generated by replacing the face image of a person to be replaced, whose privacy needs to be protected, with the generated face image P, while leaving the non-subject person, whose privacy needs not be protected, as is.

[0054] Furthermore, by continuously identifying and tracking people in a video and replacing them with the same generated face image P, unnatural situations such as faces changing midway through the video can be avoided, and natural videos with privacy protection can be obtained. For example, even if the scene changes within a video, the same person can be identified, so the facial image of the same person is always replaced with the same generated facial image P. Therefore, a natural video with privacy protection can be obtained.

[0055] Furthermore, since the facial image of a person is replaced with a generated facial image P that corresponds to the person's attributes such as age, gender, race, etc., the privacy of the person can be protected without changing the atmosphere (age group, gender, race) of the person included in the original video and the person included in the video after image processing, that is, while maintaining the scene of the live-action video.

[0056] Furthermore, since the facial image of a person is replaced with a generated facial image P according to environmental attributes that affect the facial image of the person, the privacy of the person can be protected without changing the atmosphere (lighting conditions) of the original video and the atmosphere (lighting conditions) of the video after image processing.

[0057] Furthermore, since the configuration involves generating and replacing realistic facial images using the facial image generation trained model 15c, the privacy of people included in the video can be protected while maintaining the scenery of the live-action video.

[0058] Furthermore, by replacing the facial image P with a generated face image P having the same facial direction or expression as the facial direction or expression of a person included in the video, the privacy of the person can be protected without changing the facial direction or expression of the person. Using generative AI technology, it is possible to generate just a few generated face images P for one person, allowing for smooth replacement even if the direction or expression of the face shown in the video changes.

[0059] Generative AI technology can seamlessly replace only the visible parts even if the face is occluded or out of frame. Person tracking technology, which is an elemental technology of person identification technology, can predict the position of the face even if it is occluded or out of frame, so replacement can be performed with high accuracy.

[0060] In the first embodiment, an example has been described in which a face image of a person included in a moving image is replaced with a generated face image P to realize privacy protection, but the technology according to the present embodiment can also be applied to a still image. The same is true for the other embodiments.

[0061] Also, while an example has been described in which processing unit 11 generates a facial image based on the attributes of a person when processing frame images of a video, generated facial images P may be created in advance for typical combinations of attributes and stored in storage unit 15. Processing unit 11 may be configured to replace the facial image of the person with generated facial image P stored in storage unit 15 when there is a generated facial image P that corresponds to the attributes of a person included in a frame image, and to generate and replace a facial image using facial image generation trained model 15c when there is no generated facial image P that corresponds to the attribute.

[0062] (Embodiment 2) The image processing device 1 according to the second embodiment differs from the first embodiment in that, when replacement of a facial image fails, the image processing device 1 protects privacy by performing a conventional masking process on the facial image. Since other configurations of the image processing device 1 are similar to those of the image processing device 1 according to the first embodiment, similar parts are denoted by the same reference numerals and detailed description will be omitted.

[0063] 8 is a flowchart showing the processing procedure of an image processing method according to embodiment 2. The processing unit 11 of the image processing device 1 executes the same processes as steps S111 to S117 in embodiment 1 (steps S211 to S217).

[0064] However, when there is a face image that is missing due to heavy occlusion or framing out, a sudden change in face direction, a low-quality face image due to distance, a low-quality face image due to quick movement, or a face image with low illumination, the face image of the person may be replaced unnaturally. Therefore, the replaced image is verified by the face recognition trained model 15b, and if the face image cannot be detected, it is automatically determined that the replacement has failed. In that case, the original face image is replaced by mosaic processing of the conventional technology to protect privacy. Specifically, the image processing device 1 executes the following process.

[0065] The processing unit 11 performs face recognition processing using the face recognition trained model 15b on the frame image after image processing in which the face image has been replaced, and verifies whether the replaced face image of the person to be replaced is recognized as a face image (step S218), and determines whether there are any face images for which face recognition has failed (step S219).

[0066] When it is determined that there is a face image for which face recognition has failed (step S219: YES), the processing unit 11 returns the face image of the person for which recognition has failed to the original face image, and executes masking processing such as mosaic processing, blurring processing, and superimposing a predetermined image (step S220). Note that the processing unit 11 may be configured to execute masking processing without returning the face image to the original face image.

[0067] Next, if it is determined that there is no replaced face image that has not been recognized as a face image (step S219: NO), or if the processing of step S220 has been completed, the processing unit 11 outputs a frame image in which the face image has been replaced with the generated face image P or has been masked (step S221).

[0068] Fig. 9 is a conceptual diagram showing an image processing method for privacy protection using a combination of replacement processing with generated face image P and masking processing. The upper diagram in Fig. 9 shows a frame image in which a face image has been replaced with generated face image P, but the person on the left edge is out of frame, and the replacement processing of the face image has failed. In this case, as shown in the lower diagram in Fig. 9, the face image of the person whose replacement processing has failed is subjected to masking processing such as mosaic processing and blurring. The face images of the other people to be replaced are replaced with generated face images P3 and P4, and privacy is protected.

[0069] According to the image processing device 1 according to the second embodiment configured as described above, when a person's face image is unnaturally replaced due to various circumstances, privacy protection is realized by performing a conventional masking process on the face image that has failed to be replaced. Therefore, even if the face image is severely occluded or framed out, privacy can be protected while maintaining the view as much as possible.

[0070] (Embodiment 3) The image processing device 1 according to the third embodiment differs from the first embodiment in that, when there is a possibility that the replacement of the facial image will fail, the image processing device 1 protects privacy by a conventional masking process without replacing the facial image. Since the other configurations of the image processing device 1 are similar to those of the image processing device 1 according to the first embodiment, the same reference numerals are used for the similar parts and detailed description will be omitted.

[0071] 10 is a flowchart showing the processing procedure of an image processing method according to embodiment 3. The processing unit 11 of the image processing device 1 executes the same processes as steps S111 to S114 in embodiment 1 (steps S311 to S314).

[0072] Next, the processing unit 11 determines whether or not there is a possibility of failure in the process of replacing the face image of the person detected in step S312 with the generated face image P (step S315). For example, the processing unit 11 determines that there is a possibility of failure in the replacement process when the face image is missing due to severe occlusion or framing out. The processing unit 11 also determines that there is a possibility of failure in the replacement process when there is a sudden change in the face direction. Furthermore, the processing unit 11 determines that there is a possibility of failure in the replacement process when the face image is of low image quality below a predetermined image quality due to distance, or when the face image is of low image quality due to quick movement. Furthermore, the processing unit 11 determines that there is a possibility of failure in the replacement process when there is a face image with low illuminance.

[0073] If it is determined that there is a possibility of failure (step S315: YES), a masking process such as mosaic processing, blurring processing, or superimposing a predetermined image is executed on the face image that may fail to be replaced (step S316).If it is determined that there is no possibility of failure (step S315: NO), as for the face images of the other persons, face images are generated based on the attribute information of one or more persons to be replaced included in the frame image, as in the first embodiment (step S317).

[0074] Thereafter, the same processes as steps S116 to S118 in the first embodiment are executed to replace the face image with the generated face image P, and a privacy-protected frame image is output (steps S318 to S320).

[0075] According to the image processing device 1 according to the third embodiment configured as described above, when there is a possibility that replacement of a person's face image will fail due to various circumstances, privacy protection is realized by executing a conventional masking process on the face image that may have failed to be replaced. Therefore, even if severe occlusion or framing out occurs, privacy can be protected while maintaining the view as much as possible.

[0076] (Embodiment 4) The image processing device 1 according to the fourth embodiment differs from the first embodiment in that when an image contains many people, the number of generated face images P to be created is limited, and the same generated face image P is used to replace face images of multiple people with the same attributes. Since the other configurations of the image processing device 1 are similar to those of the image processing device 1 according to the first embodiment, similar parts are given the same reference numerals and detailed description will be omitted.

[0077] 11 is a flowchart showing the processing procedure of the image processing method according to the fourth embodiment. The processing unit 11 of the image processing device 1 executes the same processes as steps S111 to S114 in the first embodiment (steps S411 to S414). Next, the processing unit 11 generates face images based on attribute information of a plurality of persons to be replaced included in the frame image, within the range of the upper limit number of generated face images P set for each attribute (step S415).

[0078] The upper limit number of generated images is, for example, 10 for male generated face images P and 10 for female generated face images P. Note that 10 is just an example, and the upper limit number may be appropriately determined depending on image processing capabilities. Also, an upper limit on the number of generated face images P may be set for each age of person. For example, the upper limit on the number of generated face images P for men in their twenties may be 5, for men in their thirties 5, for men in their forties 5, and for men over 50 5. The upper limit on the number of images for women may be set in a similar manner. Of course, an upper limit on the number of generated face images P may be set for each combination of sex, age, and race. Furthermore, regardless of the attributes of the person, an upper limit may be set on the total number of generated face images P.

[0079] Next, the processing unit 11 assigns the generated generated face image P to the replacement target person for whom the upper limit of the generation upper limit number has been reached and no face image has been generated (step S416). In other words, the same generated face image P is assigned to multiple people in a redundant manner.

[0080] Thereafter, the same processes as steps S116 to S118 in the first embodiment are executed to replace the face image with the generated face image P, and a privacy-protected frame image is output (steps S417 to S419).

[0081] FIG. 12 is a schematic diagram showing frame images of a moving image in which face images are replaced by the image processing method according to the fourth embodiment, and FIG. 13 is a conceptual diagram showing a tracking table when a generated face image P is used in a duplicated manner. The frame images shown in FIG. 12 include 14 people to be replaced. All of them are women, and the number of images exceeds the upper limit of 10 to be generated. For this reason, as shown in FIG. 13, generated face images P1 to P10 are assigned individually to person numbers 1 to 10, respectively, but generated face images P7 to P10 that have already been created are assigned in duplicate to people to be replaced with person numbers 11 to 14. In this way, the generated face images P generated within the range of 10 images, which is the upper limit of the number of images to be generated, are used in duplicate to replace the face images of the 14 people to be replaced, as shown in FIG. 12. In FIG. 12, only the generated face images P7 to P10 that have been replaced in a duplicated manner are assigned with reference numbers, and the reference numbers of the other generated face images P are omitted.

[0082] According to the image processing device 1 according to the fourth embodiment configured as described above, it is possible to protect privacy by replacing many people included in a moving image with generated face images P while suppressing the load of image processing. Therefore, even if a moving image includes many people, privacy can be protected while maintaining the scene as much as possible.

[0083] Means for solving the problems of the present disclosure are described below. (Appendix 1) Acquire a real-life image that includes a person, Identifying attributes of the person included in the acquired real-life image; The facial image of the person is replaced with a realistic facial image generated according to the identified attributes. A computer program that causes a computer to carry out processing. The actual images include moving images and still images. (Appendix 2) The live-action image includes a still image 2. The computer program of claim 1. (Appendix 3) The live-action image is a video, Tracking the person in the video; Continually replacing the facial image of the person being tracked with the generated facial image. 3. A computer program product according to claim 1 or 2, which causes a computer to carry out a process. When tracking the person included in the video, it is preferable to track the person by playing the video forward and to track the person by playing the video backward. (Appendix 4) The attribute is: At least one of the person's age, sex, and race 4. A computer program according to any one of claims 1 to 3. (Appendix 5) The attribute is: The facial features of the person, the environmental attributes of the person, or the human body features of the person. 5. A computer program according to any one of claims 1 to 4. The facial feature amount of a person includes image feature amounts of the face of the person, such as landmarks and key points, and also includes whether the person is wearing glasses or a mask. The environmental attributes for the person include lighting conditions and the like. The feature amounts of the person's body include a bounding box (center coordinate position and vertical and horizontal sizes) surrounding the image of the person's body, image feature amounts of the person's body, and the like. (Appendix 6) By inputting information including the specified attributes to a trained model that outputs a realistic generated face image according to the attributes when information including the attributes is input, the realistic generated face image according to the attributes is generated. 6. A computer program product according to any one of claims 1 to 5, configured to cause a computer to carry out a process. The trained models include a model that outputs a realistic facial image generated according to attributes when attributes are directly input, and a model that outputs a realistic facial image generated according to attributes when a prompt including attributes and describing the characteristics of an image to be generated in text is input. (Appendix 7) The real-life image includes a plurality of people, The face images of the plurality of persons are replaced with the generated face images generated according to the attributes specified for each of the plurality of persons. 7. A computer program product according to any one of claims 1 to 6, which causes a computer to carry out a process. (Appendix 8) The real-life image includes a plurality of people, replacing face images of the plurality of persons, excluding a specific target person, with the generated face image; The facial image of the target person is not replaced. 8. A computer program product according to any one of claims 1 to 7, which causes a computer to carry out a process. (Appendix 9) The face image is replaced with the generated face image according to the face direction or expression of the person. 9. A computer program product according to any one of claims 1 to 8, which causes a computer to carry out a process. The face image is replaced with the generated face image according to the face direction and expression of the person. 9. A computer program product according to any one of claims 1 to 8, which causes a computer to carry out a process. There are two main methods for generating the generated face image according to the facial orientation or expression of the person. The first method is to generate in advance generated facial images corresponding to multiple orientations or facial expressions based on the attributes of the identified person, and to select and replace the generated facial image corresponding to the facial orientation or facial expression of the person contained in the real-life image to be processed. The second method is a method of correcting or regenerating the feature amount of a generated face image facing a predetermined direction according to the feature amount of the person included in the real-life image to be processed, and replacing it with the corrected or regenerated generated face image. More specifically, based on the information on the image feature amount (landmark, key point, etc.) of the face of the person to be replaced and the information on the image feature amount (landmark, key point, etc.) of the face in the generated face image, the position and orientation of the image feature amount (landmark, key point, etc.) of the face of the person to be replaced is adjusted in accordance with the position and orientation of the image feature amount (landmark, key point, etc.) of the person to be replaced, thereby regenerating a generated face image according to the facial orientation or expression of the person to be replaced. (Appendix 10) determining whether the generated face image after the replacement can be detected as a face image by a face recognition trained model; If the generated face image after replacement is not detected as a face image, a masking process is performed on the face image of the person. 10. A computer program product according to any one of claims 1 to 9, which causes a computer to carry out a process. (Appendix 11) determining whether or not there is a possibility of failure in the process of replacing the face image of the person with the generated face image; If there is a possibility of failure, a masking process is performed on the face image of the person. 11. A computer program product according to any one of claims 1 to 10, configured to cause a computer to carry out a process. (Appendix 12) The real-life image includes a plurality of people, The generated face images are generated within a range of an upper limit number of the generated face images set for each attribute, and the facial images of some of the persons are replaced with the same generated face images. 12. A computer program product according to any one of claims 1 to 11, which causes a computer to carry out a process. (Appendix 13) After replacing the face image of the person, Cutting or masking the edges of the actual image 13. A computer program product according to any one of claims 1 to 12, configured to cause a computer to carry out a process. Specifically, the above process involves cutting out a certain width from the edges of the actual image, for example the edges of a video frame, on all four sides (top, bottom, left, right, etc.), or covering it with black or the like. (Appendix 14) Acquire a real-life image that includes a person, Identifying attributes of the person included in the acquired real-life image; The facial image of the person is replaced with a realistic facial image generated according to the identified attributes. Image processing methods. (Appendix 15) An image processing device including a processing unit that acquires a real-life image including a person and executes image processing, The processing unit includes: Identifying attributes of the person included in the acquired real-life image; The facial image of the person is replaced with a realistic facial image generated according to the identified attributes. Image processing device. [Explanation of symbols]

[0084] 1: Image processing device 1a: Recording medium 11: Processing section 11a: Human body / face detection unit 11b: Attribute extraction part 11c: Person Identification Section 11d: Replacement target person selection section 11e: Face image generation unit 11f: Face image replacement section 12: Display section 13:Operation section 14: Communications Department 15: Storage section 15a: Computer Programs 15b: Face recognition trained model 15c: Face image generation trained model 15d: Face image replacement trained model P: Generated face image

Claims

1. Acquire a real-life image that includes a person, Identifying attributes of the person included in the acquired real-life image; The facial image of the person is replaced with a realistic facial image generated according to the identified attributes. A computer program that causes a computer to carry out processing.

2. The live-action image includes a still image 2. The computer program product of claim 1.

3. The live-action image includes a video; Tracking the person in the video; Continually replacing the facial image of the person being tracked with the generated facial image. The computer program product of claim 1 , which causes a process to be executed by the computer.

4. The attribute is: At least one of the person's age, sex, and race A computer program according to any one of claims 1 to 3.

5. The attribute is: The facial features of the person, the environmental attributes of the person, or the human body features of the person. A computer program according to any one of claims 1 to 3.

6. By inputting information including the specified attribute to a trained model that outputs a realistic generated face image according to the attribute when information including the attribute is input, the realistic generated face image according to the attribute is generated. A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

7. The real-life image includes a plurality of people, The face images of the plurality of persons are replaced with the generated face images generated according to the attributes specified for each of the plurality of persons. A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

8. The real-life image includes a plurality of people, replacing face images of the plurality of persons, excluding a specific target person, with the generated face image; The facial image of the target person is not replaced. A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

9. The face image is replaced with the generated face image according to the face direction or expression of the person. A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

10. determining whether the generated face image after the replacement can be detected as a face image by a face recognition trained model; If the generated face image after replacement is not detected as a face image, a masking process is performed on the face image of the person. A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

11. determining whether or not there is a possibility of failure in the process of replacing the face image of the person with the generated face image; If there is a possibility of failure, a masking process is performed on the face image of the person. A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

12. The real-life image includes a plurality of people, The generated face images are generated within a range of an upper limit number of the generated face images set for each attribute, and the facial images of some of the persons are replaced with the same generated face images. A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

13. After replacing the face image of the person, Cutting or masking the edges of the actual image A computer program product according to any one of claims 1 to 3, which causes a computer to execute a process.

14. Acquire a real-life image that includes a person, Identifying attributes of the person included in the acquired real-life image; The facial image of the person is replaced with a realistic facial image generated according to the identified attributes. Image processing methods.

15. An image processing device including a processing unit that acquires a real-life image including a person and executes image processing, The processing unit includes: Identifying attributes of the person included in the acquired real-life image; The facial image of the person is replaced with a realistic facial image generated according to the identified attributes. Image processing device.

Citation Information

Patent Citations

  • Face recognition device, face recognition method and computer program

    JP2008257425A

  • Image processing apparatus, image processing system, image processing method, and computer program

    JP2014067131A

  • Information processing device and program

    JP2014085796A

  • Information processing device, information processing system, information processing method, program and storage medium

    JP2020091770A

  • Image processing apparatus

    JP2020086921A