Image processing apparatus, image processing method, and computer program
The image processing device uses AI to detect and attribute persons, superimposing illustrations to protect privacy while allowing attribute recognition, addressing the limitations of existing methods.
Patent Information
- Application Number
- JP2024073169
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2025-11-07
AI Technical Summary
Existing image processing methods for privacy protection, such as blurring or mosaic, may not adequately conceal individuals, especially acquaintances, and methods that replace persons can be misleading or limited in applications with multiple images. Additionally, these methods do not allow for distinguishing individual features.
An image processing device that detects persons, determines their attributes, and superimposes illustrations based on these attributes onto the detected areas, using AI to generate images that protect privacy while allowing attribute recognition.
The device generates images that intuitively convey person attributes while ensuring privacy, adaptable to various applications and scenarios, including switching between illustration display modes.
Smart Images

Figure 2025168051000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, a computer program, and the like. [Background technology]
[0002] When using video footage captured by security cameras, etc., techniques for processing images to protect the privacy of people in the video are widely used. For example, methods are known for processing the image to make it difficult to identify individuals by blurring or mosaic the entire video or the area of a person in the video.
[0003] Patent Document 1 discloses a processing technique for masking the area of a person detected from a video with a mask image according to the attributes of the person. Patent Document 2 discloses a technique for replacing a person in an original video with an image of another person having the same attributes as the person in the video. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 5834193 [Patent Document 2] Patent Publication No. 2020-91770 [Non-patent literature]
[0005] [Non-Patent Document 1] E. Mansimov et al. “Generating Images from Captions with Attention”, ICLR 2016 [Non-patent document 2] O. Vinyals et al. “Show and Tell: A Neural Image Caption Generator”, CVPR 2015 [Non-patent document 3] P. Isola et al. “Image-to-Image Translation with Conditional Adversarial Networks”, CVPR 2017 [Non-patent document 4] J. Redmon, A. Farhadi, “YOLO9000: Better Faster Stronger”, CVPR 2016 Summary of the Invention [Problem to be solved by the invention]
[0006] Blurring and mosaic processing may not adequately protect privacy if the processing is done to a level where the presence of a person can be discerned, for example, if the person is an acquaintance, the individual may be identifiable even in the processed image. Masking of person areas may not be suitable for applications that require individual distinction, as the person's features cannot be directly read from the processed image.
[0007] Furthermore, the method of replacing an image with an image of another person has the problem that the person used for replacement may be mistaken for a real person, and that the method is limited to applications where there is no problem even if multiple images of the same person are placed on the screen.
[0008] An object of the present invention is to provide an image processing device that can generate an image that makes it easy to grasp a person's attributes while protecting the person's privacy. [Means for solving the problem]
[0009] An image processing device according to one aspect of the present invention includes: An image acquisition means; person detection means for detecting a person from the image acquired by the image acquisition means; attribute determination means for determining attributes of the person detected by the person detection means; an illustration acquisition means for acquiring an illustration of the person by inputting a prompt based on the attribute determined by the attribute determination means; an image synthesis means for superimposing the illustration acquired by the illustration acquisition means on the area of the person; The present invention is characterized by having the following. [Effects of the Invention]
[0010] According to the present invention, it is possible to realize an image processing device that can generate an image that makes it easy to grasp a person's attributes while protecting the person's privacy. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a block diagram showing an example of the configuration of an image processing device 100 according to a first embodiment. [Figure 2] 1 is a functional block diagram showing an example of the functional configuration of an image processing device 100 according to a first embodiment. [Figure 3] 10 is a flowchart illustrating an example of a video processing process using a generation AI unit 270 according to the first embodiment. [Figure 4] 10(A) to 10(D) are diagrams illustrating examples of image processing according to the first embodiment. [Figure 5] 10A and 10B are flowcharts illustrating an example of a video processing process using a generation AI 500 according to the first embodiment. [Figure 6] 6A and 6B are diagrams showing an example of an illustration generation additional information screen 610 for inputting text information to be added to a prompt, which is displayed on a display device 600 in the second embodiment. [Figure 7] 10A and 10B are diagrams illustrating examples of results of image processing according to the second embodiment. [Figure 8] 8A and 8B are diagrams showing an example of a person of interest display setting screen 810 for setting the display of a person of interest, which is displayed on the display device 600 in the second embodiment. [Figure 9] 10A to 10C are diagrams illustrating an example of image processing according to the second embodiment. [Figure 10] 10A to 10C are flowcharts showing an example of processing performed by the image processing device 100 for switching to an illustration display according to the third embodiment. [Figure 11] 10(A) to 10(E) are diagrams showing examples of operation screens for setting a person illustration display according to the third embodiment. [Figure 12] 10(A) to 10(H) are diagrams illustrating examples of representations of people according to the fourth embodiment. [Figure 13] FIG. 13 is a diagram showing an example of an operation screen for setting an abstraction level of a person according to the fourth embodiment. [Figure 14] 10A and 10B are diagrams showing examples of an operation screen for performing person abstraction setting according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not limited to the following embodiments. In each drawing, the same members or elements are designated by the same reference numerals, and duplicate descriptions will be omitted or simplified.
[0013] [Embodiment 1] 1 is a block diagram showing an example of the configuration of an image processing device 100 according to embodiment 1. The image processing device 100 includes a CPU 101, a memory 102, a storage unit 103, a communication I / F unit 104, a display control unit 105, an operation control unit 106, and the like, which are communicatively connected via a system bus 107.
[0014] A CPU (Central Processing Unit) 101 functions as a computer that controls the entire image processing apparatus 100. The CPU 101 controls the operation of each functional unit connected via a system bus 107, for example.
[0015] The memory 102 stores data, computer programs, etc. that are used for processing by the CPU 101. The memory 102 also functions as a main memory, work area, etc. for the CPU 101. The CPU 101 executes processing based on the computer programs stored in the memory 102, thereby realizing a series of processes in this embodiment.
[0016] The storage unit 103 stores an OS (operating system), computer programs for the CPU 101 to execute processes, and various data such as data used in analysis and analysis result data.
[0017] The communication I / F unit 104 is an interface that connects the image processing device 100 to the network 400. The display control unit 105 controls the display of screen display data from the image processing device 100 on the display device 600.
[0018] The operation control unit 106 performs control to input user instructions input via the operation device 700 to the image processing device 100. The imaging device 300 is, for example, a network camera, and transmits captured video and associated data to the image processing device 100 via the network 400.
[0019] The generation AI 500 has an internal CPU as a computer and a memory as a storage medium. The generation AI 500 generates an image using AI (Artificial Intelligence) based on text information received from the image processing device 100 via the network 400, and transmits the generated image to the image processing device 100. The technology for generating an image from input text information can be realized by applying the technology described in Non-Patent Document 1, for example.
[0020] The generation AI 500 also has a function of generating text information representing attributes of people and the like appearing in an image received from the image processing device 100 via the network 400 using AI, based on the image, and transmitting the text information to the image processing device 100. The technology for generating text information from an input image can be realized by applying the technology described in Non-Patent Document 2, for example.
[0021] Furthermore, the generation AI 500 also has a function of generating an image by using AI based on an image received from the image processing device 100 via the network 400, from which people appearing in the image have been erased, and transmitting the generated image to the image processing device 100. The technology for generating an image by converting an input image can be realized by applying the technology described in Non-Patent Document 3, for example.
[0022] The display device 600 has a display member such as a liquid crystal display, and displays images captured by the imaging device 300, images processed by the image processing device 100, etc. via a display control unit 105. The operation device 700 has operation members such as a mouse, a touch panel, buttons, etc., and inputs user operations to the image processing device 100 via an operation control unit 106.
[0023] Fig. 2 is a functional block diagram showing an example of the functional configuration of the image processing device 100 according to the first embodiment. Note that some of the functional blocks shown in Fig. 2 are realized by causing a CPU or the like serving as a computer included in the image processing device to execute a computer program stored in a memory serving as a storage medium.
[0024] However, some or all of these functions may be implemented by hardware. Examples of hardware that can be used include dedicated circuits (ASICs) and processors (reconfigurable processors, DSPs). Furthermore, the functional blocks shown in Figure 2 do not have to be built into the same housing, and may be configured as separate devices connected to each other via signal paths.
[0025] The image processing device 100 has a video acquisition unit 210, a person detection unit 220, a person area determination unit 230, a person attribute determination unit 240, a person area processing unit 250, an image synthesis unit 260, a generation AI unit 270, an AI input / output unit 280, and a video output unit 290.
[0026] The video acquisition unit 210 functions as a video acquisition means for acquiring video. In this embodiment, the video to be processed for privacy protection is acquired from the imaging device 300 via the communication I / F unit 104. Hereinafter, the video acquired by the video acquisition unit 210 will be referred to as "input video." Furthermore, when describing various processes for each frame of the input video, the processing target will be referred to as an "image."
[0027] The person detection unit 220 functions as a person detection unit that detects people from the video acquired by the video acquisition unit, and outputs the detection result using a machine learning model. The machine learning model has been trained in advance so that it can detect people, etc., included in the image.
[0028] The detection of people and the like can be realized by applying, for example, the technology described in the following Non-Patent Document 4. The person detection process in the person detection unit 220 is not limited to the technology disclosed in Non-Patent Document 4, and various other technologies can be applied as long as they are capable of detecting people and the like from an image.
[0029] The person area determination unit 230 determines an area to be processed in order to protect the privacy of a person within the image based on the output of the person detection unit 220. Here, the person area determination unit 230 functions as a person area determination means that determines a person area on the video that is occupied by a person detected by the person detection unit 220 as a person detection means.
[0030] For example, the position information of the person detected by the person detection unit 220 is used to obtain the outline of the person through conventional image processing such as binarization and labeling, and the interior of the outline is determined as the person area. Alternatively, a technology for recognizing the person area by semantic segmentation using deep learning may be applied.
[0031] The person attribute determination unit 240 estimates person attribute information for each person detected by the person detection unit 220. The person attribute information here may be information about age, gender, clothing, emotions, and the like.
[0032] It may also be information on whether or not a specific item such as a mask is worn. Person attribute determination involves calculating scores related to age and gender based on facial features, correcting the scores based on predetermined conditions, and then selecting and outputting the most probable age and gender as the attribute estimation result for the person.
[0033] Here, the person attribute determining section 240 functions as an attribute determining means for determining the attributes of the person detected by the person detecting section 220 as a person detecting means.
[0034] The person area processing unit 250 processes the area to be processed determined by the person area determination unit 230 so that the individual cannot be identified. Processing methods include, for example, uniformly filling in the area or performing mosaic processing. Alternatively, an image may be generated that appears as if no person were captured by interpolating using information about the area surrounding the area to be processed.
[0035] The image synthesis unit 260 functions as an image synthesis means, and synthesizes the image processed by the person area processing unit 250 with the output image of the generation AI unit 270, which will be described later.
[0036] The generation AI unit 270 functions as an illustration acquisition means, and generates and outputs an image based on input text information, etc. The generation AI unit 270 receives input of text information, etc., which has been converted into a predetermined prompt by the AI input / output unit 280 based on the person attribute information output by the person attribute determination unit 240 and additional information such as text information input by the user via the operation device 700.
[0037] The generation AI unit 270 does not necessarily have to be included in the image processing device 100, and similar processing may be performed by any generation AI 500 connected to the network 400. Furthermore, both the generation AI unit 270 and the generation AI 500 may be configured to generate images from text information or generate text information from images.
[0038] The AI input / output unit 280 inputs text information to the generation AI unit 270 and inputs the image generated and output by the generation AI unit 270. The AI input / output unit 280 filters the output information from the person attribute determination unit 240 according to predetermined conditions, and creates text information to be input to the generation AI unit 270 by combining it with text information input by the user from the operation device 700.
[0039] Furthermore, the AI input / output unit 280 performs data conversion necessary for input / output to the generated AI 500 connected to the image processing device 100 via the network 400. The video output unit 290 functions as a video output means, and outputs video processed by the image processing device 100. The video output here is displayed on the display device 600 via the display control unit 105.
[0040] Next, the processing performed by the image processing device 100 will be described with reference to Fig. 3 and Fig. 4. Fig. 3 is a flowchart illustrating an example of video processing using the generation AI unit 270 according to embodiment 1. Note that the operation of each step in the flowchart in Fig. 3 is performed sequentially by a CPU or the like serving as a computer in the image processing device 100 executing a computer program stored in memory.
[0041] In step S301, the video acquisition unit 210 functions as a video acquisition step for acquiring a video to be processed. Figures 4(A) to 4(D) are diagrams for explaining an example of image processing according to the first embodiment, and Figure 4(A) shows a frame of the video acquired in step S301. From this point onwards, processing is repeated for each frame of the acquired video, but for simplicity, processing for one frame image in the video will be explained.
[0042] In step S302, the person detection unit 220 detects people in the image. In the case of the image in Fig. 4(A), three people are detected. Here, step S302 functions as a person detection step for detecting people from the video acquired in the video acquisition step.
[0043] In step S303, the person attribute determination unit 240 determines the attributes of each person in the image. Here, step S303 functions as an attribute determination step for determining the attributes of the person detected in the person detection step.
[0044] The person attribute determination unit 240 can determine multiple types of attributes, and for example, in the case of the person in the center of Fig. 4(A), it outputs person attribute determination results such as "gender" = "male," "age" = "30s," "mask presence / absence" = "mask present," and "clothing" = "casual." Note that in step S303, the person attribute determination unit 240 may also identify the person's name, ID, etc. as the person's attributes by image recognition.
[0045] The system may be configured to always execute all types of attribute determination processing and output the results, or it may be configured to narrow down the attributes to be determined according to predetermined conditions. For example, if only age needs to be determined depending on the application, other attribute determination processing may be omitted.
[0046] In step S304, the generation AI unit 270 generates and outputs an illustration of a person from the text information based on the person attribute determination result in step S303. The prompt to be input to the generation AI unit 270 is created by applying the text information of the person attribute determination result to a standard phrase that is stored in advance by the AI input / output unit 280 and matches the person attributes to be used.
[0047] For example, if the person attribute determination results are "male," "in his 30s," "wearing a mask," and "dressed casually," the AI input / output unit 280 will create a prompt such as, "Please draw a full-body illustration of a man in his 30s dressed casually and wearing a mask."
[0048] When a similarly created prompt is input for each person in the image, the generation AI unit 270 generates and outputs an illustration of the person according to the instruction. Figure 4(B) shows an example of an illustration corresponding to each of the three people in Figure 4(A).
[0049] Here, step S304 functions as an illustration acquisition step (illustration acquisition means) that acquires an illustration of a person by inputting a prompt based on the attributes determined in the attribute determination step (attribute determination means). The illustration acquisition means acquires an illustration generated by the generation AI based on the prompt.
[0050] In step S303, the person attribute determination unit 240 identifies a person's attribute such as a name or ID, and if a corresponding prompt is input, an illustration such as an avatar registered in advance for each person may be acquired in step S304. In addition, the person himself may be able to register an illustration such as an avatar.
[0051] Meanwhile, in step S305, the person area determination unit 230 determines and determines a person area that covers the entire person detected by the person detection unit 220. The dashed line in Fig. 4(C) shows an example of the determined person area.
[0052] In step S306, the person area determined in step S305 is processed by person area processing unit 250 to fill in the person area uniformly. Note that this processing may be other than filling in the person area uniformly, and may be any processing that makes the features of the person unrecognizable in order to protect the privacy of the person in the image, for example.
[0053] Here, step S306 functions as a person area processing step (person area processing means) that processes the area determined in step S305 as a person area determination step (person area determination means) to erase the person from the video.
[0054] Note that, instead of the processing by person area processing unit 250, for example, generation AI unit 270 or generation AI 500 may be configured to perform processing to turn the person area into background. That is, a background image other than the person may be prepared in advance, and the background image corresponding to the person area determined in step S305 may be used in the image synthesis processing in the next step S307 instead of the processing in step S306.
[0055] That is, the process of erasing the person in step S306 can be any of the following: a process of filling the person area uniformly or with another image, a process of interpolating the person area using information from adjacent areas of the person area, or a process of turning the person area into a background using video information.
[0056] In step S307, the image synthesis unit 260 synthesizes the illustration of the person (FIG. 4(B)) generated in step S304 with the image in which the person area has been filled in, for example, in step S306.
[0057] Here, a process is performed in which an illustration corresponding to each person in the image is overlaid on the original person area by adjusting the size. Figure 4(D) is a diagram showing an example of the result of image synthesis. Here, step S307 functions as an image synthesis step (image synthesis means) that overlays the illustration acquired in the illustration acquisition step (illustration acquisition means) on the person area.
[0058] In step S308, the video output unit 290 outputs and displays the image that has been subjected to the image synthesis process in step S306 to the display device 600 via the display control unit 105. As a result, even if the video is from a network camera installed in a public place, it is possible to display a video with privacy protection by replacing people with illustrations, as shown in Fig. 4(D).
[0059] Here, step S308 functions as a video output step (video output means) that outputs the video synthesized in step S307 as an image synthesis step (image synthesis means).
[0060] 5A and 5B are flowcharts illustrating an example of video processing using the generation AI 500 according to embodiment 1. The CPUs of the image processing device 100 and the generation AI 500 execute programs stored in memory, thereby sequentially performing the operations of the steps in FIGS.
[0061] FIG. 5A is a flowchart showing an example of processing when the generation AI 500 performs image processing, and steps S301 to S303 are the same as the processing described with reference to FIG.
[0062] In step S501, based on the person attribute determination result output in step S303, the AI input / output unit 280 creates a prompt, which is text information to be sent to the generated AI 500. The prompt created here is created by storing a predefined phrase that matches the person attributes to be used, and applying the text information of the person attribute determination result to the predefined phrase.
[0063] For example, if the person attribute determination results are "male" and "age 30s," the AI input / output unit 280 creates a prompt such as, "After deleting the person, please draw an illustration of a man in his 30s in the same location." The created prompt and the image acquired in step S301 are sent to the generation AI 500 via the network 400.
[0064] In step S502, the generation AI 500 receives the prompt created in step S501 and the image acquired in step S301 as input, and generates an AI image in which the generated illustration is superimposed in place of the person in the original image.
[0065] Then, in step S502, the AI image with the illustration superimposed thereon is transmitted to the image processing device 100. Next, in step S308, as described in FIG. 3, the video output unit 290 of the image processing device 100 displays the AI image in which the person has been converted into an illustration on the display device 600 via the display control unit 105.
[0066] 5(B) is a flowchart showing an example in which the generation AI 500 further performs analysis processing on people in an image, and in step S301, the video acquisition unit 210 acquires an image as described in Fig. 3. Furthermore, in step S503, the AI input / output unit 280 creates a prompt, for example, "Please convert the people included in the input image into illustrations that do not identify individuals," and transmits this to the generation AI 500 together with the image acquired in step S301.
[0067] In this case, if gender and age are to be reflected as person attributes in the illustration, the prompt may be, for example, "Please convert the people in the input image into illustrations that do not allow individuals to be identified without changing their gender or age."
[0068] In step S504, the generation AI 500 inputs the prompt created in step S503 and the image acquired in step S301, and converts the person in the original image into an illustration that cannot be used to identify the individual without changing their gender or age. Then, it generates an AI image with the illustration superimposed. Furthermore, in step S504, it returns the AI image with the illustration superimposed to the image processing device 100.
[0069] After that, in step S308, the image in which the relevant person has been converted into an illustration is displayed on the display device 600 by the video output unit 290 of the image processing device 100 via the display control unit 105, as described with reference to FIG.
[0070] In the first embodiment, a person's features are converted into personal attribute information, a simple format with a very small amount of information, making it impossible to identify the individual. The person in the original image is then replaced with an illustration generated based on the personal attribute information. This converts the image shown in Fig. 4(A) into the image shown in Fig. 4(D), providing an image that makes it easy to intuitively grasp the person's general features while reliably protecting the privacy of the photographed person.
[0071] [Embodiment 2] In the first embodiment, an example of processing was described in which an illustration is generated based on person attribute information and displayed in place of the photographed person in order to reliably protect the privacy of the photographed person. In the second embodiment, a method of adding specific features to the generated person illustration and using it will be described. In the following description, the same reference numerals will be used for components common to the first embodiment, and description thereof will be omitted.
[0072] 3 in the first embodiment, text information based on the person attribute information determined in step S303 was used as the prompt input to the generation AI unit 270. In this embodiment, when creating a prompt, in addition to the person attribute information, information for adding characteristics to the person illustration is used.
[0073] 6(A) and (B) are diagrams showing an example of an illustration generation additional information screen 610 for inputting text information to be added to the prompt, which is displayed on the display device 600 in embodiment 2. Here, the illustration generation additional information screen 610 functions as an operation means for specifying additional information to be input in step S304 as an illustration acquisition step (illustration acquisition means).
[0074] The illustration generation additional information screen 610 includes a preview area 611 that displays a preview image of the illustration generated with the current settings, an additional information input area 612, an apply button 613, a cancel button 614, and an OK button 615.
[0075] The user inputs text information specifying features to be added when generating an illustration using the operation device 700. The text information being input is displayed in the additional information input area 612 in a format that allows the user to recognize that input is in progress.
[0076] 6(A), the fact that "Halloween costume" is being input is indicated by the colored background of the text. When the user finishes inputting the text information to be added and presses the apply button 613, the AI input / output unit 280 creates a prompt to be input to the generation AI unit 270.
[0077] If the person attribute determination result at this time is, for example, "male" and "age 30s," the AI input / output unit 280 creates a prompt such as, "Please draw a full-body illustration of a man in his 30s dressed in a Halloween costume."
[0078] 6(B) shows a screen in which the generated illustration replaces the person in the original image, and is displayed in preview area 611. At this time, the text information entered in additional information input area 612 has already been applied to the illustration generation, so the background color of the text indicating that input is in progress is cleared.
[0079] If the user views the screen displayed in the preview area 611 and decides to accept it, the user presses the OK button 615 to complete the illustration generation addition settings and closes the illustration generation addition information screen 610. The cancel button 614 is used to exit the illustration generation addition information screen 610 without applying the text information being entered.
[0080] 7 is a diagram for explaining an example of the result of image processing according to the second embodiment, showing an example of a processed image when "Halloween costume" is specified as additional information. In this way, by generating and applying an illustration of a person to which specific characteristics have been added in addition to personal attribute information, it is possible to provide a display image that is more suitable for the purpose.
[0081] For example, if a commercial facility uses network camera footage to display a screen showing congestion levels, it can add decorations to the illustrations that match the day's event, thereby increasing the effectiveness of publicizing the event.
[0082] Next, an example will be described in which, when a particular person in an image is focused on, a display is performed to make the person easier to identify on the screen.
[0083] 8A and 8B are diagrams showing an example of a person of interest display setting screen 810 for setting the display of a person of interest, which is displayed on the display device 600 in the second embodiment.
[0084] The notable person display setting screen 810 has a notable person designation area 811, a landmark item setting area 812, a create button 813, a created item display area 814, a cancel button 815, and an OK button 816 arranged thereon.
[0085] The user designates a person to whom he or she wishes to add a marker, for example, by touching the screen or clicking with the mouse, from among the people displayed in the person of interest designation area 811, for reasons such as wanting to keep an eye on their subsequent movements.
[0086] 8(A) and (B), the person who is not designated among the three people shown is displayed in a dimmed light, indicating that the person on the left end is designated. In this way, the person of interest display setting screen 810 functions as a person selection means for selecting one or more people from those detected by the person detection means.
[0087] In this embodiment, in order to add a distinctive feature to the designated person, the designated person is equipped with an item designated in the landmark item setting area 812. In Figures 8(A) and (B), a "bouquet of flowers" is selected as the landmark item.
[0088] In this state, when the user presses the generate button 813, the specified "bouquet" illustration is generated by the generation AI unit 270 or the generation AI 500 and displayed in the generated item display area 814 (FIG. 8(A)).
[0089] When the user presses the generate button 813 again, another "bouquet" illustration is generated and displayed in the generated item display area 814, as shown in FIG. 8(B). When the user presses the OK button 816, the settings on this screen are confirmed and the notable person display setting screen 810 is closed. In this way, in this embodiment, it is possible to specify additional information for the person selected by the person selection means.
[0090] Then, the illustration displayed in the generated item display area 814 is saved in the storage unit 103, and the prompt for generating an illustration for the person designated in the notable person designation area 811 is updated.
[0091] For example, the prompt is updated to "Please draw a full-body illustration of a man in his 30s wearing casual clothing and a mask, holding a bouquet of flowers." The cancel button 815 is used to exit the notable person display setting screen 810 without applying the settings on this screen.
[0092] 9 is a diagram illustrating an example of image processing according to the second embodiment, in which an illustration of a person holding a bouquet of flowers is used as a marker for a person of interest. Note that the marker for a person of interest is not limited to designating an item for the person to wear, and any other format may be used, such as displaying the illustration in a different color from the others, or displaying a shape such as an arrow that matches the person.
[0093] The additional information as a landmark may represent, for example, at least one of the following: what the person in the illustration is wearing, a building, a landscape, a time of day, a season, a country, a region, a related event, etc. The additional information may also be at least one of a specific item the person is wearing, a specific graphic to be displayed in proximity to the person, text information to be displayed in proximity to the person, and a color or image to decorate the person.
[0094] Furthermore, the landmark item setting area 812 may be operated in such a way that the user can freely specify text information, rather than simply selecting from predetermined formats and items. Alternatively, the user may be allowed to additionally register illustrations he or she has drawn himself or herself, or illustrations downloaded from the internet, and then select from a plurality of additionally registered illustrations.
[0095] In this way, by applying a person illustration that adds distinctive features to a specified person, it is possible to provide a display image that is suitable for use in applications such as observing the purchasing behavior of one or more specific people while protecting their privacy.
[0096] [Embodiment 3] In the third embodiment, a method for allowing the user to specify whether to display or hide the generated illustration, and a method for periodically switching between these modes, will be described. In the following description, the same reference numerals will be used for components common to the first and second embodiments, and their description will be omitted.
[0097] 10A to 10C are flowcharts showing an example of processing performed by the image processing device 100 for switching to an illustration display according to embodiment 3. The CPU 101 of the image processing device 100 executes a program stored in the memory 102 and the storage unit 103, thereby sequentially performing the operations of the steps in FIG.
[0098] In step S1010, CPU 101 determines whether or not the user has issued an instruction to display an illustration from operation device 700. If it is determined that an instruction to display an illustration has not been issued, the process proceeds to step S1020, where normal display processing, which will be described later, is executed. If it is determined that an instruction to display an illustration has been issued, the process proceeds to step S1030, where illustration display processing, which will be described later, is executed.
[0099] 11(A) to 11(E) are diagrams showing examples of operation screens for setting the person illustration display according to the third embodiment, and FIG. 11(A) is an operation screen 1110 when there is no instruction to display an illustration (referred to as normal display mode).
[0100] When the person illustration display button 1112 is off, it is determined in step S1010 of Fig. 10A that "No" has been returned, i.e., no instruction to display an illustration has been issued. In this state, the normal display mode is entered, and an image in which the person in the image is filled in is displayed in the image display area 1111.
[0101] Fig. 10B is a flowchart showing an example of the processing flow in the normal display mode, showing a detailed example of step S1020 in Fig. 10A. Steps S301 and S302 are the same as those in the first embodiment.
[0102] In step S1021, the person area determination unit 230 determines the area occupied by the person detected by the person detection unit in step S302, and determines the area to be processed thereafter. At this time, because no illustration is displayed, additional processing is performed, such as adjusting the boundary position between the person and the background, compared to the first embodiment, so that the person can be easily recognized as a silhouette.
[0103] In step S1022, the person area determined in step S1021 is uniformly filled in by person area processing unit 250. In this processing, a method other than uniform filling may be used as long as the processing makes the original features of the person unrecognizable.
[0104] If the average brightness of the background is equal to or greater than a predetermined threshold V1, the background is painted with a brightness of V2 or less, and if the average brightness of the background is less than the predetermined threshold V1, the background is painted with a brightness of V3 or more.<V1、V3-β> V1, α, and β are set to predetermined contrast differences. Alternatively, the area may be filled with a color whose saturation differs by a predetermined value or more from the background color. Also, each person may be filled with a different color.
[0105] In step S1023, the video output unit 290 displays the image in which the person area has been filled in in step S1022 on the display device 600 via the display control unit 105.
[0106] In step S1024, CPU 101 determines whether the user has issued an instruction to display an illustration from operation device 700. If it determines that an instruction to display an illustration has not been issued, the process returns to step S301 and processes the next image. If it determines in step S1024 that an instruction to display an illustration has been issued, the flowchart of the illustration display process in Figure 10(C) is executed. Note that Figure 10(C) shows a detailed example of the process in step S1030 in Figure 10(A).
[0107] Fig. 11(B) shows an operation screen 1110 when an instruction to display an illustration has been given (referred to as illustration display mode). When the person illustration display button 1112 is on, it is determined in step S1010 of Fig. 10(A) that "Yes", i.e., that an instruction to display an illustration has been given. In this state, the illustration display mode is entered, and an image with a person illustration superimposed in place of the person in the image is displayed in the image display area 1111.
[0108] In this way, the person illustration display button 1112 functions as a display mode operation means for switching between a normal display mode, which displays an image with people removed, and an illustration display mode, which displays an image with an illustration superimposed.
[0109] Fig. 10C is a diagram showing an example of processing in the illustration display mode, and shows a detailed example of processing in step S1030 in Fig. 10A. Steps S301 to S308 are the same as those in the first embodiment.
[0110] In step S1031, CPU 101 determines whether or not an instruction to update the illustration has been issued. That is, if the user presses update button 1113 in Fig. 11(B), the illustration update instruction is validated, and a determination of "Yes" is made in step S1031. Here, update button 1113 functions as update operation means for specifying an image update in the illustration display mode.
[0111] Also, the illustration update instruction can be automatically enabled at a predetermined interval. Figure 11(D) shows the screen for automatically updating an illustration, with the automatic update button 1114 enabled. When the user presses the automatic update button 1114, the automatic illustration update setting screen 1120 shown in Figure 11(E) is displayed.
[0112] When the user inputs the period for updating the person illustration in the update period setting area 1121 and presses the OK button 1123, the illustration update instruction becomes valid at the set period. Here, the update period setting area 1121 functions as an automatic update period designation means for designating the automatic update period of the image in the illustration display mode. Note that 1122 is a cancel button.
[0113] If the illustration update instruction is valid, that is, if the determination in step S1031 is "Yes," the process returns to image acquisition in step S301, and illustration display processing for the next image is executed. Figure 11(C) is a diagram showing an example of the operation screen after update button 1113 is pressed in Figure 11(B). As shown in Figure 11(C), the image displayed in image display area 1111 is updated.
[0114] If the illustration update instruction is invalid, that is, if the determination in step S1031 is "No", the process proceeds to step S1032, where CPU 101 determines whether the user has used operation device 700 to instruct the illustration display to end.
[0115] If it is determined in step S1032 that an instruction to end the illustration display has been issued, i.e., if the answer is "Yes," the process proceeds to the normal display process (step S1020) shown in Fig. 10(B). On the other hand, if it is determined in S1032 that an instruction to end the illustration display has not been issued, i.e., if the answer is "No," the process returns to step S1031 to determine whether an instruction to update the illustration has been issued.
[0116] It is also possible to switch the display / non-display of the generated person illustration at a predetermined cycle. If switching at a predetermined cycle is enabled on an operation screen (not shown), the CPU 101 switches the illustration display instruction according to that cycle in step S1010 of Fig. 10(A). A GUI screen that allows the user to specify this cycle may also be provided.
[0117] As described above, in the third embodiment, a method has been described in which the user can, for example, specify whether to display or hide the generated illustration. This allows an image in which people are usually blacked out to protect privacy, and an illustration that expresses the person's attributes is displayed only when necessary. This makes it possible to provide an image that allows users to intuitively grasp the general features of a person while reducing the processing cost of generating an illustration while reliably protecting privacy.
[0118] [Embodiment 4] In the first to third embodiments, a method has been described in which an illustration of a person is generated based on person attribute information and the illustration is superimposed to replace the person in the original image. In this case, there are no limitations on the person attribute information used to generate the illustration or the type of illustration to be generated.
[0119] In contrast, in the fourth embodiment, the degree to which the expression of a person in an image is abstracted is referred to as the abstraction level, and an example will be described in which the abstraction level is switched depending on predetermined conditions such as the purpose, scene, region, etc. Furthermore, an example will be described in which person attribute information that is the basis for generating an illustration is narrowed down (filtered) depending on predetermined conditions.
[0120] In the fourth embodiment, the degree to which illustrations, icons, etc. representing people reflect the original person's information is changed depending on the abstraction level or other settings and conditions. For example, depending on the abstraction level, generated illustrations or pre-stored illustrations, icons, and symbols are used. Furthermore, the person attribute information used to generate the illustration is filtered according to conditions.
[0121] Table 1 divides the abstraction levels of people into six stages and shows examples of the corresponding display formats, uses, and situations for each. [Table 1]
[0122] In Table 1, at abstraction level 0, the display format is the original image, and the people are displayed as they are. At abstraction level 1, an illustration generated using the attribute information of the people in the image, as shown in embodiments 1 to 3, is displayed instead of the person image. By applying different illustrations to each person, it is possible to provide an image suitable for purposes such as expressing the bustle of an event venue.
[0123] At abstraction level 2, the types of person attribute information used to generate the illustrations are reduced. Furthermore, if the specified person attributes are the same, the generated illustrations are used commonly for different people. As the abstraction level increases, attributes are expressed using pre-defined icons or symbols instead of illustration generation, and at abstraction level 5, person attribute information is not used and the person is expressed using a single symbol or icon.
[0124] 12(A) to 12(H) are diagrams illustrating examples of representation formats of people according to embodiment 4. Fig. 12(A) is a display format applied to abstraction level 1, and shows an example of displaying an illustration generated for each person using the person's attribute information.
[0125] Figure 12(B) shows an example of a display format applied to abstraction level 2, in which illustrations generated using person attribute information are commonly applied to people with the same attributes. In this case, the only person attribute information used is age, and two adults and one child are depicted in the illustration.
[0126] In addition, if the attributes to be used are set to a finite number in advance, instead of generating illustrations using the generation AI 500 or the generation AI unit 270, multiple illustrations corresponding to each attribute may be created in advance, stored in the memory unit 103, and used for image synthesis.
[0127] Figure 12(C) is a diagram showing an example of a display format applied to abstraction level 3, which shows an example of person icons of two adults and one child. Figures 13(D) and (E) are diagrams showing examples of display formats applied to abstraction level 4, which show two adults (ADL: Adult) and one child (CHL: Child) using symbols that allow identification of attributes.
[0128] Here, symbols include graphics and text information. Figures 12(F), (G), and (H) show examples of display formats applied to abstraction level 5, in which people are represented by a single symbol (HM: Human) or icon, and attribute information is not included.
[0129] In addition, the user may be able to set the abstraction level, etc. Figure 13 is a diagram showing an example of an operation screen for setting the abstraction level of a person in embodiment 4, and shows an operation screen 1300 for setting the person abstraction level.
[0130] An operation screen 1300 in FIG. 13 has an abstraction level bar 1301 that indicates levels from "OFF" (no abstraction) to "HIGH" abstraction level, and a slider 1302 that specifies the setting value of the abstraction level.
[0131] Also provided are a display 1303 showing a number of abstracted display examples, a cancel button 1304, and an OK button 1305. The user can operate the slider 1302 to set the abstraction level.
[0132] Here, the operation screen 1300 functions as an abstraction level designation means for designating the abstraction level when illustrating a person, and the attributes input to the illustration acquisition means can be changed according to the abstraction level designated by the abstraction level designation means.
[0133] Note that, although Fig. 13 shows an example of an operation screen where the user directly sets the abstraction level, the present invention is not limited to this example of an operation screen. For example, the operation screen may be provided with options corresponding to the "Use / Scene" in Table 1, and when the user selects an option according to the use or scene of the image display, the abstraction level associated with that option may be set.
[0134] Alternatively, the generation AI 500 may analyze the scene based on the acquired video, generate text information that expresses the scene, and automatically determine the most appropriate abstraction level from the generated text information.
[0135] Table 2 shows examples of "usage areas" that indicate the country or region where the image is used, and "usable personal attribute information" that corresponds to each use area. For personal information that is handled differently depending on the country or region, data corresponding to Table 2 may be stored in advance in the storage unit 103, and the user may specify which abstraction ID to apply in the settings at the time of use. Alternatively, it may be automatically set which use area data to use based on GPS data, etc. [Table 2]
[0136] 14(A) and (B) are diagrams showing examples of operation screens for performing person abstraction settings according to the fourth embodiment, and show examples of operation screens for a user to edit the data shown in Table 2. Fig. 14(A) shows a person abstraction setting screen 1400, on which an abstraction button 1401 corresponding to the abstraction ID in Table 2 is displayed.
[0137] The name of the abstraction button may be the same character string as the "usage area" in Table 2, but as will be described later in Fig. 14(B), the user can freely edit it on the detailed setting screen 1410. When any one of the abstraction buttons 1401 is selected and the OK button 1404 is pressed, the corresponding abstraction ID in Table 2 is applied.
[0138] In the state shown in Fig. 14(A), "Custom 1" is highlighted, and pressing the OK button 1404 in this state applies "Custom 1," i.e., abstraction ID 6 in Table 2. When the Details button 1402 is pressed on this screen, a details setting screen 1410 for the abstraction button 1401 selected at that time is displayed, as shown in Fig. 14(B).
[0139] On the detailed setting screen 1410, it is possible to set the name of the abstraction button 1401 and the person attributes that can be used with the abstraction ID corresponding to that button. A person attribute list 1412 displays a list of person attribute items, and an item with a check mark in the check box 1413 for that item is set as usable with that abstraction ID.
[0140] 14 and Table 2 show a method in which usable person attribute information is stored as data in advance, and the user edits and selects the abstraction ID to be used, but this is not limiting. For example, information about the country or region where the image will be used may be obtained in some way, input to the generation AI 500, and the type of person attribute information permitted in that country or region may be obtained, thereby internally setting the person attribute information to be used for generating the illustration.
[0141] That is, an attribute filtering means may be provided that filters attributes according to at least one of country, region, and purpose, and the filtered attributes may be input to the illustration acquisition means.
[0142] In the above, in the fourth embodiment, a method for determining illustrations, icons, and symbols to be used in processing to edit people appearing in an image, and a method for determining person attribute information to be used in generating an illustration, have been described. According to the fourth embodiment, it is possible to provide an edited image for appropriate privacy protection according to conditions such as the purpose, scene, and location of use.
[0143] The present invention has been described in detail above based on its preferred embodiments, but the present invention is not limited to the above embodiments, and various modifications and combinations of the above embodiments are possible based on the spirit of the present invention, and these are not excluded from the scope of the present invention.
[0144] The present invention also includes those that realize the functions of the above embodiments using, for example, at least one processor such as a CPU, memory, or circuit (for example, ASIC). Also, multiple processors may be used to perform distributed processing.
[0145] In order to realize part or all of the control in the above-described embodiments, a computer program that realizes the functions of the above-described embodiments may be supplied to an image processing device or the like via a network or various storage media. Then, a computer (or a CPU, MPU, or the like) in the image processing device or the like may read and execute the program. In this case, the program and the storage medium storing the program constitute the present invention. The present invention also includes the following combinations.
[0146] (Configuration 1) An image processing device comprising: a video acquisition means; a person detection means for detecting a person from a video acquired by the video acquisition means; an attribute determination means for determining the attributes of the person detected by the person detection means; an illustration acquisition means for acquiring an illustration of the person by inputting a prompt based on the attribute determined by the attribute determination means; and an image synthesis means for superimposing the illustration acquired by the illustration acquisition means on an area of the person.
[0147] (Configuration 2) The image processing device according to Configuration 1, characterized in comprising: a person area determination means for determining a person area on the image occupied by the person detected by the person detection means; and a person area processing means for processing the area determined by the person area determination means to erase the person from the image.
[0148] (Configuration 3) The image processing device according to Configuration 2, further characterized in that the process of erasing the person by the person area processing means is any one of a process of filling the person area uniformly or with another image, a process of interpolating the inside of the person area using information on areas adjacent to the person area, and a process of making the person area a background using video information.
[0149] (Configuration 4) The image processing device according to configuration 2 or 3, further comprising a display mode operation means for switching between a normal display mode in which the image is displayed with the person erased and an illustration display mode in which the image is displayed with the illustration superimposed.
[0150] (Configuration 5) The image processing device according to configuration 4, further comprising update operation means for specifying update of the image in the illustration display mode.
[0151] (Configuration 6) The image processing device according to configuration 5, further comprising an automatic update period designation means for designating an automatic update period of the image in the illustration display mode.
[0152] (Configuration 7) The image processing device according to any one of configurations 1 to 6, wherein the illustration acquisition means acquires the illustration generated by a generation AI based on the prompt.
[0153] (Configuration 8) The image processing device according to any one of configurations 1 to 7, further comprising an operation means for specifying additional information to be input to the illustration acquisition means.
[0154] (Configuration 9) The image processing device described in Configuration 8, wherein the additional information represents at least one of what the person in the illustration is wearing, a building, a landscape, a time of day, a season, a country, a region, and a related event.
[0155] (Configuration 10) An image processing device according to configuration 8 or 9, characterized in that it has a person selection means for selecting one or more of the people detected by the person detection means, and the operation means is capable of specifying the additional information for the people selected by the person selection means.
[0156] (Configuration 11) The image processing device described in any one of configurations 8 to 10, characterized in that the operation means can specify at least one of the following as the additional information: a specific item worn by the person, a specific graphic to be displayed in close proximity to the person, text information to be displayed in close proximity to the person, and a color or image to decorate the person.
[0157] (Configuration 12) An image processing device according to any one of configurations 1 to 11, characterized in that it has an abstraction level designation means for designating an abstraction level when illustrating the person, and changes the attributes input to the illustration acquisition means according to the abstraction level designated by the abstraction level designation means.
[0158] (Configuration 13) An image processing device according to any one of configurations 1 to 12, characterized in that it has an attribute filtering means for filtering the attributes according to at least one of country, region, and purpose, and the filtered attributes are used as the attributes to be input to the illustration acquisition means.
[0159] (Method) An image processing method comprising: an image acquisition step; a person detection step for detecting a person from the image acquired in the image acquisition step; an attribute determination step for determining the attributes of the person detected in the person detection step; an illustration acquisition step for acquiring an illustration of the person by inputting a prompt based on the attribute determined in the attribute determination step; and an image synthesis step for superimposing the illustration acquired in the illustration acquisition step on the area of the person.
[0160] (Program) A computer program for controlling each means of the image processing device according to any one of configurations 1 to 13 by a computer. [Explanation of symbols]
[0161] 100: Image processing device 220: Person detection unit 230: Human area determination section 240: Person attribute determination section 250: Human area processing section 260: Image synthesis unit 270:Generation AI Department 500: Generation AI
Claims
1. An image acquisition means; person detection means for detecting a person from the image acquired by the image acquisition means; attribute determination means for determining attributes of the person detected by the person detection means; an illustration acquisition means for acquiring an illustration of the person by inputting a prompt based on the attribute determined by the attribute determination means; an image synthesis means for superimposing the illustration acquired by the illustration acquisition means on the area of the person; 1. An image processing device comprising:
2. a person area determination means for determining a person area on the video image occupied by the person detected by the person detection means; a person area processing means for processing the area determined by the person area determination means to erase the person from the image; 2. The image processing device according to claim 1, further comprising:
3. The process of erasing the person by the person area processing means includes: Filling the person area with a uniform or other image; a process of interpolating the inside of the person region using information on the region adjacent to the person region; A process of converting the person area into a background using video information 3. The image processing device according to claim 2, further characterized in that:
4. a normal display mode in which the image is displayed with the person removed; an illustration display mode in which the image on which the illustration is superimposed is displayed; and 3. The image processing apparatus according to claim 2, further comprising:
5. 5. The image processing apparatus according to claim 4, further comprising update operation means for specifying update of said image in said illustration display mode.
6. 6. The image processing apparatus according to claim 5, further comprising automatic update period designation means for designating an automatic update period of the image in the illustration display mode.
7. The image processing device according to claim 1 , wherein the illustration acquisition means acquires the illustration generated by a generation AI based on the prompt.
8. 2. The image processing apparatus according to claim 1, further comprising an operation means for specifying additional information to be input to said illustration acquisition means.
9. The additional information is 9. The image processing device according to claim 8, wherein the illustration represents at least one of an item worn by the person, a building, a landscape, a time period, a season, a country, a region, and a related event.
10. a person selection means for selecting one or more people from the people detected by the person detection means, 9. The image processing apparatus according to claim 8, wherein the operation means is capable of specifying the additional information for the person selected by the person selection means.
11. The operation means may include, as the additional information: specific items worn by said person; a specific figure to be displayed in a position close to the person; text information to be displayed in proximity to the person; colors or images decorating said person; 9. The image processing apparatus according to claim 8, wherein at least one of the following can be specified.
12. an abstraction level designation means for designating an abstraction level when illustrating the person; 2. The image processing apparatus according to claim 1, wherein the attribute input to the illustration obtaining means is changed in accordance with the abstraction level designated by the abstraction level designating means.
13. 2. The image processing device according to claim 1, further comprising an attribute filtering means for filtering the attributes according to at least one of country, region, and purpose, and the filtered attributes are input to the illustration acquisition means.
14. an image acquisition step; a person detection step of detecting a person from the image acquired in the image acquisition step; an attribute determination step of determining attributes of the person detected in the person detection step; an illustration acquisition step of acquiring an illustration of the person by inputting a prompt based on the attribute determined in the attribute determination step; an image synthesis step of superimposing the illustration acquired in the illustration acquisition step on the area of the person; An image processing method comprising:
15. A computer program for controlling each unit of the image processing apparatus according to any one of claims 1 to 13 by a computer.
Citation Information
Patent Citations
CLR2016
Display plate for clock
JP1983034193A
Information processing device, information processing system, information processing method, program and storage medium
JP2020091770A