Image editing device, image editing method, and program
The image editing device accurately applies visual effects to target objects in moving images by recognizing and considering display attributes, addressing unintended applications and aesthetic issues in conventional methods.
Patent Information
- Application Number
- JP2025098400
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Conventional techniques often fail to apply visual effects to target objects as intended, leading to unintended application on non-target objects, aesthetic impairment, and reduced readability of videos.
An image editing device that recognizes target objects, applies visual effects by generating and overlaying target object images, and considers display attributes to harmonize with other image elements.
Enables precise application of visual effects in line with user intentions, maintaining aesthetic unity and readability of moving images.
Smart Images

Figure 0007799121000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for adding visual effects to moving images. [Background technology]
[0002] Conventionally, techniques have been proposed for generating and editing video using generative AI (Artificial Intelligence). For example, ChatGPT-4o, a large-scale language model, has the function of editing target video based on instructions entered in natural language. Also, a technique for generating video from text information is known (see, for example, Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] "Adobe Firefly Video Model: Bringing the power of generative AI to video," [online], [Retrieved May 12, 2025], Internet, <URL:https: / / blog.adobe.com / jp / publish / 2024 / 09 / 13 / cc-video-bringing-gen-ai-to-video-adobe-firefly-video-model-coming-soon> Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventional techniques sometimes fail to edit a target video as intended by a user. For example, conventional techniques sometimes apply visual effects to unintended objects. To avoid this, it is necessary to individually specify which objects to apply visual effects to for each video. Furthermore, conventional techniques sometimes apply visual effects to parts other than the target object, even when a user wants to apply a visual effect only to a specific object. Furthermore, conventional techniques sometimes fail to apply visual effects to a target object in a way that harmonizes with other parts, potentially impairing the aesthetics and readability of the video.
[0005] In view of the above circumstances, an object of the present invention is to provide a technique that can impart visual effects to moving images in a manner that is more in line with the user's intentions. [Means for solving the problem]
[0006] One aspect of the present invention is an image editing device comprising a target image input unit that inputs a target image, which is an image of a target to which a visual effect is to be applied; an object recognition unit that recognizes one or more objects from the target image; a target object recognition unit that recognizes a target object, which is an object to which the visual effect is to be applied, from the one or more objects recognized by the object recognition unit; and a visual effect application unit that applies the visual effect to the target object included in the target image.
[0007] One aspect of the present invention is the above-mentioned image editing device, wherein the visual effect imparting unit generates a target object image, which is an image of the target object to which the visual effect has been imparted, by imparting the visual effect to an image of the target object, and imparts the visual effect to the target image by overlaying the target object image on the target image.
[0008] In one aspect of the present invention, in the image editing device described above, the generation of the target object image is performed by generating a program for displaying the image of the target object in a manner that has a visual effect.
[0009] One aspect of the present invention is the above-mentioned image editing device, wherein the visual effect imparting unit removes an image of the target object from the target image and overlays the target object image on the target image from which the image of the target object has been removed.
[0010] One aspect of the present invention is the above-mentioned image editing device, further comprising an attribute information acquisition unit that acquires attribute information regarding the display attributes of the object from the target image, and the visual effect imparting unit imparts a visual effect to the target object in a manner corresponding to the display attributes of the object based on the attribute information.
[0011] One aspect of the present invention is the above-mentioned image editing device, wherein the object is text within the target image, and the attribute information includes one or more of the number of characters, direction, position, line break position, and area shape of the string of characters that make up the text.
[0012] One aspect of the present invention is the above-mentioned image editing device, wherein the object is text in the target image, the attribute information includes attributes of a decorative object associated with the text, and the visual effect imparting unit imparts a visual effect to the target object in a manner corresponding to the attributes of the decorative object.
[0013] One aspect of the present invention is the above-mentioned image editing device, wherein the object is text within the target image, and the target object recognition unit performs a layout analysis of a group of text within the target image and selects target text, which is text to which a visual effect is to be applied, based on the results of the layout analysis.
[0014] One aspect of the present invention is the above-mentioned image editing device, wherein the attribute information includes display attributes for multiple objects included in the target image, and the visual effect imparting unit imparts visual effects to the target objects in a manner corresponding to the display attributes of the multiple objects.
[0015] One aspect of the present invention is the above-mentioned image editing device, wherein the visual effect imparting unit imparts a visual effect to the target object based on the attribute information and rules regarding the imparting of the visual effect so that the rules are applied between the display attributes of objects other than the target object and the display attributes of the target object.
[0016] One aspect of the present invention is the above-mentioned image editing device, wherein the object is a foreground object in the target image, and the attribute information includes one or more of the position, category, and area shape of the foreground object.
[0017] In one aspect of the present invention, in the image editing device, the target image is an advertising image for advertising a product, and the object is the product.
[0018] One aspect of the present invention is an image editing method for the above-mentioned image editing device, comprising: a target image input step for inputting a target image, which is an image to which a visual effect is to be applied; an object recognition step for recognizing one or more objects from the target image; a target object recognition step for recognizing a target object, which is an object to which the visual effect is to be applied, from the one or more objects recognized in the object recognition step; and a visual effect application step for applying the visual effect to the target object included in the target image.
[0019] One aspect of the present invention is the above-mentioned image editing device, which is a program for causing a computer to execute the following steps: a target image input step and unit for inputting a target image, which is an image to which a visual effect is to be applied; an object recognition step for recognizing one or more objects from the target image; a target object recognition step for recognizing a target object, which is an object to which the visual effect is to be applied, from the one or more objects recognized in the object recognition step; and a visual effect application step for applying the visual effect to the target object included in the target image. [Effects of the Invention]
[0020] According to the present invention, it is possible to impart visual effects to moving images in a manner more in line with the user's intentions. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a diagram illustrating an example of a functional configuration of an advertising image editing device 100A according to a first embodiment. [Figure 2] 10 is a flowchart showing an example of a processing flow in which advertising image editing device 100A of the first embodiment imparts a visual effect to a target image. [Figure 3] FIG. 10 is an image diagram illustrating an outline of a processing flow for adding a visual effect to target text in a target image. [Figure 4] FIG. 10 is a diagram illustrating an example of a functional configuration of an advertising image editing device 100B according to a second embodiment. [Figure 5] 10 is a flowchart showing an example of a processing flow in which advertising image editing device 100B according to the second embodiment imparts a visual effect to a target image. [Figure 6] FIG. 10 is a diagram illustrating an example of a functional configuration of an advertising image editing device 100C according to a third embodiment. [Figure 7] 11 is a flowchart showing an example of a processing flow in which an advertising image editing device 100C according to the third embodiment imparts a visual effect to a target image. [Figure 8]FIG. 10 is an image diagram illustrating an outline of a processing flow for adding a visual effect to a target object in a target image. DETAILED DESCRIPTION OF THE INVENTION
[0022] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0023] First Embodiment FIG. 1 is a diagram illustrating an example of the functional configuration of an advertising image editing device 100A according to the first embodiment. The advertising image editing device 100A is a device that has a function of applying visual effects to a target advertising image. The advertising image editing device 100A includes, for example, a target image input unit 110, a text recognition unit 120, a target text recognition unit 130, a visual effect applying unit 140A, and a storage unit 150. The advertising image editing device 100A includes, for example, a processor such as a CPU (Central Processing Unit) and a memory. The advertising image editing device 100A functions as a device including the target image input unit 110, the text recognition unit 120, the target text recognition unit 130, the visual effect applying unit 140A, and the storage unit 150 by the processor executing a program. All or part of the functions of advertising image editing device 100A may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The above program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs: Solid State Drives), as well as storage devices such as hard disks and semiconductor storage devices built into computer systems. The above program may be transmitted via a telecommunications line.
[0024] The target image input unit 110 inputs an advertising image (hereinafter referred to as "target image") to which a visual effect is to be applied. The target image input unit 110 includes, for example, a network interface and receives data of the target image (target image data) from another device on the network. Furthermore, for example, the target image input unit 110 may include a connection interface for connecting to a storage device and acquire the target image data from the storage device.
[0025] The text recognition unit 120 recognizes text in a target image. The text recognition unit 120 can recognize the text in the target image by performing OCR (Optical Character Recognition) processing on the target image. The text recognition unit 120 stores text information 151 related to the text recognized from the target image in the storage unit 150.
[0026] The target text recognition unit 130 recognizes text within the target image that is to be given a visual effect (hereinafter referred to as "target text"). The target text recognition unit 130 recognizes the target text based on the target image and text information. The target text recognition unit 130 stores target text information 152 relating to the recognized target text in the storage unit 150.
[0027] The visual effect applying unit 140A applies a visual effect to the target image. More specifically, the visual effect applying unit 140A applies visual information to the target text in the target image based on the target image and the target text information 152. Examples of visual effects include fade-in, fade-out, zoom-in, zoom-out, blinking, and rotation. The visual effect may be a combination of these. The visual effect applying unit 140A may be configured using an API (Application Program Interface) that provides an image editing function, or may be configured using a generation AI (Artificial Intelligence) with a video generation function, such as ChatGPT-4o (registered trademark). The visual effect applying unit 140A stores the target image after the visual effect has been applied (edited target image) in the storage unit 150.
[0028] The storage unit 150 is configured using, for example, a magnetic storage device such as an HDD, or a semiconductor storage device such as an SSD (Solid State Drive), flash memory, etc. The storage unit 150 stores, for example, text information 151, target text information 152, and a post-edit target image 153.
[0029] 2 is a flowchart showing an example of the processing flow in which the advertising image editing device 100A of the first embodiment applies a visual effect to a target image. In the advertising image editing device 100A, the target image input unit 110 first inputs the target image (S101). Next, the text recognition unit 120 performs text recognition processing on the target image input in S101 to recognize text within the target image (S102). For example, the text recognition unit 120 performs OCR processing to recognize the coordinates of text within the target image and character strings indicating the text, and outputs text information 151 including the recognized coordinates and character strings.
[0030] Next, the target text recognition unit 130 performs a process of recognizing the target text in the target image based on the text information 151 output in S102 (S103). For example, the target text recognition unit 130 can analyze the category of the text group in the target image based on the text information 151, and select, as the target text, a text having a category corresponding to the target text category. Here, the target text category is a category of text selected as the target text, and is set in advance in the advertising image editing device 100A.
[0031] The category analysis may be a rule-based method for recognizing categories based on text character strings, or may be implemented using an LLM (Large Language Model) such as ChatGPT-4o (registered trademark). For example, the target text recognition unit 130 may provide text information 151 recognized from the target image to the LLM and have the LLM determine to which target text category the character string in the text information 151 belongs. As a result, for example, if the target text category "product name" is set, the target text recognition unit 130 can select, as the target text, text that includes the product name from among the text groups in the target image. The advertising image editing device 100A may be configured as a device that includes an LLM, or may be configured to use an LLM service that can communicate via a network.
[0032] The target text recognition unit 130 may also provide text information 151 to the LLM, causing it to directly determine which text to apply visual effects to without going through categories. In this case, the LLM may be provided with information indicating criteria for selecting the text to which the visual effect is to be applied (target text). The target text recognition unit 130 may also be configured to accept a text designation operation by a user and recognize the designated text as the target text.
[0033] The target text recognition unit 130 may also be configured to perform a layout analysis of text groups in the target image based on text information 151 recognized from the target image and select target text based on the results of the layout analysis. For example, the target text recognition unit 130 may be configured to recognize text with a large font size as target text, or to recognize, among multiple recognized text regions, a text region whose layout meets specific conditions as target text. For example, the target text recognition unit 130 may be configured to recognize, among multiple text regions, a text region located in the center as target text, or a text region located at the top or bottom as target text. The target text recognition unit 130 may also be configured to infer the attributes of each text region by inputting the results of the layout analysis into a model that has learned the relationship between text layout and the attributes of the text, and to recognize text regions with specific attributes, such as "headline" or "product name," as target text. The target text recognition unit 130 outputs target text information 152 including the coordinates and character strings of the recognized target text.
[0034] Next, the visual effect applying unit 140A applies a visual effect to the target text in the target image based on the target text information 152 output in S103 (S104).
[0035] FIG. 3 is an image diagram illustrating an outline of the processing flow for applying visual effects to target text in a target image. First, the text recognition unit 120 recognizes text TX1 and TX2 by performing text recognition processing on the target image IM11. Next, the target text recognition unit 130 recognizes text TX2 of the recognized texts TX1 and TX2 as target text. Next, the visual effect applying unit 140A applies visual effects to the recognized target text TX2. More specifically, the visual effect applying unit 140A generates, for example, information (visual effect information) for applying visual effects to the target text TX2.
[0036] For example, when a visual effect is applied using a function of a web browser, the visual effect applying unit 140A can generate source code for applying the visual effect as visual effect information. Fig. 3 illustrates an example of source code as visual effect information, in which style information CD11 such as CSS (Cascading Style Sheet) and HTML (Hyper-Text Markup Language) information CD12 are generated. For example, by defining multiple display patterns for the target text TX2 in the style information CD11 and describing a process for switching between the display patterns in the HTML information CD12, a target image IM13 with a visual effect applied to the target text TX2 can be displayed.
[0037] More specifically, the visual effect imparting unit 140A generates visual effect information while also generating a target image IM12 in which the target text TX2 is deleted from the target image IM11. For example, the visual effect imparting unit 140A can generate the target image IM12 by replacing the image of the target text TX2 in the target image IM11 with an image similar to the surrounding background. This type of image replacement is called image inpainting, and any image inpainting technique may be used by the visual effect imparting unit 140. The visual effect imparting unit 140A executes the style information CD11 and the HTML information CD12 to display the target text TX2 on the target image IM12 and periodically switch the display mode of the target text TX2. This allows a target image IM13 having a visual effect on the target text TX2 to be displayed.
[0038] Here, a method for displaying a target image IM13 to which a visual effect is applied using style information CD11 and HTML information CD12 has been described. In this case, the target image IM12 from which the target text TX2 has been deleted, the style information CD11, and the HTML information CD12 may be stored in the storage unit 150 as the edited target image 153. On the other hand, the visual effect applying unit 140A may be configured to generate an image of the target text TX2 with a visual effect and combine it with the target image IM12 to generate the image with the visual effect (for example, an image in PNG format) itself as the edited target image 153. In this way, by previously deleting the target to which the visual effect is applied from the target image before combining, the visual effect can be applied in a more natural manner when a moving animation is applied as a visual effect.
[0039] According to the advertising image editing device 100A of the first embodiment described above, it is possible to recognize target text from a target image and impart visual effects to the recognized target text, thereby preventing unintended visual effects from being imparted to text other than the target text.
[0040] Second Embodiment 4 is a diagram showing an example of the functional configuration of the advertising image editing device 100B of the second embodiment. The advertising image editing device 100B of the second embodiment differs from the advertising image editing device 100A of the first embodiment in that it includes a visual effect adding unit 140B instead of the visual effect adding unit 140A, and in that it further includes an attribute information acquiring unit 160B. The rest of the configuration is the same as that of the advertising image editing device 100A. Therefore, the following will comprehensively describe the components of the advertising image editing device 100B that are different from those of the advertising image editing device 100A, and will omit a description of the common components as much as possible.
[0041] The attribute information acquisition unit 160B acquires attribute information 154 related to the text in the target image. For example, the attribute information 154 is information indicating attributes related to the display mode of the text. For example, the attribute information 154 may include information on decorative objects displayed in association with the text. Here, decorative objects are objects that decorate the text, such as underlines, side marks, emphasis marks, and side lines. The attribute information 154 may include information such as the position, size, and color of the decorative objects, and the correspondence between the decorative objects and the text.
[0042] For example, decorative objects can be detected using an object detection model trained for detecting decorative objects. Examples of object detection models include R-CNN, YOLO, SSD, and the like, which are based on a convolutional neural network (CNN). Furthermore, decorative object detection can be achieved by generating a binary image indicating decorative objects and other portions using a semantic segmentation model trained for detecting decorative objects. Furthermore, the correspondence between decorative objects and text can be determined based on, for example, the distance between the decorative objects and the text.
[0043] The attribute information 154 may also include, for example, text attributes (such as orientation, size, character color, and font type). The text attributes may be included in the text information 151 instead of in the attribute information 154. The attribute information 154 may also include, for example, attributes of foreground objects other than text. For example, the attributes of a foreground object may be the position or type of the foreground object. Examples of foreground objects include a "person," a "product," and a "logo." The text orientation can be determined, for example, using an image recognition model trained for text recognition. The text orientation may also be determined based on the shape of a text region. For example, an object detection model trained for foreground object detection can be used to detect foreground objects other than text. The attribute information acquisition unit 160B stores the acquired attribute information 154 in the storage unit 150.
[0044] The visual effect applying unit 140B applies a visual effect to the target text in the target image. The visual effect applying unit 140B applies visual information to the target text in the target image based on the target image, the target text information 152, and the attribute information 154. The visual effect applying unit 140B stores the target image after the visual effect has been applied (edited target image) in the storage unit 150.
[0045] 5 is a flowchart showing an example of the flow of processing in which advertising image editing device 100B of the second embodiment imparts a visual effect to a target image. In FIG. 5, S101 to S103 are the same as in FIG. 2, and therefore description thereof will be omitted. Following S103, attribute information acquisition unit 160B acquires attribute information 154 related to text in the target image (S201). For example, attribute information acquisition unit 160B recognizes the position and character string of text in the target image based on text information 151 output by text recognition unit 120, and outputs attribute information 154 by performing a recognition process on the display attributes of characters on the image of the recognized text area.
[0046] Next, the visual effect imparting unit 140B imparts a visual effect to the target text in the target image in a manner corresponding to the display attributes of the other text, based on the target text information 152 output in S103 and the attribute information 154 output in S201 (S202). For example, the visual effect imparting unit 140B generates visual effect information for imparting a visual effect to the target text in the target image so as to maintain the display attributes of the target text. Furthermore, for example, the visual effect imparting unit 140B may recognize text other than the target text based on the text information 151 and the target text information 152, and generate visual effect information such that the target text is displayed in a manner corresponding to the display attributes of the other text.
[0047] More specifically, visual effect applying section 140B generates visual effect information that takes into consideration the following factors, for example, regarding the display mode of text. If the target text has decorations such as underlines or brackets, the display mode of the decorations will change in conjunction with changes in the display mode of the target text (fading, moving, etc.). -Applies the same visual effect to target text within a specified range (for example, within the same sentence). -Prevent text with visual effects from obscuring other text.
[0048] [Example of attribute information and visual effects] For example, when adding a visual effect that involves movement, such as fading, the "orientation" of the text can be used as attribute information. In this case, for example, the visual effect adding unit 140B may recognize the orientation of the target text based on the attribute information and generate visual effect information that adds vertical movement to vertically written text and horizontal movement to horizontally written text. In this way, it is possible to prevent a decrease in visibility from occurring when the movement of the characters caused by the visual effect does not match the movement of the eyes reading the characters.
[0049] Furthermore, in addition to the text orientation, the "number of line breaks" in the text can be used as attribute information. Generally, the more line breaks there are in the text, the more difficult it tends to be to visually determine whether the text is written vertically or horizontally. Therefore, when adding a movement such as a fade to text with many line breaks, the direction in which the visual effect is added becomes particularly important. For this reason, the visual effect adding unit 140B may be configured to give priority to text with many line breaks and add visual effects so that the direction of the text and the direction of the movement more closely match.
[0050] Furthermore, for example, the "shape of the text area" can be used as attribute information. As in the case of the "number of line breaks" described above, when the shape of the text area is closer to a square, it tends to be more difficult to visually determine whether the text is written vertically or horizontally compared to when the shape is vertically or horizontally long. Therefore, as in the case of the "number of line breaks," the visual effect imparting unit 140B may be configured to impart visual effects preferentially to text areas closer to a square in shape, so that the orientation of the text and the direction of movement more closely match.
[0051] Furthermore, for example, the "number of characters" in the text can be used as attribute information. The more characters in the text, the more likely it is that there will be line breaks. Therefore, similar to the above-described "number of line breaks" and "shape of the text area," the more characters there are in the text, the less visible the orientation of the text becomes. Therefore, similar to the "number of line breaks" and "shape of the text area," the visual effect imparting unit 140B may be configured to impart visual effects preferentially to text with a larger number of characters so that the orientation of the text and the direction of movement more closely match.
[0052] According to the advertising image editing device 100B of the second embodiment described above, it is possible to recognize the target text and the display attributes of the target text from the target image, and to impart visual effects to the recognized target text while maintaining the display attributes. This makes it possible to impart intended visual effects to the target text while suppressing the imparting of visual effects to text other than the target text.
[0053] Furthermore, according to the second embodiment of the advertising image editing device 100B, visual effects can be applied to the target text in a manner that corresponds to the display manner of objects other than the target text in the target image, so that a unified visual effect can be applied to the target image without impairing the aesthetics or readability of the target image.
[0054] Third Embodiment FIG. 6 is a diagram illustrating an example of the functional configuration of an advertising image editing device 100C according to the third embodiment. The advertising image editing device 100C according to the third embodiment differs from the advertising image editing device 100B according to the second embodiment in that it includes a foreground object recognition unit 121 instead of the text recognition unit 120, a target object recognition unit 131 instead of the target text recognition unit 130, an attribute information acquisition unit 160C instead of the attribute information acquisition unit 160B, and a visual effect addition unit 140C instead of the visual effect addition unit 140B. The remaining configuration is the same as that of the advertising image editing device 100B. Therefore, the following comprehensively describes the components of the advertising image editing device 100C that are different from those of the advertising image editing device 100B, and omits descriptions of common components as much as possible.
[0055] The foreground object recognition unit 121 has an object recognition function that recognizes a foreground area in which an object is captured within a target image and recognizes the type of object in the foreground area (hereinafter referred to as a "foreground object"). Any object detection model may be used for the object recognition function. For example, R-CNN, YOLO, SSD, etc., based on a CNN (Convolutional Neural Network), can be cited as examples of object detection models. The foreground object recognition unit 121 stores foreground object information 155 relating to the foreground object recognized from the target image in the storage unit 150.
[0056] The target object recognition unit 131 recognizes, from among foreground objects in a target image, a foreground object to which a visual effect is to be applied (hereinafter referred to as a "target object"). The target object recognition unit 131 recognizes the target object based on the target image and foreground object information 155. The target object recognition unit 131 stores, in the storage unit 150, target object information 156 relating to the recognized target object.
[0057] Attribute information acquisition unit 160C acquires attribute information 157 related to a foreground object in a target image. For example, attribute information 157 is information indicating attributes related to the display mode of the foreground object. For example, attribute information 157 includes the state of the foreground object (such as the traveling direction) and information on other objects in contact with the foreground object. Attribute information acquisition unit 160C stores the acquired attribute information 157 in storage unit 150.
[0058] The visual effect applying unit 140C applies a visual effect to a target object in a target image. The visual effect applying unit 140C applies visual information to the target object in the target image based on the target image, target object information 156, and attribute information 157. The visual effect applying unit 140C stores the target image after the visual effect has been applied in the storage unit 150 as an edited target image 158.
[0059] 7 is a flowchart showing an example of the flow of processing in which advertising image editing device 100C of the third embodiment applies a visual effect to a target image. In advertising image editing device 100C, first, target image input unit 110 inputs a target image (S301). Next, foreground object recognition unit 121 performs object recognition processing for recognizing foreground objects in the target image input in S301 (S302). For example, foreground object recognition unit 121 recognizes the coordinates and type of foreground objects in the target image using an object detection model, and outputs foreground object information 155 including the recognized coordinates and type.
[0060] Next, target object recognition unit 131 performs a process of recognizing a target object in the target image based on foreground object information 155 output in S302 (S303). For example, target object recognition unit 131 can analyze the category of a group of foreground objects in the target image based on foreground object information 155, and select a foreground object having a category corresponding to the target object category as the target object. Here, the target object category is a category of the foreground object selected as the target object, and is set in advance in advertising image editing device 100C.
[0061] The category analysis may be a rule-based category analysis based on the type of foreground object, or may be implemented by an LLM such as ChatGPT-4o (registered trademark). For example, the target object recognition unit 131 may provide the foreground object information 155 recognized from the target image to the LLM and have the LLM determine which target object category the type of foreground object belongs to. As a result, for example, if the target object category "vegetables" is set, the target object recognition unit 131 can select, as the target object, an object that corresponds to a vegetable from among the foreground objects in the target image.
[0062] Furthermore, the target object recognition unit 131 may provide foreground object information 155 to the LLM, causing it to directly determine which foreground object a visual effect should be applied to, without going through a category. In this case, the LLM may be provided with information indicating criteria for selecting a foreground object (target object) to which a visual effect is to be applied. Furthermore, the target object recognition unit 131 may be configured to accept a foreground object designation operation by a user, and recognize the designated foreground object as the target object.
[0063] Furthermore, the target object recognition unit 131 may be configured to perform a layout analysis of foreground objects in the target image based on foreground object information 155 recognized from the target image, and to select a target object based on the results of the layout analysis. For example, the target object recognition unit 131 may be configured to recognize and compare the sizes of multiple recognized foreground objects and recognize the largest foreground object as the target object. This method is suitable for images in which the larger an object is to be emphasized, such as advertising images, for example. Furthermore, for example, the target object recognition unit 131 may be configured to recognize, among the multiple recognized foreground objects, one whose layout meets specific conditions as the target object. For example, the target object recognition unit 131 may be configured to recognize, as the target object, one foreground object located in the center of the multiple foreground objects, or one foreground object located at the top or bottom. Furthermore, the target object recognition unit 131 may be configured to infer the attributes of each foreground object by inputting the results of the layout analysis into a model that has learned the relationship between the arrangement of foreground objects and the attributes of the foreground objects, and to recognize foreground objects that have specific attributes such as "headline" or "product name" as target objects. The target object recognition unit 131 outputs target object information 156 that includes the coordinates and type of the target object recognized in this way.
[0064] Next, attribute information acquisition unit 160C acquires attribute information 157 related to the foreground object in the target image (S304). For example, attribute information acquisition unit 160C recognizes the position and type of the foreground object in the target image based on foreground object information 155 output by target object recognition unit 131, and outputs attribute information 157 by performing a display attribute recognition process on the image of the recognized foreground object.
[0065] Next, visual effect imparting unit 140C imparts a visual effect to the target object in the target image based on target object information 156 output in S303 and attribute information 157 output in S304 (S304). For example, visual effect imparting unit 140C generates visual effect information for imparting a visual effect to the target object in the target image so as to maintain the display attributes of the target object. Furthermore, for example, visual effect imparting unit 140C may recognize foreground objects other than the target object based on foreground object information 155 and target object information 156, and generate visual effect information such that the target object is displayed in a manner according to the display attributes of the other foreground objects.
[0066] 8 is an image diagram illustrating an outline of the processing flow for applying a visual effect to a target object in a target image. First, the foreground object recognition unit 121 recognizes foreground objects BS1 and BS2 by performing object recognition processing on the target image IM21. Next, the target object recognition unit 131 recognizes the foreground object BS2 of the recognized foreground objects BS1 and BS2 as the target object. Next, the visual effect applying unit 140C applies a visual effect to the recognized target object BS2. More specifically, the visual effect applying unit 140C generates, for example, information (visual effect information) for applying a visual effect to the target object BS2.
[0067] For example, when a visual effect is applied using a function of a web browser, the visual effect applying unit 140C can generate source code for applying the visual effect as visual effect information. Fig. 8 illustrates an example of source code as visual effect information, in which HTML information CD21 and script information CD22 such as JavaScript are generated. For example, by describing in the HTML information CD21 a process for combining an image of a target object BS2 with a target image, and by describing in the script information CD22 a process for applying a visual effect to the target object BS2, it is possible to display a target image IM23 to which the visual effect has been applied to the target object BS2.
[0068] More specifically, the visual effect imparting unit 140C generates visual effect information, while also generating a target image IM22 in which the target object BS2 has been deleted from the target image IM21, and an image IM24 in which everything except the target object BS2 has been erased (e.g., made transparent) from the target image IM21 (hereinafter referred to as the "target object image"). For example, the visual effect imparting unit 140C can generate the target image IM22 by replacing the image of the target object BS2 in the target image IM21 with an image similar to the surrounding background. Then, the visual effect imparting unit 140C executes the HTML information CD21 and the script information CD22 to combine the target object image IM24 with the target image IM22 and impart a visual effect to the target object image IM24. This allows the display of a target image IM23 with a visual effect on the target object BS2.
[0069] Furthermore, in this case, the visual effect imparting unit 140C can impart to the target object BS2 an appropriate visual effect that corresponds to the display mode of the foreground objects other than the target object BS2, based on the attribute information of the foreground object. For example, the visual effect imparting unit 140C can impart a visual effect to the target object BS2 while suppressing imparting of an unintended visual effect to the foreground object BS1. Furthermore, for example, the visual effect imparting unit 140C can impart a visual effect to the target object BS2 so as not to adversely affect the display of the foreground object BS1 (for example, so that the foreground object BS1 is not hidden by the target object BS2).
[0070] Here, a method for displaying a target image IM23 to which a visual effect is applied using HTML information CD21 and script information CD22 has been described. In this case, the target image IM22 from which the target object BS2 has been deleted, the target object image IM24 from which everything other than the target object BS2 has been erased, the HTML information CD21, and the script information CD22 may be stored in the storage unit 150 as the edited target image 158. On the other hand, the visual effect applying unit 140C may be configured to generate an image of the target object BS2 with a visual effect and combine it with the target image IM22 to generate the image with the visual effect (for example, an image in PNG format) itself as the edited target image 158.
[0071] According to the advertising image editing device 100C of the third embodiment described above, the target object and the display attributes of the target object are recognized from the target image, and a visual effect is imparted to the recognized target object while maintaining the display attributes, thereby making it possible to impart the intended visual effect to the target object while suppressing the imparting of visual effects to foreground objects other than the target object.
[0072] Note that the advertising image editing device 100C of the third embodiment differs from the advertising image editing device 100B of the second embodiment in that the target to which the visual effect is applied is a foreground object rather than text. However, the advertising image editing device 100C is common to both the second embodiment in that it recognizes objects (text and foreground objects) in a target image, recognizes from among them an object to which the visual effect is applied (hereinafter referred to as a "target object"), and applies the visual effect to the recognized target object. Therefore, like the advertising image editing device 100A of the first embodiment, the advertising image editing device 100C of the third embodiment may be configured as a device that does not include the attribute information acquisition unit 160C. Furthermore, the details of the device configuration in this case can be understood in the same way by replacing "text" with "foreground object" in the first embodiment.
[0073] <Modification> When multiple objects (text or foreground objects) can be recognized from the target image, the advertising image editing devices 100A, 100B, and 100C (hereinafter collectively referred to as the advertising image editing device 100) according to the embodiments may store information indicating rules for applying visual effects in advance. In this case, the advertising image editing device 100 may be configured to apply a visual effect to the target object based on attribute information and the rules so that the rules are applied between the display attributes of recognized objects other than the target object and the display attributes of the target object. For example, when applying a visual effect involving movement to the target object, the advertising image editing device 100 may be configured to apply a visual effect to the target object that prevents the target object from obscuring other objects while moving. Applying visual effects that take such consideration into account can improve the appeal of the target image while suppressing a decrease in visibility.
[0074] In the above embodiment, the advertising image editing devices 100A to 100C are described as applying visual effects to advertising images. However, the advertising image editing devices 100A to 100C of the embodiment may be configured as devices that apply visual effects to any image other than advertising images. Furthermore, in the above embodiment, the advertising image editing devices 100A to 100C are described as applying visual effects to images. However, the advertising image editing devices 100A to 100C of the embodiment may be configured as devices that apply visual effects to moving images by applying visual effects to time-series images that constitute the moving image. Furthermore, the moving images to which visual effects are applied may be advertising moving images or moving images other than advertising moving images.
[0075] According to at least one of the embodiments described above, one or more objects are recognized from a target image, and from the one or more recognized objects, a target object to which a visual effect is to be applied is recognized, and the visual effect is applied to the recognized target object, thereby making it possible to apply visual effects to a moving image in a manner that is more in line with the user's intentions.
[0076] For example, in the past, when trying to apply a visual effect to a specific object (text or foreground object), it was necessary to individually specify which object to apply the visual effect to for each video, which was time-consuming. In contrast, in this embodiment, the target object to which the visual effect is applied is automatically selected, so the user can accurately apply the visual effect to the appropriate target object.
[0077] In addition, for example, in the past, there were cases where visual effects were applied to objects other than the target object. In contrast to this, in the present embodiment, visual effects can be applied to target objects after taking into consideration the display attributes of the objects in the target image, so that visual effects with a sense of unity can be applied to the target moving image without impairing aesthetics or readability.
[0078] Furthermore, according to this embodiment, users can easily edit and remake existing videos. For example, when editing an advertising video, users can reduce advertising production costs by reusing existing advertisements. Furthermore, according to this embodiment, even when creating a new advertisement with visual effects, the process of adding visual effects can be shortened, thereby reducing advertising production costs. Furthermore, according to this embodiment, new videos can be easily and accurately generated based on existing videos. This allows users to prototype more advertisements and select advertisements that better suit their purposes from a wider range of options.
[0079] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Industrial Applicability]
[0080] The present invention is applicable to applications in which visual effects are applied to objects in moving images. [Explanation of symbols]
[0081] 100A, 100B, 100C Advertising image editing device 110 Target image input unit 120 Text Recognition Unit 121 Foreground object recognition unit 130 Target text recognition unit 131 Target object recognition unit 140A, 140B, 140C Visual effects section 150 Storage section 151 Text Information 152 Target text information 153, 158 Edited images 154, 157 Attribute information 155 Foreground object information 156 Target object information 160B, 160C Attribute information acquisition section
Claims
1. a target image input unit for inputting a target image to which a visual effect is to be applied; an object recognition unit that recognizes one or more objects from the target image; a target object recognition unit that recognizes a target object to which the visual effect is to be applied from among the one or more objects recognized by the object recognition unit; a visual effect applying unit that applies the visual effect to the target object included in the target image; an attribute information acquisition unit that acquires attribute information related to display attributes of the object from the target image; Equipped with the visual effect applying unit applies the visual effect to the target object based on the attribute information so as to maintain the display attribute of the target object; Image editing equipment.
2. a target image input unit for inputting a target image to which a visual effect is to be applied; an object recognition unit that recognizes one or more objects from the target image; a target object recognition unit that recognizes a target object to which the visual effect is to be applied from among the one or more objects recognized by the object recognition unit; a visual effect applying unit that applies the visual effect to the target object included in the target image; Equipped with the visual effect imparting unit generates a target object image, which is an image of the target object to which the visual effect has been imparted, by imparting the visual effect to the image of the target object, and imparts the visual effect to the target image by superimposing the target object image on the target image; The generation of the target object image is to generate a program for displaying the image of the target object in a manner that has a visual effect. Image editing equipment.
3. the visual effect applying unit applies a visual effect accompanied by movement to the target object included in the target image; 3. The image editing device according to claim 1.
4. an attribute information acquisition unit that acquires attribute information related to display attributes of the object from the target image; the visual effect applying unit applies the visual effect to the target object based on the attribute information so as to maintain the display attribute of the target object; 3. The image editing device according to claim 2.
5. the visual effect applying unit removes the image of the target object from the target image, and superimposes the target object image on the target image from which the image of the target object has been removed.
3. The image editing device according to claim 2.
6. the object is text in the target image, The attribute information includes one or more of the number of characters, the direction, the position, the position of a line break, and the shape of an area of a character string that constitutes the text. The image editing device according to claim 1 .
7. the object is text in the target image, the attribute information includes attributes of decorative objects associated with the text; the visual effect applying unit applies a visual effect to the target object in a manner corresponding to an attribute of the decorative object; The image editing device according to claim 1 .
8. the object is text in the target image, the target object recognition unit performs a layout analysis of a group of texts in the target image, and selects target text, which is text to which a visual effect is to be applied, based on a result of the layout analysis. The image editing device according to claim 1 .
9. the attribute information includes display attributes of the plurality of objects included in the target image, the visual effect applying unit applies a visual effect to the target object in a manner according to display attributes of the plurality of objects; The image editing device according to claim 1 .
10. the visual effect applying unit applies a visual effect to the target object based on the attribute information and a rule regarding the application of the visual effect, such that the rule is applied between a display attribute of an object other than the target object among the objects and a display attribute of the target object. The image editing device according to claim 9.
11. the object is a foreground object in the target image; the attribute information includes one or more of a position, a category, and a shape of a region of the foreground object; The image editing device according to claim 1 .
12. the target image is an advertising image for advertising a product, the object is the product; 12. The image editing device according to claim 1, 2, 9, 10, or 11.
13. a target image input step of inputting a target image to which a visual effect is to be applied; an object recognition step of recognizing one or more objects from the target image; a target object recognition step of recognizing a target object to which the visual effect is to be applied from among the one or more objects recognized in the object recognition step; a visual effect applying step of applying the visual effect to the target object included in the target image; an attribute information acquisition step of acquiring attribute information relating to display attributes of the object from the target image; 1. An image editing method comprising: the visual effect applying step applies the visual effect to the target object based on the attribute information so as to maintain the display attribute of the target object; How to edit images.
14. a target image input step of inputting a target image to which a visual effect is to be applied; an object recognition step of recognizing one or more objects from the target image; a target object recognition step of recognizing a target object to which the visual effect is to be applied from among the one or more objects recognized in the object recognition step; a visual effect applying step of applying the visual effect to the target object included in the target image; 1. An image editing method comprising: the visual effect applying step generates a target object image, which is an image of the target object to which the visual effect has been applied, by applying the visual effect to an image of the target object, and applies the visual effect to the target image by superimposing the target object image on the target image; The generation of the target object image is to generate a program for displaying the image of the target object in a manner that has a visual effect. How to edit images.
15. a target image input step of inputting a target image to which a visual effect is to be applied; an object recognition step of recognizing one or more objects from the target image; a target object recognition step of recognizing a target object to which the visual effect is to be applied from among the one or more objects recognized in the object recognition step; a visual effect applying step of applying the visual effect to the target object included in the target image; an attribute information acquisition step of acquiring attribute information relating to display attributes of the object from the target image; A program for causing a computer to execute the above, the visual effect applying step applies the visual effect to the target object based on the attribute information so as to maintain the display attribute of the target object; program.
16. a target image input step of inputting a target image to which a visual effect is to be applied; an object recognition step of recognizing one or more objects from the target image; a target object recognition step of recognizing a target object to which the visual effect is to be applied from among the one or more objects recognized in the object recognition step; a visual effect applying step of applying the visual effect to the target object included in the target image; A program for causing a computer to execute the above, the visual effect applying step generates a target object image, which is an image of the target object to which the visual effect has been applied, by applying the visual effect to an image of the target object, and applies the visual effect to the target image by superimposing the target object image on the target image; The generation of the target object image is to generate a program for displaying the image of the target object in a manner that has a visual effect. program.
Citation Information
Patent Citations
Virtual gift special effect processing method, virtual gift special effect processing device and live broadcast system
CN110493630A
Information processing apparatus
JP2004112112A
Document production system, document production method, program and storage medium
JP2007172573A
Image generation server using real-time enhancement synthesis technology, image generation system, and method
JP2019009754A
Information processing device, control method, and storage medium
WO2016072117A1