Information processing device, control method for information processing device, and program
The information processing device enhances image modification by identifying subject objects and generating form parameters from user utterances, accurately reflecting user intentions and reducing the time needed for image alteration.
Patent Information
- Application Number
- JP2023204878
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2043-12-04
AI Technical Summary
Existing image generation technologies either fail to accurately reflect user intentions or require time-consuming sequential text input for image modification.
An information processing device that acquires user utterances to identify subject objects and form modification factors, generating parameters for modifying image features based on neural networks, and outputting the modified image to a destination.
The device effectively reproduces the intended shape of an object indicated by user statements, improving accuracy and reducing the time required for image modification.
Smart Images

Figure 0007757379000001 
Figure 0007757379000002 
Figure 0007757379000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, a control method for an information processing device, and a program. [Background technology]
[0002] In recent years, video conferencing has become increasingly popular. One of the benefits of video conferencing is that it can improve the accuracy of information transmission by visually sharing images between multiple users. In such use cases, participants may use the shared image as a basis for further discussion, discussing through conversation what changes should be made to the subject of discussion represented by the image, and then reconcile their views. Against this background, various methods have been proposed for creating new images or modifying existing images based on information from conversations between participants or text information entered by participants. Non-Patent Document 1 discloses a technology that accepts input of text called a prompt, and creates and outputs a new image based on semantic information of the text. Non-Patent Document 2 discloses a technology that accepts input of a base image and text information that modifies the image, and outputs an image that has undergone style conversion so that the image is modified according to the text information. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] R. Rombach, “High-Resolution Image Synthesis with Latent Diffusion Models”, CVPR 2021. [Non-patent document 2] O. Patashnik, “StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery”, ICCV 2021. Summary of the Invention [Problem to be solved by the invention]
[0004] On the other hand, the technology disclosed in Non-Patent Document 1 creates an image that is plausibly modified based on input text information, and therefore does not necessarily create an image that accurately reflects the user's intention. One example of a technology for solving this problem is the technology disclosed in Non-Patent Document 2. However, the technology disclosed in Non-Patent Document 2 requires the user to sequentially prepare text information for modifying an image, which is time-consuming for the user.
[0005] In view of the above problems, the present invention has an object to make it possible to reproduce the shape of an object indicated by the content of a user's statement in a more suitable manner. [Means for solving the problem]
[0006] The information processing device according to the present invention includes: an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance content, and form modification factor information that is information for modifying the form of the object included in the utterance content; Based on the shape modifier information, Among one or more objects included in the image that is the target of the user's remark, an object that corresponds to the subject object information acquired by the acquisition means generating means for generating a form modification parameter indicating a form feature of the The shape modification Parameters Based on To the object a modifying means for making modifications; and an output control means for controlling the output of the result of the modifications made to the object by the modifying means to a predetermined output destination; The feature modification factor information includes at least one of an estimated likelihood of the object being the subject of the comment, whether or not a change can be made to the object, and information on a user who made a comment indicated by the comment content. It is characterized by: [Effects of the Invention]
[0007] According to the present invention, it is possible to reproduce the shape of an object indicated by the content of a user's statement in a more suitable manner. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 illustrates an example of a hardware configuration of an information processing device. [Figure 2] FIG. 2 is a diagram illustrating an example of a functional configuration of an information processing device. [Figure 3] 10 is a flowchart illustrating an example of processing by an information processing device. [Figure 4] FIG. 2 is a diagram illustrating an example of a display screen of an information processing device. [Figure 5] FIG. 2 is a diagram showing an example of a display state of a display screen. [Figure 6] 10 is a flowchart illustrating an example of processing by an information processing device. [Figure 7] FIG. 2 is a diagram showing an example of a display state of a display screen. [Figure 8] 10 is a flowchart illustrating an example of processing by an information processing device. [Figure 9] 10 is a flowchart illustrating an example of processing by an information processing device. [Figure 10] 10 is a flowchart illustrating an example of processing by an information processing device. [Figure 11] FIG. 10 is a schematic diagram showing an example of a detection likelihood map. [Figure 12] FIG. 2 is a diagram illustrating an example of a display screen of an information processing device. [Figure 13] FIG. 2 is a diagram illustrating an example of a functional configuration of an information processing device. [Figure 14] 10 is a flowchart illustrating an example of processing by an information processing device. [Figure 15] FIG. 10 is a diagram showing an example of a result of updating the shape of a subject object. [Figure 16] FIG. 2 is a diagram illustrating an example of a display screen of an information processing device. [Figure 17] FIG. 2 is a diagram illustrating an example of a display screen of an information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0009] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0010] First Embodiment A first embodiment of the present disclosure will be described below. Fig. 1 is a diagram showing an example of the hardware configuration of an information processing device according to this embodiment. The information processing device 100 includes a CPU (Central Processing Unit) 104, a RAM (Random Access Memory) 105, and a ROM (Read Only Memory) 106. The information processing device 100 also includes an input unit 101, a display unit 102, an image input unit 103, and an HDD (Hard Disk Drive) 107. The above-described components of the information processing device 100 are connected via a data bus 108 so as to be able to transmit and receive data to and from each other.
[0011] The CPU 104 reads out a control computer program stored in the ROM 106, loads it into the RAM 105, and executes various control processes based on the program. The RAM 105 is used as an area for loading the program executed by the CPU 104, a temporary storage area such as a work memory, etc. The image input unit 103 serves as an interface for receiving image data from the outside. The image data can be received, for example, from an external device such as an imaging device via a transmission path such as a cable, from another device via a network such as the Internet, or from screen information displayed on a display unit. The HDD 107 stores various data such as image data and setting parameters, as well as various programs.
[0012] Image data received via the image input unit 103 is transmitted to the CPU 104, RAM 105, and ROM 106 via a data bus 108. Furthermore, the CPU 104 executes an information processing program stored in the ROM 106 or HDD 107, thereby realizing information processing on input data (for example, image data). Furthermore, data received from an external device via the image input unit 103 may be stored in the HDD 107 . The input unit 101 serves as an input interface for receiving information input from a user. The input unit 101 may include, for example, input devices such as a keyboard, a pointing device such as a mouse, and a touch panel. The input unit 101 may also include an audio input device such as a microphone. The display unit 102 serves as an output interface for presenting information to the user and may include a display device such as a liquid crystal display.
[0013] An example of the functional configuration of the information processing device according to this embodiment will be described with reference to Fig. 2. The information processing device 100 according to this embodiment includes an image acquisition unit 201, an utterance data acquisition unit 202, a text generation unit 203, a subject object extraction unit 204, a feature parameter generation unit 205, a feature update unit 206, and an output control unit 207.
[0014] The image acquisition unit 201 acquires image data to be processed. The image data to be processed may be image data acquired by an external device such as an imaging device, image data stored in a storage device such as a hard disk, or image data received via a network such as the Internet. The image acquisition unit 201 outputs the acquired image data to the subject object extraction unit 204.
[0015] The utterance data acquisition unit 202 acquires utterance data indicating the content of utterances made by users (speakers). The utterance data includes information indicating the users (speakers) as metadata (hereinafter also referred to as user information), and is data indicating the content of utterances made by the users as text information. There is no particular limit to the number of users from which utterance data is acquired, as long as there is one or more users. Furthermore, if there are multiple target users, it is preferable to acquire utterance data for each user. In this case, it is possible to identify, based on the user information, which user's utterance content each piece of utterance data indicates. The method for acquiring text information indicating the content of a user's utterances included in utterance data is not particularly limited. For example, the text information indicating the content of a user's utterances may be acquired based on text data input by the user via an input device such as a keyboard. As another example, the user's voice input via a sound collection device such as a microphone may be converted into text data, and the text information indicating the content of a user's utterances may be acquired based on the text data. Furthermore, the text information indicating the user's utterances using the various input interfaces exemplified above may be acquired in real time in synchronization with the information processing device, or may be acquired by reading out pre-stored data. Furthermore, the method of acquiring data (e.g., text data, audio data, etc.) from which the text information indicating the user's utterances is acquired is not particularly limited. For example, the target data may be acquired from an external device connected via a transmission path such as a cable, or may be acquired from an external device via a network. Furthermore, in addition to text information indicating the content of a user's utterance and user information, the utterance data may also include, as metadata, information indicating the position of a target utterance in a series of utterances (for example, a position in the time series or a position in the context, etc.). In this way, by including, as metadata, information indicating the position of a target utterance in a series of utterances, it becomes possible to divide the content of a user's utterance indicated by the utterance data into phrases and manage them together with the order of each phrase. Note that, hereinafter, for convenience, various explanations will be given assuming that the utterance time is applied as information indicating the position of a target utterance in a series of utterances. The utterance data acquisition unit 202 outputs the acquired utterance data to the text generation unit 203 .
[0016] The text generating unit 203 generates main object information and form modifier information based on the utterance data. The main object information and form modifier information will be explained in more detail below.
[0017] The subject object information is information about an object (e.g., an object in an image) mentioned as the subject of the utterance content indicated by the utterance data. The subject object information may be configured as structured data including information such as the name of the object that is the subject of the utterance content, the estimated likelihood of the object being the subject, and the corresponding utterance time in the utterance data. As another example, the subject object information may be configured as a list including one or more of the above structured data. Note that, as will be described in detail later, when the subject object extraction unit 204 uniquely determines a subject object, it is preferable to refer to the estimated likelihood of the object being the subject in the context indicating the user's utterance content.
[0018] The form modification factor information is information about modifications to the form of the subject object mentioned in the utterance content indicated by the utterance data, and may be, for example, information about suggestions for modifying the form of the subject object. Modifying a form corresponds to, for example, adding, duplicating, deleting, moving, transforming, or other processing related to the shape or appearance such as color or texture. The feature modifier information may include, for example, a summary name of the feature modifier, the name of the linked subject object, an estimated likelihood as a feature modifier, the time of the corresponding utterance in the utterance data, user information (the proposer), a status regarding whether or not the feature modification can be implemented, attributes of the feature modifier, etc. The feature modifier information may also be a list made up of structured data including each of the above-mentioned information. A feature modifier is associated with one subject object. A subject object may also be associated with multiple feature modifiers. For example, it is generally expected that changes will be made to multiple locations on a feature. In such cases, feature modifiers corresponding to the changes made to each of the multiple locations will be associated with the feature of the subject object.
[0019] The estimated likelihood of a form modifier is the likelihood that the target expression (the stated information) is an expression generally related to modifying a form. For example, if the form modifier is the expression "round the corners," it will be output as a higher likelihood value because it is generally interpreted as an expression related to changing a form. On the other hand, if the form modifier is the expression "turn quietly," it will be output as a lower likelihood value because it is not generally interpreted as an expression that describes or modifies a form.
[0020] The state regarding whether or not a form modification can be implemented is information indicating whether or not a form modification can be implemented in response to a statement regarding the propriety of a form modification in a series of statements (e.g., a series of statements made in a discussion between users). For example, suppose that a first user proposes a form change to a certain subject object, and then a second user makes a statement indicating that the form change is not acceptable. In such a situation, the text generator 203 extracts a form modification factor indicating the form change based on the statement by the first user, and also extracts a judgment of whether or not the form change is appropriate based on the statement by the second user. The state regarding whether or not a form modification can be implemented may be a boolean value of 0 / 1, or may be a numerical value.
[0021] The attribute of a feature modification factor is information that indicates which of multiple modifications, including addition, deletion, and transformation, the target modification process will make to the feature of the subject object. The attribute of a feature modification factor is referenced, for example, when the feature parameter generation unit 205, which will be described later, guides the process to different modifications for each attribute.
[0022] The text generator 203 may be implemented, for example, by a natural language generation model based on a neural network that can process a sufficiently long context. By applying such a configuration, it becomes possible to perform a text-to-text conversion task, such as generating topic object information and form modifier information from input utterance data. Of course, the above is merely an example, and the method is not particularly limited as long as it is possible to extract and generate information corresponding to topic object information and form modifier information from the utterance content indicated by the input utterance data. Of the generated series of data, the text generating unit 203 outputs the main object information to the main object extracting unit 204, and outputs the feature modifier information to the feature parameter generating unit.
[0023] The subject object extraction unit 204 extracts an object indicated by the subject object information from the image represented by the image data, based on the image data acquired by the image acquisition unit 201 and the subject object information generated by the text generation unit. In the following description, the object indicated by the subject object information will also be referred to as a subject object for convenience. The subject object extraction result by the subject object extraction unit 204 includes, for example, an image of the subject object included in the image represented by the image data (a partial image of the subject object's area) and location information of the location of the subject object within the image. The method for extracting the image of the subject object from the image represented by the image data is not particularly limited. For example, a method of cutting out the subject object along the outline of the region by segmentation may be applied, or a method of cutting out a rectangle containing the region of the subject object may be applied. The position information of the subject object may be specified as the position information of a rectangle containing the region of the subject object, for example.
[0024] When the subject object information is a list consisting of multiple structured data, the subject object extraction unit 204 uniquely identifies subject object information that is most likely to be included as metadata in the subject object information. The subject object extraction unit 204 may then output a subject object extraction result based on the identified subject object information. By applying this control, when there are multiple subject object candidates, it becomes possible to extract the object that is most likely to be the subject object from among the multiple candidates, and filter out other objects (exclude them from the extraction targets). For example, the subject object extraction unit 204 may determine that there is no subject object if all estimated likelihoods included in the subject object information are smaller than a certain value (threshold). As another example, the subject object extraction unit 204 may determine that there is no subject object if the statistical value of likelihoods at the time of segmentation in the subject object extraction result is smaller than a certain value (threshold). This process is designed to handle cases where no subject object is included in the utterance data, and may apply to situations where there is no relevance (or very low relevance) between the target image data and the user's utterance. As a specific example, this may apply to situations where a conversation that is unrelated to the target image data is taking place between multiple users from which utterance data is to be acquired. The subject object extraction unit 204 may be implemented, for example, by a neural network-based model capable of searching for and extracting objects corresponding to specified text from within an image. Furthermore, the model applied to the subject object extraction unit 204 is preferably a model trained on a large-scale set of natural language data and image data pairs, in order to achieve high extraction accuracy even with zero-shot extraction for a wide variety of subject objects. "Zero-shot" refers to executing a task for labels with new classes that the model has never learned before. The main subject object extraction unit 204 outputs the extraction results of the main subject object to the feature parameter generation unit 205 and the feature update unit 206 .
[0025] The feature parameter generating unit 205 generates feature modification parameters based on the feature modification factor information generated by the text generating unit 203 and the result of extraction of the subject object by the subject object extracting unit 204 . A form modification parameter is a parameter that represents a form feature based on modification factor information. Specifically, a form modification parameter is structured data that includes geometric information such as the position, size, shape, and orientation of a form feature that is generated based on a form modification factor, as well as appearance information such as color and texture. Furthermore, a form modification parameter may inherit some or all of the structured information of the form modification factor from which it is generated. The feature parameter generator can be realized by a multimodal input neural network that receives, for example, text and images as input and outputs the structured data described above.
[0026] Among the form modification parameters that express form features, the position, size, etc. are restricted by the shape of the subject object to be input, particularly by the contour information, so it is desirable to control the parameters using the subject object image. This point will be explained below with specific examples. Assume that a feature modification factor "add bolt holes" is input for both a subject object with a width of 30 mm and a subject object with a width of 1000 mm. Because the feature modification factor does not include a specific modifier related to size, the size of the feature modification parameter is set to, for example, "normal," which indicates a generally applicable size. Meanwhile, a user typically expects the diameter of a "bolt hole" to be approximately 10 mm or less for a feature with a width of 30 mm and approximately 50 mm or less for a feature with a width of 1000 mm. However, it is difficult for the feature update unit 206 (described later) to accommodate the variation in the size of the feature feature to be generated, between 10 mm and 50 mm, in response to the size instruction of the feature modification parameter "normal." In consideration of this situation, the feature parameter generation unit 205 may control the feature modification parameters to be generated using information such as the subject object image.
[0027] The feature parameter generation unit 205 may execute different tasks depending on the attributes of the input feature modification factor information. For example, suppose feature modification factor information having an attribute such as "add bolt holes to flat plate portion" indicating the addition of a feature feature to the feature of the subject object is input. In this case, the feature parameter generation unit 205 may execute, for example, a task of detecting flat plate portions from the subject object image and a task of outputting map information of locations where bolt holes can be added, and determine feature modification parameters such as the position and size of the bolt holes to be added. As another example, suppose feature modification factor information having an attribute such as "remove bolt holes from flat plate portion" indicating the removal of a feature feature from the feature of the subject object is input. In this case, the feature parameter generation unit 205 may execute, for example, only the task of detecting bolt holes in flat plate portions from the subject object image. The feature parameter generating unit 205 outputs the generated feature modification parameters to the feature updating unit 206 .
[0028] The feature update unit 206 performs at least one of a number of modifications, including adding, modifying, and deleting feature features, on the subject object extracted by the subject object extraction unit 204 based on the feature modification parameters generated by the feature parameter generation unit 205. The feature updater 206 may be implemented, for example, by a diffusion neural network model for image generation that is trained to minimize the distance in feature space between the original image and the generated image in order to preserve the feature features of the original subject object. The feature update unit 206 then outputs to the output control unit 207 the feature of the main subject object that has been updated by making changes to the main subject object based on the feature modification parameters (hereinafter also referred to as the updated feature).
[0029] The output control unit 207 controls so that information indicating the form of the main subject object changed by the form update unit 206 (that is, the updated form of the main subject object) is output to a predetermined output destination. For example, the output control unit 207 may display the updated form of the main subject object in a predetermined display area (e.g., the display unit 102). At this time, the output control unit 207 may display, in the predetermined display area, an image indicated by the image data acquired by the image acquisition unit 201, the utterance data acquired by the utterance data acquisition unit 202, and the updated form of the main subject object. In this case, the output control unit 207 may selectively switch whether or not to display each of the above-mentioned pieces of information based on a user instruction. The output control unit 207 may also perform interactive processing based on a user instruction via the input unit 101, targeting the updated form of the main subject object displayed in the predetermined display area. An example of this interactive processing will be described in detail later. Furthermore, the above is merely an example and does not limit the destination to which the information indicating the updated form of the subject object is output by the output control unit 207. As a specific example, the output control unit 207 may output information indicating the updated form of the subject object to a device that performs various image processing or various analytical processing on the target data, as the target of the image processing or analytical processing.
[0030] An example of processing by the information processing device 100 according to this embodiment will be described with reference to FIG. In S101, the image acquisition unit 201 acquires image data to be processed. In S102, the utterance data acquiring unit 202 acquires utterance data indicating the content of a user's utterance input via a predetermined input interface. Note that the utterance data acquiring unit 202 is not limited to acquiring utterance data based on the utterances of a single user, but may acquire utterance data based on the utterances of each of a plurality of users. In S103, the text generating unit 203 generates main object information and form modifier information based on the utterance data acquired in S102. In S104, the subject object extraction unit 204 extracts the subject object indicated by the subject object information in the image indicated by the image data, based on the image data acquired in S101 and the subject object information generated in S103. In S105, the feature parameter generating unit 205 generates feature modification parameters based on the feature modification factor information generated in S103 and the result of the subject object extraction in S104. In S106, the feature update unit 206 modifies the subject object extracted in S104 by adding, changing, or deleting features based on the feature modification parameters generated in S105, thereby updating the feature of the subject object. In S107, the output control unit 207 controls so that information indicating the form of the main subject object changed in S106 (i.e., the updated form of the main subject object) is output to a predetermined output destination. As a specific example, the output control unit 207 may present the updated form of the main subject object to the user by displaying the updated form of the main subject object in a predetermined display area.
[0031] 4, an example of a screen that the information processing device 100 according to this embodiment presents to the user via the display unit 102 will be described. The display screen 400 shown in FIG. 4 includes a target image display unit 410, a main object image display unit 420, an update result display unit 430, a comment content display unit 440, and a form modification factor display unit 450.
[0032] The target image display section 410 is a display area where an image indicated by the image data acquired by the image acquisition section 201 is displayed. In the example shown in Fig. 4, four objects, namely, an electrical box 301a, a sphere 301b, a reinforcing plate 301c, and a work desk 301d, are captured as subjects in the image. For convenience, in the example shown in Fig. 4, it is assumed that the electrical box 301a has been determined as the main object based on a utterance from a user whose input was previously accepted. In addition, in the example shown in Fig. 4, in order to clearly indicate the extraction result of the main object, a rectangular box 310 (a so-called bounding box) is displayed in the image displayed in the target image display section 410 so as to encompass the area of the electrical box 301a. The target image display section 410 corresponds to an example of a first partial region.
[0033] The utterance content display unit 440 is a display area in which the utterance content of the user indicated by the utterance data acquired by the utterance data acquisition unit 202 is displayed. In the example shown in Fig. 4, utterances 341a and 341b corresponding to two pieces of utterance data, respectively, are displayed in a chronological order, with the older one being presented at the top and the newer one being presented below. As described above, the utterance data includes text information indicating the utterance content and information indicating the speaker of the target utterance. Each of the utterances 341a and 341b is displayed based on the text information indicating the utterance content included in the target utterance data. In the example shown in Figure 4, three proposals are made in comment 341a: "Add the drawn shape to the center of the surface as large as possible in the space created by moving the round hole," "Move the rightmost round hole as far to the right as possible," and "Add a hemming bend to the lower bend." In response to these proposals, comment 341b agrees with the first and second proposals, but rejects the third proposal.
[0034] The form modifier display section 450 is a display area where form modifiers indicated by form modifier information generated based on utterance data are displayed. First, the subject object name included in the subject object information determined as the target is displayed in the display area 351 in the form modifier display section 450. In addition, the summary name included in the form modifier information related to the subject object is displayed in the form modifier display section 450. 4, summary names 352a, 352b, and 352c corresponding to three pieces of form modification factor information are displayed. Also, the results of estimation by the text generating unit 203 regarding whether or not form modifications (changes made to an object) can be made for each of the three pieces of form modification factor information are reflected in check boxes 452a, 452b, and 452c displayed in association with each piece of form modification factor information. In the example shown in FIG. 4, the proposal for "add a hemming bend to the bottom bend" indicated by abstract name 352c has been rejected by comment 341b. Therefore, the implementation status of the form modifier for this proposal is set to "No," and check box 452c is displayed as OFF. On the other hand, the proposals corresponding to abstract names 352a and 352b have been agreed upon by comment 341b. Therefore, the implementation status of the form modifier for these proposals is set to "Yes," and check boxes 452a and 452b are displayed as ON. The ON / OFF state of each of the check boxes 452a, 452b, and 452c can be arbitrarily changed based on an instruction from the user via the input unit 101. A change in the state of the form modification factors via these check boxes regarding whether they can be implemented may be used as a trigger to execute a process of modifying the updated form of the subject object displayed in the update result display unit 430.
[0035] The subject object image display section 420 is a display area in which an image of the subject object extracted from the acquired image data is displayed based on the subject object extraction result. In the example shown in Fig. 4, an image of an electrical box determined as the subject object is extracted from the image represented by the acquired image data and displayed. The subject object image display section 420 corresponds to an example of a second partial area.
[0036] The update result display section 430 is a display area that displays an image of the shape of the main subject object (the updated shape of the main subject object) that has been modified by adding, changing, or deleting shape features from the main subject object based on the shape modification factors. In the example shown in Fig. 4, the results of two modifications to the main subject object image 320, "moving the rightmost round hole as far to the right as possible" and "adding an aperture shape as large as possible to the center of the surface," are displayed as the updated shape 330. The update result display section 430 corresponds to an example of a third partial area.
[0037] It should be noted that the display state of each of the components (display sections denoted by reference numerals 410 to 450) that make up the display screen 400 may be individually and arbitrarily switched ON / OFF in response to an instruction from the user.
[0038] Fig. 5 is a schematic diagram showing an example of the display state of target image display section 410 on display screen 400 shown in Fig. 4. Specifically, Fig. 5 shows a schematic diagram of a situation in which, in addition to objects 301a to 301d that are candidates for the subject of each user's comment, a finger 302 is being imaged as another object. Using the state shown in Fig. 5 as an example, an example of the processing by the subject object extraction unit 204 will be described below with reference to Fig. 6. In order to further improve the estimation accuracy of the subject object in the image, the subject object extraction unit 204 uses feature amounts that can be extracted from the image, such as spatial feature amounts and modality feature amounts.
[0039] In S111, the image acquisition unit 201 receives input of image data to be processed from a user. The utterance data acquisition unit 202 acquires utterance data indicating the content of the user's utterance. The text generation unit 203 generates subject object information and form modification factor information based on the utterance data acquired by the utterance data acquisition unit 202.
[0040] In S112, the subject object extraction unit 204 extracts the subject object indicated by the subject object information from the image indicated by the image data (hereinafter also referred to as the input image) based on the image data and subject object information acquired in S111. For example, in the example shown in Fig. 5, if the name of the acquired subject object is "box," the subject object extraction unit 204 estimates each of objects 301a and 301c as a subject object candidate with a higher likelihood than other objects. In such a case, the subject object extraction unit 204 may use various feature amounts, such as spatial feature amounts and modality feature amounts, to more accurately and uniquely determine the subject object. In the example shown in Fig. 6, the subject object extraction unit 204 uses spatial feature amounts and modality feature amounts to uniquely determine the subject object.
[0041] The spatial feature is a feature based on the idea that the main object is larger, more central, and more clearly (out of focus) in the image. In the example shown in Fig. 6, the main object extraction unit 204 is assumed to have a spatial feature encoder (not shown). In S113, the subject object extraction unit 204 inputs subject object information to the spatial feature encoder, thereby acquiring spatial features as the output of the spatial feature encoder.
[0042] The modality feature is a feature based on the idea that when an indicator object that indicates a subject object exists in an image, the modality of the indicator object is used to accurately identify the subject object. In the example shown in Fig. 6, the subject object extraction unit 204 is assumed to have a modality feature encoder (not shown). In S114, the subject object extraction unit 204 detects the finger 302 as the pointing object. In this case, there is a high possibility that a subject object exists at the position or direction pointed to by the finger 302 detected as the pointing object. Therefore, in S115, the subject object extraction unit 204 uses a modality feature encoder to acquire, as a modality feature, a subject object existence probability map that outputs a high score for a specific position or direction based on the finger gesture, for example.
[0043] In S116, the subject object extraction unit 204 uniquely determines a subject object based on the subject object information acquired in S111, the spatial feature amount acquired in S113, and the modality feature amount acquired in S115. 5, object 301a is larger than object 301c in image capture and is located in an area where the probability of the subject object being present is higher according to the modality of fingers, so it is more likely to be the subject object. Therefore, in this case, of objects 301a and 301c, which are subject object candidates, the subject object extraction unit 204 determines object 301a to be the subject object. The above is merely an example, and the feature quantities used to determine the main object are not particularly limited as long as they can be extracted from the target image. As a specific example, the image quality of each of a series of objects included in the image (e.g., a feature quantity that serves as an index for evaluating image quality) may be used as the feature quantity. In this case, for example, of the series of objects included in the image, the object with the highest image quality may be determined as the main object.
[0044] In S117, the subject object extraction unit 204 outputs the extraction result of the subject object (object 301a) determined in S116, including the position information of the subject object, to a predetermined output destination. Note that the position information of the subject object may be output as, for example, the position information of a rectangle (bounding box) that contains the subject object.
[0045] By applying the above-described control, it is expected that the accuracy of extracting the subject object will be further improved.
[0046] 7, another example of the display state of the target image display section 410 on the display screen 400 will be described. In addition to the objects 301a to 301d, rectangles 311a and 311b containing objects 301a and 301c, respectively, which are candidates for the subject object based on the subject object information, are displayed in the image displayed on the target image display section 410. The user can select one of the objects 301a and 301c, which are candidates for the subject object, as the subject object by using a point 401 via the input section. Using the state shown in Figure 7 as an example, and referring to Figure 8, we will explain another example of the processing of the subject object extraction unit 204, where the subject object extraction unit 204 uniquely determines a subject object via user input.
[0047] In S121, the image acquisition unit 201 receives input of image data to be processed from a user. The utterance data acquisition unit 202 acquires utterance data indicating the content of the user's utterance. The text generation unit 203 generates subject object information and form modification factor information based on the utterance data acquired by the utterance data acquisition unit 202.
[0048] In S122, the subject object extraction unit 204 extracts the subject object indicated by the subject object information from the image indicated by the image data (hereinafter also referred to as the input image) based on the image data and subject object information acquired in S121. For example, in the example shown in Figure 7, if the name of the acquired subject object is "box", the subject object extraction unit 204 estimates each of objects 301a and 301c as a candidate subject object with a higher likelihood than other objects.
[0049] In S123, the output control unit 207 causes the target image display unit 410 to display rectangles 311a and 311b containing the objects 301a and 301c, respectively, that are the candidate main object, based on the estimation result of the candidate main object in S122. In S124, the subject object extraction unit 204 receives from the user a selection of either the rectangles 311a or 311b displayed in the target image display area 410 by the user operating the point 401 via the input unit. In S125, the subject object extraction unit 204 identifies the object corresponding to the rectangle selected in S124 as the subject object, and outputs the extraction result of the subject object, including the position information of the subject object, to a predetermined output destination. Note that the position information of the subject object may be output as, for example, the position information of a rectangle (bounding box) that contains the subject object.
[0050] By applying the above-described control, it becomes possible to more reliably identify the subject object intended by the user.
[0051] An example of the processing of the text generation unit 203 will be described, focusing particularly on the processing related to the extraction of form modifiers, with reference to Fig. 9. In the example shown in Fig. 9, the text generation unit 203 filters out form modifiers with low likelihood of being more faithful to the user's intention.
[0052] In S131, the text generation unit 203 receives utterance data (for example, utterance data acquired by the utterance data acquisition unit 202) as input. In S132, the text generating unit 203 generates subject object information based on the utterance data received as input in S131. In S133, the text generation unit 203 generates form modifier information based on the utterance data received as input in S131. The generated form modifier information includes an estimated likelihood as a form modifier. As a specific example, if two form modifiers, "move the round hole to the right" and "turn it quietly," are extracted from the utterance data, the former will have a higher likelihood as a form modifier, and the latter will have a lower likelihood. In S134, the text generating unit 203 compares the likelihood included in each piece of form modification factor information generated in S133 with a preset threshold, and deletes form modification factor information whose likelihood is lower than the threshold. In S135, if there is any form modification factor information that has not been deleted as a result of the processing in S134, the text generating unit 203 outputs the form modification factor information to a predetermined output destination.
[0053] As described above, filtering of form modification factor information by the text generation unit 203 omits subsequent processing of form modification factor information deleted from the processing target, which is expected to reduce the processing load and improve the processing speed. Furthermore, the information displayed in the form modification factor display section 450 on the display screen 400 is simplified, which is expected to improve user convenience.
[0054] On the other hand, the above-described method may also extract words and phrases from utterance data that are less relevant to the topic but generally modify the shape. For example, suppose the topic object is an "electrical box," and the text generator 203 generates shape modification factor information corresponding to two shape modification factors: "move the round hole to the right" and "add tapered edges to the gear." In this example, even if the topic is an "electrical box" and the shape modification factor "add tapered edges to the gear" is unrelated to the topic, a shape modification factor derived from a word clearly related to shape modification may be extracted with a higher likelihood. In light of this situation, another example of a method for filtering shape modification factors is proposed below.
[0055] Fig. 10 is a flowchart showing an example of the processing of the feature update unit 206 according to this embodiment. In the example shown in Fig. 10, the feature update unit 206 filters feature modifiers with a lower estimated likelihood as feature modifiers in order to update the feature of the subject object more faithfully to the user's intention (to make changes to the feature). FIG. 11 is a schematic diagram showing an example of a detection likelihood map generated by the feature update unit 206 according to this embodiment. The detection likelihood map indicates where in an image a detection target exists and with what probability. The detection likelihood map will be described in detail later. In the example shown in FIG. 11, a subject object image 320 includes round holes 321a to 321c and flat surfaces 322a and 321b. In the example shown in FIG. 11, the subject object is named "electrical box," and FIGS. 11(a), 11(b), and 11(c) show summary names of different feature modification parameters. Specifically, FIG. 11(a) corresponds to an example of a feature modification parameter indicating "moving a hole," FIG. 11(b) corresponds to "adding a taper to the center of a surface," and FIG. 11(c) corresponds to "adding a taper to the edge of a gear."
[0056] In S141, the feature update unit 206 receives the feature object extraction result including the feature object image and the feature modification parameters as input. The feature modification parameters inherit the feature modification factor information and also include the summary name and attributes included in the feature modification factor information.
[0057] In S142, the feature update unit 206 executes a detection task for the feature to be processed, regardless of whether the attribute of the feature modification parameter is added, changed, or deleted, and generates a detection likelihood map based on the execution result of the detection task. For example, in the example shown in Fig. 11(a), the feature update unit 206 detects each of the round holes 321a to 321c from the subject object image for the feature modification parameter "move hole." At this time, the detection likelihood of the areas corresponding to the round holes 321a to 321c is set to a higher value than that of other areas. In the figure, hatched areas indicate areas with a higher detection likelihood (e.g., areas where the detection likelihood is equal to or greater than a threshold), and unhatched areas indicate areas with a lower detection likelihood (e.g., areas where the detection likelihood is less than a threshold). 11(b), the feature update unit 206 detects each of the flat surfaces 322a and 322b from the subject object image in response to the feature modification parameter "add aperture shape to center of surface." At this time, the detection likelihood of the areas corresponding to the flat surfaces 322a and 322b is set to a higher value than that of other areas. 11(c), the feature update unit 206 executes a process to detect gears from the main object image for the feature modification parameter "add tapered edges of gears." On the other hand, in the example shown in FIG. 11, since no gears exist, a lower value is set as the detection likelihood across the entire main object image.
[0058] In S143, the feature update unit 206 calculates a predetermined statistical quantity (e.g., a maximum value) based on each of the series of detection likelihood maps generated in S142, and if the statistical quantity is less than or equal to a threshold value, deletes the feature modification parameter corresponding to the detection likelihood map. In S144, if there are any shape modification parameters that have not been deleted as a result of the processing in S143, the shape update unit 206 updates the shape of the subject object based on those shape modification parameters and outputs the results of the update to a specified output destination.
[0059] Through the above series of processes, the subject object image is matched with the feature modification parameters, and a filtering process is performed using the statistics of the feature modification parameter detection likelihood map. This prevents the output of a feature to which feature modification factors not originally intended by the user have been applied, and makes it possible to update the feature of the subject object more faithfully to the user's intention.
[0060] Next, an example of a method for generating form modification parameters that more accurately reflect the user's intentions through interactive processing via the display unit 102 and the input unit 101 will be described below.
[0061] FIG. 12 is a diagram showing another example of a screen presented to a user via the display unit 102 by the information processing device 100 according to this embodiment. The display screen 400 shown in each of FIGS. 12(a) to 12(d) includes an update result display unit 430 and a feature modification factor display unit 450. A pointer 401 is displayed in the feature modification factor display unit 450, and the pointer 401 is used to designate a portion of the display screen 400 based on a user operation received via the input unit 101. In the example shown in FIG. 12, three feature modification factors 352a to 352c are displayed in the feature modification factor display unit 450. The feature modification factor 352a indicates a feature modification of "moving the rightmost hole." The feature modification factor 352b indicates a feature modification of "adding a large drawn shape to the center of the surface." The feature modification factor 352c indicates a feature modification of "adding a hemming bend to the lower bend."
[0062] In the example shown in FIG. 12(a), the user can select any feature modification factor in the feature modification factor display section 450. For example, in the example shown in FIG. 12(a), it is assumed that the user selects a feature modification factor 352a indicating the feature modification "move the rightmost hole." In this case, a command list 353a for executing predefined processes for the selected feature modification factor 352a is displayed. In the example shown in FIG. 12(a), commands for applying processes such as "move," "duplicate," "delete," and "change shape" to feature features based on the feature modification factor are displayed in the command list 353a. Furthermore, more detailed conditions may be specified for at least some commands. For example, in the example shown in Fig. 12(a), a text box 354a is displayed for the "change shape" command to accept specification of a new shape as text information. With this configuration, it becomes possible to enjoy usability that is substantially the same as when using the information processing device in which shapes are automatically generated based on the above-mentioned utterance data.
[0063] The example shown in FIG. 12(b) is a schematic diagram illustrating an example of a state in which a command indicating "move" is selected by the user for a shape modification factor 352a indicating a shape modification of "move the rightmost hole." When the command indicating "move" is selected, the position of the round hole 331, which is a shape feature based on the target shape modification factor 352a, can be moved by a drag operation using the pointer 401, for example. In the example shown in FIG. 12(b), the round hole 331' of the shape 330 of the updated subject object shows the round hole 331 after it has been moved by the above operation.
[0064] The example shown in Figure 12(c) is a schematic diagram illustrating an example of a state in which a command indicating "change shape" has been selected by the user for a shape modification factor 352a indicating a shape modification of "move the rightmost hole." When the command indicating "change shape" is selected, the outline of the round hole 331, which is a shape feature based on the target shape modification factor 352a, can be moved by a drag operation using the pointer 401, for example. In the example shown in Figure 12(c), the round hole 331'' of the shape 330 of the updated subject object shows the round hole 331 after its size has been changed by pulling the outline of the round hole 331 outward.
[0065] The example shown in FIG. 12(d) schematically illustrates another example of a state in which a command indicating "change shape" is selected by the user for a shape modification factor 352a indicating a shape modification of "move the rightmost hole." When the command indicating "change shape" is selected, an instruction to change the shape of the round hole 331 can be given via the text box 354a. In the example shown in FIG. 12(d), an instruction to double the size of the round hole 331 is given via the text box 354a. Furthermore, the round hole 331''' of the updated subject object shape 330 shows the round hole 331 after its size has been changed based on the instruction entered in the text box 354a.
[0066] By applying the above-described controls, the user can set geometric feature information, such as the position, size, shape, and orientation of feature information based on feature modification factors, to the desired state. While the above description focuses primarily on the case where geometric feature information is modified, the subject of modification is not limited to geometric feature information. As a specific example, appearance information, such as color and texture, can also be modified by applying substantially the same controls as those described above.
[0067] <Second embodiment> A second embodiment of the present disclosure will be described below. In this embodiment, an example of a configuration will be described in which changes made to a subject object based on form modification factor information are stored in a predetermined storage area, and the information stored in the storage area is made available for subsequent use. In addition, this embodiment will be described focusing on parts that are particularly different from the first embodiment described above, and detailed description of parts that are substantially similar to the first embodiment will be omitted.
[0068] An example of the functional configuration of the information processing device according to this embodiment will be described with reference to Fig. 13. The information processing device according to this embodiment differs from the information processing device according to the first embodiment described with reference to Fig. 2 in that it includes a storage unit 208. The storage unit 208 is a storage area for storing various types of data. The data stored in the storage unit 208 includes two main types of data. The first type of data is learning data for learning the correlation between form modification factors and form modification parameters for each user so that the information processing device 100 can more accurately reflect the user's intentions when making changes to the subject object. The second type of data is history data indicating the results of updates that are saved as history so that the user can quickly check the results of updates to the form of the subject object when form features are added, changed, or deleted from the subject object.
[0069] First, an example of the processing of the information processing device according to this embodiment will be described, focusing on the processing related to saving the learning data, with reference to Fig. 14. Fig. 14 shows an example of a series of processing executed after the processing related to generating feature parameters by the feature parameter generating unit 205.
[0070] In S211, the feature parameter generating unit 205 generates feature modification parameters based on the subject object image and feature modification factor information received as input. In S212, the feature update unit 206 modifies the feature of the subject object so that it has features according to the feature modification parameters generated in S211. In S213, the output control unit 207 displays the result of updating the form of the main subject object in a predetermined display area in S212. As a specific example, the output control unit 207 may cause the display unit 102 to display the display screen 400, and cause the update result display area 450 of the display screen 400 to display the result of updating the form of the main subject object.
[0071] In S214, the feature parameter generation unit 205 accepts a user operation related to modification of the feature characteristics of the subject object as a result of updating the feature of the subject object displayed in the predetermined display area in S213. At this time, the feature parameter generation unit 205 may accept an operation related to modification of the feature characteristics of the subject object from the user through interactive processing via the display unit 102 and the input unit 101, as described above with reference to Fig. 12 . In S215, the feature parameter generation unit 205 executes processing related to the modification of the feature modification parameters generated in S211 so as to follow the operation related to the modification of the feature features of the subject object received from the user in S214. As a specific example, suppose that a feature "add hole" is added, the "shape" item of the feature modification parameter in the corresponding feature modification factor information is round, and the user modifies the shape of the target hole to a rectangle. In this case, the feature parameter generating unit 205 changes the "shape" item of the feature modification parameter to a rectangle.
[0072] In S216, the form parameter generating unit 205 determines whether the user operation is completed. If the feature parameter generation unit 205 determines in S216 that the user operation has not been completed, it proceeds to S214. In this case, the processes of S214 and S215 are executed again. In this way, the feature parameter generation unit 205 accepts user operations related to modification of the feature features of the subject object until the user operation is completed, and each time, it executes processing related to modification of the feature modification parameters in accordance with the user operation. Then, if the form parameter generating unit 205 determines in S216 that the user operation has been completed, it advances the processing to S217.
[0073] In S217, the output control unit 207 stores in the storage unit 208, as learning data, a set of the feature modification parameters that have been corrected by the processes of S214 to S216 and the feature modification factors that are the source of generation of the feature modification parameters. The learning data generated in the above manner is divided into data for each user and then used for learning by the feature parameter generation unit 205, which enables the feature parameter generation unit 205 to generate feature modification parameters that are close to the user's intentions.
[0074] The process described with reference to FIG. 14 will now be described in more detail with reference to FIG. 15, using a specific example. Figure 15(a) schematically shows a subject object image to be processed. It is assumed that a user makes a statement instructing "add aperture" to the subject object image shown in Figure 15(a). In this case, based on the statement data indicating the content of the user's statement, the result of changes made to the shape of the subject object is displayed in a predetermined display area, as shown in Figure 15(b). Then, the user further performs an operation to modify the aperture shape from the state shown in Figure 15(b), and the result of the modification of the aperture shape reflected in the shape of the subject object is displayed, as shown in Figure 15(c). In Fig. 15(b), the form parameter generating unit 205 generates form modification parameters indicating that the aperture shape is round, and the result of modifying the form characteristics indicated by the form modification parameters for the form of the subject object is displayed. On the other hand, since the form characteristics have been modified as shown in Fig. 15(c), it is clear that the aperture shape intended by the user was rectangular in the example shown in Fig. 15. In response to this result, the storage unit 208 stores the subject object image, form modification factors, and form modification parameters (form modification parameters reflecting modifications based on user instructions) as learning data. By using this learning data for training the feature parameter generation unit 205, it is possible to have the feature parameter generation unit 205 learn that when an instruction is given to add an aperture shape to a box-shaped part, the shape of the feature modification parameter is rectangular.
[0075] An example of a display screen of the information processing device according to this embodiment will be described with reference to Fig. 16. It is assumed here that the feature parameter generating unit 205 that has been trained based on the training data described with reference to Fig. 15 is applied. 16, the display screen 400 includes an update result display section 430 and a feature modifier display section 450. The update result display section 430 displays the subject object "reinforcement plate," and the feature modifier display section 450 displays an instruction to "add aperture." In response to the feature modifier instruction, the update result display section 430 also displays a marker 336 indicating the center position for generating the feature, and a feature candidate window 335 for selecting the shape, size, and orientation of the feature to be applied from among the candidates. The feature candidate window 335 also displays, for each feature candidate to be selected, a likelihood corresponding to an estimated result of the likelihood that the candidate will be selected by the target user. In the example shown in FIG. 16, the form parameter generation unit 205 has been trained using the training data generated based on the conditions shown in FIG. 15, and therefore in response to the instruction to "add aperture", the estimated likelihood of an aperture having a rectangular shape is the highest. In this way, candidates with higher likelihoods estimated by the feature parameter generation unit 205, which has been trained using learning data for each user, are displayed with higher priority in the feature candidate window 335. By applying this type of control, the user can more easily select a feature candidate that is closest to their intention from the series of candidates presented in the feature candidate window 335.
[0076] Next, the history data stored in the storage unit 208 according to the update results of the form of the subject object will be described. FIG. 17 is a diagram showing an example of a display screen of the information processing device according to this embodiment. FIGS. 17(a) to 17(c) are diagrams showing the states of the display screen 400 in chronological order, with FIG. 17(a) showing the oldest state and FIG. 17(c) showing the newest state. The display screen 400 shown in FIG. 17 includes an update result display section 430 and a history display section 460.
[0077] First, the state shown in Fig. 17(a) will be described. In the state shown in Fig. 17(a), only the form modification factor 361a is selected as the target to be applied. Therefore, the update result display section 430 displays the form 330a of the subject object that has been updated by reflecting the form feature indicated by the form modification factor 361a. At this time, the history display section 460 also reflects information indicating that the form modification factor 361a has been applied. In this case, the updated form 330a of the subject object and the form modification factor 361a are linked and stored in the storage section 208 as history data. 17, the selection state of each form modifying factor or form modifying parameter is displayed by the selection state of a check box in the history display section 460. Furthermore, by operating the check box, the state can be switched between a selected state and a non-selected state, thereby selectively switching whether or not the target form modifying factor or form modifying parameter is applied.
[0078] Next, the state shown in FIG. 17(b) will be described. In the state shown in FIG. 17(b), the form modification factors 361a and 361b and the morphological modification parameters 361a' and 361a'' have been selected as the objects to be applied. Therefore, the update result display section 430 displays a form 330c of the subject object that has been updated by reflecting the morphological features indicated by the form modification factors 361a and 361b and the morphological modification parameters 361a' and 361a''. At this time, the history display section 460 also reflects information indicating that the form modification factors 361a and 361b and the morphological modification parameters 361a' and 361a'' have been applied. In this case, the updated feature 330c of the subject object, the feature modification factors 361a and 361b, and the feature modification parameters 361a' and 361a'' are linked together and stored in the storage unit 208 as history data.
[0079] Next, the state shown in FIG. 17(c) will be described. In the state shown in FIG. 17(c), the checkboxes for the form modifiers and form modifier parameters other than the form modifier 361a are deselected. Therefore, the update result display unit 430 displays a form 330a' of the subject object that has been updated by reflecting the form features of only the form modifier 361a among the series of form modifiers and form modifier parameters presented in the history display unit 460. At this time, the output control unit 207 checks whether history data of the updated form 330a to which only the form modifier 361a is linked is stored in the storage unit 208. Note that in the example shown in FIG. 17(a), history data of the updated form 330a to which only the form modifier 361a is linked is stored in the storage unit 208. Therefore, the output control unit 207 displays the updated feature 330 a on the update result display unit 430 based on the history data stored in the storage unit 208 without going through the feature update unit 206 .
[0080] A possible use case of this embodiment is when a user compares and examines the shape of a target object by selectively switching on and off the application of various previously specified shape modifiers, and then confirms the updated shape of the object. Even in such a case, this embodiment eliminates the need to regenerate the corresponding shape if history data is saved, thereby speeding up the process of displaying the updated shape of the subject object. Furthermore, because the updated shape of the subject object is displayed based on the history data, it is possible to ensure the reproducibility of the shape of the subject object according to the selected state of the shape modifiers.
[0081] <Other embodiments> Although the exemplary embodiments have been described above in detail, the present invention can be embodied as, for example, a system, an apparatus, a method, a program, a recording medium (storage medium), etc. Specifically, the present invention may be applied to a system made up of multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may be applied to an apparatus made up of a single device. Needless to say, the object of the present invention can be achieved by the following: A recording medium (or storage medium) on which program code (computer program) of software that realizes the functions of the above-described embodiments is recorded is supplied to a system or device. The storage medium is, of course, a computer-readable storage medium. The computer (or CPU or MPU) of the system or device then reads and executes the program code stored on the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-described embodiments, and the recording medium on which the program code is recorded constitutes the present invention.
[0082] The disclosure of this embodiment also includes the following configurations, methods, and programs. (Configuration 1) An acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance, and form modification factor information which is information for modifying the form of the object included in the utterance; a modification means for modifying, based on the form modification factor information, an object corresponding to the subject object information acquired by the acquisition means, among one or more objects included in an image that is the subject of the user's utterance; and an output control means for controlling the results of the modifications made to the object by the modification means to be output to a predetermined output destination. An information processing device comprising: (Configuration 2) An information processing device as described in Configuration 1, characterized in that it has a generation means for generating form modification parameters indicating form characteristics of the object based on the form modification factor information, and the modification means makes modifications to the object corresponding to the subject object information based on the form modification parameters. (Configuration 3) The information processing device described in Configuration 2, characterized in that the modifications made by the modification means include at least one of adding the geometric feature to the object corresponding to the subject object information based on the geometric modification parameters, transforming at least a part of the object based on the geometric feature, and deleting the geometric feature from the object. (Configuration 4) The information processing device described in configuration 2 or 3, characterized in that the shape modification parameters include at least one of shape information including at least one of position, size, shape, and orientation, and appearance information including at least one of color and texture. (Configuration 5) The information processing device described in any one of configurations 2 to 4, characterized in that the form modification factor information includes, as an attribute, information indicating which of multiple modifications, including addition, deletion, and transformation, is to be made to the object corresponding to the subject object information, and the generation means controls the configuration of the form modification parameters to be generated according to the attribute included in the form modification factor information. (Configuration 6) An information processing device described in any one of configurations 2 to 5, characterized in that the form modification factor information includes at least one of the name of the object that is the subject of the comment content, the estimated likelihood of the object being the subject, whether or not changes can be made to the object, and information about the user who made the comment indicated by the comment content. (Configuration 7) The information processing device described in Configuration 6, characterized in that the modification means suppresses modifications based on the form modification factor information to the object corresponding to the subject object information when the estimated likelihood contained in the form modification factor information is less than a threshold value. (Configuration 8) The information processing device described in Configuration 6 is characterized in that the modification means is a likelihood map indicating where and with what probability an object corresponding to the subject object information exists in the image that is the subject of the user's remarks, and when a predetermined statistical quantity obtained from the likelihood map based on the estimated likelihood contained in the form modification factor information is less than a threshold, the information processing device suppresses modifications to the object corresponding to the subject object information based on the form modification factor information. (Configuration 9) An information processing device as described in Configuration 2, characterized in that it has a storage means for correlating form modification factor information with form modification parameters generated based on the form modification factor information and storing them as learning data, and the generation means generates form modification parameters based on user information included in the form modification factor information received as input and the learning data stored by the storage means. (Configuration 10) The information processing device described in any one of configurations 2 to 9, characterized in that the output control means controls the results of changes made to the object by the change means to be displayed in a predetermined display area, the display area including a first partial area in which an image that is the subject of the user's remarks is displayed, a second partial area in which an image of an object corresponding to the subject object information is displayed, and a third partial area in which an image corresponding to the results of changes made to the object by the change means is displayed, and the information processing device is configured to be able to selectively switch whether or not to display images in each of the first partial area, the second partial area, and the third partial area. (Configuration 11) The information processing device described in Configuration 10 is characterized in that the third partial area is configured to be able to accept instructions from the user regarding changes to the object corresponding to the subject object information, and the generation means controls the form modification parameters applied to make changes to the object corresponding to the subject object information in accordance with the instructions accepted by the third partial area from the user. (Configuration 12) The information processing device described in configuration 10 or 11, characterized in that the display area includes a fourth partial area configured to be able to accept selection of at least one of the acquired series of form modification factor information, and the generation means controls whether or not to add changes to the object corresponding to the subject object information based on each of the series of form modification factor information, depending on the selection state of each of the series of form modification factor information accepted via the fourth partial area. (Configuration 13) The fourth partial area is configured to be able to accept designation of a process for modifying a form feature to be applied to the selected form modification factor information, The information processing device described in configuration 12, characterized in that the generation means controls the feature modification parameters applied to make changes to the object corresponding to the subject object information based on the feature modification factor information selected via the fourth partial area and the processing specified for the feature modification factor information. (Configuration 14) An information processing device as described in any one of configurations 2 to 13, characterized in that it comprises a storage means for sequentially storing as a history the results of changes made to the object by the change means in accordance with the results of acquisition of the form modification factor information by the acquisition means, and a reception means for receiving instructions regarding whether or not to apply each of the series of form modification factor information acquired by the acquisition means, and the output control means, when the results of changes made to the object by the change means based on one or more form modification factor information for which an instruction to apply has been accepted by the reception means, is stored as the history, controls so that the results indicated in the history are output to a predetermined output destination. (Configuration 15) The information processing device described in Configuration 14 is characterized in that the receiving means receives instructions regarding whether to apply each of one or more processes for modifying a geometric feature, which are specified for at least one of the one or more pieces of shape modification factor information for which an instruction to apply has been received, and the output control means, when the results of changes made to the object by the modifying means based on the one or more pieces of shape modification factor information for which an instruction to apply has been received by the receiving means and one of the one or more processes for which an instruction to apply has been received, controls so that the results indicated in the history are output to a predetermined output destination. (Configuration 16) The information processing device described in Configuration 15 is characterized in that, if the results of changes made to the object by the change means based on the one or more pieces of form modification factor information for which an instruction to apply has been accepted by the acceptance means and a form modification parameter from the series of form modification parameters for which an instruction to apply has been accepted, are not stored as the history, the change means makes changes to the object corresponding to the subject object information based on the one or more pieces of form modification factor information and the form modification parameter for which an instruction to apply has been accepted, and the output control means controls so that the results of changes made to the object by the change means are output to a predetermined output destination. (Configuration 17) An information processing device described in any one of configurations 1 to 16, characterized in that the utterance data includes at least one of text data input by the user and text data generated based on the recognition results of the speech uttered by the user. (Configuration 18) An information processing device described in any one of configurations 1 to 17, characterized in that the modification means extracts an object corresponding to the subject object information from the image based on at least one of spatial features, image quality, and modality of one or more objects contained in the image that is the subject of the user's remarks, and makes modifications to the extracted object based on the form modification factor information. (Configuration 19) The information processing device described in Configuration 18 is characterized in that, when there are multiple candidate objects corresponding to the subject object information, the change means extracts at least some of the multiple candidate objects as objects corresponding to the subject object information based on instructions from the user. (Method 1) A control method for an information processing device, comprising: an acquisition step of acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance content, and form modification factor information which is information that modifies the form of the object included in the utterance content; a modification step of making changes based on the form modification factor information to an object corresponding to the subject object information acquired in the acquisition step, among one or more objects included in an image that is the subject of the user's utterance; and an output control step of controlling so that the results of the changes made to the object in the modification step are output to a predetermined output destination. (Program 1) A program for causing a computer to function as an information processing device, characterized by having an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating the object that is the subject of the utterance, and form modification factor information which is information that modifies the form of the object included in the utterance; a modification means for making changes based on the form modification factor information to one or more objects included in an image that is the subject of the user's utterance, which object corresponds to the subject object information acquired by the acquisition means; and an output control means for controlling the results of the changes made to the object by the modification means to be output to a predetermined output destination. [Explanation of symbols]
[0083] 100 Information processing device 203 Text Generation Unit 206 Feature update section 207 Output control section
Claims
1. an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance content, and form modification factor information which is information that modifies the form of the object included in the utterance content; a generating means for generating, based on the form modification factor information, form modification parameters indicating form features of an object corresponding to the subject object information acquired by the acquiring means, among one or more objects included in the image that is the target of the user's remark; modifying means for modifying the object based on the feature modification parameters; an output control means for controlling the output of the result of the modification made to the object by the modification means to a predetermined output destination; and The form modification factor information includes at least one of an estimated likelihood of the object being the subject of the discussion, whether or not a change can be made to the object, and information on the user who made the comment indicated by the comment content. An information processing device characterized by:
2. The modification performed by the modification means is based on the feature modification parameters, adding said feature to an object corresponding to said subject object information; deforming at least a portion of the object based on the geometrical features; and Removing said feature from the object. Contains at least one of the following:
2. The information processing device according to claim 1, wherein:
3. The information processing device according to claim 1 , characterized in that the form modification parameters include at least one of form information including at least one of position, size, shape, and orientation, and appearance information including at least one of color and texture.
4. The feature modifier information includes, as an attribute, information indicating which of a plurality of modifications, including addition, deletion, and transformation, is to be made to the object corresponding to the subject object information; The generating means controls the configuration of the feature modification parameters to be generated in accordance with the attributes included in the feature modification factor information.
2. The information processing device according to claim 1, wherein:
5. The information processing device according to claim 1 , characterized in that the modification means suppresses modifications based on the form modification factor information to the object corresponding to the subject object information when the estimated likelihood contained in the form modification factor information is less than a threshold value.
6. The information processing device described in claim 1, characterized in that the modification means is a likelihood map indicating where and with what probability an object corresponding to the subject object information exists in the image that is the subject of the user's comment, and if a predetermined statistical amount obtained from the likelihood map based on the estimated likelihood included in the form modification factor information is less than a threshold, the modification means suppresses modifications to the object corresponding to the subject object information based on the form modification factor information.
7. a storage means for storing, as learning data, form modification factor information and form modification parameters generated based on the form modification factor information in association with each other; The generating means generates morphological modification parameters based on user information included in the morphological modification factor information received as input and the learning data stored by the storing means.
2. The information processing device according to claim 1, wherein:
8. the output control means controls the display of the result of the change made to the object by the change means in a predetermined display area; The display area is a first partial area in which an image that is the subject of the user's utterance is displayed, a second partial area in which an image of an object corresponding to the subject object information is displayed, and a third partial area in which an image according to the result of the change made to the object by the change means is displayed, The display of an image in each of the first partial area, the second partial area, and the third partial area can be selectively switched.
2. The information processing device according to claim 1, wherein:
9. the third partial area is configured to be able to receive, from the user, an instruction relating to a change to an object corresponding to the subject object information; The generating means controls the feature modification parameters to be applied to the third partial region to modify the object corresponding to the subject object information in accordance with an instruction received from the user.
9. The information processing device according to claim 8, wherein:
10. the display area includes a fourth partial area configured to be able to accept a selection of at least one of the acquired series of feature modification factor information, The generating means controls whether or not to add a change to the object corresponding to the subject object information based on each of the series of feature modifier information, depending on the selection state of each of the series of feature modifier information received via the fourth partial area.
9. The information processing device according to claim 8, wherein:
11. the fourth partial area is configured to be able to accept designation of a process to be applied to the selected feature modification factor information for modifying a feature feature; The generating means controls the feature modification parameters to be applied to make changes to the object corresponding to the subject object information, based on the feature modification factor information selected via the fourth partial area and the processing specified for the feature modification factor information. The information processing device according to claim 10 .
12. a storage means for sequentially storing, as a history, the results of modifications made to the object by the modification means in accordance with the results of acquisition of the form modification factor information by the acquisition means; a receiving means for receiving an instruction regarding whether or not each of the series of feature modification factor information acquired by the acquiring means is applicable; and When a result of a change made to the object by the change means based on one or more pieces of form modification factor information for which an instruction to apply has been accepted by the acceptance means is stored as the history, the output control means controls so that the result indicated by the history is output to a predetermined output destination.
2. The information processing device according to claim 1, wherein:
13. the receiving means receives an instruction regarding whether or not to apply one or more processes for modifying a feature, the processes being specified for at least one of the one or more feature modification factor information for which an instruction to apply the process has been received; When a result of a change made to the object by the change means based on the one or more pieces of form modification factor information for which an instruction to apply has been accepted by the acceptance means and a process for which an instruction to apply has been accepted among the one or more processes is stored as the history, the output control means controls so that the result indicated by the history is output to a predetermined output destination.
13. The information processing device according to claim 12.
14. The change means is If the result of the modification made to the object by the modifying means based on the one or more pieces of feature modification factor information for which an instruction to apply has been accepted by the accepting means and a feature modification parameter for which an instruction to apply has been accepted from the series of feature modification parameters is not stored as the history, modifying the object corresponding to the subject object information based on the one or more feature modifier information and the feature modifier parameters for which the instruction to apply has been received; The output control means controls the output of the result of the change made to the object by the change means to a predetermined output destination.
14. The information processing device according to claim 13,
15. 2. The information processing device according to claim 1, wherein the utterance data includes at least one of text data input by the user and text data generated based on a recognition result of a voice uttered by the user.
16. The information processing device according to claim 1, characterized in that the modification means extracts an object corresponding to the subject object information from the image based on at least one of spatial characteristics, image quality, and modality of one or more objects included in the image that is the subject of the user's comment, and modifies the extracted object based on the form modification factor information.
17. 17. The information processing device according to claim 16, wherein the modification means, when there are multiple candidate objects corresponding to the subject object information, extracts at least some of the multiple candidate objects as the object corresponding to the subject object information based on instructions from a user.
18. A control method for an information processing device, comprising: an acquisition step of acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance content, and form modification factor information that is information that modifies the form of the object included in the utterance content; a generating step of generating, based on the feature modification factor information, feature modification parameters indicating feature features of an object corresponding to the subject object information acquired in the acquiring step, among one or more objects included in the image that is the target of the user's remarks; a modifying step for modifying the object based on the feature modification parameters; an output control step of controlling the result of the change made to the object in the change step to be output to a predetermined output destination; Including, The form modification factor information includes at least one of an estimated likelihood of the object being the subject of the discussion, whether or not a change can be made to the object, and information on the user who made the comment indicated by the comment content.
10. A method for controlling an information processing device, comprising:
19. Computer, an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance content, and form modification factor information which is information that modifies the form of the object included in the utterance content; a generating means for generating, based on the form modification factor information, form modification parameters indicating form features of an object corresponding to the subject object information acquired by the acquiring means, among one or more objects included in the image that is the target of the user's remark; modifying means for modifying the object based on the feature modification parameters; an output control means for controlling the output of the result of the modification made to the object by the modification means to a predetermined output destination; and The form modification factor information includes at least one of an estimated likelihood of the object being the subject of the discussion, whether or not a change can be made to the object, and information on the user who made the comment indicated by the comment content. A program for causing an information processing device to function as the information processing device.
Citation Information
Patent Citations
Text-based real image editing with diffusion models
JP2024154427A