Information processing device, method for controlling information processing device, and program

The information processing apparatus addresses the challenge of accurately reflecting user intentions in video conferencing by extracting subject object information and form modification factors from speech data and applying them to images in real-time, enhancing the alignment of visual changes with spoken instructions.

WO2025121215A1PCT designated stage expired Publication Date: 2025-06-12CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041927
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-11-27
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing video conferencing techniques struggle to accurately reflect a user's intentions in modifying images based on spoken content, with existing methods requiring sequential text input for image modification.

Method used

An information processing apparatus that acquires speech data to extract subject object information and form modification factors, and then applies these factors to an image to modify the subject object, allowing for real-time visual representation of intended changes.

Benefits of technology

Enables more accurate and efficient reproduction of image modifications intended by the user, improving the alignment of visual changes with spoken instructions without the need for sequential text input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041927_12062025_PF_FP_ABST
    Figure JP2024041927_12062025_PF_FP_ABST
Patent Text Reader

Abstract

In the present invention, a text generation unit 203 acquires, on the basis of speech data indicating user speech content, theme object information indicating an object that is the theme of the speech content, and feature modification factor information, which is information for modifying a feature of the object included in the speech content. On the basis of the feature modification factor information, a feature update unit 206 adds a change to an object corresponding to the theme object information acquired by the text generation unit 203, among one or more objects included in an image toward which the speech of the user is directed. An output control unit 207 performs control so that a result obtained by the addition of the change to the object by the feature update unit 206 is output to a predetermined output destination.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, control method for information processing device, and program

[0001] The present disclosure relates to an information processing device, a control method for an information processing device, and a program.

[0002] In recent years, video conferencing has become increasingly popular. One advantage of using video conferencing is that it can be expected to improve the accuracy of information transmission by visually sharing images between multiple users. In such use cases, participants may use a shared image as a basis for further discussion, discussing what changes to make to the subject of discussion represented by the image, and then reconcile their views. Against this background, various methods have been proposed for creating new images or modifying existing images based on information from conversations between participants or text information entered by participants. Non-Patent Document 1 discloses a technology that accepts input of text called a prompt, creates a new image based on the semantic information of the text, and outputs it. Non-Patent Document 2 discloses a technology that accepts input of a base image and text information to modify the image, and outputs an image that has been style-converted so that the image is modified according to the text information.

[0003] R. Rombach, “High-Resolution Image Synthesis with Latent Diffusion Models”, CVPR 2021. O. Patashnik, “StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery”, ICCV 2021.

[0004] On the other hand, the technology disclosed in Non-Patent Document 1 creates an image that is plausibly modified based on input text information, and therefore does not necessarily create an image that accurately reflects the user's intention. One example of a technology for solving this problem is the technology disclosed in Non-Patent Document 2. However, the technology disclosed in Non-Patent Document 2 requires the user to sequentially prepare text information for modifying an image, which is time-consuming for the user.

[0005] In view of the above problems, the present invention has an object to make it possible to reproduce the shape of an object indicated by the content of a user's statement in a more suitable manner.

[0006] The information processing device of the present invention is characterized by having an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating the object that is the subject of the utterance, and form modification factor information which is information that modifies the form of the object included in the utterance; a modification means for making changes based on the form modification factor information to one or more objects included in an image that is the subject of the user's utterance, which object corresponds to the subject object information acquired by the acquisition means; and an output control means for controlling the results of the changes made to the object by the modification means to be output to a predetermined output destination.

[0007] According to the present invention, it is possible to reproduce the shape of an object indicated by the content of a user's statement in a more suitable manner.

[0008] 1 is a diagram showing an example of a hardware configuration of an information processing device. FIG. 1 is a diagram showing an example of a functional configuration of an information processing device. FIG. 2 is a flowchart showing an example of processing of the information processing device. FIG. 2 is a diagram showing an example of a display screen of the information processing device. FIG. 3 is a diagram showing an example of a display state of the display screen. FIG. 4 is a flowchart showing an example of processing of the information processing device. FIG. 4 is a diagram showing an example of a display state of the display screen. FIG. 5 is a flowchart showing an example of processing of the information processing device. FIG. 6 is a flowchart showing an example of processing of the information processing device. FIG. 7 is a schematic diagram showing an example of a detection likelihood map. FIG. 8 is a schematic diagram showing an example of a detection likelihood map. FIG. 9 is a diagram showing an example of a display screen of the information processing device. FIG. 10 is a diagram showing an example of a display screen of the information processing device. FIG. 11 is a diagram showing an example of a display screen of the information processing device. FIG. 12 is a diagram showing an example of a display screen of the information processing device. FIG. 13 is a diagram showing an example of a functional configuration of the information processing device. FIG. 14 is a flowchart showing an example of processing of the information processing device. FIG. 15 is a diagram showing an example of an update result of the shape of a subject object. FIG. 16 is a diagram showing an example of a display screen of the information processing device. FIG. 17 is a diagram showing an example of a display screen of the information processing device.

[0009] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0010] First Embodiment A first embodiment of the present disclosure will be described below. Fig. 1 is a diagram showing an example of the hardware configuration of an information processing apparatus according to this embodiment.

[0011] The information processing device 100 includes a CPU (Central Processing Unit) 104, a RAM (Random Access Memory) 105, and a ROM (Read Only Memory) 106. The information processing device 100 also includes an input unit 101, a display unit 102, an image input unit 103, and an HDD (Hard Disk Drive) 107. The components of the information processing device 100 described above are connected via a data bus 108 so as to be able to transmit and receive data to and from each other.

[0012] The CPU 104 reads out a control computer program stored in the ROM 106, loads it into the RAM 105, and executes various control processes based on the program. The RAM 105 is used as an area for loading the program executed by the CPU 104, a temporary storage area such as a work memory, etc.

[0013] The image input unit 103 serves as an interface for receiving image data from an external device. The image data may be received, for example, from an external device such as an imaging device via a transmission path such as a cable, from another device via a network such as the Internet, or from screen information displayed on a display unit.

[0014] The HDD 107 stores various data such as image data and setting parameters, as well as various programs.

[0015] Image data received via the image input unit 103 is transmitted to the CPU 104, RAM 105, and ROM 106 via a data bus 108.

[0016] Furthermore, the CPU 104 executes an information processing program stored in the ROM 106 or HDD 107, thereby realizing information processing of input data (for example, image data).

[0017] Furthermore, data received from an external device via the image input unit 103 may be stored in the HDD 107 .

[0018] The input unit 101 serves as an input interface for receiving information input from a user. The input unit 101 may include, for example, input devices such as a keyboard, a pointing device such as a mouse, and a touch panel. The input unit 101 may also include an audio input device such as a microphone.

[0019] The display unit 102 serves as an output interface for presenting information to the user, and may include a display device such as a liquid crystal display.

[0020] An example of the functional configuration of the information processing device according to this embodiment will be described with reference to Fig. 2. The information processing device 100 according to this embodiment includes an image acquisition unit 201, an utterance data acquisition unit 202, a text generation unit 203, a main subject object extraction unit 204, a feature parameter generation unit 205, a feature update unit 206, and an output control unit 207.

[0021] The image acquisition unit 201 acquires image data to be processed. The image data to be processed may be image data acquired by an external device such as an imaging device, image data stored in a storage device such as a hard disk, or image data received via a network such as the Internet. The image acquisition unit 201 outputs the acquired image data to the subject object extraction unit 204.

[0022] The utterance data acquisition unit 202 acquires utterance data indicating the content of utterances made by users (speakers). The utterance data includes information indicating the users (speakers) as metadata (hereinafter also referred to as user information), and is data indicating the content of utterances made by the users as text information. There is no particular limit to the number of users from which utterance data is acquired, as long as there is one or more users. Furthermore, if there are multiple target users, it is preferable to acquire utterance data for each user. In this case, it is possible to identify, based on the user information, which user's utterance content each piece of utterance data indicates.

[0023] The method for acquiring text information indicating the content of a user's utterances included in utterance data is not particularly limited. For example, the text information indicating the content of a user's utterances may be acquired based on text data input by the user via an input device such as a keyboard. As another example, the user's voice input via a sound collection device such as a microphone may be converted into text data, and the text information indicating the content of a user's utterances may be acquired based on the text data.

[0024] Furthermore, the text information indicating the user's utterances using the various input interfaces exemplified above may be acquired in real time in synchronization with the information processing device, or may be acquired by reading out pre-stored data. Furthermore, the method of acquiring data (e.g., text data, audio data, etc.) from which the text information indicating the user's utterances is acquired is not particularly limited. For example, the target data may be acquired from an external device connected via a transmission path such as a cable, or may be acquired from an external device via a network.

[0025] Furthermore, in addition to text information indicating the content of a user's utterance and user information, the utterance data may also include, as metadata, information indicating the position of a target utterance in a series of utterances (e.g., a chronological position, a position in a context, etc.), such as the time of utterance. By including, as metadata, information indicating the position of a target utterance in a series of utterances, the content of a user's utterance indicated by the utterance data can be divided into phrases and managed together with the order of each phrase. Hereinafter, for convenience, various explanations will be given assuming that the time of utterance is used as information indicating the position of a target utterance in a series of utterances.

[0026] The utterance data acquisition unit 202 outputs the acquired utterance data to the text generation unit 203 .

[0027] The text generator 203 generates main object information and form modifier information based on the utterance data. The main object information and form modifier information will be described in more detail below.

[0028] The subject object information is information about an object (e.g., an object in an image) mentioned as the subject of the utterance content indicated by the utterance data. The subject object information may be configured as structured data including information such as the name of the object that is the subject of the utterance content, an estimated likelihood of the object being the subject, and the corresponding utterance time in the utterance data. As another example, the subject object information may be configured as a list including one or more of the above structured data. Note that, as will be described in detail later, when the subject object extraction unit 204 uniquely determines a subject object, it is preferable to refer to the estimated likelihood of the object being the subject within the context indicating the user's utterance content.

[0029] The form modification factor information is information about modifications to the form of the subject object mentioned in the utterance content indicated by the utterance data, and may be, for example, information about suggestions for modifying the form of the subject object. Modifying a form corresponds to, for example, adding, duplicating, deleting, moving, transforming, or other processing related to the shape or appearance such as color or texture.

[0030] The feature modifier information may include, for example, a summary name of the feature modifier, the name of the linked subject object, an estimated likelihood of the feature modifier being a feature modifier, the time of the corresponding utterance in the utterance data, user information (the person who proposed it), a status regarding whether or not the feature modification can be implemented, attributes of the feature modifier, etc. The feature modifier information may also be a list made up of structured data including each of the above-mentioned information.

[0031] A feature modifier is linked to one subject object. A subject object may also be linked to multiple feature modifiers. For example, it is generally expected that changes will be made to multiple locations on a feature. In such cases, feature modifiers corresponding to the changes made to each of the multiple locations will be linked to the feature of the subject object.

[0032] The estimated likelihood as a form modifier is the likelihood that the target expression (uttered information) is an expression generally related to modifying a form. For example, if the form modifier is the expression "round the corners," it is generally interpreted as an expression related to changing a form, and therefore the likelihood will be output as a higher value. On the other hand, if the form modifier is the expression "turn quietly," it is generally interpreted as not an expression that describes or modifies a form, and therefore the likelihood will be output as a lower value.

[0033] The state regarding whether or not a form modification can be implemented is information indicating whether or not a form modification can be implemented in response to a statement regarding the propriety of a form modification in a series of statements (e.g., a series of statements made in a discussion between users). For example, suppose a first user proposes a form change to a certain subject object, and then a second user makes a statement indicating that the form change is unacceptable. In such a situation, the text generator 203 extracts a form modification factor indicating the form change based on the statement by the first user, and also extracts a judgment regarding whether or not the form change is acceptable based on the statement by the second user. The state regarding whether or not a form modification can be implemented may be a boolean value of 0 / 1 or a numerical value.

[0034] The attribute of a feature modification factor is information that indicates which of multiple modifications, including addition, deletion, and transformation, the target modification process will make to the feature of the subject object. The attribute of a feature modification factor is referenced, for example, when the feature parameter generation unit 205, which will be described later, guides the process to different modifications for each attribute.

[0035] The text generator 203 may be implemented, for example, by a natural language generation model based on a neural network that can process a sufficiently long context. By applying such a configuration, it becomes possible to perform a text-to-text conversion task, such as generating topic object information and form modifier information from input utterance data. Of course, the above is merely an example, and any method can be used as long as it is possible to extract and generate information corresponding to topic object information and form modifier information from the utterance content indicated by the input utterance data.

[0036] Of the generated series of data, the text generating unit 203 outputs the main object information to the main object extracting unit 204, and outputs the feature modifier information to the feature parameter generating unit.

[0037] The subject object extraction unit 204 extracts an object indicated by the subject object information from the image represented by the image data, based on the image data acquired by the image acquisition unit 201 and the subject object information generated by the text generation unit. In the following description, for convenience, the object indicated by the subject object information will also be referred to as a subject object. The subject object extraction result by the subject object extraction unit 204 includes, for example, an image of the subject object included in the image represented by the image data (a partial image of the subject object's area) and location information of the location of the subject object within the image.

[0038] The method for extracting the image of the subject object from the image represented by the image data is not particularly limited. For example, a method of cutting out the subject object along the outline of the region by segmentation may be applied, or a method of cutting out a rectangle containing the region of the subject object may be applied. The position information of the subject object may be specified as the position information of a rectangle containing the region of the subject object, for example.

[0039] When the subject object information is a list consisting of multiple structured data, the subject object extraction unit 204 uniquely identifies subject object information that is most likely to be included as metadata in the subject object information. The subject object extraction unit 204 may then output a subject object extraction result based on the identified subject object information. By applying such control, when there are multiple subject object candidates, it becomes possible to extract the object that is most likely to be the subject object from among the multiple candidates, and filter out other objects (exclude them from the extraction targets).

[0040] For example, the subject object extraction unit 204 may determine that there is no subject object if all estimated likelihoods included in the subject object information are smaller than a certain value (threshold). As another example, the subject object extraction unit 204 may determine that there is no subject object if the statistical value of likelihoods at the time of segmentation in the subject object extraction result is smaller than a certain value (threshold). This process assumes a case where no subject object is included in the utterance data, and may apply to a situation where there is no relevance (or very low relevance) between the target image data and the user's utterance. As a specific example, this may apply to a situation where a conversation that is unrelated to the target image data is taking place between multiple users from whom utterance data is to be acquired.

[0041] The subject object extraction unit 204 may be implemented, for example, by a neural network-based model capable of searching for and extracting objects corresponding to specified text from within an image. Furthermore, the model applied to the subject object extraction unit 204 is preferably a model trained on a large-scale set of natural language data and image data pairs, in order to achieve high extraction accuracy even with zero-shot extraction for a wide variety of subject objects. "Zero-shot" refers to executing a task for a label with a new class that the model has never learned.

[0042] The main subject object extraction unit 204 outputs the extraction results of the main subject objects to the feature parameter generation unit 205 and the feature update unit 206 .

[0043] The feature parameter generating unit 205 generates feature modification parameters based on the feature modification factor information generated by the text generating unit 203 and the result of extraction of the subject object by the subject object extracting unit 204 .

[0044] A form modification parameter is a parameter that represents a form feature based on modification factor information. Specifically, a form modification parameter is structured data that includes geometric information such as the position, size, shape, and orientation of a form feature that is generated based on a form modification factor, as well as appearance information such as color and texture. Furthermore, a form modification parameter may inherit some or all of the structured information of the form modification factor from which it is generated.

[0045] The feature parameter generator can be realized by a multimodal input neural network that receives, for example, text and images as input and outputs the structured data described above.

[0046] Among the form modification parameters that express form features, the position, size, etc. are restricted by the shape of the subject object to be input, particularly by the contour information, so it is desirable to control the parameters using the subject object image. This point will be explained below with specific examples.

[0047] Assume that a feature modification factor "add bolt holes" is input for each of the target subject object widths of 30 mm and 1000 mm. Because the feature modification factor does not include a specific modifier related to size, the size of the feature modification parameter is set to, for example, "normal," which indicates a generally applicable size. Meanwhile, a user typically expects the diameter of the "bolt holes" to be approximately 10 mm or less for a 30 mm-wide feature and approximately 50 mm or less for a 1000 mm-wide feature. However, it is difficult for the feature update unit 206 (described later) to accommodate variations in the size of the feature feature to be generated, such as 10 mm or 50 mm, in response to a size instruction using the feature modification parameter "normal." In consideration of this situation, the feature parameter generation unit 205 may control the feature modification parameters to be generated using information such as the subject object image.

[0048] The feature parameter generating unit 205 may execute different tasks depending on the attributes of the input feature modification factor information. For example, suppose feature modification factor information having an attribute such as "add bolt holes to flat plate portion" indicating the addition of a feature feature to the feature of the subject object is input. In this case, the feature parameter generating unit 205 may execute, for example, a task of detecting flat plate portions from the subject object image and a task of outputting map information of locations where bolt holes can be added, and determine feature modification parameters such as the position and size of the bolt holes to be added. As another example, suppose feature modification factor information having an attribute such as "remove bolt holes from flat plate portion" indicating the removal of a feature feature from the feature of the subject object is input. In this case, the feature parameter generating unit 205 may execute, for example, only the task of detecting bolt holes in flat plate portions from the subject object image.

[0049] The feature parameter generating unit 205 outputs the generated feature modification parameters to the feature updating unit 206 .

[0050] The feature update unit 206 performs at least one of a number of modifications, including adding, modifying, and deleting feature features, on the subject object extracted by the subject object extraction unit 204 based on the feature modification parameters generated by the feature parameter generation unit 205.

[0051] The feature updater 206 may be implemented, for example, by a diffusion neural network model for image generation that is trained to minimize the distance in feature space between the original image and the generated image in order to preserve the feature features of the original subject object.

[0052] Then, the feature update unit 206 outputs the feature of the subject object (hereinafter also referred to as the updated feature) that has been updated by making changes to the subject object based on the feature modification parameters to the output control unit 207.

[0053] The output control unit 207 controls the output of information indicating the feature of the main subject object changed by the feature update unit 206 (i.e., the updated feature of the main subject object) to a predetermined output destination.

[0054] For example, the output control unit 207 may display the updated form of the main subject object in a predetermined display area (e.g., the display unit 102). In this case, the output control unit 207 may display, in the predetermined display area, an image indicated by the image data acquired by the image acquisition unit 201, the utterance data acquired by the utterance data acquisition unit 202, and the updated form of the main subject object. In this case, the output control unit 207 may selectively switch whether or not to display each of the above-mentioned pieces of information based on a user instruction. The output control unit 207 may also perform interactive processing based on a user instruction via the input unit 101, targeting the updated form of the main subject object displayed in the predetermined display area. An example of this interactive processing will be described in detail later.

[0055] Furthermore, the above is merely an example and does not limit the destination to which the information indicating the updated form of the subject object is output by the output control unit 207. As a specific example, the output control unit 207 may output information indicating the updated form of the subject object to a device that performs various image processing or various analytical processing on the target data, as the target of the image processing or analytical processing.

[0056] An example of processing by the information processing device 100 according to this embodiment will be described with reference to Fig. 3. In S101, the image acquisition unit 201 acquires image data to be processed. In S102, the utterance data acquisition unit 202 acquires utterance data indicating the content of a user's utterance input via a predetermined input interface. Note that the utterance data acquisition unit 202 is not limited to acquiring utterance data based on utterances from a single user, and may acquire utterance data based on utterances from multiple users.

[0057] In S103, the text generating unit 203 generates main object information and form modifier information based on the utterance data acquired in S102.

[0058] In S104, the subject object extraction unit 204 extracts the subject object indicated by the subject object information in the image indicated by the image data based on the image data acquired in S101 and the subject object information generated in S103.

[0059] In S105, the feature parameter generating unit 205 generates feature modification parameters based on the feature modification factor information generated in S103 and the result of the subject object extraction in S104.

[0060] In S106, the feature update unit 206 modifies the main subject object extracted in S104 by adding, changing, or deleting features based on the feature modification parameters generated in S105, thereby updating the feature of the main subject object.

[0061] In S107, the output control unit 207 controls so that information indicating the form of the main subject object changed in S106 (i.e., the updated form of the main subject object) is output to a predetermined output destination. As a specific example, the output control unit 207 may present the updated form of the main subject object to the user by displaying the updated form of the main subject object in a predetermined display area.

[0062] 4, an example of a screen that the information processing device 100 according to this embodiment presents to the user via the display unit 102 will be described. The display screen 400 shown in FIG. 4 includes a target image display unit 410, a main object image display unit 420, an update result display unit 430, a comment content display unit 440, and a form modification factor display unit 450.

[0063] The target image display section 410 is a display area where an image indicated by the image data acquired by the image acquisition section 201 is displayed.

[0064] In the example shown in Fig. 4, four objects, an electrical box 301a, a sphere 301b, a reinforcing plate 301c, and a work desk 301d, are captured as subjects in the image. For convenience, in the example shown in Fig. 4, it is assumed that the electrical box 301a has been determined as the main object based on a user's previously accepted input. In addition, in the example shown in Fig. 4, in order to clearly indicate the extraction result of the main object, a rectangular box 310 (a so-called bounding box) is displayed in the image displayed in the target image display section 410 so as to encompass the area of ​​the electrical box 301a. The target image display section 410 corresponds to an example of a first partial region.

[0065] The utterance content display unit 440 is a display area in which the content of a user's utterance indicated by the utterance data acquired by the utterance data acquisition unit 202 is displayed. In the example shown in Fig. 4, utterances 341a and 341b corresponding to two pieces of utterance data, respectively, are displayed in chronological order, with the older one being presented at the top and the newer one being presented below. As described above, the utterance data includes text information indicating the utterance content and information indicating the speaker of the target utterance. Each of the utterances 341a and 341b is displayed based on the text information indicating the utterance content included in the target utterance data.

[0066] 4, three proposals are made in comment 341a: "Add the drawn shape to the center of the surface as large as possible in the space created by moving the round hole," "Move the rightmost round hole as far to the right as possible," and "Add a hemming bend to the lower bend." In response to these proposals, comment 341b agrees with the first and second proposals, but rejects the third proposal.

[0067] The form modifier display section 450 is a display area where form modifiers indicated by form modifier information generated based on utterance data are displayed. First, the subject object name included in the subject object information determined as the target is displayed in the display area 351 in the form modifier display section 450. In addition, the summary name included in the form modifier information related to the subject object is displayed in the form modifier display section 450.

[0068] 4, summary names 352a, 352b, and 352c corresponding to three pieces of form modification factor information are displayed. Also, the results of estimation by the text generator 203 regarding whether or not form modifications (changes made to an object) can be made for each of the three pieces of form modification factor information are reflected in check boxes 452a, 452b, and 452c displayed in association with each piece of form modification factor information.

[0069] 4, the proposal for "add a hemming bend to the lower bend" indicated by abstract name 352c has been rejected by comment 341b. Therefore, the implementation status of the form modification factor for this proposal is set to "No," and check box 452c is displayed in an OFF state. On the other hand, the proposals corresponding to abstract names 352a and 352b have been agreed upon in comment 341b. Therefore, the implementation status of the form modification factor for these proposals is set to "Yes," and check boxes 452a and 452b are displayed in an ON state.

[0070] The ON / OFF state of each of the check boxes 452a, 452b, and 452c can be changed arbitrarily based on an instruction from the user via the input unit 101. A change in the state regarding whether or not to implement a form modification factor via these check boxes may be used as a trigger to execute a process of making changes to the updated form of the subject object displayed in the update result display unit 430.

[0071] The main object image display section 420 is a display area in which an image of the main object extracted from the acquired image data is displayed based on the extraction result of the main object. In the example shown in Fig. 4, an image of an electrical box determined as the main object is extracted from the image represented by the acquired image data and displayed. The main object image display section 420 corresponds to an example of a second partial area.

[0072] The update result display section 430 is a display area that displays an image of the feature of the subject object (the updated feature of the subject object) that has been modified by adding, changing, or deleting feature features from the subject object based on the feature modification factors. In the example shown in Fig. 4, the results of two modifications to the subject object image 320, "moving the rightmost round hole as far to the right as possible" and "adding an aperture shape as large as possible to the center of the surface," are displayed as the updated feature 330. The update result display section 430 corresponds to an example of a third partial area.

[0073] It should be noted that the display state of each component (each display unit indicated by reference numerals 410 to 450) that makes up the display screen 400 may be individually switched ON / OFF as desired in response to an instruction from the user.

[0074] Fig. 5 is a schematic diagram showing an example of the display state of the target image display section 410 on the display screen 400 shown in Fig. 4. Specifically, Fig. 5 shows a schematic diagram of a situation in which, in addition to objects 301a to 301d that are candidates for the subject of each user's comment, a finger 302 is also imaged as another object.

[0075] Using the state shown in Fig. 5 as an example, an example of the processing by the subject object extraction unit 204 will be described below with reference to Fig. 6. In order to further improve the estimation accuracy of the subject object in the image, the subject object extraction unit 204 uses feature amounts extractable from the image, such as spatial feature amounts and modality feature amounts.

[0076] In S111, the image acquisition unit 201 receives input of image data to be processed from a user. The utterance data acquisition unit 202 acquires utterance data indicating the content of the user's utterance. The text generation unit 203 generates subject object information and form modification factor information based on the utterance data acquired by the utterance data acquisition unit 202.

[0077] In S112, the subject object extraction unit 204 extracts the subject object indicated by the subject object information from the image indicated by the image data (hereinafter also referred to as the input image) based on the image data and subject object information acquired in S111.

[0078] For example, in the example shown in Figure 5, if the name of the acquired subject object is "box," the subject object extraction unit 204 estimates each of objects 301a and 301c as a subject object candidate with a higher likelihood than other objects. In such a case, the subject object extraction unit 204 may use various feature amounts, such as spatial feature amounts and modality feature amounts, to more accurately and uniquely determine the subject object. Note that in the example shown in Figure 6, the subject object extraction unit 204 uses spatial feature amounts and modality feature amounts to uniquely determine the subject object.

[0079] The spatial feature is a feature based on the idea that the main object is larger, more central, and more clearly (out of focus) in the image. In the example shown in Fig. 6, the main object extraction unit 204 is assumed to have a spatial feature encoder (not shown).

[0080] In S113, the main subject object extraction unit 204 inputs main subject object information to the spatial feature encoder, thereby acquiring spatial features as the output of the spatial feature encoder.

[0081] The modality feature is a feature based on the idea that, when a pointing object pointing to a subject object is present in an image, the modality of the pointing object is utilized to accurately identify the subject object. In the example shown in FIG. 6 , the subject object extraction unit 204 includes a modality feature encoder (not shown). In S114, the subject object extraction unit 204 detects a finger 302 as the pointing object. In this case, the subject object is likely to exist at the position or direction indicated by the finger 302 detected as the pointing object. Therefore, in S115, the subject object extraction unit 204 uses the modality feature encoder to acquire, as a modality feature, a subject object existence probability map that outputs a high score for a specific position or direction based on the finger gesture.

[0082] In S116, the subject object extraction unit 204 uniquely determines a subject object based on the subject object information acquired in S111, the spatial feature amount acquired in S113, and the modality feature amount acquired in S115.

[0083] 5, object 301a is larger than object 301c in the image and is located in an area where the probability of the subject object being present is higher according to the finger modality, so object 301a is more likely to be the subject object. Therefore, in this case, of objects 301a and 301c, which are subject object candidates, the subject object extraction unit 204 determines object 301a to be the subject object.

[0084] The above is merely an example, and the features used to determine the main object are not particularly limited as long as they can be extracted from the target image. As a specific example, the image quality of each of a series of objects included in the image (e.g., a feature that serves as an index for evaluating image quality) may be used as the feature. In this case, for example, of the series of objects included in the image, the object with the highest image quality may be determined as the main object.

[0085] In S117, the subject object extraction unit 204 outputs the extraction result of the subject object (object 301 a) determined in S116, including the position information of the subject object, to a predetermined output destination. Note that the position information of the subject object may be output as, for example, the position information of a rectangle (bounding box) that contains the subject object.

[0086] By applying the above-described control, it is expected that the accuracy of extracting the subject object will be further improved.

[0087] 7, another example of the display state of the target image display section 410 on the display screen 400 will be described. In addition to the objects 301a to 301d, rectangles 311a and 311b containing objects 301a and 301c, respectively, which are candidate subject objects based on the subject object information, are displayed in the image displayed on the target image display section 410. The user can select one of the candidate subject objects 301a and 301c as the subject object by using a point 401 via the input section.

[0088] Using the state shown in Figure 7 as an example, and referring to Figure 8, we will explain another example of the processing of the subject object extraction unit 204, where the subject object extraction unit 204 uniquely determines a subject object via user input.

[0089] In S121, the image acquisition unit 201 receives input of image data to be processed from a user. The utterance data acquisition unit 202 acquires utterance data indicating the content of the user's utterance. The text generation unit 203 generates subject object information and form modification factor information based on the utterance data acquired by the utterance data acquisition unit 202.

[0090] In S122, the subject object extraction unit 204 extracts the subject object indicated by the subject object information from the image indicated by the image data (hereinafter also referred to as the input image) based on the image data and subject object information acquired in S121.

[0091] For example, in the example shown in Figure 7, if the name of the acquired subject object is "box", the subject object extraction unit 204 estimates each of objects 301a and 301c as a candidate subject object with a higher likelihood than other objects.

[0092] In S123, based on the estimation result of the candidate subject objects in S122, the output control unit 207 displays rectangles 311a and 311b containing the candidate subject objects 301a and 301c, respectively, on the target image display unit 410. In S124, the subject object extraction unit 204 receives from the user a selection of either rectangle 311a or 311b displayed on the target image display unit 410 by the user operating a point 401 via the input unit.

[0093] In S125, the main object extraction unit 204 identifies the object corresponding to the rectangle selected in S124 as the main object, and outputs the extraction result of the main object, including the position information of the main object, to a predetermined output destination. Note that the position information of the main object may be output as the position information of a rectangle (bounding box) that contains the main object, for example.

[0094] By applying the above-described control, it becomes possible to more reliably identify the subject object intended by the user.

[0095] An example of the processing performed by the text generator 203 will be described with particular focus on the processing related to the extraction of form modifiers, with reference to Fig. 9. In the example shown in Fig. 9, the text generator 203 filters out form modifiers with low likelihood of being more faithful to the user's intention.

[0096] In S131, the text generating unit 203 receives utterance data (for example, utterance data acquired by the utterance data acquiring unit 202) as input.

[0097] In S132, the text generating unit 203 generates subject object information based on the utterance data received as input in S131.

[0098] In S133, the text generation unit 203 generates form modification factor information based on the utterance data received as input in S131. The generated form modification factor information includes an estimated likelihood as a form modification factor. As a specific example, if two form modification factors, "move the round hole to the right" and "turn it quietly," are extracted from the utterance data, the former will have a higher likelihood as a form modification factor, and the latter will have a lower likelihood.

[0099] In S134, the text generating unit 203 compares the likelihood included in each piece of form modification factor information generated in S133 with a preset threshold, and deletes form modification factor information whose likelihood is lower than the threshold.

[0100] In S135, if there is any form modification factor information that has not been deleted as a result of the processing in S134, the text generating unit 203 outputs the form modification factor information to a predetermined output destination.

[0101] As described above, filtering of form modification factor information by the text generation unit 203 eliminates subsequent processing of form modification factor information deleted from the processing target, which is expected to reduce the processing load and improve the processing speed. Furthermore, the information displayed in the form modification factor display section 450 on the display screen 400 is simplified, which is expected to improve user convenience.

[0102] On the other hand, the above-described method may also extract words and phrases from utterance data that are less relevant to the topic but generally modify the shape. For example, assume that the topic object is an "electrical box," and the text generator 203 generates shape modification factor information corresponding to two shape modification factors: "move the round hole to the right" and "add tapered edges to the gear." In this example, even if the topic is an "electrical box" and the shape modification factor "add tapered edges to the gear" is unrelated to the topic, a shape modification factor derived from a word clearly related to shape modification may be extracted with a higher likelihood. In light of this situation, another example of a method for filtering shape modification factors is proposed below.

[0103] 10 is a flowchart showing an example of the processing of the feature update unit 206 according to this embodiment. In the example shown in Fig. 10, the feature update unit 206 filters feature modifiers with a lower estimated likelihood as feature modifiers in order to update the feature of the subject object more faithfully to the user's intention (to make changes to the feature).

[0104] 11A to 11C are schematic diagrams showing an example of a detection likelihood map generated by the feature update unit 206 according to this embodiment. The detection likelihood map is a map indicating where in the image a detection target exists and with what probability. Details of the detection likelihood map will be described separately below. In the example shown in FIGS. 11A to 11C, the subject object image 320 includes round holes 321a to 321c and flat surfaces 322a and 321b. In the example shown in FIGS. 11A to 11C, the subject object is named "electrical box," and FIGS. 11A, 11B, and 11C each show a summary name of a different feature modification parameter. Specifically, FIG. 11A corresponds to an example of a feature modification parameter indicating "moving a hole," FIG. 11B corresponds to an example of a feature modification parameter indicating "adding a taper to the center of the surface," and FIG. 11C corresponds to an example of a feature modification parameter indicating "adding a taper to the edge of the gear."

[0105] In step S141, the feature update unit 206 receives as input the feature object extraction result including the feature object image and the feature modification parameters. The feature modification parameters inherit the feature modification factor information and also include the summary name and attributes included in the feature modification factor information.

[0106] In S142, the feature update unit 206 executes a detection task for the feature to be processed, regardless of whether the attribute of the feature modification parameter is added, changed, or deleted, and generates a detection likelihood map based on the execution result of the detection task.

[0107] 11A, the feature update unit 206 detects each of the round holes 321a to 321c from the subject object image for the feature modification parameter "move hole." At this time, the detection likelihood of the areas corresponding to the round holes 321a to 321c is set to a higher value than that of the other areas. In the figure, hatched areas indicate areas with a higher detection likelihood (e.g., areas where the detection likelihood is equal to or greater than a threshold), and unhatched areas indicate areas with a lower detection likelihood (e.g., areas where the detection likelihood is less than a threshold).

[0108] 11B, the feature update unit 206 detects each of the flat surfaces 322a and 322b from the subject object image in response to the feature modification parameter "add aperture shape to center of surface." At this time, the detection likelihood of the areas corresponding to the flat surfaces 322a and 322b is set to a higher value than that of other areas.

[0109] 11C, the feature update unit 206 executes a process to detect gears from the main object image for the feature modification parameter "add tapered edges of gears." On the other hand, in the examples shown in Figures 11A to 11C, since no gears are present, a lower value is set as the detection likelihood across the entire main object image.

[0110] In S143, the feature update unit 206 calculates a predetermined statistical quantity (e.g., a maximum value) based on each of the series of detection likelihood maps generated in S142, and if the statistical quantity is below a threshold value, deletes the feature modification parameter corresponding to the detection likelihood map.

[0111] In S144, if there are any feature modification parameters that have not been deleted as a result of the processing in S143, the feature update unit 206 updates the feature of the subject object based on those feature modification parameters and outputs the results of the update to a specified output destination.

[0112] Through the above series of processes, the subject object image is matched with the feature modification parameters, and a filtering process is performed using the statistics of the feature modification parameter detection likelihood map. This prevents the output of a feature to which feature modification factors not originally intended by the user have been applied, and makes it possible to update the feature of the subject object more faithfully to the user's intention.

[0113] Next, an example of a method for generating form modification parameters that more accurately reflect the user's intentions through interactive processing via the display unit 102 and the input unit 101 will be described below.

[0114] 12A to 12D are diagrams showing another example of a screen that the information processing apparatus 100 according to this embodiment presents to the user via the display unit 102. The display screen 400 shown in each of FIGS. 12A to 12D includes an update result display unit 430 and a feature modification factor display unit 450. A pointer 401 is also displayed in the display screen 400, which is used to designate a portion of the display screen 400 based on a user operation received via the input unit 101. In the example shown in FIGS. 12A to 12D, three feature modification factors 352a to 352c are displayed in the feature modification factor display unit 450. The feature modification factor 352a indicates a feature modification of "moving the rightmost hole." The feature modification factor 352b indicates a feature modification of "adding a large draw shape to the center of the surface." The feature modification factor 352c indicates a feature modification of "adding a hemming bend to the lower bend."

[0115] In the example shown in Fig. 12A, the user can select any feature modification factor in the feature modification factor display section 450. For example, in the example shown in Fig. 12A, it is assumed that the user selects a feature modification factor 352a indicating the feature modification "move the rightmost hole." In this case, a command list 353a for executing predefined processes for the selected feature modification factor 352a is displayed. In the example shown in Fig. 12A, commands for applying processes such as "move," "duplicate," "delete," and "change shape" to feature features based on the feature modification factor are displayed in the command list 353a.

[0116] Furthermore, more detailed conditions may be specified for at least some commands. For example, in the example shown in Fig. 12A, a text box 354a is displayed for the "Change Shape" command to accept specification of a new shape as text information.

[0117] With this configuration, it becomes possible to enjoy usability that is substantially the same as when using the information processing device in which shapes are automatically generated based on the above-mentioned utterance data.

[0118] The example shown in Figure 12B is a schematic diagram illustrating an example of a state in which a command indicating "move" has been selected by the user for a feature modifier 352a indicating a feature modification of "move the rightmost hole." When the command indicating "move" has been selected, the position of the round hole 331, which is a feature based on the target feature modifier 352a, can be moved by a drag operation using the pointer 401, for example. In the example shown in Figure 12B, the round hole 331' of the feature 330 of the updated subject object shows the round hole 331 after it has been moved by the above operation.

[0119] The example shown in Figure 12C is a schematic diagram illustrating an example of a state in which a command indicating "change shape" has been selected by the user for a shape modifier 352a indicating a shape modification of "move the rightmost hole." When the command indicating "change shape" is selected, the outline of the round hole 331, which is a shape feature based on the target shape modifier 352a, can be moved by a drag operation using the pointer 401, for example. In the example shown in Figure 12C, the round hole 331'' of the shape 330 of the updated subject object shows the round hole 331 after its size has been changed by pulling the outline of the round hole 331 outward.

[0120] The example shown in FIG. 12D is a schematic diagram illustrating another example of a state in which a command indicating "change shape" is selected by the user for a shape modification factor 352a indicating a shape modification of "move the rightmost hole." When the command indicating "change shape" is selected, an instruction to change the shape of the round hole 331 can be entered via the text box 354a. In the example shown in FIG. 12D, an instruction to double the size of the round hole 331 is entered via the text box 354a. Furthermore, the round hole 331''' of the updated subject object shape 330 shows the round hole 331 after its size has been changed based on the instruction entered in the text box 354a.

[0121] By applying the above-described controls, the user can set geometric feature information, such as the position, size, shape, and orientation of feature information based on feature modification factors, to the desired state. While the above description focuses primarily on the case where geometric feature information is modified, the subject of modification is not limited to geometric feature information. As a specific example, appearance information, such as color and texture, can also be modified by applying substantially the same controls as those described above.

[0122] Second Embodiment A second embodiment of the present disclosure will be described below. In this embodiment, an example of a configuration will be described in which changes made to a subject object based on form modification factor information are stored in a predetermined storage area, and the information stored in the storage area is made available for subsequent use. In addition, this embodiment will be described focusing on parts that are particularly different from the first embodiment described above, and detailed description of parts that are substantially similar to the first embodiment will be omitted.

[0123] An example of the functional configuration of the information processing device according to this embodiment will be described with reference to Fig. 13. The information processing device according to this embodiment differs from the information processing device according to the first embodiment described with reference to Fig. 2 in that it includes a storage unit 208.

[0124] The storage unit 208 is a storage area for storing various types of data. The data stored in the storage unit 208 includes two main types of data. The first type of data is learning data for learning the correlation between form modification factors and form modification parameters for each user, so that the information processing device 100 can more accurately reflect the user's intentions when making changes to the subject object. The second type of data is history data indicating the results of updates that are saved as history so that the user can quickly check the results of updates to the form of the subject object when form features are added, changed, or deleted from the subject object.

[0125] First, an example of the processing of the information processing device according to this embodiment will be described, focusing on the processing related to saving the learning data, with reference to Fig. 14. Fig. 14 shows an example of a series of processing executed after the processing related to generating feature parameters by the feature parameter generating unit 205.

[0126] In S211, the feature parameter generating unit 205 generates feature modification parameters based on the subject object image and feature modification factor information received as input.

[0127] In S212, the feature update unit 206 modifies the feature of the subject object so that it has features according to the feature modification parameters generated in S211.

[0128] In S213, the output control unit 207 displays the result of updating the form of the main subject object in a predetermined display area in S212. As a specific example, the output control unit 207 may display the display screen 400 on the display unit 102, and display the result of updating the form of the main subject object in the update result display area 450 of the display screen 400.

[0129] In S214, the feature parameter generation unit 205 accepts a user operation related to modification of the feature characteristics of the main subject object displayed in the predetermined display area in S213. At this time, the feature parameter generation unit 205 may accept an operation related to modification of the feature characteristics of the main subject object from the user through interactive processing via the display unit 102 and the input unit 101, as described above with reference to Figures 12A to 12D.

[0130] In S215, the feature parameter generation unit 205 executes processing related to the modification of the feature modification parameters generated in S211 so as to follow the operation related to the modification of the feature features of the subject object received from the user in S214.

[0131] As a specific example, suppose that a feature "add hole" is added, the "shape" item of the feature modification parameter in the corresponding feature modification factor information is round, and the user modifies the shape of the target hole to a rectangle. In this case, the feature parameter generating unit 205 changes the "shape" item of the feature modification parameter to a rectangle.

[0132] In S216, the form parameter generating unit 205 determines whether the user operation is completed.

[0133] If the feature parameter generation unit 205 determines in S216 that the user operation has not been completed, it proceeds to S214. In this case, the processes of S214 and S215 are executed again. In this way, the feature parameter generation unit 205 accepts user operations related to modification of the feature features of the subject object until the user operation is completed, and each time, it executes processing related to modification of the feature modification parameters in accordance with the user operation.

[0134] If the form parameter generating unit 205 determines in S216 that the user operation has been completed, the process proceeds to S217.

[0135] In S217, the output control unit 207 stores in the storage unit 208 as learning data a set of the form modification parameters corrected by the processes in S214 to S216 and the form modification factors from which the form modification parameters were generated.

[0136] The learning data generated in the above manner is divided into data for each user and then used for learning by the feature parameter generation unit 205, thereby enabling the feature parameter generation unit 205 to generate feature modification parameters that are close to the user's intentions.

[0137] The process described with reference to FIG. 14 will now be described in more detail with reference to FIG. 15 using a specific example.

[0138] Figure 15(a) schematically shows a subject object image to be processed. It is assumed that a user makes a statement indicating an instruction to "add an aperture" to the subject object image shown in Figure 15(a). In this case, based on the statement data indicating the content of the user's statement, the result of changes made to the shape of the subject object is displayed in a predetermined display area, as shown in Figure 15(b). Then, the user further performs an operation to modify the aperture shape from the state shown in Figure 15(b), and the result of the modification of the aperture shape reflected in the shape of the subject object is displayed, as shown in Figure 15(c).

[0139] In Fig. 15(b), the feature parameter generating unit 205 generates feature modification parameters indicating that the aperture shape is round, and the result of modifying the feature characteristics indicated by the feature modification parameters for the feature of the main object is displayed. On the other hand, since the feature characteristics have been modified as shown in Fig. 15(c), it is clear that the aperture shape intended by the user was rectangular in the example shown in Fig. 15. In response to these results, the storage unit 208 stores the target main object image, feature modification factors, and feature modification parameters (the feature modification parameters reflecting modifications based on user instructions) as learning data.

[0140] By using this learning data for training the feature parameter generation unit 205, it is possible to have the feature parameter generation unit 205 learn that when an instruction is given to add an aperture shape to a box-shaped part, the shape of the feature modification parameter is rectangular.

[0141] An example of a display screen of the information processing device according to this embodiment will be described with reference to Fig. 16. It is assumed here that the feature parameter generating unit 205 that has been trained based on the training data described with reference to Fig. 15 is applied.

[0142] 16 , the display screen 400 includes an update result display section 430 and a feature modifier display section 450. The update result display section 430 displays the subject object "reinforcement plate," and the feature modifier display section 450 displays an instruction to "add aperture." In response to the feature modifier instruction, the update result display section 430 also displays a marker 336 indicating the center position for generating a feature, and a feature candidate window 335 for selecting the shape, size, and orientation of the feature to be applied from among the candidates. The feature candidate window 335 also displays, for each feature candidate to be selected, a likelihood corresponding to an estimated result of the likelihood that the candidate will be selected by the target user.

[0143] In the example shown in Figure 16, the form parameter generation unit 205 has been trained using the training data generated based on the conditions shown in Figure 15, and therefore in response to the instruction to "add aperture", the estimated likelihood of an aperture having a rectangular shape is highest.

[0144] In this way, candidates with higher likelihoods estimated by the feature parameter generation unit 205, which has been trained using learning data for each user, are displayed with higher priority in the feature candidate window 335. By applying this type of control, the user can more easily select a feature candidate that is closest to their intention from the series of candidates presented in the feature candidate window 335.

[0145] Next, the history data stored in the storage unit 208 according to the results of updating the form of the subject object will be described. FIGS. 17A to 17C are diagrams showing an example of a display screen of an information processing device according to this embodiment. FIGS. 17A to 17C are diagrams showing the states of a display screen 400 in chronological order, with FIG. 17A showing the oldest state and FIG. 17C showing the newest state. The display screen 400 shown in FIGS. 17A to 17C also includes an update result display section 430 and a history display section 460.

[0146] First, the state shown in Fig. 17A will be described. In the state shown in Fig. 17A, only the feature modification factor 361a is selected as the target for application. Therefore, the update result display section 430 displays the feature 330a of the main object that has been updated by reflecting the feature characteristics indicated by the feature modification factor 361a. At this time, the history display section 460 also reflects information indicating that the feature modification factor 361a has been applied. In this case, the updated feature 330a of the main object and the feature modification factor 361a are linked and stored in the storage section 208 as history data.

[0147] 17A to 17C, the selection state of each form modification factor or each form modification parameter is displayed by the selection state of a check box in the history display section 460. Furthermore, by operating the check box, the state can be switched between a selected state and a non-selected state, thereby selectively switching whether or not the target form modification factor or form modification parameter is applied.

[0148] Next, the state shown in FIG. 17B will be described. In the state shown in FIG. 17B, the form modification factors 361a and 361b and the morphological modification parameters 361a' and 361a'' have been selected as the objects to be applied. Therefore, the update result display section 430 displays a form 330c of the main subject object that has been updated by reflecting the form features indicated by the form modification factors 361a and 361b and the morphological modification parameters 361a' and 361a''. At this time, the history display section 460 also reflects information indicating that the form modification factors 361a and 361b and the morphological modification parameters 361a' and 361a'' have been applied. In this case, the updated form 330c of the main subject object, the form modification factors 361a and 361b, and the morphological modification parameters 361a' and 361a'' are linked to each other and stored in the storage unit 208 as history data.

[0149] Next, the state shown in FIG. 17C will be described. In the state shown in FIG. 17C, the check boxes for the form modifiers and form modifier parameters other than the form modifier 361a are deselected. Therefore, the update result display section 430 displays the subject object's form 330a', which has been updated by reflecting the form features of only the form modifier 361a among the series of form modifiers and form modifier parameters presented in the history display section 460. At this time, the output control section 207 checks whether history data for the updated form 330a, to which only the form modifier 361a is linked, is stored in the storage section 208. Note that in the example shown in FIGS. 17A to 17C, history data for the updated form 330a, to which only the form modifier 361a is linked, is stored in the storage section 208 in the state shown in FIG. 17A. Therefore, the output control unit 207 displays the updated feature 330 a on the update result display unit 430 based on the history data stored in the storage unit 208 without going through the feature update unit 206 .

[0150] A possible use case of this embodiment is when a user compares and examines the shape of a target object by selectively switching on and off the application of various previously specified shape modifiers, and then confirms the updated shape of the object. Even in such a case, this embodiment eliminates the need to regenerate the corresponding shape if history data is saved, thereby speeding up the process of displaying the updated shape of the subject object. Furthermore, because the updated shape of the subject object is displayed based on the history data, it is possible to ensure the reproducibility of the shape of the subject object according to the selected state of the shape modifiers.

[0151] Other Embodiments Although the exemplary embodiments have been described above in detail, the present invention can be embodied, for example, as a system, a device, a method, a program, a recording medium (storage medium), etc. Specifically, the present invention may be applied to a system configured from multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may be applied to an apparatus consisting of a single device.

[0152] Needless to say, the object of the present invention can be achieved by the following: A recording medium (or storage medium) on which software program code (computer program) that realizes the functions of the above-described embodiments is recorded is supplied to a system or device. The storage medium is, of course, a computer-readable storage medium. The computer (or CPU or MPU) of the system or device then reads and executes the program code stored on the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-described embodiments, and the recording medium on which the program code is recorded constitutes the present invention.

[0153] The disclosure of this embodiment also includes the following configurations, methods, and programs.

[0154] (Configuration 1) An information processing device characterized by having an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance content, and form modification factor information which is information that modifies the form of the object included in the utterance content; a modification means for making changes based on the form modification factor information to an object corresponding to the subject object information acquired by the acquisition means, among one or more objects included in an image that is the subject of the user's utterance; and an output control means for controlling the results of the changes made to the object by the modification means to be output to a predetermined output destination.

[0155] (Configuration 2) The information processing device according to Configuration 1, further comprising a generating means for generating a form modification parameter indicating a form characteristic of the object based on the form modification factor information, and the modifying means for modifying the object corresponding to the subject object information based on the form modification parameter.

[0156] (Configuration 3) The information processing device according to Configuration 2, characterized in that the modification performed by the modification means includes at least one of adding the geometric feature to the object corresponding to the subject object information based on the geometric modification parameters, transforming at least a part of the object based on the geometric feature, and deleting the geometric feature from the object.

[0157] (Configuration 4) The information processing device according to Configuration 2 or 3, wherein the form modification parameters include at least one of form information including at least one of position, size, shape, and orientation, and appearance information including at least one of color and texture.

[0158] (Configuration 5) An information processing device described in any one of configurations 2 to 4, characterized in that the form modification factor information includes, as an attribute, information indicating which of multiple modifications, including addition, deletion, and transformation, is to be made to the object corresponding to the subject object information, and the generation means controls the configuration of the form modification parameters to be generated in accordance with the attribute included in the form modification factor information.

[0159] (Configuration 6) An information processing device described in any one of configurations 2 to 5, characterized in that the form modification factor information includes at least one of the name of the object that is the subject of the comment content, an estimated likelihood of the object being the subject, whether or not changes can be made to the object, and information about the user who made the comment indicated by the comment content.

[0160] (Configuration 7) The information processing device described in Configuration 6, characterized in that the change means suppresses changes based on the form modification factor information to the object corresponding to the subject object information when the estimated likelihood included in the form modification factor information is less than a threshold value.

[0161] (Configuration 8) The information processing device described in Configuration 6, wherein the modification means is a likelihood map indicating where and with what probability an object corresponding to the subject object information exists in the image that is the subject of the user's remarks, and when a predetermined statistical quantity obtained from the likelihood map based on the estimated likelihood included in the form modification factor information is less than a threshold, the information processing device suppresses modifications to the object corresponding to the subject object information based on the form modification factor information.

[0162] (Configuration 9) An information processing device as described in Configuration 2, further comprising a storage means for storing form modification factor information and form modification parameters generated based on the form modification factor information in association with each other as learning data, wherein the generation means generates form modification parameters based on user information included in the form modification factor information received as input and the learning data stored by the storage means.

[0163] (Configuration 10) The information processing device described in any one of configurations 2 to 9, characterized in that the output control means controls the results of changes made to the object by the change means to be displayed in a predetermined display area, the display area including a first partial area in which an image that is the subject of the user's utterance is displayed, a second partial area in which an image of an object corresponding to the subject object information is displayed, and a third partial area in which an image corresponding to the results of changes made to the object by the change means is displayed, and the information processing device is configured to be able to selectively switch whether or not to display images in each of the first partial area, the second partial area, and the third partial area.

[0164] (Configuration 11) The information processing device described in Configuration 10, wherein the third partial area is configured to be able to accept instructions from the user regarding changes to the object corresponding to the main object information, and the generation means controls the form modification parameters applied to make changes to the object corresponding to the main object information in accordance with the instructions accepted by the third partial area from the user.

[0165] (Configuration 12) The information processing device described in configuration 10 or 11, characterized in that the display area includes a fourth partial area configured to be able to accept selection of at least any one of the acquired series of form modification factor information, and the generation means controls whether or not to add changes to the object corresponding to the subject object information based on each of the series of form modification factor information, depending on the selection state of each of the series of form modification factor information accepted via the fourth partial area.

[0166] (Configuration 13) The information processing device described in Configuration 12 is characterized in that the fourth partial area is configured to be able to accept specification of a process to be applied to the selected form modification factor information to make changes to form features, and the generation means controls the form modification parameters to be applied to make changes to the object corresponding to the subject object information based on the form modification factor information selected via the fourth partial area and the process specified for the form modification factor information.

[0167] (Configuration 14) An information processing device as described in any one of configurations 2 to 13, characterized in that it comprises a storage means for sequentially storing as a history the results of changes made to the object by the change means in accordance with the results of acquisition of the form modification factor information by the acquisition means, and a reception means for receiving instructions regarding whether or not to apply each of the series of form modification factor information acquired by the acquisition means, and the output control means, when the results of changes made to the object by the change means based on one or more form modification factor information for which an instruction to apply has been accepted by the reception means, is stored as the history, controls so that the results indicated in the history are output to a predetermined output destination.

[0168] (Configuration 15) The information processing device described in Configuration 14, characterized in that the receiving means receives instructions regarding whether to apply one or more processes for modifying a geometric feature, which are specified for at least any of the one or more pieces of shape modification factor information for which an instruction to apply has been received, and the output control means, when the results of changes made to the object by the modifying means based on the one or more pieces of shape modification factor information for which an instruction to apply has been received by the receiving means and one of the one or more processes for which an instruction to apply has been received, is stored as the history, controls so that the results indicated in the history are output to a predetermined output destination.

[0169] (Configuration 16) The information processing device described in Configuration 15 is characterized in that, if the results of changes made to the object by the change means based on the one or more pieces of form modification factor information for which an instruction to apply has been accepted by the acceptance means and a form modification parameter from the series of form modification parameters for which an instruction to apply has been accepted, are not stored as the history, the change means makes changes to the object corresponding to the subject object information based on the one or more pieces of form modification factor information and the form modification parameter for which an instruction to apply has been accepted, and the output control means controls so that the results of changes made to the object by the change means are output to a predetermined output destination.

[0170] (Configuration 17) The information processing device according to any one of configurations 1 to 16, wherein the utterance data includes at least one of text data input by the user and text data generated based on a recognition result of a speech uttered by the user.

[0171] (Configuration 18) The information processing device described in any one of configurations 1 to 17, characterized in that the modification means extracts an object corresponding to the subject object information from the image based on at least one of spatial features, image quality, and modality of one or more objects included in the image that is the subject of the user's remarks, and makes modifications to the extracted object based on the form modification factor information.

[0172] (Configuration 19) The information processing device described in Configuration 18, characterized in that, when there are multiple candidate objects corresponding to the subject object information, the change means extracts at least some of the multiple candidate objects as objects corresponding to the subject object information based on instructions from a user.

[0173] (Method 1) A control method for an information processing device, comprising: an acquisition step of acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance content, and form modification factor information which is information that modifies the form of the object included in the utterance content; a modification step of making changes based on the form modification factor information to an object corresponding to the subject object information acquired in the acquisition step, among one or more objects included in an image that is the subject of the user's utterance; and an output control step of controlling so that the results of the changes made to the object in the modification step are output to a predetermined output destination.

[0174] (Program 1) A program for causing a computer to function as an information processing device, characterized by having an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating the object that is the subject of the utterance, and form modification factor information which is information that modifies the form of the object included in the utterance; a modification means for making changes based on the form modification factor information to one or more objects included in an image that is the subject of the user's utterance, which object corresponds to the subject object information acquired by the acquisition means; and an output control means for controlling the results of the changes made to the object by the modification means to be output to a specified output destination.

[0175] The present invention is not limited to the above-described embodiments, and various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, the following claims are appended to apprise the public of the scope of the present invention.

[0176] This application claims priority based on Japanese Patent Application No. 2023-204878, filed December 4, 2023, the entire contents of which are incorporated herein by reference.

[0177] 100 Information processing device 203 Text generation unit 206 Shape update unit 207 Output control unit

Claims

1. An information processing device comprising: an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance, and form modification factor information which is information for modifying the form of the object included in the utterance; a modification means for making modifications based on the form modification factor information to an object corresponding to the subject object information acquired by the acquisition means, among one or more objects included in an image which is the subject of the user's utterance; and an output control means for controlling the results of the modifications made to the object by the modification means to be output to a specified output destination.

2. An information processing device as described in claim 1, further comprising a generating means for generating a form modification parameter indicating a form characteristic of the object based on the form modification factor information, and the modifying means for making modifications to an object corresponding to the subject object information based on the form modification parameter.

3. The information processing device of claim 2, characterized in that the modifications made by the modification means include at least one of the following based on the geometrical modification parameters: adding the geometrical feature to an object corresponding to the subject object information; transforming at least a part of the object based on the geometrical feature; and deleting the geometrical feature from the object.

4. An information processing device as described in claim 2, characterized in that the shape modification parameters include at least one of shape information including at least one of position, size, shape, and orientation, and appearance information including at least one of color and texture.

5. The information processing device of claim 2, characterized in that the form modification factor information includes, as an attribute, information indicating which of a number of modifications, including addition, deletion, and transformation, is to be made to the object corresponding to the subject object information, and the generation means controls the configuration of the form modification parameters to be generated in accordance with the attribute included in the form modification factor information.

6. The information processing device of claim 2, characterized in that the form modification factor information includes at least any one of the name of the object that is the subject of the comment content, an estimated likelihood of the object being the subject of the comment content, whether or not changes can be made to the object, and information about the user who made the comment indicated by the comment content.

7. The information processing device described in claim 6, characterized in that the modification means suppresses modifications based on the form modification factor information to the object corresponding to the subject object information when the estimated likelihood contained in the form modification factor information is less than a threshold value.

8. The information processing device described in claim 6, wherein the modification means is a likelihood map indicating where and with what probability an object corresponding to the subject object information exists in the image that is the subject of the user's comment, and when a predetermined statistical amount obtained from the likelihood map based on the estimated likelihood contained in the form modification factor information is less than a threshold value, the modification to the object corresponding to the subject object information based on the form modification factor information is suppressed.

9. An information processing device as described in claim 2, further comprising a storage means for storing form modification factor information and form modification parameters generated based on the form modification factor information in association with each other as learning data, wherein the generating means generates form modification parameters based on user information contained in the form modification factor information received as input and the learning data stored by the storage means.

10. The information processing device of claim 2, characterized in that the output control means controls so that the results of the changes made to the object by the modification means are displayed in a specified display area, the display area including: a first partial area in which an image that is the subject of the user's remarks is displayed; a second partial area in which an image of an object corresponding to the subject object information is displayed; and a third partial area in which an image according to the results of the changes made to the object by the modification means is displayed; and the information processing device is configured to be selectively configured to enable or disable display of images in each of the first partial area, the second partial area, and the third partial area.

11. The information processing device of claim 10, characterized in that the third partial area is configured to be capable of receiving instructions from the user regarding changes to the object corresponding to the subject object information, and the generation means controls the form modification parameters applied to make changes to the object corresponding to the subject object information in accordance with the instructions received by the third partial area from the user.

12. The information processing device of claim 10, wherein the display area includes a fourth sub-area configured to accept a selection of at least any one of the acquired series of form modification factor information, and the generation means controls whether or not to add changes to the object corresponding to the subject object information based on each of the series of form modification factor information depending on the selection state of each of the series of form modification factor information accepted via the fourth sub-area.

13. The information processing device of claim 12, characterized in that the fourth sub-area is configured to accept specification of a process for modifying a geometric feature to be applied to the selected geometric modification factor information, and the generation means controls the geometric modification parameters applied to modify an object corresponding to the subject object information based on the geometric modification factor information selected via the fourth sub-area and the process specified for the geometric modification factor information.

14. An information processing device as described in claim 2, comprising: a storage means for sequentially storing, as a history, the results of changes made to the object by the modification means in accordance with the results of acquisition of the form modification factor information by the acquisition means; and a receiving means for receiving an instruction as to whether or not to apply each of the series of form modification factor information acquired by the acquisition means; wherein, when the results of changes made to the object by the modification means based on one or more pieces of form modification factor information for which an instruction to apply has been accepted by the receiving means are stored as the history, the output control means controls so that the results indicated by the history are output to a specified output destination.

15. The information processing device according to claim 14, characterized in that the receiving means receives instructions regarding whether or not to apply one or more processes for modifying a geometric feature, which are designated for at least any of the one or more pieces of geometric modification factor information for which an instruction to apply has been received, and the output control means, when a result of a change made to the object by the modification means based on the one or more pieces of geometric modification factor information for which an instruction to apply has been received by the receiving means and a process among the one or more processes for which an instruction to apply has been received, is stored as the history, controls so that the result indicated by the history is output to a specified output destination.

16. The information processing device according to claim 15, characterized in that, if the results of changes made to the object by the modification means based on the one or more pieces of form modification factor information for which an instruction to apply has been accepted by the accepting means and a form modification parameter from the series of form modification parameters for which an instruction to apply has been accepted, are not retained as the history, the modification means makes changes to the object corresponding to the subject object information based on the one or more pieces of form modification factor information and the form modification parameter for which an instruction to apply has been accepted, and the output control means controls the results of changes made to the object by the modification means to be output to a specified output destination.

17. The information processing device according to claim 1, wherein the utterance data includes at least one of text data input by the user and text data generated based on the recognition results of the voice spoken by the user.

18. The information processing device described in claim 1, characterized in that the modification means extracts an object corresponding to the subject object information from the image based on at least one of spatial characteristics, image quality, and modality of one or more objects contained in the image that is the subject of the user's comment, and applies modifications to the extracted object based on the form modification factor information.

19. The information processing device according to claim 18, characterized in that, when there are multiple candidates for an object corresponding to the subject object information, the modification means extracts at least a portion of the multiple candidate objects as the object corresponding to the subject object information based on instructions from a user.

20. A control method for an information processing device, comprising: an acquisition step of acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance, and form modification factor information which is information for modifying the form of the object included in the utterance; a modification step of making modifications based on the form modification factor information to an object corresponding to the subject object information acquired in the acquisition step, among one or more objects included in an image which is the subject of the user's utterance; and an output control step of controlling so that the results of the modifications made to the object in the modification step are output to a predetermined output destination.

21. A program for causing a computer to function as an information processing device, comprising: an acquisition means for acquiring, based on utterance data indicating the content of a user's utterance, subject object information indicating an object that is the subject of the utterance, and form modification factor information which is information for modifying the form of the object contained in the utterance; a modification means for making modifications based on the form modification factor information to an object corresponding to the subject object information acquired by the acquisition means, among one or more objects contained in an image which is the subject of the user's utterance; and an output control means for controlling the results of the modifications made to the object by the modification means to be output to a specified output destination.

Citation Information

Patent Citations

  • Information processor receiving voice operation instruction

    JP1996221245A

  • CAD command input device

    JP2018049350A

  • Control device

    WO2020067256A1