Information processing device, method for controlling information processing device, and program

The information processing apparatus addresses the challenges of accurately reflecting user intentions in image modification by extracting subject object and form modification information from speech data and applying these modifications to images in real-time, enhancing user experience and simplifying the process.

JP2025089908AActive Publication Date: 2025-06-16CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023204878
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-16
Estimated Expiration
2043-12-04

AI Technical Summary

Technical Problem

Existing technologies for image modification based on user speech often fail to accurately reflect the user's intentions, and require cumbersome sequential preparation of text information for image modification.

Method used

An information processing apparatus that acquires speech data to extract subject object information and form modification factor information, and then modifies the corresponding object in an image based on this information, allowing for real-time and intuitive image modification.

Benefits of technology

Enables more accurate and intuitive reproduction of the form of an object indicated by user speech, improving user experience by simplifying the image modification process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025089908000001_ABST
    Figure 2025089908000001_ABST
Patent Text Reader

Abstract

To reproduce the shape of an object indicated by contents of a user's speech in a more preferable manner.SOLUTION: A text generation unit 203 acquires, on the basis of speech data indicating a user's speech contents, theme object information indicating an object that is the theme of the speech contents, and shape modification factor information, which is information for modifying a shape of the object included in the speech contents. On the basis of the shape modification factor information, a shape update unit 206 adds a change to an object corresponding to the theme object information acquired by the text generation unit 203, among one or more objects included in an image toward which the speech of the user is directed. An output control unit 207 performs control so that a result obtained by the addition of the change to the object by the shape update unit 206 is output to a predetermined output destination.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, a control method for the information processing apparatus, and a program.

Background Art

[0002] In recent years, the use of video conferencing has been increasing. As an advantage of using video conferencing, there is an expected effect of improving the accuracy of information transmission by visually sharing images among multiple users. In such a use case, based on the shared image, image alignment may be performed among the persons in charge while supplementing in conversation what kind of changes should be made to the subject of the discussion shown in the image. Against this background, various techniques have been proposed for creating new images or modifying existing images from the information of conversations made among the persons in charge and the text information input by the persons in charge. Non-Patent Document 1 discloses a technique for receiving an input of text called a prompt and newly creating and outputting an image based on the semantic information of the text. Non-Patent Document 2 discloses a technique for receiving an input of a base image and text information for modifying the image, and outputting an image in which style conversion is performed so that the modification indicated by the text information is applied to the image.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] On the other hand, in the technology disclosed in Non-Patent Document 1, an image with a plausible modification is created based on the input text information. Therefore, an image that exactly reflects the user's intention is not always created. As an example of a technology for solving such problems, the technology disclosed in Non-Patent Document 2 can be cited. However, in the technology disclosed in Non-Patent Document 2, since text information for modifying an image is sequentially prepared, it is troublesome for the user.

[0005] In view of the above problems, an object of the present invention is to be able to reproduce the form of an object indicated by the content spoken by a user in a more suitable manner.

Means for Solving the Problems

[0006] An information processing apparatus according to the present invention includes: an acquisition unit that acquires, based on speech data indicating the speech content of a user, subject object information indicating an object that is the subject of the speech content, and form modification factor information that is information for modifying the form of the object included in the speech content; a change unit that makes a change to an object corresponding to the subject object information acquired by the acquisition unit among one or more objects included in an image that is the target of the user's speech, based on the form modification factor information; and an output control unit that controls so that a result of the change made to the object by the change unit is output to a predetermined output destination.

Effects of the Invention

[0007] According to the present invention, it becomes possible to reproduce the form of an object indicated by the content spoken by a user in a more suitable manner.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0009] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted.

[0010] <First Embodiment> The first embodiment of the present disclosure will be described below. FIG. 1 is a diagram showing an example of the hardware configuration of the information processing apparatus according to the present embodiment. The information processing apparatus 100 includes a CPU (Central Processing Unit) 104, a RAM (Random Access Memory) 105, and a ROM (Read Only Memory) 106. The information processing apparatus 100 also includes an input unit 101, a display unit 102, an image input unit 103, and an HDD (Hard Disk Drive) 107. Each component of the information processing apparatus 100 described above is connected via a data bus 108 so as to be able to transmit and receive data to and from each other.

[0011] The CPU 104 reads out the control computer program stored in the ROM 106 and expands it in the RAM 105, and executes various control processes based on the program. The RAM 105 is used as an area for expanding the program executed by the CPU 104 and a temporary storage area such as a work memory. The image input unit 103 serves as an interface for receiving image data from the outside. As a method of receiving image data, for example, a method of receiving from an external device such as an imaging device via a transmission path such as a cable, a method of receiving from another device via a network such as the Internet, a method of receiving screen information displayed on the display unit, etc. can be applied. The HDD 107 stores various data such as image data and setting parameters, and various programs.

[0012] The image data received via the image input unit 103 is transmitted to the CPU 104, the RAM 105, and the ROM 106 via the data bus 108. Also, by the CPU 104 executing the information processing program stored in the ROM 106 or the HDD 107, information processing on input data (for example, image data) is realized. Also, the data received from an external device via the image input unit 103 may be stored in the HDD 107. The input unit 101 serves as an input interface for receiving input of information from the user. The input unit 101 may include, for example, input devices such as a keyboard, a pointing device such as a mouse, and a touch panel. Also, the input unit 101 may include a voice input device such as a microphone. The display unit 102 serves as an output interface for presenting information to the user. The display unit 102 may include, for example, a display device such as a liquid crystal display.

[0013] Referring to FIG. 2, an example of the functional configuration of the information processing apparatus according to the present embodiment will be described. The information processing apparatus 100 according to the present embodiment includes an image acquisition unit 201, a speech data acquisition unit 202, a text generation unit 203, a subject object extraction unit 204, a body parameter generation unit 205, a body update unit 206, and an output control unit 207.

[0014] The image acquisition unit 201 acquires image data to be processed. The image data to be processed may be image data acquired by an external device such as an imaging device, image data stored in a storage device such as a hard disk, or image data received via a network such as the Internet. The image acquisition unit 201 outputs the acquired image data to the subject object extraction unit 204.

[0015] The speech data acquisition unit 202 acquires speech data indicating the speech content of a user (speaker). The speech data includes information indicating the user (speaker) as metadata (hereinafter also referred to as user information), and is data in which the speech content of the user is shown as text information. The number of users for which speech data is to be acquired is not particularly limited as long as there is one or more. Also, when there are a plurality of target users, it is preferable that speech data is acquired for each user. In this case, it is possible to identify which user's speech content each speech data indicates based on the above user information. Regarding the text information indicating the speech content of the user included in the speech data, the acquisition method is not particularly limited. For example, based on text data input by the user via an input device such as a keyboard, text information indicating the speech content of the user may be acquired. Also, as another example, by converting the speech of the user input via a sound collection device such as a microphone into text data, text information indicating the speech content of the user may be acquired based on the text data. Also, regarding the acquisition of text information indicating the speech content of the user using the various input interfaces exemplified above, it may be acquired in real time in synchronization with the information processing device, or may be acquired by reading out pre-stored data. Also, the acquisition method of the data (for example, text data, voice data, etc.) that is the source of the text information indicating the speech content of the user is not particularly limited. For example, the target data may be acquired from an external device connected via a transmission path such as a cable, or the target data may be acquired from an external device via a network. In addition to the text information indicating the user's speech content and the user information, the speech data may also include, as metadata, information indicating the position where the target speech was made in a series of speeches, such as the speech time (e.g., the position in time series, the position in context, etc.). In this way, by including, as metadata, information indicating the position where the target speech was made in a series of speeches, it becomes possible to manage the speech content of the user indicated by the speech data by dividing it into clauses and in accordance with the order of each clause. Hereinafter, for the sake of convenience, various explanations will be made assuming that the speech time is applied as information indicating the position where the target speech was made in a series of speeches. The speech data acquisition unit 202 outputs the acquired speech data to the text generation unit 203.

[0016] Based on the speech data, the text generation unit 203 generates subject object information and morphological modification factor information. Hereinafter, each of the subject object information and the morphological modification factor information will be described in more detail.

[0017] The subject object information is information regarding an object (e.g., an object in an image, etc.) mentioned as the subject of the speech content indicated by the speech data. The subject object information can be configured as structured data including, for example, information such as the name of the object that is the subject of the speech content, the estimated likelihood as the subjectiveness of the object, and the corresponding speech time in the speech data. Also, as another example, the subject object information may be configured as a list including one or more of the above structured data. Although details will be described later, when the subject object extraction unit 204 uniquely determines the subject object, it is preferable to refer to the estimated likelihood as the subjectiveness in the context indicating the user's speech content.

[0018] The physical modification factor information is information regarding the modification of the form of the subject object mentioned in the speech content indicated by the speech data. For example, information regarding the proposal for modifying the form of the subject object may be applicable. Modifying the form means, for example, adding, duplicating, deleting, moving, deforming, etc. to the appearance such as shape, color, or texture. The physical modification factor information may include, for example, the summary name of the physical modification factor, the name of the associated subject object, the estimated likelihood as the physical modification factor likelihood, the corresponding speech time in the speech data, user information (proposer), the status regarding the feasibility of implementing the physical modification, the attributes of the physical modification factor, etc. Also, the physical modification factor information may be a list consisting of structured data including the above-mentioned each information. The physical modification factor is associated with one subject object. Also, a plurality of physical modification factors may be associated with the subject object. For example, it can generally be assumed that multiple changes are made to one form. Therefore, in such a case, for the form of the subject object, physical modification factors corresponding to the changes for each of the multiple locations are associated.

[0019] The estimated likelihood as the physical modification factor likelihood is the likelihood indicating whether the target expression (spoken information) is generally an expression related to modifying the form. For example, if the physical modification factor is an expression such as "round the corners", it is generally interpreted as an expression related to form change, so a higher value of likelihood is output. On the other hand, if the physical modification factor is an expression such as "turn gently", it is generally interpreted as not being an expression for describing / modifying the form, so a lower value of likelihood is output.

[0020] The state regarding the feasibility of form modification is information indicating the feasibility of the corresponding form modification according to a statement when there is a statement regarding the feasibility of the target form modification in a series of statements (for example, a series of statements made in a discussion among users). For example, after the first user proposes a change to a certain form for a certain subject object, assume that the second user makes a statement indicating that the change to the form cannot be allowed. In such a situation, the text generation unit 203 extracts a form modification factor indicating the change to the form based on the above statement by the first user, and also extracts a determination as to whether the implementation of the change to the form is affirmative or negative based on the above statement by the second user. The state regarding the feasibility of form modification may be a true / false value of 0 / 1 or a numerical value.

[0021] The attribute of the form modification factor is information indicating which of a plurality of changes including addition, deletion, and deformation to the form of the subject object the target modification process performs. Regarding the attribute of the form modification factor, for example, it is referred to when the form parameter generation unit 205 described later induces different processes for each attribute.

[0022] The text generation unit 203 can be realized, for example, by a natural language generation model based on a neural network that can process a context of sufficient length. By applying such a configuration, for example, it becomes possible to process a text-to-text conversion task such as generating subject object information and form modification factor information from the input speech data. Of course, the above is merely an example, and the method is not particularly limited as long as it is possible to extract or generate information corresponding to the subject object information and form modification factor information from the speech content indicated by the input speech data. Among the series of generated data, the text generation unit 203 outputs the subject object information to the subject object extraction unit 204 and outputs the form modification factor information to the form parameter generation unit.

[0023] Based on the image data acquired by the image acquisition unit 201 and the subject object information generated by the text generation unit, the subject object extraction unit 204 extracts the object indicated by the subject object information from the image shown by the image data. In the following description, for convenience, the object indicated by the subject object information is also referred to as the subject object. The extraction result of the subject object by the subject object extraction unit 204 includes, for example, the image of the subject object included in the image shown by the image data (the partial image of the area of the subject object) and the position information of the position where the subject object exists in the image. The method for extracting the image of the subject object from the image shown by the image data is not particularly limited. For example, a method of cutting out along the contour of the area of the subject object by segmentation may be applied, or a method of cutting out a rectangle enclosing the area of the subject object may be applied. Regarding the position information of the subject object, for example, it may be defined as the position information of the rectangle enclosing the area of the subject object.

[0024] When the subject object information is a list composed of a plurality of pieces of structured data, the subject object extraction unit 204 uniquely identifies the subject object information with a higher likelihood included as metadata in the subject object information. Moreover, the subject object extraction unit 204 may output the extraction result of the subject object based on the identified subject object information. By applying such control, when there are multiple candidates for the subject object, it is possible to extract the object with the highest likelihood of being the subject object from among the multiple candidates and filter out (exclude from the extraction target) other objects. For example, when all the estimated likelihoods included in the subject object information are smaller than a certain value (threshold value), the subject object extraction unit 204 may determine that there is no subject object. As another example, when the statistic of the likelihood at the time of segmentation in the extraction result of the subject object is smaller than a certain value (threshold value), the subject object extraction unit 204 may determine that there is no subject object. This process is for the case where the speech data does not include a subject object, and a situation where there is no (or extremely low) relevance between the target image data and the user's speech may apply. As a specific example, a situation where a conversation is made among a plurality of users who are acquisition targets of the speech data and there is no relevance to the target image data may apply. The subject object extraction unit 204 can be realized by, for example, a model based on a neural network capable of executing a task of searching for and extracting an object corresponding to text specified from within an image. Further, for the model applied to the subject object extraction unit 204, in order to exhibit high extraction accuracy even in the zero-shot case for a variety of subject objects, it is desirable that the model be learned based on a pair of large-scale natural language data and image data. Zero-shot corresponds to executing a task for a label having a new class that the model has never learned. The subject object extraction unit 204 outputs the extraction result of the subject object to the form parameter generation unit 205 and the form update unit 206.

[0025] The form parameter generation unit 205 generates form modification parameters based on the form modification factor information generated by the text generation unit 203 and the extraction result of the subject object by the subject object extraction unit 204. The shape modification parameter is a parameter that represents the shape characteristics based on the modification factor information. Specifically, the shape modification parameter is structured data that includes geometric shape information such as the position, size, shape, and orientation of the shape characteristics generated based on the shape modification factor, and appearance information such as color and texture. Further, the shape modification parameter may inherit some or all of the structured information of the shape modification factor that is the generation source. The shape parameter generation unit can be realized, for example, by a neural network for multi-modal input that outputs the above-described structured data with text and images as inputs.

[0026] Among the shape modification parameters that represent the shape characteristics, regarding the position, size, etc., since they are restricted by the shape of the subject object to be input, especially the contour information, it is desirable that the parameters be controlled by the subject object image. A specific example will be described below with respect to this point. Suppose that for the case where the width of the subject object is 30 mm and the case where it is 1000 mm, an input of the shape modification factor "add bolt holes" is made. Since the shape modification factor does not particularly include a modifier regarding size, as the size of the shape modification parameter, for example, a setting of "ordinary" indicating a generally applicable size will be made. On the other hand, generally, it is assumed that the user assumes a size of about 10 mm or less for the shape with a width of 30 mm and about 50 mm or less for the shape with a width of 1000 mm regarding the diameter of the "bolt hole". However, in the shape update unit 206 described later, it is difficult to correspond to the variation in the size of the shape characteristics to be generated to 10 mm or 50 mm for the size indication by the shape modification parameter of "ordinary". In view of such a situation, the shape parameter generation unit 205 may control the shape modification parameter to be generated using information such as the subject object image.

[0027] The shape parameter generation unit 205 may execute different tasks depending on the attributes of the input shape modification factor information. For example, assume that shape modification factor information having an attribute indicating the addition of a shape feature to the shape of the subject object, such as "adding bolt holes to a flat plate portion", is input. In this case, the shape parameter generation unit 205 may execute, for example, a task of detecting the flat plate portion from the subject object image and a task of outputting map information of locations where bolt holes can be added, and determine shape modification parameters such as the position and size for adding bolt holes. As another example, assume that shape modification factor information having an attribute indicating the deletion of a shape feature from the shape of the subject object, such as "deleting bolt holes from a flat plate portion", is input. In this case, the shape parameter generation unit 205 may execute, for example, only a task of detecting bolt holes in the flat plate portion from the subject object image. The shape parameter generation unit 205 outputs the generated shape modification parameters to the shape update unit 206.

[0028] The shape update unit 206 performs at least any one of a plurality of changes including addition, deformation, and deletion of shape features on the subject object extracted by the subject object extraction unit 204 based on the shape modification parameters generated by the shape parameter generation unit 205. The shape update unit 206 can be realized, for example, by a diffusion neural network model that performs image generation and is trained to minimize the distance in the feature space between the original image and the generated image in order to retain the shape features of the original subject object. Then, the shape update unit 206 outputs to the output control unit 207 the shape of the subject object (hereinafter, also referred to as the updated shape) updated by making changes to the subject object based on the shape modification parameters.

[0029] The output control unit 207 controls so that information indicating the shape of the subject object (that is, the updated shape of the subject object) modified by the shape update unit 206 is output to a predetermined output destination. For example, the output control unit 207 may display the updated form of the subject object in a predetermined display area (e.g., the display unit 102). At this time, the output control unit 207 may display the image indicated by the image data acquired by the image acquisition unit 201, the speech data acquired by the speech data acquisition unit 202, and the updated form of the subject object in the predetermined display area. Also, in this case, the output control unit 207 may selectively switch the displayability of each of the above-exemplified information based on a designation from the user. Further, the output control unit 207 can also execute an interactive process based on an instruction from the user via the input unit 101 for the updated form of the subject object displayed in the predetermined display area. Details of an example of the interactive process will be described separately later. Also, the above is merely an example, and does not limit the output destination of the information indicating the updated form of the subject object by the output control unit 207. As a specific example, the output control unit 207 may output information indicating the updated form of the subject object as an object of the image processing and the analysis processing to a device that performs various image processing and various analysis processing on the target data.

[0030] With reference to FIG. 3, an example of the processing of the information processing apparatus 100 according to the present embodiment will be described. In S101, the image acquisition unit 201 acquires image data to be processed. In S102, the speech data acquisition unit 202 acquires speech data indicating the speech content of the user input via a predetermined input interface. Note that the speech data acquisition unit 202 is not limited to only one user, and may acquire speech data based on the speech of each of a plurality of users. In S103, the text generation unit 203 generates subject object information and form modification factor information based on the speech data acquired in S102. In S104, based on the image data acquired in S101 and the subject object information generated in S103, the subject object extraction unit 204 extracts the subject object indicated by the subject object information from the image shown in the image data. In S105, based on the form modification factor information generated in S103 and the extraction result of the subject object in S104, the form parameter generation unit 205 generates form modification parameters. In S106, the form update unit 206 makes a change to the subject object extracted in S104 by adding, changing, or deleting form features based on the form modification parameters generated in S105. As a result, the form of the subject object is updated. In S107, the output control unit 207 controls so that the information indicating the form of the subject object (that is, the form of the subject object after update) changed in S106 is output to a predetermined output destination. As a specific example, the output control unit 207 may present the updated form to the user by displaying the form of the subject object after update in a predetermined display area.

[0031] Referring to FIG. 4, an example of a screen presented by the information processing apparatus 100 according to the present embodiment to the user via the display unit 102 will be described. The display screen 400 shown in FIG. 4 includes a target image display unit 410, a subject object image display unit 420, an update result display unit 430, a speech content display unit 440, and a form modification factor display unit 450.

[0032] The target image display unit 410 is a display area in which the image shown by the image data acquired by the image acquisition unit 201 is displayed. In the example shown in FIG. 4, four objects, namely the electrical equipment box 301a, the sphere 301b, the reinforcing plate 301c, and the workbench 301d, are imaged as subjects in the image. In the example shown in FIG. 4, for the sake of convenience, it is assumed that the electrical equipment box 301a has been determined as the subject object based on the speech from the user whose input was previously received. Also, in the example shown in FIG. 4, in order to clarify the extraction result of the subject object, a rectangular box 310 (so-called bounding box) is displayed so as to enclose the area of the electrical equipment box 301a in the image displayed on the target image display unit 410. The target image display unit 410 corresponds to an example of the first partial area.

[0033] The speech content display unit 440 is a display area where the speech content of the user indicated by the speech data acquired by the speech data acquisition unit 202 is displayed. In the example shown in FIG. 4, the speeches 341a and 341b corresponding to the two pieces of speech data respectively are arranged and displayed such that the older one is presented higher and the newer one is presented lower in chronological order. As described above, the speech data includes text information indicating the speech content and information indicating the speaker of the target speech. Each of the speeches 341a and 341b is displayed based on the text information indicating the speech content included in the target speech data. In the example shown in FIG. 4, in speech 341a, three proposals are made: "Add as large as possible to the space created by moving the round hole to the center of the surface of the aperture shape", "Move the rightmost round hole as far to the right as possible", and "Add hemming bending to the lower bend". Regarding these proposals, in speech 341b, agreement is reached on the first and second proposals, and the third proposal is rejected.

[0034] The form modification factor display unit 450 is a display area where the form modification factor indicated by the form modification factor information generated based on the speech data is displayed. First, the name of the subject object included in the subject object information determined as the target is displayed in the display area 351 in the form modification factor display unit 450. Also, the summary name included in the form modification factor information related to the subject object is displayed on the form modification factor display unit 450. In the example shown in FIG. 4, summary names 352a, 352b, and 352c corresponding to the three form modification factor information are displayed. Also, the estimation result of the state regarding the feasibility of performing form modification (changing the object) for each of the three form modification factor information by the text generation unit 203 is reflected in the check boxes 452a, 452b, and 452c that are displayed in association with each form modification factor information. Also, in the example shown in FIG. 4, regarding the proposal of "adding hemming bending to the lower bending" shown as the summary name 352c, it has been rejected by the statement 341b. Therefore, for this proposal, the status of the feasibility of the form modification factor is set to "no", and the check box 452c is displayed in the OFF state. On the other hand, regarding the proposals corresponding to the summary names 352a and 352b respectively, an agreement has been reached in the statement 341b. Therefore, for these proposals, the status of the feasibility of the form modification factor is set to "yes", and the check boxes 452a and 452b are each displayed in the ON state. Each of the check boxes 452a, 452b, and 452c can be arbitrarily changed between the ON / OFF states based on an instruction from the user via the input unit 101. Based on the change in the state regarding the feasibility of the form modification factor via these check boxes as a trigger, a process of changing the updated form of the subject object displayed on the update result display unit 430 may be executed.

[0035] The subject object image display unit 420 is a display area where an image of the subject object extracted from the acquired image data is displayed based on the extraction result of the subject object. In the example shown in FIG. 4, an image of the electrical equipment box determined as the subject object is extracted and displayed from the image shown by the acquired image data. The subject object image display unit 420 corresponds to an example of the second partial area.

[0036] The update result display unit 430 is a display area where an image of the form of the subject object (the form of the subject object after update) with changes made by adding, changing, or deleting form features to the subject object based on the form modification factor is displayed. In the example shown in FIG. 4, for the subject object image 320, the results of two changes, "move the rightmost round hole as far to the right as possible" and "add the aperture shape as large as possible to the center of the surface", are displayed as the updated form 330. The update result display unit 430 corresponds to an example of the third partial area.

[0037] Note that for each component (each display unit indicated by reference numerals 410 to 450) constituting the display screen 400, it may be possible to arbitrarily switch the ON / OFF of the display state individually according to an instruction from the user.

[0038] FIG. 5 is a schematic diagram showing an example of the display state of the target image display unit 410 in the display screen 400 shown in FIG. 4. Specifically, FIG. 5 schematically shows a situation where in addition to objects 301a to 301d that are candidates for the subject of the speech content of each user, a finger 302 is imaged as another object. Taking the state shown in this FIG. 5 as an example, with reference to FIG. 6, an example of the process of the subject object extraction unit 204 will be described below. The subject object extraction unit 204 uses feature amounts that can be extracted from an image, such as spatial feature amounts and modality feature amounts, in order to further improve the estimation accuracy of the subject object in the image.

[0039] In S111, the image acquisition unit 201 receives input of image data to be processed from the user. Also, the speech data acquisition unit 202 acquires speech data indicating the speech content of the user. The text generation unit 203 generates subject object information and form modification factor information based on the speech data acquired by the speech data acquisition unit 202.

[0040] In S112, based on the image data and the subject object information obtained in S111, the subject object extraction unit 204 extracts the subject object indicated by the subject object information from the image shown in the image data (hereinafter also referred to as the input image). For example, in the example shown in FIG. 5, when the name of the obtained subject object is "box", the subject object extraction unit 204 estimates each of the objects 301a and 301c as a candidate for the subject object with a higher likelihood than other objects. In such a case, the subject object extraction unit 204 may use various feature amounts such as spatial feature amounts and modality feature amounts in order to more accurately and uniquely determine the subject object. In the example shown in FIG. 6, it is assumed that the subject object extraction unit 204 uniquely determines the subject object by using the spatial feature amount and the modality feature amount.

[0041] The spatial feature amount is a feature amount based on the idea that the object that is the subject appears larger, more centrally, and more clearly (without out-of-focus) in the image. In the example shown in FIG. 6, it is assumed that the subject object extraction unit 204 has a spatial feature amount encoder (not shown). In S113, the subject object extraction unit 204 inputs the subject object information to the spatial feature amount encoder, and obtains the spatial feature amount as the output of the spatial feature amount encoder.

[0042] The modality feature amount is a feature amount based on the idea that when there is an indication object indicating the subject object in the image, the modality of the indication object is used to accurately identify the subject object. In the example shown in FIG. 6, it is assumed that the subject object extraction unit 204 has a modality feature amount encoder (not shown). In S114, the subject object extraction unit 204 detects the finger 302 as an instruction object. In this case, there is a high possibility that the subject object exists at the position and in the direction indicated by the finger 302 detected as the instruction object. Therefore, in S115, the subject object extraction unit 204 uses the modality feature encoder to obtain, as modality features, a subject object existence probability map that outputs a high score at a specific position and in a specific direction based on, for example, the gesture of the finger.

[0043] In S116, the subject object extraction unit 204 uniquely determines the subject object based on the subject object information acquired in S111, the spatial features acquired in S113, and the modality features acquired in S115. In the case of the example shown in FIG. 5, since the object 301a is imaged larger than the object 301c and is located in a region with a higher subject object existence probability based on the modality of the finger, the possibility of being the subject object is higher. Therefore, in this case, the subject object extraction unit 204 determines the object 301a as the subject object among the objects 301a and 301c that are candidates for the subject object. Note that the above is merely an example, and the features used for determining the subject object are not particularly limited as long as they are features that can be extracted from the target image. As a specific example, the image quality of each of a series of objects included in the image (for example, a feature serving as an index for evaluating the image quality) may be used as a feature. In this case, for example, among a series of objects included in the image, an object with higher image quality may be determined as the subject object.

[0044] In S117, the subject object extraction unit 204 outputs the extraction result of the subject object (object 301a), which includes the position information of the subject object determined in S116, to a predetermined output destination. Note that the position information of the subject object can be output, for example, as the position information of a rectangle (bounding box) that encloses the subject object.

[0045] By applying the control as described above, an effect of further improving the accuracy of extracting the subject object can be expected.

[0046] Referring to FIG. 7, another example of the display state of the target image display unit 410 in the display screen 400 will be described. In the image displayed on the target image display unit 410, in addition to objects 301a to 301d, rectangles 311a and 311b that enclose objects 301a and 301c, which are candidates for the subject object based on the subject object information, are displayed. The user can select, via the input unit, either object 301a or 301c, which is a candidate for the subject object, as the subject object by point 401. Taking the state shown in this FIG. 7 as an example, referring to FIG. 8, as another example of the processing of the subject object extraction unit 204, the processing when the subject object extraction unit 204 uniquely determines the subject object via user input will be described.

[0047] In S121, the image acquisition unit 201 receives input of image data to be processed from the user. Also, the speech data acquisition unit 202 acquires speech data indicating the speech content of the user. The text generation unit 203 generates subject object information and morphological modification factor information based on the speech data acquired by the speech data acquisition unit 202.

[0048] In S122, based on the image data and the subject object information obtained in S121, the subject object extraction unit 204 extracts the subject object indicated by the subject object information from the image indicated by the image data (hereinafter also referred to as the input image). For example, in the example shown in FIG. 7, when the name of the obtained subject object is "box", the subject object extraction unit 204 estimates each of the objects 301a and 301c as a candidate for the subject object with a higher likelihood than other objects.

[0049] In S123, based on the estimation result of the candidates for the subject object in S122, the output control unit 207 causes the target image display unit 410 to display rectangles 311a and 311b each enclosing the objects 301a and 301c that are candidates for the subject object. In S124, the subject object extraction unit 204 accepts from the user a selection of either of the rectangles 311a and 311b displayed on the target image display unit 410 by a point 401 operated by the user via the input unit. In S125, the subject object extraction unit 204 specifies the object corresponding to the rectangle selected in S124 as the subject object, and outputs the extraction result of the subject object including the position information of the subject object to a predetermined output destination. Note that the position information of the subject object can be output, for example, as the position information of a rectangle (bounding box) enclosing the subject object.

[0050] By applying the control as described above, it becomes possible to more reliably specify the subject object intended by the user.

[0051] Referring to FIG. 9, an example of the processing of the text generation unit 203 will be described with a focus on the processing related to the extraction of the shape modification factor. In the example shown in FIG. 9, the text generation unit 203 performs filtering of shape modification factors with low likelihood that are faithful to the user's intention.

[0052] In S131, the text generation unit 203 receives speech data (for example, the speech data acquired by the speech data acquisition unit 202) as input. In S132, the text generation unit 203 generates subject object information based on the speech data received as input in S131. In S133, the text generation unit 203 generates body modification factor information based on the speech data received as input in S131. The generated body modification factor information includes the estimated likelihood as the body modification factor likeness. As a specific example, assuming that two body modification factors, "move the round hole to the right" and "turn quietly", are extracted from the speech data, the likelihood as the body modification factor is higher for the former and lower for the latter. In S134, the text generation unit 203 compares the likelihood included in each body modification factor information generated in S133 with a preset threshold value, and deletes the body modification factor information whose likelihood is lower than the threshold value. In S135, if there is body modification factor information remaining without being deleted as a result of the process in S134, the text generation unit 203 outputs the body modification factor information to a predetermined output destination.

[0053] As described above, by filtering the body modification factor information by the text generation unit 203, subsequent processing for the body modification factor information deleted from the processing target is omitted, and a reduction in processing load and an improvement in processing speed can be expected. In addition, since the information displayed in the body modification factor display unit 450 in the display screen 400 is simplified, an effect of improving the convenience for the user can be expected.

[0054] On the other hand, in the method exemplified above, a situation may be assumed where, although the relevance to the subject is low, generally, phrases that modify the form are extracted from the speech data. For example, assume that the subject object is an "electrical equipment box", and the text generation unit 203 generates form modification factor information corresponding to two form modification factors, namely, "move the round hole to the right side" and "add a taper to the edge of the gear". In such an example, even if the subject is an "electrical equipment box" and the form modification factor "add a taper to the edge of the gear" is irrelevant to the subject, form modification factors derived from words related to form modification may be extracted with a higher likelihood. In view of such a situation, another example of a method related to filtering form modification factors is proposed below.

[0055] FIG. 10 is a flowchart showing an example of the processing of the form update unit 206 according to the present embodiment. In the example shown in FIG. 10, the form update unit 206 filters form modification factors with a lower estimated likelihood as the likelihood of being a form modification factor in order to faithfully update (make a change to) the form of the subject object according to the user's intention. Further, FIG. 11 is a schematic diagram showing an example of a detection likelihood map generated by the form update unit 206 according to the present embodiment. The detection likelihood map is a map showing where and with what probability the detection target exists in the image. Details of the detection likelihood map will be described separately later. In the example shown in FIG. 11, the subject object image 320 includes round hole portions 321a to 321c and flat portions 322a and 321b. Also, in the example shown in FIG. 11, the name of the subject object is an "electrical equipment box", and each of FIGS. 11(a), 11(b), and 11(c) shows a summary name of different form modification parameters. Specifically, FIG. 11(a) corresponds to an example of a form modification parameter showing "move the hole", FIG. 11(b) corresponds to an example of a form modification parameter showing "add a constriction shape to the center of the surface", and FIG. 11(c) corresponds to an example of a form modification parameter showing "add a taper to the edge of the gear".

[0056] In S141, the shape update unit 206 receives, as inputs, the extraction result of the subject object including the subject object image and the shape modification parameter. The shape modification parameter inherits the shape modification factor information and also includes the summary name and attributes included in the shape modification factor information.

[0057] In S142, regardless of whether the attribute of the shape modification parameter is addition, change, or deletion, the shape update unit 206 executes a detection task for the shape to be processed and generates a detection likelihood map based on the execution result of the detection task. For example, in the case of the example shown in FIG. 11(a), the shape update unit 206 detects each of the round hole portions 321a to 321c from the subject object image for the shape modification parameter of "move the hole". At this time, the detection likelihood of the location corresponding to each of the round hole portions 321a to 321c is set to a higher value than other locations. In the figure, the hatched portions indicate regions with a higher detection likelihood (for example, regions where the detection likelihood is equal to or greater than the threshold), and the non-hatched portions indicate regions with a lower detection likelihood (for example, regions where the detection likelihood is less than the threshold). Also, in the case of the example shown in FIG. 11(b), the shape update unit 206 detects each of the flat portions 322a and 322b from the subject object image for the shape modification parameter of "add a constriction shape to the center of the surface". At this time, the detection likelihood of the location corresponding to each of the flat portions 322a and 322b is set to a higher value than other locations. Also, in the case of the example shown in FIG. 11(c), the shape update unit 206 executes a process of detecting a gear from the subject object image for the shape modification parameter of "add a taper to the edge of the gear". On the other hand, in the example shown in FIG. 11, since there is no gear, a lower value is set as the detection likelihood over the entire subject object image.

[0058] In S143, the shape update unit 206 calculates a predetermined statistic (e.g., the maximum value) based on each of the series of detection likelihood maps generated in S142, and deletes the shape modification parameter corresponding to the detection likelihood map when the statistic is less than or equal to a threshold value. In S144, if there are shape modification parameters that remain without being deleted as a result of the process in S143, the shape update unit 206 updates the shape of the subject object based on the shape modification parameters, and outputs the result of the update to a predetermined output destination.

[0059] Through the series of processes as described above, the subject object image is collated with the shape modification parameters, and a filtering process using the statistic of the detection likelihood map of the shape modification parameters is executed. This prevents the occurrence of a situation where a shape to which a shape modification factor not originally intended by the user is applied is output, and it becomes possible to update the shape of the subject object more faithfully according to the user's intention.

[0060] Next, an example of a method for generating shape modification parameters that more accurately reflect the user's intention through interactive processing via the display unit 102 and the input unit 101 will be described below.

[0061] FIG. 12 is a diagram showing another example of a screen presented by the information processing apparatus 100 according to the present embodiment to the user via the display unit 102. The display screens 400 shown in FIGS. 12(a) to 12(d) each include an update result display unit 430 and a shape modification factor display unit 450. Also, a pointer 401 used to specify a part within the display screen 400 is displayed based on a user operation received via the input unit 101. In the example shown in FIG. 12, three shape modification factors 352a to 352c are displayed in the shape modification factor display unit 450. The shape modification factor 352a indicates a shape modification of "move the rightmost hole". The shape modification factor 352b indicates a shape modification of "add a large aperture shape to the center of the surface". The shape modification factor 352c indicates a shape modification of "add a hemming bend to the lower bend".

[0062] In the example shown in FIG. 12(a), the user can select any form modification factor in the form modification factor display unit 450. For example, in the example shown in FIG. 12(a), assume that the user selects a form modification factor 352a indicating a form modification of "move the rightmost hole". In this case, a command list 353a for executing a predefined process for the selected form modification factor 352a is displayed. In the example shown in FIG. 12(a), commands for applying processes such as "move", "copy", "delete", and "change shape" to the form features based on the form modification factor are displayed for the command list 353a. Also, more detailed conditions may be specifiable for at least some of the commands. For example, in the example shown in FIG. 12(a), a text box 354a for receiving a specification of a new shape as text information is displayed for the "change shape" command. With such a configuration, it becomes possible to enjoy substantially the same usability as when using an information processing apparatus in which a form is automatically generated based on the above-described speech data.

[0063] In the example shown in FIG. 12(b), an example of a state in which a command indicating "move" is selected by the user for the form modification factor 352a indicating a form modification of "move the rightmost hole" is schematically shown. In a state where the command indicating "move" is selected, the arrangement of the round hole portion 331, which is a form feature based on the target form modification factor 352a, can be moved by a drag operation with the pointer 401 or the like. In the example shown in FIG. 12(b), the round hole portion 331' included in the form 330 of the updated subject object indicates the round hole portion 331 after being moved by the above operation.

[0064] In the example shown in FIG. 12(c), an example of a state in which a command indicating "change of shape" is selected by the user for the morphological modification factor 352a indicating "move the rightmost hole" is schematically shown. In a state where a command indicating "change of shape" is selected, the contour line of the round hole portion 331, which is a morphological feature based on the target morphological modification factor 352a, can be moved by a drag operation or the like using the pointer 401. In the example shown in FIG. 12(c), the round hole portion 331'' of the morphology 330 of the updated subject object indicates the round hole portion 331 after its size is changed by pulling the contour line of the round hole portion 331 outward.

[0065] In the example shown in FIG. 12(d), another example of a state in which a command indicating "change of shape" is selected by the user for the morphological modification factor 352a indicating "move the rightmost hole" is schematically shown. In a state where a command indicating "change of shape" is selected, an instruction regarding the change of the shape of the round hole portion 331 can be given via the text box 354a. In the example shown in FIG. 12(d), an instruction to double the size of the round hole portion 331 is given via the text box 354a. Also, the round hole portion 331''' of the morphology 330 of the updated subject object indicates the round hole portion 331 after its size is changed based on the instruction input to the text box 354a.

[0066] By applying the control as described above, the user can set the geometric morphological information such as the position, size, shape, and orientation of the morphological features based on the morphological modification factor in the manner intended by the user. Also, although the above mainly focused on the case where the geometric morphological information is corrected, the correction target is not limited to only the geometric morphological information. As a specific example, for appearance information such as color and texture, it is also possible to make it a correction target by applying control substantially the same as the above-described content.

[0067] <Second Embodiment> A second embodiment of the present disclosure will be described below. In this embodiment, an example of a configuration will be described in which, after storing, in a predetermined storage area, the result of making a change to a subject object based on form modification factor information, the information stored in the storage area can be used afterwards. Also, in this embodiment, the description will focus on the parts that are particularly different from the above-described first embodiment, and detailed description of the parts that are substantially the same as the first embodiment will be omitted.

[0068] Referring to FIG. 13, an example of the functional configuration of the information processing apparatus according to this embodiment will be described. The information processing apparatus according to this embodiment is different from the information processing apparatus according to the first embodiment described with reference to FIG. 2 in that it has a storage unit 208. The storage unit 208 is a storage area for storing various data. Mainly two types of data are listed as the data stored in the storage unit 208. The first type of data is learning data for learning the correlation between form modification factors and form modification parameters for each user in order for the information processing apparatus 100 to more accurately reflect the user's intention when making a change to a subject object. The second type of data is history data indicating the update result, which is stored as a history so that the user can quickly confirm the result of updating the form of the subject object when the form features of the subject object are added, changed, or deleted.

[0069] First, referring to FIG. 14, an example of the processing of the information processing apparatus according to this embodiment will be described, focusing on the processing related to the storage of the above learning data. FIG. 14 shows an example of a series of processes executed after the process related to the generation of form parameters by the form parameter generation unit 205.

[0070] In S211, the form parameter generation unit 205 generates form modification parameters based on the subject object image received as input and the form modification factor information. In S212, the shape update unit 206 changes the shape of the subject object so as to have a shape feature according to the shape modification parameter generated in S211. In S213, the output control unit 207 causes the update result of the shape of the subject object in S212 to be displayed in a predetermined display area. As a specific example, the output control unit 207 may cause the display unit 102 to display the display screen 400, and cause the update result display unit 450 of the display screen 400 to display the update result of the shape of the subject object.

[0071] In S214, the shape parameter generation unit 205 receives a user operation related to the modification of the shape feature with respect to the update result of the shape of the subject object displayed in a predetermined display area in S213. At this time, as described above with reference to FIG. 12, the shape parameter generation unit 205 may receive an operation related to the modification of the shape feature of the subject object from the user through the interactive process via the display unit 102 and the input unit 101. In S215, the shape parameter generation unit 205 executes a process related to the modification of the shape modification parameter generated in S211 so as to follow the operation related to the modification of the shape feature of the subject object received from the user in S214. As a specific example, it is assumed that the addition of a shape feature of "add a hole" is made, and in the situation where the "shape" item of the shape modification parameter in the corresponding shape modification factor information is circular, the shape of the target hole is modified to a rectangle by the user. In this case, the shape parameter generation unit 205 changes the "shape" item of the shape modification parameter to a rectangle.

[0072] In S216, the shape parameter generation unit 205 determines whether the user operation has been completed. When the shape parameter generation unit 205 determines in S216 that the user operation has not been completed, the process proceeds to S214. In this case, the processes of S214 and S215 are executed again. As described above, until the user operation is completed, the shape parameter generation unit 205 accepts the user operation related to the modification of the shape characteristics of the subject object, and each time, executes the process related to the modification of the shape modification parameter according to the user operation. When the shape parameter generation unit 205 determines in S216 that the user operation has been completed, the process proceeds to S217.

[0073] In S217, the output control unit 207 stores, in the storage unit 208, as learning data, a set of the shape modification parameters modified by the processes of S214 to S216 and the shape modification factor that is the source of the shape modification parameters. The learning data generated as described above is classified into data for each user and then used for the learning of the shape parameter generation unit 205, so that the shape parameter generation unit 205 can generate shape modification parameters closer to the user's intention.

[0074] Here, with reference to FIG. 15, a specific example will be given to describe in more detail the process described with reference to FIG. 14. FIG. 15(a) schematically shows a subject object image to be processed. Assume that a user makes a statement indicating an instruction to "add a diaphragm" to the subject object image shown in FIG. 15(a). In this case, based on the speech data indicating the content of the user's statement, as shown in FIG. 15(b), the result of changing the shape of the subject object is displayed in a predetermined display area. Moreover, from the state illustrated in FIG. 15(b), the user further performs an operation related to the modification of the diaphragm shape, and as shown in FIG. 15(c), the result in which the modification of the diaphragm shape is reflected in the shape of the subject object is displayed. In FIG. 15(b), the form parameter generation unit 205 generates a form modification parameter indicating that the aperture shape has a circular shape, and the result of changing the form feature indicated by the form modification parameter with respect to the form of the subject object is displayed. On the other hand, since the form feature has been changed as illustrated in FIG. 15(c), it is clear that the aperture shape intended by the user is a rectangle in the example shown in FIG. 15. In response to such a result, in the storage unit 208, the target subject object image, the form modification factor, and the form modification parameter (the form modification parameter reflecting the correction based on the instruction from the user) are stored as learning data. By using this learning data for the learning of the form parameter generation unit 205, it becomes possible to teach the form parameter generation unit 205 that when an instruction to add an aperture shape to a box-shaped part is given, the shape of the form modification parameter is a rectangle.

[0075] With reference to FIG. 16, an example of the display screen of the information processing apparatus according to the present embodiment will be described. Here, it is assumed that the form parameter generation unit 205 that has been learned based on the learning data described with reference to FIG. 15 is applied. In the example shown in FIG. 16, the display screen 400 includes an update result display unit 430 and a form modification factor display unit 450. In the update result display unit 430, a "reinforcement plate" which is the subject object is displayed, and in the form modification factor display unit 450, an instruction of "add an aperture" is shown. Further, in the update result display unit 430, corresponding to the instruction of the form modification factor, a marker 336 indicating the center position for generating the form and a form candidate window 335 for selecting the shape, size, and orientation of the applied form from among the candidates are displayed. Further, in the form candidate window 335, for each candidate of the form to be selected, a likelihood corresponding to the estimated result of the probability that the candidate will be selected by the target user is displayed. In the example shown in FIG. 16, since the form parameter generation unit 205 has been learned by the learning data generated based on the conditions shown in FIG. 15, the estimated likelihood of a rectangular-shaped aperture is the highest for the instruction of "add an aperture". In this way, in the form candidate window 335, candidates with higher likelihood estimation results by the form parameter generation unit 205 learned from the learning data for each user are preferentially displayed. By applying such control, the user can more easily select a form candidate closer to their intention from among a series of candidates presented in the form candidate window 335.

[0076] Next, the history data stored in the storage unit 208 according to the update result of the form of the subject object will be described. FIG. 17 is a diagram showing an example of the display screen of the information processing apparatus according to the present embodiment. FIGS. 17(a) to 17(c) are diagrams showing the state of the display screen 400 in chronological order, where FIG. 17(a) shows the oldest state and FIG. 17(c) shows the newest state. Further, the display screen 400 shown in FIG. 17 includes an update result display unit 430 and a history display unit 460.

[0077] First, the state shown in FIG. 17(a) will be described. In the state shown in FIG. 17(a), only the form modification factor 361a is selected as the application target. Therefore, in the update result display unit 430, the form 330a of the subject object updated by reflecting the form characteristics indicated by the form modification factor 361a is displayed. At this time, information indicating that the form modification factor 361a has been applied is reflected in the history display unit 460. Also, in this case, the updated form 330a of the subject object and the form modification factor 361a are associated and stored in the storage unit 208 as history data. In the example shown in FIG. 17, in the history display unit 460, the selection state of each form modification factor and each form modification parameter is displayed by the selection state of the check box. Also, by switching the state to either the selected state or the non-selected state by operating the check box, it is selectively switched whether the target form modification factor or form modification parameter is applied.

[0078] Next, the state shown in FIG. 17(b) will be described. In the state shown in FIG. 17(b), the shape modification factors 361a and 361b and the form modification parameters 361a' and 361a'' are selected as the application targets. Therefore, in the update result display unit 430, the shape 330c of the updated subject object is displayed by reflecting the shape characteristics indicated by each of the shape modification factors 361a and 361b and each of the form modification parameters 361a' and 361a''. At this time, in the history display unit 460, information indicating that each of the shape modification factors 361a and 361b and each of the form modification parameters 361a' and 361a'' have been applied is reflected. Also, in this case, the updated shape 330c of the subject object, each of the shape modification factors 361a and 361b, and each of the form modification parameters 361a' and 361a'' are associated and stored in the storage unit 208 as history data.

[0079] Next, the state shown in FIG. 17(c) will be described. In the state shown in FIG. 17(c), the selection states of the check boxes for the form modification factors and shape modification parameters other than the shape modification factor 361a are released. Therefore, in the update result display unit 430, the shape 330a' of the updated subject object is displayed by reflecting the shape characteristics only for the shape modification factor 361a among the series of form modification factors and shape modification parameters presented in the history display unit 460. At this time, the output control unit 207 checks whether the history data of the updated shape 330a associated only with the shape modification factor 361a is stored in the storage unit 208. In the example shown in FIG. 17, in the state shown in FIG. 17(a), the history data of the updated shape 330a associated only with the shape modification factor 361a is stored in the storage unit 208. Therefore, the output control unit 207 causes the updated shape 330a to be displayed on the update result display unit 430 based on the history data stored in the storage unit 208 without passing through the shape update unit 206.

[0080] As a use case of this embodiment, in order for a user to compare and consider the form of an object of interest, a case can be envisioned where, while selectively switching the application or non-application of various previously specified form modification factors, the updated form of the object is confirmed. Even in such a case, according to this embodiment, when historical data is stored, it is not necessary to regenerate the corresponding form, so it is possible to speed up the processing related to the display of the form of the updated subject object. Also, since the form of the updated subject object is displayed based on the historical data, it is possible to ensure the reproducibility of the form of the subject object according to the selection state of the form modification factor.

[0081] <Other Embodiments> As described above, the embodiment examples have been detailed, but the present invention can be implemented, for example, as an embodiment in a system, device, method, program, or recording medium (storage medium), etc. Specifically, it may be applied to a system composed of a plurality of devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or it may also be applied to a device consisting of a single device. Also, it goes without saying that the object of the present invention is achieved by the following. That is, a recording medium (or storage medium) recording a software program code (computer program) that realizes the functions of the above-described embodiment is supplied to a system or device. Needless to say, such a storage medium is a computer-readable storage medium. Then, a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-described embodiment, and the recording medium recording the program code constitutes the present invention.

[0082] Also, the disclosure of this embodiment includes the following configurations, methods, and programs. (Configuration 1) Acquisition means for acquiring, based on utterance data indicating the utterance content of a user, subject object information indicating an object that is the subject of the utterance content, and morphological modification factor information that is information for modifying the form of the object included in the utterance content; among one or more objects included in an image that is the target of the user's utterance, for the object corresponding to the subject object information acquired by the acquisition means, modification means for making a change based on the morphological modification factor information; and output control means for controlling so that the result of the change made to the object by the modification means is output to a predetermined output destination. An information processing apparatus, characterized by having the above. (Configuration 2) Having generation means for generating a morphological modification parameter indicating a morphological feature of the object based on the morphological modification factor information, wherein the modification means makes a change to the object corresponding to the subject object information based on the morphological modification parameter. The information processing apparatus according to Configuration 1. (Configuration 3) The change made by the modification means includes at least one of adding the morphological feature to the object corresponding to the subject object information based on the morphological modification parameter, deforming at least a part of the object based on the morphological feature, and deleting the morphological feature from the object. The information processing apparatus according to Configuration 2. (Configuration 4) The morphological modification parameter includes at least one of morphological information including at least one of position, size, shape, and orientation, and appearance information including at least one of color and texture. The information processing apparatus according to Configuration 2 or 3. (Configuration 5) The morphological modification factor information includes, as an attribute, information indicating which of a plurality of changes including addition, deletion, and deformation to the object corresponding to the subject object information is to be performed, and the generation means controls the configuration of the morphological modification parameter to be generated according to the attribute included in the morphological modification factor information. The information processing apparatus according to any one of Configurations 2 to 4. (Configuration 6) The form modification factor information includes at least any one of the name of the object that is the subject of the speech content, the estimated likelihood as the subjectivity of the object, the feasibility of making a change to the object, and the information of the user who made the speech indicated by the speech content. The information processing apparatus according to any one of Claims 2 to 5, characterized in that it is included. (Configuration 7) The changing means suppresses the change based on the form modification factor information to the object corresponding to the subject object information when the estimated likelihood included in the form modification factor information is less than a threshold value. The information processing apparatus according to Configuration 6, characterized in that it is included. (Configuration 8) The changing means is a likelihood map indicating where and with what probability the object corresponding to the subject object information exists in the image that is the target of the user's speech, and is obtained from the likelihood map based on the estimated likelihood included in the form modification factor information. When a predetermined statistic is less than a threshold value, the information processing apparatus according to Configuration 6, characterized in that the change based on the form modification factor information to the object corresponding to the subject object information is suppressed. (Configuration 9) It has a storage means for storing, as learning data, the form modification factor information and the form modification parameters generated based on the form modification factor information in association with each other, and the generation means includes the user information included in the form modification factor information received as an input, and the learning data stored by the storage means. The information processing apparatus according to Configuration 2, characterized in that a form modification parameter is generated based on the above. (Configuration 10) The output control means controls such that the result of the change made to the object by the change means is displayed in a predetermined display area. The display area includes a first partial area in which an image that is the subject of the user's speech is displayed, a second partial area in which an image of the object corresponding to the subject object information is displayed, and a third partial area in which an image corresponding to the result of the change made to the object by the change means is displayed. The information processing apparatus according to any one of Configurations 2 to 9, characterized in that the displayability of the image to each of the first partial area, the second partial area, and the third partial area can be selectively switched. (Configuration 11) The third partial area is configured to be able to receive an instruction related to a change to the object corresponding to the subject object information from the user. The generation means controls the form modification parameters applied to make a change to the object corresponding to the subject object information according to the instruction received by the third partial area. The information processing apparatus according to Configuration 10, characterized in that. (Configuration 12) The display area includes a fourth partial area configured to be able to receive a selection of at least any one of the acquired series of form modification factor information. The generation means controls the addability of a change to the object corresponding to the subject object information based on each of the series of form modification factor information according to the selection state of each of the series of form modification factor information received via the fourth partial area. The information processing apparatus according to Configuration 10 or 11, characterized in that. (Configuration 13) The fourth partial area is configured to be able to receive a designation of a process for changing the form feature to be applied to the selected form modification factor information. The generation means controls the form modification parameters applied to make a change to the object corresponding to the subject object information based on the form modification factor information selected via the fourth partial area and the process designated for the form modification factor information. The information processing apparatus according to Configuration 12, characterized in that. (Configuration 14) A holding means for sequentially holding, as a history, the result of a change made to the object by the changing means according to the acquisition result of the form modification factor information by the acquisition means; and a reception means for receiving an instruction regarding whether or not to apply each of a series of form modification factor information acquired by the acquisition means. The output control means controls such that, when the result of a change made to the object by the changing means based on one or more form modification factor information for which an application instruction has been received by the reception means is held as the history, the result indicated by the history is output to a predetermined output destination. The information processing apparatus according to any one of Claims 2 to 13, characterized in that. (Configuration 15) The reception means receives an instruction regarding whether or not to apply each of one or more processes for changing a form feature, which is specified for at least any one of the one or more form modification factor information for which an application instruction has been received. The output control means controls such that, when the result of a change made to the object by the changing means based on the one or more form modification factor information for which an application instruction has been received by the reception means and the process for which an application instruction has been received among the one or more processes is held as the history, the result indicated by the history is output to a predetermined output destination. The information processing apparatus according to Claim 14, characterized in that. (Configuration 16) When the result of a change made to the object by the changing means based on the one or more form modification factor information for which an application instruction has been received by the reception means and the form modification parameter for which an application instruction has been received among the series of form modification parameters is not held as the history, the changing means makes a change to the object corresponding to the subject object information based on the one or more form modification factor information and the form modification parameter for which the application instruction has been received. The output control means controls such that the result of the change made to the object by the changing means is output to a predetermined output destination. The information processing apparatus according to Claim 15, characterized in that. (Configuration 17) The speech data includes at least one of text data input by the user and text data generated based on the recognition result of the speech spoken by the user. The information processing apparatus according to any one of Configurations 1 to 16 is characterized by this. (Configuration 18) When there are a plurality of candidate objects corresponding to the subject object information, the changing means extracts at least a part of the plurality of candidate objects based on an instruction from the user as the object corresponding to the subject object information. The information processing apparatus according to Configuration 18 is characterized by this. (Configuration 19) Based on the speech data indicating the speech content of the user, an acquisition step of acquiring subject object information indicating an object that is the subject of the speech content and body modification factor information that is information for modifying the body of the object included in the speech content, and among one or more objects included in the image that is the target of the user's speech, a change step of making a change to the object corresponding to the subject object information acquired in the acquisition step based on the body modification factor information, and an output control step of controlling so that the result of the change made to the object in the change step is output to a predetermined output destination. The control method of the information processing apparatus is characterized by including these steps. (Method 1) A control method of an information processing apparatus, including: an acquisition step of acquiring subject object information indicating an object that is the subject of the speech content and body modification factor information that is information for modifying the body of the object included in the speech content based on speech data indicating the speech content of the user; a change step of making a change to the object corresponding to the subject object information acquired in the acquisition step among one or more objects included in the image that is the target of the user's speech based on the body modification factor information; and an output control step of controlling so that the result of the change made to the object in the change step is output to a predetermined output destination. The control method of the information processing apparatus is characterized by this. (Program 1) A program for causing a computer to function as an information processing apparatus, the program comprising: an acquisition means for acquiring, based on speech data indicating the speech content of a user, subject object information indicating an object that is the subject of the speech content and morphological modification factor information that is information for modifying the form of the object included in the speech content; a modification means for making a modification to an object corresponding to the subject object information acquired by the acquisition means among one or more objects included in an image that is the target of the user's speech, based on the morphological modification factor information; and an output control means for controlling so that a result of the modification made to the object by the modification means is output to a predetermined output destination.

Explanation of Signs

[0083] 100 Information processing apparatus 203 Text generation unit 206 Form update unit 207 Output control unit

Claims

1. Acquisition means for acquiring, based on speech data indicating the speech content of a user, subject object information indicating an object that is the subject of the speech content, and shape modification factor information that is information for modifying the shape of the object included in the speech content; Modification means for modifying, based on the shape modification factor information, an object corresponding to the subject object information among one or more objects included in an image that is the subject of the user's speech; Output control means for controlling so that the result of modifying the object by the modification means is output to a predetermined output destination; An information processing apparatus, characterized by comprising:

2. It has generation means for generating a shape modification parameter indicating the shape characteristics of the object based on the shape modification factor information, The modification means modifies the object corresponding to the subject object information based on the shape modification parameter The information processing apparatus according to claim 1, characterized in that.

3. The modification performed by the modification means is based on the shape modification parameter, Adding the shape characteristics to the object corresponding to the subject object information, Deforming at least a part of the object based on the shape characteristics, and Deleting the shape characteristics from the object including at least any one of The information processing apparatus according to claim 2, characterized in that.

4. The shape modification parameter includes at least any one of shape information including at least any one of position, size, shape, and orientation, and appearance information including at least any one of color and texture. The information processing apparatus according to claim 2, characterized in that it includes at least any one of them.

5. The shape modification factor information includes, as an attribute, information indicating which of a plurality of changes including addition, deletion, and deformation to the object corresponding to the subject object information is to be performed, The generation means controls the configuration of the shape modification parameters to be generated according to the attribute included in the shape modification factor information. The information processing apparatus according to claim 2, characterized in that.

6. The shape modification factor information includes at least any one of the name of the object that is the subject of the speech content, the estimated likelihood as the subjectness of the object, the feasibility of making changes to the object, and the information of the user who made the speech indicated by the speech content. The information processing apparatus according to claim 2, characterized in that.

7. The change means suppresses a change based on the shape modification factor information to the object corresponding to the subject object information when the estimated likelihood included in the shape modification factor information is less than a threshold value. The information processing apparatus according to claim 6, characterized in that.

8. The change means is a likelihood map indicating where in the image that is the target of the user's speech the object corresponding to the subject object information exists with what probability, and when a predetermined statistic obtained from the likelihood map based on the estimated likelihood included in the shape modification factor information is less than a threshold value, the change based on the shape modification factor information to the object corresponding to the subject object information is suppressed. The information processing apparatus according to claim 6, characterized in that.

9. It has a storage means for associating and storing shape modification factor information and shape modification parameters generated based on the shape modification factor information as learning data. The generation means generates shape modification parameters based on the user information included in the shape modification factor information received as an input and the learning data stored by the storage means. The information processing apparatus according to claim 2, characterized in that...

10. The output control means controls such that the result of the change made to the object by the change means is displayed in a predetermined display area. The display area is... composed of a first partial area in which an image that is the subject of the user's speech is displayed, a second partial area in which an image of an object corresponding to the subject object information is displayed, and a third partial area in which an image corresponding to the result of the change made to the object by the change means is displayed. The display of images in the first partial area, the second partial area, and the third partial area can be selectively switched. The information processing apparatus according to claim 2, characterized in that...

11. The third partial area is configured to be able to receive an instruction from the user regarding a change to the object corresponding to the subject object information. The generation means controls the form modification parameters applied to make a change to the object corresponding to the subject object information in accordance with the instruction received by the third partial area from the user. The information processing apparatus according to claim 10, characterized in that...

12. The display area includes a fourth partial area configured to be able to receive a selection of at least any one of the acquired series of form modification factor information. The generation means controls the addition of changes to the object corresponding to the subject object information based on each of the series of form modification factor information according to the selection state of each of the series of form modification factor information received via the fourth partial area. The information processing apparatus according to claim 10, characterized in that...

13. The fourth partial area is configured to be able to receive a designation of a process for changing a physical feature to be applied to the selected physical modification factor information. Based on the physical modification factor information selected via the fourth partial area and the process designated for the physical modification factor information, the generation means controls the physical modification parameters applied to change the object corresponding to the subject object information. The information processing apparatus according to claim 12, characterized in that.

14. A holding means for sequentially holding, as a history, the result of changing the object by the changing means according to the acquisition result of the physical modification factor information by the acquisition means; A receiving means for receiving an instruction regarding the applicability of each of a series of physical modification factor information acquired by the acquisition means; and having When the result of changing the object by the changing means based on one or more physical modification factor information for which an application instruction has been received by the receiving means is held as the history, the output control means controls so that the result indicated by the history is output to a predetermined output destination. The information processing apparatus according to claim 2, characterized in that.

15. The receiving means receives an instruction regarding the applicability of each of one or more processes for changing a physical feature, which are designated for at least any one of the one or more physical modification factor information for which an application instruction has been received. When the result of changing the object by the changing means based on the one or more physical modification factor information for which an application instruction has been received by the receiving means and the process for which an application instruction has been received among the one or more processes is held as the history, the output control means controls so that the result indicated by the history is output to a predetermined output destination. The information processing apparatus according to claim 14, characterized in that.

16. The changing means If the result of the change made to the object by the changing means based on the one or more shape modification factor information for which an application instruction has been received by the receiving means and the shape modification parameter for which an application instruction has been received among the series of shape modification parameters is not retained as the history, Based on the one or more shape modification factor information and the shape modification parameter for which the application instruction has been received, a change is made to the object corresponding to the subject object information, The output control means controls so that the result of the change made to the object by the changing means is output to a predetermined output destination. The information processing apparatus according to claim 15, characterized in that.

17. The speech data includes at least any one of text data input by the user and text data generated based on the recognition result of the voice spoken by the user. The information processing apparatus according to claim 1.

18. The changing means extracts an object corresponding to the subject object information from the image based on at least any one of the spatial features, image quality, and modality feature amounts of one or more objects included in the image that is the target of the user's speech, and makes a change based on the shape modification factor information to the extracted object. The information processing apparatus according to claim 1, characterized in that.

19. When there are a plurality of candidates for the object corresponding to the subject object information, the changing means extracts at least a part of the plurality of object candidates based on an instruction from the user as the object corresponding to the subject object information. The information processing apparatus according to claim 18, characterized in that.

20. A control method for an information processing apparatus, An acquisition step of acquiring, based on utterance data indicating the utterance content of a user, subject object information indicating an object that is the subject of the utterance content and morphological modification factor information that is information for modifying the form of the object included in the utterance content; A change step of making a change to an object corresponding to the subject object information acquired in the acquisition step among one or more objects included in an image that is the target of the user's utterance, based on the morphological modification factor information; An output control step of controlling so that a result of the change made to the object in the change step is output to a predetermined output destination; A control method for an information processing apparatus, characterized by including the above.

21. A computer, An acquisition means for acquiring, based on utterance data indicating the utterance content of a user, subject object information indicating an object that is the subject of the utterance content and morphological modification factor information that is information for modifying the form of the object included in the utterance content; A change means for making a change to an object corresponding to the subject object information acquired by the acquisition means among one or more objects included in an image that is the target of the user's utterance, based on the morphological modification factor information; An output control means for controlling so that a result of the change made to the object by the change means is output to a predetermined output destination; A program for causing an information processing apparatus to function, characterized by having the above.

Citation Information

Patent Citations

  • Text-based real image editing with diffusion models

    JP2024154427A