Data output method and data output device

The data output method and device efficiently identify candidate changes by targeting key areas of the target object contributing to a specified index, reducing calculation time and enhancing processing speed.

WO2025169478A1PCT designated stage Publication Date: 2025-08-14NISSAN MOTOR CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/004596
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing data output methods take a long time to identify candidate changes needed to bring an image closer to a specified index due to extensive calculation processes.

Method used

A data output method and device that acquires data on a target object, specifies an index, extracts key areas contributing to that index, identifies candidate changes, and outputs data for these changes using a trained model to reduce calculation time.

Benefits of technology

Reduces the time required for calculating candidate modifications by focusing on specific areas of the target object that contribute most to the specified index, thereby speeding up the data output process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024004596_14082025_PF_FP_ABST
    Figure JP2024004596_14082025_PF_FP_ABST
Patent Text Reader

Abstract

A controller (12): acquires data of a target object; acquires a prescribed index for the target object; extracts a target region of the target object that has a greater contribution to the prescribed index than other regions; identifies at least one candidate for a change to be made to change a feature of the target region to a different feature related to the feature of the target region; and outputs data of the identified at least one candidate for the change, or data obtained by changing the change target on the basis of the at least one candidate for the change.
Need to check novelty before this filing date? Find Prior Art

Description

Data output method and data output device

[0001] The present invention relates to a data output method and a data output device.

[0002] A technology is known in which a target latent code is obtained based on a latent code obtained by encoding an image in the space of a generative adversarial network and a latent code obtained by mapping a text code based on target language image pre-training in the space, and a target image is generated based on the target latent code (Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2022-180519

[0004] However, the technology described in Patent Document 1 has the problem that when there are many candidate changes to make to an image to bring the image closer to the indicator indicated in the text description, the calculation process to identify candidate changes to make to the image takes a long time.

[0005] The problem that the present invention aims to solve is to provide a data output method and a data output device that can shorten the calculation processing time for identifying candidate changes to be made to data when outputting data that has been changed to bring it closer to a specified index.

[0006] The present invention solves the above problem by acquiring data on a target object, acquiring a specified index for the target object, extracting a target area in the target object that contributes more to the specified index than other areas, identifying at least one candidate change for changing the characteristics of the target area to another characteristic related to the characteristics of the target area, and outputting data for the identified candidate change or data for which changes have been made to the target to be changed based on the candidate change.

[0007] According to the present invention, when data is output that has been modified so as to approach a predetermined index, the time required for calculation processing to identify candidates for modifications to be made to the data can be reduced.

[0008] Fig. 1 is a block diagram showing an example of a data output system including a data output device according to an embodiment of the present invention. Fig. 2 is a diagram showing an example of a method for generating candidate data for modification executed by the data output device according to the embodiment. Fig. 3 is a flowchart showing an example of a control procedure for executing the data output method according to the embodiment.

[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS A data output method and a data output device according to an embodiment of the present invention will be described below with reference to the accompanying drawings.

[0010] FIG. 1 is a block diagram illustrating an example of a data output system including a data output device according to the present embodiment. As illustrated in FIG. 1 , the data output system 100 includes a sensor 2, an input device 3, a gaze estimation device 4, a database 5, and a data output device 10. The sensor 2, the input device 3, the gaze estimation device 4, the database 5, and the data output device 10 are mounted on a vehicle 1 and exchange information with each other via an in-vehicle network such as a CAN. In the data output system 100, the data output device 10 outputs data to a user. The user is a user of the data output device, such as the owner of the vehicle 1. The user is not limited to the owner of the vehicle 1, but may also be the designer of the vehicle 1, for example. Note that, in the present embodiment, an example is described in which the user who is the owner of the vehicle 1 uses the data output device 10 mounted on the vehicle 1, but the data output device 10 is not limited to a device mounted on the vehicle 1. For example, the data output device 10 may be located in a room, such as a workshop, where the user who is the designer of the vehicle 1 performs design work.

[0011] The output data is data of candidate changes to be made to the change target, or data resulting from changes made to the change target based on the candidate changes. The data of the change target is data of the target to which the user wishes to make changes, such as image or audio data. The image is a 2D image or 3D data. The image is, for example, an image including a vehicle. In this embodiment, the image is described as including a vehicle (vehicle 1) owned by the user, but the image may be other images, such as an image including clothing or an image including a face. The audio includes, for example, music. The changes made to the change target are, for example, vehicle parts. If the image is an image including clothing or an image including a face, the changes made to the change target are clothing parts, makeup, etc. If the change target is music, the changes made to the change target are changes made to certain parts of the music. In the following embodiment, the change target is described as an image including a vehicle, but is not limited to this and may be any target to which the user wishes to make changes, such as other images or audio.

[0012] In this embodiment, when a user considers customizing a vehicle, the data output system 100 outputs an image including candidate modifications to the vehicle to be modified, and proposes customization of the vehicle 1. This allows the user to check the image of the vehicle's appearance after modifications are made before actually customizing the vehicle. The candidate modifications include parts, paint, interior parts, or accessories to be added or changed to the vehicle's exterior.

[0013] The user identifies a target vehicle. The target vehicle is an object that serves as a target for the changes the user makes to the change object, and has the characteristics the user wants to add to the change object. The user identifies a reference target object to indicate how to change the change object. The user also specifies a predetermined index for the target vehicle. The predetermined index is an index for evaluating the object's appearance characteristics, and indicates, for example, the impression the object's appearance gives to the user. For example, the predetermined index indicates an impression of the object's appearance, such as "red" or "cute." The user can specify how the impression of the change object should be changed by specifying the target vehicle and a predetermined index indicating the impression the vehicle gives. For example, if the user is driving vehicle 1 and feels that another vehicle traveling ahead is "cute," the user specifies the other vehicle as the target vehicle and specifies the text "cute" as the predetermined index. In this embodiment, the data output device 10 identifies candidate changes to be made to the change object so as to bring the impression of the target object closer to "cute." As a result, the data output device 10 identifies candidate changes to make the target vehicle into a "cute" vehicle like the target vehicle. The data output device 10 outputs the identified candidate changes or data in which changes have been made to the target vehicle based on the candidate changes.

[0014] In the case of conventional techniques, when an image is generated in which the characteristics of a vehicle to be modified are modified to approximate a predetermined index, if there are many candidate modifications to be made to the vehicle, the calculation process for identifying the candidate modifications may take a long time. In contrast, in the present embodiment, the data output device 10 can reduce the calculation process time for identifying candidate modifications by limiting which areas of the target vehicle are to be modified. The method for identifying candidate modifications will be described in detail below.

[0015] The sensor 2 detects the surrounding environment including surrounding objects located around the user. The sensor 2 may be, for example, a camera installed in the vehicle 1, and may be a camera equipped with an imaging element such as a CCD or CMOS. The sensor 2 acquires surrounding images capturing the surrounding environment including the surrounding objects. For example, the surrounding objects include automobiles (other vehicles) other than the subject vehicle, motorcycles, and bicycles. In this embodiment, the sensor 2 acquires images capturing a vehicle that is a change target, a target target, or a learning target. Note that the sensor 2 is not limited to a device installed in the vehicle 1, but may also be a device installed in a mobile terminal or the like owned by the user.

[0016] When gaze information including the user's gaze direction is input from the gaze estimation device 4, the sensor 2 points the camera in the user's gaze direction and captures an image of the surrounding environment including surrounding objects in the user's gaze direction to acquire a surrounding image. Note that, although the sensor 2 acquires an image in this embodiment, it is not limited to this and may also acquire a video. The surrounding image detected by the sensor 2 is output to the data output device 10.

[0017] The input device 3 is an input interface that accepts input from the user. The input device 3 is composed of one or more devices. For example, the input device 3 is composed of a microphone that allows the user to input voice data. For example, the input device 3 inputs voice data including a predetermined indicator specified by the user using the microphone. The input device 3 may also be equipped with a touch panel or a keyboard. The user can input text data into an input form. The input text data includes the predetermined indicator. Note that the input device 3 is not limited to a device installed in the vehicle 1, and may also be a device installed in a mobile terminal or the like owned by the user.

[0018] The gaze estimation device 4 performs a user gaze estimation process. The gaze estimation device 4 estimates the gaze direction in which the user is directing their gaze toward the surrounding environment, and outputs gaze information including the gaze direction to the sensor 2. The gaze estimation device 4 has a gaze measurement camera that captures an image of the user's pupils, and estimates the user's gaze direction based on the captured image of the user's pupils.

[0019] The database 5 is a database that stores various information. The database 5 stores trained models. The stored trained models are trained to output target numerical values ​​based on input data including target object data and predetermined indicators. The target numerical values ​​will be described later. The trained models are, for example, CLIP (Contrastive Language-Image Pre-Training) models. The database 5 may also store other trained models, such as trained models trained for image generation. Examples of trained models for image generation include Stable Diffusion and StyleGAN. In this embodiment, the controller 12 generates a trained model using the learning unit 102 and stores it in the database 5. Note that this embodiment is not limited to this, and the controller 12 may acquire a trained model from an external server via the communication device 13.

[0020] The database 5 also stores a training dataset. The training dataset associates data of a training object with a predetermined index for the training object. For example, the training object and the predetermined index are image and text data. The dataset includes a plurality of images and text corresponding to each image. Each image-text pair is associated with identification information (e.g., an identification ID) for identifying the pair. In this embodiment, the image is described as an image of a training vehicle, and the text data includes a predetermined index corresponding to the image. The predetermined index indicates a user's evaluation of the training vehicle.

[0021] The database 5 may be a database that stores data of candidate modifications, such as images of vehicle parts and other candidate modifications.

[0022] The data output device 10 is a device that outputs data to a user. The data is data related to candidate changes to be made to a change target. The output data includes candidate changes or data resulting from changes to the change target based on the candidate changes. In this embodiment, the output data is, for example, an image including candidate changes to be made to a vehicle, or an image of a vehicle that has been changed based on the candidate changes. The data output device 10 includes an output device 11, a controller 12, and a communication device 13. The controller 12 acquires information from the sensor 2, the input device 3, the database 5, and the communication device 13, and identifies candidate changes based on the acquired information. The data output device 10 outputs data based on the identified candidate changes to the user via the output device 11, thereby suggesting vehicle customization to the user.

[0023] The output device 11 is an output interface that outputs data to the user. The data output to the user includes data on candidate changes, or data on the target to be changed based on the candidate changes. For example, an image including candidate changes to be made to the vehicle, or an image of the vehicle changed based on the candidate changes, is output. The data generated by the data output device 10 is displayed on a screen and / or output as audio. The output device 11 is composed of one or more devices. For example, the output device 11 is composed of a display that displays on a screen and a speaker that outputs audio data. Examples of displays include a liquid crystal panel and an organic EL panel. Note that if a touch panel display is used as the output device 11, it can also be used as the input device 3.

[0024] The communication device 13 is a device that exchanges information with the outside of the vehicle 1 via a network. The communication device 13 receives a trained model from, for example, an external server.

[0025] The controller 12 includes a computer having hardware and software. The computer includes a ROM storing a program, a CPU that executes the program stored in the ROM, and a RAM that functions as an accessible storage device. Note that an MPU, DSP, ASIC, FPGA, etc. can be used as the operating circuit instead of or in addition to the CPU. As shown in FIG. 1 , the controller 12 includes functional blocks, such as an acquisition unit 101, a learning unit 102, a calculation unit 103, a generation unit 104, and an output unit 105. Software for executing each process cooperates with the hardware to perform each function. Note that in this embodiment, the functions of the controller 12 are divided into five blocks, and the functions of each functional block are described. However, the functions of the controller 12 do not necessarily have to be divided into five blocks. Furthermore, in this embodiment, one controller has multiple functional units. However, this is not limited to this, and multiple controllers may each have each functional unit. Furthermore, each functional unit does not have to be provided in an on-board controller, but may also be provided in an external cloud.

[0026] The acquisition unit 101 performs an acquisition process to acquire various data. In the acquisition process, the acquisition unit 101 acquires data from the sensor 2 and the input device 3. In this embodiment, the acquisition process is executed in the following three situations. That is, in the first situation, the data output device 10 acquires learning data for performing machine learning of a model. In the second situation, the data output device 10 acquires target data. In the third situation, the data output device 10 acquires data to be changed.

[0027] The data acquired from the sensor 2 includes surrounding information including the user's surrounding environment. The acquisition unit 101 acquires, from the sensor 2, the surrounding information detected by the sensor 2. The surrounding information is, for example, a surrounding image including surrounding objects. In this embodiment, the acquisition unit 101 acquires, from the sensor 2, an image of a vehicle that is a learning object, a target object, or a change object.

[0028] The data acquired from the input device 3 is input data including a predetermined index for the object. The input data including the predetermined index is user voice data acquired by a microphone or text data input by the user. In this embodiment, the text data including the predetermined index is text data including a predetermined index for a vehicle, which is a learning object or a target object. The predetermined index is identified based on voice data including a user's statement such as "cute" about the vehicle, or text data in which "cute" is input.

[0029] Here, an example of the acquisition process of the acquisition unit 101 will be described. The first example is an example of a situation in which the data output device 10 acquires learning data for performing machine learning on a model. First, an acquisition process in which learning data is acquired based on user voice data will be described. When the microphone of the input device 3 picks up the user's voice, the acquisition unit 101 acquires the user's voice data from the input device 3. The voice data includes a predetermined indicator for a vehicle that is a learning target. The acquisition unit 101 then acquires gaze information including the user's gaze direction at the time the voice data was acquired from the gaze estimation device 4. Based on the gaze information, the acquisition unit 101 causes the sensor 2 to capture an image of the user's gaze direction and acquires, from the sensor 2, a surrounding image captured in the user's gaze direction. The surrounding image includes the vehicle that is a learning target. The acquisition unit 101 converts the user's voice data into text data. The acquisition unit 101 associates the image of the vehicle that is a learning target acquired from the sensor 2 with text data including the predetermined indicator, and stores a dataset in which the image and the text data are paired in the database 5. For example, when a user is viewing an exhibition, an image of a vehicle (the vehicle being used for learning) in the direction of the user's line of sight and audio data containing the user's thoughts about the vehicle (such as "cute") are acquired, and the acquired data are stored as a pair in database 5.

[0030] Next, an acquisition process for acquiring learning data based on text data entered by a user will be described. The acquisition unit 101 acquires gaze information including the user's gaze direction from the gaze estimation device 4. Based on the gaze information, the acquisition unit 101 causes the sensor 2 to capture an image of the user's gaze direction and acquires, from the sensor 2, a surrounding image captured in the user's gaze direction. The surrounding image includes a vehicle that is a learning target. The acquisition unit 101 acquires text data entered by the user into an input form from the input device 3. The text data includes a predetermined indicator for the vehicle that is a learning target. The acquisition unit 101 associates an image of the vehicle that is a learning target acquired from the sensor 2 with text data including the predetermined indicator, and stores a data set in which the image and the text data are paired in the database 5. For example, when a user is entering text into an input form such as a questionnaire survey, the acquisition unit 101 acquires an image of a vehicle (the learning target vehicle) in the direction the user is looking and text data of the user's input for the vehicle, and stores the acquired data paired in the database 5.

[0031] As an example of the acquisition process of the acquisition unit 101, an example of a situation in which the data output device 10 acquires data of a target object will be described. First, an acquisition process in which data of a target object is acquired based on user voice data will be described. When the microphone of the input device 3 picks up the user's voice, the acquisition unit 101 acquires the user's voice data from the input device 3. Then, the acquisition unit 101 determines whether the user's voice data includes a predetermined indicator for the target object, such as a keyword used by the user to evaluate the target object. The keywords for evaluating the target object are stored in a keyword list in advance. The acquisition unit 101 determines that the user's voice data includes a predetermined indicator for the target object by recognizing a keyword in the keyword list from the user's voice data. If the user's voice data includes a predetermined indicator for the target object, the acquisition unit 101 acquires gaze information including the user's gaze direction from the gaze estimation device 4. Based on the gaze information, the acquisition unit 101 causes the sensor 2 to capture an image of the user's gaze direction and acquires a surrounding image capturing the user's gaze direction from the sensor 2. The surrounding image includes the target object. The acquisition unit 101 converts the user's voice data into text data. The acquisition unit 101 associates an image of a vehicle as a target acquired from the sensor 2 with text data including predetermined indicators, and stores a data set in which the image and the text data are paired in the database 5.

[0032] Next, an acquisition process for acquiring data on a target object based on text data input by a user will be described. The acquisition unit 101 acquires gaze information including the user's gaze direction from the gaze estimation device 4, causes the sensor 2 to capture an image of the user's gaze direction based on the gaze information, and acquires a surrounding image capturing the user's gaze direction from the sensor 2. The surrounding image includes the target object. The acquisition unit 101 acquires text data entered by the user into the input form. The acquisition unit 101 associates the image of the vehicle, which is the target object, acquired from the sensor 2 with text data including predetermined indicators, and stores a dataset in which the image and the text data are paired in the database 5.

[0033] Next, an acquisition process for acquiring data to be changed will be described. The acquisition unit 101 estimates the imaging conditions under which the target object was imaged. The imaging conditions include, for example, viewpoint and distance. The acquisition unit 101 causes the sensor 2 to capture an image of the vehicle to be changed under the same imaging conditions as the imaging conditions, and acquires an image of the vehicle to be changed from the sensor 2. Note that the data to be changed may be acquired from the input device 3 when the user inputs the data to the input device 3.

[0034] The learning unit 102 learns a model to generate a trained model. In this embodiment, the learning unit 102 generates a trained model that is trained to output output data including a value obtained by quantifying the predetermined index for the target based on input data including target data and a predetermined index for the target. In this embodiment, using the trained model, data including target object data and the predetermined index for the target object is used as input data, and output data including a target numerical value obtained by quantifying the predetermined index for the target object is output. The value obtained by quantifying the predetermined index for the target is, for example, a value indicating the similarity between the target and the predetermined index. The similarity is the degree to which the target and the predetermined index are semantically related. The trained model is, for example, a CLIP model. The CLIP model calculates the degree to which an image is semantically related to text data including the predetermined index as the similarity. For example, the trained model is trained based on a dataset in which data of the training target is associated with text data including the predetermined index for the training target. In this embodiment, the data of the training target is an image of a vehicle that is the training target acquired by the acquisition unit 101.

[0035] Here, an example of a learning method by the learning unit 102 will be described. First, the learning unit 102 acquires a dataset of image and text pairs from the database 5 through batch processing. For example, the learning unit 102 acquires each pair of image and text (image 1 and text 1, image 2 and text 2, ..., image N and text N). The learning unit 102 organizes each image (image 1, image 2, ..., image N) as an image batch and each text (text 1, text 2, ..., text N) as a text batch. The learning unit 102 inputs the image batch and the text batch into the learning model, and outputs an image vector batch (image vector 1, image vector 2, ..., image vector N) and a text vector batch (text vector 1, text vector 2, ..., text vector N). The learning unit 102 calculates an inner product matrix based on each image vector and each text vector, calculates cross-entropy for each row to obtain an average value 1, and calculates cross-entropy for each column to obtain an average value 2. The learning unit 102 obtains a loss value by calculating the average value of the obtained average values ​​1 and 2. The learning unit 102 optimizes the learning model by updating the parameters of the learning model so as to minimize the loss value. The learning unit 102 stores the optimized learning model in the database 5 as a trained model.

[0036] The calculation unit 103 calculates a target numerical value by quantifying a predetermined index for a target object based on data of the target object and a predetermined index. The predetermined index is identified by text data. The target numerical value is a value indicating the similarity between the target object and the predetermined index. The similarity between the target object and the predetermined index is the degree to which the target object and the predetermined index are semantically related. The closer the target object is to the predetermined index, the larger the calculated target numerical value. For example, the calculation unit 103 calculates the target numerical value using a trained model. The trained model is a model trained to output output data including the target numerical value based on input data including data of the target object and the predetermined index, and is, for example, a CLIP model. The calculation unit 103 inputs the data of the target object and the predetermined index into the trained model and causes the trained model to output the target numerical value, thereby calculating the target numerical value. When the target object data is an image including a target vehicle, the calculation unit 103 inputs input data including an image including the target vehicle and text data including a predetermined indicator for the vehicle into the trained model, and calculates a target numerical value indicating the similarity between the input image and the text data based on the trained model. Specifically, the calculation unit 103 calculates the target numerical value by using the trained model to calculate an inner product based on the vector of the image of the target object and the vector of the text data including the predetermined indicator.

[0037] The calculation unit 103 also calculates a heat map indicating the proportion of each region in the target object that contributes to a predetermined index. The heat map emphasizes regions in the target object that contribute to the predetermined index more than other regions. The emphasized regions are displayed, for example, in a highlighted manner. The calculation of the heat map is performed using a trained model. The calculation unit 103 inputs text data including an image of the target object and the predetermined index into the trained model, and calculates, based on the feature map of the trained model, an image that emphasizes regions in the image of the target object that the trained model focuses on as characteristic parts related to the predetermined index, as a heat map.

[0038] In this embodiment, the calculation unit 103 uses an image of a target vehicle as an input image and calculates a heat map in which areas of the target vehicle that contribute more to a predetermined index than other areas are highlighted. For example, if the predetermined index indicates an impression of the target vehicle being "cute," areas of the target vehicle that contribute to improving the impression of "cute" are highlighted in the heat map.

[0039] Here, an example of heat map calculation will be described. The calculation unit 103 calculates a heat map using a CLIP model as the trained model. The heat map is calculated from a query and a key of an image encoder (e.g., Vision Transformer) of the CLIP model. The calculation unit 103 applies Softmax to the inner product of the query and the key to output an Attention weight (output value A). The calculation unit 103 calculates a gradient value for the output value A and calculates the product of each element of a matrix based on the gradient value and the output value A. The calculation unit 103 excludes negative values ​​of the product and calculates the average of the calculated vector on the first axis as the heat map. Note that these calculation methods are merely examples, and other calculation methods may be used. For example, the average of the output value A on the first axis may be directly calculated as the heat map.

[0040] Here, an example of a heat map calculated by the data output device according to this embodiment will be described with reference to FIG. 2 . FIG. 2 is a diagram illustrating an example of a procedure for generating candidate data for modification executed by the data output device according to this embodiment. For example, if the predetermined index is “cute,” the calculation unit 103 calculates a higher target numerical value the closer the image of the target vehicle to “cute.” Furthermore, the calculation unit 103 calculates a heat map indicating the proportion of each region of the target vehicle that contributes to “cute.” For example, as shown in FIG. 2 , if a user utters, “This car is cute,” the calculation unit 103 calculates a heat map (C) based on text data (A) of “This car is cute” and an image (B) including the target vehicle TV. Assume that the target part TP of the target vehicle TV is a part that gives the user the impression of “cute.” In this case, the heat map highlights regions of the image of the target vehicle that contribute more to “cute” than other regions. In the example of heat map (C), an area A including the target part TP in the vehicle TV, which is the target object, is highlighted as an area that contributes greatly to the "cute" feeling.

[0041] The generation unit 104 performs a generation process to generate data of candidate changes and / or data in which changes have been made to the change target based on the candidate changes. In the generation process, the generation unit 104 first extracts a target region in the target target that contributes more to a predetermined index than other regions. For example, the generation unit 104 extracts the target region using a heat map calculated by the calculation unit 103. The target region is extracted as a region highlighted in the heat map. In the example of FIG. 2 , the generation unit 104 extracts region A as the target region using heat map (C). By extracting the target region, the region searched for candidate changes can be limited, thereby reducing the time required for the calculation process to identify candidate changes. Note that in this embodiment, the extraction of the target region is not limited to the heat map, and other methods may be used as long as the region in the target target that contributes more to a predetermined index can be identified.

[0042] Next, the generation unit 104 identifies at least one change candidate for changing the feature of the target region to another feature related to the feature of the target region. Specifically, the generation unit 104 extracts an image included in the target region and acquires the feature of the extracted image. The acquisition of the feature of the image included in the target region is performed, for example, using SIFT (Scale Invariant Feature Transform) features or intermediate output of a trained convolutional neural network. In the example of FIG. 2 , the generation unit 104 acquires the feature of the image included in the target region, with region A as the target region. For example, if region A includes a target part TP, the feature of the target part TP is extracted. The generation unit 104 acquires a change candidate having another feature related to the feature of the image of the target region from a database that stores change candidates. The another feature is, for example, a feature similar to the feature of the image of the target region, for example, a feature that gives the user the same impression as the feature of the image of the target region. For example, the generation unit 104 compares the features of the image of the target region with the features of images containing candidate changes stored in the database, and identifies candidate changes having features whose feature amounts are within a certain distance. In the example of Figure 2, the candidate changes are images (E) containing accessories or the like similar to the target part TP.

[0043] Furthermore, the generation unit 104 is not limited to identifying data of change candidates, and may generate data in which changes have been made to the change target based on the change candidates. The generation unit 104 extracts a corresponding area corresponding to a target area of ​​the target target in the image of the change target acquired by the acquisition unit 101. The generation unit 104 generates an image in which changes have been made to the corresponding area in the change target based on the change candidates. The image is generated using Stable Diffusion, StyleGAN, or the like. In the example of FIG. 2 , the generation unit 104 acquires an image (D) including the vehicle CV that is the change target, and generates an image (F) in which a change candidate CP has been added to the corresponding area A' in the image (D) of the vehicle CV. The change candidate CP is, for example, another part having features similar to the features of the target part TP.

[0044] Furthermore, the generation unit 104 may acquire change candidates based on predetermined additional conditions. The predetermined additional conditions include, for example, conditions such as color, budget, etc. The generation unit 104 acquires change candidates that satisfy the predetermined additional conditions from among the change candidates.

[0045] After generating data based on the change candidates, the generation unit 104 compares the candidate numerical values ​​with the target numerical values ​​to identify the change candidates to be output to the user. The candidate numerical values ​​are values ​​that quantify predetermined indicators for the change candidates or the data to be changed after the changes have been made. For example, the generation unit 104 calculates the candidate numerical values ​​by using an image including the change candidates and predetermined indicators as input data and outputting output data including the candidate numerical values ​​from the trained model. The generation unit 104 identifies the change candidates based on the calculated candidate numerical values ​​and target numerical values. For example, the generation unit 104 identifies the change candidates whose candidate numerical values ​​are close to the target numerical value. A candidate numerical value close to the target numerical value is a candidate numerical value that is within a predetermined range from the target numerical value.

[0046] Furthermore, the generation unit 104 may calculate a candidate numerical value by quantifying a predetermined index for the target image to be changed based on the image to be changed and a predetermined index. Specifically, the generation unit 104 calculates the candidate numerical value by using the image to be changed based on the candidate change and the predetermined index as input data and outputting output data including the candidate numerical value from the trained model. The generation unit 104 identifies the target image to be changed based on the target numerical value and the candidate numerical value. For example, the generation unit 104 identifies an image whose candidate numerical value is close to the target numerical value.

[0047] Furthermore, in this embodiment, when there are multiple target regions in the image of the target object and the multiple target regions are in a line-symmetric relationship, the generation unit 104 may change corresponding regions in the image to be changed that correspond to the multiple target regions that are in a line-symmetric relationship, based on the change candidate. For example, in the case of an image including a vehicle, the vehicle is symmetrical with respect to the center line. If the installation areas of headlights and taillights are the target regions, the respective lights are in a line-symmetric relationship, and therefore the left and right light portions are changed as corresponding regions in the image including the vehicle to be changed, based on the change candidate.

[0048] The output unit 105 outputs the data generated by the generation unit 104. For example, the output unit 105 transmits a control instruction to the output device 11, causing the output device 11 to output data. Specifically, the output unit 105 causes the output device 11 to output an image including a modified vehicle that has been modified based on the modification candidates. The output unit 105 may also output store data. The store data is data about stores that can modify an actual vehicle to be similar to the modified vehicle included in the image. Stores include general dealerships, customization specialty stores, etc. Furthermore, if the vehicle can be modified by DIY rather than at a store, the output unit 105 may output a notification indicating that DIY is possible.

[0049] Next, the control procedure of the data output method executed by the data output device 10 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the control procedure for executing the data output method according to this embodiment. In this embodiment, when target data is input, the controller 12 starts the flow from step S101.

[0050] In step S101, the controller 12 acquires data of a target object. In step S102, the controller 12 acquires a predetermined index for the target object. In step S103, the controller 12 extracts a target region in the target object that has a higher contribution to the predetermined index than other regions. In step S104, the controller 12 identifies at least one change candidate for changing the characteristics of the target region to another characteristic related to the characteristics of the target region. In step S105, the controller 12 outputs data of the identified change candidate, or data in which a change has been made to the change target based on the change candidate.

[0051] As described above, in the data output method and data output device according to this embodiment, the controller acquires data of a target object, acquires a predetermined index for the target object, extracts a target area of ​​the target object that contributes more to the predetermined index than other areas, identifies at least one candidate change for changing the characteristics of the target area to another characteristic related to the characteristics of the target area, and outputs data of the identified candidate change or data in which changes have been made to the change object based on the candidate change. This reduces the computational processing time required to identify candidate changes to make to the data when outputting data in which changes have been made to bring it closer to the predetermined index.

[0052] Furthermore, in the data output method and data output device according to this embodiment, the controller calculates a target numerical value by quantifying a predetermined indicator for the target object, calculates candidate numerical values ​​by quantifying predetermined indicators for a plurality of candidate changes, and identifies candidate changes based on the target numerical value and candidate numerical values. This makes it possible to identify candidate changes based on the predetermined indicators for the target object.

[0053] In the data output method and device according to the present embodiment, the controller identifies candidate changes whose candidate values ​​are close to the target values, thereby identifying candidate changes that are close to a predetermined indicator for the target object.

[0054] Furthermore, in the data output method and data output device according to this embodiment, the controller inputs target object data and a predetermined index into a trained model and causes the trained model to output the target numerical value, thereby calculating the target numerical value, and the trained model is a model trained to output output data including the target numerical value based on input data including the target object data and the predetermined index. This makes it possible to obtain the predetermined index for the target object as a numerical value.

[0055] Furthermore, in the data output method and data output device according to this embodiment, the controller extracts target regions using a heat map that indicates the proportion of each region of the target object that contributes to a predetermined index, thereby making it possible to extract regions of the target object that have a greater impact on the predetermined index.

[0056] Furthermore, in the data output method and data output device according to this embodiment, when there are multiple target regions in the image of the target object and the multiple target regions are in a line-symmetric relationship, the controller changes the regions in the image to be changed that correspond to the multiple target regions in a line-symmetric relationship based on the change candidates, thereby making it possible to change the line-symmetric regions in the image to be changed based on the change candidates.

[0057] Furthermore, in the data output method and data output device according to this embodiment, the predetermined index is identified based on text data, which allows the predetermined index to be specified by text.

[0058] Furthermore, in the data output method and data output device according to this embodiment, the text data is input by the user, which allows the user to specify the predetermined index as desired.

[0059] In the data output method and data output device according to the present embodiment, the image to be modified is an image including a vehicle, and the modification candidates include parts, paint, interior parts, or accessories to be added or modified to the exterior of the vehicle, thereby making it possible to suggest modification candidates for customizing the vehicle to the user.

[0060] In the data output method and data output device according to this embodiment, the image to be modified is an image including a vehicle, and the controller outputs an image including the modified vehicle modified based on the modification candidates, and also outputs data of shops that can modify an actual vehicle to be similar to the modified vehicle included in the image, thereby allowing the user to know shops that can customize vehicles.

[0061] It should be noted that the above-described embodiments have been described to facilitate understanding of the present invention, and are not intended to limit the present invention. Therefore, each element disclosed in the above-described embodiments is intended to include all design modifications and equivalents that fall within the technical scope of the present invention.

[0062] REFERENCE SIGNS LIST 100: Data output system 2: Sensor 3: Input device 4: Gaze estimation device 5: Database 10: Data output device 11: Output device 12: Controller 101: Acquisition unit 102: Learning unit 103: Calculation unit 104: Generation unit 105: Output unit 13: Communication device

Claims

1. A data output method executed by a controller, wherein the controller: acquires data of a target object; acquires a predetermined index for the target object; extracts a target area of the target object that contributes more to the predetermined index than other areas; identifies at least one candidate change for changing the characteristics of the target area to another characteristic related to the characteristics of the target area; and outputs data of the identified candidate change, or data in which the change target has been changed based on the candidate change.

2. A data output method as described in claim 1, wherein the controller calculates a target numerical value by quantifying the specified indicator for the target object, calculates candidate numerical values by quantifying the specified indicator for each of a plurality of candidate changes, and identifies the candidate changes based on the target numerical value and the candidate numerical values.

3. A data output method according to claim 2, wherein the controller identifies the candidate change such that the candidate value is close to the target value.

4. A data output method as claimed in claim 2 or 3, wherein the controller inputs the target object data and the specified indicator into a trained model and calculates the target numerical value by outputting the target numerical value from the trained model, and the trained model is a model trained to output output data including the target numerical value based on input data including the target object data and the specified indicator.

5. A data output method according to any one of claims 1 to 4, wherein the controller extracts the target area using a heat map that indicates the proportion of each area in the target object that contributes to the specified index.

6. A data output method according to any one of claims 1 to 5, wherein the controller, when there are multiple target regions in the image of the target object and the multiple target regions are in a line-symmetric relationship, modifies regions in the image to be modified that correspond to the multiple target regions that are in a line-symmetric relationship based on the modification candidate.

7. A data output method according to any one of claims 1 to 6, wherein the predetermined indicator is identified based on text data.

8. A data output method according to claim 7, wherein the text data is input by a user.

9. A data output method according to any one of claims 1 to 8, wherein the image to be changed is an image including a vehicle, and the candidate changes include parts, paint, interior parts, or accessories to be added to or changed in the exterior of the vehicle.

10. A data output method according to any one of claims 1 to 9, wherein the image to be modified is an image including a vehicle, and the controller outputs an image including the modified vehicle after modifications have been made based on the candidate modifications, and outputs data on a shop that can modify the actual vehicle to the same extent as the modified vehicle.

11. A data output device having a controller, wherein the controller: acquires data of a target object; acquires a predetermined index for the target object; extracts a target area of the target object that contributes more to the predetermined index than other areas; identifies at least one candidate change for changing the characteristics of the target area to another characteristic related to the characteristics of the target area; and outputs data of the identified candidate change, or data in which the change target has been changed based on the candidate change.

Citation Information

Patent Citations

  • Image processing server

    JP2004326280A

  • Evaluation device, evaluation method, and evaluation program

    JP2018195078A

  • Data extension system, data extension method and program

    JP2021120914A