Learning device and learning method

The learning device enhances model robustness by generating three-dimensional model-based images and incorporating user feedback to handle varying imaging angles, improving vehicle evaluation accuracy.

JP7679746B2Active Publication Date: 2025-05-20NISSAN MOTOR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021154492
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-05-20
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

Existing learning models trained on images taken at different angles risk reduced robustness due to variations in imaging angles.

Method used

A learning device that identifies vehicle types in images, generates corresponding three-dimensional model-based images, and learns user evaluation comments to enhance model robustness by increasing the number of images at varied angles using virtual imaging.

Benefits of technology

Generates a highly robust learning model capable of accurately analyzing vehicle quality from varied imaging angles by integrating three-dimensional model data and user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007679746000001
    Figure 0007679746000001
  • Figure 0007679746000002
    Figure 0007679746000002
  • Figure 0007679746000003
    Figure 0007679746000003
Patent Text Reader

Abstract

To provide a learning device and a learning method capable of generating a highly robust model.SOLUTION: A learning device 1 includes a storage device 30 that stores a first image of a vehicle in association with a user evaluation comment on the first image, and a controller 20. The controller 20 identifies a type of the vehicle that appears in the first image, generates a second image different from the first image using three-dimensional model data corresponding to the identified type of the vehicle, learns a model related to the user evaluation comment using the first image, the second image, and the user evaluation comments as input data and stores the learned model in the storage device 30.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a learning device and a learning method. [Background technology]

[0002] Conventionally, there has been known an invention that uses evaluation results of the presentation position of food shown in an image (of plated food as a subject), inputs an image of plated food as a subject, and uses the evaluation results of the presentation position of food shown in the image as training data to obtain a learning model (Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2020-181436 A Summary of the Invention [Problem to be solved by the invention]

[0004] However, when there is a difference in the number of images taken at different imaging angles, there is a risk that the robustness of the model will be reduced if the model is trained as is.

[0005] The present invention has been made in consideration of the above problems, and has an object to provide a learning device and a learning method capable of generating a highly robust model. [Means for solving the problem]

[0006] A learning device according to one embodiment of the present invention identifies the type of vehicle depicted in a first image, generates a second image different from the first image using three-dimensional model data corresponding to the identified type of vehicle, learns a model related to user evaluation comments using the first image, the second image, and the user evaluation comments as input data, and stores the learned model in a storage device. Effect of the Invention

[0007] According to the present invention, it is possible to generate a highly robust model. [Brief description of the drawings]

[0008] [Figure 1] FIG. 1 is a configuration diagram of a learning device 1 according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a flowchart illustrating an example of the operation of the learning device 1. [Diagram 3] FIG. 3 is a diagram illustrating an example of an imaging angle. [Figure 4] FIG. 4 is a flowchart illustrating an example of the operation of the learning device 1. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the description of the drawings, the same parts are given the same reference numerals and the description will be omitted.

[0010] An example of the configuration of the learning device 1 will be described with reference to Fig. 1. As shown in Fig. 1, the learning device 1 includes a communication I / F 10, a controller 20, and a storage device 30. The learning device 1 is mounted on a general-purpose computer, as an example.

[0011] The communication I / F 10 is implemented as hardware such as a network adapter, various communication software, or a combination of these, and is configured to realize wired or wireless communication via a network. The communication I / F 10 also has functions as an input unit and an output unit for transmitting and receiving data. In this embodiment, the communication I / F 10 will be described as performing Internet communication.

[0012] The storage device 30 is composed of a hard disk drive (HDD), a solid state drive (SSD), etc. A plurality of databases are stored in the storage device 30. As shown in Fig. 1, the plurality of databases include an image database 31, an angle data set 32, a three-dimensional vehicle body model data set 33, and a learning model database 34. Details of each database will be explained together with each function of the controller 20.

[0013] The controller 20 is an electronic control unit (ECU) having a central processing unit (CPU), a read only memory (ROM), a random access memory (RAM), a controller area network (CAN) communication circuit, and the like. A computer program for functioning as the learning device 1 is installed in the controller 20. By executing the computer program, the controller 20 functions as a plurality of information processing circuits included in the learning device 1. Note that, here, an example is shown in which the plurality of information processing circuits included in the learning device 1 are realized by software, but it is of course possible to configure the information processing circuits by preparing dedicated hardware for executing each information processing shown below. In addition, the plurality of information processing circuits may be configured by individual hardware. The controller 20 includes a data acquisition unit 21, an angle estimation unit 22, a classification unit 23, a comparison unit 24, a generation unit 25, and a learning unit 26 as a plurality of information processing circuits.

[0014] Next, the functions of the data acquisition unit 21, the angle estimation unit 22, and the classification unit 23 will be described with reference to FIGS.

[0015] The data acquisition unit 21 acquires vehicle images (two-dimensional images) from the Internet via the communication I / F 10 (step S101 in FIG. 2). When acquiring vehicle images, the data acquisition unit 21 also acquires user comments on the images (sometimes called evaluation comments). In order to acquire user comments on the images, the data acquisition unit 21 acquires vehicle images and user comments on the images mainly from word-of-mouth sites, SNS (Social Networking Service), or sites that publish survey results. As shown in FIG. 1, the data acquisition unit 21 associates the acquired images with the comments and stores them in the image database 31.

[0016] The data acquisition unit 21 also acquires the angle at which the vehicle was imaged (the angle seen from the camera with respect to the vehicle) and the distance from the camera to the vehicle at which the vehicle was imaged. Hereinafter, the "angle at which the vehicle was imaged" may be simply referred to as the "image capture angle," and the "distance from the camera to the vehicle at which the vehicle was imaged" may be simply referred to as the "image capture distance." In this embodiment, the "image capture angle" is defined as an angle when the vehicle is imaged from directly in front as shown in FIG. 3, and the angle increases clockwise. After 180 degrees, the positive and negative signs are reversed. Note that such a definition of the image capture angle is an example, and other definitions may be used. The data acquisition unit 21 associates the acquired image, comment, image capture angle, and image capture distance, and stores them in the image database 31. Since the method of associating data is well known, a description thereof will be omitted.

[0017] The angle estimation unit 22 reads out the image stored in the image database 31 and estimates the imaging angle and imaging distance (step S103 in FIG. 2). In the above description, it has been described that the imaging angle and imaging distance are acquired by the data acquisition unit 21. However, when an image is acquired, there are cases where the imaging angle and imaging distance cannot be acquired together. For example, this corresponds to a case where only the image and the user's comment are made public. In such a case, the angle estimation unit 22 estimates the imaging angle and imaging distance. More specifically, the angle estimation unit 22 estimates the imaging angle and imaging distance for an image that is not associated with the imaging angle and imaging distance among the images acquired by the data acquisition unit 21. Therefore, if the imaging angle and imaging distance are associated with all the images acquired by the data acquisition unit 21, the angle estimation unit 22 becomes unnecessary. The method of estimating the imaging angle and imaging distance from the image is not particularly limited, but for example, a method of estimating using a machine learning model can be mentioned. The angle estimation unit 22 associates the estimated imaging angle and imaging distance with the image and stores them in the image database 31.

[0018] The classification unit 23 classifies the images stored in the image database 31 into predetermined angles according to the imaging angle (step S105 in FIG. 2). The predetermined angle is, for example, 45 degrees as shown in FIG. 3. The reference numeral 40 in FIG. 3 indicates a range of -22.5 degrees to 22.5 degrees. Similarly, the reference numeral 41 indicates a range of 22.5 degrees to 67.5 degrees, the reference numeral 42 indicates a range of 67.5 degrees to 112.5 degrees, the reference numeral 43 indicates a range of 112.5 degrees to 157.5 degrees, the reference numeral 44 indicates a range of 157.5 degrees to -157.5 degrees, the reference numeral 45 indicates a range of -157.5 degrees to -112.5 degrees, the reference numeral 46 indicates a range of -112.5 degrees to -67.5 degrees, and the reference numeral 47 indicates a range of -67.5 degrees to -22.5 degrees. For example, when the imaging angle is 0 degrees, it is classified into class 40. In this way, the classification unit 23 classifies the images stored in the image database 31 into eight classes according to the imaging angle. The classification unit 23 stores the classification results in the angle dataset 32. The classification unit 23 also stores the number of images classified into each class (classes 40 to 47) in the angle dataset 32 ​​(step S107 in FIG. 2). Comments, imaging angles, and imaging distances are also associated with the classified images.

[0019] Next, the functions of the comparison unit 24 and the generation unit 25 will be described with reference to FIG. 4. The number of pairs of images and comments (the number of images classified into each class) may differ depending on the imaging angle. For classes with a small number of pairs, data needs to be added. In step S201, the comparison unit 24 acquires classified data by referring to the angle data set 32. The process proceeds to step S203, where the comparison unit 24 compares the number of pairs of each class. If there is a difference in the number of pairs of each class, the classes are classified into the class with the largest number of pairs and the other classes. There may be a plurality of classes with the largest number of pairs. The comparison unit 24 outputs the comparison result to the generation unit 25. Note that there may be no difference in the number of pairs of each class. Here, it is assumed that the class with the largest number of pairs is class 40. In other words, it is assumed that the number of pairs of classes 41 to 47 is smaller than the number of pairs of class 40 (YES in step S205).

[0020] The generating unit 25 generates data based on the result by the comparing unit 24. The generating unit 25 generates data until the number of pairs of classes 41 to 47 reaches the number of pairs of class 40. An example of a data generating method will be described with reference to steps S209 to S217. In step S207, the generating unit 25 refers to the angle dataset 32 ​​and acquires a set of pairs (a series of datasets) from the class with the largest number of pairs (class 40 in this case). The "series of datasets" refers to a dataset in which an image of a vehicle, a user's comment on the image, an imaging angle, and an imaging distance are linked. The process proceeds to step S209, where the generating unit 25 extracts a vehicle region from the acquired image. The "car region" refers to an area in the image in which a vehicle is captured.

[0021] The process proceeds to step S211, where the generation unit 25 divides the vehicle region into its individual parts using semantic segmentation, and extracts the shape features of each part. Semantic segmentation is a deep learning algorithm that associates a label or category with every pixel in an image, and is a well-known technique. By using semantic segmentation, for example, shape features such as a grille and headlights are extracted from a front view image (an image classified as class 40).

[0022] The process proceeds to step S213, where the generation unit 25 searches for a three-dimensional model by referring to the three-dimensional body model data set 33. The three-dimensional body model data set 33 stores three-dimensional model data of the body for each vehicle type. Furthermore, the three-dimensional body model data set 33 stores shape features for each part of each vehicle type. The generation unit 25 compares the shape features extracted in step S211 by referring to the three-dimensional body model data set 33, and outputs three-dimensional model data with the highest similarity. Note that although the three-dimensional body model data set 33 stores three-dimensional model data of the body for each vehicle type, shape features for each part of each vehicle type may not be stored. In this case, the generation unit 25 may output three-dimensional model data with the highest similarity by the following process. Since the generation unit 25 acquires a series of data sets by referring to the angle data set 32, it knows the imaging angle and imaging distance. The generation unit 25 takes a photo of the three-dimensional model data for each vehicle type (a plurality of different three-dimensional model data) using the imaging angle and imaging distance. This process means taking a virtual photo from a virtual viewpoint. As a result, a virtual image captured at a predetermined distance (imaging distance) and a predetermined angle (imaging angle) is obtained for each of the three-dimensional model data for each vehicle type. The generation unit 25 compares each of the multiple virtual images with the vehicle images related to the series of data sets, and outputs the three-dimensional model data related to the virtual image with the highest similarity as the three-dimensional model data with the highest similarity. Note that the similarity includes at least one of pixel value similarity, edge similarity, and shape feature similarity.

[0023] The process proceeds to step S215, where the generator 25 uses the three-dimensional model data output in step S213 to take pictures from a predetermined distance for the angle classes with the fewest number of pairs (classes 41 to 47 in this case). This allows virtual images to be obtained for classes 41 to 47. Note that the angle may be any angle for each class. For example, when acquiring a virtual image for class 41, this means that the imaging angle is not limited as long as it is between 22.5 degrees and 67.5 degrees.

[0024] The process proceeds to step S217, where the generator 25 processes the virtual image. Specifically, the generator 25 acquires features such as color and brightness of images related to a series of data sets acquired from the class with the largest number of pairs (class 40 in this example), and assigns the acquired color, brightness, and the like to the virtual image. The process proceeds to step S219, where the generator 25 associates a comment, an imaging angle, and an imaging distance randomly selected from the same class with the smallest number of pairs with the processed virtual image, and stores them in the angle data set 32. As a result, the number of pairs in classes 40 to 47 becomes equal. If step S205 is NO, the generator 25 does not perform the process.

[0025] Next, the function of the learning unit 26 will be described. The learning unit 26 acquires a series of data sets for each class (classes 40 to 47) with reference to the angle data set 32 ​​to generate a learning model. The learning unit 26 generates a learning model (evaluation comment generation model) using three pieces of data, namely, an image of a vehicle, an imaging angle, and an imaging distance, from the series of data sets as input data, and an evaluation comment on the image of the vehicle as output data. The model generation method is not particularly limited, and a well-known method is used. The learning unit 26 compares the evaluation comment generated by the evaluation comment generation model with the evaluation comment of an actual user, and calculates an error. Examples of the method of calculating the error include a method of comparing keywords, and a method of inputting the data into a machine learning language model represented by Transformer, BERT, etc., calculating the difference in the features of the output sentence, and outputting it as an error. The learning unit 26 performs backward processing using the calculated error, and updates the parameters of the learning model by a backpropagation method. The learning unit 26 stores the updated parameters in the learning model database 34. The learning unit 26 repeats learning hundreds of times for all data sets stored in the angle data set 32, and searches for parameters of a model with a small error.

[0026] When an arbitrary image of a vehicle, the angle at which the image was captured, and the distance from the camera to the vehicle at the time the image was captured are input to the learning model (trained model) generated by the learning unit 26, an evaluation comment for the image is output. In order to output an image of the entire vehicle, the controller 20 generates images of the vehicle at different angles from a three-dimensional model, and generates evaluation comments for the images at each angle using the trained model. The controller 20 extracts adjectives related to keywords of specific parts from the generated evaluation comments, and judges the quality of the parts. However, if an adjective is extracted even though there is no keyword for the part, the controller 20 records it as the quality of the entire vehicle. The controller 20 maps the evaluation of the entire vehicle and the evaluation of each part to the three-dimensional model data.

[0027] (Action and effect) As described above, the learning device 1 according to this embodiment provides the following advantageous effects.

[0028] The learning device 1 includes a storage device 30 that associates and stores a first image of a vehicle with a user's evaluation comment on the first image, and a controller 20. The first image is an image acquired from the Internet via a communication I / F 10. The controller 20 identifies the vehicle type depicted in the first image. The controller 20 generates a second image different from the first image using three-dimensional model data corresponding to the identified vehicle type. The controller 20 learns a model related to the user's evaluation comment using the first image, the second image, and the user's evaluation comment as input data, and stores the learned model in the storage device 30. According to the learning device 1, it is possible to increase the number of images by generating images at a small number of angles using a three-dimensional model of the identified vehicle type in images captured at a small number of angles. Then, since the model is learned after the number of images is increased, it is possible to generate a learning model with high robustness.

[0029] The storage device 30 stores three-dimensional model data of the vehicle body for each vehicle model, and shape features for each part of each vehicle model. The controller 20 extracts shape features of each part of the vehicle shown in the first image. The controller 20 compares the extracted shape features with shape features stored in the storage device 30, and outputs three-dimensional model data with the highest similarity. The similarity includes at least one of pixel value similarity, edge similarity, and shape feature similarity. The learning device 1 searches for three-dimensional model data of the vehicle body based on the shape features of the parts, making it possible to identify the vehicle model from the features of small parts such as the grille and headlights.

[0030] The storage device 30 stores three-dimensional model data of the vehicle body for each vehicle model. The controller 20 extracts the shape characteristics of each part of the vehicle shown in the first image. If the shape characteristics of each part of each vehicle model are not stored in the storage device 30, the controller 20 acquires the angle at which the first image was captured and the distance from the camera to the vehicle at the time the first image was captured. The controller 20 acquires a virtual image for each vehicle model by virtually taking a photograph using the imaging angle and imaging distance for the three-dimensional model data for each vehicle model. The controller 20 compares each of the multiple virtual images with the first image and outputs the three-dimensional model data with the highest similarity. The similarity includes at least one of the similarity of pixel values, the similarity of edges, and the similarity of shape characteristics. This makes it possible to output the three-dimensional model data with the highest similarity even if the shape characteristics of each part of each vehicle model are not stored in the storage device 30.

[0031] A plurality of vehicle images including a first image are stored in the storage device 30. The controller 20 classifies the plurality of images according to the angles at which the images were captured. As an example, the controller 20 classifies the images into classes 40 to 47 according to the angles (see FIG. 3). The controller 20 obtains the color and brightness from the image classified into the angle with the largest number of images, and processes the second image corresponding to an angle other than the angle with the largest number of images using the color and brightness. This makes it possible to increase the variety of the added images. An example of the processing is the addition of color and brightness.

[0032] The controller 20 generates an image with an angle different from that of the first image using the output three-dimensional model data, and generates an evaluation comment from the model learned for each angle. This makes it possible to analyze the quality of the evaluation comment.

[0033] The controller 20 maps the evaluation of the entire vehicle and the evaluation of each part onto the three-dimensional model data. This makes it possible to intuitively grasp the quality of the entire vehicle and each part for each vehicle model.

[0034] Each of the functions described in the above embodiments may be implemented by one or more processing circuits. A processing circuit includes a programmed processing device, such as a processor including electrical circuitry. A processing circuit also includes devices such as application specific integrated circuits (ASICs) or circuit components arranged to perform the described functions.

[0035] As described above, the embodiment of the present invention has been described, but the description and drawings forming a part of this disclosure should not be understood as limiting this invention. From this disclosure, various alternative embodiments, examples and operating techniques will become apparent to those skilled in the art. [Explanation of symbols]

[0036] 1 Learning device, 20 Controller, 30 Storage device

Claims

1. a storage device that stores a first image of a vehicle and a user's evaluation comment on the first image in association with each other; A controller, The controller: Identifying the type of the vehicle shown in the first image; generating a second image different from the first image using three-dimensional model data corresponding to the identified vehicle type; Learning a model regarding the user's evaluation comments using the first image, the second image, and the user's evaluation comments as input data; The learned model is stored in the storage device. A learning device characterized by:

2. The storage device stores three-dimensional model data of a vehicle body for each vehicle type, and shape characteristics of each part of each vehicle type, The controller: Extracting shape features of each part of the vehicle shown in the first image; comparing the extracted shape features with shape features stored in the storage device, and outputting the three-dimensional model data having the highest similarity; The similarity includes at least one of the similarity of pixel values, the similarity of edges, and the similarity of the shape features.

2. The learning device according to claim 1 .

3. The storage device stores three-dimensional model data of a vehicle body for each vehicle type, The controller: Extracting shape features of each part of the vehicle shown in the first image; If the storage device does not store the shape characteristics of each part of each vehicle model, the angle at which the first image was captured and the distance from the camera to the vehicle at which the first image was captured are acquired; A virtual image of each vehicle model is obtained by virtually taking a photograph using the angle and the distance for the three-dimensional model data of each vehicle model; comparing each of the plurality of virtual images with the first image, and outputting three-dimensional model data having the highest similarity; The similarity includes at least one of the similarity of pixel values, the similarity of edges, and the similarity of the shape features.

2. The learning device according to claim 1 .

4. A plurality of images of the vehicle including the first image are stored in the storage device; The controller: classifying the plurality of images, including the image of the vehicle, according to angles at which the images were captured; Obtain the color and brightness from the image classified into the angle with the most number of images, The second image corresponding to an angle other than the most common angle is processed using the color and brightness.

4. The learning device according to claim 1, wherein the learning device is a learning device that performs a learning process.

5. The controller: Using the outputted three-dimensional model data, an image is generated at an angle different from that of the first image, and an evaluation comment is generated for each angle from the learned model.

4. The learning device according to claim 2 or 3.

6. The controller: The evaluation of the entire vehicle and the evaluation of each part are mapped onto the three-dimensional model data.

6. The learning device according to claim 5.

7. A learning method for a learning device including a storage device that stores a first image of a vehicle and a user's evaluation comment on the first image in association with each other, and a controller, the learning method comprising: The controller: Identifying the type of the vehicle shown in the first image; generating a second image different from the first image using three-dimensional model data corresponding to the identified vehicle type; Learning a model regarding the user's evaluation comments using the first image, the second image, and the user's evaluation comments as input data; The learned model is stored in the storage device. A learning method comprising:

Citation Information

Patent Citations

  • Evaluation device, evaluation method, and evaluation program

    JP2018195078A

  • Cooking support device, learning device, cooking support method, learning method and program

    JP2020181436A