Method for generating training data, device for generating training data, program for generating training data, and system for generating training data
By simulating vehicle vibrations and integrating biological sensors, the method enhances the accuracy of AI models estimating driver emotions by training them with data that replicates real-world driving conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DENSO TEN LTD
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Collecting training data for AI models that estimate driver emotions using facial images is challenging, especially in simulated environments, as the images lack the vibrational effects present in actual vehicle use, leading to reduced output accuracy.
Generate learning data by simulating vehicle vibrations in a controlled environment, applying image fluctuation processing to facial images to mimic real-world driving conditions, and integrating biological sensors for emotion estimation.
Improves the accuracy of AI models by training them with data that replicates the effects of actual vehicle driving, ensuring the models perform as effectively as those trained with real-world data.
Smart Images

Figure 2026069896000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a learning data generation method, a learning data generation device, a learning data generation program, and a learning data generation system.
Background Art
[0002] Conventionally, there is a technology for estimating the driver's emotion from the expression of the face of the vehicle driver. In such a technology, a technology for estimating the driver's emotion from the driver's face image using an AI (Artificial Intelligence) model has been proposed. For example, Patent Document 1 discloses a technology for specifying the face direction from a face image and estimating the emotion based on the face direction.
[0003] In addition, when learning an AI model using learning data such as face images, there are a method of collecting face images of the driver in a state of actually getting in the vehicle and a method of collecting face images in a simulated environment simulating a vehicle such as a laboratory.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Collecting training data requires considerable effort, and collecting data using actual vehicles is particularly difficult. Therefore, from the perspective of reducing effort, it is desirable to collect training data in a simulated driving environment that mimics a vehicle, but there was a concern that the output accuracy of the AI model would be low. Specifically, when training an AI model using facial images collected in a simulated driving environment, the facial images collected in the simulated environment are unaffected by vibrations etc. that occur when the vehicle is running, but the images input to the AI model in the actual usage scenario are affected by such vibrations etc. Thus, because there is a difference between the facial images used for training and the facial images input to the AI model due to vehicle operation (vibrations etc.), there was a risk that the output accuracy of the AI model trained using facial images collected in a simulated environment would be reduced.
[0006] The present invention has been made in view of the above, and aims to provide a method for generating training data, a training data generation device, a training data generation program, and a training data generation system that can improve the output accuracy of an AI model. [Means for solving the problem]
[0007] To solve the above-mentioned problems and achieve the objective, the learning data generation method according to the present invention is a learning data generation method that generates learning data to be used for training a learning model that takes captured image-related data based on images of a user taken by an in-vehicle camera as input and outputs user state estimation data indicating the state of the user, and collects simulated environment subject image-related data relating to images of a subject taken in a simulated driving environment of a vehicle, calculates the image fluctuation state that is estimated to be experienced by the simulated environment subject image during actual driving based on the vehicle driving conditions in the simulated driving environment, processes the simulated environment subject image-related data based on the image fluctuation state to generate learning image-related data, collects learning user state estimation data corresponding to the user state estimation data for the subject in the simulated driving environment, and generates learning data using the learning image-related data and the learning user state estimation data. [Effects of the Invention]
[0008] According to the present invention, by applying processing to simulated driving environment subject image-related data collected in a simulated driving environment, which adds fluctuation elements that occur during actual driving, such as simulated vibration processing, it is possible to generate simulated driving environment subject image-related data that is affected by actual vehicle driving, such as vehicle vibration. Then, by training an AI model with training data based on the generated simulated driving environment subject image-related data, the AI model will learn based on images affected by actual vehicle driving, just like the captured image-related data used when the model is in use. Therefore, the AI model will achieve performance equivalent to that of an AI model trained with training data created in an actual vehicle driving environment, and the output accuracy of the AI model can be improved. [Brief explanation of the drawing]
[0009] [Figure 1A] Figure 1A is an explanatory diagram of the learning system according to an embodiment. [Figure 1B] Figure 1B is an explanatory diagram of the emotion estimation system according to the embodiment. [Figure 2] Figure 2 shows an example of the configuration of a learning data generation device according to an embodiment. [Figure 3] Figure 3 shows an example of simulated vibration machining. [Figure 4] Figure 4 shows an example of simulated vibration machining. [Figure 5] Figure 5 is a flowchart showing the processing procedure of the learning data generation process performed by the learning data generation device according to the embodiment. [Modes for carrying out the invention]
[0010] The learning data generation method, learning data generation apparatus, learning data generation program, and learning data generation system according to the embodiments will be described in detail below with reference to the attached drawings. However, the present invention is not limited to the embodiments shown below.
[0011] First, an overview of the learning system and emotion estimation system according to the embodiment will be described using Figures 1A and 1B. Figure 1A is an explanatory diagram of the learning system LS according to the embodiment. Figure 1B is an explanatory diagram of the emotion estimation system ES according to the embodiment.
[0012] In the learning system LS shown in Figure 1A, a learning model is trained to output biometric information for estimating the driver's emotions. In the emotion estimation system ES shown in Figure 1B, facial images of the driver while actually driving are input to the learning model to estimate emotional information. In this disclosure, a learning model is given as an example that takes the driver's facial image itself as input and outputs the driver's emotions (biometric information). However, a learning model may also take image features extracted from the driver's facial image (such as gaze, coordinates of facial parts (eyes, nose, mouth, ears, etc.)) as input and output the driver's emotions, or emotional indicators used for emotion estimation (such as arousal level, activity level), or basic biometric information such as electroencephalogram and heart rate.
[0013] The learning system LS shown in Figure 1A includes a learning data generation device 1, a learning device 200, and a simulation device 100. The simulation device 100 is, for example, a driving simulator and provides a simulated driving environment that simulates a vehicle. Specifically, the simulation device 100 has vehicle driving control components (steering wheel, accelerator pedal, brake pedal) and a display unit, and changes the display on the screen of the display unit (simulated screen of the vehicle's surroundings) in response to the operation of the driving control components (simulated driving operation) by the subject, who is the driver. As a result, the subject can experience the feeling of actually driving in the simulated driving space.
[0014] The simulation device 100 outputs image fluctuation data that is presumed to affect the driver's (subject's) facial image during actual driving. This image fluctuation data could include vehicle vibrations based on road surface conditions (road surface irregularities), and vehicle acceleration (linear acceleration and angular acceleration due to turning) generated in the vehicle based on the subject's braking, acceleration, and steering wheel operations. Note that vehicle acceleration is the acceleration applied to the subject and will affect the subject's facial image. In this example, the image fluctuation data will be explained as vehicle vibration data (including vehicle acceleration (a type of vibration)).
[0015] The simulation device 100 outputs simulated vehicle vibration data D10 (simulation data D10) during simulated driving, based on a simulated screen of the vehicle's surroundings and the subject's operation of the driving control components. Specifically, the simulation device 100 calculates vehicle vibrations generated during vehicle movement based on the road surface conditions (road surface irregularities) of the simulated driving route on the simulated screen, and also calculates vehicle acceleration based on the subject's operation of the driving control components. Based on these vehicle vibrations and accelerations, it calculates vibration data. For the generation of simulated images (videos), the simulation device 100 calculates the vehicle's movement (velocity, acceleration) based on the amount of operation on the driving control components using predetermined calculation formulas and table data, and moves (creates) background images such as roads in accordance with the calculated vehicle movement. This vehicle movement (velocity, acceleration) data is used as acceleration, etc., for creating vibration data. Furthermore, since the simulation device 100 has road surface image data (part of the background data) for generating simulated images (videos), it can generate vehicle vibration data through analysis processing of this road surface image data (recognition of unevenness by image recognition processing), etc.
[0016] The learning data generation device 1 receives simulation data D10 (vibration data D10) from the simulation device 100. If the simulation device 100 is designed as a simulation device for the learning system LS, it can have the above functions. However, when using a simulation device in general, the learning data generation device 1 must receive simulation data D10, such as image data and, if possible, driving and operation data, from the simulation device 100, and generate vibration data by performing processing such as the image analysis described above. In other words, the learning data generation device 1 is equipped with a vibration data generation unit 11, which generates vibration data D11 that represents vibrations that occur during actual vehicle operation based on the simulation data D10.
[0017] The learning data generation device 1 is equipped with a camera 14 and captures images of a subject HS who is simulating the operation of the simulation device 100. Specifically, the camera 14 is attached to the simulation device 100 and captures video of the subject during the simulated operation, outputting subject image D14 (without vehicle vibration component). In other words, the learning data generation device 1 uses the camera 14 to capture time-series images of the subject in the simulated environment (a series of still images with a frame period, i.e., a video). The face image processing unit 15 of the learning data generation device 1 takes subject image D14 as input and performs various image processing such as image analysis to extract the subject's face image, noise reduction, contrast / brightness adjustment, etc., to produce and output face image data D15 (without vehicle vibration component). The face image processing unit 15 takes face image data D15 as input and vibration data D11 as input, and performs processing to add a vibration component to face image data D15 to produce and output face image data D16 (with vehicle vibration component). For example, facial image data D16 is a moving image in which time-varying blurring occurs in the image due to vibration compared to facial image data D15.
[0018] In addition, the learning data generation device 1 includes a biological sensor 12 (detecting brain waves, heartbeats, etc. for emotion estimation), and acquires the biological information D12 of the subject HS wearing the biological sensor 12. The biological signal type emotion estimation unit 13 inputs the biological information D12, estimates the emotion of the subject HS based on the biological information D12, and outputs emotion data D13 indicating the emotion. Regarding the emotion estimation method, various known emotion estimation methods can be applied.
[0019] Then, the learning data generation unit 17 inputs the emotion data D13 and the face image data D16, generates and outputs learning data D17 with the emotion data D13 as the target variable and the face image data D16 (with vehicle vibration components) as the explanatory variable. Note that the emotion data D13 and the face image data D16 constituting the emotion data D13 are synchronized data. That is, the detection period of the biological signal from which the emotion data D13 constituting the emotion data D13 is derived and the shooting period of the camera image from which the face image data D16 is derived are of the same timing (the sampling time lengths are set to appropriate time lengths respectively). Thereby, the face image data D16 in the learning data D17 becomes face image data simulating the influence of vibrations generated during actual vehicle driving.
[0020] Then, the learning device 200 sequentially acquires the learning data D17 output by the learning data generation device 1 (learning data generation unit 17), accumulates the learning data D17 in the amount required for learning to generate a learning data set. Then, using the generated learning data set, a pre-learning (emotion estimation) model MB is learned by a learning method such as the error backpropagation method.
[0021] Next, the emotion estimation system ES will be described using FIG. 1B. The emotion estimation system ES includes a camera 21 and images a driver DR driving a vehicle. Specifically, the camera 214 is a camera installed on the front window of the vehicle or the like and images the driver DR, and for example, a camera such as a drive recorder can be used. Also, the emotion estimation system ES may be configured to be built into a drive recorder. The camera 21 captures a video of the driver DR during driving and outputs driver image data D21 (including vehicle vibration components because it is during actual vehicle driving). The face image processing unit 22 inputs the driver image data D21 and performs various image processes such as image analysis processing to cut out the face image of the driver DR, noise removal processing, and appropriate processing to make the image suitable for emotion estimation data such as contrast / brightness adjustment, and generates and outputs driver face image data D22 (including vehicle vibration components). The learned model MA is a model obtained by training the pre-trained model MB. Then, the learned model MA inputs the driver face image data D22 and outputs emotion data D23. Note that the driver face image data D22 input to the learned model MA is preferably an image in the same format as the face image data D16 in the training data D17, with the same sampling period, resolution, and frame rate of the moving image, and also preferably the same processing in the image processing. And the emotion data output by the emotion estimation system ES (learned model MA) is used for control in various devices that perform control using emotion data, such as a driving support device or the like.
[0022] Note that in the above-described learning system LS and emotion estimation system ES, various processes are performed using the image data itself, but it is also possible to process the image data obtained by converting the face image data into feature data indicating the features of the face, for example, the image data consisting of the positions (coordinates) of each part (eyes, nose, mouth, etc.) constituting the face and the data group of the line of sight. Also, the expression of image-related data is used as an expression including the image data itself and the feature data of the image.
[0023] Next, we will explain the learning process and the estimation process, which are performed in the learning system LS and the emotion estimation system ES, using specific examples. In the following explanation, we will use an example that uses (face) image features obtained by converting face image data into data that represents the features of the face.
[0024] (Learning process) In the learning process (executed by the learning system LS), a processing operation called simulated vibration processing is applied to image features (face image data D15), which are data related to the simulated environment subject image (subject image D14) extracted from the simulated environment subject image. Specifically, first, the learning data generation device 1 acquires time-series simulated environment subject images (i.e., video) from the camera 14 over a predetermined period length suitable for learning data, and also acquires time-series vibration data from the simulation device 100. Then, the learning data generation device 1 sets a detection frame Rn (see Figure 3) to be extracted as a face image in the learning data from each frame image (still image) in the simulated environment subject image (video) over the predetermined period length. Note that the target frame images to be processed are not all frame images that make up the video, but may be frame images at predetermined intervals determined considering the accuracy of emotion estimation, processing load, etc. In this case, the detection frame Rn is the position moved based on the vibration data, specifically the vibration displacement value (amplitude and direction) at the time the target frame image is captured. For example, as shown in Figure 3, when there is no vehicle vibration (images P1a, P3a), detection frames R1 and R3 are set to a predetermined reference area (the area around the subject's face). On the other hand, when there is upward vehicle vibration (riding over a step) (image P2a), detection frame R2 moves upward. Figure 3 shows the image processing details, but the processing of image features is as follows. Image features are, for example, the coordinates of the subject's gaze and facial features while driving, as described above, and can be detected using known methods with facial images. In the processing of image features, processing is performed such as adding the vibration displacement value at the corresponding timing in the vibration data to the image feature data of each frame image to be processed in the video. In other words, simulated vibration processing is an image processing or numerical correction process that simulates the effect (image fluctuation state) that the image actually receives due to vehicle vibration. That is, since the subject image (image features) in the simulated environment does not include image shaking due to vehicle vibration, the learning data generation device 1 applies simulated vibration processing so that the image (image features) includes the component of the above-mentioned image shaking (image fluctuation state).
[0025] Specifically, the learning data generation device 1 acquires acceleration information (estimated value: detected by analysis of simulated driving operation data, simulated driving images, etc.) from the simulation device 100, and calculates the image fluctuation state that is estimated to occur to the subject image in the simulated environment during actual driving based on the acquired acceleration (vehicle driving conditions). Then, the learning data generation device 1 applies simulated vibration processing to the subject image-related data in the simulated environment according to the image fluctuation state. For example, if the learning data generation device 1 acquires acceleration information indicating that the vehicle has driven over a step (details will be described later), it applies simulated vibration processing according to the effect (image fluctuation state) that occurs in the image when the vehicle actually drives over a step. For example, as simulated vibration processing when the vehicle drives over a step, the learning data generation device 1 adds vibration components (vibration displacement values) to the facial feature data (position coordinates of each facial part). More specifically, if the vibration displacement is n pixels upward, the learning data generation device 1 applies a correction (movement of the cropping position in the case of image data, or numerical correction processing (addition of n pixels) to the position coordinates of each face part in the case of face feature data) by shifting the coordinates of the face image detection frame R (face image cropping frame) upward by n pixels as a simulated vibration processing. In other words, when the vehicle goes over a bump, the learning data generation device 1 moves the entire vehicle (camera position) relatively upward relative to the driver's face position (instantaneously). Because an effect occurs (vibration in the vertical direction), (humans have elastic and vibration-absorbing properties, and the effect of vehicle vibration on human movement is reduced), a simulated vibration processing is applied to mimic this effect. This allows the learning data generation device 1 to include the vibration component that occurs when the vehicle goes over a bump in the face image data (face image-related data) in the learning data. The learning data generation device 1 then generates learning data using the learning image-related data that has undergone simulated vibration processing and the corresponding learning user state estimation data (emotion data). The learning device 200 then trains the pre-training model MB using a dataset consisting of a suitable amount of learning data for training.
[0026] (Inference process) The emotion estimation system ES then uses the trained model MA, which was trained during the learning process, to estimate the emotions of the driver actually operating the vehicle (inference process). This trained model MA is installed (stored) in an in-vehicle device such as a dashcam or navigation system. The emotion estimation system ES extracts image features from the driver's face image (video: driver image data D21) captured by the camera 21 installed in the vehicle for a predetermined period of time (the same period as the learning process) (driver face image (related) data D22). In other words, the emotion estimation system S extracts image features for a predetermined period of time that are affected by actual vehicle vibrations (acceleration effect). The emotion estimation system S inputs the extracted (captured) face image related data for the predetermined period of time into the trained model MA. When the trained model MA receives the (captured) face image related data for the predetermined period of time, it outputs emotion data (user state estimation data). Furthermore, if the trained MA model was trained using facial images, then the facial image-related data will be facial image data, and this facial image data will be input to the trained MA model.
[0027] In the example above, the learning model was one that estimated emotions from facial image-related data, but it could also be a model that estimates biometric information or various biometric values used to estimate emotions. For example, emotions can be estimated based on the level of arousal of the central nervous system and the activity level of the autonomic nervous system. The level of arousal can be derived, for example, from electroencephalogram (EEG) data. The activity level can be derived, for example, from heart rate data. The level of arousal is calculated from the beta / alpha waves of the EEG. The activity level is calculated from the standard deviation of the heart rate LF (Low Frequency) component (low frequency component of the heart rate waveform signal). Therefore, it is also possible to construct an emotion estimation system ES that estimates the level of arousal and activity using a model that estimates the level of arousal and activity based on facial image-related data, and then estimates emotions based on the estimated level of arousal and activity. In this case, the target variable in the learning data would be the level of arousal or activity, and the level of arousal or activity calculated from biometric information would be used.
[0028] Thus, in this disclosure, by applying simulated vibration processing to subject image-related data (subject image D14) collected in a simulated driving environment, image features affected by vehicle vibration (training image-related data (face image data D16)) can be generated. Furthermore, in this disclosure, by training a learning model using the training image-related data (image features) that has undergone simulated vibration processing as training data, the training image-related data used for training can be made to be similar in nature (affected by vehicle vibration) to the image-related data used during inference. In other words, since the training image-related data during training can be made closer to the image-related data during inference, the output accuracy of the learning model can be improved.
[0029] In this disclosure, we have given an example where the learning model is a model for estimating emotional information, but it can be applied to models that perform some kind of estimation or control based on images of vehicle occupants (or animals). For example, it could be a device that controls the vehicle's driving characteristics in autonomous driving according to the facial expressions of the vehicle occupants (for example, a control that performs gentle autonomous driving when the occupant has an anxious expression).
[0030] Next, an example of the configuration of the learning data generation device 1 according to the embodiment will be described using Figure 2. Figure 2 is a diagram showing an example of the configuration of the learning data generation device 1 according to the embodiment. As shown in Figure 2, the learning data generation device 1 is connected to the simulation device 100 and the learning device 200. The learning device 200 is a device that trains a pre-trained model using a dataset generated from a quantity of learning data suitable for training generated by the learning data generation device 1, and generates a trained model. Note that the learning data generation device 1 may be configured integrally with at least one of the simulation device 100 and the learning device 200.
[0031] The learning data generation device 1 comprises a controller 2 and a storage unit 3.
[0032] Controller 2 is implemented, for example, by a processor such as a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing various programs stored in a memory device (such as memory unit 3) inside the learning data generation device 1, using RAM or the like as a working area. Alternatively, Controller 2 may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array), or GPGPU (General Purpose Graphic Processing Unit).
[0033] The memory unit 3 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as hard disks and optical discs.
[0034] The controller 2 executes the corresponding program stored in the memory unit 3 to realize the functions of each functional unit shown in Figure 1A (vibration data generation unit 11, biosignal type emotion estimation unit 13, face image processing unit 15, face image data processing unit 16, and learning data generation unit 17). The program can be installed in the learning data generation device 1 (memory unit 3) using a portable recording medium such as an optical disc or memory card, or via online distribution.
[0035] Next, we will explain an example of simulated vibration processing using Figures 3 and 4. Figures 3 and 4 are diagrams illustrating an example of simulated vibration processing. In Figures 3 and 4, we use as an example face images (video) P1a to P3a for a predetermined period of time, taken from the front of a subject operating in the simulation device 100. Also, Figures 3 and 4 show detection frames R1 to R3 (image cropping frames), which are the regions used as image data in the training data. The basic detection frame is set by the design and development team, for example, as a rectangular frame region that can acquire an image appropriate for emotion estimation depending on the camera installation position. The controller 2 then generates detection frames R1 to R3 by processing the basic detection frame by moving, deforming, etc., according to the simulated vibration state, and generates image data for training data by cropping the camera-captured images P1a to P3a within the detection frames R1 to R3.
[0036] First, using Figure 3, we will explain the simulated vibration processing that occurs when the vehicle hits a bump during simulated driving in the simulation device 100. In actual driving, when a vehicle hits a bump, the vehicle body momentarily lifts up, causing the camera's imaging range to momentarily shift upward. On the other hand, the driver is considerably less affected by vehicle vibration (hitting a bump) due to the elasticity and vibration absorption characteristics of the human body, as well as the contact characteristics between the human body and the vehicle (non-fixed position, influence of seat cushioning). For the sake of clarity, we will assume here that there is no effect from vehicle vibration. In other words, because the camera is temporarily shifted upward relative to the driver, an image fluctuation occurs in which the driver's face image captured by the camera momentarily shifts downward. However, in the simulation device 100, even when the vehicle hits a bump, no vibration is applied to the camera, and the camera image is unaffected. As a result, no image fluctuation occurs in the subject's face image. The same applies when various accelerations are applied to the vehicle (including phenomena caused by driving operations such as sudden acceleration and sudden deceleration). Therefore, this disclosure describes processing of camera images in response to vehicle vibrations.
[0037] Specifically, as shown in Figure 3, the controller 2 performs image processing on the simulation device 100, moving the detection frame R2 of image P2a at the time the vehicle hits a step upward compared to the detection frame R1 at the immediately preceding time. The upward movement distance of the detection frame R2 is set to a length (pixel value) corresponding to the vibration fluctuation amount (distance) (this amount corresponds to the optical characteristics of the camera). Then, the detection frame R3 at subsequent times is returned to the same position as the original position of detection frame R1. As a result, the processed image P2b of the subject cut out by detection frame R2 in the simulated environment subject image data that has undergone simulated vibration processing will be data that is instantaneously shifted downward compared to the processed images P1b and P3b immediately before and after (processed image P2b). In other words, the effect that occurs in the image data when the vehicle hits a step is reflected in the simulated environment subject image data. More specifically, the controller 2 performs image processing to shift the position of the driver's face downward within the detection frame R2. As a result, the coordinates of the facial features extracted from within the detection frame R2 are shifted downwards. In other words, the image fluctuation state can be reproduced. This allows the controller 2 to reduce the discrepancy with the image data obtained when the vehicle actually drives over the subject by using the subject image-related data in the simulated environment with simulated vibration processing as training data, thereby improving the output accuracy of the learning model. The controller 2 may also extract image features from image P2b (image P2b cut out by changing the position of the detection frame R2) which has been processed according to the image fluctuation state, or it may apply numerical correction processing according to the image fluctuation state to the image features extracted from the detection frame R2.
[0038] Next, using Figure 4, we will explain the simulated vibration processing that occurs when a subject applies the brakes suddenly while driving in the simulation device 100. In Figure 4, it is assumed that the processed images P1b to P3b are set to be smaller than the image frame size Pmax. Here, in actual driving, when a driver applies the brakes suddenly, the driver's body instantaneously moves forward (leaning forward) due to inertia, and then the body moves backward due to the reaction of the brakes. In other words, the driver instantaneously moves closer to the camera due to the brakes, and then moves away due to the reaction of the brakes, resulting in an image fluctuation state. That is, the image of the subject's face captured by the camera instantaneously becomes larger and then smaller. However, in the simulated driving using the simulation device 100, when a subject applies the brakes suddenly, no acceleration is actually applied to the subject, and the changes in the captured image that occur during actual driving do not occur. Therefore, in this disclosure, the controller 2 acquires brake operation information from the simulation device 100, estimates the acceleration applied to the subject during actual driving based on the acquired brake operation information, and performs simulated vibration processing according to that acceleration.
[0039] Specifically, as shown in Figure 4, the controller 2 reduces the detection frame R2 used to extract the face portion from image P2a at the time of the sudden braking in the simulation device 100 to a smaller size (shrink) than the detection frame R1 at the previous time. Then, the controller 2 enlarges the processed image P2b of detection frame R2, from which the face portion has been extracted, to approximately the same size as the image frame size Pmax. As a result, the face portion in processed image P2b is enlarged compared to the previous processed image P1b. Subsequently, the controller 2 increases the detection frame R3 of image P3a to a larger size (shrink) than detection frame R1 (the detection frame immediately preceding detection frame R2). Then, the controller 2 reduces the processed image P3b of detection frame R3, from which the face portion has been extracted, to a smaller size than the processed image P1b. As a result, the face portion in processed image P2b is reduced compared to the previous processed image P1b. The position of the face (center of gravity) remains unchanged. The degree of reduction and enlargement of the face portion is such that, for example, the greater the acceleration, the greater the reduction and enlargement ratios. As a result, the subject image data under simulated vibration processing, which does not include facial image change components that occur with acceleration during actual driving before processing, will include these facial image change components after processing.
[0040] More specifically, when sudden braking occurs, controller 2 performs image processing by enlarging and then shrinking the driver's (subject's) face. As a result, the coordinates of the facial parts extracted from within detection frames R2 and R3 shift in position according to the enlargement and shrinkage. In other words, the image fluctuation state can be reproduced. This allows controller 2 to reduce the discrepancy with image data obtained when the driver actually performs sudden braking by using subject image-related data in a simulated environment with simulated vibration processing as training data, thereby improving the output accuracy of the learning model. Controller 2 may also extract image features from the image processed according to the image fluctuation state (images cropped by enlarging and shrinking detection frame R), or it may apply numerical correction processing according to the image fluctuation state to the image features extracted from detection frame R. It is preferable to set the size of the image cropping frame and the size of the image for training data to an appropriate size in order to properly perform the image processing (so that the entire face image exists within the training data image (frame) in the processed image).
[0041] In Figures 3 and 4, we have given an example of simulated vibration processing in which the detection frame R is moved, enlarged, and reduced. However, it may also be simulated vibration processing in which the entire image obtained from the simulation device 100 is moved, enlarged, and reduced.
[0042] Next, the processing procedure for the learning data generation process executed by the learning data generation device 1 according to the embodiment will be explained using Figure 5. Figure 5 is a flowchart showing the processing procedure for the learning data generation process executed by the learning data generation device 1 (controller 2) according to the embodiment. The learning data generation process shown in Figure 5 is executed when a generation request is received from the learning device 200 or a generation instruction is received from the learning model administrator.
[0043] As shown in Figure 5, in step S101, the controller 2 acquires image data related to the subject image in the simulated environment from the camera 14 and proceeds to step S102. Specifically, the controller 2 acquires image features extracted from the face image (still image constituting the face video image). In step S102, the controller 2 acquires the subject's biological information from the biosensor 12 and proceeds to step S103. For example, the controller 2 acquires biological information such as the subject's electroencephalogram and heart rate. In step S103, the controller 2 generates vibration data based on the simulation data acquired from the simulation device 100 and proceeds to step S104. For example, the controller 2 calculates vibration data indicating the displacement amount (displacement direction, distance) of the captured image based on the acceleration data in the simulated vehicle driving data from the simulation device 100. In step S104, the controller 2 determines whether the acquired simulated environment subject image-related image data, biometric information, and vibration data have been acquired for a predetermined period of time, which is a time length suitable for a continuous data set suitable for emotion estimation processing units. If they have been acquired (Yes), the controller proceeds to step S105; otherwise, it returns to step S101.
[0044] In step S105, Controller 2 estimates emotions based on acquired biometric information and proceeds to step S106. For example, Controller 2 calculates arousal and activity levels from acquired electroencephalogram and heart rate data, and estimates emotions (types) based on these arousal and activity levels. In step S106, Controller 2 applies simulated vibration processing to the acquired subject image data in the simulated environment using synchronized vibration data, and proceeds to step S107. For example, Controller 2 adds the displacement amount of the vibration data to the position coordinate values of the face parts in the image features. In step S107, Controller 2 generates training data consisting of simulated vibration processing data of the subject image in the simulated environment from the same acquisition period, and emotion data estimated based on biometric information. In step S108, Controller 2 determines whether the acquisition of image data related to the simulated environment subject images has been completed (either the data collection period has ended, or the processing of all previously collected image data related to the simulated environment subject images has been completed). If it is completed (Yes), it proceeds to step S109; otherwise, it returns to step S101. In step S109, Controller 2 aggregates the generated training data to create a training dataset and completes the processing. This generated training dataset will be used to train the pre-trained model MB.
[0045] As described above, the image data generation method according to the embodiment is a learning data generation method that generates learning data to be used for training a learning model (trained model MA) that outputs user state estimation data (estimated emotion) indicating the user's state (emotion) by taking image-related data based on images of the user taken by an in-vehicle camera as input, and collects simulated environment subject image-related data (face image data (no vibration) D15) related to simulated environment subject images (subject image D14 (no vibration)) which are images of the subject taken in a simulated driving environment (simulation device 100), and simulated driving environment Based on the vehicle driving conditions at the boundary (simulation data D10), the image fluctuation state (vibration data D11) that is presumed to be experienced by the subject's image in the simulated environment during actual driving is calculated. Based on the image fluctuation state, processing is applied to the subject's image-related data in the simulated environment to generate training image-related data (face image data (with vibration) D16). Training user state estimation data (emotion data D13) corresponding to the user state estimation data for the subject in the simulated driving environment is collected, and training data (training data D17) is generated using the training image-related data and the training user state estimation data.
[0046] According to this disclosure, by applying simulated vibration processing to subject image-related data collected in a simulated driving environment, it is possible to generate subject image-related data in a simulated environment that is affected by vehicle vibrations during actual vehicle driving. Therefore, it is possible to train a learning model with training data that is close to the training data generated based on data from actual vehicle driving, thereby improving the estimation accuracy of the learning model. In other words, since the image data during training can be made closer to the images used when the model is in use, the output accuracy of the learning model can be improved.
[0047] Furthermore, to generate a learning model that takes image features extracted from the driver's facial image (such as gaze direction and coordinates of facial features (eyes, nose, mouth, ears, etc.)) as input and outputs biometric information such as the driver's emotions, the configurations of the learning system LS shown in Figure 1A and the emotion estimation system ES shown in Figure 1B should be modified as follows.
[0048] (1) The facial image processing unit 15 converts the subject image D14 captured by the camera 14 into facial image features (numerical data such as gaze direction and coordinates of facial parts (eyes, nose, mouth, ears, etc.)) through image analysis processing.
[0049] (2) The face image data processing unit 16 adds vibration data D11 (image displacement data (numerical values) due to vibration) to the image features generated by the face image processing unit 15 to generate face image features (numerical data such as gaze direction and coordinates of face parts (eyes, nose, mouth, ears, etc.)) that have been affected by vibration.
[0050] (3) The face image processing unit 22 converts the driver image data D21 captured by the camera 21 into face image features (numerical data such as gaze direction and coordinates of face parts (eyes, nose, mouth, ears, etc.)) through image analysis processing to generate driver face image data D22, which is then output to the trained model MA.
[0051] Furthermore, to generate a learning model that outputs emotional indicators (arousal level, activity level) used to estimate the driver's emotions, the configuration of the learning system LS shown in Figure 1A and the emotion estimation system ES shown in Figure 1B should be changed as follows.
[0052] (1) The biosignal-type emotion estimation unit 13 calculates emotion indicators based on biometric information D12 from the biosensor and outputs them to the learning data generation unit 17. For example, the biosignal-type emotion estimation unit 13 calculates emotion indicators such as arousal level and activity level based on the biometric information D12 from the biosensor, which are brain waves and heart rate, and outputs them to the learning data generation unit 17.
[0053] (2) Add an emotion estimation model to the emotion estimation system ES that estimates emotions based on emotion indicators output by the trained model MA. For example, add an emotion estimation model to the emotion estimation system ES that estimates emotions based on arousal and activity levels output by the trained model MA.
[0054] Furthermore, to generate a learning model that outputs basic biometric information (electroencephalogram, heart rate) used for estimating the driver's emotions, the configuration of the learning system LS shown in Figure 1A and the emotion estimation system ES shown in Figure 1B should be changed as follows.
[0055] (1) The biosignal-type emotion estimation unit 13 outputs the bioinformation D12 from the biosensor directly to the learning data generation unit 17. For example, the biosignal-type emotion estimation unit 13 outputs electroencephalogram and heart rate, which are bioinformation D12 from the biosensor, to the learning data generation unit 17.
[0056] (2) Add an emotion estimation model to the emotion estimation system ES that estimates emotions based on the brain waves and heart rate output by the trained model MA. For example, add an emotion estimation model to the emotion estimation system ES that calculates the level of arousal and activity based on the brain waves and heart rate output by the trained model MA, and estimates emotions based on these calculated levels of arousal and activity.
[0057] Further effects and modifications can be readily derived by those skilled in the art. Therefore, broader aspects of the present invention are not limited to the specific details and representative embodiments expressed and described above. Accordingly, various modifications are possible without departing from the spirit or scope of the overall concept of the invention as defined by the appended claims and their equivalents. [Explanation of symbols]
[0058] 1. Training data generation device 2 Controllers 3 Storage section 11. Vibration data generation unit 12. Biosensors 13. Biosignal-type emotion estimation unit 14 Cameras 15. Facial Image Processing Unit 16. Facial Image Data Processing Section 17. Training Data Generation Unit 21 Cameras 22 Face Image Processing Unit 100 Simulation devices 200 Learning Devices 214 Cameras D10 Simulation Data D11 Vibration Data D12 Biological Information D13 Emotional Data D14 Subject Image D15 Face image data D16 Face image data D17 Training Data D21 Driver Image Data D22 Driver Face Image Data D23 Emotional Data DR Driver ES Emotion Estimation System HS subjects LS Learning System MA Trained Model MB pre-trained model S Emotion Estimation System
Claims
1. A method for generating training data to be used to train a learning model that outputs user state estimation data indicating the state of the user, using image-related data based on images of the user taken by an in-vehicle camera as input, We collect data related to simulated environment subject images, which are images of subjects taken in a simulated vehicle driving environment. Based on the vehicle driving conditions in the simulated driving environment, the image fluctuation state that is estimated to occur in the subject's image under the simulated environment during actual driving is calculated. Based on the aforementioned image fluctuation state, the subject image-related data under the simulated environment is processed to generate training image-related data. For subjects in the simulated driving environment, training user state estimation data corresponding to the user state estimation data is collected. Training data is generated using the aforementioned training image-related data and training user state estimation data. The method used by the controller to generate training data.
2. The aforementioned captured image-related data and simulated environment subject image-related data are image data. The processing described above is an image processing process of the subject image under the simulated environment based on the image fluctuation state. The method for generating training data according to claim 1.
3. The aforementioned captured image-related data and simulated environment subject image-related data are numerical data extracted from images, and are image feature data that represent image features. The processing described above is a numerical correction process for the image feature data of the subject image under the simulated environment, based on the image fluctuation state. The method for generating training data according to claim 1.
4. The aforementioned images of the subject in the simulated environment are data collected from a subject-capture camera in a simulation device in which the subject operates in a simulated environment. The vehicle driving conditions in the simulated driving environment are data collected from the simulation device. The method for generating training data according to claim 1.
5. A learning data generation device that generates learning data used to train a learning model that outputs user state estimation data indicating the user's state, taking as input image-related data based on images of the user taken by an in-vehicle camera, and comprising a controller, The aforementioned controller, We collect data related to simulated environment subject images, which are images of subjects taken in a simulated vehicle driving environment. Based on the vehicle driving conditions in the simulated driving environment, the image fluctuation state that is estimated to occur in the subject's image under the simulated environment during actual driving is calculated. Based on the aforementioned image fluctuation state, the subject image-related data under the simulated environment is processed to generate training image-related data. For subjects in the simulated driving environment, training user state estimation data corresponding to the user state estimation data is collected. Training data is generated using the aforementioned training image-related data and training user state estimation data. A device for generating training data.
6. A training data generation program that generates training data to be used to train a learning model that outputs user state estimation data indicating the state of the user, taking as input image-related data based on images of the user taken by an in-vehicle camera, We collect data related to simulated environment subject images, which are images of subjects taken in a simulated vehicle driving environment. Based on the vehicle driving conditions in the simulated driving environment, the image fluctuation state that is estimated to occur in the subject's image under the simulated environment during actual driving is calculated. Based on the aforementioned image fluctuation state, the subject image-related data under the simulated environment is processed to generate training image-related data. For subjects in the simulated driving environment, training user state estimation data corresponding to the user state estimation data is collected. Training data is generated using the aforementioned training image-related data and training user state estimation data. A training data generation program executed by the controller.
7. A simulation device in which subjects perform simulated driving of a vehicle in a simulated driving environment, A learning data generation system comprising: a learning data generation device that generates learning data used to train a learning model that outputs user state estimation data indicating the user's state, taking as input image-related data based on images of the user taken by an in-vehicle camera; The aforementioned simulation device is Images of the subject in the simulated environment are acquired from a camera that films the subject driving the vehicle in the simulated driving environment. Sensors that measure the state of a subject driving a vehicle in a simulated driving environment are used to acquire subject state data in a simulated environment. The learning data generation device is provided with the subject images in the simulated environment, the subject state data in the simulated environment, and the vehicle driving status data in the simulated environment, which shows the vehicle driving status in the simulated driving environment. The training data generation device is From the simulation device, the subject image in the simulated environment, the subject state data in the simulated environment, and the vehicle driving status data in the simulated environment are acquired. Based on the vehicle driving conditions data under the simulated environment, the image fluctuation state that is estimated to occur in the subject's image under the simulated environment during actual driving is calculated. Based on the image fluctuation state, processing is applied to the simulated environment subject image-related data relating to the simulated environment subject image to generate training image-related data. Training data is generated using the aforementioned training image-related data and the corresponding simulated environment subject state data. A system for generating training data.
Citation Information
Patent Citations
automated teller machine
JP7358956B2