Information processing device, information processing method, and information processing program
The information processing device uses image and time-series data to model occupant behavior and tire performance, addressing the challenge of identifying sensory evaluation bases for tire performance by deriving associations between occupant field of view and tire evaluation.
Patent Information
- Application Number
- JP2024042787
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-10-01
AI Technical Summary
Existing technologies struggle to identify the basis for sensory evaluations of tire performance, such as steering stability, as they rely on vehicle occupant feedback, making it difficult to determine what information contributes and at what timing.
An information processing device and method that utilizes a series of image information and time-series data to generate feature information and derive associations between occupant field of view and tire evaluation, employing variational autoencoders and recurrent neural networks to model occupant behavior and tire performance.
Enables estimation of tire evaluation by accounting for occupant inputs and outputs, including field of view changes, thereby improving the understanding of tire performance factors and their timing impacts.
Smart Images

Figure 2025143070000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] There is known an information processing device that uses time-series data from a sensor to calculate the importance of the output of a model that uses the time-series data as input (see, for example, Patent Document 1). This technology acquires time-series data related to a processing target, and then extracts time-series data of a predetermined time width from the acquired time-series data, resulting in multiple time-series data with different time widths, which are then set as groups. Then, for each of the groups, a score is calculated that indicates the importance of the output value of a model that uses the time-series data as input values.
[0003] Also known is a generation device that selects feature vectors for each time interval from a sequence of feature vectors for each frame to generate a representative feature vector (see, for example, Patent Document 2). This technology includes a means for selecting, for each time interval, feature vectors of multiple frames included in the time interval from the sequence of feature vectors for each frame, and a means for selecting, for each time interval, feature vectors of different dimensions from the feature vectors of different frames within the selected time interval, and generating a representative feature vector that represents the time interval. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2020-166442 [Patent Document 2] International Publication No. 2010 / 087125 Summary of the Invention [Problem to be solved by the invention]
[0005] Incidentally, the performance of a tire mounted on a vehicle, for example, steering stability performance, may be subject to a sensory evaluation by the vehicle occupant while the vehicle is in motion. However, although this sensory evaluation is obtained as a result of the vehicle occupant's driving, it has been difficult to identify the basis for the results of the sensory evaluation, for example, what information contributes and at what timing. Furthermore, with recent technological advances, models have been developed to predict the results of sensory evaluation. These models make it possible to predict the results of the sensory evaluation. However, even when these models are used, it has been difficult to identify the basis for the predicted results of the sensory evaluation, for example, what information contributes and at what timing. Therefore, there is room for improvement in evaluating the performance of tires mounted on a vehicle.
[0006] The present disclosure aims to provide an information processing device, an information processing method, and an information processing program that can estimate information related to tire evaluation by taking into account input and output to a vehicle occupant, including the field of view of the vehicle occupant, which changes as the vehicle travels. [Means for solving the problem]
[0007] A first aspect of the present disclosure is an acquisition unit that acquires a series of multiple image information related to the field of view of an occupant of the vehicle, the image information changing over time in response to the running of the vehicle on which the tire is mounted and indicating inputs and outputs to the occupant of the vehicle; a generation unit that uses an image feature model that receives the plurality of pieces of image information acquired by the acquisition unit as an input and outputs feature information related to the occupant's field of view to generate image feature information that is feature information related to the occupant's field of view in each of the plurality of pieces of image information; an extraction unit that uses the image feature information generated by the generation unit as an input and extracts a plurality of pieces of feature information that indicate features related to the occupant's field of view in each of a plurality of driving scenes obtained by dividing the driving of the vehicle into a plurality of time intervals, using a first model that is capable of extracting information that indicates features that change over time with respect to the input information; a derivation unit that uses a second model that receives the plurality of pieces of feature information extracted by the extraction unit as an input and receives evaluation information indicating an evaluation of the tire as an output, to derive association information indicating an association between the evaluation information and the field of view of the occupant in each of the plurality of driving scenes; The information processing device is provided with:
[0008] A second aspect is the information processing device of the first aspect, The extraction unit extracts, as the plurality of pieces of feature information, information indicating an internal state of the first model at a plurality of different predetermined timings corresponding to each of the plurality of driving scenes.
[0009] A third aspect is the information processing device of the first or second aspect, The derivation unit derives, as the related information, an importance indicating the degree of relevance of the occupant's field of view to the evaluation information, based on information indicating the internal state of the second model in each of the plurality of driving scenes.
[0010] A fourth aspect is an information processing device according to any one of the first to third aspects, The acquisition unit further acquiring a plurality of series of time-series information that changes over time in response to the running of a vehicle on which the tire is mounted and that indicates inputs and outputs to an occupant of the vehicle other than inputs and outputs related to the field of view of the occupant; The extraction unit using the plurality of sets of image information and the plurality of sets of time-series information acquired by the acquisition unit as inputs, and using the first model, further extracting information indicating features of input and output of the occupant in each of the plurality of driving scenes as the plurality of sets of feature information; The lead-out portion is Using the second model, which receives the plurality of pieces of feature information extracted by the extraction unit as input and outputs the evaluation information indicating an evaluation of the tire, information indicating the relationship between the evaluation information and input / output for the occupant in each of the plurality of driving scenes is derived as the related information.
[0011] A fifth aspect is an information processing device according to any one of the first to fourth aspects, the image feature model is a variational autoencoder model; the first model is a recurrent neural network model; The second model is a random forest model.
[0012] A sixth aspect is an information processing device according to any one of the fourth to fifth aspects, the time-series information includes arm information indicating inputs and outputs related to the arms of the occupant, waist information indicating inputs to the waist of the occupant, and head information indicating inputs to the head of the occupant; The arm information is data indicating at least one of the steering angle and the steering torque, the waist information is data indicating the center of gravity position of the occupant on the seat, and the head information is data indicating the head acceleration and angular velocity.
[0013] The seventh aspect is The computer acquiring a series of multiple image information related to the field of view of an occupant of the vehicle, the image information changing over time in response to the running of the vehicle on which the tire is mounted and indicating inputs and outputs to and from the occupant of the vehicle; generating image feature information, which is feature information regarding the visual field of the occupant in each of the plurality of pieces of image information, using an image feature model that receives the plurality of pieces of image information as an input and outputs feature information regarding the visual field of the occupant; extracting a plurality of pieces of feature information indicating features related to the occupant's field of view in each of a plurality of driving scenes obtained by dividing the driving of the vehicle into a plurality of time intervals using a first model that uses the generated image feature information as an input and is capable of extracting information indicating features that change over time with respect to the input information; deriving association information indicating an association between the evaluation information and the field of view of the occupant in each of the plurality of driving scenes using a second model that receives the extracted plurality of pieces of feature information as an input and outputs evaluation information indicating an evaluation of the tire; This is an information processing method for processing the above.
[0014] The eighth aspect is To the computer acquiring a series of multiple image information related to the field of view of an occupant of the vehicle, the image information changing over time in response to the running of the vehicle on which the tire is mounted and indicating inputs and outputs to and from the occupant of the vehicle; generating image feature information, which is feature information regarding the visual field of the occupant in each of the plurality of pieces of image information, using an image feature model that receives the plurality of pieces of image information as an input and outputs feature information regarding the visual field of the occupant; extracting a plurality of pieces of feature information indicating features related to the occupant's field of view in each of a plurality of driving scenes obtained by dividing the driving of the vehicle into a plurality of time intervals using a first model that uses the generated image feature information as an input and is capable of extracting information indicating features that change over time with respect to the input information; deriving association information indicating an association between the evaluation information and the field of view of the occupant in each of the plurality of driving scenes using a second model that receives the extracted plurality of pieces of feature information as an input and outputs evaluation information indicating an evaluation of the tire; It is an information processing program that processes the following:
[0015] The time-series information may include vehicle information indicating the behavior of the vehicle. [Effects of the Invention]
[0016] According to the present disclosure, it is possible to estimate information related to the evaluation of a tire by taking into account input and output to a vehicle occupant, including the field of view of the vehicle occupant, which changes as the vehicle travels. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a block diagram showing a schematic configuration of an estimation device according to an embodiment. [Figure 2] FIG. 1 is a conceptual diagram showing the correlation between an occupant and a vehicle. [Figure 3] FIG. 10 is a diagram illustrating an example of the arrangement of measurement sensors. [Figure 4] FIG. 1 is a diagram schematically illustrating an example of vehicle travel. [Figure 5] 10A and 10B are diagrams illustrating an example of a change in field of view and a change in steering angle when changing lanes. [Figure 6] FIG. 1 is a diagram illustrating a conceptual configuration of an estimation model. [Figure 7] FIG. 2 is a diagram illustrating a conceptual configuration of an image feature model. [Figure 8] FIG. 2 is a diagram illustrating a conceptual configuration of a first model. [Figure 9] FIG. 10 is a diagram illustrating a conceptual configuration of a second model. [Figure 10] FIG. 2 is a diagram illustrating an example of an electrical configuration of the estimation device. [Figure 11] 10 is a flowchart illustrating an example of the flow of an estimation process. DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, embodiments for realizing the technology of the present disclosure will be described in detail with reference to the drawings. In addition, components and processes that perform the same actions and functions are given the same reference numerals throughout the drawings, and duplicated explanations may be omitted as appropriate. Furthermore, the present disclosure is not limited to the following embodiments, and can be implemented with appropriate modifications within the scope of the purpose of the present disclosure.
[0019] In this disclosure, a tire is a concept that refers to an elastically deformable member attached to a vehicle, which is an example of a moving body, and includes a member that can be attached to or detached from the vehicle. Furthermore, the steering stability performance of a tire is an example of tire performance and is also an example of a characteristic related to the vehicle's maneuverability. Tire performance is also a concept that includes characteristics related to the performance of the vehicle to which the tire is attached. In this disclosure, a plurality of series of image information is a concept that refers to information on a group of images showing the occupant's field of view or changes in field of view that change over time as the vehicle travels. Time-series image information can be applied to the plurality of series of image information, and an example of a group of consecutive images captured continuously by a camera or the like can be cited as an example. In this disclosure, feature information is a concept that refers to features related to the occupant's field of view that are related to the tire's evaluation of the tire's running, among the occupant's input and output. In this disclosure, related information is a concept that refers to a relationship between the tire evaluation and the occupant's field of view or changes in field of view. The related information is a concept that refers to the importance of the feature information, and information indicating the degree of association between the tire evaluation and the occupant's input and output can be applied as the importance.
[0020] The feature information may include features of the occupant's input / output other than the occupant's field of view. The related information may also include associations with the occupant's input / output other than the occupant's field of view or changes in field of view. Time-series information representing the occupant's behavior may be applied to the information indicating the input / output to the occupant. Examples of the information indicating the input / output include data indicating the steering angle and steering torque as information indicating the input / output related to the occupant's arms, data indicating the center of gravity position on the seat as information indicating the input to the occupant's waist, and head acceleration and angular velocity as information indicating the input to the occupant's head.
[0021] In addition, in this disclosure, a vehicle is a concept that includes a moving body such as an automobile equipped with tires.
[0022] <Estimation device> 1 is a diagram showing an example of the configuration of an estimation device 1 capable of executing an estimation process for estimating information related to tire evaluation. The estimation device 1 includes an estimation unit 5. The estimation device 1 can be realized by a computer including a CPU as an execution device that executes the processes described below.
[0023] The estimation unit 5 receives, as input data, information indicating inputs and outputs to and from the occupant, including at least one of the occupant's field of view and a change in field of view, which is the behavior of the vehicle occupant that changes depending on the driving state 3 indicating the driving state of the vehicle 2. The information indicating inputs and outputs to and from the occupant includes information indicating the behavior of the occupant due to the movement of multiple parts of the occupant and the exchange of forces. In this embodiment, image information indicating the occupant's field of view, etc. will be mainly described. The image information indicating the occupant's field of view, etc., is a series of multiple image information in time series, and measurement data 4 measured by a measurement sensor that measures the behavior of the occupant is applied. The measurement data 4, which is image information, is acquired by a camera 24 (FIG. 3), which is an example of a measurement sensor that functions as the detection unit 118 (FIG. 10). Note that information indicating inputs and outputs to and from the occupant other than the image information indicating the occupant's field of view, etc., is a series of time series information, and time series measurement data acquired by measurement sensors 21, 22, and 23 (FIG. 3), which function as the detection unit 118 (FIG. 10), as another example of the measurement sensor, can be applied.
[0024] The estimation unit 5 estimates information related to the evaluation of a tire during vehicle travel using an estimation model 51 based on a series of input image information (i.e., a plurality of time-series image data) through an estimation process described below, and outputs the information as output data. In this embodiment, a case will be described in which evaluation data serving as first output data indicating the evaluation of a tire and associated data serving as second output data relating to the tire evaluation and the field of view of an occupant, etc., are applied as the information related to the evaluation of a tire during vehicle travel. That is, the estimation unit 5 can estimate evaluation data 6 indicating the evaluation of a tire during vehicle travel using time-series measurement data 4 (i.e., time-series image data), and can also estimate associated data 7 related to the tire evaluation. The estimation model 51 includes a first model 51A and a second model 51B. The estimation model 51 including the first model 51A and the second model 51B will be described later.
[0025] Specifically, the estimation unit 5 is a functional unit that estimates information related to the evaluation of a tire. The estimation unit 5 is connected to a vehicle 2 equipped with a measurement sensor (FIG. 3), and a series of image information (i.e., time-series image data) is input as measurement data 4. The estimation unit 5 estimates and outputs output data related to the evaluation of a tire using an estimation model 51 based on the input measurement data 4, which is a plurality of time-series image data. The field of view, etc., of the occupant of the vehicle 2 changes depending on the running state 3 indicating the running of the vehicle 2. The estimation unit 5 outputs, as estimation results using the estimation model 51, evaluation data 6 indicating the evaluation of the tire and associated data 7 relating the occupant's field of view, etc. and the tire evaluation as first output data.
[0026] The estimation unit 5 uses an estimation model 51 to estimate information related to the evaluation of tires during vehicle travel, including a series of input time-series information (i.e., time-series measurement data measured by the measurement sensor 21, etc.), through an estimation process described below, and can output the information as output data (details will be described later).
[0027] Here, the correlation between the occupant and the vehicle will be explained. When evaluating tire performance that affects steering stability and the like, a sensory evaluation by a vehicle occupant may be performed. The sensory evaluation by the occupant is assumed to be based (e.g., dominant) on the amount felt by the occupant (hereinafter referred to as sensory amount). For this reason, in this embodiment, image data related to the occupant's field of view is used as the series of image information (i.e., multiple pieces of image data in time series) input to the estimation unit 5. Note that the information input to the estimation unit 5 may include, in addition to image data, time-series measurement data measured by the measurement sensor 21 or the like as a series of time-series information.
[0028] Fig. 2 is a conceptual diagram showing the correlation between an occupant and the vehicle 2. Fig. 2 shows an example of the correlation regarding input and output to the occupant, including the occupant's field of view, which is the occupant's behavior when steering the vehicle 2.
[0029] When the occupant drives the vehicle 2, the occupant visually checks, for example, the outside of the vehicle, considering the driving conditions assuming the target position, direction, speed, etc. That is, in driving situations such as turning right or left or changing lanes as the vehicle 2 drives, the occupant visually checks the surroundings of the vehicle 2 while steering or performing other operations such as steering. This visual state can be measured as a sensory amount related to the occupant's field of view due to the occupant's line of sight, changes in field of view, etc.
[0030] Furthermore, the occupant steers the vehicle 2 by operating the steering wheel or the like while assuming a target driving state. The steering is output from the occupant by operating the steering wheel or the like and is reflected in the steering angle. In the vehicle 2, the behavior of the vehicle 2 is reflected according to the steering angle by the occupant, and in response, a steering torque is generated in the steering wheel or the like and input to the occupant. These inputs and outputs to the occupant can be measured as sensory quantities in the occupant's arms.
[0031] Furthermore, the vehicle behavior changes according to the steering by the occupant, which is reflected in the actual driving conditions and affects the occupant's behavior. In the example shown in the figure, the occupant's head and waist are used as examples of parts of the occupant's behavior that are affected. Specifically, the occupant's posture and position change occur in the seat depending on the actual driving conditions, and this is input to the occupant. This occupant's posture and position change can be measured as the occupant's center of gravity position and / or the change in center of gravity position as a sensory quantity of the occupant's waist. Note that the occupant's center of gravity position may be measured as the moving speed and / or the moving acceleration of the center of gravity position.
[0032] Furthermore, the behavior of the occupant's head changes depending on the actual driving conditions. For example, a change in the position of the occupant's head occurs and is input to the occupant. In this embodiment, the acceleration and angular velocity of the occupant's head are applied as the behavior of the occupant's head. The acceleration and angular velocity of the occupant's head, which are input to the occupant, can be measured as sensory quantities of the occupant's head. At least one of the acceleration and angular velocity of the occupant's head can be applied as the behavior of the occupant's head. Note that the behavior of the occupant's head may be applied as a measurement value obtained by measuring the position and / or position change of the occupant's head. Note that examples of the behavior of the occupant's head may include the above-mentioned line of sight and field of view of the occupant.
[0033] The target driving state assumed by the occupant can be considered as a driving condition. That is, the target driving state can be applied as a driving condition indicating that the steering should be performed assuming the driving state desired by the occupant.
[0034] The occupant performs a sensory evaluation according to the occupant's behavior, which is sensed both actively and passively, while the vehicle 2 is traveling, and provides the result as a sensory evaluation value.
[0035] Next, the measurement sensors will be described. The measurement sensors can detect at least a series of image information as measurement data 4, which is a series of image data. The measurement sensors can also acquire time-series measurement data 4, which is a series of time-series information. The measurement sensors are arranged in accordance with the type of occupant behavior to be detected, and measure the behavior of the occupant in the vehicle 2.
[0036] 3 is a diagram showing an example of the arrangement of measurement sensors. In this embodiment, as an example of a measurement sensor, at least a camera 24 that detects the field of view of an occupant is provided. In addition, measurement sensors 21, 22, and 23 corresponding to the types of detection of the behavior of other occupants are provided.
[0037] The camera 24 is a sensor capable of capturing an image corresponding to the occupant's field of view, and obtains information (images) indicating input and output related to the occupant's eyeballs in the field of view as measurement data. In this embodiment, the camera 24 is attached to a head-mounted member 25 such as glasses worn by the occupant, and is configured to capture an image in front of the occupant. Note that the camera 24 can be configured so that when the line of sight relative to the occupant's head changes, the imaging axis changes in accordance with the change in the line of sight.
[0038] The measurement sensor 21 is a sensor that measures the behavior of the occupant's arm, and the measurement sensor 21 can obtain measurement data that indicates information about input and output related to the occupant's arm. Data that indicates a steering angle and / or steering torque can be applied to the measurement data measured by the measurement sensor 21. For example, the measurement sensor 21 may be a sensor that measures a steering angle and / or steering torque.
[0039] The measurement sensor 22 is a sensor that measures the behavior of the occupant's waist, and the measurement sensor 22 can obtain measurement data that indicates information about input related to the occupant's waist. The measurement data measured by the measurement sensor 22 can be data that indicates the load and center of gravity of the occupant on the seat. For example, the measurement sensor 22 can be a sensor that measures the load and center of gravity of the occupant on the seat. One example of such a sensor is a six-component force gauge, and for example, multiple (for example, four) six-component force gauges can be installed on the seat or the like, and the load and center of gravity can be determined from the values obtained.
[0040] The measurement sensor 23 is a sensor that measures the behavior of the occupant's head, and the measurement sensor 23 can obtain measurement data that indicates information about inputs related to the occupant's head. The measurement data measured by the measurement sensor 23 can be data that indicates the acceleration and / or angular velocity of the occupant's head. For example, the measurement sensor 23 can be a sensor that measures the acceleration and / or angular velocity of the occupant's head. One example of such a sensor is a glasses-type device that can be used as a motion sensor, and the acceleration and angular velocity of the motion sensor can be measured.
[0041] The measurement sensors 21, 22, and 23 may be dedicated measurement sensors such as position detection sensors, motion sensors, and torque sensors, or may measure from images of the occupant. The measurement sensors are not limited to directly detecting the behavior of the occupant, but may also acquire data detectable by the vehicle 2, such as steering angle and steering torque, and use the data as measurement data. In other words, the measurement data 4 as time-series information can include vehicle information indicating the behavior of the vehicle 2.
[0042] Next, the traveling state of the vehicle 2 will be described. Incidentally, an occupant performs a sensory evaluation of the running state of the vehicle 2, which changes from moment to moment, depending on the behavior of the vehicle 2 as it runs. However, it has been difficult to identify factors that contribute to the result of the sensory evaluation. Therefore, in this embodiment, a plurality of different timings related to the result of the occupant's sensory evaluation (e.g., the basis for the judgment) in the running state of the vehicle 2 are considered.
[0043] Fig. 4 is a diagram showing a schematic example of the driving (driving state) of the vehicle 2. The example in the figure shows the driving state when changing lanes from one lane (the right lane in the figure) to the other lane (the left lane in the figure) on a two-lane road. Fig. 5 is a diagram showing the change in steering angle and the change in field of view when changing lanes as shown in Fig. 4.
[0044] Specifically, the occupant changes lanes by rotating the vehicle's steering wheel (not shown) clockwise and counterclockwise. During the lane change, the occupant's field of view changes from moment to moment. Furthermore, a lane change can be classified into three driving scenes, for example, an early stage, a middle stage, and a late stage. A driving scene is a driving state that extracts at least a portion of the driving of the vehicle 2. The early stage driving scene is a driving state at the stage when the lane change is initiated. The middle stage driving scene is a driving state at the stage when the vehicle 2 moves due to the lane change. The late stage driving scene is a driving state at the stage when the posture of the vehicle 2 is adjusted in the later stage of the lane change. It is preferable that the early stage driving scene and the middle stage driving scene include inflection points of maximum or minimum values in the steering angle characteristics.
[0045] In this embodiment, a time interval up to a predetermined timing is defined as a section corresponding to a driving scene. Specifically, the time interval up to timing ts is defined as a section Ts corresponding to an early driving scene, the time interval up to timing tm is defined as a section Tm corresponding to a middle driving scene, and the time interval up to timing te is defined as a section Te corresponding to a later driving scene. Hereinafter, the sections Ts, Tm, and Te may be referred to as driving scenes Ts, Tm, and Te. Note that the above-mentioned predetermined timings may be determined, for example, from experimental values or statistical values, as timings including inflection points of the maximum and minimum values in the steering angle characteristics described above. These timings may be changed using behavior values of the vehicle 2, such as vehicle speed, as parameters, assuming that they may vary depending on the environment during driving of the vehicle 2, for example, vehicle speed.
[0046] In this way, the driving state of the vehicle 2 is a part of a series of driving situations in which the vehicle travels, and is a state in which different driving situations continue in a time series. Therefore, as will be described in detail later, by using the measurement data 4 in each driving situation, it becomes possible to identify factors that contribute to the results of the sensory evaluation in each driving situation.
[0047] The driving scenes are not limited to the three driving scenes of the early, middle, and late stages described above, but may be, for example, two or more driving scenes. Furthermore, the multiple driving scenes may be arranged so that adjacent driving scenes partially overlap, or adjacent driving scenes may be separated. Furthermore, the above-mentioned timings ts, tm, and te may be timings that represent each driving scene, and are not limited to the above-mentioned timings ts, tm, and te. The timings ts, tm, and te are examples of multiple different timings in the technology of the present disclosure.
[0048] <Estimation model> Next, the estimation model 51 will be described. The estimation model 51 is a machine learning model that models an algorithm related to the sensory evaluation of an occupant riding in the vehicle 2. That is, the estimation model 51 is a machine learning model that estimates output data (evaluation data 6 and related data 7) that indicates information related to the evaluation of the tires in the traveling state of the vehicle from the behavior (for example, field of view images) of the occupant of the vehicle 2 indicated by the input time-series measurement data 4. Note that in this embodiment, a machine learning model that estimates output data including the behavior of the occupant that indicates input and output to and from the occupant other than the field of view images is applied to the estimation model 51.
[0049] FIG. 6 is a diagram showing the conceptual configuration of the estimation model 51. The estimation model 51 receives as input at least time-series measurement data 4 (visual field images) relating to the visual field as occupant behavior, and outputs evaluation data 6 indicating a sensory evaluation value as first output data. The estimation model 51 can also output related data 7 as second output data.
[0050] 6 shows an example of an estimation model 51 in which time-series measurement data 4 of the occupant's arms, waist, and head are also input and evaluation data 6 is output. The estimation model 51 can also output related data 7 as second output data.
[0051] The evaluation data 6 is an example of evaluation information of the present disclosure, and the related data 7 is an example of related information of the present disclosure. The time-series measurement data 4 (field of view image) relating to the field of view as the occupant's behavior is an example of a series of multiple image information of the present disclosure. Furthermore, the time-series measurement data 4 of the occupant's arms, waist, and head are examples of arm information, waist information, and head information of the present disclosure.
[0052] The estimation model 51 takes into account the time-series measurement data 4. In this embodiment, a neural network is applied to the estimation model 51.
[0053] The estimation model 51 uses the visual field image of the occupant, and therefore includes an image feature model 51V that extracts features of the visual field image.
[0054] FIG. 7 is a diagram showing the conceptual configuration of a VAE (Variational Autoencoder model) applied to the image feature model 51V.
[0055] Since VAE is a well-known technology, detailed description will be omitted. However, VAE is a generative model formed including a trainable decoder and an encoder configured by a neural network. The VAE is a generative model that learns pre-prepared learning image data as training data and generates an image that approximates the input image. The encoder receives an input image x as input and outputs a latent variable z, and the decoder receives the latent variable z output from the encoder as input and outputs an approximate image x that approximates the input image x. The distribution of the latent variable z in the VAE reflects the characteristics of the input image. In this embodiment, to extract the characteristics of the field of view image, the value of the latent variable z in a fully trained VAE is extracted as a feature. That is, the image feature model 51V generates and outputs image feature data (image feature information), which is a feature indicating the characteristics of the input field of view image. The feature amount indicating the feature of the field of view image is an example of image feature information of the present disclosure.
[0056] The VAE is a generative model that generates an image that is similar to an input image. However, the image feature model 51V in this embodiment does not use the function of generating an approximate image. Specifically, the image feature data (image feature information) generated by the VAE is passed to the first model 51A as an output.
[0057] Incidentally, for time-series continuous data such as measurement data of time-series visual field images, past history is important. Therefore, the estimation model 51 of this embodiment includes a recurrent neural network model using a recurrent neural network (RNN) as a model that enables the use of past history of data. Hereinafter, a case where an ESN (Echo State Network), which is an example of an RNN, is used will be described. The ESN is also an example of a neural network called reservoir computing.
[0058] As shown in FIG. 6 , the estimation model 51 includes a first model 51A and a second model 51B. The first model 51A and the second model 51B are linked to form the estimation model 51. A model capable of extracting information indicating characteristics that change over time with respect to input information can be applied to the first model. In this embodiment, an ESN is applied as an example of the first model. Note that, as will be described in detail later, the first model 51A linked to the second model 51B is part of the ESN. A well-known recurrent network may also be applied to the first model. In this case, the first model may be formed so that certain information is input and an evaluation corresponding to the certain information, i.e., evaluation information indicating an evaluation of the tire, is output.
[0059] 8 is a diagram showing the conceptual configuration of an ESN applied to the first model 51A. The ESN is expressed as a collection of information on the weights (strengths) of connections between nodes (neurons) that make up a neural network, and the internal state of the ESN can be obtained as state data indicating the value of each node (neuron). Hereinafter, a node may be referred to as a neuron.
[0060] The ESN applicable to the first model 51A can be trained to receive image feature data and measurement data 4 indicating other inputs and outputs as inputs, and to output values indicating tire evaluations, which are sensory evaluation values. That is, the ESN applicable to the first model 51A is a model that models the sensory evaluations of occupants. In this embodiment, three indicators, good (appropriate), average (medium), and poor (unsuitable), are used as examples of output values indicating tire evaluations.
[0061] The ESN can be applied as a model that models the occupant's sensory evaluation, which takes the measurement data 4 as input and outputs a sensory evaluation value that indicates the tire evaluation. The first model 51A uses a part of the ESN, as will be described later. In this embodiment, three indicators, good (appropriate), normal (average), and poor (unsuitable), are used as examples of output values that indicate the tire evaluation.
[0062] Specifically, the ESN applied to the first model 51A is composed of an input layer 510, an output layer 514, and a hidden layer called a reservoir layer 512. The reservoir layer 512 includes randomly connected neurons. As is known in reservoir computing, the weights Wres of the reservoir layer 512 and the connection weights Wn between the input layer 510 and the reservoir layer 512 are not subject to learning but are randomly initialized in advance. This random structure allows for the propagation of past information when appropriate parameters such as the number of neurons are set, thereby enabling the representation of long-term dependencies. Note that, as will be described later, in this embodiment, the ESN applied to the first model 51A does not use the output layer 514 but uses the neuron configuration up to the reservoir layer 512 (specifically, only neuron values indicating the internal state).
[0063] The above ESN (first model 51A) can be expressed by the following equations (1) and (2). x(t) = (1-α) x(t-1) + α f(Win u(t) + Wres x(t-1)) -(1) y(t) = fout(Wout·x(t)) -(2)
[0064] Here, u(t) is the image feature data or measurement data 4, which is the input data at time t. In this embodiment, each measurement value, which is the input data 4, is treated as a scalar. y(t) is the output data, expressed by the function fout. If the number of neurons in the reservoir layer 512 is N, x(t) is the state data indicating the state quantity held by the neurons in the reservoir layer 512 at time t, and is an N-dimensional vector. W is the connection weight between the input layer 510 and the reservoir layer 512, and is an N-dimensional vector. W is the weight of the reservoir layer 512, and is an NxN-dimensional matrix representing the connection quantity between neurons within the reservoir layer 512. α is called the leaky integrator (LI) and controls the rate at which internal parameters are successively updated. In other words, α controls the forgetting speed of the neuron. Therefore, the ESN maintains its internal state at each time step, and the internal state evolves over time according to equation (1), which includes the above function f.
[0065] In this embodiment, an ESN is constructed for each type of measurement data 4 indicating image feature data and other inputs and outputs, and a first model 51A is formed.
[0066] The model applied to the first model 51A described above is not limited to the ESN, but other RNNs that are regression models that can utilize past history based on the time-series measurement data 4 may also be applied.
[0067] 6, the estimation model 51 includes a memory 52 that stores state data at any timing for the ESN, which is the first model 51A. That is, the estimation model 51 can extract the internal state of the ESN at any timing and store it in the memory 52. In this embodiment, the internal state of the ESN is extracted at each of the timings ts, tm, and te that define the above-mentioned driving scenes, for example, the early driving scene, the middle driving scene, and the late driving scene of a lane change, and stored in the memory 52. Therefore, the internal state of the ESN is extracted for each type of measurement data 4 and stored in the memory 52.
[0068] The state data indicating the internal state of the ESN at these timings corresponds to the feature quantities indicating the characteristics of the image feature data in each driving scene and the measurement data 4 indicating other inputs and outputs (view image information, which is a series of image information, and time-series information). The feature quantities, which are state data, correspond to information that quantifies the relationship between the image feature data and the measurement data 4 indicating other inputs and outputs, and the evaluation value for the tire, which is a sensory evaluation value. The feature amount is an example of feature information of the present disclosure.
[0069] Incidentally, when estimating an output using an ESN, the output layer 514 and its weight Wout are generally used. However, the first model 51A to which the ESN is applied in this embodiment does not use the output layer 514 and its weight Wout. Specifically, state data extracted at a predetermined timing and stored in the memory 52 is passed as output to the second model 51B. That is, in this embodiment, the output layer 514 of the ESN and its weight Wout are not used, and the internal state of the neurons up to the reservoir layer 512 is used.
[0070] Next, the second model 51B of the estimation model 51 will be described.
[0071] The second model 51B receives the state data of the first model 51A as input and outputs evaluation data 6, which is the occupant's sensory evaluation of the tire. The second model 51B is also a model that models the occupant's sensory evaluation. In this embodiment, three indicators, good (appropriate), normal (average), and poor (inappropriate), are used as examples of the evaluation data 6.
[0072] 9 is a diagram showing a conceptual configuration of the second model 51 B. In this embodiment, a case will be described in which a random forest model is applied to the first model 51 B as an example of a classification model.
[0073] The random forest model is a well-known technology and will not be described in detail here, but it is an ensemble learning algorithm that uses bagging and decision trees as learning devices. In this random forest model, multiple decision trees trained in parallel make predictions for an input, and the results of majority voting or average calculations are used as the output.
[0074] For example, the second model 51B includes functional units that perform grouping, modeling, provisional estimation, and final estimation.
[0075] In the second model 51B, first, data is randomly sampled from the original data (state data) using bootstrap (a process called sampling with replacement) to generate a predetermined number (here, M) of data groups (grouping). Next, a decision tree 516 is generated for each of the M groups (modeling). The decision tree 516 is a model that selects an output using conditional branching (for example, if-then-). Next, an estimate is made using the decision tree for each of the M groups (provisional estimation). Then, a majority vote (or average) of the M groups is taken to make a final prediction of the output data (evaluation value) (final prediction).
[0076] In the second model 51B employing this random forest model, for example, each value of N pieces of original data (state data) is used as a feature, and some of the feature values are randomly selected and used to divide the decision tree 516 in the modeling. Then, by slightly modifying the original data by bagging and slightly modifying the feature values by randomization, multiple decision trees 516 with low correlation with each other are generated. The generation of the decision tree 516 involves repeated division using a conditional branching division method and a predetermined index (e.g., error or Gini impurity) that evaluates the division point. This index is a value that indicates the appropriateness of the conditional branching at the division point. Specifically, the smaller the index, the more appropriate the conditional branching. Therefore, the smaller the index, the higher the importance of the feature. Therefore, the importance of the feature can be derived by calculating the degree of decrease in index (e.g., sum) for each feature value in the decision tree 516 upon division. In other words, it is possible to derive the importance of the feature, which indicates the degree of relevance of the input / output to the occupant, including at least the visual field image, to the evaluation data. The importance of a feature is an example of related information of the present disclosure.
[0077] The second model 51B of this embodiment includes a derivation unit 518 that derives the importance of the above-mentioned feature. Specifically, the value of the original data (state data) is used as the feature, and the degree of decrease in the index for each feature in the decision tree 516 is calculated to derive the importance of the feature. The derivation unit 518 is an example of the derivation unit of the present disclosure.
[0078] In this way, a random forest model is applied to the second model 51B, which receives as input state data extracted from the internal state of the first model 51A (ESN). After image feature data of a series of field-of-view images and a series of time-series data have been input to the first model 51A (ESN), the second model 51B constructs a random forest model which receives as input state data indicating the internal state of the ESN extracted corresponding to each driving scene and outputs evaluation data 6 indicating a sensory evaluation value. The second model 51B also includes a derivation unit 518 which derives the importance of the above-mentioned feature amount.
[0079] Therefore, by using state data at each timing indicating the driving scene at a predetermined time (e.g., the early, middle, and late stages of steering) during a driving state (e.g., when changing lanes), it becomes possible to output information indicating which data (e.g., which field of view image) at which timing is important to the results of the sensory evaluation.
[0080] Note that a deep forest model may be applied to the first model 51B instead of the random forest model described above.
[0081] As described above, the first model 51A of the estimation model 51 is constructed to take as input at least a field of view image that changes depending on the driving state of the vehicle 2, and to output state data indicating the internal state value of the first model 51A (ESN) at a predetermined timing corresponding to the driving scene.
[0082] Furthermore, the second model 51B of the estimation model 51 is constructed such that state data indicating the internal state values of the first model 51A (ESN) are input and output data is output. The input state data indicates the internal state values of the first model 51A (ESN) at a predetermined timing corresponding to a driving scene that progresses at least by a field of view image that changes depending on the driving state of the vehicle 2. The output data of the second model 51B includes evaluation data 6 and related data 7. The evaluation data 6 is data corresponding to the result of the occupant's sensory evaluation (i.e., the sensory evaluation value). The related data 7 is data corresponding to a driving scene (e.g., timing) that contributes to the result of the occupant's sensory evaluation (e.g., the basis for judgment) and a feature value (i.e., the importance of the feature value) that is considered important in that driving scene. In other words, the related data 7 is related to the result of the occupant's sensory evaluation and corresponds to information indicating the feature of the field of view image that is considered important for each driving scene or input / output to the occupant.
[0083] As described above, the estimation model 51, i.e., the first model 51A and the second model 51B, are stored in a memory (not shown) of the estimation device 1. By using the estimation model 51, the estimation device 1 can estimate a result corresponding to the occupant's sensory evaluation regarding the traveling of the vehicle 2, and can also estimate information indicating a field of view image, its features, and input / output to the occupant that is important for each traveling scene and that is related to the result of the occupant's sensory evaluation.
[0084] <Configuration of the estimation device> Next, an example of a specific configuration of the above-mentioned estimation device 1 will be further described. Fig. 10 is a diagram showing an example of the electrical configuration of the estimation device 1. The estimation device 1 shown in Fig. 9 is configured to include a computer as an execution device that executes processes to realize the various functions described above. The estimation device 1 described above can be realized by causing the computer to execute a program representing each of the functions described above.
[0085] The computer functioning as the estimation device 1 includes a computer main unit 100. The computer main unit 100 includes a CPU 102, a RAM 104 such as a volatile memory, a ROM 106, an auxiliary storage device 108 such as a hard disk drive (HDD), and an input / output interface (I / O) 110. The CPU 102, RAM 104, ROM 106, auxiliary storage device 108, and input / output I / O 110 are connected via a bus 112 so that data and commands can be exchanged among them. The input / output I / O 110 is also connected to a communication unit 114 for communicating with external devices, an operation display unit 116 such as a display and keyboard, and a detection unit 118 including the measurement sensor described above. The detection unit 118 functions to acquire image data indicating a visual field image of a passenger aboard the vehicle 2 and measurement data 4 via a wireless or wired connection. Note that the data including the image data may be acquired via the communication unit 114.
[0086] The auxiliary storage device 108 stores an estimation program 108P for causing the computer main body 100 to function as an example of an information processing device of the present disclosure. The CPU 102 reads the estimation program 108P from the auxiliary storage device 108, loads it into the RAM 104, and executes processing. As a result, the computer main body 100 that has executed the estimation program 108P operates as the estimation device 1. The estimation program 108P is an example of an information processing program of the present disclosure.
[0087] The auxiliary storage device 108 stores the estimation model 51 including the first model 51A and the second model 51B, and data 108D including various data. The estimation program 108P may be provided by a recording medium such as a CD-ROM.
[0088] <Estimation process> Next, the estimation process in the estimation device 1 implemented by a computer will be further described. Fig. 11 is a diagram showing an example of the flow of estimation processing by the estimation program 108P executed by the computer main body 100. The estimation processing shown in Fig. 11 is executed by the CPU 102 when the computer main body 100 is powered on. The CPU 102 reads the estimation program 108P from the auxiliary storage device 108, loads it into the RAM 104, and executes the processing. The estimation process shown in FIG. 11 includes a process corresponding to a processing flow that can realize the information processing method of the present disclosure.
[0089] First, the CPU 102 reads the estimation model 51 from the estimation model 108M in the auxiliary storage device 108 and expands it in the RAM 104, thereby acquiring the estimation model 51 (step S200). Specifically, by expanding the estimation model 51 (see FIG. 7) in the RAM 104, the estimation model 51 including the image feature model 51V, the first model 51A, and the second model 51B is constructed.
[0090] Next, CPU 102 acquires, via detection unit 118, measurement data 4 that is continuous in time series and includes image data showing the occupant's field of view (step S201). The acquired measurement data includes at least image data showing the occupant's field of view. Next, CPU 102 inputs the image data showing the occupant's field of view to image feature model 51V of estimation model 51 (step S202). CPU 102 generates image feature data (image feature information), which is latent data, as a feature amount showing the characteristics of the occupant's field of view image using image feature model 51V (step S203). Then, CPU 102 sequentially inputs the generated image feature data and measurement data 4, which is another input / output, to first model 51A of estimation model 51 (step S204). Thus, the image feature data and other time-series measurement data 4 are applied to ESN, which is first model 51A, and the neuron value (neuron value) is updated. The camera 24, the measurement sensors 21, 22, and 23, and the detection unit 118 are examples of the acquisition unit in the technology of the present disclosure. The process of step S201 is an example of the process by the acquisition unit in the technology of the present disclosure. The process of step S203 is an example of the process by the generation unit in the technology of the present disclosure.
[0091] Next, CPU 102 stores state data indicating the internal state of first model 51A at a predetermined timing (step S206). The predetermined timing is a timing determined in advance corresponding to the driving scene, as described above. At the timing determined in advance corresponding to these driving scenes, state data indicating the internal state of first model 51A, i.e., the neuron values of ESN, is extracted. In addition, state data is extracted for each type of measurement data 4. The process of step S206 is an example of a process performed by the extraction unit of the present disclosure.
[0092] Next, CPU 102 inputs state data indicating the internal state of first model 51A stored in memory to second model 51B (step S208). For example, when a series of image feature data and other measurement data of the vehicle 2's travel, i.e., from the initial travel scene to the later travel scene, has been input to first model 51A (ESN), state data is input to second model 51B, i.e., the random forest model. This state data is the internal state of first model 51A (ESN) for each type of field of view image and measurement data 4, and is data extracted at timing corresponding to each travel scene.
[0093] When these data have been input into the first model 51A (ESN), a second model 51B (random forest model) is constructed, which takes state data indicating the internal state of the first model 51A as input and outputs evaluation data 6 indicating the sensory evaluation value.
[0094] Next, the CPU 102 estimates the evaluation data 6 (step S210). Specifically, the CPU 102 obtains the output of the constructed second model 51B (random forest model) to obtain the evaluation data 6 of the estimation result.
[0095] Next, CPU 102 estimates related data 7 (step S212). Specifically, derivation unit 518 calculates the index for each feature amount of decision tree 516 of second model 51B as described above, thereby deriving the importance of the feature amount. For example, in second model 51B, the degree of decrease in the index for the feature amount is calculated, that is, the degree of decrease in the index for evaluating the conditional branch of decision tree 516 is calculated (for example, the sum).
[0096] The derived related data 7 is the importance of the feature, and corresponds to the importance of the type of input / output to the occupant that contributes to the evaluation data 6 indicating the sensory evaluation value. Specifically, from the perspective of the occupant's field of view, it corresponds to the importance of the feature of the field of view image showing the occupant's field of view in each driving scene included in the vehicle's driving. Therefore, in each driving scene included in the vehicle's driving, the degree of contribution to the evaluation data 6 increases as the importance of the feature for at least the occupant's field of view image increases, corresponding to the type of input / output to the occupant. Therefore, it is possible to identify a feature whose importance in a driving scene is maximum or exceeds a predetermined threshold, i.e., the occupant's field of view image, as information that contributes to the evaluation data 6 in the corresponding driving scene.
[0097] In addition, when inputs and outputs to the occupant other than the visual field are included, it is also possible to identify the types of inputs and outputs to the occupant as information that contributes to the evaluation data 6 in the relevant driving scene. In this case, the importance of the features of the visual field image that indicates the occupant's visual field may be identified in combination with the types of other inputs and outputs. The process of step S210 and the process of step S212 are examples of processes performed by the derivation unit of the present disclosure.
[0098] Next, CPU 102 outputs the output data (evaluation data 6 and related data 7) of the estimation results obtained in steps S210 and S212 (step S214), and ends this processing routine. The output data may be displayed on operation display unit 116, or may be output to an external device via communication unit 114.
[0099] In this way, the estimation device 1 can estimate information related to the result of the occupant's sensory evaluation (for example, the basis for the judgment) regarding vehicle driving. Specifically, the estimation device 1 uses the internal state (neuron information) of the model at the timing indicating the driving scene regarding vehicle driving such as lane changes, and thereby can output information indicating which information, for example, which feature of the field of view image (and also the type of other input / output to the occupant) is important at what timing for the result of the sensory evaluation.
[0100] In the above, the occupant's arms, waist, and head have been described as examples of occupant behavior and as examples of parts that affect input and output to the occupant other than the occupant's field of view. However, the present disclosure is not limited to this. For example, the occupant's legs or other parts may be included. Furthermore, each part may be subdivided, or multiple parts that combine multiple of these parts may be used as application parts.
[0101] In the above, a case where information (image data and measurement data) related to the occupant's field of view and other inputs and outputs to the occupant is used as an example of occupant behavior has been described, but the technology of the present disclosure is not limited to this. Specifically, the above-described estimation process may be performed using only the occupant's field of view image. In this case, the above-described measurement sensors 21, 22, and 23 (FIG. 3) are not required, and it is also not necessary to input other measurement data to the first model 51A in the estimation model 51 (FIG. 7).
[0102] Furthermore, the technical scope of the present disclosure is not limited to the scope described in the above embodiments. Various modifications or improvements can be made to the above embodiments without departing from the gist of the present disclosure, and such modifications or improvements are also included in the technical scope of the present disclosure.
[0103] In addition, in the above embodiment, the estimation process is described as being realized by a software configuration using a flowchart, but this is not limited to this, and for example, each process may be realized by a hardware configuration.
[0104] Furthermore, a part of the estimation device, for example, a neural network such as an estimation model, may be configured as a hardware circuit.
[0105] Although the technology of the present disclosure has been described above using embodiments, the technical scope of the present disclosure is not limited to the scope described in the above embodiments. Various modifications and improvements can be made to the above embodiments without departing from the gist of the disclosure, and such modifications and improvements are also included in the technical scope of the disclosure.
[0106] In the above embodiment, the processing is performed by executing a program stored in an auxiliary storage device, but at least a part of the processing of the program may be realized by hardware. The processing flow of the program described in the above embodiment is also an example, and unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within the scope of the gist.
[0107] Furthermore, in order to execute the processing in the above-described embodiment by a computer, a program in which the above-described processing is written in code that can be processed by a computer may be stored on a storage medium such as an optical disk and distributed.
[0108] In the above-described embodiment, a CPU is used as an example of a general-purpose processor. However, in the above-described embodiment, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).
[0109] Furthermore, the operation of the processor in the above-described embodiments may not only be performed by a single processor, but may also be performed by multiple processors working together, or may be performed by multiple processors located in physically separate locations working together.
[0110] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[0111] The SDGs have been proposed to realize a sustainable society. One embodiment of the present invention is thought to be a technology that can contribute to "No. 9" and other goals. [Explanation of symbols]
[0112] 1 Estimation device 2 vehicles 3. Driving condition 4. Measurement data 5 Estimation part 6. Evaluation Data 7 Related Data 21, 22, 23 Measurement sensors 24 Camera 51 Estimation Model 51V Image Feature Model 51A 1st model 51B 2nd model 52 memory 100 Computer main body 108 Auxiliary storage 108P Estimation Program 114 Communications Department 116 Operation display section 118 Detector 510 Input Layer 512 Reservoir Layer 516 Decision Tree 518 Derivation part
Claims
1. an acquisition unit that acquires a series of multiple image information related to the field of view of an occupant of the vehicle, the image information changing over time in response to the running of the vehicle on which the tire is mounted and indicating inputs and outputs to the occupant of the vehicle; a generation unit that uses an image feature model that receives the plurality of pieces of image information acquired by the acquisition unit as an input and outputs feature information related to the occupant's field of view to generate image feature information that is feature information related to the occupant's field of view in each of the plurality of pieces of image information; an extraction unit that uses the image feature information generated by the generation unit as an input and extracts a plurality of pieces of feature information that indicate features related to the occupant's field of view in each of a plurality of driving scenes obtained by dividing the driving of the vehicle into a plurality of time intervals, using a first model that is capable of extracting information that indicates features that change over time with respect to the input information; a derivation unit that uses a second model that receives the plurality of pieces of feature information extracted by the extraction unit as an input and outputs evaluation information indicating an evaluation of the tire, to derive association information indicating an association between the evaluation information and the field of view of the occupant in each of the plurality of driving scenes; An information processing device comprising:
2. the extraction unit extracts, as the plurality of pieces of feature information, information indicating an internal state of the first model at a plurality of predetermined different timings corresponding to each of the plurality of driving scenes; The information processing device according to claim 1 .
3. the derivation unit derives, as the relevance information, an importance indicating a degree of relevance of the visual field of the occupant to the evaluation information, based on information indicating an internal state of the second model in each of the plurality of driving scenes. The information processing device according to claim 1 .
4. The acquisition unit further acquiring a plurality of series of time-series information that changes over time in response to the running of a vehicle on which the tire is mounted and that indicates inputs and outputs to an occupant of the vehicle other than inputs and outputs related to the field of view of the occupant; The extraction unit using the plurality of sets of image information and the plurality of sets of time-series information acquired by the acquisition unit as inputs, and using the first model, further extracting information indicating features of inputs and outputs of the occupant in each of the plurality of driving scenes as the plurality of sets of feature information; The lead-out portion is deriving, as the association information, information indicating an association between the evaluation information and input / output for the occupant in each of the plurality of driving scenes using the second model that receives the plurality of pieces of feature information extracted by the extraction unit as an input and outputs the evaluation information indicating an evaluation of the tire; The information processing device according to claim 1 .
5. the image feature model is a variational autoencoder model; the first model is a recurrent neural network model; The second model is a random forest model. The information processing device according to claim 1 .
6. the time-series information includes arm information indicating inputs and outputs related to the arms of the occupant, waist information indicating inputs to the waist of the occupant, and head information indicating inputs to the head of the occupant, The arm information is data indicating at least one of a steering angle and a steering torque, the waist information is data indicating a center of gravity position of the occupant on the seat, and the head information is data indicating a head acceleration and an angular velocity. The information processing device according to claim 4 .
7. The computer acquiring a plurality of series of image information relating to the field of view of an occupant of the vehicle, the image information changing over time in response to the running of the vehicle on which the tire is mounted and representing inputs and outputs to the occupant of the vehicle; generating image feature information, which is feature information regarding the visual field of the occupant in each of the plurality of pieces of image information, using an image feature model that receives the plurality of pieces of image information as an input and outputs feature information regarding the visual field of the occupant; extracting a plurality of pieces of feature information indicating features related to the occupant's field of view in each of a plurality of driving scenes obtained by dividing the driving of the vehicle into a plurality of time intervals using a first model that uses the generated image feature information as an input and is capable of extracting information indicating features that change over time with respect to the input information; deriving association information indicating an association between the evaluation information and the field of view of the occupant in each of the plurality of driving scenes using a second model that receives the extracted plurality of pieces of feature information as an input and outputs evaluation information indicating an evaluation of the tire; An information processing method that processes the following.
8. To the computer acquiring a plurality of series of image information relating to the field of view of an occupant of the vehicle, the image information changing over time in response to the running of the vehicle on which the tire is mounted and representing inputs and outputs to the occupant of the vehicle; generating image feature information, which is feature information regarding the visual field of the occupant in each of the plurality of pieces of image information, using an image feature model that receives the plurality of pieces of image information as an input and outputs feature information regarding the visual field of the occupant; extracting a plurality of pieces of feature information indicating features related to the occupant's field of view in each of a plurality of driving scenes obtained by dividing the driving of the vehicle into a plurality of time intervals using a first model that uses the generated image feature information as an input and is capable of extracting information indicating features that change over time with respect to the input information; deriving association information indicating an association between the evaluation information and the field of view of the occupant in each of the plurality of driving scenes using a second model that receives the extracted plurality of pieces of feature information as an input and outputs evaluation information indicating an evaluation of the tire; An information processing program that processes information.
Citation Information
Patent Citations
Information processing apparatus, calculation method, and calculation program
JP2020166442A
Time segment representative feature vector generation device
WO2010087125A1