Information processing device, information processing method, and program
The information processing device improves autonomous driving accuracy by recognizing road environments through targeted resource allocation and model retraining, addressing the challenge of maintaining accurate road maps in dynamic conditions.
Patent Information
- Application Number
- JP2024068298
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-19
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2044-04-19
AI Technical Summary
Existing autonomous driving systems face challenges in maintaining accurate road environment recognition while minimizing resource allocation and cost, particularly when road structures change or lanes are closed, leading to inconsistencies in road map updates.
An information processing device performs a first process to recognize predetermined objects using an in-vehicle camera and a second process to improve recognition accuracy for regions where the vehicle is predicted to travel, allocating additional resources to these areas based on vehicle data and machine learning model retraining.
This approach enhances road environment recognition accuracy while optimizing resource use, ensuring stable and cost-effective autonomous driving by improving recognition in critical areas as the vehicle moves.
Smart Images

Figure 0007910593000001 
Figure 0007910593000002 
Figure 0007910593000003
Abstract
Description
Technical Field
[0001] This disclosure relates to vehicle technology.
Background Art
[0002] There is a technology for generating road map data in real time while sensing the road environment. In this regard, for example, Patent Document 1 discloses an apparatus that performs weighted correction according to the driving environment on an image recognition result and recognizes road division lines based on the correction result.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] <� This disclosure aims to achieve both high recognition accuracy of the road environment and low cost.
Means for Solving the Problems
[0005] One aspect of this disclosure is performing a first process of recognizing a predetermined object based on an image acquired by an in-vehicle camera of a first vehicle, and performing a second process for improving recognition accuracy on a first area including at least a road area where the first vehicle is predicted to travel, among the areas included in the image, and having a control unit that executes the processes.
[0006] One aspect of this disclosure is An information processing method performed by an information processing device capable of communicating with a first vehicle, The information processing method includes: performing a first process to recognize a predetermined object based on an image acquired by an on-board camera of the first vehicle; and performing a second process to improve the recognition accuracy for a first region of the image that includes at least a road region on which the first vehicle is expected to travel.
[0007] Other embodiments include a program for causing a computer to execute the above-described information processing method, or a computer-readable storage medium that non-temporarily stores the program. [Effects of the Invention]
[0008] According to this disclosure, it is possible to achieve both high accuracy in recognizing the road environment and low costs. [Brief explanation of the drawing]
[0009] [Figure 1] A diagram illustrating the issues in this disclosure. [Figure 2] A diagram illustrating the issues in this disclosure. [Figure 3] A diagram illustrating the configuration of the in-vehicle device 10. [Figure 4] A diagram illustrating the outline of the additional processing in the first embodiment. [Figure 5] A diagram illustrating the data flow in the in-vehicle device 10. [Figure 6] A flowchart of the process performed by the in-vehicle device 10 in the first embodiment. [Figure 7] A diagram illustrating the overview of the process in the second embodiment. [Figure 8] A diagram illustrating the overview of the process in the second embodiment. [Figure 9] A flowchart of the process performed by the in-vehicle device 10 in the second embodiment. [Figure 10] A diagram illustrating the outline of the process in a modified example of the second embodiment. [Figure 11]Flowchart of the process executed by the in-vehicle device 10 in the third embodiment.
Embodiments for Carrying Out the Invention
[0010] In recent years, research has been underway on an autonomous driving system in which a vehicle autonomously travels along a preset route. In autonomous driving, the vehicle determines its own position and orientation by comparing a pre-stored road map with the result of sensing the road environment.
[0011] However, in such a system, there is a problem that the road map must always be kept up-to-date. For example, when a building or structure along the road is demolished, an inconsistency with the road map occurs, and there is a risk that the vehicle may not be able to correctly recognize its own position. The same problem also occurs when some lanes are closed due to construction or the like. There is also a method of updating the road map based on the information collected by a probe car, but since a time lag occurs, this problem cannot be completely solved.
[0012] In order to address this, technologies have been studied in which the vehicle travels while recognizing the road environment in real time without holding a road map on the vehicle side. For example, the vehicle holds only data for route guidance and determines the line to travel based on the result of recognizing the road area (drivable area) in real time. In this form, it is necessary to accurately determine the road area based on the image acquired by the in-vehicle camera.
[0013] In order to accurately determine the road area, for example, resources allocated to image processing may be increased, such as improving the resolution and frame rate. However, in-vehicle devices have limited resources compared to stationary computers. Therefore, a resource allocation method with better cost performance is required. The information processing device according to the present disclosure solves such problems.
[0014] An information processing apparatus according to one aspect of the present disclosure performs a first process of recognizing a predetermined object based on an image acquired by an in-vehicle camera of a first vehicle, and a second process for improving recognition accuracy for a first region including at least a road region in which the first vehicle is predicted to travel, among the regions included in the image. The information processing apparatus has a control unit that executes the above processes.
[0015] The information processing apparatus according to the present disclosure may be an in-vehicle device mounted on a first vehicle, or may be a server device that performs processing based on an image acquired by the first vehicle and provides information to the first vehicle. The control unit executes a first process of recognizing a predetermined object based on an image acquired by the in-vehicle camera. The predetermined object may be, for example, a road on which the first vehicle is traveling. For example, the control unit can recognize a region (road region) where the host vehicle can travel based on the image and travel while recognizing it.
[0016] Further, the control unit performs a second process for improving recognition accuracy for the first region in the image. The first region is a region including at least a road region in which the host vehicle is predicted to travel. The road region in which the host vehicle is predicted to travel may be any road region where the host vehicle may travel. The second process is an additional process for improving recognition accuracy. That is, the control unit allocates more resources to perform recognition processing for regions where the host vehicle may travel in the future. In other words, the control unit does not execute an additional process for improving recognition accuracy for regions where the host vehicle has no possibility of traveling. According to such a configuration, additional resources can be allocated to more important regions in the image.
[0017] Note that the first region may at least include a region that is separated from the first vehicle by a predetermined value or more. Areas that are more than a predetermined distance from the first vehicle have lower resolution compared to other areas, resulting in relatively lower recognition accuracy. Therefore, applying additional processing to such areas can improve overall recognition accuracy.
[0018] Furthermore, the control unit may determine the first region in the image based on the planned route of the first vehicle. The planned route of the first vehicle may be determined, for example, based on information obtained from a navigation system or a control device that manages autonomous driving. This makes it possible to determine where the vehicle is headed in the image and to appropriately determine the first region.
[0019] Furthermore, the control unit may determine the first region in the image based on vehicle data acquired from the first vehicle. Vehicle data refers to data that indicates the state and behavior of the vehicle, for example. Vehicle data may include data indicating the steering angle and the operation status of the turn signals. Based on such data, the control unit can estimate the vehicle's planned route.
[0020] Furthermore, the control unit may set the first region to include a second road region corresponding to the lane in which a second vehicle located near the first vehicle is traveling.
[0021] In some cases, improved recognition accuracy is required not only for the lane in which the vehicle is traveling, but also for lanes in the vicinity of the vehicle in which other vehicles are traveling. Therefore, the first area may be set to include the lane in which the second vehicle is traveling.
[0022] As a second process, for example, one could describe a process that increases the resources used for image recognition compared to the first process. For example, this could take the form of "performing recognition without downsampling the resolution" or "performing recognition without reducing the frame rate."
[0023] Furthermore, if the first process is one that is performed using a machine learning model, the second process may be one in which the machine learning model is retrained. For example, if image recognition is performed at each time step, as an object in the first region approaches the first vehicle, it becomes possible to recognize the object more accurately. Therefore, the recognition results obtained in later time steps may be used as training data for the input data in previous time steps, and the machine learning model may be retrained. This can improve the accuracy of object recognition.
[0024] Furthermore, the second process may be a process that corrects the recognition result of the road area included in the first area by using information sources not used in the first process. For example, the second process may involve correcting the recognition result of the road area included in the first area by using data output from sensors other than image sensors (e.g., radar, LiDAR, etc.) or data received from outside the vehicle.
[0025] Embodiments of this disclosure will be described below with reference to the drawings. The configurations of the following embodiments are illustrative, and this disclosure is not limited to the configurations of these embodiments.
[0026] (First embodiment) An overview of the vehicle system according to the first embodiment will be described. The vehicle system according to this embodiment comprises a vehicle 1, an on-board device 10 mounted on the vehicle 1, and a camera 20 mounted on the vehicle 1.
[0027] Refer to Figures 1 and 2 to explain the problems the system solves. The onboard device 10 mounted on vehicle 1 recognizes the road area around the vehicle based on the image captured by the camera 20, and uses the recognition results to perform control that allows vehicle 1 to drive autonomously. The in-vehicle device 10 stores data for guiding the vehicle to its destination (guidance data) and drives along the recognized road area according to the guidance data.
[0028] A road area is typically the area on which a vehicle 1 can travel. A road area can be recognized by detecting the road boundary (road edge), but the object of recognition is not limited to the road edge. For example, lane boundary lines, lane centerlines, and road centerlines may also be objects of recognition.
[0029] In the example shown in Figure 1, we assume there are two road segments (segments A and B) composed of curves, and that vehicle 1 travels along them. Segment A is a curve with radius r1, and segment B is a curve with radius r2, where r2 is smaller than r1. In other words, segment B is a sharper curve than segment A.
[0030] Figure 2 shows an example of an image captured by camera 20 on vehicle 1 just before entering the curve shown in the figure. This image captures both segment A and segment B. However, areas far from the vehicle tend to have lower image resolution, resulting in worse recognition accuracy compared to areas closer to the vehicle. For example, at the time shown in the diagram, the on-board device 10 may recognize that "a curve with radius r1 continues (both segments A and B have a radius of r1)." In this case, the on-board device 10 generates a driving trajectory based on the premise that "a curve with radius r1 continues," as indicated by reference numeral 1001 in Figure 1.
[0031] However, in reality, the curve in segment B has a smaller radius than the curve in segment A. As the image resolution increases for closer road areas, the on-board device 10 becomes able to correctly recognize the curve radius of segment B as vehicle 1 moves. Consequently, the driving trajectory indicated by reference numeral 1001 may be corrected during driving, resulting in a sudden change in the turning rate in the middle of a curve, which is undesirable for the vehicle's driving stability.
[0032] Therefore, the in-vehicle device 10 according to this embodiment performs additional processing (hereinafter referred to as "additional processing") to improve recognition accuracy for areas further away from the vehicle (for example, area 2002 in Figure 2). By performing additional processing only partially in this way, it is possible to improve recognition accuracy while minimizing the increase in cost.
[0033] [Device configuration] Figure 3 shows an example of the configuration of vehicle 1. The in-vehicle device 10 can be configured as a computer having a processor (CPU, GPU, etc.), main memory (RAM, ROM, etc.), and auxiliary storage (EPROM, hard disk drive, removable media, etc.). The auxiliary storage contains an operating system (OS), various programs, various tables, etc., and by executing the programs stored therein, various functions (software modules) that meet predetermined purposes, as described later, can be realized. However, some or all of the functions may be, for example, It may also be implemented as a hardware module using hardware circuits such as ASICs and FPGAs.
[0034] The in-vehicle device 10 is comprised of a control unit 11, a storage unit 12, a communication unit 13, and an input / output unit 14.
[0035] The control unit 11 is a computing unit that realizes various functions of the in-vehicle device 10 by executing a predetermined program. The control unit 11 can be implemented by a hardware processor such as a CPU. The control unit 11 may also be configured to include RAM, ROM (Read Only Memory), cache memory, etc.
[0036] The control unit 11 is composed of four software modules: a recognition unit 111, a generation unit 112, a correction unit 113, and a driving control unit 114. Each software module may be implemented by the control unit 11 (CPU, etc.) executing a program stored in the storage unit 12, which will be described later.
[0037] The recognition unit 111 acquires an image from the camera 20 (described later) and recognizes the road area contained in the image. In this embodiment, the recognition unit 111 converts the acquired image into features and inputs the obtained features into a machine learning model stored in the memory unit 12. The machine learning model is a model specialized for recognizing road areas (hereinafter referred to as the recognition model). This makes it possible to estimate the road area in the image. The road area recognition result is transmitted to the generation unit 112.
[0038] The generation unit 112 generates map data based on the road area recognition results performed by the recognition unit 111. Map data is a two-dimensional or three-dimensional map showing the drivable area (road area) within the space where the vehicle is located. Based on the recognition results received from the recognition unit 111, the generation unit 112 identifies the drivable area (road area) within the space and generates data representing the identified road area (road area data). It also generates map data, which is a collection of road area data. The map data may be deleted each time, or the map data generated once may be stored and reused.
[0039] In this embodiment, additional processing may be performed to improve the accuracy of road areas that have been recognized once. In other words, the map data that has been generated may be updated as a vehicle travels through it.
[0040] The correction unit 113 performs additional processing to improve the accuracy of road areas that have been recognized once. As mentioned above, the resolution of distant road areas is low, while the resolution of nearby road areas is high. In other words, the recognition accuracy of distant road areas is initially low, and as vehicle 1 moves, the recognition accuracy gradually improves. In other words, the correct answer for the result of recognizing a road area at one time step may be determined at a later time step. Therefore, the correction unit 113 retrains the machine learning model for recognizing the road area when the target road area approaches the vehicle.
[0041] Figure 4 is a diagram illustrating the process. At time t1 in the figure, assume that there is a road area B near the vehicle and a road area A in the distance. The recognition model of the in-vehicle device 10 recognizes the distant road area A and the nearby road area B separately. In this embodiment, the distant area refers to the area that includes the road area that the vehicle is expected to travel through, and which is at a distance of a predetermined value or more from the vehicle. Here, road area B is closer to the vehicle. Therefore, the recognition accuracy is higher than that of road area A.
[0042] Meanwhile, as vehicle 1 moves forward and time t2 arrives, the on-board device 10 becomes able to recognize road area A more accurately. In this embodiment, at this timing, the correction unit 113 retrains the recognition model. Specifically, the recognition result of road region A at time t2 (i.e., the highly accurate recognition result) is treated as the correct answer for the features used for recognition at time t1. In other words, the recognition model is retrained using the features corresponding to road region A at time t1 as input data and the recognition result at time t2 as training data.
[0043] The driving control unit 114 controls the autonomous driving of the vehicle based on the generated map data. The driving control unit 114 detects obstacles around the vehicle based on the image captured by the camera 20, and controls the autonomous driving of the vehicle by using the data obtained as a result of the detection (hereinafter referred to as environmental data) in combination with the generated map data. Environmental data may include, but is not limited to, the number and location of vehicles present around the vehicle, and the number and location of obstacles present around the vehicle (e.g., pedestrians, bicycles, structures, buildings, etc.). Anything necessary for autonomous driving may be detected. Environmental data is generated by a process different from the road area recognition process.
[0044] The driving control unit 114 drives the vehicle along the route indicated by the guidance data and ensures that no obstacles enter a predetermined safety area centered on the vehicle. Known methods can be used for autonomous driving of the vehicle. The guidance data is data for providing directions, and typically includes information for guiding the driver to right or left turns, interchanges to be used, etc., on the way to the destination.
[0045] The memory unit 12 is a means for storing information and is composed of storage media such as RAM, magnetic disks, and flash memory. The memory unit 12 stores programs executed by the control unit 11, data used by those programs, and so on.
[0046] The memory unit 12 stores the aforementioned map data, guidance data, and recognition model, etc.
[0047] The communication unit 13 is a wireless communication interface for connecting the in-vehicle device 10 to the in-vehicle network.
[0048] The input / output unit 14 is a unit that receives input operations performed by the vehicle occupants and presents information to the occupants. In this embodiment, it consists of a single touch panel display. That is, it is composed of a liquid crystal display and its control means, and a touch panel and its control means.
[0049] The specific hardware configuration of the in-vehicle device 10 can be appropriately omitted, replaced, and added depending on the embodiment. For example, the control unit 11 may include multiple hardware processors. The hardware processors may consist of microprocessors, FPGAs, GPUs, etc. In addition, input / output devices other than those exemplified (e.g., optical drives, etc.) may be added. Furthermore, the in-vehicle device 10 may be composed of multiple computers. In this case, the hardware configurations of each computer may or may not be the same.
[0050] [Details of additional processing] Next, we will describe the details of the additional processing performed by the correction unit 113.
[0051] Figure 5 is a chart that shows in more detail the data flow when the in-vehicle device 10 retrains its recognition model. First, at time t1, an image (referred to as the camera image) is acquired from the camera 20. The recognition unit 111 converts the camera image into features and recognizes road areas using a recognition model. At this time, the recognition unit 111 divides the features into road areas located near the vehicle (hereinafter referred to as the nearby area) and road areas that are more than a predetermined distance from the vehicle (hereinafter referred to as the far area), and recognizes each area individually. The nearby area and the far area can be set at any position in the image, but both the nearby area and the far area include road areas that the vehicle is expected to travel on, with the nearby area being closer to the vehicle. The recognition unit 111 may, for example, recognize the lane the vehicle is currently traveling in and then set the nearby area and the far area on that lane. For example, in the example in Figure 2, code 2001 can be the nearby area and code 2002 can be the far area. As a result, data including the recognition results (road area data) is generated for both the nearby and distant areas.
[0052] Next, let's consider the case where the vehicle continues to travel until time t2. In this example, time t2 is the point at which the road area that was far away from the vehicle at time t1 comes close to the vehicle. At time t2, the recognition unit 111 similarly performs road area recognition for both the nearby and distant regions. Here, it is assumed that the recognition result for the nearby region at time t2 (code 502) and the recognition result for the distant region at time t1 (code 501) are the result of recognizing the same road area. However, in reality, code 501 has lower recognition accuracy. Therefore, the correction unit 113 compares both code 501 and code 502, and if the error is greater than or equal to a predetermined value, it performs retraining of the recognition model.
[0053] Retraining is performed using the feature vector corresponding to the far region at time t1 (code 503) as input data and the recognition result at time t2 (code 502) as training data. This improves the recognition model's ability to recognize the far region at time t1.
[0054] The decision of whether or not to perform retraining may be dynamically determined based on the results of comparing the road area data indicated by reference numeral 501 and the road area data indicated by reference numeral 502. For example, if the differences in shape, area, curve radius, etc., between the road areas obtained as a result of recognition are greater than or equal to a predetermined value, it may be determined that retraining is necessary. Furthermore, even if the discrepancy in the recognition results is large, if it does not impede driving safety, retraining may be omitted.
[0055] [Processing flow] Next, we will explain the processing flow performed by the in-vehicle device 10. Figure 6 is a flowchart of the processing performed by the in-vehicle device 10. The illustrated processing is performed periodically while the vehicle 1 is in motion.
[0056] First, in step S11, the recognition unit 111 acquires an image from the camera 20.
[0057] Next, in step S12, the recognition unit 111 determines the road area in the image where the vehicle is expected to travel. The road area where the vehicle is expected to travel can be determined based on information such as the following. • The vehicle's lane, determined by image recognition. • Location information obtained by the GPS module • Direction of travel (vehicle direction) obtained from the gyroscope. • Steering angle Based on the combination of this information, the recognition unit 111 can determine the road area in the image that the vehicle is predicted to be heading towards.
[0058] The recognition unit 111 may also estimate that the vehicle will change its course within a predetermined period based on other information. For example, if the vehicle's turn signal is activated, the recognition unit 111 may estimate that the vehicle will change lanes in the direction indicated by the turn signal within a predetermined time. Such determinations may be made, for example, based on CAN data flowing through the in-vehicle network. This data is an example of "vehicle data".
[0059] Next, in step S13, the recognition unit 111 sets a nearby area and a far-away area in the image acquired from the camera 20. Both the nearby area and the far-away area are set within the road area where the vehicle is expected to travel. For example, both the nearby area and the far-away area can be areas that include the lane in which the vehicle is currently traveling. The recognition unit 111 may, for example, recognize the lane in which the vehicle is traveling and then set the nearby area and the far-away area on that lane. For example, in the example in Figure 2, reference numeral 2001 is the nearby area and reference numeral 2002 is the far-away area. The nearby area is the area where the road area is recognized by normal processing, while the distant area is the area to which additional processing is performed in addition to normal processing. Therefore, it is preferable to set the distant area to an area where recognition accuracy is expected to be low. For example, the distant area can be an area that is more than a predetermined distance away from the vehicle and has a relatively low resolution. By defining the nearby and far regions in this step, data containing the recognition results (road region data) is generated for each of the nearby and far regions.
[0060] Next, in step S14, the recognition unit 111 recognizes the road area for the processing area and generates road area data as a result. The road area data can be data that represents the area (road area) in which the vehicle can travel in a two-dimensional or three-dimensional space. If the processing area is divided into a nearby area and a distant area, the recognition unit 111 generates road area data for each.
[0061] Next, in step S15, the correction unit 113 determines whether there are any areas to be reprocessed among the road areas processed in step S15. Areas to be reprocessed are areas where additional processing should be performed to improve recognition accuracy. For example, as explained with reference to Figures 4 and 5, if the recognition accuracy of the road area performed for the distant area in a previous time step was low, it can be determined that reprocessing should be performed for that road area in the current time step. Low accuracy of the road area recognized in a previous time step can be determined, for example, by comparing road area data corresponding to the same road area between different time steps, as shown by reference numerals 501 and 502 in Figure 5. Furthermore, if there are no areas where the recognition accuracy is determined to be low, or no areas where low recognition accuracy would pose a safety problem, the correction unit 113 may determine that "there are no areas to be reprocessed."
[0062] If it is determined in step S15 that there is an area to be reprocessed, the process proceeds to step S16, where the correction unit 113 performs retraining on the target area. In step S16, for example, the recognition model is retrained using the feature quantities (code 503) used for estimation in past time steps as input data and the road area recognition result (code 502) in the current time step as training data.
[0063] As described above, the in-vehicle device 10 according to the first embodiment performs additional processing to improve recognition accuracy for areas included in the image acquired by the in-vehicle camera that include at least the road area where the vehicle is expected to travel. For road areas where there is no possibility of recognition, processing to improve recognition accuracy is not performed. This allows for improved recognition accuracy while minimizing the increase in costs.
[0064] (Second Embodiment) In the first embodiment, the recognition accuracy was improved in step S16 by retraining the recognition model. Specifically, the additional processing involved retraining the recognition model using both the data acquired in past time steps and the data acquired in the current time step. On the other hand, the additional processing may be a process that improves recognition accuracy using only the data obtained in the current time step.
[0065] The in-vehicle device 10 according to the second embodiment determines the road area where the vehicle is expected to travel from the camera image, and allocates more processing resources to that road area in real time.
[0066] Here, we will provide an additional example of a method for determining the road area in which the vehicle is expected to travel. The first method involves estimating the vehicle's path and then using the estimation result to determine the "area in which the vehicle is expected to travel." For example, if the vehicle is traveling in a lane where lane changes are not permitted, it can be estimated that "the vehicle will continue to travel in the same lane." In this case, for example, the area indicated by reference numeral 701 in Figure 7 can be considered the area in which the vehicle is expected to travel.
[0067] Furthermore, for example, if it is determined that the vehicle will change lanes within a predetermined period based on route information acquired in advance or vehicle information acquired in real time during driving (e.g., the status of the turn signals), the area corresponding to the adjacent lane may be designated as the area where the vehicle is expected to travel. Route information may be acquired from a navigation system or the like, or from an ECU that controls autonomous or semi-autonomous driving.
[0068] The second method involves dividing the area included in the camera image into areas where the vehicle may travel and areas where it is not likely to travel, and treating the area where the vehicle may travel as the "area where the vehicle's travel is predicted."
[0069] Figure 8 is a diagram illustrating this method. In the illustrated example, the image captured by the in-vehicle camera is divided into the region indicated by code 801 and the region indicated by code 802. Of these, the region indicated by code 802 is an area where the vehicle will not travel, so a low recognition accuracy is acceptable for this region. In this way, the region included in the camera image can be divided into an area where the vehicle may travel and an area where it does not, and the former can be treated as an area where the vehicle's travel is predicted.
[0070] In the second embodiment, the correction unit 113 performs additional processing in real time. For example, if the camera outputs images at a frame rate of 60 frames per second, it is possible to "perform recognition processing on region 802 at 30 frames per second, and on region 801 at 60 frames per second." Also, if the resolution of the images output by the camera is downsampled, it is possible to "perform recognition processing on region 802 after downsampling, and on region 801 without downsampling."
[0071] Furthermore, the correction unit 113 may use information other than the camera image to correct the road area recognition result. For example, the correction unit 113 may use sensors other than the vehicle-mounted image sensor to perform processing to correct the recognition result.
[0072] Figure 9 is a flowchart of the processes performed by the in-vehicle device 10 in the second embodiment. The illustrated processes are executed periodically while the vehicle 1 is in motion.
[0073] First, in step S21, the recognition unit 111 acquires an image from the camera 20. Next, in step S22, the recognition unit 111 determines the road area in the image where the vehicle is expected to travel. The road area where the vehicle is expected to travel can be determined by the same method as in step S12. The road area where the vehicle is expected to travel is the road area for which recognition accuracy should be improved. In the following explanation, the road area where the vehicle is expected to travel will be referred to as the "planned travel area," and all other areas will be referred to as the "non-travel area."
[0074] Next, in step S23, it is determined whether or not to improve the recognition accuracy for the planned driving area. For example, if there is a lot of traffic and it is preferable to take a large safety margin, it is preferable to improve the recognition accuracy for the planned driving area. Also, if the likelihood (confidence level) output by the recognition model is below a predetermined value, it is preferable to improve the recognition accuracy. If the result in this step is negative, the process proceeds to step S24. If the result in this step is positive, the process proceeds to step S25.
[0075] In step S24, the recognition unit 111 recognizes the road area in the processing area by the same process as in step S14, and generates road area data as a result. In this embodiment, since the processing area is not divided into a nearby area and a distant area, the recognition unit 111 generates a single road area data.
[0076] In step S25, in addition to the processing described in step S24, additional processing is performed to improve the recognition accuracy of the planned driving area. The additional processing may involve recognizing the road area using a different method than the processing performed in step S24, and correcting the recognition result based on the result.
[0077] Furthermore, the additional processing may involve correcting the road area recognition result using information sources not used in step S24. For example, if messages transmitted from surrounding vehicles or roadside devices are available, these messages may be used as additional information sources to correct the recognition result. Alternatively, the recognition result may be corrected based on data output by sensors other than the on-board camera. In either case, additional resources are allocated to recognize the road area in the planned driving area compared to other areas.
[0078] As explained above, in the second embodiment, the road area where the vehicle is expected to travel is divided into areas other than the road area, and additional processing is performed to improve the recognition accuracy of the road area where the vehicle is expected to travel. This makes it possible to allocate resources for improving the recognition accuracy of the road area only to the areas where the vehicle is likely to travel.
[0079] In the example shown in Figure 8, additional processing was performed on all areas (region 801) where the vehicle may travel. On the other hand, additional processing may be performed only on areas where the recognition accuracy is relatively low. For example, as shown in Figure 10, the areas where the vehicle may travel can be divided into areas closer to the vehicle (region 801A) and areas further away from the vehicle (region 801B). Region 801A is an area where recognition can be performed with a predetermined accuracy without additional processing. In contrast, region 801B is a region further away from region 801A. This is an area with low recognition accuracy. In such cases, the in-vehicle device 10 may perform additional processing only for the area 801B located at a distance. This can save resources required for additional processing.
[0080] (Third embodiment) In the first and second embodiments, additional processing to improve recognition accuracy was performed only for the road area where the vehicle's travel is predicted (the area where the vehicle is scheduled to travel). However, the road area where improving recognition accuracy is beneficial is not necessarily limited to the lane in which the vehicle is traveling. For example, if another vehicle is traveling alongside in an adjacent lane, additional processing to improve recognition accuracy may be performed for the lane in which the other vehicle is traveling. This would make it possible, for example, to accurately predict the behavior of the other vehicle.
[0081] In the third embodiment, in addition to the processes described in the first and second embodiments, the in-vehicle device 10 recognizes other vehicles located near its own vehicle and performs additional processing on the road area (second road area) corresponding to the lane in which the other vehicle is traveling. Figure 11 is a flowchart of this process. The flowchart shown is executed after the process shown in Figure 6, after the process shown in Figure 9, or at any arbitrary timing.
[0082] First, in step S31, the recognition unit 111 detects the presence of another vehicle traveling in a lane other than the lane in which the vehicle is traveling. The presence of the other vehicle may be detected via the camera 20, or it may be detected using other on-board sensors (for example, ultrasonic sensors or LiDAR). Next, in step S32, the correction unit 113 determines whether or not to perform additional processing on the lane in which the detected other vehicle is traveling. For example, if the in-vehicle device 10 has sufficient resources, or if the position of the other vehicle is close to the own vehicle and it is preferable to predict its behavior with higher accuracy, then this step will result in a positive determination.
[0083] If the determination in step S32 is positive, the process proceeds to step S33, where the recognition unit 111 performs additional processing on the area corresponding to the lane in which the detected other vehicle is traveling. If the determination in step S32 is negative, the process ends.
[0084] According to the third embodiment, additional processing is performed to improve the recognition accuracy of lanes in which other vehicles are traveling. This has the effect of enabling accurate prediction of the behavior of those other vehicles.
[0085] The in-vehicle device 10 may also detect all other vehicles from the image acquired by the camera 20 and then perform additional processing for all lanes where other vehicles are present.
[0086] (Other variations) The embodiments described above are merely examples, and this disclosure may be modified as appropriate without departing from its essence. For example, the processes and means described in this disclosure can be freely combined and implemented, as long as no technical inconsistencies arise.
[0087] Furthermore, while the description of the embodiment mentions an example where the in-vehicle device 10 recognizes the road area, the recognition of the road area may be performed by a server device installed in a location different from the vehicle. In this case, the camera image may be transmitted from the in-vehicle device to the server device, and the result of the road area recognition may be transmitted from the server device to the in-vehicle device.
[0088] Furthermore, in the description of the embodiment, the road itself was used as an example of the object of recognition, but the object of recognition The target doesn't necessarily have to be the road itself. For example, objects on the road or objects moving on the road could also be used as the target of recognition.
[0089] Furthermore, in the description of the embodiment, the road area where the vehicle is expected to travel was determined, and additional processing was performed on that road area. However, the road area to which the additional processing is applied does not necessarily have to be an area where travel is expected, as long as it is an area where the vehicle may travel. For example, even if the turn signal is not activated, additional processing may be performed on lanes adjacent to the lane in which the vehicle is currently traveling.
[0090] Furthermore, a process described as being performed by a single device may be divided and executed by multiple devices. Conversely, a process described as being performed by different devices may be executed by a single device. In a computer system, the hardware configuration (server configuration) by which each function is implemented can be flexibly changed.
[0091] The present disclosure can also be realized by supplying a computer program implementing the functions described in the embodiments above to a computer, and having one or more processors in the computer read and execute the program. Such a computer program may be provided to the computer by a non-temporary computer-readable storage medium that can be connected to the computer's system bus, or it may be provided to the computer via a network. Non-temporary computer-readable storage mediums include, for example, any type of disk such as magnetic disks (floppy disks, hard disk drives (HDDs), etc.), optical disks (CD-ROMs, DVDs, Blu-ray discs, etc.), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, flash memory, optical cards, and any type of medium suitable for storing electronic instructions. [Explanation of Symbols]
[0092] 1. Vehicle 10...In-vehicle equipment 11. Control Unit 12...Storage section 13. Communications Department 14...Input / output section
Claims
1. Based on images acquired by the onboard camera of the first vehicle, a first process is performed to recognize a predetermined object using a machine learning model, A second process is performed to improve recognition accuracy for a first region of the image that includes at least the road region where the first vehicle is expected to travel. It has a control unit that performs the following: The second process is a process of retraining the machine learning model using data acquired at the time when an object included in the first region approaches the first vehicle. Information processing device.
2. The first region includes at least a region that is at least a predetermined distance from the first vehicle. The information processing apparatus according to claim 1.
3. The control unit determines the first region in the image based on the planned route of the first vehicle. The information processing apparatus according to claim 1.
4. The control unit determines the first region in the image based on the vehicle data acquired from the first vehicle. The information processing apparatus according to claim 1.
5. The control unit sets the first region to include a second road region corresponding to the lane in which a second vehicle located near the first vehicle is traveling. The information processing apparatus according to claim 1.
6. The second process is a process that increases the resources for image recognition compared to the first process. The information processing apparatus according to claim 1.
7. The first process described above is a process for recognizing the road area, The second process is a process that corrects the recognition result of the road area included in the first area by using information sources not used in the first process. The information processing apparatus according to claim 1.
8. An information processing method performed by an information processing device capable of communicating with a first vehicle, Based on the image acquired by the onboard camera of the first vehicle, a first process is performed to recognize a predetermined object using a machine learning model. A second process is performed to improve recognition accuracy for a first region of the image that includes at least the road region where the first vehicle is expected to travel. Includes, The second process is a process of retraining the machine learning model using data acquired at the time when an object included in the first region approaches the first vehicle. Information processing methods.
9. The first region includes at least a region that is at least a predetermined distance from the first vehicle. The information processing method according to claim 8.
10. Based on the planned route of the first vehicle, the first region in the image is determined. The information processing method according to claim 8.
11. Based on the vehicle data obtained from the first vehicle, the first region in the image is determined. The information processing method according to claim 8.
12. The first area is defined to include a second road area corresponding to the lane in which a second vehicle located near the first vehicle is traveling. The information processing method according to claim 8.
13. The second process is a process that increases the resources for image recognition compared to the first process. The information processing method according to claim 8.
14. The first process described above is a process for recognizing the road area, The second process is a process that corrects the recognition result of the road area included in the first area by using information sources not used in the first process. The information processing method according to claim 8.
15. A program for causing a computer to execute the information processing method described in any one of claims 8 to 14.
Citation Information
Patent Citations
Stereo camera device
JP2011191905A
External environment recognition method and device, and vehicle system
JP2013061919A
Object recognition device
JP2013161187A
Feature image recognition system, feature image recognition method, and computer program
JP2016194815A
Map providing device
JP2020118890A