Learning program, learning device, and learning method
By extracting partial image data and using machine learning to generate a learning model, the learning program reduces the processing requirements for obstacle detection in senior cars, eliminating the need for expensive GPUs and enabling cost-effective obstacle avoidance.
Patent Information
- Application Number
- JP2021018896
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-25
- Filing Date
- 2021-02-09
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2041-02-09
AI Technical Summary
The existing obstacle detection systems in senior cars, which rely on deep learning models like YOLO or SSD, require expensive GPUs for processing, making them costly and impractical for widespread adoption in senior cars.
A learning program and method that reduces the processing requirements for obstacle detection by extracting partial image data from fixed positions in the image data, adding identification information, and using machine learning to generate a learning model, thereby reducing the need for high-powered GPUs.
This approach allows for efficient obstacle detection in senior cars without the need for expensive hardware, reducing processing time and costs while maintaining effective obstacle avoidance.
Smart Images

Figure 0007672034000005 
Figure 0007672034000006 
Figure 0007672034000007
Abstract
Description
[Technical field]
[0001] The present invention relates to a learning program, a learning device, and a learning method. [Background technology]
[0002] In recent years, the use of electric carts (hereafter referred to as senior cars) aimed at supporting the daily activities of the elderly has become widespread. For example, by riding in a senior car when going out for shopping, the elderly can reduce the physical burden associated with going out.
[0003] Here, there is a possibility that the senior car as described above may fall over while being driven due to, for example, a bad road, etc. In this case, the elderly person may not be able to stand up by himself / herself.
[0004] Therefore, for example, a senior car detects the presence of an obstacle on the travel route that may cause the vehicle to fall, and travels while avoiding the detected obstacle. This enables the senior car to automatically avoid the risk of falling (see Patent Documents 1 and 2). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2016-186703 A [Patent Document 2] JP 2018-156408 A Summary of the Invention [Problem to be solved by the invention]
[0006] The above-mentioned obstacle detection is generally performed by using deep learning models such as YOLO (You Only Live Once) and SSD (Single Shot Multibox Detector). In this case, the need to perform processing using the deep learning model requires the senior car to be equipped with expensive equipment such as a high-performance GPU (Graphics Processing Unit).
[0007] This makes it possible for the senior car to carry out processing using a deep learning model that requires a large amount of processing within the required time (for example, within the time it takes for the senior car to avoid an obstacle it detects).
[0008] However, when the sales price range of senior cars is taken into consideration, it may not always be appropriate to install such expensive equipment in senior cars. Therefore, in the field of senior cars, there is a demand for a method that can reduce the amount of processing required to detect obstacles on the travel route.
[0009] SUMMARY OF THE PRESENT EMBODIMENT An object of the present invention is to provide a learning program, a learning device, and a learning method that make it possible to reduce the amount of processing required to detect an obstacle on a travel route. [Means for solving the problem]
[0010] In order to achieve the above-mentioned object, the learning program of the present invention is characterized in that it causes a computer to execute a process of extracting multiple training partial image data from training image data captured by an imaging device, each of which corresponds to multiple fixed positions in the training image data, generating multiple training data by adding identification information indicating the type of object reflected in each of the training partial image data to color information in each of the multiple training partial image data, and generating a learning model by performing machine learning using the multiple training data.
[0011] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that, for each of the multiple learning partial image data, the multiple learning data are generated by adding identification information indicating the type of the object reflected in each learning partial image data to the color information in each learning partial image data and distance information from the imaging device to the object reflected in each learning partial image data.
[0012] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that the multiple fixed positions include a multiple first fixed positions having a first size and a multiple second fixed positions having a second size smaller than the first size.
[0013] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that the multiple learning partial image data corresponding to the multiple first fixed positions are learning partial image data that depict an area closer to the imaging device than the multiple learning partial image data corresponding to the multiple second fixed positions.
[0014] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that the process of extracting the multiple learning partial image data, the process of generating the multiple learning data, and the process of generating the learning model are performed for each of the learning image data contained in video data captured by the imaging device.
[0015] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that it causes a computer to execute a process of extracting, from detection image data captured by an imaging device, a plurality of detection partial image data corresponding to each of the plurality of fixed positions in the detection image data, and outputting, for each of the plurality of detection partial image data, the identification information indicated by the value output from the learning model in response to input of color information in each detection partial image data.
[0016] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that, for each of the multiple detection partial image data, the identification information indicated by the value output from the learning model is output in response to input of the color information in each detection partial image data and distance information from the imaging device to the object reflected in each detection partial image data.
[0017] In addition, the learning program of the present invention for achieving the above-mentioned object is characterized in that, for each of a plurality of training image data captured by an imaging device, identification information indicating the type of object reflected in the plurality of training partial image data corresponding to each of the training image data is added to color information for each of a plurality of training partial image data corresponding to a plurality of fixed positions in each of the training image data, thereby generating a plurality of training data corresponding to each of the plurality of training image data, and performing machine learning using the plurality of training data to generate a learning model, by causing a computer to execute a process.
[0018] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that it causes a computer to execute a process of outputting, for each of a plurality of detection image data captured by an imaging device, each of the identification information indicated by a plurality of values output from the learning model in response to input of color information for each of a plurality of detection partial image data corresponding to the plurality of fixed positions in each detection image data.
[0019] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that, in the process of generating the learning model, the learning model is generated having a plurality of output nodes corresponding to each combination of the plurality of learning partial image data and the type of the target object, and in the process of outputting the identification information, for each of the plurality of detection partial image data included in each of the plurality of detection image data, the identification information corresponding to a value that satisfies a predetermined condition among a plurality of values output from the plurality of output nodes corresponding to each detection partial image data is output as the identification information corresponding to each detection partial image.
[0020] In addition, in one aspect, the learning program of the present invention for achieving the above-mentioned object is characterized in that, in the process of outputting the identification information, for each of the multiple detection partial image data contained in each of the multiple detection image data, the identification information corresponding to the maximum value among the multiple values output from the multiple output nodes corresponding to each detection partial image data is output as the identification information corresponding to each detection partial image.
[0021] In addition, in order to achieve the above-mentioned object, the learning device of the present invention is characterized in having a partial image extraction unit that extracts a plurality of learning partial image data corresponding to each of a plurality of fixed positions in the learning image data from the learning image data captured by an imaging device, a learning data generation unit that generates a plurality of learning data by adding identification information indicating the type of object reflected in each of the learning partial image data to color information in each of the plurality of learning partial image data, and a model generation unit that generates a learning model by performing machine learning using the plurality of learning data.
[0022] In addition, in order to achieve the above-mentioned object, the learning device of the present invention is characterized in that it has a learning data generation unit that generates multiple learning data corresponding to each of multiple training image data captured by an imaging device by adding identification information indicating the type of object reflected in the multiple training partial image data corresponding to each of the training image data to color information for each of multiple training partial image data corresponding to multiple fixed positions in each of the training image data, and a model generation unit that generates a learning model by performing machine learning using the multiple training data.
[0023] In addition, the learning method of the present invention for achieving the above-mentioned object is characterized in that it has a computer execute a process of extracting multiple training partial image data from training image data captured by an imaging device, each of which corresponds to multiple fixed positions in the training image data, generating multiple training data by adding identification information indicating the type of object reflected in each training partial image data to color information in each of the multiple training partial image data, and performing machine learning using the multiple training data to generate a learning model.
[0024] In addition, the learning method of the present invention for achieving the above-mentioned object is characterized in that, for each of a plurality of training image data captured by an imaging device, identification information indicating the type of object reflected in the plurality of training partial image data corresponding to each of the training image data is added to color information for each of a plurality of training partial image data corresponding to a plurality of fixed positions in each of the training image data, thereby generating a plurality of training data corresponding to each of the plurality of training image data, and performing machine learning using the plurality of training data to generate a learning model, by having a computer execute a process. Effect of the Invention
[0025] According to the learning program, learning device, and learning method of the present invention, it is possible to reduce the amount of processing required to detect an obstacle on a travel route. [Brief description of the drawings]
[0026] [Figure 1] FIG. 1 is a diagram showing an example of a configuration of an information processing device 1 according to the first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of the configuration of the detecting terminal 2 in the first embodiment. [Diagram 3] FIG. 3 is a diagram for explaining an outline of the learning process in the first embodiment. [Figure 4] FIG. 4 is a diagram for explaining an outline of the inference process in the first embodiment. [Diagram 5] FIG. 5 is a flowchart illustrating details of the learning process in the first embodiment. [Figure 6] FIG. 6 is a flowchart illustrating details of the learning process in the first embodiment. [Figure 7] FIG. 7 is a flowchart illustrating details of the inference process in the first embodiment. [Figure 8] FIG. 8 is a flowchart illustrating details of the inference process in the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating a specific example of the process of S21. [Figure 10] FIG. 10 is a diagram illustrating a specific example of the process of S26. [Figure 11] FIG. 11 is a diagram for explaining a specific example of the inference process. [Figure 12] FIG. 12 is a flowchart illustrating details of the learning process in the second embodiment. [Figure 13] FIG. 13 is a flowchart illustrating details of the learning process in the second embodiment. [Figure 14] FIG. 14 is a flowchart illustrating details of the inference process in the second embodiment. [Figure 15] FIG. 15 is a diagram illustrating a specific example of the process of S54 and the process of S55. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0027] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the technical scope of the present invention is not limited to these embodiments.
[0028] First, a description will be given of a configuration example of an information processing device 1 (hereinafter also referred to as a learning device 1) in the first embodiment. Fig. 1 is a diagram showing a configuration example of the information processing device 1 in the first embodiment.
[0029] The information processing device 1 is a computer device, for example, a general-purpose PC (Personal Computer). The information processing device 1 performs a learning process (hereinafter, also simply referred to as a learning process) of a learning model used for detecting an obstacle on a travel route, for example.
[0030] 1, the information processing device 1 has a hardware configuration of a general-purpose computer device, and includes, for example, a CPU 101 which is a processor, a memory 102, a communication interface 103, and a storage medium 104. Each unit is connected to each other via a bus 105.
[0031] The storage medium 104 has, for example, a program storage area (not shown) for storing a program (not shown) for performing a learning process.
[0032] The storage medium 104 also includes a storage unit 110 (hereinafter also referred to as a storage area 110) that stores information used when performing the learning process. The storage medium 104 may be, for example, a hard disk drive (HDD) or a solid state drive (SSD).
[0033] The CPU 101 executes a program loaded from the storage medium 104 to the memory 102 to perform learning processing.
[0034] The communication interface 103 communicates with the detection terminal 2 via a network NW such as the Internet.
[0035] Next, a description will be given of a configuration example of the detecting terminal 2 in the first embodiment. Fig. 2 is a diagram showing a configuration example of the detecting terminal 2 in the first embodiment.
[0036] The detection terminal 2 is a computer device, for example, a mobile terminal such as a smartphone. The detection terminal 2 is, for example, a device attached near the front of the senior car in the traveling direction, and the learning model generated by the information processing device 1 operates.
[0037] The detection terminal 2 has the hardware configuration of a general-purpose computer device, and for example, as shown in Fig. 2, has a CPU 201 which is a processor, a memory 202, a communication interface 203, and a storage medium 204. Each part is connected to each other via a bus 205.
[0038] The storage medium 204 has, for example, a program storage area (not shown) that stores a program (not shown) for performing inference processing (hereinafter also simply referred to as inference processing) by using a learning model generated by the information processing device 1.
[0039] The storage medium 204 also includes a storage unit 210 (hereinafter also referred to as a storage area 210) that stores information used when performing an inference process using a learning model generated by the information processing device 1. The storage medium 204 may be, for example, an HDD or an SSD.
[0040] The CPU 201 executes a program loaded from the storage medium 204 to the memory 202 to perform inference processing.
[0041] The communication interface 203 communicates with the information processing device 1 via a network NW such as the Internet. Note that the transfer of information between the information processing device 1 and the detection terminal 2 may be performed manually by an operator using a storage medium such as a USB memory.
[0042] Specifically, for example, while the senior car is traveling, the detection terminal 2 continuously inputs image data included in video data about the traveling route captured by an imaging device (not shown) such as a camera to the learning model previously received from the information processing device 1. Then, the detection terminal 2 continuously detects obstacles existing on the traveling route by using values output from the learning model.
[0043] As a result, when an obstacle on the travel route is detected, the detection terminal 2 notifies, for example, the driver (elderly person) of the senior car of information on the obstacle whose presence has been detected (for example, the position of the obstacle whose presence has been detected). In this case, the detection terminal 2 automatically controls the travel of the senior car so that, for example, the senior car does not come into contact with the obstacle.
[0044] The detection terminal 2 may have a built-in imaging device that captures video data about the travel route.
[0045] In addition, although the following describes a case where the learning process is performed in the information processing device 1, the learning process may be performed in the detection terminal 2. In other words, the detection terminal 2 may perform inference processing by using a learning model generated in the detection terminal 2 itself.
[0046] [Outline of the first embodiment] Next, an overview of the learning process and the inference process in the first embodiment will be described.
[0047] First, an outline of the learning process in the first embodiment will be described. Fig. 3 is a diagram for explaining an outline of the learning process in the first embodiment.
[0048] The image acquisition unit 111 of the information processing device 1 acquires, for example, image data used to generate a learning model (hereinafter, also referred to as learning image data). Specifically, the image acquisition unit 111 acquires, for example, a plurality of image data constituting video data (video data captured in advance by an imaging device) stored in advance in the storage area 110 by an operator.
[0049] Then, the partial image extraction unit 112 of the information processing device 1 extracts, from the image data acquired by the image acquisition unit 111, a plurality of partial image data (hereinafter also referred to as learning partial image data) corresponding to each of the plurality of fixed positions in the image data.
[0050] Next, the learning data generation unit 113 of the information processing device 1 generates multiple pieces of learning data by adding identification information (correct answer label) indicating the type of object depicted in each piece of partial image data to color information (e.g., RGB information) in each piece of partial image data extracted by the partial image extraction unit 112.
[0051] In addition, the learning data generation unit 113 may generate multiple learning data by, for example, adding identification information indicating the type of object reflected in each partial image data to color information in each partial image data and distance information from the imaging device to the object (obstacle) reflected in each partial image data for each partial image data extracted by the partial image extraction unit 112.
[0052] Thereafter, the model generation unit 114 of the information processing device 1 performs machine learning using the plurality of pieces of learning data generated by the learning data generation unit 113, thereby generating a learning model.
[0053] Next, an outline of the inference process in the first embodiment will be described below. Fig. 4 is a diagram for explaining an outline of the inference process in the first embodiment.
[0054] The image acquisition unit 211 of the detection terminal 2 acquires image data (hereinafter also referred to as detection image data) included in video data captured by an imaging device (not shown), for example. Specifically, the image acquisition unit 211 may receive image data transmitted from the detection terminal 2, for example.
[0055] Then, the partial image extraction unit 212 of the detecting terminal 2 extracts, from the image data acquired by the image acquisition unit 211, a plurality of partial image data (hereinafter also referred to as detection partial image data) corresponding to each of the plurality of fixed positions in the image data.
[0056] Next, the value acquisition unit 213 of the detection terminal 2 acquires, for each of the partial image data extracted by the partial image extraction unit 212, a value output from the learning model in response to input of color information (e.g., RGB information) in each partial image data and distance information from the imaging device to the object (obstacle) depicted in each partial image data.
[0057] Thereafter, the information output unit 214 of the detection terminal 2 outputs, for example, the identification information corresponding to the value acquired by the value acquisition unit 213 to an operation terminal (not shown) of the worker.
[0058] In other words, the information processing device 1 in this embodiment generates a learning model by inputting each of partial image data corresponding to a plurality of predetermined fixed positions, thereby making it possible to generate a smaller-scale learning model compared to the case where a learning model is generated by inputting the learning image data as is (for example, a learning model using YOLO or SSD).
[0059] This enables the information processing device 1 to reduce the amount of processing involved in inference processing using a learning model. Therefore, the information processing device 1 can complete the execution time of the inference processing in the detection terminal 2 within a required time (for example, within the time it takes for the senior car to avoid an obstacle detected) without mounting an expensive device such as a GPU.
[0060] Furthermore, by using a learning model generated by inputting partial image data corresponding to a plurality of predetermined fixed positions, the detection terminal 2 is able to recognize the type of object reflected in each partial image data for each of the partial image data extracted from the detection image data. Therefore, the detection terminal 2 does not need to determine whether or not an object exists in each partial image data (detection of the object) before recognizing the type of object reflected in each partial image data. In other words, the detection terminal 2 is able to reduce the detection of the object reflected in each partial image data and the recognition of the type of the object to a recognition problem. Therefore, the detection terminal 2 is able to further reduce the amount of processing involved in the execution of the inference process, and is able to further shorten the execution time of the inference process.
[0061] Furthermore, the detection terminal 2 recognizes the type of object reflected in each partial image data extracted from the detection image data, thereby making it possible to recognize each of the objects even if the detection image data contains multiple obstacles.
[0062] [Details of the first embodiment] Next, the learning process and the inference process in the first embodiment will be described in detail. Figures 5 to 8 are flow charts for explaining the learning process and the inference process in the first embodiment. Figures 9 to 11 are diagrams for explaining the learning process and the inference process in the first embodiment.
[0063] [Learning process details] First, details of the learning process in the first embodiment will be described below. Figures 5 and 6 are diagrams for explaining the details of the learning process.
[0064] 5, the image acquisition unit 111 waits until the learning timing (NO in S11), for example. The learning timing may be the timing when the operator inputs information to start the learning process of the learning model via an operation terminal (not shown), for example.
[0065] Then, when the learning timing arrives (YES in S11), the image acquisition unit 111 acquires video data about RGB information (hereinafter also referred to as first video data) from the video data stored in the storage area 110, for example (S12).
[0066] In this case, the image acquisition unit 111 acquires video data on distance information (hereinafter also referred to as second video data) from the video data stored in the storage area 110, for example (S13).
[0067] That is, the image acquisition section 111 acquires second moving image data obtained by capturing the same area as the first moving image data acquired in the process of S12 at the same timing.
[0068] The first and second moving image data may be moving image data captured by different imaging devices.
[0069] Then, the image acquisition unit 111 acquires one piece of image data (hereinafter also referred to as first image data) from the first moving image data acquired in the process of S12 (S14).
[0070] Specifically, the image acquisition section 111 acquires, in chronological order, one image data constituting the first moving image data acquired in the process of S12.
[0071] Furthermore, the image acquisition unit 111 acquires one piece of image data (hereinafter also referred to as second image data) corresponding to the first image data acquired in the process of S14 from the second moving image data acquired in the process of S13 (S15).
[0072] 6, the partial image extraction unit 112 extracts a plurality of partial image data (hereinafter, also referred to as first partial image data) corresponding to each fixed position in the first image data acquired in the process of S14 (S21). A specific example of the process of S21 will be described below.
[0073] [Specific example of S21 processing] Fig. 9 is a diagram for explaining a specific example of the process of S21. Specifically, Fig. 9 is a diagram for explaining a plurality of partial image data corresponding to a plurality of regions of interest ROI in the first image data DT1. Note that, in the following explanation, the number of each region of interest ROI will be explained as corresponding to the number of each region shown in Fig. 9.
[0074] For example, the operator determines in advance the fixed positions and sizes of the regions of interest ROI before starting the learning process by the information processing device 1. Specifically, for example, as shown in Fig. 9, the operator determines six regions of interest ROI (regions of interest ROI1 to ROI6) each consisting of a rectangle of a first size, and 16 regions of interest ROI (regions of interest ROI7 to ROI22) each consisting of a rectangle of a second size larger than the first size.
[0075] Then, in the process of S21, the partial image extraction unit 112 extracts, for example, from the first image data DT1, each of the partial image data corresponding to the 22 regions of interest ROI determined by the operator.
[0076] Specifically, the partial image extraction unit 112 extracts partial image data (eight partial image data consisting of rectangles of the second size) obtained by dividing the lower end area of the first image data DT1 into eight equal parts along the horizontal direction, as shown in, for example, FIG. 9, as partial image data corresponding to the regions of interest ROI15, ROI16, ROI17, ROI18, ROI19, ROI20, ROI21 and ROI22 (hereinafter, these will be collectively referred to simply as the first regions of interest ROI).
[0077] In addition, the partial image extraction unit 112 extracts partial image data (eight partial image data consisting of rectangles of the second size) obtained by dividing the area adjacent to the upper end of the first region of interest ROI in the first image data DT1 into eight equal parts along the horizontal direction, as shown in, for example, FIG. 9, as partial image data corresponding to the regions of interest ROI7, ROI8, ROI9, ROI10, ROI11, ROI12, ROI13 and ROI14 (hereinafter, these will be collectively referred to simply as the second region of interest ROI).
[0078] Furthermore, as shown in FIG. 9, the partial image extraction unit 112 extracts partial image data (six partial image data consisting of rectangles of a first size) obtained by dividing an area adjacent to the upper end of the second region of interest ROI in the first image data DT1 and near the center in the horizontal direction into six equal parts along the horizontal direction, as partial image data corresponding to the regions of interest ROI1, ROI2, ROI3, ROI4, ROI5, and ROI6 (hereinafter, these are collectively referred to simply as the third region of interest ROI).
[0079] In the example shown in FIG. 9, the worker determines the second size so that, for example, an area 1 m to 2 m ahead of the senior car corresponds to the first region of interest ROI and the second region of interest ROI. That is, for example, when the senior car runs at a speed of 6 km per hour, the worker determines the second size so that an area on the route where the senior car runs in about 1 second corresponds to the first region of interest ROI and the second region of interest ROI. Also, in the example shown in FIG. 9, the worker determines the first size so that, for example, an area 2 m to 4 m ahead of the senior car corresponds to the third region of interest ROI. That is, for example, when the senior car runs at a speed of 6 km per hour, the worker determines the first size so that an area on the route where the senior car runs in about 2 seconds corresponds to the third region of interest ROI.
[0080] This enables the detection terminal 2 to detect, for example, an obstacle present in an area on the route through which the senior car will travel in approximately 1 second (hereinafter also referred to as a danger area), and an obstacle present in an area on the route through which the senior car will travel in approximately 2 seconds (hereinafter also referred to as a notice area), as described below.
[0081] In the following, a case where an operator determines a region of interest ROI of two different sizes will be described, but the operator may determine a region of interest ROI of three or more different sizes. In addition, in the following, a case where an operator determines a rectangular region of interest ROI will be described, but the operator may determine a region of interest ROI of a shape other than a rectangle (for example, a regular hexagon).
[0082] Returning to FIG. 6, the partial image extraction unit 112 extracts a plurality of partial image data (hereinafter also referred to as second partial image data) corresponding to each fixed position in the second image data acquired in the process of S15 (S22).
[0083] Specifically, the partial image extraction unit 112 extracts, for example, second partial image data corresponding to each of the 22 regions of interest ROI described in FIG. 9 from the second image data extracted in the process of S15.
[0084] Thereafter, the learning data generation unit 113 waits until it receives input of a correct label (identification information) corresponding to each combination of the first and second partial image data acquired in the processes of S21 and S22 (NO in S23).
[0085] Then, when the input of the correct label corresponding to each combination is received (YES in S23), the learning data generation unit 113 generates learning data including the partial image data included in each combination and the correct label received in the process of S23 for each combination of the first and second partial image data acquired in the process of S21 and the process of S22 (S24). After that, the learning data generation unit 113 stores the generated learning data in the memory area 110, for example.
[0086] That is, for each combination of the first and second partial image data obtained in the processes of S21 and S22, the worker inputs a correct answer label for the type of object shown in each of the image data included in each combination. Then, the learning data generation unit 113 generates learning data by associating the combinations of the first and second partial image data obtained in the processes of S21 and S22 with the correct answer labels corresponding to the combinations.
[0087] Thereafter, the learning data generating unit 113 determines whether or not all of the first image data and the second image data have been acquired in the processes of S14 and S15 (S25).
[0088] As a result, when it is determined that all of the first image data and second image data have not been acquired (NO in S25), the image acquisition unit 111 performs the processes from S14 onwards again.
[0089] On the other hand, if it is determined that all of the first image data and second image data have been acquired (YES in S25), the model generation unit 114 generates a learning model by performing machine learning using the learning data generated in the processing of S24 (S26).
[0090] The learning data generation unit 113 may normalize the first and second partial image data included in the generated learning data in the process of S24. Specifically, the learning data generation unit 113 may adjust the size of each partial image data so that the first and second partial image data included in the generated learning data are all the same size. Then, the model generation unit 114 may generate a learning model by performing machine learning using learning data including the normalized first and second partial image data in the process of S26.
[0091] This enables the information processing device 1 to perform inference processing by using, for example, the same learning model. A specific example of the processing of S26 will be described below.
[0092] [Specific example of S26 processing] Fig. 10 is a diagram for explaining a specific example of the process of S26. Specifically, Fig. 10 is a specific example for explaining a case where a learning model is generated by applying a distillation technique.
[0093] In the example shown in FIG. 10, the model generation unit 114 generates a teacher model MDt by using learning data DT4 including image data DT2 (first image data DT2) regarding RGB information and image data DT3 (second image data DT3) regarding distance information.
[0094] Then, the model generation unit 114 generates a student model MDs by using, for example, a smaller amount of training data DT4 than when the teacher model MDt was generated, the error between the output from the student model MDs and the output from the teacher model MDt (hereinafter also referred to as the first error), and the error between the output from the student model MDs and the correct label included in the training data DT4 (hereinafter also referred to as the second error).
[0095] Specifically, the model generating unit 114 generates the student model MDs in accordance with, for example, the following formula (1) so that the sum of the first error and the second error is small.
[0096]
number
[0097] In the above equation (1), CE is the cross-entropy loss function, σ is a softmax function parameterized by temperature T, α is a hyperparameter to balance the two losses, z_b is the logit of the teacher model MDt, z_s is the logit of the student model MDs, and y is the ground truth label.
[0098] This enables the information processing device 1 to generate a smaller-scale learning model (student model MDs) in the process of S26.
[0099] [Details of inference process] Next, the details of the inference process in the first embodiment will be described below. Figures 7 and 8 are diagrams for explaining the details of the inference process.
[0100] As shown in Fig. 7, the image acquisition unit 211 waits until it is the inference timing (NO in S31), for example. The inference timing may be, for example, the timing when image data is captured by an imaging device (not shown) mounted on the senior car while the car is moving. In other words, the inference timing may be the timing that occurs every time the imaging device mounted on the senior car captures image data (frame) about the area ahead in the traveling direction. Specifically, when the number of frames of video data captured by the imaging device is 30, the inference timing may be the timing that occurs 30 times per second.
[0101] Then, when the inference timing arrives (YES in S31), the image acquisition unit 211 acquires image data (hereinafter also referred to as third image data) regarding RGB information captured by the imaging device (S32).
[0102] Furthermore, the image acquisition unit 211 acquires image data (hereinafter also referred to as fourth image data) regarding distance information captured by the imaging device (S33).
[0103] Next, the partial image extraction unit 212 extracts a plurality of partial image data (hereinafter also referred to as third partial image data) corresponding to each fixed position in the third image data acquired in the process of S32 (S34).
[0104] Furthermore, the partial image extraction unit 212 extracts a plurality of partial image data (hereinafter also referred to as fourth partial image data) corresponding to each fixed position in the fourth image data acquired in the process of S33 (S35).
[0105] Then, as shown in FIG. 8, for each combination of the third and fourth partial image data acquired in the processes of S34 and S35, the value acquisition unit 213 acquires a value output from the learning model (the learning model generated in the process of S26) in response to the input of the third and fourth partial image data corresponding to each combination (S41).
[0106] Specifically, the value acquisition unit 213, for example, for each combination of the third and fourth partial image data acquired in the processing of S34 and the processing of S35, normalizes the third and fourth partial image data corresponding to each combination, and acquires values output from the learning model in response to the input of the normalized third and fourth image data.
[0107] Then, the information output unit 214 outputs the identification information corresponding to the value acquired in the process of S41 for each combination of the third and fourth partial image data acquired in the process of S34 and the process of S35 (S42). A specific example of the inference process will be described below.
[0108] [Specific examples of inference processing] FIG. 11 is a diagram for explaining a specific example of the inference process.
[0109] 11, the value acquisition unit 213 generates a plurality of partial image data DT5a by dividing image data DT5 (third image data DT5) for RGB information, and generates a plurality of partial image data DT6a by dividing image data DT6 (fourth image data DT6) for distance information.
[0110] Then, for each combination of multiple partial image data DT5a and multiple partial image data DT6a, the value acquisition unit 213 acquires a value to be output by inputting the partial image data DT5a and partial image data DT6a included in each combination into the student model MDs.
[0111] Specifically, for each combination of multiple partial image data DT5a and multiple partial image data DT6a, the value acquisition unit 213 normalizes the partial image data DT5a and partial image data DT6a included in each combination, and acquires the value to be output by inputting the normalized partial image data DT5a and partial image data DT6a into the student model MDs.
[0112] After that, the value acquisition unit 213 specifies the identification information indicated by the value output from the student model MDs for each combination of the partial image data DT5a and the partial image data DT6a.
[0113] Then, the information output unit 214 outputs image data DT7 generated by superimposing each piece of identification information specified by the value acquisition unit 213 on the image data DT5 to the operator's operation terminal (not shown).
[0114] Specifically, the information output unit 214 outputs "creature" as identification information of partial image data corresponding to the regions of interest ROI7, ROI8, ROI15, and ROI6, among the partial image data included in the image data DT5, as shown in image data DT7 in Fig. 11. Also, the information output unit 214 outputs "weed" as identification information of partial image data corresponding to the regions of interest ROI4, ROI5, ROI6, ROI13, and ROI4, among the partial image data included in the image data DT5, as shown in Fig. 11.
[0115] Therefore, in this case, the detection terminal 2 determines that there are obstacles near the left and right ends of the route along which the senior car will be traveling, and notifies the driver of the senior car (elderly person), for example, that he or she needs to travel near the center of the road.
[0116] This makes it possible for the detection terminal 2 to prevent the senior car from falling over due to contact with an obstacle or the like.
[0117] In addition, the operator may prepare multiple student models MDs for the inference process. The value acquisition unit 213 may perform in parallel inference processes corresponding to each combination of multiple partial image data DT5a and multiple partial image data DT6a. This enables the detection terminal 2 to detect an obstacle (inference process) more quickly.
[0118] [Outline of the second embodiment] Next, an overview of the learning process and the inference process in the second embodiment will be described.
[0119] In the learning process in the first embodiment, a learning model is generated by inputting learning data including partial image data corresponding to each region of interest ROI. In the inference process in the first embodiment, the partial image data included in the image data is input one by one in order to infer identification information corresponding to each partial image data.
[0120] That is, in the inference processing in the first embodiment, when there are multiple regions of interest ROI corresponding to image data, it is necessary to execute the inference processing multiple times by inputting partial image data corresponding to each region of interest ROI.
[0121] In contrast to this, in the learning process in the second embodiment, a learning model is generated that has an output node for each region of interest ROI and for each type of object to be recognized.
[0122] As a result, in the second embodiment, even if there are multiple regions of interest ROI corresponding to the image data, it is possible to infer the type of each object reflected in each of the multiple partial image data included in the image data by inputting the image data itself and performing the inference process only once. Therefore, in the second embodiment, it is possible to reduce the time required to perform the inference process.
[0123] In the following, the region of interest ROI is also called a cell of interest COI (Cell of interest).
[0124] [Details of the second embodiment] Next, the learning process and the inference process in the second embodiment will be described in detail. Fig. 12 to Fig. 14 are flow charts for explaining the learning process and the inference process in the second embodiment in detail. Fig. 15 is a diagram for explaining the learning process and the inference process in the second embodiment in detail. Note that only the points different from the learning process and the inference process in the first embodiment will be described below.
[0125] [Learning process details] First, the learning process in the second embodiment will be described in detail below. Figures 12 and 13 are diagrams for explaining the details of the learning process.
[0126] As shown in FIG. 12, the image acquisition unit 111 waits until the learning timing, for example (NO in S31).
[0127] Then, when the learning timing arrives (YES in S31), the image acquisition unit 111 acquires video data on RGB information (hereinafter also referred to as first video data) from the video data stored in the storage area 110, for example (S32).
[0128] In this case, the image acquisition unit 111 acquires video data on distance information (hereinafter also referred to as second video data) from the video data stored in the storage area 110 (S33).
[0129] That is, the image acquisition section 111 acquires second moving image data obtained by capturing the same area as the first moving image data acquired in the process of S32 at the same timing.
[0130] Then, the image acquisition unit 111 acquires one piece of image data (hereinafter also referred to as first image data) from the first moving image data acquired in the process of S32 (S34).
[0131] Specifically, the image acquisition unit 111 acquires, in chronological order, one image data constituting the first moving image data acquired in the process of S32.
[0132] Furthermore, the image acquisition unit 111 acquires one piece of image data (hereinafter also referred to as second image data) corresponding to the first image data acquired in the process of S34 from the second moving image data acquired in the process of S33 (S35).
[0133] Next, the learning data generation unit 113 waits until it receives input of a correct label (identification information) corresponding to each combination of multiple partial image data corresponding to each fixed position in the first image data obtained in the processing of S34 and the second image data obtained in the processing of S35 (NO in S41).
[0134] Then, when receiving input of a correct label corresponding to each combination (YES in S41), the learning data generation unit 113 generates learning data including each combination and the correct label received in the process of S41 for each combination of the first and second image data acquired in the processes of S34 and S35 (S42). Furthermore, the learning data generation unit 113 stores the generated learning data in the memory area 110, for example.
[0135] In addition, in the process of S42, the learning data generation unit 113 may generate learning data including each first image data and a correct label corresponding to each first image data for each first image data acquired in the process of S34. In other words, the learning data generation unit 113 may generate learning data without using, for example, the second image data acquired in the process of S35.
[0136] Thereafter, the learning data generating unit 113 determines whether or not all of the first image data and the second image data have been acquired in the processes of S34 and S35 (S43).
[0137] As a result, when it is determined that all of the first image data and second image data have not been acquired (NO in S43), the image acquisition unit 111 performs the processes from S34 onwards again.
[0138] On the other hand, if it is determined that all of the first image data and second image data have been acquired (YES in S43), the model generation unit 114 generates a learning model by performing machine learning using the learning data generated in the processing of S42 (S44).
[0139] Specifically, the model generation unit 114 generates a learning model having a number of output nodes determined according to, for example, the following formula (2).
[0140]
number
[0141] In the above formula (2), N o denotes the number of output nodes, and N cell indicates the number of cells of interest COI included in the first image data and the second image data, and N c indicates the number of types of objects to be recognized.
[0142] Here, the partial image data corresponding to each cell of interest COI may not include an object that needs to be recognized. Therefore, when the partial image data does not include an object, it is preferable for the learning model to output information indicating that the object does not exist in order to clearly indicate that the partial image data does not include the object. Therefore, in the above formula (2), 1, which corresponds to the case where the object is not included, is multiplied by N. c The value calculated by adding cell By multiplying with N o is calculated.
[0143] Furthermore, in the process of S44, the model generation unit 114 generates a learning model so that the cross-entropy loss function shown in the following formula (3) is minimized.
[0144]
number
[0145] In the above formula (3), J represents a cross-entropy loss function, σ represents a sigmoid activation function, i represents a variable that identifies each output node, and y represents a value output from each output node.
[0146] The model generation unit 114 may calculate the determination accuracy of the learning model generated in the process of S44, for example, according to the following formula (4).
[0147]
number
[0148] In the above formula (4), Accuracy indicates the judgment accuracy of the learning model, and y ik indicates the correct label corresponding to the kth cell of interest COI contained in the ith image data, and p ik indicates the predicted label (value output from the learning model) corresponding to the kth cell of interest COI included in the i-th image data.
[0149] Then, for example, if the calculated judgment accuracy is below a predetermined threshold, the model generation unit 114 may determine that the judgment accuracy of the learning model is insufficient, and may further perform learning processing, for example, by using new learning data.
[0150] [Details of inference process] Next, the details of the inference process in the second embodiment will be described below. Fig. 14 is a diagram for explaining the details of the inference process.
[0151] As shown in FIG. 14, the image acquisition unit 211 waits until the inference timing, for example (NO in S51).
[0152] Then, when the inference timing arrives (YES in S51), the image acquisition unit 211 acquires image data (hereinafter also referred to as third image data) regarding RGB information captured by the imaging device (S52).
[0153] Furthermore, the image acquisition unit 211 acquires image data (hereinafter also referred to as fourth image data) regarding distance information captured by the imaging device (S53).
[0154] Next, the value acquisition unit 213 acquires multiple values output from each of the output nodes of the learning model (the learning model generated in the processing of S44) in response to the input of the third and fourth image data acquired in the processing of S52 and S53 (S54).
[0155] Specifically, the value acquisition unit 213 acquires a plurality of values for each type of object to be recognized, for each cell of interest COI included in the third and fourth image data acquired in the processes of S52 and S53.
[0156] In addition, in the processing of S44, if the learning model is generated by using learning data including the first image data and the correct label corresponding to the first image data (learning data not including the second image data and the correct label corresponding to the second image data), the value acquisition unit 213 inputs the fourth image data acquired in the processing of S52 to the learning model in the processing of S54.
[0157] Then, the information output unit 214 outputs each of the identification information corresponding to the multiple values acquired in the process of S54 (S55). Specific examples of the process of S54 and the process of S55 will be described below.
[0158] [Specific examples of the processing in S54 and S55] FIG. 15 is a diagram illustrating a specific example of the process of S54 and the process of S55.
[0159] For example, if the number of cells of interest COI is 22, the number of types of objects to be recognized is 9, and the above formula (2) is followed, the number of output nodes included in the learning model will be 220. Hereinafter, the values output from each of these 220 output nodes will be referred to as "C1" to "C 220 " is also called.
[0160] Specifically, in the example shown in FIG. 10 " indicates the value output from the output node corresponding to the first cell of interest COI. 211 " to "C 220 " indicates the value output from the output node corresponding to the 22nd cell of interest COI.
[0161] Furthermore, in the example shown in FIG. 11 ", "C 21 ", "C 211 " indicates a value indicating whether the type of object included in each partial image data is the first type. 12 ", "C 22 ", "C 212 " indicates a value indicating whether the type of object included in each partial image data is the second type. 10 ", "C 20 ", "C 30 ", "C 220 " indicates a value indicating whether or not an object is not captured in each partial image data.
[0162] So, for example, "C1" to "C 10 " are "0", "0", "1", "0", "0", "0", "0", "0", "0", and "0", the information output unit 214 outputs information indicating that the type of object appearing in the partial image data corresponding to the first cell of interest COI is the third type (e.g., "creature"), since the value output from "C3" is "1".
[0163] Also, for example, "C1" to "C 10 " are "0", "0", "0", "0", "0", "0", "0", "1", "0" and "0", the information output unit 214 outputs information indicating that the type of object appearing in the partial image data corresponding to the first cell of interest COI is the eighth type (e.g., "weed"), since the value output from "C8" is "1".
[0164] Furthermore, for example, "C1" to "C 10 When the values output from each of "C" and "D" are "0", "0", "0", "0", "0", "0", "0", "0", "0", and "1", the information output unit 214 outputs "C 10 " is "1", so information is output indicating that no object is captured in the partial image data corresponding to the first cell of interest COI.
[0165] That is, the information output unit 214 outputs, for example, 10 The type corresponding to the output node that outputs "1" out of "1" is output as the type of object appearing in the partial image data corresponding to the first cell of interest COI.
[0166] This enables the information generating unit 214 to output information indicating the type of object reflected in the partial image data corresponding to each cell of interest COI.
[0167] For example, the information output unit 214 outputs the information from "C1" to "C 10 " may be output as the type of the object appearing in the partial image data corresponding to the first cell of interest COI.
[0168] This enables the information output unit 214 to identify and output the type of object reflected in the partial image data corresponding to each cell of interest COI, even if a value other than "0" or "1" is output from each output node.
[0169] In this way, in the second embodiment, even if there are multiple cells of interest COI corresponding to the image data, it is possible to limit the number of times the inference process is executed to one. Therefore, in the second embodiment, it is possible to reduce the time required to execute the inference process. Therefore, in the inference process in the second embodiment, it is possible to suppress an increase in the execution time of the inference process even if, for example, there are a large number of cells of interest COI.
[0170] That is, in the second embodiment, for example, it is possible to reduce the maximum time required to detect an obstacle while the senior car is traveling. Therefore, in the second embodiment, it is possible to prevent failure to detect an obstacle due to the execution of the inference process not being completed while the senior car is traveling. [Explanation of symbols]
[0171] 1: Information processing device 2: Detection terminal 101:CPU 102: Memory 103: Communication interface 104:Storage medium 105: Bus
Claims
1. extracting a plurality of partial learning image data corresponding to a plurality of fixed positions in the learning image data captured by an imaging device from the learning image data; generating a plurality of learning data by adding, to color information in each of the plurality of learning partial image data, identification information indicating a type of object depicted in each of the learning partial image data; generating a learning model by performing machine learning using the plurality of learning data; The process is executed by a computer, The type of the object is a type of an obstacle present on a travel path of the electric cart. A learning program characterized by:
2. In claim 1, In the process of generating the plurality of learning data, the plurality of learning data are generated by adding, for each of the plurality of learning partial image data, the identification information indicating the type of the object shown in each learning partial image data to the color information in each learning partial image data and distance information from the imaging device to the object shown in each learning partial image data. A learning program characterized by:
3. In claim 1, the plurality of fixed locations includes a plurality of first fixed locations having a first size and a plurality of second fixed locations having a second size smaller than the first size; A learning program characterized by:
4. In claim 3, The plurality of learning partial image data corresponding to the plurality of first fixed positions are learning partial image data that show an area closer to the imaging device than the plurality of learning partial image data corresponding to the plurality of second fixed positions. A learning program characterized by:
5. In claim 1, The process of extracting the plurality of learning partial image data, the process of generating the plurality of learning data, and the process of generating the learning model are performed for each of the learning image data included in the video data captured by the imaging device. A learning program characterized by:
6. In claim 1, further comprising: extracting a plurality of detection partial image data corresponding to the plurality of fixed positions in the detection image data captured by an imaging device from the detection image data captured by the imaging device; outputting the identification information indicated by a value output from the learning model in response to input of color information in each of the plurality of partial detection image data; A learning program that causes a computer to execute a process.
7. In claim 6, further comprising: In the process of outputting the identification information, the identification information indicated by a value output from the learning model is output in response to input of the color information in each of the plurality of partial detection image data and distance information from the imaging device to the object shown in each of the partial detection image data. A learning program characterized by:
8. For each of a plurality of pieces of training image data captured by an imaging device, identification information indicating the type of object reflected in the plurality of pieces of training partial image data corresponding to each of the plurality of training image data is added to color information for each of the plurality of pieces of training partial image data corresponding to a plurality of fixed positions in each of the training image data, thereby generating a plurality of pieces of training data corresponding to each of the plurality of training image data; generating a learning model by performing machine learning using the plurality of learning data; The process is executed by a computer, The type of the object is a type of an obstacle present on a travel path of the electric cart. A learning program characterized by:
9. In claim 8, further comprising: For each of a plurality of pieces of detection image data captured by an imaging device, outputting each of the identification information indicated by a plurality of values output from the learning model in response to input of color information for each of a plurality of pieces of detection partial image data corresponding to the plurality of fixed positions in each of the detection image data. A learning program that causes a computer to execute a process.
10. In claim 9, In the process of generating the learning model, the learning model is generated having a plurality of output nodes corresponding to each combination of the plurality of learning partial image data and the type of the object; In the process of outputting each of the identification information, for each of the plurality of partial detection image data included in each of the plurality of detection image data, the identification information corresponding to a value that satisfies a predetermined condition among a plurality of values output from the plurality of output nodes corresponding to each of the partial detection image data is output as the identification information corresponding to each partial detection image. A learning program characterized by:
11. In claim 10, In the process of outputting each of the identification information, for each of the plurality of partial detection image data included in each of the plurality of detection image data, the identification information corresponding to a maximum value among a plurality of values output from the plurality of output nodes corresponding to each of the partial detection image data is output as the identification information corresponding to each partial detection image. A learning program characterized by:
12. a partial image extraction unit that extracts a plurality of learning partial image data corresponding to a plurality of fixed positions in the learning image data captured by an imaging device from the learning image data; a learning data generating unit that generates a plurality of learning data by adding identification information indicating a type of an object reflected in each of the plurality of learning partial image data to color information in each of the plurality of learning partial image data; a model generation unit that generates a learning model by performing machine learning using the plurality of learning data; The type of the object is a type of an obstacle present on a travel path of the electric cart. A learning device characterized by:
13. a learning data generating unit that generates a plurality of learning data corresponding to each of the plurality of learning image data captured by an imaging device by adding, to color information for each of a plurality of learning partial image data corresponding to a plurality of fixed positions in each of the plurality of learning image data, each of the plurality of pieces of identification information that indicate a type of object reflected in the plurality of learning partial image data corresponding to each of the plurality of learning image data; a model generation unit that generates a learning model by performing machine learning using the plurality of learning data; The type of the object is a type of an obstacle present on a travel path of the electric cart. A learning device characterized by:
14. extracting a plurality of partial learning image data corresponding to a plurality of fixed positions in the learning image data captured by an imaging device from the learning image data; generating a plurality of learning data by adding, to color information in each of the plurality of learning partial image data, identification information indicating a type of object depicted in each of the learning partial image data; generating a learning model by performing machine learning using the plurality of learning data; The computer executes the process, The type of the object is a type of an obstacle present on a travel path of the electric cart. A learning method comprising:
15. For each of a plurality of pieces of training image data captured by an imaging device, identification information indicating the type of object reflected in the plurality of pieces of training partial image data corresponding to each of the plurality of training image data is added to color information for each of the plurality of pieces of training partial image data corresponding to a plurality of fixed positions in each of the training image data, thereby generating a plurality of pieces of training data corresponding to each of the plurality of training image data; generating a learning model by performing machine learning using the plurality of learning data; The computer executes the process, The type of the object is a type of an obstacle present on a travel path of the electric cart. A learning method comprising:
Citation Information
Patent Citations
Image processor, information processing method and program
JP2016099734A
Image recognition method, image recognition device, and image recognition program
JP2016186703A
Information processing apparatus, information processing method, and program
JP2018017103A
Recognition device and program
JP2018073308A
Image recognizing and capturing apparatus
JP2018156408A