Method and device for determining vehicle accessibility area for driving video using artificial neural network
An artificial neural network system segments and classifies driving images to automate the generation of boarding area information, addressing the lack of accurate vehicle access area information and enhancing safety and efficiency.
Patent Information
- Application Number
- JP2023540450
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-08
- Filing Date
- 2021-09-08
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2041-09-08
AI Technical Summary
Existing technologies fail to provide drivers and pedestrians with accurate information about vehicle boarding areas, leading to potential accidents and inefficiencies in selecting safe boarding locations.
An artificial neural network-based system that segments driving images into multiple strips, classifies boarding areas, and generates accurate boarding area information using activation maps and image segmentation, enabling automated generation of rideable area information.
Provides safer and more accurate boarding area information, reducing waiting times and enhancing the safety of pedestrian vehicle access by automating the identification of accessible areas.
Smart Images

Figure 0007783279000001 
Figure 0007783279000002 
Figure 0007783279000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and device for determining a vehicle's occupancy area for a driving image using an artificial neural network. More specifically, the present invention relates to a technology for classifying a vehicle's driving image into occupancy areas and non-occupancy areas using an artificial neural network, and then providing the driver with occupancy area information. [Background technology]
[0002] Generally, there are areas where pedestrians can board vehicles on the sidewalk without any obstacles between the roadway on which vehicles travel and the footpath on which pedestrians walk, and areas where pedestrians can board vehicles because there are structures such as fences or trees to prevent accidents. Therefore, when calling a call taxi and boarding a taxi or boarding a friend's vehicle, selecting an area where boarding is possible can be problematic unless accurate geographic information for the area is known.
[0003] Furthermore, if a driver wants to stop on a roadway adjacent to a sidewalk to pick up an acquaintance, and does not know in advance exactly where the vehicle can be boarded, the driver may end up wandering around the area looking for an area where boarding is possible, and an accident may occur if the driver tries to stop on the roadway to pick up an acquaintance even though the area is not one where boarding is possible.
[0004] Therefore, if a driver of a vehicle or a pedestrian intending to board a vehicle could know in advance the location where they can board the vehicle, the problems described above would not occur. However, the reality is that there is no technology that provides drivers with information about areas on the roadway where pedestrians can safely board, or that provides pedestrians with information about areas in the surrounding area where they can board. Summary of the Invention [Problem to be solved by the invention]
[0005] Therefore, the method and apparatus for determining vehicle accessibility areas for driving footage using an artificial neural network according to one embodiment is an invention devised to solve the above-described problems, and relates to a technology that uses an artificial neural network to provide information on areas where pedestrians can access vehicles.
[0006] More specifically, the purpose is to enable pedestrians to board vehicles more safely by using an artificial neural network module to generate information about boarding areas based on input information about the image of a vehicle traveling, and displaying the generated information about boarding areas on the image of the vehicle traveling or providing the driver or pedestrian with information mapped on a map.
[0007] Another purpose is to provide users with map information that maps the results of analyzing the boarding area for pre-recorded driving footage, allowing them to call a taxi based on the boarding area more accurately, or to enable them to share the boarding location with the driver in advance when inviting an acquaintance to board the vehicle.
[0008] Another object of the present invention is to make it easier to produce maps that include rideable area information by providing rideable area information that is automatically generated based on driving video. [Means for solving the problem]
[0009] According to one embodiment, an apparatus for determining a vehicle's boarding area for a driving image using an artificial neural network may include an image segmentation module that acquires a driving image of a vehicle in a driving direction from a camera module and segments the driving image into a plurality of image strips; a pre-trained boarding area classification artificial neural network module that uses the image strips as input information and boarding area information for the image strips as output information; a feature extraction module that extracts an activation map including feature information for the image strip from the boarding area classification artificial neural network module; and an area information generation module that generates boarding area information for the image strip based on the feature information included in the activation map.
[0010] The device for determining a vehicle accessible area for a driving image using an artificial neural network further includes a pre-trained segmentation artificial neural network module that receives the image strip as input information and outputs image segmentation information for the image strip, and the area information generation module can generate accessible area information for the image strip based on the activation map and the image segmentation information.
[0011] The area information generation module may generate a corrected activation map using reliability information for each coordinate included in the image segmentation information based on a representative point of the activation map, and generate boarding area information for the image strip based on the corrected activation map.
[0012] The loss of the segmentation artificial neural network module includes a representative point loss, which is the difference in distance between representative point information in the reference information for the image strip and representative point information in the image segmentation information, and a boundary line loss, which is the difference in distance between boundary line information in the reference information for the image strip and boundary line information in the image segmentation information, and the segmentation artificial neural network module can perform learning in a direction to reduce the loss of the segmentation artificial neural network module.
[0013] The image division module may divide the traveling image into rectangular image strips in which the vertical length of the image strip is longer than the horizontal length.
[0014] According to one embodiment, a method for determining a vehicle boarding area for a driving image using an artificial neural network may include an image division step of acquiring a driving image of a vehicle in a driving direction from a camera module and dividing the driving image into a plurality of image strips; a step of extracting an activation map including feature information for the image strips using a pre-trained boarding availability classification artificial neural network module that uses the image strips as input information and boarding availability information for the image strips as output information; and a step of generating boarding availability area information for the image strips based on the feature information included in the activation map. [Effects of the Invention]
[0015] According to one embodiment of the present invention, a method and apparatus for determining vehicle accessibility areas for driving images using an artificial neural network provides information to vehicle drivers or pedestrians about areas where vehicles can be accessed on the current road, thereby enabling pedestrians to access vehicles more safely.
[0016] In addition, while providing information about a vehicle's boarding area requires passive human intervention in the prior art, the present invention utilizes an artificial neural network to automatically generate boarding area information ahead of a currently moving vehicle, thereby providing safer and more accurate boarding area information.
[0017] In addition, since the location of the boarding area can be accurately known, the waiting time for call taxi service can be significantly reduced.
[0018] In addition, when creating a map including rideable area information, according to conventional technology, the rideable area had to be passively labeled on the map, but according to the present invention, rideable area information can be automatically generated based on acquired driving video, which has the effect of making it easier to create a map including rideable area information. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a block diagram illustrating some components of a device for determining a vehicle access area for a driving video using an artificial neural network according to an embodiment. [Figure 2] 10 is a diagram illustrating input information input to and output information output from an artificial neural network module for classifying boarding availability according to an embodiment; [Figure 3] FIG. 10 illustrates the relationship between an artificial neural network module for boarding availability classification and a feature extraction module according to an embodiment. [Figure 4] FIG. 10 shows activation maps extracted from a feature extraction module for an input image strip. [Figure 5] 10 is a diagram illustrating the relationship between an artificial neural network module 120 for classifying boarding availability, a feature extraction module, and a region information module according to an embodiment. FIG. [Figure 6] FIG. 10 is a diagram illustrating input information input to and output information output from a segmentation artificial neural network module according to an embodiment. [Figure 7] FIG. 10 is a diagram illustrating input information input to and output information output from a segmentation artificial neural network module according to an embodiment. [Figure 8] FIG. 10 is a diagram illustrating the relationship between a boarding availability classification artificial neural network module, a feature extraction module, a region information generation module, and a segmentation artificial neural network module according to an embodiment. [Figure 9] 10A and 10B are diagrams illustrating the principle of how a region information generating module corrects a region using reliability information for each coordinate according to an embodiment. [Figure 10] 10 is a diagram for explaining the process in which the principle described with reference to FIG. 9 is applied to the present invention. [Figure 11] FIG. 10 is a diagram illustrating a method for training a segmentation artificial neural network module according to an embodiment. [Figure 12] 10A and 10B are diagrams illustrating a method for performing learning using output information and reference information in a segmentation artificial neural network module according to an embodiment. [Figure 13] FIG. 10 is a schematic diagram illustrating joint activation map generation of the feature extraction module according to another exemplary embodiment of the present invention. [Figure 14] FIG. 10 is a schematic diagram illustrating the generation of a nulling activation map according to another embodiment of the present invention. [Figure 15] FIG. 10 is a schematic diagram illustrating the generation of a nulling activation map according to another embodiment of the present invention. [Figure 16] FIG. 10 is a schematic diagram illustrating generation of an integrated activation map according to another embodiment of the present invention. [Figure 17] FIG. 10 is a diagram showing a screen on which a boarding area is mapped on a map according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. When assigning reference numerals to components in each drawing, it should be noted that the same components are designated by the same numerals whenever possible, even if they are displayed in different drawings. Furthermore, in describing embodiments of the present invention, if a detailed description of related disclosed configurations or functions is deemed to hinder understanding of the embodiments of the present invention, the detailed description thereof will be omitted. Furthermore, although embodiments of the present invention will be described below, the technical concept of the present invention is not limited thereto, and may be modified and implemented in various ways by those skilled in the art.
[0021] Furthermore, the terms used in this specification are used to describe the embodiments and are not intended to limit and / or restrict the disclosed invention. A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0022] In this specification, the terms "comprises," "includes," "has," and the like are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, but do not preclude the possible presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0023] Furthermore, throughout the specification, when any part is "connected" to another part, this includes not only "directly connected" but also "indirectly connected" with an additional element interposed therebetween, and terms including ordinal numbers such as "first" and "second" used in this specification may be used to describe various components, but the above components are not limited to the above terms.
[0024] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily carry out the invention, and portions not related to the description will be omitted in order to clearly explain the present invention.
[0025] 1 is a block diagram showing some components of a vehicle accessibility area determination device 100 for a driving video using an artificial neural network according to an embodiment of the present invention, and FIG. 2 is a diagram showing input information input to and output information output from an artificial neural network module for classifying whether or not a vehicle is accessible. Hereinafter, for convenience of explanation, the vehicle accessibility area determination device 100 for a driving video using an artificial neural network will be referred to as the vehicle accessibility area determination device 100.
[0026] Referring to FIG. 1, the vehicle accessibility area determination device 100 includes an image segmentation module 110, an accessibility classification artificial neural network module 120, a feature extraction module 130, an area information generation module 140, and a segmentation artificial neural network module 150.
[0027] The image segmentation module 110 collects driving video images 5 of the front and sides of the vehicle captured by a camera, segments the collected driving video images 5 into a plurality of image strips 10, and then transmits the segmented plurality of image strips 10 to the boarding availability classification artificial neural network module 120.
[0028] Specifically, the moving image 5 is divided into a plurality of image strips 10 each having a certain size, and the plurality of image strips 10 may be divided into a plurality of image strips 10 that share overlapping areas as shown in Fig. 2. Conversely, the plurality of image strips 10 may be divided into a plurality of image strips 10 that do not overlap each other.
[0029] Furthermore, when dividing the driving video image 5 to generate the plurality of image strips 10, the image division module 110 may generate the image strips 10 in a rectangular shape in which the vertical length is longer than the horizontal length.
[0030] Generally, obstacles that affect pedestrians' ability to board a vehicle are those that are horizontally long (e.g., flower beds) that have a greater impact than those that are vertically long (e.g., telephone poles). Therefore, accurate analysis of horizontal information can accurately generate boarding availability information. Therefore, the image segmentation module 110 of the present invention segments the image strips 10 into a form in which the vertical length is longer than the horizontal length when generating the multiple image strips 10, thereby more efficiently determining the area in which pedestrians can board a vehicle. This will be described in more detail later.
[0031] The boarding availability classification artificial neural network module 120 is an artificial neural network module that receives the multiple image strips 10 sent by the image segmentation module 110 as input information and outputs boarding availability information 20 for each image strip as output information. The boarding availability classification artificial neural network module 120 includes a learning session 121 that performs learning based on the input information and output information, and an inference session 122 that infers output information based on the input information. Here, classification refers to what is generally called classification in an artificial neural network module.
[0032] The learning session 121 of the boarding availability classification artificial neural network module 120 is a session that learns based on input information and output information, and the inference session 122 analyzes the image strip 10 input in real time using the artificial neural network module, and then outputs boarding availability information 20, which includes information on whether a boarding area exists for each image and its reliability information, as output information.
[0033] For example, as shown in FIG. 2, when five images input sequentially or simultaneously are input to the boarding availability classification artificial neural network module 120, it can output the presence or absence of a boarding availability area and its reliability information for each of the five images.
[0034] For example, the boarding availability classification artificial neural network module 120 may output information indicating the presence of a boarding area and a reliability of 0.6 as first output information 21 for the first image, information indicating the presence of a boarding area and a reliability of 0.7 as second output information 22 for the second image, information indicating the presence of a boarding area and a reliability of 0.9 as third output information 23 for the third image, information indicating the presence of a boarding unavailable area and a reliability of 0.7 as fourth output information 24 for the fourth image, and information indicating the presence of a boarding unavailable area and a reliability of 0.6 as fifth output information 25 for the fifth image.
[0035] 2, the boarding availability classification artificial neural network module 120 is shown as outputting both boarding availability classification information and boarding ineligible classification information, but the present invention is not limited thereto, and the boarding availability classification artificial neural network module 120 may output only boarding availability classification information, or conversely, may output only boarding ineligible classification information. Meanwhile, the specific process and structure of the boarding availability classification artificial neural network module 120 may be borrowed from a conventionally disclosed image classification artificial neural network module.
[0036] The feature extraction module 130 is a module that outputs an activation map 30 containing feature information for the image strip 10. Regarding the feature extraction module 130, FIG. 3 is a diagram illustrating the relationship between the boarding availability classification artificial neural network module 120 and the feature extraction module 130 according to an embodiment of the present invention, and FIG. 4 is a diagram illustrating an activation map extracted from the feature extraction module 130 for the image strip, which is input information.
[0037] As shown in FIG. 3, the feature extraction module 130 refers to a module that outputs an activation map 30 containing feature information in a layer 123 before the fully connected layer 124 and the output layer in a pre-trained image classification artificial neural network including a ConvNet.
[0038] Specifically, the feature extraction module 130 analyzes the input image strip 10 to extract an activation map 30 containing information about regions where an object exists and regions where an object does not exist in the image. Since the boarding availability classification artificial neural network module 120 is an artificial neural network module that outputs boarding availability information 20 for an input image as described above, the information input to the pre-output layer 123 contains information about whether boarding is possible for each region of the image input to the boarding availability classification artificial neural network module 120. Therefore, the feature extraction module 130 can apply various filters published in the pre-output layer 123 to extract an activation map 30 containing boarding availability information for each region.
[0039] For example, when the image shown in FIG. 4(a) is input to the boarding availability classification artificial neural network module 120 as the image strip 10, the feature extraction module 130 can extract the activation map 30 shown in FIG. 4(b) from the fully connected layer 124 and the layer 123 before the output layer.
[0040] In Figure 4(b), white areas indicate areas that are determined to be boardable, and black areas indicate areas that are not, and this information is called feature information. That is, when the feature extraction module 130 applies a filter to the layer 123 to extract the activation map 30 before output, it can classify and display probability information for each area as a white area or a black area based on probability information regarding whether or not a boarding area is possible based on the information input to the layer 123 before output.
[0041] Furthermore, the feature extraction module 130 may display only areas that have a probability above a certain standard in white based on vector information containing position information for each area and probability information, and display areas that do not have a certain standard in black.
[0042] The area information generating module 140 can generate rideable area information for the image strip 10 based on the feature information included in the activation map 30 .
[0043] FIG. 5 is a diagram illustrating the relationship between the boarding availability classification artificial neural network module 120, the feature extraction module 130, and the area information generation module 140 according to an embodiment of the present invention.
[0044] Specifically, the region information generation module 140 may generate boarding area information 40 based on feature information for obstacle areas and non-obstacle areas included in the activation map 30 extracted by the feature extraction module 130, as shown in Fig. 5. As described above, the activation map 30 extracted using a filter includes coordinate information for each area and information on whether or not boarding is possible as vector information, so the region information generation module 140 may determine the boarding area based on such information.
[0045] Specifically, the area information generating module 140 may generate the boarding area information 40 based on information about the white area and the black area shown in Fig. 4(b). As described above, the white area is an area without obstacles, and the black area is an area with obstacles, so the boarding area information 40 may be generated based on the white area. However, even if a white area is free of obstacles, it does not necessarily mean that a passenger can board a vehicle. Therefore, the area information generating module 140 generates the boarding area information 40 for a vehicle based on information about whether the white area is formed to a certain size or larger that allows a person to move, information about whether the white area is connected to the ground, information about whether the white area is connected to a vehicle, etc., and the boarding area information 40 may be generated including coordinate information.
[0046] Therefore, the area information generation module 140 generates image information based on the boarding available area information 40 having coordinate information for boarding availability, and then displays the image information in a separate area on the image strip 10 that is input to the boarding available / unboardable classification artificial neural network module 120. When the boarding available area information 40 is displayed in a separate area on the image strip 10, it is possible to intuitively grasp the area where the driver or passenger can board.
[0047] In addition, after generating the boarding area information 40, the area information generation module 140 may display the boarding area information 40 in an overlapping manner on the image strip 10 as shown in the lower left part of Fig. 5, or may generate the boarding area information 40 as coordinate information as shown in the lower right part of Fig. 5. The coordinate information refers to coordinate information for the boarding area within the image strip 10, and the coordinate information is used as information for mapping the boarding area on a map together with position information in a driving image. Therefore, when the boarding area information 40 is generated as coordinate information as in the present invention, it is advantageous in that it is easier to map the boarding area on a map.
[0048] Additionally, the area information generation module 140 may generate the boarding area information 40 based on the image segmentation information 50 output from the segmentation artificial neural network module 150. The segmentation artificial neural network module 150 will be described in detail below.
[0049] 6 and 7 are diagrams illustrating input information input to and output information output from a segmentation artificial neural network module according to an embodiment.
[0050] The segmentation artificial neural network module 150 is an artificial neural network module that receives as input the multiple image strips 10 sent by the image segmentation module 110 and outputs image segmentation information 50 for each image strip. The segmentation artificial neural network module 150 includes a learning session 151 that performs learning based on the input information and output information, and an inference session 152 that infers output information based on the input information.
[0051] Here, the image segmentation information 50 refers to an image in which an image is classified into several pixel sets, and also refers to information in which coordinates having the same or similar characteristics are grouped into several groups for each coordinate. Generally, regions displayed in the same image refer to regions having similar characteristics such as color, brightness, or material, and regions displayed in different images refer to regions having different characteristics.
[0052] The learning session 151 of the segmentation artificial neural network module 150 is a session that can perform learning based on input information and output information, and the inference session 152 can analyze the image strip 10 input in real time using the artificial neural network module and infer image segmentation information 50 that includes classification information that classifies the objects present in each image.
[0053] Therefore, when an image including a roadway, foot traffic, pedestrians, trees, etc. is input as input information to the segmentation artificial neural network module 150 as shown in FIG. 7, image segmentation information 50 in which the roadway, foot traffic, pedestrians, trees, etc. are clustered and displayed as different images can be output as output information.
[0054] Therefore, the segmentation artificial neural network module 150 includes an artificial neural network capable of outputting image segmentation information for an input image, and a CNN neural network can be representatively used. The segmentation artificial neural network module 150 according to the present invention can be any of the publicly known artificial neural networks that can output the image segmentation information 50 described above for an input image.
[0055] In the present specification, the neural network applied to the segmentation artificial neural network module 150 of the present invention has been described based on CNN. However, the neural network structure applied to the segmentation artificial neural network module 150 of the present invention is not limited to CNN. Various publicly known artificial neural network models useful for image detection, such as Google Mobile Net v2, VGGNet16, and ResNet50, may be applied. The previously described boardability classification artificial neural network module 120 may also use artificial neural networks such as CNN, Google Mobile Net v2, VGGNet16, and ResNet50.
[0056] FIG. 8 is a diagram illustrating the relationship between a boarding availability classification artificial neural network module, a feature extraction module, a region information generation module, and a segmentation artificial neural network module according to an embodiment.
[0057] As shown in FIG. 8, the area information generation module 140 can determine a boarding area using image segmentation information 50 for the activation map 30 extracted from the feature extraction module 130, and display the determined boarding area information 40 in a separate area on the image strip 10.
[0058] In order for rideable area information to be accurately output, feature information, which is information about obstacles and other information determined by the activation map 30, must be accurate. Therefore, the area information generation module 140 can improve the accuracy of the feature information based on the clustered information for each area output from the segmentation artificial neural network module 150.
[0059] Specifically, the area information generation module 140 generates a corrected activation map by performing segmentation correction on the activation map 30 based on the image segmentation information 50 output from the segmentation artificial neural network module 150 and its reliability information, and can determine the boarding area based on the generated corrected activation map.
[0060] As described above, the segmentation artificial neural network module 150 outputs clustered information for each specific object, and since the output information is clustered into different images, the information includes information on whether or not an obstacle exists in each area. Therefore, the region information generation module 140 corrects the activation map based on the clustered information for each area output from the segmentation artificial neural network module 150, and can generate an activation map that accurately includes information on whether or not an obstacle exists and the boarding area for each area. A method for generating an activation map will now be described in detail.
[0061] The segmentation correction process for an activation map will be described with reference to Figures 9 and 10. Figure 9 is a diagram illustrating the principle by which a region information generation module corrects a region using reliability information for each coordinate according to an embodiment, and Figure 10 is a diagram illustrating the process by which the principle described with reference to Figure 9 is applied to the present invention.
[0062] As explained above, segmentation refers to determining the relationship between feature points and surrounding areas and classifying them as the same area if they are determined to be in the same group, and classifying them as a different area if they are not. Generally, as shown in Figure 10(a), after determining a representative point (red point) on an activation map, clustering is performed based on the red point.
[0063] Clustering is generally performed by determining the relationship between a representative point, which is the starting point, and the surrounding area, but in the prior art, it was common to determine whether or not the surrounding area belonged to the same classification group based on the RGB (Red, Green, Blue) values of the representative point and the surrounding area. However, unlike the prior art, the region information generation module 140 according to the present invention has a feature in that it performs a segmentation correction process based on reliability information for each coordinate included in the image segmentation information 50.
[0064] Specifically, if the reliability value in the image segmentation information 50 for a region around a representative point of the activation map 30 is higher than a preset value, the region information generation module 140 groups the region in the same classification group as the representative point and expands its range. However, if the reliability value in the image segmentation information 50 for a region around the representative point is lower than a preset value, the region is determined to belong to a different classification group from the representative point and the range is not expanded. Therefore, the yellow points shown in Figures 10(b) and 10(c) are determined to be regions with high reliability based on the red point, which is the representative point, and are displayed in the same classification group (A), while the purple points are determined to be regions with low reliability based on the representative point and are displayed in a different classification group (B).
[0065] When this method is applied to the activation map 30 as shown in FIG. 10, the points displayed in yellow in FIG. 10 are classified into the same classification group (A) as the representative point and are therefore areas that do not affect the activation map 30. However, the areas displayed in purple are areas that are classified into the same area as the representative point in the activation map extracted from the feature extraction module 130, or areas that are determined to be in a different classification group (B) when classified based on the image segmentation information 50. Therefore, correction of these areas is necessary based on the image segmentation information 50.
[0066] Therefore, the area information generation module 140 generates a corrected activation map in which areas corresponding to purple points, etc. are corrected to black areas rather than white areas, as shown in FIG. 10(c), and based on this, determines the boarding area and displays boarding area information 40 on the image strip 10 as shown in FIG. 8.
[0067] That is, the area information generation module 140 according to one embodiment increases the accuracy of the activation map containing feature information using reliability information contained in the image segmentation information 50, and then generates boarding area information based on the increased accuracy, thereby providing the advantage of being able to generate more accurate boarding area information.
[0068] In addition, since the corrected activation map is generated based on the image segmentation information 50 output from the segmentation artificial neural network module 150, the more successful the learning of the segmentation artificial neural network module 150, the more reliable the boarding area information becomes. Therefore, the segmentation artificial neural network module 150 can perform learning using reference information 60. This will be described below with reference to FIGS. 12 and 13.
[0069] FIG. 11 is a diagram illustrating a learning method of the segmentation artificial neural network module 150 according to an embodiment, and FIG. 12 is a diagram illustrating a learning method of the segmentation artificial neural network module according to an embodiment using output information and reference information.
[0070] Referring to FIG. 11, the segmentation artificial neural network module 150 receives the image strip 10 as input information, outputs image segmentation information 50 including representative point information 51 and boundary line information 52, and uses reference information 60 including actual representative point reference information 61 and actual boundary line reference information 62 as ground truth information. A loss function can be configured to update the parameters of the segmentation artificial neural network module 150 in a direction that reduces the difference between the output information and the reference information.
[0071] Since the accuracy of segmentation depends on the boundary line between the representative point where the clustering starts and a different area, the segmentation artificial neural network module 150 of the present invention can learn along the representative point and the boundary line independently or together.
[0072] Specifically, the loss function of the segmentation artificial neural network module 150 may include a representative point loss function and a boundary line loss function, where the representative point loss function means the difference between representative point information 51 corresponding to the center point of the inferred boarding area and representative point reference information 61 corresponding to the center point of the actual boarding area, and the boundary line loss function means the difference between boundary line information 52 corresponding to the boundary line of the inferred boarding area and boundary line reference information 62 corresponding to the boundary line information 62 of the actual boarding area, and the segmentation artificial neural network module 150 may update the parameters of the segmentation artificial neural network module 150 in a direction that reduces the value of the total loss function, which is the sum of the representative point loss function and the boundary line loss function.
[0073] In FIG. 11, the explanation is based on a loss function that combines the representative point loss function and the boundary line loss function, but the segmentation artificial neural network module 150 may also perform learning based on only one of the representative point loss function or the boundary line loss function.
[0074] Referring to Figure 12, Figure 12(a) is the image segmentation information 50 output from the segmentation artificial neural network module 150 (reliability not shown), Figure 12(b) is a diagram showing the area classified as boardable in the image segmentation information 50 in an actual image, and Figure 12(c) is the actual image serving as reference information.
[0075] The segmentation artificial neural network module 150 generates coordinate information based on the inferred center point C1 of the boarding area as described above, and also generates coordinate information for the actual center point C2 of the boarding area, and can then learn about the center point in the direction that minimizes the value between the generated coordinate information, or can learn in the direction where C1 coincides with C2 based on C2.
[0076] Similarly, the segmentation artificial neural network module 150 generates boundary line information based on the boundary line L1 of the boarding area as a reference, and also generates boundary line information for the boundary line L2 of the actual boarding area, and then learns about the center point in the direction that minimizes the value between the generated information, or learns about the direction that L1 coincides with L2 as a reference.
[0077] 13 is a schematic diagram illustrating the generation of an integrated activation map by the feature extraction module 130 according to another exemplary embodiment of the present invention. As shown in FIG. 13, the feature extraction module 130 according to another exemplary embodiment of the present invention may be configured to extract multiple nulling activation maps 31 using the boarding availability classification artificial neural network module 120 and integrate the extracted multiple nulling activation maps 31 to generate a single activation map 30.
[0078] The nulling activation map 31 refers to an activation map output by applying nulling to a part of the convolution filter (Conv. Filter) of the boarding availability classification artificial neural network module 120. That is, in general, when extracting an activation map using a filter, it means using a filter to which nulling is applied, and applying nulling to a filter means randomly setting the feature to 0 in a part of the filter.
[0079] Therefore, after applying nulling to a part of the convolution filter (Conv. Filter) of the boarding feasibility classification artificial neural network module 120, the artificial neural network module is trained in a direction to reduce the loss between the boarding feasibility information 20 output from the output layer of the boarding feasibility classification artificial neural network module 120 and the actual boarding feasibility information, which is the ground truth.The activation map extracted by the feature extraction module 130 in the Conv. Layer before the FC (Fully Connected Layer) of the boarding feasibility classification artificial neural network module 120 is defined as a nulling activation map 31, and therefore the nulling activation map is an activation map 31 having different vector values.
[0080] To explain in detail the method for generating the nulling activation map 31, FIGS. 14 and 15 are schematic diagrams illustrating the generation of a nulling activation map according to another embodiment of the present invention.
[0081] In the case of Figure 14, in relation to the nulling method, the convolution operation may be performed based on a method of applying random selection in which the stride of the convolution filter is set to 1 and nulling is applied arbitrarily to the convolution filter each time the window slides (the parts marked with x in the figure are the parts to which nulling is applied). When performing a convolution operation by applying nulling through random selection, the parts to be nulled in the filter are not fixed, but the parts to be nulled in the filter change randomly each time the sliding window is performed.
[0082] For example, when the filter window slides for the first time, a convolution operation is performed based on a filter to which a random selection has been applied, as shown by an x in the red filter 51, and thereafter a convolution operation is performed based on a filter to which a random selection has been applied, as shown by an x in the blue filter 52. In other words, the filter slides the feature map, and at this time, a feature is randomly selected for each sliding window.
[0083] Therefore, the first nulling activation map 31a can be extracted by applying nulling in the form of a red filter to the convolution filter (Conv. Filter) of the boarding availability classification artificial neural network module 120, and the second nulling activation map 31b can be extracted by applying nulling in the form of a blue filter to the convolution filter (Conv. Filter) of the boarding availability classification artificial neural network module 120. Therefore, by repeating these steps, the feature extraction module 130 can extract multiple nulling activation maps 31.
[0084] In performing the convolution operation, as another embodiment of the present invention, a randomly selected feature map 12 may be applied to an enlarged image strip 11 obtained by enlarging an input image strip 10, and then the convolution operation may be performed based on the enlarged image strip 11, as shown in Fig. 15. In Fig. 15, the enlarged image strip 11 refers to an image strip enlarged by multiplying the size of the input image strip 10 by s, which is the number of strides applied when performing the convolution operation.
[0085] In comparison with Figure 14, in Figure 14, when performing the convolution operation, random selection is performed on the filter itself while sliding the filter (the stride is 1 in Figure 14), while in Figure 15, a random selection feature map 12 corresponding to the size of the enlarged image strip 11 is applied to the enlarged image strip 11 (the part marked x in Figure 15 shows the state where random selection is applied), and then the convolution operation can be performed on the applied image.
[0086] That is, in FIG. 14, random selection is applied while the filter performs a sliding window on the image strip 10, while in FIG. 15, pre-randomly selected features are expanded and enlarged to fit the size of the image strip 10, and then a convolution operation can be performed by the filter.
[0087] Therefore, in the case of FIG. 15, after applying nulling to a part of the convolution layer (Conv. Layer) of the boarding feasibility classification artificial neural network module 120, the boarding feasibility classification artificial neural network module 120 is trained in a direction that reduces the loss between the boarding feasibility information 20 output from the output layer (output layer) of the boarding feasibility classification artificial neural network module 120 and the actual boarding feasibility information, which is the ground truth, and the feature extraction module 130 can be configured to extract the nulling activation map 31 in the Conv. Layer before the FC (Fully Connected Layer).
[0088] When the convolution operation is performed in the manner shown in Figure 15, the extracted result is the same as that shown in Figure 14, but since random selection is not required while sliding and the convolution operation with random selection applied is performed once for the entire enlarged image, the nulling activation map 31 can be generated more quickly than the method shown in Figure 14. That is, in the case of Figure 14, the sliding window must be randomly extracted by performing convolution, but if this is to be realized by a separate processor, it is difficult to realize this in a general deep learning framework (e.g., Pytorch, Tensorflow). However, when the convolution operation is performed in the manner shown in Figure 15, the convolution operation with random selection applied can be performed relatively quickly, resulting in an effect of shortening the overall process of the processor.
[0089] In addition, although not shown in the drawings, when applying dropout to a feature map extracted after convolution, the present invention can apply dropout in a manner that does not remove features corresponding to the center of each sliding window, instead of a general dropout method. General dropout has the disadvantage of making it difficult to grasp the correlation between each sliding window, and the efficiency of the artificial neural network decreases because spatial information is mixed in when the artificial neural network performs learning. However, when dropout is performed based on a method that includes a feature corresponding to the center, information about the center is preserved, allowing the artificial neural network module to properly learn based on spatial information.
[0090] The integrated activation map 32 is an activation map in which multiple nulling activation maps 31 extracted by the feature extraction module 130 are combined into a single matrix. Figure 16 is a schematic diagram illustrating the generation of the integrated activation map 32 according to another embodiment of the present invention. As shown in Figure 16, the feature extraction module 130 can be configured to generate the integrated activation map 32 by combining the extracted multiple nulling activation maps 31 into a single matrix.
[0091] The integrated activation map 32 generated in this manner is input to the area information generation module 140, and as described above with reference to FIG. 5, the area information generation module 140 determines a boarding area on the image strip 10 based on information regarding areas with obstacles and areas without obstacles included in the integrated activation map 32, and can display the determined boarding area information 40 in a separate area on the image strip 10.
[0092] In another embodiment, the region information generation module 140 may determine the boarding area using the image segmentation information 50 for the integrated activation map 32 extracted from the feature extraction module 130, as described with reference to Fig. 8, and display the determined boarding area information 40 in a separate area on the image strip 10. The method by which the region information generation module 140 determines the boarding area using the image segmentation information 50 has been described above, and therefore will not be described again.
[0093] FIG. 17 is a diagram showing an embodiment in which boarding area information according to an embodiment of the present invention is displayed on a map.
[0094] In the present invention, information about a rideable area on a road generated based on a vehicle driving video is mapped onto a map, and information about the nearest rideable area based on the user's current location can be provided to the user.
[0095] Specifically, the area information generation module 140 can generate boarding area information 70 by mapping the coordinate information of the boarding area information 40 on a map based on the boarding area information 40 and the current position information and driving direction information in the driving image.
[0096] Therefore, the vehicle boarding area determination device 100 according to one embodiment of the present invention can determine the current location of the user terminal through an application module installed in the user terminal, and control the application module installed in the user terminal so that the nearest boarding area information 70 based on the current location of the user terminal is displayed on the map of the application module of the user terminal.
[0097] As an example, if it is determined via the user terminal that the current user location is at the location indicated by the symbol 80 in FIG. 17, the vehicle boarding area determination device 100 can control the application module so that the boarding area information 70 closest to the user terminal is displayed on the map of the application module installed on the user terminal, as shown in FIG. 17.
[0098] According to this, when a user calls a taxi, the user can call the taxi based on the boarding area information 70, or when the user needs to board a friend's vehicle, the location of the boarding area information 70 can be sent to the driver, which has the effect of allowing the user to board the vehicle more safely.
[0099] So far, a method and apparatus for determining a vehicle access area for a driving image using an artificial neural network according to an embodiment have been described in detail.
[0100] According to an embodiment of the present invention, a method and apparatus for determining a vehicle accessibility area for a driving video using an artificial neural network provides a driver or pedestrian with information about an area where a vehicle can be accessed on a current road, thereby enabling pedestrians to access vehicles more safely. While providing information about a vehicle accessibility area requires manual human intervention in the prior art, the present invention utilizes an artificial neural network to generate information about an accessibility area ahead of a currently moving vehicle in an automated manner, thereby providing safer and more accurate information about the accessibility area. This allows users to accurately determine their location relative to the accessibility area, significantly reducing waiting times for call taxi services.
[0101] In addition, when creating a map including boarding area information, according to the prior art, it was necessary to manually label the boarding area on the map, but according to the present invention, boarding area information is automatically generated based on the acquired driving video, which has the effect of making it easier to create a map including boarding area information.
[0102] However, components, units, modules, elements, etc. described herein as "modules" may be shared or separate, but may be implemented separately as interoperable logic devices. Depictions of different characteristics for modules, units, etc. are intended to highlight different functional embodiments and do not necessarily imply that they must be implemented by separate hardware or software components. Rather, functionality associated with one or more modules or units may be performed by separate hardware or software components or integrated within shared or separate hardware or software components.
[0103] A computer program (also known as a program, software, software application, script or code) may be written in any form of programming language, including compiled or interpreted languages, and a priori or procedural languages, and may take any form, including independent programs, modules, components, subroutines, or other units suitable for use in a computing environment.
[0104] Additionally, the logic flow and structural block diagrams described in this patent document describe corresponding functions supported by the disclosed structural means, corresponding acts supported by the steps, and / or specific methods, and may be used to construct corresponding software structures and algorithms and their equivalents.
[0105] The processes and logic flows described herein are executable by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output.
[0106] The written description sets forth the best mode of the invention and provides examples to explain the invention and to enable one skilled in the art to make and use the invention. The specification so written is not intended to limit the invention to the specific terms set forth.
[0107] Although the present invention has been described above with reference to preferred embodiments, it will be understood by those skilled in the art or those having ordinary knowledge in the art that various modifications and changes can be made to the present invention without departing from the spirit and technical scope of the present invention as set forth in the claims below. Therefore, the technical scope of the present invention should not be limited to the contents described in the detailed description of the specification, but should be determined by the claims.
Claims
1. an image division module that acquires a driving image in a driving direction of the vehicle from the camera module and divides the driving image into a plurality of image strips; a pre-trained boarding availability classification artificial neural network module that receives the image strip as input information and outputs boarding availability information for the image strip; a feature extraction module for extracting an activation map containing feature information for the image strip in the boarding availability classification artificial neural network module; a pre-trained segmentation artificial neural network module that receives the image strip as input information and generates image segmentation information for the image strip as output information; an area information generating module that generates boarding area information for the image strip based on the feature information and the image segmentation information included in the activation map; A vehicle accessibility area determination device for a driving video using an artificial neural network, comprising:
2. 2. The device for determining a vehicle accessible area for a driving image using an artificial neural network according to claim 1, wherein the area information generation module generates a corrected activation map using reliability information for each coordinate included in the image segmentation information based on a representative point of the activation map, and generates accessible area information for the image strip based on the corrected activation map.
3. The loss of the segmentation artificial neural network module is a representative point loss, which is a difference in distance between representative point information in the reference information for the image strip and representative point information in the image segmentation information, and a boundary line loss, which is a difference in distance between boundary line information in the reference information for the image strip and boundary line information in the image segmentation information; 2. The device for determining a vehicle access area for a driving video using an artificial neural network according to claim 1, wherein the segmentation artificial neural network module performs learning in a direction to reduce a loss of the segmentation artificial neural network module.
4. 2. The device for determining a vehicle-ridable area for a driving image using an artificial neural network according to claim 1, wherein the image division module divides the driving image into rectangular image strips whose vertical length is longer than their horizontal length.
5. an image division step of acquiring a driving image in a driving direction of the vehicle from a camera module and dividing the driving image into a plurality of image strips; extracting an activation map including feature information from the image strip using a pre-trained boarding availability classification artificial neural network module, which uses the image strip as input information and boarding availability information for the image strip as output information; obtaining image segmentation information for the image strip through a pre-trained segmentation artificial neural network module that receives the image strip as input information; generating rideable area information for the image strip based on the feature information and the image segmentation information included in the activation map; A method for determining a vehicle-ridable area for a driving video using an artificial neural network, comprising:
Citation Information
Patent Citations
Image recognition device capable of changing arrangement and combination of windows used in image recognition in accordance with configuration information
JP2016133878A
System for determining stop place included in captured image, method, and program
JP2019079149A