Method and device for predicting three-dimensional occupancy state by using neural network model, and computer readable recording medium
By using a neural network model that learns through multiple tasks and utilizes multi-view image inputs, the problem of lack of distance information in camera images is solved, enabling accurate prediction of 3D occupancy status and improving the accuracy of autonomous driving systems.
Patent Information
- Application Number
- CN202380099470.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-01
- Filing Date
- 2023-10-31
- Publication Date
- 2026-01-13
AI Technical Summary
Existing technologies struggle to accurately predict 3D occupancy status from camera images, and lack direct distance information.
A neural network model employing a multi-task learning approach extracts two-dimensional image features from multi-view image inputs and converts them into three-dimensional space to predict occupancy status.
It enables accurate extraction of 3D information from camera images and low-error prediction of occupancy status, thereby improving the accuracy of autonomous driving systems.
Smart Images

Figure CN121336244A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method, apparatus, and computer-readable recording medium for predicting three-dimensional occupancy status using a neural network model. Background Technology
[0002] In autonomous driving systems, predicting the occupancy status around the vehicle based on images acquired from one or more sensors can be an important problem.
[0003] At this point, camera sensors offer the following advantages: low setup cost, and the ability to capture high-resolution images that include detailed features such as the color and texture of objects. However, since camera images do not directly provide distance information, predicting occupancy status based on camera images can be a challenging task.
[0004] Therefore, in order to accurately predict the 3D occupancy status from real-time acquired camera images, a robust neural network model is needed that can accurately extract 3D information from 2D images and predict the occupancy status based on the 3D information with low error.
[0005] The aforementioned background technology refers to technical information that the inventors possessed or obtained during the derivation of this invention, and is not necessarily publicly known technology disclosed to the general public before the application for this invention. Summary of the Invention
[0006] The problem the invention aims to solve
[0007] According to some embodiments of this disclosure, a method, apparatus, and computer-readable recording medium for predicting three-dimensional occupancy states using a neural network model are intended to be provided. The problems to be solved by this invention are not limited to those described above; other problems and advantages not mentioned in this invention can be understood through the following description, and the invention will be more clearly understood through the embodiments. Furthermore, it can be seen that the problems and advantages to be solved by this invention can be achieved by the means and combinations thereof pointed out in the claims.
[0008] means for solving problems
[0009] As a technical means to solve the above-mentioned technical problems, a first aspect of this disclosure provides a method for predicting occupancy status using a neural network model. The method includes: acquiring multiview images from a camera; and training the neural network model by using the multiview images as input data to the model, so that the neural network model outputs a three-dimensional prediction result related to the occupancy status using the multiview images. The training step of the neural network model includes: training the neural network model to perform a first task, a second task, and a third task, wherein the first task includes extracting two-dimensional image features for each of the multiview images, the second task includes converting the two-dimensional image features into a three-dimensional space, and the third task includes predicting the occupancy status in the three-dimensional space.
[0010] A second aspect of this disclosure provides an apparatus for predicting occupancy status using a neural network model, the apparatus comprising: a memory storing at least one program; and a processor executing the at least one program. The processor is configured to: acquire multi-view images from a camera, and to train the neural network model by using the multi-view images as input data to the neural network model, so that the neural network model outputs a three-dimensional prediction result related to the occupancy status using the multi-view images.
[0011] A third aspect of this disclosure is to provide a computer-readable recording medium having a program recorded thereon for executing the method according to the first aspect on a computer.
[0012] In addition, other methods, other systems for implementing the present invention, and computer-readable recording media storing computer programs for performing said methods may also be provided.
[0013] Other aspects, features, and advantages besides those described above will become apparent from the following drawings, claims, and detailed description of the invention.
[0014] Invention Effects
[0015] According to one embodiment of the present disclosure, a robust neural network model can be trained to accurately extract three-dimensional information from a two-dimensional image and predict occupancy status based on the three-dimensional information with low error.
[0016] According to one embodiment of this disclosure, a neural network model can be used to accurately predict the three-dimensional occupancy status from real-time acquired camera images. Attached Figure Description
[0017] Figure 1 This is a diagram illustrating an example of a method for predicting the occupancy status around a vehicle according to an embodiment of an autonomous driving system.
[0018] Figure 2 and Figure 3 This is a diagram illustrating an autonomous driving method according to one embodiment.
[0019] Figure 4 This is a diagram illustrating an example of a method for enabling a neural network model to learn according to one embodiment.
[0020] Figure 5 This is a diagram illustrating another example of a method for enabling a neural network model to learn according to one embodiment.
[0021] Figure 6 This is a diagram illustrating yet another example of a method for enabling a neural network model to learn according to one embodiment.
[0022] Figure 7 This is a diagram illustrating yet another example of a method for enabling a neural network model to learn according to one embodiment.
[0023] Figure 8 This is a diagram illustrating yet another example of a method for enabling a neural network model to learn according to one embodiment.
[0024] Figure 9 This is a diagram illustrating an example of a method for predicting occupancy status using a neural network model according to an embodiment.
[0025] Figure 10 This is a flowchart illustrating an example of a method for predicting occupancy status using a neural network model according to an embodiment.
[0026] Figure 11 This is a configuration diagram illustrating an example of the internal structure of a device for predicting occupancy status using a neural network model according to an embodiment. Detailed Implementation
[0027] According to one embodiment of this disclosure, a method for predicting occupancy status using a neural network model may include: acquiring multi-view images from a camera; and training the neural network model by using the multi-view images as input data, so that the neural network model outputs a three-dimensional prediction result related to the occupancy status using the multi-view images. The training step may include: performing a first task, a second task, and a third task by training the neural network model, wherein the first task includes extracting two-dimensional image features for each of the multi-view images, the second task includes converting the two-dimensional image features into a three-dimensional space, and the third task includes predicting the occupancy status in the three-dimensional space.
[0028] The advantages and features of the present invention, as well as methods for implementing them, will become apparent from the accompanying drawings and detailed description of the embodiments. However, the invention is not limited to the embodiments presented below, but can be implemented in various different forms, and should be understood to include all modifications, equivalents, and substitutions within the spirit and scope of the invention. The embodiments described below are intended to complete the disclosure of the invention, and the invention is provided to fully inform those skilled in the art of its scope. In describing the invention, detailed descriptions of known techniques will be omitted if it is determined that such specific descriptions might obscure the gist of the invention.
[0029] The terminology used in this application is for illustrative purposes only and is not intended to limit the invention. Unless expressly stated herein, the singular form should be understood to include the plural form. It should be understood that in this application, terms such as "comprising" or "having" are intended to indicate the presence of features, numbers, steps, actions, constituent elements, components, or combinations thereof described in the specification, and do not preclude the presence or additional possibilities of one or more other features, numbers, steps, actions, constituent elements, components, or combinations thereof.
[0030] Some embodiments of this disclosure can be represented by functional block configurations and various processing steps. Some or all of these functional blocks can be implemented in any number of hardware and / or software configurations performing a specific function. For example, the functional blocks of this disclosure can be implemented by more than one microprocessor or by a circuit configuration for a predetermined function. Furthermore, for example, the functional blocks of this disclosure can be implemented in various programming or scripting languages. The functional blocks can be implemented by algorithms running on more than one processor. In addition, this disclosure can employ conventional techniques for electronic environment setup, signal processing and / or data processing, etc. Terms such as “mechanism,” “element,” “device,” and “configuration” can be used broadly and are not limited to mechanical and physical configurations.
[0031] Furthermore, the connecting lines or connecting members between the components shown in the figures are merely illustrative of functional connections and / or physical or electrical connections. In actual devices, the connections between multiple components can be represented by various alternative or additional functional connections, physical connections, or electrical connections.
[0032] In this disclosure, "vehicle" can refer to any type of means of transport that has a power unit and is used to transport people or goods, including automobiles, buses, motorcycles, electric scooters or trucks.
[0033] In this disclosure, a "neural network model" refers to an algorithmic technique capable of autonomously classifying and / or learning features of input data. Its constituent techniques utilize machine learning algorithms such as deep learning to simulate the cognitive and judgmental functions of the human brain. The neural network model in embodiments of this disclosure may, for example, have a deep neural network structure. The neural network model can learn using training data based on one or more nodes and the operational rules between nodes. The structure of nodes, the structure of layers, and the operational rules between nodes may vary depending on the embodiment. The neural network model may include one or more hardware resources such as processors, memory, registers, addition processing units, parallel processing units, or multiplication processing units, and cause these hardware resources to operate according to a set of parameters applied to each hardware resource. Therefore, the processor used to cause the neural network model to operate can perform tasks for allocating hardware resources or resource management processing for each action of the neural network model. The neural network model may, for example, have structures such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and Long Short-Term Memory (LSTM).
[0034] The present embodiment will now be described in detail with reference to the accompanying drawings. However, the embodiment may be implemented in various different forms and is not limited to the examples described herein.
[0035] Figure 1 This is a diagram illustrating an example of a method for predicting the occupancy status around a vehicle according to an embodiment of an autonomous driving system.
[0036] Reference Figure 1 The autonomous driving system 100 may include a vehicle 110 equipped with an autonomous driving device (not shown).
[0037] The autonomous driving device may include various sensors used by the autonomous driving system 100 to identify the environment surrounding the vehicle 110. For example, the autonomous driving device may include various sensors such as LiDAR, cameras, and ultrasonic sensors. The sensors can collect information required for autonomous driving, such as the position of the vehicle 110, the position and speed of surrounding vehicles, the driving path, and the position of surrounding objects.
[0038] The autonomous driving device may include computing hardware that drives software for controlling the vehicle 110. For example, the autonomous driving device may identify the environment or conditions around the vehicle 110 based on data collected from sensors, determine the behavior of the vehicle 110, and drive the software for controlling the vehicle 110.
[0039] In this disclosure, an autonomous driving device can predict the occupancy status around the vehicle 110 based on images acquired from one or more sensors.
[0040] More specifically, the autonomous driving device according to embodiments of the present disclosure can utilize a neural network model to obtain a three-dimensional prediction result 130 related to the occupancy state from multi-view images 120 acquired by a camera.
[0041] Camera sensors offer several advantages: low setup costs and the ability to capture detailed features such as color and texture of objects in high-resolution images. However, since camera images do not directly provide distance information, predicting occupancy status based on camera images can be a challenging task.
[0042] Therefore, in order to accurately predict the 3D occupancy status from real-time acquired camera images, a robust neural network model is needed that can accurately extract 3D information from 2D images and predict the occupancy status based on the 3D information with low error.
[0043] To achieve the above objectives, this disclosure provides a method for learning a neural network model using a multi-task learning approach, and for obtaining a three-dimensional prediction result 130 related to the occupancy state using the learned neural network model.
[0044] For example, according to the multi-task learning method, multi-view images acquired from a camera can be used as input data for learning a neural network model, which can then learn to perform multiple tasks by utilizing the multi-view images.
[0045] The neural network model according to embodiments of the present disclosure learns to perform multiple tasks simultaneously, thereby enabling the sharing and utilization of acquired information during the learning process to learn a general representation and prevent overfitting for individual tasks.
[0046] On the other hand, according to embodiments of this disclosure, the specific method for enabling the neural network model to learn will refer to the following. Figures 4 to 8 Please provide an explanation.
[0047] Figure 2 and Figure 3 This is a diagram illustrating an autonomous driving method according to one embodiment.
[0048] Reference Figure 2 According to an embodiment of the present invention, an autonomous driving device can be installed in a vehicle to realize an autonomous driving vehicle 10. The autonomous driving device installed on the autonomous driving vehicle 10 may include various sensors for collecting information about the surrounding environment. As an example, the autonomous driving device can detect the movement of a vehicle 20 traveling in front of it by means of an image sensor and / or an event sensor installed in front of the autonomous driving vehicle 10. The autonomous driving device may also include sensors for detecting the front of the autonomous driving vehicle 10, as well as other vehicles 30 traveling in the adjacent lane and pedestrians around the autonomous driving vehicle 10.
[0049] According to one embodiment, the autonomous driving device may include multiple cameras for capturing images of the surroundings of the autonomous vehicle 10 from multiple angles. For example, the multiple cameras may be respectively configured at the front, sides, and rear of the autonomous vehicle 10, thereby generating images of the surroundings of the autonomous vehicle 10 from different viewpoints or angles. Furthermore, the autonomous driving device may acquire multi-view images related to the surroundings of the autonomous vehicle 10 based on the images captured by the multiple cameras.
[0050] On the other hand, at least one of the sensors used to collect information about the situation around autonomous vehicles, such as Figure 1 As shown, it can have a specified field of view (FoV). For example, when a sensor mounted in front of the autonomous vehicle 10 has such... Figure 1 At the field of view (FoV) shown, the information detected from the center of the sensor may be of relatively high importance. This is because most of the information detected from the center of the sensor likely includes information related to the motion of the vehicle 20 ahead.
[0051] The autonomous driving device can process information collected by the sensors of the autonomous vehicle 10 in real time to control the movement of the autonomous vehicle 10, while storing at least a portion of the information collected by the sensors in a memory device.
[0052] ReferenceFigure 3 The autonomous driving device 40 may include a sensor unit 41, a processor 46, a memory system 47, and a vehicle control module 48, etc. The sensor unit 41 may include multiple sensors (reference numerals 42 to 45), which may include image sensors, event sensors, light sensors, GPS devices, acceleration sensors, etc.
[0053] Data collected by multiple sensors (reference numerals 42 to 45) can be transmitted to processor 46. Processor 46 can store the data collected by the multiple sensors (reference numerals 42 to 45) in memory system 47, and control vehicle control module 48 based on the data collected by the multiple sensors (reference numerals 42 to 45) to determine the movement of the vehicle. Memory system 47 may include two or more memory devices and a system controller for controlling the memory devices. Each memory device may be provided as a single semiconductor chip.
[0054] In addition to the system controller of the memory system 47, the memory devices included in the memory system 47 may each include a memory controller, and the memory controller may include artificial intelligence (AI) computing circuitry such as neural networks. The memory controller can assign predetermined weights to data received from multiple sensors (reference numerals 42 to 45) or processor 46 to generate computational data and store the computational data in the memory chip.
[0055] Figure 4 This is a diagram illustrating an example of a method for enabling a neural network model to learn according to one embodiment.
[0056] The apparatus for predicting occupancy status using a neural network model according to embodiments of the present disclosure (hereinafter referred to as the "apparatus") enables the neural network model to learn so that the neural network model outputs a three-dimensional prediction result related to the occupancy status using multi-view images.
[0057] According to one embodiment, the device can acquire multi-view images 410 from a camera and use the multi-view images 410 to construct a learning dataset for a neural network model 420 to learn from. For example, the device can use the multi-view images 410 as input data to the neural network model 420 and use the neural network model 420 to take the data that is expected to be output for each task as the ground truth data corresponding to the input data.
[0058] On the other hand, although Figure 1 The text describes devices that use neural network models to predict occupancy states as autonomous driving devices, but it is not limited to this. As an example, Figure 4 The device can be with Figure 1The same as the autonomous driving device, or it may be a device included in an autonomous driving device. As another example, Figure 4 The device can be an external server or external device with its own computing power, installed outside the vehicle 110, for exchanging information with the autonomous driving device. For example, the device can acquire multi-view images from a camera installed on the vehicle 110 to predict the occupancy status and transmit the prediction results to the autonomous driving device.
[0059] Reference Figure 4 The device enables the neural network model 420 to learn so that the neural network model 420 can extract two-dimensional image features 430 for each multi-view image 410 as its first task.
[0060] According to one embodiment, the device can be configured such that the neural network model 420 includes an image encoder module. For example, the image encoder can take at least one image from a multi-view image as input data. Furthermore, the image encoder can extract features from the two-dimensional multi-view images and generate a two-dimensional feature map. For example, the image encoder can be constructed using a Residual Network (ResNet) and a Feature Pyramid Network (FPN).
[0061] Figure 5 This is a diagram illustrating another example of a method for enabling a neural network model to learn according to one embodiment.
[0062] Reference Figure 5 The first task may include receiving, in addition to receiving multi-view images 510, receiving lidar point clouds 520 acquired from a lidar sensor, and outputting two-dimensional prediction results 54 related to the occupancy status.
[0063] In other words, the device can utilize not only the multi-view image 510, but also the lidar point cloud 520 as input data for the neural network model 530 to learn, and can enable the neural network model 530 to learn so that the neural network model 530 can perform the first task of using the multi-view image 510 and the lidar point cloud 520 to input a two-dimensional prediction result 540 related to the occupancy state.
[0064] The lidar point cloud 520 may include information acquired when the lidar sensor emits light towards an object. For example, the lidar point cloud 520 may include three-dimensional information related to the surrounding environment, such as the distance to the object, the height of the object, and the intensity of reflection, which can be obtained using the light reflected back from the object.
[0065] According to one embodiment, the device can be configured such that the neural network model 530 includes a prediction module. For example, the prediction module can receive at least one image from a multi-view image and a LiDAR point cloud 520, and output a two-dimensional prediction result 540 related to the occupancy status.
[0066] According to one embodiment, the ground truth data corresponding to the two-dimensional prediction result 540 can be an image that projects the LiDAR points 520 onto a multi-view image. For example, the device can calculate the loss between the two-dimensional prediction result 540 and the projection image as the loss for the first task.
[0067] Figure 6 This is a diagram illustrating yet another example of a method for enabling a neural network model to learn according to one embodiment.
[0068] Reference Figure 6 The device enables the neural network model 620 to learn so that the neural network model 620 can perform a second task based on the multi-view image 610, converting it into a three-dimensional space 630.
[0069] According to one embodiment, the second task may include the step of measuring depth information 640 on a two-dimensional image. In other words, the device can enable a neural network model 620 to learn and perform the second task of measuring depth information 640 on a two-dimensional image, in addition to converting the multi-view image 610 into a three-dimensional space 630.
[0070] According to one embodiment, the apparatus can be configured such that the neural network model 620 includes a view transformer module. For example, the view transformer may refer to a module that uses two-dimensional image features (e.g., image features for multi-view images) obtained from an image encoder and camera parameter information to measure depth information 640 of the two-dimensional image and estimate context information. Furthermore, the view transformer can perform voxel pooling operations based on the depth information 640 and the context information to convert the two-dimensional multi-view image into a three-dimensional space 630.
[0071] According to one embodiment, the ground truth data of the depth information 640 measured by the neural network model 620 can be ground truth depth information. For example, the device can calculate the loss between the measured depth information 640 and the ground truth depth information as a loss related to the second task described above.
[0072] Figure 7 This is a diagram illustrating yet another example of a method for enabling a neural network model to learn according to one embodiment.
[0073] Reference Figure 7 The device enables the neural network model 720 to learn so that the neural network model 720 can perform a third task based on multi-view images 710 to predict the three-dimensional spatial occupancy state 730.
[0074] According to one embodiment, the third task may include: the step of extracting three-dimensional features for a three-dimensional space; and the step of predicting whether a specified object is occupied based on the three-dimensional features, and classifying the three-dimensional space according to the label of the specified object.
[0075] In other words, the third task performed by the neural network model 720 can be to determine which regions in three-dimensional space a given object occupies and to classify those regions. For example, when no object occupies a region in three-dimensional space, the neural network model 720 can classify that region into the first group of categories. As another example, when a given object occupies a region in three-dimensional space, the neural network model 720 can classify that region into the second group of categories based on the object's label. For example, the first group of categories could include "Free Class." Furthermore, the second group of categories could include categories such as "barrier," "bicycle," "bus," "car," "motorcycle," "sidewalk," "pedestrian," "drivable road," "non-drivable road," and "vegetation."
[0076] According to one embodiment, the device can be configured such that the neural network model 720 includes a 3D encoder module and an occupancy prediction head. For example, the 3D encoder module may refer to a module that extracts 3D features from a 3D space obtained from a view transformer. For example, the 3D encoder module may include a 3D residual network (ResNet3D) module and a 3D feature pyramid network (FPN 3D) module. Furthermore, the occupancy prediction head may refer to a module that predicts whether a specified object is occupied based on 3D features and performs category classification in the 3D space based on the label of the specified object.
[0077] According to one embodiment, the neural network model 720 can use ground truth class data for the categories obtained from classification in three-dimensional space. For example, the device can calculate the loss between the classified categories and the ground truth class as the loss related to the third task mentioned above.
[0078] Figure 8 This is a diagram illustrating yet another example of a method for enabling a neural network model to learn according to one embodiment.
[0079] As described above, the apparatus according to an embodiment of the present disclosure enables a neural network model to learn so that the neural network model performs a first task, a second task, and a third task, and can calculate a loss associated with the output of each of the first task, the second task, and the third task.
[0080] According to one embodiment, the device can enable a neural network model to learn based on a first loss, a second loss, and a third loss calculated respectively for a first task, a second task, and a third task. For example, the device can calculate a total loss based on the first loss, the second loss, and the third loss, and enable the neural network model to learn in a direction that reduces the total loss. For example, the device can assign weights to each of the first loss, the second loss, and the third loss and sum them to calculate the total loss.
[0081] On the other hand, refer to Figure 8 The first loss may include a loss calculated based on a two-dimensional prediction result 830 related to the occupancy state. In other words, the device can input the multi-view image 810 as input data into the neural network model 820, obtain the two-dimensional prediction result 830 related to the occupancy state from the neural network model 820 as output data, and calculate the first loss based on the two-dimensional prediction result 830.
[0082] Furthermore, the second loss may include a loss calculated based on the depth measurement 840 of the multi-view image. For example, the device may input the multi-view image 810 as input data into the neural network model 820, obtain the depth measurement 840 of the multi-view image 810 as output data from the neural network model 820, and calculate the second loss based on the depth measurement 840.
[0083] According to one embodiment, the third loss may include a loss 850 calculated using class balance. More specifically, the loss associated with the third task may include a loss calculated based on the classification of categories in three-dimensional space and the frequency of those categories. For example, the learning dataset utilized when the neural network model 820 learns to perform the third task may be a class-imbalanced dataset. In other words, regions included in the multi-view image 810 that are typically classified as "road class" or "free class" may be more numerous than regions classified as "bicycle class" or "motorcycle class." Therefore, when calculating the loss associated with the third task, the device may give more consideration to the frequency of the categories, assigning lower weights to categories with higher frequencies and higher weights to categories with lower frequencies, thereby calculating the third loss. For example, the third loss may be calculated using weighted cross-entropy or Dice loss.
[0084] Next, refer to Figure 8 The device can assign weights to the first loss calculated based on the two-dimensional prediction result 830, the second loss calculated based on the depth measurement 840, and the third loss calculated considering class balance 850, and sum them to calculate the total loss. Furthermore, the device can enable the neural network model 820 to learn in a direction that reduces the total loss, thereby outputting a three-dimensional prediction result 870 related to the occupancy state.
[0085] Figure 9 This is a diagram illustrating an example of a method for predicting occupancy status using a neural network model according to an embodiment.
[0086] Reference Figure 9 The device can input multi-view images 910 as input data to the reference device. Figures 4 to 8 The neural network model 920 learned by the method described.
[0087] In addition, the device can acquire three-dimensional prediction results 940 related to the occupancy status as output data of the neural network model 920.
[0088] More specifically, the device can extract two-dimensional image features from multi-view images and convert them into three-dimensional space. Furthermore, the device can extract three-dimensional features from the three-dimensional space, predict whether a specified object occupies each region of the three-dimensional space based on these features, and classify each region according to the specified object's label, thereby obtaining a three-dimensional prediction result 940 related to the occupancy status.
[0089] However, as Figure 1As stated above, since the images acquired from the camera do not include distance information, ambiguity issues may arise based on the distances of objects included in the image. Therefore, in the inference steps of the neural network model according to an embodiment of this disclosure, an ambiguity suppression process can be further performed.
[0090] According to one embodiment, the device can determine the category based on the confidence level of the categories for each region in three-dimensional space. The confidence level of the category can be calculated based on the distance between each region in three-dimensional space and the ego car. For example, when a region in three-dimensional space is classified as a second group, the device can determine the category of that region as a first group based on the category confidence level. As an example, when a region in three-dimensional space is classified as "car class," but the calculated confidence level is low because the region is far from the ego car, the region can be determined as "free class."
[0091] Figure 10 This is a flowchart illustrating an example of a method for predicting occupancy status using a neural network model according to an embodiment.
[0092] Reference Figure 10 The method for predicting occupancy status using a neural network model may include steps 1010 to 1050. However, even if the following is omitted, refer to Figures 1 to 9 The content described can also be applied to Figure 10 A method for predicting occupancy status using a neural network model.
[0093] In step 1010, the device can acquire multi-view images captured by the camera.
[0094] In step 1030, the device can enable the neural network model to learn by using multi-view images as input data, so that the neural network model can output a three-dimensional prediction result related to the occupancy state using the multi-view images.
[0095] For example, the device can enable a neural network model to perform the following tasks through learning: first, extracting two-dimensional image features for each multi-view image; second, converting the two-dimensional image features into a three-dimensional space; and third, predicting the occupancy status of the three-dimensional space.
[0096] According to one embodiment, the first task may include the steps of receiving, in addition to receiving multi-view images, receiving lidar point clouds acquired from lidar sensors, and additionally outputting two-dimensional prediction results related to the occupancy status.
[0097] According to one embodiment, the second task may further include the step of measuring depth information of the multi-view images.
[0098] According to one embodiment, the third task may include: the step of extracting three-dimensional features for a three-dimensional space; and the step of predicting whether a specified object occupies each region of the three-dimensional space based on the three-dimensional features, and classifying each region according to the label of the specified object.
[0099] According to one embodiment, the apparatus can enable a neural network model to learn based on a first loss, a second loss, and a third loss calculated for each of the first, second, and third tasks. The first loss may include a loss calculated based on two-dimensional prediction results related to occupancy status; the second loss may include a loss calculated based on depth measurements of multi-view images; and the third loss may include a loss calculated based on the classification of categories in three-dimensional space and the frequency of those categories.
[0100] In step 1050, the device can input multi-view images as input data into the learned neural network model and obtain the three-dimensional prediction results related to the occupancy state as the output data of the neural network model.
[0101] According to one embodiment, the device can determine a category based on the confidence level of the categories for each region in three-dimensional space. The confidence level can be calculated based on the distance between each region in three-dimensional space and the vehicle. For example, when a region in three-dimensional space is classified as a second group, the device can determine the category as a first group for that region based on the confidence level of the category.
[0102] Figure 11 This is a configuration diagram illustrating an example of the internal structure of a device for predicting occupancy status using a neural network model according to an embodiment.
[0103] Reference Figure 11 The device 1100 may include a processor 1110, a memory 1120, an input / output interface 1130, and a communication module 1140. For ease of explanation, Figure 11 Only the constituent elements relevant to this invention are shown. Therefore, except for... Figure 11 In addition to the multiple components shown, the device 1100 may also include other general components. Furthermore, those skilled in the art will know that... Figure 4 to 9 The processor 1110, memory 1120, input / output interface 1130 and communication module 1140 shown can be implemented by separate devices.
[0104] Processor 1110 can process computer program instructions by performing basic arithmetic, logic, and input / output operations. These instructions can be provided by memory 1120 or external devices. Furthermore, processor 1110 can control the operation of other components included in device 1100 as a whole.
[0105] Processor 1110 can be implemented as an array of logic gates, or as a combination of a general-purpose microprocessor and a memory storing a program executable by the microprocessor. For example, processor 1110 may include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, and a state machine. In some environments, processor 1110 may also include an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), etc. For example, processor 1110 can refer to a combination of processing devices, such as a combination of a digital signal processor (DSP) and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors integrated with a DSP core, or any other combination of such configurations.
[0106] Memory 1120 may include any non-transitory computer-readable recording medium. As an example, memory 1120 may include a permanent mass storage device such as random access memory (RAM), read-only memory (ROM), a disk drive, a solid-state drive (SSD), or flash memory. As another example, a non-volatile mass storage device (e.g., read-only memory (ROM), solid-state disk (SSD), flash memory, disk drive, etc.) may be a separate persistent storage device distinct from the memory. Additionally, memory 1120 may store an operating system (OS) and at least one program code (e.g., code for causing processor 1110 to execute references). Figure 4 to 9 The code for the action described.
[0107] These software components can be loaded from a computer-readable recording medium independent of memory 1120. Such a separate computer-readable recording medium can be a recording medium capable of being directly connected to device 1100, such as a floppy disk drive, magnetic disk, magnetic tape, DVD / CD-ROM drive, memory card, or other computer-readable recording medium. Alternatively, the software components can also be loaded into memory 1120 via communication module 1140 instead of a computer-readable recording medium. For example, at least one program can be a computer program (e.g., for causing processor 1110 to execute a reference) installed via file provided by the developer or a file distribution system distributing installation files of an application through communication module 1140. Figure 11 The computer program (or similar program) for the described action is loaded into memory 1120.
[0108] The input / output interface 1130 may be a device for interface connection with a device (e.g., keyboard, mouse, etc.) that can be connected to or included in device 1100 for input or output. In this context, the input / output interface 1130 is shown as a component configured independently of the processor 1110, but is not limited thereto; the input / output interface 1130 may also be configured to be included in the processor 1110.
[0109] The communication module 1140 can provide the device 1100 with the configuration or function to communicate with other external devices via a network. For example, control signals, instructions, data, etc., provided by the processor 1110 can be sent to a server and / or external devices via the communication module 1140 and the network.
[0110] As one embodiment, device 1100 can be a mobile electronic device. For example, device 1100 can be a smartphone, tablet computer, personal computer (PC), smart TV, personal digital assistant (PDA), laptop computer, media player, navigator, device equipped with camera, and other mobile electronic devices. Furthermore, device 1100 can also be a wearable device with communication and data processing functions, such as a watch, glasses, hairband, and ring.
[0111] As another embodiment, device 1100 may be an electronic device embedded in a vehicle. For example, device 1100 may be an electronic device installed in a vehicle after the manufacturing process by tuning.
[0112] As another embodiment, device 1100 may be a server located outside the vehicle. The server may be implemented as a computer device or multiple computer devices that communicate via a network to provide commands, code, files, content, services, etc. The server may receive data required for predicting occupancy status from the on-board device and predict the three-dimensional occupancy status based on the received data.
[0113] As another embodiment, the process performed in device 1100 may be performed by at least a portion of a mobile electronic device, an electronic device embedded in a vehicle, and a server located outside the vehicle.
[0114] According to embodiments of the present invention, the program can be implemented in the form of a computer program, which can be executed by various components of a computer, and the computer program described above can be recorded on a computer-readable medium. In this case, the medium may include hardware devices specifically configured to store and execute program instructions, such as magnetic media like hard disks, floppy disks, and magnetic tapes; optical media like CD-ROMs and DVDs; magneto-optical media like floppy disks; and memories such as ROMs, RAMs, and flash memory.
[0115] On the other hand, the computer program may be specifically designed and configured for this invention, or it may be known and available to those skilled in the art of computer software. Examples of computer programs may include not only machine language code generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.
[0116] According to one embodiment, methods according to various embodiments of this disclosure may be provided in a computer program product. The computer program product, as a commodity, can be traded between a seller and a buyer. The computer program product may be distributed (e.g., downloaded or uploaded) in the form of a device-readable storage medium (e.g., a CD-ROM, compact disc read-only memory) or through an application store (e.g., the Play Store™) or directly or online between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0117] The steps constituting the method according to the invention can be performed in any suitable order unless the order is explicitly stated or there is a contrary description. The invention is not necessarily limited to the order of the steps described above. Any examples or exemplary terms used in the embodiments (e.g., etc.) are merely for the purpose of describing the embodiments in detail, and the scope of the embodiments is not limited by the above examples or exemplary terms unless limited by the scope of the patent claims. Furthermore, those skilled in the art will understand that various modifications, combinations, and variations can be made according to design conditions and factors within the scope of the appended claims or their equivalents.
[0118] Therefore, the concept of the present invention should not be limited to the above embodiments, and all scopes that are equivalent to or have been modified from the scope of these claims, not only the scope of the patent claims described below, fall within the scope of the concept of the present invention.
Claims
1. A method for predicting occupancy status using a neural network model, characterized in that, The method includes: The steps to acquire multi-view images from a camera, and The steps of using the multi-view images as input data to enable the neural network model to learn, so that the neural network model can output a three-dimensional prediction result related to the occupancy state using the multi-view images; The steps for enabling the neural network model to learn include: The steps include enabling the neural network model to learn and perform a first task, a second task, and a third task, wherein the first task includes extracting two-dimensional image features for each of the multi-view images, the second task includes converting the two-dimensional image features into a three-dimensional space, and the third task includes predicting the occupancy state in the three-dimensional space.
2. The method for predicting occupancy status using a neural network model according to claim 1, characterized in that, The first task includes: In addition to the multi-view images, the process also includes receiving lidar point clouds obtained from lidar sensors and outputting additional two-dimensional prediction results related to the occupancy status.
3. The method for predicting occupancy status using a neural network model according to claim 1, characterized in that, The second task also includes: The step of measuring depth information of the multi-view image.
4. The method for predicting occupancy status using a neural network model according to claim 1, characterized in that, The third task includes: The steps for extracting three-dimensional features from the three-dimensional space; and The steps are: predicting whether a specified object occupies each region of the three-dimensional space based on the three-dimensional features, and classifying the categories of each region based on the labels of the specified objects.
5. The method for predicting occupancy status using a neural network model according to claim 1, characterized in that, The steps for enabling the neural network model to learn include: The steps involve enabling the neural network model to learn based on a first loss, a second loss, and a third loss calculated respectively for the first task, the second task, and the third task.
6. The method for predicting occupancy status using a neural network model according to claim 5, characterized in that, The first loss includes a loss calculated based on two-dimensional prediction results related to the occupancy state. The second loss includes a loss calculated based on depth measurements of the multi-view images. The third loss includes a loss calculated based on the classification of categories in the three-dimensional space and the frequency of those categories.
7. The method for predicting occupancy status using a neural network model according to claim 1, characterized in that, The method further includes: The step of inputting the multi-view images as input data for the neural network model learned by the method according to claim 1; and The step of obtaining the three-dimensional prediction result related to the occupancy state as the output data of the neural network model.
8. The method for predicting occupancy status using a neural network model according to claim 7, characterized in that, The method further includes: The step of determining the category based on the confidence level of the categories classified for each region of the three-dimensional space.
9. The method for predicting occupancy status using a neural network model according to claim 8, characterized in that, The confidence level is calculated based on the distance between each region in the three-dimensional space and the vehicle.
10. The method for predicting occupancy status using a neural network model according to claim 8, characterized in that, The steps for determining the category include: When a region in the three-dimensional space is classified into the second group, the step of determining the region into the first group based on the confidence level of the category.
11. An apparatus for predicting occupancy status using a neural network model, characterized in that, The device includes: The memory stores at least one program, and The processor executes the at least one program; The processor is configured to: Acquire multi-view images from the camera. The neural network model is trained by using the multi-view images as input data, so that the neural network model can output a three-dimensional prediction result related to the occupancy state using the multi-view images.
12. A computer-readable recording medium, characterized in that, The computer-readable recording medium contains a program for performing the method according to claim 1 on a computer.