Three dimensional object recognition method and device
The method integrates visual, point cloud, and language data using an encoding and decoding model to accurately detect 3D objects, addressing the inefficiencies of current recognition techniques.
Patent Information
- Application Number
- JP2024102458
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-06-26
- Publication Date
- 2025-05-23
AI Technical Summary
Current methods for three-dimensional object recognition lack efficiency in processing and integrating visual, point cloud, and language data to accurately detect and classify objects in 3D space.
A method and apparatus that utilize an encoding model to generate features from input images, point clouds, and languages, followed by a multimodal decoding model and object detection model to select target video features and detect 3D bounding boxes corresponding to target objects.
This approach enables accurate recognition and detection of 3D objects by effectively integrating multiple data modalities, improving the generalization and accuracy of object recognition in complex 3D environments.
Smart Images

Figure 2025080215000001_ABST
Abstract
Description
[Technical field]
[0001] The following embodiments relate to a method and apparatus for three-dimensional object recognition. [Background technology]
[0002] The technical automation of the recognition process is realized, for example, through a neural network model embodied in a processor as a specialized computational structure, which can provide computationally intuitive mapping between input patterns and output patterns after considerable training. The trained ability to generate such mapping is called the learning ability of the neural network model. Furthermore, due to specialized training, such a specialized trained neural network model may have a generalization ability (or generalization ability) to generate relatively accurate output for, for example, untrained input patterns. Summary of the Invention [Problem to be solved by the invention]
[0003] The following embodiments aim to provide a method and apparatus for recognizing a three-dimensional object. [Means for solving the problem]
[0004] According to one embodiment, a 3D object recognition method includes the steps of receiving an input image related to a 3D space, an input point cloud related to the 3D space, and an input language related to a target object in the 3D space; generating candidate video features for a subregion of the input image, point cloud features of the input point cloud, and language features of the input language using an encoding model; selecting a target video feature corresponding to the language feature from the candidate video features based on a similarity score of similarity between the candidate video features and the language features; operating a multimodal decoding model based on the target video features and the point cloud features to generate a decoding output; and operating an object detection model based on the decoding output to detect a 3D bounding box corresponding to the target object.
[0005] An electronic device according to one embodiment includes one or more processors and a memory for storing instructions, and the instructions are configured to: receive, by the one or more processors, an input image relating to a three-dimensional space, an input point cloud relating to the three-dimensional space, and an input language relating to a target object in the three-dimensional space; generate candidate video features for a subregion of the input image, point cloud features of the input point cloud, and language features of the input language using an encoding model; select a target video feature corresponding to the language feature among the candidate video features based on a similarity score of similarity between the candidate video features and the language features; operate a multimodal decoding model based on the target video features and the point cloud features to generate a decoding output; and operate an object detection model based on the decoding output to detect a 3D bounding box corresponding to the target object.
[0006] A vehicle in one embodiment includes a camera that generates an input image related to three-dimensional space, a lidar sensor that generates an input point cloud related to the three-dimensional space, one or more processors that receive the input image related to the three-dimensional space, the input point cloud related to the three-dimensional space, and an input language related to a target object in the three-dimensional space, generate candidate video features for subregions of the input image, point cloud features of the input point cloud, and language features of the input language using an encoding model, select target video features from the candidate video features corresponding to the language features based on a similarity score of similarity between the candidate video features and the language features, operate a multimodal decoding model based on the target video features and the point cloud features to generate a decoding output, operate an object detection model based on the decoding output, and detect a three-dimensional bounding box corresponding to the target object, and a control system that controls the vehicle based on the three-dimensional bounding box. Effect of the Invention
[0007] According to the embodiment, a method and apparatus for recognizing a three-dimensional object can be provided. [Brief description of the drawings]
[0008] [Figure 1] FIG. 2 is a diagram illustrating an example of a configuration of a three-dimensional object recognition model according to an embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an exemplary configuration of a vision-language model according to an embodiment. [Diagram 3] 13A-13C are exemplary illustrations of objects of the same class with different geometric properties according to one embodiment. [Figure 4] FIG. 2 is a diagram illustrating an exemplary operation of a multi-modal decoding model according to an embodiment. [Diagram 5] FIG. 2 is a diagram illustrating an exemplary training process of a vision-language model according to an embodiment. [Figure 6]FIG. 2 is a diagram illustrating an exemplary training process of a 3D object recognition model according to an embodiment. [Figure 7] 1 is a flowchart illustrating an exemplary 3D object recognition method according to an embodiment. [Figure 8] FIG. 1 is a block diagram illustrating a configuration of an electronic device according to an embodiment. [Figure 9] 1 is a block diagram illustrating an example configuration of a vehicle according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] The specific structural or functional description of the embodiments is disclosed for the purpose of illustration only, and may be modified in various forms. Therefore, the embodiments are not limited to the specific disclosed forms, and the scope of the present specification includes modifications, equivalents, or alternatives within the technical spirit.
[0010] Although terms such as first or second may be used to describe multiple components, such terms should be construed only for the purpose of distinguishing one component from the other components. For example, a first component may be named a second component, and similarly, a second component may be named a first component.
[0011] When any component is referred to as being "connected" to another component, it is directly connected or connected to the other component, but it should be understood that there may be other components in between.
[0012] The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, the terms "comprise" or "have" and the like indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood as not precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0013] As used herein, phrases such as "at least one of A or B" and "at least one of A, B, or C" may include any one or all possible combinations of the items listed with that phrase.
[0014] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the present invention belongs. Commonly used predefined terms should be interpreted as having a meaning consistent with the meaning they have in the context of the relevant art, and should not be interpreted as ideal or overly formal unless expressly defined in this specification.
[0015] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. In the description with reference to the drawings, the same components are given the same reference numerals regardless of the reference numerals, and the duplicated description thereof will be omitted.
[0016] 1 is a diagram illustrating an example of a configuration of a 3D object perception model according to an embodiment. Referring to FIG. 1, the 3D object perception model 100 includes a vision-language model 110, a point cloud encoding model 120, a multi-modal decoding model 130, and an object detection model 140.
[0017] The 3D object recognition model 100 receives an input video 101 for a 3D space, an input language 102 for a target object in the 3D space, and an input point cloud 103 for the 3D space. The target object means an object that is a target of recognition. The 3D object recognition model 100 can detect a 3D bounding box 141 corresponding to the target object using a vision-language model 110, a point cloud encoding model 120, a multimodal decoding model 130, and an object detection model 140.
[0018] The three-dimensional space indicates an actual space. The input image 101 is a result of capturing at least a part of the three-dimensional space using a camera. For example, the input image 101 may be a color image such as an RGB image. The camera may include a sub-camera of another view. The input image 101 may include a sub-image of another view captured by a sub-camera. For example, a sub-image of a six-direction view may be generated through a sub-camera of a six-direction view. The input point cloud 103 is a result of detecting at least a part of the three-dimensional space using a lidar (light detection and range) sensor. The scene shown in the input image 101 and the scene shown in the input point cloud 103 may at least partially overlap. For example, the camera (or sub-camera) and the lidar sensor may be provided on a vehicle and may capture the surroundings of the vehicle at an angle of 360 degrees. A target object may be included in the overlapping portion. A virtual space corresponding to the three-dimensional space is defined by the input image 101 and the input point cloud 103. A three-dimensional bounding box 141 may be formed in the virtual space.
[0019] The vision-language model 110 can learn the relationship between vision information (e.g., video information) and language information (e.g., text information, speech information) and solve problems based on the relationship. The vision-language model 110 learns a large number of objects through vision information and language information, and is used for open vocabulary object detection (VOD) according to the learning characteristics of the vision-language model 110. Open VOD refers to the ability to detect not only objects that were the learning target, but also objects that were not the learning target.
[0020] The 3D object recognition model 100 uses the encoding models to generate candidate video features for subregions of the input video 101, language features for the input language 102, and point cloud features 121 for the input point cloud 103. The encoding models include a video encoding model of the vision-language model 110, a language encoding model of the vision-language model 110, and a point cloud encoding model 120. The video encoding model, the language encoding model, and the point cloud encoding model 120 are neural network models that can generate features related to input data, such as a convolutional neural network (CNN) or a transformer encoder.
[0021] The vision-language model 110 may include a language encoding model, a video encoding model, and a region proposal model. The video encoding model is executed based on the input video 101 to generate video features of the input video 101, and the region proposal model may determine candidate video features corresponding to partial regions of the input video from the video features. The language encoding model is executed based on the input language 102 to generate language features corresponding to the input language 102. The vision-language model 110 determines a similarity score of similarity between the candidate video features and the language features, and selects a target video feature 111 corresponding to the language feature from the candidate video features based on the similarity score. The target video feature 111 may be called an attended image feature in that it is selected via the language feature. The point cloud encoding model 120 may be executed based on the input point cloud 103 to generate a point cloud feature 121 corresponding to the input point cloud 103.
[0022] The multimodal decoding model 130 may be executed based on the target image features 111 and the point cloud features 121 to generate a decoding output. The multimodal decoding model 130 analyzes associations between the target image features 111 and the point cloud features 121 of different modalities. The multimodal decoding model 130 may grasp a shape of a target object from the target image features 111 corresponding to 2D image information, identify points corresponding to the shape of the target object from the point cloud features 121, and select a detection position corresponding to the position of the target object from among detection position candidates of the detection guide information 104 based on the position of the point.
[0023] The multimodal decoding model 130 is used to detect various objects based on the open VOD of the vision-language model 110. A general object recognition model can add new classes by predetermining and learning classes of target objects and going through a new learning process. The open VOD of the vision-language model 110 can generate target image features 111 of various objects including new classes without such a new learning process. The multimodal decoding model 130 can grasp the shapes of various objects including new objects from such target image features 111 and generate a decoding output for determining a 3D bounding box 141 of the corresponding object.
[0024] The multimodal decoding model 130 may be a transformer decoder. The multimodal decoding model 130 extracts correlations from query data (Q), key data (K), and value data (V) to perform decoding. The key data and value data are determined based on the target video feature 111 and the point cloud feature 121, and the query data is determined based on the detection guide information 104. The detection guide information 104 indicates detection position candidates where a target object may be detected in a three-dimensional space.
[0025] An object detection model 140 is run based on the decoding output to detect a 3D bounding box 141 corresponding to a target object. The object detection model 140 may be a neural network model such as a multi-layer perceptron (MLP).
[0026] 2 is a diagram illustrating an example of a configuration of a vision-language model according to an embodiment. Referring to FIG. 2, the vision-language model 200 includes a language encoding model 220, a video encoding model 240, and a region proposal model 250. The language encoding model 220, the video encoding model 240, and the region proposal model 250 may be neural network models.
[0027] The language encoding model 220 generates language features 261 based on the input language 210. For example, the input language 210 may include textual and / or speech information. Extended expressions 211 may be generated based on the input language 210. The extended expressions 211 include a position field 2111 and a class field 2112, each of which indicates a geometric characteristic of a target object. The class field 2112 has other values for other classes (e.g., vehicle, human, traffic light, lane, etc.).
[0028] The geometric characteristics of the position field 2111 include a geometric position, a geometric shape, etc. The position field 2111 has other values for other positions (e.g., near, far, left side of the image, right side of the image, top of the image, bottom of the image, top left side of the image, bottom left side of the image, top right side of the image, bottom right side of the image, etc.). Objects of the same class with other geometric characteristics are differentiated based on the position field 2111. For example, a vehicle in the top of the image and a vehicle in the bottom of the image may be differentiated based on the position field 2111. As another example, a vehicle in a complete state without occlusion and a vehicle in a state with occlusion may be differentiated based on the position field 2111.
[0029] The position field 2111 has a learnable characteristic. As described again below, the value of the position field 2111 can be determined in a training process (or a training process or a learning process) of the multimodal decoding model. The input language 210 corresponds to a prompt. Through situation optimization of the training process, the value of the position field 2111 is optimized so that various geometric characteristics of each object can be distinguished and detected. Objects of the same class having various geometric characteristics are presented, and the multimodal decoding model distinguishes and learns objects of the same class having various geometric characteristics, thereby improving the detection accuracy of the multimodal decoding model for objects of the corresponding class.
[0030] The video encoding model 240 generates video features corresponding to the input video 230. The region proposal model 250 determines candidate video features 262 corresponding to subregions of the input video 230 based on the video features. The subregions are regions where an object is likely to exist.
[0031] The score table 260 includes a similarity score 263 of the degree of similarity between the linguistic feature 261 and the candidate video feature 262. For example, the similarity score 263 may be determined based on the Euclidean distance.
[0032] The detail features SF_1, SF_2, and SF_3 of the linguistic features 261 correspond to the extended representation 211. For example, the first detail feature SF_1 corresponds to the first extended representation of the extended representation 211, the second detail feature SF_2 corresponds to the second extended representation of the extended representation 211, and the third detail feature SF_3 corresponds to the third extended representation of the extended representation 211.
[0033] Once a target object is determined, an augmented representation 211 having the target object in a class field 2112 may be constructed. For example, if a vehicle is determined to be the target object, an augmented representation 211 of various locations having the vehicle as a class may be constructed. The augmented representation 211 of each class may be determined during the training process of the multimodal decoding model. If the input language 210 specifies multiple classes, an augmented representation 211 and language features 261 are constructed for each class.
[0034] Among the candidate video features 262, at least some of the candidate video features having the highest similarity scores to the detail features SF_1, SF_2, and SF_3 of the language feature 261 are selected as target video features. For example, when the third candidate video feature CIF_3 shows the highest similarity score to the first detail feature SF_1, the third candidate video feature CIF_3 is selected as the target video feature for the first detail feature SF_1, when the fifth candidate video feature CIF_5 shows the highest similarity score to the second detail feature SF_2, the fifth candidate video feature CIF_5 is selected as the target video feature for the second detail feature SF_2, and when the first candidate video feature CIF_1 shows the highest similarity score to the third detail feature SF_3, the first candidate video feature CIF_1 may be selected as the target video feature for the third detail feature SF_3. In this way, the target video features for each detail feature are selected, and a three-dimensional bounding box corresponding to each target video feature can be determined.
[0035] FIG. 3 exemplarily illustrates objects of the same class having different geometric characteristics according to an embodiment. Referring to FIG. 3, an input image 310 may include a vehicle 301 in a complete state at the bottom left of the input image 310, a vehicle 302 in a complete state at the top of the input image 310, and a vehicle 303 in a blocked state at the bottom right of the input image 310. If only the vehicle class can be distinguished through the input language, all of the vehicles 301, 302, and 303 with various characteristics may not be recognized. According to the embodiment, not only the vehicle class but also various geometric characteristics are distinguished through the input language, so that all of the vehicles 301, 302, and 303 with various characteristics can be accurately recognized.
[0036] 4 is a diagram illustrating an example of the operation of a multimodal decoding model according to an embodiment. Referring to FIG. 4, a target video feature 401 may be divided into image tokens 402, and a point cloud feature 403 may be divided into point cloud tokens 404. Position information 405 includes first position information indicating the relative positions of the image tokens 402 and second position information indicating the relative positions of the point cloud tokens 404. The relative positions indicate the positions of each token in each video.
[0037] The key data and value data are determined based on the video token 402, the point cloud token 404, the first location information of the location information 405, and the second location information of the location information 405. The key data and value data may be the same. For example, corresponding pairs of video tokens 402 and point cloud tokens 404 based on their relative positions may be combined (e.g., concatenated) and the first location information and / or the second location information may be added to the combined corresponding pairs to in turn form the key data and value data.
[0038] The detection guide information 406 indicates possible detection location candidates at which a target object may be detected in a three-dimensional space. The detection guide information 406 constitutes the query data. The detection location candidates indicate non-uniform locations. As will be described again below, the detection location candidates may be optimized by the detection guide model during the training of the multimodal decoding model 410. The detection location candidates may indicate non-uniform locations as a result of the optimization.
[0039] The multimodal decoding model 410 corresponds to a transformer decoder. The multimodal decoding model 410 is executed based on the query data, the key data, and the value data to generate a decoding output. The multimodal decoding model 410 can extract associations from the target video features 401, the point cloud features 403, and the detection guide information 406 based on the query data, the key data, and the value data to generate a decoding output. The object detection model 420 is executed based on the decoding output to determine a 3D bounding box 421. The 3D bounding box 421 corresponds to one of the detection position candidates.
[0040] 5 is a diagram illustrating a training process of a vision-language model according to an embodiment of the present invention. Referring to FIG. 5, the vision-language model 500 includes a language encoding model 520, a video encoding model 540, and a region proposal model 550.
[0041] The language encoding model 520 generates language features 561 based on the input language 510. The input language 510 identifies a respective class. The input language 510 may not include separate location information such as the location field 2111 shown in FIG. 2. For example, the input language 510 may be a prompt describing an object shown in the input image 530. The object corresponds to the target object.
[0042] The video encoding model 540 generates video features corresponding to the input video 530. The region proposal model 550 determines candidate video features 562 corresponding to subregions of the input video 530 based on the video features. The subregions may be regions where an object is likely to exist. The region location loss 571 indicates the difference between the object region proposed by the region proposal model 550 and a ground truth (GT). The region proposal model 550 has the ability to propose regions where an object is likely to exist by training based on the region location loss 571.
[0043] The language feature 561 corresponds to the input language 510. For example, the first language feature LF_1 corresponds to the first input language of the input language 510, the second language feature LF_2 corresponds to the second input language of the input language 510, and the third language feature LF_3 corresponds to the third input language of the input language 510.
[0044] The candidate region features 562 may be aligned to match the language features. For example, a video feature corresponding to a first language feature class may be determined as a first candidate video feature CIF_1, a video feature corresponding to a second language feature class may be determined as a second candidate video feature CIF_2, and a video feature corresponding to a third language feature class may be determined as a third candidate video feature CIF_3.
[0045] The score table 560 may include a similarity score 563 of the similarity between the language feature 561 and the candidate video feature 562. For example, the similarity score 563 may be determined based on a Euclidean distance. The similarity score 563 may be trained based on an alignment loss 572 so that diagonal elements have high values and off-diagonal elements have low values. For example, in the score table 560, SS_11, SS_22, and SS_33 correspond to the diagonal elements, and SS_12, SS_13, and SS_21. SS_23, SS_31, and SS_32 correspond to the off-diagonal elements. The alignment loss 572 may increase the similarity score value of the diagonal elements and decrease the similarity score value of the off-diagonal elements. As a result of the training, the score table 260 may be a diagonal matrix or approach a diagonal matrix.
[0046] Based on the alignment loss 572-based training, pairs of language features and candidate video features related to the same object (e.g., LF_1 and CIF_1, LF_2 and CIF_2, LF_3 and CIF_3) have similar feature values, and the language encoding model 520 and the video encoding model 540 are trained to generate language features 561 and candidate video features 562 such that the corresponding pairs have similar feature values. The vision-language model 500 has the ability to solve the open VOD problem through such alignment loss 572-based training.
[0047] 6 is a diagram illustrating an example of a training process of a 3D object recognition model according to an embodiment. Referring to FIG. 6, the 3D object recognition model 600 generates a 3D bounding box 651 corresponding to an input image 601, an input language 602, an input point cloud 603, and guide coordinate information 604 based on a vision-language model 610, a point cloud encoding model 620, a multimodal decoding model 630, a detection guide model 640, and an object detection model 650.
[0048] The vision-language model 610 generates target video features 611 based on the input video 601 and the input language 602. The point cloud encoding model 620 generates point cloud features 621 based on the input point cloud 603. The multimodal decoding model 630 generates a decoding output based on the target video features 611, the point cloud features 621, and the detection guide information. The detection guide information corresponds to the output of the detection guide model 640. The object detection model 650 determines a 3D bounding box 651 based on the decoding output.
[0049] The 3D object recognition model 600 may be trained based on a box position loss 661 of the 3D bounding box 651. The box position loss 661 corresponds to the difference between the 3D bounding box 651 and the GT. Here, the position field 6021, the multimodal decoding model 630, the detection guide model 640, and the object detection model 650 may be trained in a state where the vision-language-model 610 (e.g., a language encoding model, a video encoding model, a region proposal model), the class field 6022, and the point cloud encoding model 620 are frozen. For example, the vision-language model 610 may be loaded onto the 3D object recognition model 600 in a pre-trained state based on the scheme shown in FIG. 5.
[0050] The multimodal decoding model 630 is used to detect various objects based on the open VOD of the vision-language model 610. A general object recognition model can add new classes by predetermining and learning classes of target objects and going through a new learning process. The open VOD of the vision-language model 110 can generate target image features 611 of various objects including new classes without such a new learning process. The multimodal decoding model 630 can grasp the shapes of various objects including new objects from such target image features 611 and generate a decoding output for determining a 3D bounding box 651 of the corresponding object.
[0051] The input language 602 may be expanded into an expanded representation that includes a location field 6021 and a class field 6022. The location field 6021 has learnable properties and the class field 6022 is frozen. For example, the location field 6021 may include learnable parameters of a certain magnitude, and the parameters may be trained to values that reduce the box location loss 661 during the training of the multimodal decoding model 630.
[0052] The location field 6021 may be initialized to an arbitrary value, and a box location loss 661 may be used for contextual optimization to include spatial information. As a result of learning, the location field 6021 and the class field 6022 may be paired to provide a spatial identity to each object of the same class with other geometric characteristics of the same class. This training method does not encode spatial information in an extrinsic manner, but allows the 3D object recognition model to learn spatial information in an intrinsic manner using 3D labels, and many queries can be initialized to statistically meaningful positions by the detection guide model 640, compared to the conventional use of equally spaced anchors.
[0053] The detection guide model 640 generates detection guide information based on the guide coordinate information 604. The detection guide information indicates potential detection positions where a target object of the input language 602 may be detected in a three-dimensional space. The guide coordinate information 604 indicates a uniform position for the three-dimensional space, and the detection guide information indicates a non-uniform position for the three-dimensional space. The uniform positions of the guide coordinate information 604 may be adjusted to the non-uniform positions of the detection guide information by training the detection guide model 640. The detection guide model 640 may adjust the uniform positions to non-uniform positions where a target object is more likely to be detected based on the training dataset.
[0054] The detection guide model 640 may be a neural network model, such as a 1×1 convolutional neural network. The detection guide model 640 may adjust the dimensions of the guide coordinate information 604 to the input dimensions of the multimodal decoding model 630.
[0055] 7 is a flowchart illustrating an exemplary 3D object recognition method according to an embodiment. Referring to FIG. 7, the electronic device receives an input video related to a 3D space, an input point cloud related to the 3D space, and an input language related to a target object in the 3D space in step 710, generates candidate video features of a partial region of the input video, point cloud features of the input point cloud, and language features of the input language using an encoding model in step 720, selects a target video feature corresponding to the language feature among the candidate video features based on a similarity score between the candidate video features and the language feature in step 730, operates a multimodal decoding model based on the target video features and point cloud features to generate a decoding output in step 740, and operates an object detection model based on the decoding output to detect a 3D bounding box corresponding to the target object in step 750.
[0056] Step 720 includes the steps of running a language encoding model based on an input language to generate language features corresponding to the input language, running a video encoding model and a region proposal model based on the input video to generate candidate video features corresponding to subregions of the input video, and running a point cloud encoding model based on the input point cloud to generate point cloud features corresponding to the input point cloud.
[0057] An augmented representation is generated based on the input language, the augmented representation including a position field indicating a geometric characteristic of the target object and a class field indicating a class of the target object, and linguistic features can be generated based on the augmented representation. Objects of the same class of other geometric characteristics are classified based on the position field. The position field has learnable characteristics.
[0058] Step 740 includes the steps of segmenting the target video features to generate video tokens, segmenting the point cloud features to generate point cloud tokens, generating first position information indicating the relative positions of the video tokens, generating second position information indicating the relative positions of the point cloud tokens, and running a multimodal decoding model on key data and value data based on the video tokens, the point cloud tokens, the first position information, and the second position information.
[0059] Step 740 includes running a multi-modal decoding model on the query data based on detection guide information indicating potential detection locations where a target object may be detected in three-dimensional space. The potential detection locations indicate non-uniform locations. The multi-modal decoding model can extract associations from the target video features, the point cloud features, and the detection guide information to generate a decoding output.
[0060] Fig. 8 is a block diagram exemplarily illustrating a configuration of an electronic device according to an embodiment. Referring to Fig. 8, the electronic device 800 includes a processor 810 and a memory 820. Although not shown in Fig. 8, the electronic device 800 may further include other devices such as a camera, a lidar sensor, a storage device, an input device, an output device, and a network device. For example, the electronic device 800 is any device that can recognize the surrounding environment using a camera or a sensor (e.g., a lidar), such as a vehicle, a robot, a mobile device, etc.
[0061] The processor 810 may be one or more processors. The memory 820 is coupled to the processor 810 and stores instructions executable by the processor 810, data operated on by the processor 810, data processed by the processor 810, or a combination thereof. The memory 820 may include a non-transitory computer-readable storage medium, such as a high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., a disk storage device, a flash memory device, or other non-volatile solid-state memory device).
[0062] The processor 810 may execute commands for performing the operations of FIGS. 1 to 7 and 9. For example, when the commands are executed by the processor 810, the electronic device 800 may receive an input image related to a three-dimensional space, an input point cloud related to the three-dimensional space, and an input language related to a target object in the three-dimensional space, generate candidate image features of partial regions of the input image, point cloud features of the input point cloud, and language features of the input language using an encoding model, select a target image feature corresponding to the language feature among the candidate image features based on a similarity score of similarity between the candidate image features and the language feature, operate a multimodal decoding model based on the target image feature and the point cloud feature to generate a decoding output, and operate an object detection model based on the decoding output to detect a three-dimensional bounding box corresponding to the target object. The input image may be generated by a camera, and the input point cloud may be generated by a lidar sensor. The input language may be set by a user or may be set in advance.
[0063] Fig. 9 is a block diagram illustrating an example of a configuration of a vehicle according to an embodiment. Referring to Fig. 9, a vehicle 900 includes a camera 910, a lidar sensor 920, a processor 930, and a control system 940. Although not shown in Fig. 9, the vehicle 900 may further include other devices such as a storage device, a memory, an input device, an output device, a network device, and a drive system.
[0064] The camera 910 generates an input image related to the three-dimensional space, and the lidar sensor 920 generates an input point cloud related to the three-dimensional space. The input language may be set by a user or may be set in advance. The processor 930 may execute commands for performing the operations shown in FIGS. 1 to 8. For example, the processor 930 may receive an input image related to the three-dimensional space, an input point cloud related to the three-dimensional space, and an input language related to a target object in the three-dimensional space, generate candidate image features of a partial region of the input image, point cloud features of the input point cloud, and language features of the input language using an encoding model, select a target image feature corresponding to the language feature from the candidate image features based on a similarity score of the similarity between the candidate image features and the language feature, operate a multimodal decoding model based on the target image feature and the point cloud feature to generate a decoding output, and operate an object detection model based on the decoding output to detect a three-dimensional bounding box corresponding to the target object. The control system 940 may control the vehicle 900 and / or the drive system based on the object recognition result. For example, the control system 940 may control the speed and / or steering of the vehicle 900 based on the recognition results of a three-dimensional object (eg, a three-dimensional bounding box).
[0065] The embodiments described above may be implemented using hardware components, software components, or a combination of hardware and software components. For example, the devices and components described herein may be implemented using one or more general-purpose or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable array (FPA), a programmable logic unit (PLU), a microprocessor, or other devices that execute and respond to instructions. The processing device executes an operating system (OS) and one or more software applications that run on the operating system. The processing device also accesses, stores, manipulates, processes, and generates data in response to the execution of the software. For ease of understanding, the processing device may be described as being one in which only one processing device is used, but one skilled in the art will appreciate that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, the processing device may include multiple processors or one processor and one controller. Other processing configurations, such as parallel processors, are also possible.
[0066] The software may include computer programs, codes, instructions, or any combination of one or more thereof, capable of configuring or instructing a processing device to operate as desired, either independently or in combination. The software and / or data may be embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or transmitted signal wave, either permanently or temporarily, to be interpreted by the processing device or to provide instructions or data to the processing device. The software may be distributed across computer systems coupled to a network, and may be stored and executed in a distributed manner. The software and data may be stored on one or more computer readable recording media.
[0067] The method according to the present invention may be embodied in the form of program instructions to be executed by various computer means and recorded on a computer-readable recording medium. The recording medium may include program instructions, data files, data structures, and the like, alone or in combination. The recording medium and the program instructions may be specially designed and constructed for the purposes of the present invention, or may be well known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROMs, RAMs, flash memories, and the like. Examples of program instructions include not only machine language code, such as that generated by a compiler, but also high-level language code executed by a computer using an interpreter, etc.
[0068] The hardware devices described above may be configured to operate as one or more software modules to perform the operations illustrated in the present invention, and vice versa.
[0069] Although the embodiments have been described above with reference to limited drawings, those skilled in the art may apply various technical modifications and variations based on the above description. For example, the described techniques may be performed in a different order than described, and / or the components of the described systems, structures, devices, circuits, etc. may be combined or combined in a different form than described, and may be replaced or substituted with other components or equivalents to achieve appropriate results.
[0070] Accordingly, other implementations, embodiments, and equivalents of the claims are intended to be within the scope of the following claims.
Claims
1. receiving an input image relating to a three-dimensional space, an input point cloud relating to the three-dimensional space, and an input language relating to a target object in the three-dimensional space; generating candidate video features for subregions of the input video, point cloud features for the input point cloud, and language features for the input language using an encoding model; selecting a target video feature corresponding to the language feature from among the candidate video features based on a similarity score between the candidate video features and the language feature; operating a multi-modal decoding model based on the target video features and the point cloud features to generate a decoding output; operating an object detection model based on the decoding output to detect a 3D bounding box corresponding to the target object; A method for three-dimensional object recognition comprising:
2. The step of generating the candidate video features, the point cloud features, and the language features comprises: operating a language encoding model based on the input language to generate linguistic features corresponding to the input language; operating a video encoding model and a region proposal model based on the input video to generate candidate video features corresponding to subregions of the input video; operating a point cloud encoding model on the input point cloud to generate point cloud features corresponding to the input point cloud; The three-dimensional object recognition method of claim 1 , comprising:
3. an augmented representation is generated based on the input language, the augmented representation including a position field indicating a geometric characteristic of the target object and a class field indicating a class of the target object; The method of claim 1 , wherein the linguistic features are generated based on the extended representation.
4. The method of claim 3, wherein objects of the same class but other geometric characteristics are classified based on the position field.
5. The method of claim 3 , wherein the position field has learnable characteristics.
6. The step of generating a decoding output comprises: segmenting the target video features to generate video tokens; segmenting the point cloud features to generate point cloud tokens; generating first position information indicative of a relative position of the video token; generating second position information indicative of a relative position of the point cloud tokens; running the multi-modal decoding model on key data and value data based on the video tokens, the point cloud tokens, the first location information, and the second location information; The three-dimensional object recognition method of claim 1 , comprising:
7. 7. The method of claim 6, wherein generating the decoding output comprises operating the multi-modal decoding model with query data based on detection guide information indicating candidate detection positions at which the target object may be detected in the three-dimensional space.
8. The method of claim 7 , wherein the candidate detection positions indicate non-uniform positions.
9. The method of claim 7 , wherein the multi-modal decoding model extracts associations from the target video features, the point cloud features, and the detection guide information to generate the decoding output.
10. A computer program causing a computer to execute the three-dimensional object recognition method according to any one of claims 1 to 9.
11. 1. An electronic device comprising: one or more processors; A memory for storing instruction words; Including, The instructions are executed by the one or more processors: receiving an input image relating to a three-dimensional space, an input point cloud relating to the three-dimensional space, and an input language relating to a target object in the three-dimensional space; generating candidate video features for subregions of the input video, point cloud features for the input point cloud, and language features for the input language using an encoding model; selecting a target video feature corresponding to the language feature from among the candidate video features based on a similarity score between the candidate video features and the language feature; operating a multi-modal decoding model based on the target image features and the point cloud features to generate a decoding output; An electronic device configured to operate an object detection model based on the decoding output to detect a three-dimensional bounding box corresponding to the target object.
12. The instructions are transmitted by the one or more processors to: operating a language encoding model based on the input language to generate linguistic features corresponding to the input language; operating a video encoding model and a region proposal model based on the input video to generate candidate video features corresponding to subregions of the input video; The electronic device of claim 11 , further configured to operate a point cloud encoding model based on the input point cloud to generate point cloud features corresponding to the input point cloud.
13. an augmented representation is generated based on the input language, the augmented representation including a position field indicating a geometric characteristic of the target object and a class field indicating a class of the target object; The electronic device of claim 11 , wherein the linguistic features are generated based on the augmented representation.
14. The electronic device of claim 13 , wherein objects of the same class are separated based on the position field and other geometric characteristics.
15. The electronic device of claim 13 , wherein the location field has learnable characteristics.
16. The instructions are transmitted by the one or more processors to: Segmenting the target video features to generate video tokens; Segmenting the point cloud features to generate point cloud tokens; generating first position information indicative of a relative position of the video token; generating second position information indicative of a relative position of the point cloud tokens; 12. The electronic device of claim 11, further configured to operate the multi-modal decoding model on key data and value data based on the video tokens, the point cloud tokens, the first location information, and the second location information.
17. The instructions are transmitted by the one or more processors to:
17. The electronic device of claim 16, further configured to operate the multi-modal decoding model on query data based on detection guide information indicative of potential detection locations at which the target object may be detected in the three-dimensional space.
18. The electronic device of claim 17 , wherein the potential detection locations indicate non-uniform locations.
19. The electronic device of claim 17 , wherein the multi-modal decoding model extracts associations from the target video features, the point cloud features, and the detection guide information to generate the decoding output.
20. A vehicle, A camera for generating an input image relating to a three-dimensional space; a lidar sensor that generates an input point cloud relating to the three-dimensional space; receiving the input video relating to the three-dimensional space, the input point cloud relating to the three-dimensional space, and an input language relating to a target object in the three-dimensional space; generating candidate video features for subregions of the input video, point cloud features for the input point cloud, and language features for the input language using an encoding model; selecting a target video feature corresponding to the language feature from among the candidate video features based on a similarity score between the candidate video features and the language feature; operating a multi-modal decoding model based on the target image features and the point cloud features to generate a decoding output; one or more processors for operating an object detection model based on the decoding output to detect a 3D bounding box corresponding to the target object; a control system that controls the vehicle based on the three-dimensional bounding box; Including, vehicles.