Method and apparatus utilizing three-dimensional object perception

By receiving multimodal input data in three-dimensional space, using the coding model to generate features and integrate similarity scores, and performing multimodal decoding models and object detection models, the shortcomings of three-dimensional object bounding box detection in the existing technology are solved, and efficient multimodal data integration and accurate detection effects are achieved.

CN119992533APending Publication Date: 2025-05-13SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410574545.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-13
Filing Date
2024-05-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect bounding boxes of three-dimensional objects, especially in terms of information integration between multimodal data such as images and point clouds.

Method used

By receiving input images, point clouds and languages ​​in three-dimensional space, using the encoding model to generate candidate image features, point cloud features and language features, select target image features based on similarity scores, perform multimodal decoding model to generate decoded output, and detect the three-dimensional bounding box through the object detection model.

Benefits of technology

Accurate detection of three-dimensional object bounding boxes is achieved, the ability to integrate multimodal data is improved, and it can handle untrained input patterns to provide relatively accurate output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992533A_ABST
    Figure CN119992533A_ABST
Patent Text Reader

Abstract

A method and apparatus for three-dimensional (3D) object detection are provided. The method includes: receiving an input image relative to a 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space; using the encoding model to generate candidate image features of a partial region of the input image, point cloud features of the input point cloud, and language features of the input language; selecting a target image feature corresponding to the linguistic feature from among the candidate image features based on a similarity score of a similarity between the candidate image features and the linguistic feature; generating a decoded output by executing a multi-modal decoding model based on the target image features and the point cloud features; and detecting a 3D bounding box corresponding to the target object by executing an object detection model based on the decoded output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority from Korean Patent Application No. 10-2023-0156519 filed in the Korean Intellectual Property Office on November 13, 2023, the disclosure of which is incorporated herein by reference for all purposes. Technical Field

[0003] The following description relates to methods and apparatus for utilizing three-dimensional object perception. Background Art

[0004] Technical automation of the perception process has been achieved, for example, by a neural network model implemented by a processor configured as a special computing structure, which provides an intuitive mapping for calculations between input patterns and output patterns after extensive training. Such training capabilities to generate mappings can be referred to as the learning capabilities of the neural network model. In addition, the neural network model trained and specialized through special training has a generalization capability to provide relatively accurate outputs for untrained input patterns, for example. Summary of the invention

[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0006] In a general aspect, a method for detecting a three-dimensional (3D) object includes: receiving an input image relative to a 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space; using an encoding model to generate candidate image features of a partial area of ​​the input image, point cloud features of an input point cloud, and language features of the input language; selecting a target image feature corresponding to the language feature from among the candidate image features based on a similarity score of a similarity between the candidate image features and the language features; generating a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features; and detecting a 3D bounding box corresponding to the target object by executing an object detection model based on the decoding output.

[0007] Generating candidate image features, point cloud features, and language features may include: generating language features corresponding to the input language by performing inference on the input language through a language coding model; generating candidate image features corresponding to partial regions of the input image by executing an image coding model and a region proposal model based on the input image; and generating point cloud features corresponding to the input point cloud by executing a point cloud coding model based on the input point cloud.

[0008] The method may also include generating extended expressions, each extended expression including (i) a position field indicating a geometric characteristic of the target object based on the input language, and including (ii) a category field indicating a category of the target object, and wherein the language feature is generated based on the extended expression.

[0009] Objects of the same category and having different geometric properties can be distinguished from each other based on the position field of the extended expression.

[0010] The location field can be learned through training.

[0011] Generating a decoding output may include: generating an image tag by segmenting a target image feature; generating a point cloud tag by segmenting a point cloud feature; generating first position information indicating a relative position of a corresponding image tag; generating second position information indicating a relative position of a corresponding point cloud tag; and executing a multimodal decoding model using key data and value data based on the image tag, the point cloud tag, the first position information and the second position information.

[0012] Generating the decoding output may also include executing the multimodal decoding model using the query data based on the detection guidance information, the detection guidance information indicating detection position candidates having a likelihood of detecting the target object in the 3D space.

[0013] The detection position candidates may be distributed non-uniformly.

[0014] The multimodal decoding model can generate decoding output by extracting correlations from target image features, point cloud features, and detection guidance information.

[0015] A non-transitory computer-readable storage medium may store instructions, which, when executed by a processor, cause the processor to perform any one of the methods described.

[0016] In another general aspect, an electronic device includes: one or more processors; and a memory storing instructions, wherein the instructions are configured to cause the one or more processors to perform the following operations: receive an input image relative to a three-dimensional 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space; use an encoding model to generate candidate image features for a partial area of ​​the input image, point cloud features of an input point cloud, and language features of the input language; select a target image feature corresponding to the language feature from among the candidate image features based on a similarity score of a similarity between the candidate image features and the language features; generate a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features; and detect a 3D bounding box corresponding to the target object by executing an object detection model based on the decoding output.

[0017] The instructions can also be configured to cause one or more processors to perform the following operations: generate language features corresponding to the input language by performing inference on the input language through a language encoding model, generate candidate image features corresponding to a partial area of ​​the input image by executing an image encoding model and a region proposal model based on the input image, and generate point cloud features corresponding to the input point cloud by executing a point cloud encoding model based on the input point cloud.

[0018] The instructions may also be configured to cause one or more processors to generate extended expressions, each extended expression comprising (i) a location field indicating geometric characteristics of a target object based on an input language, and (ii) a category field indicating a category of the target object, and wherein language features are generated based on the extended expression.

[0019] Objects of the same category with different geometric properties can be distinguished from each other based on the location field.

[0020] The location field can be learned through training.

[0021] The instructions can also be configured to cause one or more processors to perform the following operations: generate image tags by segmenting target image features, generate point cloud tags by segmenting point cloud features, generate first position information indicating the relative positions of corresponding image tags, generate second position information indicating the relative positions of corresponding point cloud tags, and execute a multimodal decoding model using key data and value data based on the image tags, point cloud tags, the first position information and the second position information.

[0022] The instructions may also be configured to cause the one or more processors to execute the multimodal decoding model using the query data based on detection guidance information indicating detection position candidates having a likelihood of detecting the target object in the 3D space.

[0023] The detection position candidates may be distributed non-uniformly.

[0024] The multimodal decoding model can be configured to generate decoding output by extracting correlations from target image features, point cloud features, and detection guidance information.

[0025] In another general aspect, a vehicle includes: a camera configured to generate an input image relative to a three-dimensional 3D space; a light detection and ranging (lidar) sensor configured to generate an input point cloud relative to the 3D space; one or more processors configured to: receive an input image relative to the 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space; use an encoding model to generate candidate image features for a partial area of ​​the input image, point cloud features of an input point cloud, and language features of the input language; select a target image feature corresponding to the language feature from among the candidate image features based on a similarity score of a similarity between the candidate image features and the language features; generate a decoded output by executing a multimodal decoding model based on the target image features and the point cloud features; and detect a 3D bounding box corresponding to the target object by executing an object detection model based on the decoded output; and a control system configured to control the vehicle based on the 3D bounding box.

[0026] Other features and aspects will become apparent from the following detailed description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Example configurations of three-dimensional (3D) object perception models are shown in accordance with one or more embodiments.

[0028] Figure 2 An example configuration of a visual language model is shown in accordance with one or more embodiments.

[0029] Figure 3 Examples of objects of the same category having different geometric characteristics are shown in accordance with one or more embodiments.

[0030] Figure 4 Example operations of a multimodal decoding model in accordance with one or more embodiments are shown.

[0031] Figure 5 An example training process for a visual language model in accordance with one or more embodiments is shown.

[0032] Figure 6 An example training process for a 3D object perception model in accordance with one or more embodiments is shown.

[0033] Figure 7 An example of a 3D object recognition method according to one or more embodiments is shown.

[0034] Figure 8 An example configuration of an electronic device according to one or more embodiments is shown.

[0035] Fig. 9An example configuration of a vehicle is shown in accordance with one or more embodiments.

[0036] Throughout the drawings and detailed description, unless otherwise described or provided, the same or similar reference numerals will be understood to refer to the same or similar elements, features, and structures. The drawings may not be drawn to scale, and the relative sizes, proportions, and depictions of elements in the drawings may be exaggerated for clarity, illustration, and convenience. DETAILED DESCRIPTION

[0037] The following detailed description is provided to help the reader obtain a comprehensive understanding of the method, device and / or system described herein. However, after understanding the disclosure of the present application, various changes, modifications and equivalents of the method, device and / or system described herein will be apparent. For example, the order of operations described herein is merely an example and is not limited to those order of operations set forth herein, but can be significantly changed after understanding the disclosure of the present application, except for operations that must be performed in a certain order. In addition, for greater clarity and brevity, the description of features known after understanding the disclosure of the present application may be omitted.

[0038] The features described herein may be implemented in different forms and should not be construed as being limited to the examples described herein. Rather, the examples described herein are provided merely to illustrate some of the many possible ways to implement the methods, devices and / or systems described herein, which will be apparent after understanding the disclosure of the present application.

[0039] The terms used in this article are only used to describe various examples and are not used to limit the present disclosure. Unless the context clearly indicates otherwise, the articles "one", "an" and "the" are also intended to include plural forms. As used herein, the term "and / or" includes any one of the items listed in the association and any combination of any two or more. As a non-limiting example, the term "comprises" or "comprising", "includes" or "includes", and "has" or "contains" indicates the presence of the features, quantities, operations, components, elements, and / or combinations thereof, but does not exclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0040] Throughout the specification, when a component or element is described as being "connected to," "coupled to," or "engaged to" another component or element, it may be directly "connected to," "coupled to," or "engaged to" the other component or element, or one or more other components or elements may reasonably be present in between. When a component or element is described as being "directly connected to," "directly coupled to," or "directly engaged to" another component or element, there may not be other elements in between. Similarly, for example, "between" and "directly between," as well as "adjacent to" and "immediately adjacent to" may also be interpreted as described above.

[0041] Although terms such as "first", "second" and "third", or A, B, (a), (b) may be used herein to describe various components, assemblies, regions, layers or parts, these components, assemblies, regions, layers or parts are not limited by these terms. For example, each of these terms is not used to define the nature, order or sequence of the corresponding component, component, region, layer or part, but is only used to distinguish the corresponding component, component, region, layer or part from other components, components, regions, layers or parts. Therefore, without departing from the teachings of the examples described herein, the first component, component, region, layer or part mentioned in the examples may also be referred to as the second component, component, region, layer or part.

[0042] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those commonly understood by those of ordinary skill in the art to which the present disclosure belongs, and as understood based on the understanding of the disclosure of the present application. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with the meaning in the context of the relevant technology and the disclosure of the present application, and should not be interpreted as an ideal or overly formal meaning, unless explicitly defined as such herein. In this article, the use of the term "may" with respect to examples or embodiments (e.g., with respect to what an example or embodiment may include or implement) indicates that there is at least one example or embodiment that includes or implements such features, and all examples are not limited thereto.

[0043] Figure 1 An example configuration of a three-dimensional (3D) object perception model according to one or more embodiments is shown. Figure 1 , the 3D object perception model 100 may include a visual-language model 110, a point cloud encoding model 120, a multimodal decoding model 130, and an object detection model 140. Here, "3D object perception" mainly refers to perception in three-dimensional space. As a non-limiting example, the visual-language model 110 may be implemented as a contrastive language-image pre-training (CLIP) model.

[0044] The 3D object perception model 100 may receive an input image 101 relative to a 3D space, an input language 102 relative to a target object in the 3D space, and an input point cloud 103 relative to the 3D space. The input language 102 may be a word or phrase in a human language. The target object may be an object that is the target of perception. The 3D object perception model 100 may detect a 3D bounding box 141 corresponding to the target object using a visual-language model 110, a point cloud encoding model 120, a multimodal decoding model 130, and an object detection model 140.

[0045] The 3D space may represent an actual / physical space (e.g., a space near a vehicle). The input image 101 may be the result of a camera capturing at least a portion of the 3D space. For example, the input image 101 may be a color image (e.g., an RGB image, an infrared image, a multi-band image, etc.). The camera may include sub-cameras having different views (positions and directions). The input image 101 may include sub-images corresponding to different views captured by the sub-cameras, respectively. For example, sub-images of views in six directions may be generated by sub-cameras of views in six directions. The input point cloud 103 may be the result of sensing at least a portion of the 3D space by a light detection and ranging (lidar) sensor (or RADAR, or any suitable sensor that can generate a 3D point cloud). The scene shown in the input image 101 and the scene shown in the input point cloud 103 may overlap at least partially. For example, a camera (or sub-camera) and a lidar sensor may be installed in a vehicle and may capture the surroundings of the vehicle at a 360-degree angle (without full circumferential coverage). The target object may be included in an area where the sensing areas of the camera and the lidar overlap. A virtual space corresponding to the 3D space may be defined by the input image 101 and the input point cloud 103. A 3D bounding box 141 may be formed in the virtual space.

[0046] The visual-language model 110 can learn the relationship between visual information (e.g., image information) and language information (e.g., text information or voice information), and can solve problems based on the relationship (e.g., identifying the 3D bounding box of a target object). The visual-language model 110 can learn various objects through visual information and language information, and can be used for open vocabulary object detection (VOD) according to the learning characteristics of the visual-language model 110. "Open VOD" generally refers to the ability to detect objects that have been vocabulary learned (trained) and objects that have not been vocabulary learned. In detail, the open VOD model can transfer the multimodal (e.g., image mode and point cloud mode) capabilities of the pre-trained visual-language model to object detection. Open VOD generally extends traditional object detection to open categories and eliminates the need for cumbersome annotations that often have to be done manually.

[0047] The 3D object perception model 100 can generate candidate image features of a partial area (region) of the input image 101, language features of the input language 102, and point cloud features 121 of the input point cloud 103 using the corresponding encoding models. The encoding model may include an image encoding model of the visual-language model 110, a language encoding model of the visual-language model 110, and a point cloud encoding model 120. The image encoding model, the language encoding model, and the point cloud encoding model 120 may be corresponding neural network models capable of generating features associated with the corresponding input data. For example, the encoding model may be a convolutional neural network (CNN) or a transformer encoder.

[0048] The visual-language model 110 may include a language encoding model (e.g., language encoding model 220), an image encoding model (e.g., image encoding model 240), and a region proposal model (e.g., region proposal model 250). The image encoding model may be executed based on the input image 101, and may generate (infer) image features of the input image 101, and the region proposal model may determine candidate image features corresponding to a partial region of the input image from the image features. The language encoding model may be executed based on (applied to) the input language 102, and may generate / infer language features corresponding to the input language 102. The visual-language model 110 may determine a similarity score of the similarity between the candidate image features and the language features, and may select a target image feature 111 corresponding to the language feature from the candidate image features, wherein the selection is made based on the similarity score. Because the target image feature 111 is selected by the language feature, the target image feature 111 may be referred to as a focus image feature. The point cloud encoding model 120 may be executed based on (applied to) the input point cloud 103, and may generate / infer a point cloud feature 121 corresponding to the input point cloud 103.

[0049] The multimodal decoding model 130 can be executed based on the target image features 111 (first mode) and the point cloud features 121 (second mode), and a decoding output can be generated. The multimodal decoding model 130 can analyze the correlation between the target image features 111 and the point cloud features 121, which have different modalities. The multimodal decoding model 130 can determine the shape of the target object from the target image features 111 corresponding to the two-dimensional (2D) image information, can identify points corresponding to the shape of the target object from the point cloud features 121, and can select a detection position corresponding to the position of the target object, the detection position is selected from the detection position candidates of the detection guide information 104, and the detection is based on the point position.

[0050] The multimodal decoding model 130 can be used to detect various objects based on the open VOD of the visual-language model 110. Typical prior object perception models can predetermine and learn the category of the target object (e.g., through training), and the addition of previously unlearned categories may require complete relearning of both models. That is, when learning new categories, typical object perception models require complete retraining. The open VOD of the visual-language model 110 can generate target image features 111 for various objects (including new categories) without requiring a new learning process (e.g., training). The multimodal decoding model 130 can utilize the target image features 111 to recognize the shapes of various objects, including objects of new categories, and can generate a decoding output to determine the 3D bounding box 141 of the object.

[0051] The multimodal decoding model 130 may be a transformer decoder. The multimodal decoding model 130 may perform decoding by extracting correlations from the query data, the key data, and the value data. The key data and the value data may be determined based on the target image features 111 and the point cloud features 121, and the query data may be determined based on the detection guidance information 104. The detection guidance information 104 may indicate detection position candidates having a likelihood of detecting the target object in the 3D space.

[0052] The object detection model 140 may be executed based on (applied to) the decoded output and detect a 3D bounding box 141 corresponding to the target object. The object detection model 140 may be a neural network model (eg, a multi-layer perceptron (MLP)).

[0053] Figure 2 An example configuration of a visual language model according to one or more embodiments is shown. Figure 2 , the visual-language model 200 may include a language encoding model 220, an image encoding model 240, and a region proposal model 250. The language encoding model 220, the image encoding model 240, and the region proposal model 250 may be implemented as corresponding neural network models.

[0054] The language encoding model 220 can generate language features 261 based on the input language 210. For example, the input language 210 can include text information (text or its representation) and / or speech information. The extended expression 211 can be generated based on the input language 210. The extended expressions 211 can each include (i) a position field 2111 indicating the geometric characteristics of the corresponding target object, and (ii) a category field 2112 indicating the category of the corresponding target object. The category field 2112 can have different values ​​for different categories (e.g., vehicles, people, traffic signals, and lanes). That is, in some cases, there can be multiple extended expressions 211 with the same position field 2111 but different category fields 2112.

[0055] The geometric characteristics of the position field 2111 may include a geometric position, a geometric shape, etc. The position field 2111 may have different values ​​for different positions (e.g., short distance, long distance, left side of the image, right side of the image, upper side of the image, lower side of the image, upper left side of the image, lower left side of the image, upper right side of the image, lower right side of the image). Objects of the same category with different geometric characteristics can be distinguished based on the position field 2111. For example, an object of the category "vehicle" on the upper side of the image can be distinguished from an object of the category "vehicle" on the lower side of the image based on the position field 2111. In another example, if "occlusion state" is included as part / subfield of the position field 2111, a vehicle whose position field 2111 indicates that the vehicle is in a complete state (i.e., "no occlusion") can be distinguished from a vehicle with "occlusion" in its position field 2111, even if the two vehicles are in the same position.

[0056] The position field 2111 may have a learnable property (may be trainable). Although a brief description is provided below, the value of the position field 2111 may be determined during the training process of the multimodal decoding model. The input language 210 may correspond to a prompt (supplemental input information that may guide the inference of the main input). Based on the context optimization of the training process, the value of the position field 2111 may be optimized to distinguish and detect various geometric properties of each object (such as occlusion / unocclusion). Objects of the same category with various geometric properties may appear, and since the multimodal decoding model distinguishes and learns objects of the same category with different geometric properties, the detection accuracy of the multimodal decoding model for objects of the same corresponding category may be improved.

[0057] The image encoding model 240 may generate image features corresponding to the input image 230. The region proposal model 250 may determine, based on the image features, candidate image features 262 corresponding to a partial region (area) of the input image 230. The partial region may be a region where an object is likely to exist.

[0058] Score table 260 may include similarity scores 263 between language features 261 and candidate image features 262. For example, similarity scores 263 may be determined based on a Euclidean distance between the language features and the candidate image features.

[0059] Sub-features SF_1, SF_2, and SF_3 of the language feature 261 may respectively correspond to the extended expression 211. For example, the first sub-feature SF_1 may correspond to the first extended expression of the extended expression 211, the second sub-feature SF_2 may correspond to the second extended expression of the extended expression 211, and the third sub-feature SF_3 may correspond to the third extended expression of the extended expression 211.

[0060] When a target object is determined / detected, an extended expression 211 with the target object as a category field 2112 may be configured. For example, when a vehicle is determined as a target object, an extended expression 211 at each position with the vehicle as a category may be configured. The extended expression 211 for each category may be determined during the training process of the multimodal decoding model. When the input language 210 specifies multiple categories, the extended expression 211 and the language feature 261 may be configured for each category.

[0061] Among the candidate image features 262, at least some of the candidate image features having the highest similarity scores with respect to the similarities with the sub-features SF_1, SF_2, and SF_3 of the language feature 261 may be selected as target image features. For example, when the third candidate image feature CIF_3 indicates / has the highest similarity score (SS_31) with the first sub-feature SF_1, the third candidate image feature CIF_3 may be selected as the target image feature for the first sub-feature SF_1. Similarly, when the fifth candidate image feature CIF_5 indicates / has the highest similarity score with the second sub-feature SF_2, the fifth candidate image feature CIF_5 may be selected as the target image feature for the second sub-feature SF_2. When the first candidate image feature CIF_1 indicates / has the highest similarity score with the third sub-feature SF_3, the first candidate image feature CIF_1 may be selected as the target image feature for the third sub-feature SF_3. As described above, the target image feature having the highest similarity score for each sub-feature may be selected, and a 3D bounding box corresponding to each target image feature may be determined.

[0062] Figure 3 Examples of objects of the same category having different geometric properties are shown in accordance with one or more embodiments. Figure 3, the input image 310 may include a vehicle 301 in a complete state at the lower left side of the input image 310, a vehicle 302 in a complete state at the upper side of the input image 310, and a vehicle 303 in an occluded state at the lower right side of the input image 310. When only the category of the vehicle (among other possible characteristics) is distinguishable through the input language, all of the vehicles 301, 302, and 303 having various characteristics may not be recognized. According to one or more embodiments, since not only the category “vehicle” but also various geometric characteristics can be distinguished through the input language, all of the vehicles 301, 302, and 303 having various characteristics can be accurately recognized.

[0063] Figure 4 FIG. 1 shows an example operation of a multimodal decoding model according to one or more embodiments. Figure 4 , the target image feature 401 may be segmented into image tokens 402, and the point cloud feature 403 may be segmented into point cloud tokens 404 (the tokens are in the form of extracted features). The position information 405 may include first position information indicating the relative position of the corresponding image token 402, and second position information indicating the relative position of the corresponding point cloud token 404. The relative position may be the position of each token in each image.

[0064] The key data and the value data (K, V) may be determined based on the image tag 402, the point cloud tag 404, the first position information of the position information 405, and the second position information of the position information 405. The key data may be the same as the value data. For example, matching pairs according to the relative positions of the image tag 402 and the point cloud tag 404 may be combined (e.g., concatenated), and the key data and the value data may be sequentially configured by adding the first position information and / or the second position information to the combined matching pairs.

[0065] The detection guidance information 406 may indicate a detection position candidate having a possibility of detecting a target object in a 3D space. The detection guidance information 406 may configure query data (Q). The detection position candidate may indicate a non-uniform position. Although a brief description is provided below, the detection position candidate may be optimized by the detection guidance model during the training process of the multimodal decoding model 410. The detection position candidate may indicate a non-uniform position as an optimization result.

[0066] The multimodal decoding model 410 may correspond to a transformer decoder. The multimodal decoding model 410 may be executed based on (i.e., applied to) the query data, the key data, and the value data, and may generate / infer a decoded output. The multimodal decoding model 410 may generate a decoded output by extracting correlations from the target image features 401, the point cloud features 403, and the detection guidance information 406 based on the query data, the key data, and the value data. The object detection model 420 may be executed based on (applied to) the decoded output, and may determine a 3D bounding box 421. The 3D bounding box 421 may correspond to one of the detection position candidates.

[0067] Figure 5 An example of a training process of a visual-language model according to one embodiment is shown. Figure 5 , the visual-language model 500 may include a language encoding model 520 , an image encoding model 540 , and a region proposal model 550 .

[0068] The language encoding model 520 may generate language features 561 based on the input language 510 (e.g., words / phrases in the form of text data; as used herein, "words" include short phrases). The input language 510 may specify / name corresponding categories. The input language 510 may not include, for example, Figure 2 The position field 2111 may contain position information of the input image 530. For example, the input language 510 may be a prompt describing objects appearing in the input image 530. These objects may correspond to the target object.

[0069] The image coding model 540 may generate image features corresponding to the input image 530. The region proposal model 550 may determine candidate image features 562 corresponding to partial regions (regions, which may not overlap) of the input image 530 based on the image features. The partial regions may be regions where objects are likely to exist. The region position loss 571 may represent the difference between the true (GT) object region and the object region proposed by the region proposal model 550. The region proposal model 550 may acquire the ability to propose regions where objects are likely to exist based on training based on the region position loss 571.

[0070] The language features 561 may respectively correspond to the input languages ​​510. For example, the first language feature LF_1 may correspond to a first input language of the input languages ​​510, the second language feature LF_2 may correspond to a second input language of the input languages ​​510, and the third language feature LF_3 may correspond to a third input language of the input languages ​​510.

[0071] The candidate image features 562 may be arranged to match the language features. For example, the image feature corresponding to the category of the first language feature may be determined as the first candidate image feature CIF_1, the image feature corresponding to the category of the second language feature may be determined as the second candidate image feature CIF_2, and the image feature corresponding to the category of the third language feature may be determined as the third candidate image feature CIF_3.

[0072] The score table 560 may include a similarity score 563 of the similarity between the language feature 561 and the candidate image feature 562. For example, the similarity score 563 may be determined based on the Euclidean distance. The similarity score 563 may be trained so that the diagonal elements may have a large value based on the alignment loss 572, and the non-diagonal elements may have a small value. For example, in the score table 560, SS_11, SS_22, and SS_33 may correspond to the diagonal elements, and SS_12, SS_13, SS_21, SS_23, SS_31, and SS_32 may correspond to the non-diagonal elements. The similarity score values ​​of the diagonal elements may be increased according to the training of the alignment loss 572, and the similarity score values ​​of the non-diagonal elements may be reduced. As a result of the training, the score table 560 may be a diagonal matrix, or may be close to a diagonal matrix.

[0073] According to the training based on the alignment loss 572, pairs of language features and candidate image features relative to the same object (associated with the same object) (e.g., pairs of LF_1 and CIF_1, pairs of LF_2 and CIF_2, and pairs of LF_3 and CIF_3) can have similar feature values, and the language encoding model 520 and the image encoding model 540 can be trained to generate language features 561 and candidate image features 562 so that corresponding pairs have similar feature values. Through such training, the visual-language model 500 can obtain the ability to solve the problem of open VOD through training based on the alignment loss 572.

[0074] Figure 6 An example training process of a 3D object perception model according to one or more embodiments is shown. Figure 6 The 3D object perception model 600 can generate a 3D bounding box 651 corresponding to the input image 601, input language 602, input point cloud 603 and guidance coordinate information 604 based on the visual-language model 610, the point cloud model 620, the multimodal decoding model 630, the detection guidance model 640 and the object detection model 650.

[0075] The visual-language model 610 may generate / infer target image features 611 based on the input image 601 and the input language 602. The point cloud encoding model 620 may generate / infer point cloud features 621 based on the input point cloud 603. The detection guidance model 640 may generate / infer detection guidance information from the guidance coordinate information 604. The multimodal decoding model 630 may generate / infer a decoding output based on receiving as input the target image features 611, the point cloud features 621, and the detection guidance information. The object detection model 650 may determine a 3D bounding box 651 based on the decoding output.

[0076] The 3D object perception model 600 can be trained based on the box position loss 661 of the 3D bounding box 651. The box position loss 661 can correspond to the difference between the 3D bounding box 651 and the GT bounding box. In this case, when the visual-language model 610 (e.g., language encoding model, image encoding model, and region proposal model), the category field 6022, and the point cloud encoding model 620 are frozen, the position field 6021, the multimodal decoding model 630, the detection guidance model 640, and the object detection model 650 can be trained. For example, the visual-language model 610 can be trained based on the reference Figure 5 The state of the pre-trained vision-language model 610 is installed in the 3D object perception model 600 in the described manner.

[0077] The multimodal decoding model 630 can be used to detect various objects based on the open VOD of the visual-language model 610. Typical previous object perception models can predetermine and learn the categories of target objects, and the addition of new categories may require complete retraining of the model (including training for already learned categories). The open VOD of the visual-language model 610 can generate target image features 611 for various objects including new categories without the need for such a new learning process. The multimodal decoding model 630 can utilize the target image features 611 to recognize the shapes of various objects including new objects, and can generate a decoding output to determine the 3D bounding box 651 of the object.

[0078] The input language 602 may be extended to include an expression of a corresponding position field 6021 and a category field 6022. The position field 6021 may have a learnable property, and the category field 6022 may be frozen (not learned). For example, the position field 6021 may include a learnable parameter having a predetermined size, and the parameter may be trained to reduce the box position loss 661 during the training process of the multimodal decoding model 630.

[0079] The position field 6021 can be initialized to an arbitrary value and can be contextually optimized using a box position loss 661 to include spatial information. As a result of learning, the position field 6021 and the category field 6022 can be paired, and a spatial identifier can be assigned to each object of the same category with different geometric properties. This training method can not encode spatial information in an external way, but can enable the 3D object perception model to learn spatial information in an intrinsic way using 3D labels, so that many queries can be initialized at statistically significant locations by the detection guidance model 640 compared to traditional methods where anchors are distributed at equal / regular intervals.

[0080] The detection guidance model 640 may generate detection guidance information based on the guidance coordinate information 604. The detection guidance information may indicate a detection position candidate having a possibility of detecting a target object of the input language 602 in the 3D space. The guidance coordinate information 604 may indicate a uniform position relative to the 3D space, and the detection guidance information may indicate a non-uniform position relative to the 3D space (by adjusting the uniform position by the detection guidance model 640). Based on the training of the detection guidance model 640, the uniform position of the guidance coordinate information 604 may be adjusted to a non-uniform position of the detection guidance information. The detection guidance model 640 may adjust the uniform position to a non-uniform position so that it (via an influence on the multimodal decoding model 630) provides a high possibility of detecting the target object based on the training data set.

[0081] The detection guidance model 640 may be a neural network model (eg, 1×1 CNN). The dimension of the guidance coordinate information 604 may be adjusted by the detection guidance model 640 to match the input dimension of the multimodal decoding model 630 .

[0082] Figure 7 An example of a 3D object recognition method according to one or more embodiments is shown. Figure 7 In operation 710, the electronic device may receive an input image relative to a 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space. The electronic device may (i) use a coding model to generate candidate image features of a partial area of ​​the input image, point cloud features of the input point cloud, and language features of the input language in operation 720, (ii) select a target image feature corresponding to the language feature from the candidate image features based on a similarity score between the candidate image features and the language features in operation 730, (iii) generate a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features in operation 740, and (iv) detect a 3D bounding box corresponding to the target object by executing an object detection model based on the decoding output in operation 750.

[0083] Operation 720 may include: (i) an operation of generating / inferring language features corresponding to the input language by executing a language encoding model based on (applying the model to) the input language, (ii) an operation of generating candidate image features corresponding to a partial region of the input image by executing an image encoding model and a region proposal model based on the input image, and (iii) an operation of generating point cloud features corresponding to the input point cloud by executing a point cloud encoding model based on the input point cloud.

[0084] An extended expression including a corresponding position field (indicating the geometric characteristics of the target object based on the input language) and a corresponding category field indicating the category of the target object may be generated, and a language feature may be generated based on the extended expression. Objects of the same category but with different geometric characteristics may be distinguished based on the position field. The position field may have a learnable characteristic (may be learnable).

[0085] Operation 740 may include: (i) an operation of generating an image marker by segmenting a target image feature, (ii) an operation of generating a point cloud marker by segmenting a point cloud feature, an operation of generating first position information indicating a relative position of the image marker, (iii) an operation of generating second position information indicating a relative position of the point cloud marker, and (iv) an operation of executing a multimodal decoding model using key data and value data based on input (e.g., image marker, point cloud marker, first position information and second position information).

[0086] Operation 740 may include: performing an operation of a multimodal decoding model using query data based on detection guidance information indicating detection position candidates having a likelihood of detecting a target object in 3D space. The detection position candidates may indicate non-uniform positions (non-uniformly distributed position candidates). The multimodal decoding model may generate a decoding output by extracting correlations from target image features, point cloud features, and detection guidance information.

[0087] Figure 8 An example configuration of an electronic device according to one or more embodiments is shown. Figure 8 , the electronic device 800 may include a processor 810 and a memory 820. Figure 8 Although not shown in the figure, the electronic device 800 may also include other devices, such as cameras, lidar sensors, storage devices, input devices, output devices, and network devices. For example, the electronic device 800 may correspond to any device (e.g., a vehicle, a robot, or a mobile device) that can use a camera or a sensor (e.g., a lidar) to recognize the surrounding environment.

[0088] The memory 820 may be connected to the processor 810 and may store instructions executable by the processor 810, data to be calculated by the processor 810, data processed by the processor 810, or a combination thereof. In practice, the processor 810 may be a combination of the processors mentioned below. The memory 820 may include a non-transitory computer-readable medium (e.g., a high-speed random access memory) and / or a non-volatile computer-readable medium (e.g., a disk storage device, a flash memory device, or other non-volatile solid-state storage device).

[0089] The processor 810 may execute instructions to perform the above reference Figures 1 to 7 and Fig. 9 The operations described. For example, when the instructions are executed by the processor 810, the electronic device 800 may receive an input image relative to the 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space, and may use the encoding model to generate candidate image features of a partial area of ​​the input image, point cloud features of the input point cloud, and language features of the input language, and may select a target image feature corresponding to the language feature from the candidate image features based on a similarity score between the candidate image features and the language features, and may generate a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features, and may detect a 3D bounding box corresponding to the target object by executing an object detection model based on the decoding output. The input image may be generated by a camera, and the input point cloud may be generated by a lidar sensor. The input language may be set by a user, or may be set in advance.

[0090] Fig. 9 An example configuration of a vehicle according to one or more embodiments is shown. Fig. 9 , the vehicle 900 may include a camera 910, a lidar sensor 920, a processor 930, and a control system 940. Fig. 9 Not shown, the vehicle 900 may also include other devices (eg, storage devices, memory, input devices, output devices, network devices, and drive systems).

[0091] The camera 910 may generate an input image relative to the 3D space, and the lidar sensor 920 may generate an input point cloud relative to the 3D space. The input language may be set by the user, or may be preset. The processor 930 may execute instructions to perform the above reference Figures 1 to 8Operations described. For example, the processor 930 may receive an input image relative to a 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space, and may use a coding model to generate candidate image features for a partial area of ​​the input image, point cloud features of the input point cloud, and language features of the input language, and may select a target image feature corresponding to the language feature from the candidate image features based on a similarity score between the candidate image features and the language features, and may generate a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features, and may detect a 3D bounding box corresponding to the target object by executing an object detection model based on the decoding output. The control system 940 may control the vehicle 900 and / or the drive system based on the object perception result. For example, the control system 940 may control the speed and / or steering of the vehicle 900 based on the 3D object perception result (e.g., a 3D bounding box).

[0092] Computing devices, vehicles, electronic devices, processors, memories, image sensors, vehicle / operating function hardware, ADAS / AD systems, displays, information output systems and hardware, storage devices, and the like referenced herein Figures 1 to 9Other devices, equipment, units, modules and components described are implemented or represented by hardware components. Examples of hardware components that can be used to perform the operations described in the present application include controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in the present application, where appropriate. In other examples, one or more hardware components performing the operations described in the present application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements (e.g., logic gate arrays, controllers and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve the desired result). In one example, a processor or computer includes (or is connected to) one or more memories storing instructions or software executed by a processor or computer. The hardware components implemented by a processor or a computer can execute instructions or software, for example, an operating system (OS) and one or more software applications running on the OS to perform the operations described in this application. The hardware components can also access, manipulate, process, create and store data in response to the execution of instructions or software. For the sake of brevity, the singular term "processor" or "computer" can be used in the description of the examples described in this application, but multiple processors or computers can be used in other examples, or the processor or computer can include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors or another processor and another controller. One or more processors, or a processor and a controller can implement a single hardware component, or two or more hardware components. The hardware components may have any one or more of different processing configurations, examples of which include a single processor, independent processors, parallel processors, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.

[0093] Figures 1 to 9The method for performing the operation described in the present application shown in the is performed by computing hardware, for example, by one or more processors or computers implemented as described above, implementing instructions or software to perform the operation described in the present application (the operation performed by the method). For example, a single operation, or two or more operations can be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations can be performed by one or more processors, or a processor and a controller, and one or more other operations can be performed by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller can perform a single operation, or two or more operations.

[0094] For controlling computing hardware (e.g., one or more processors or computers) to implement hardware components and perform the instructions or software of the method as described above can be written as a computer program, code segment, instruction or any combination thereof, for indicating or configuring one or more processors or computers individually or collectively to operate as a machine or a special-purpose computer to perform the operations performed by the above-mentioned hardware components and methods. In one example, the instruction or software includes a machine code directly executed by one or more processors or computers, such as a machine code generated by a compiler. In another example, the instruction or software includes a higher-level code executed by one or more processors or computers using an interpreter. Can be based on the block diagrams and flow charts shown in the accompanying drawings, and the corresponding descriptions herein (which disclose algorithms for performing the operations performed by the hardware components and the methods as described above), use any programming language to write instructions or software.

[0095] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement hardware components and perform the methods described above, as well as any associated data, data files, and data structures, may be recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), flash memory, card-type storage such as multimedia card micro or card (e.g., secure digital (SD) or extreme digital (XD)), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured as follows: store instructions or software and any associated data, data files and data structures in a non-temporary manner, and provide instructions or software and any associated data, data files and data structures to one or more processors or computers so that one or more processors or computers can execute the instructions. In one example, the instructions or software and any associated data, data files and data structures are distributed on a networked computer system so that one or more processors or computers store, access and execute the instructions and software and any associated data, data files and data structures in a distributed manner.

[0096] Although the present disclosure includes specific examples, it will be apparent after understanding the disclosure of the present application that various changes in form and detail may be made to these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein should be considered to be descriptive only and not for limiting purposes. The description of features or aspects in each example is considered to be applicable to similar features or aspects in other examples. Suitable results may be achieved if the described techniques are performed in a different order and / or if components in the described systems, architectures, devices, or circuits are combined in different ways and / or replaced or supplemented by other components or their equivalents.

[0097] Therefore, the scope of the disclosure may be defined by the claims and their equivalents in addition to the above disclosure, and all variations within the scope of the claims and their equivalents are to be construed as being included in the disclosure.

Claims

1. A method for detecting a three-dimensional 3D object, the method comprising: Receiving an input image relative to a 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space; Using the encoding model to generate candidate image features of the partial area of ​​the input image, point cloud features of the input point cloud, and language features of the input language; selecting a target image feature corresponding to the language feature from among the candidate image features based on a similarity score of a similarity between the candidate image feature and the language feature; Generating a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features; as well as A 3D bounding box corresponding to the target object is detected by executing an object detection model based on the decoded output.

2. The method according to claim 1, wherein: Generating the candidate image features, the point cloud features, and the language features includes: Inferring the input language through a language encoding model to generate the language features corresponding to the input language; Generating the candidate image features corresponding to the partial region of the input image by executing an image coding model and a region proposal model based on the input image; and The point cloud features corresponding to the input point cloud are generated by executing a point cloud encoding model based on the input point cloud.

3. The method according to claim 1 further comprises generating extended expressions, each extended expression comprising (i) a position field indicating a geometric characteristic of the target object based on the input language, and (ii) a category field indicating a category of the target object, and in, The language feature is generated based on the expanded expression.

4. The method according to claim 3, wherein: Objects of the same category and having different geometrical properties are distinguished from each other based on the position field of the extended expression.

5. The method according to claim 3, wherein: The position field is learned through training.

6. The method according to claim 1, wherein: Generating the decoded output comprises: generating an image tag by segmenting the target image features; generating a point cloud marker by segmenting the point cloud features; generating first position information indicating relative positions of corresponding image markers; generating second position information indicating relative positions of corresponding point cloud markers; and The multimodal decoding model is executed using key data and value data based on the image tags, the point cloud tags, the first position information, and the second position information.

7. The method according to claim 6, wherein: Generating the decoded output further includes executing the multimodal decoding model using query data based on detection guidance information indicating detection position candidates having a likelihood of detecting the target object in the 3D space.

8. The method according to claim 7, wherein: The detection position candidates are non-uniformly distributed.

9. The method according to claim 7, wherein: The multimodal decoding model generates the decoding output by extracting correlations from the target image features, the point cloud features, and the detection guidance information.

10. A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, cause the processor to perform the method according to claim 1.

11. An electronic device, comprising: one or more processors; as well as A memory storing instructions configured to cause the one or more processors to: Receiving an input image relative to a three-dimensional 3D space, an input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space; Using the encoding model to generate candidate image features of the partial area of ​​the input image, point cloud features of the input point cloud, and language features of the input language; selecting a target image feature corresponding to the language feature from among the candidate image features based on a similarity score of a similarity between the candidate image feature and the language feature; Generating a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features; as well as A 3D bounding box corresponding to the target object is detected by executing an object detection model based on the decoded output.

12. The electronic device according to claim 11, wherein: The instructions are further configured to cause the one or more processors to: Inferring the input language through a language encoding model to generate the language features corresponding to the input language; Generating the candidate image features corresponding to the partial region of the input image by executing an image coding model and a region proposal model based on the input image; as well as The point cloud features corresponding to the input point cloud are generated by executing a point cloud encoding model based on the input point cloud.

13. The electronic device according to claim 11, wherein: The instructions are further configured to cause the one or more processors to generate extended expressions, each extended expression comprising (i) a position field indicating a geometric characteristic of the target object based on the input language, and (ii) a category field indicating a category of the target object, and The language feature is generated based on the extended expression.

14. The electronic device according to claim 13, wherein: Objects of the same category and having different geometrical properties are distinguished from each other based on the location field.

15. The electronic device according to claim 13, wherein: The position field is learned through training.

16. The electronic device according to claim 11, wherein: The instructions are further configured to cause the one or more processors to: generating an image tag by segmenting the target image features; generating a point cloud marker by segmenting the point cloud features; generating first position information indicating relative positions of corresponding image markers; generating second position information indicating relative positions of corresponding point cloud markers; as well as The multimodal decoding model is executed using key data and value data based on the image tags, the point cloud tags, the first position information and the second position information.

17. The electronic device according to claim 16, wherein: The instructions are further configured to cause the one or more processors to execute the multimodal decoding model using query data based on detection guidance information indicating detection position candidates having a likelihood of detecting the target object in the 3D space.

18. The electronic device according to claim 17, wherein: The detection position candidates are non-uniformly distributed.

19. The electronic device according to claim 17, wherein: The multimodal decoding model is configured to generate the decoding output by extracting correlations from the target image features, the point cloud features, and the detection guidance information.

20. A vehicle comprising: A camera configured to generate an input image relative to a three-dimensional 3D space; a light detection and ranging sensor configured to generate an input point cloud relative to the 3D space; One or more processors configured to: receiving the input image relative to the 3D space, the input point cloud relative to the 3D space, and an input language relative to a target object in the 3D space; Using the encoding model to generate candidate image features of the partial area of ​​the input image, point cloud features of the input point cloud, and language features of the input language; selecting a target image feature corresponding to the language feature from among the candidate image features based on a similarity score of a similarity between the candidate image feature and the language feature; Generating a decoding output by executing a multimodal decoding model based on the target image features and the point cloud features; as well as detecting a 3D bounding box corresponding to the target object by executing an object detection model based on the decoded output; as well as A control system is configured to control the vehicle based on the 3D bounding box.

Citation Information

Patent Citations

  • Pipe thinning estimating system and method

    KR1020230156519A