Integrated model for 3D object detections and neural reconstruction of object centers
The integrated object-centric neural reconstruction method addresses limitations in single-view 3D object reconstruction by combining neural radiance field techniques with pose estimation, enhancing efficiency and generalization for applications like autonomous driving and augmented reality.
Patent Information
- Application Number
- JP2025121514
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Existing single-view 3D object reconstruction techniques are limited by sparse views and dynamic objects, relying heavily on external 3D object detection, which incurs computational overhead and reduces efficiency and generalizability.
An integrated object-centric neural reconstruction approach combining neural radiance field techniques with pose estimation to generate refined poses and textures from single-view images, reducing dependency on external detection and enhancing computational efficiency and generalization.
The integrated approach achieves more efficient and generalizable 3D reconstruction by jointly predicting pose, shape, and texture, improving performance in autonomous driving, augmented reality, and robotics applications.
Smart Images

Figure 2026015314000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to artificial intelligence (AI) techniques for image recognition and processing. [Background technology]
[0002] background Various systems are configured to perform tasks using machine learning (ML) or other artificial intelligence (AI) techniques. For example, systems configured to perform image recognition, object detection, and / or other automated tasks may implement AI techniques. As one example, image detection systems and methods use various detection models trained for object and feature detection. Summary of the Invention [Means for solving the problem]
[0003] overview A method for performing pose estimation for an image includes, in one or more processing devices, receiving an input image; generating, based on the input image, a pose code corresponding to an estimated pose of an object in the input image; generating a box code corresponding to a bounding box of the object in the input image; performing pose estimation for the input image by generating a refined pose of the object using the pose code and the box code; generating a predicted output for the object in the input image based on the input image and the refined pose; and controlling one or more functions of the device based on the predicted output.
[0004] Other embodiments include a non-transitory computer-readable storage medium configured to store instructions that, when executed by a processor included in a computing device, cause the computing device to perform any of the various steps of the above-described methods. Further embodiments include a computing device configured to perform any of the various steps of the above-described methods. Further embodiments include an apparatus configured to perform any of the various steps of the above-described methods.
[0005] Other aspects and advantages of the present invention will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the described embodiments. [Brief explanation of the drawings]
[0006] [Figure 1] 1 illustrates generally a system for training a machine learning model in accordance with the principles of the present invention; [Figure 2] FIG. 1 illustrates generally a computer-implemented method for training and implementing a machine learning model in accordance with the principles of the present disclosure. [Figure 3A] 1 illustrates generally an audio data labeling system in accordance with the principles of the present disclosure; [Figure 3B] 1 illustrates generally a portion of a data acquisition system in accordance with the principles of the present disclosure; [Figure 3C] FIG. 1 illustrates generally an alternative audio data labeling system in accordance with the principles of the present disclosure. [Figure 4A] FIG. 1 illustrates an exemplary overall processing pipeline for a visual model, in accordance with the principles of the present disclosure. [Figure 4B] FIG. 1 illustrates an exemplary pose estimation module in accordance with the principles of the present disclosure. [Figure 4C] FIG. 1 illustrates an exemplary integrated model or pipeline for both training and inference, in accordance with the principles of the present disclosure. [Figure 4D]FIG. 1 illustrates steps of an exemplary method for implementing a visual model (e.g., training with a visual model and subsequently performing pose estimation) in accordance with the principles of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram illustrating the interaction between a computer-controlled device and a control system in accordance with the principles of the present disclosure. [Figure 6] FIG. 6 is a schematic diagram illustrating the control system of FIG. 5 configured to control a vehicle, which may be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, in accordance with the principles of the present disclosure. [Figure 7] 6 is a schematic diagram illustrating the control system of FIG. 5 configured to control a manufacturing machine, such as a punch cutter, cutter, or gun drill, of a manufacturing system, such as part of a production line. [Figure 8] 6 is a schematic diagram illustrating the control system of FIG. 5 configured to control a power tool, such as a power drill or driver, having an at least partially autonomous mode. [Figure 9] FIG. 6 is a schematic diagram illustrating the control system of FIG. 5 configured to control an automated personal assistant. [Figure 10] 6 is a schematic diagram illustrating the control system of FIG. 5 configured to control a surveillance system such as an access control system or a sentry system. [Figure 11] 6 is a schematic diagram illustrating the control system of FIG. 5 configured to control an imaging system, such as an MRI machine, an X-ray imaging machine, or an ultrasound machine. DETAILED DESCRIPTION OF THE INVENTION
[0007] Detailed Description Embodiments of the present disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples, and that other embodiments may take various alternative forms. The drawings are not necessarily to scale, and some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching those skilled in the art how to employ the embodiments in various ways. As will be understood by those skilled in the art, various features shown in and described with reference to any one of the drawings can be combined with features shown in one or more other drawings to create embodiments not explicitly shown or described. The combinations of illustrated features provide representative embodiments for typical applications. However, various combinations and modifications of features consistent with the teachings of the present disclosure may be desired for particular applications or implementations.
[0008] As used herein, "a," "an," and "the" refer to both singular and plural referents unless the context clearly dictates otherwise. By way of example, a "processor" programmed to perform various functions may refer to one processor programmed to perform every respective function, or may refer to two or more processors collectively programmed to perform each of the various functions.
[0009] As used herein, “content” can refer to original content (e.g., data representing captured images, video, audio, text, etc.) or synthesized content (e.g., synthesized images, video, audio, text, etc.) corresponding to input data. In some examples, “content” can include images that can correspond to captured images, synthesized images, or a combination thereof. An image may be represented by image data. In some contexts herein, the terms “image” and “image data” may be used interchangeably and can refer to actual pixel values, color channels, vectors, and / or binary data that correspond to the visual content of an image. In one example, “image” and / or “image data” refer to a raw representation of an image, such as an array of numbers representing pixel intensities, and in some examples may include preprocessed data resulting from an image sensor. Conversely, “metadata” or “image metadata” can refer to relevant or supplemental details about an image, such as image size, format, creation date, geolocation data, etc. In various examples, “image” and “image data” can further include metadata, but do not necessarily do so.
[0010] Various systems are configured to perform tasks using machine learning (ML) or other artificial intelligence (AI) techniques (e.g., ML or other AI models). For example, systems configured to perform image recognition, object detection, and / or other automated tasks may implement AI techniques. As one example, image detection systems and methods use various detection models (e.g., vision models) trained for object and feature detection.
[0011] Some vision models are configured to generate images of three-dimensional (3D) objects from two-dimensional (2D) images (e.g., reconstruct a 3D object from a single 2D image), which may be referred to as single-view 3D object reconstruction. Single-view 3D object reconstruction is an important technology with a wide range of applications, including, but not limited to, autonomous driving, augmented reality / virtual reality (AV / VR) systems, robotics, and embodied AI. Single-view 3D object reconstruction techniques are limited by the constraints of the primary data source, which may include sparse views and dynamic objects.
[0012] The visual model can implement neural radiance field ("NeRF") techniques to perform 3D reconstruction, which offers the particular advantage of presenting a scene at fine resolution and generating novel view images from the reconstructed scene. In some embodiments, object-centric NeRF techniques further increase the flexibility of novel data synthesis. However, object-centric NeRF methods impose stringent requirements on multi-view observation and accurate object poses and / or rely heavily on third-party object detection to provide initial object poses. The dependency on external 3D object detection incurs computational overhead in both training and deployment.
[0013] Systems and methods according to the present disclosure implement a vision model configured to perform integrated object-centric reconstruction (e.g., object-centric NeRF 3D reconstruction) techniques. In particular, the systems and methods described herein combine object-centric neural reconstruction and pose estimation to obtain more efficient and generalizable reconstruction results. In one example, the vision module includes, implements, includes, and / or communicates with a posterior estimation module configured to generate updated pose data (“refined pose” or “refined pose data”) for an object in an input image from the input image and the object's input / current pose (i.e., the pose of the object shown in the input image). As used herein, “pose” refers to the position and orientation of an object in 3D space relative to a camera. As used herein, “camera” may refer to the camera that captured the image. Thus, the estimated or calculated pose or pose data may include coordinates defining the pose of the object (e.g., X, Y, and Z coordinates in a 3D coordinate space), an angular heat map, a bounding box (with rotation / rotation angles), etc. In the examples described herein, the pose or pose data includes a bounding box. A "refined" pose, as distinguished from an input or current pose, is a pose (and / or corresponding pose data) that is predicted or calculated for a different range (i.e., distance from the object), angle, orientation, etc.
[0014] FIG. 1 illustrates an exemplary system 100 for training an ML or other AI model, such as a vision model according to the present disclosure. As used herein, for brevity, a “vision” model may refer to a pose estimation model or module, a vision model, or the like, configured to perform pose estimation in accordance with the techniques of the present disclosure. The system 100 may be configured to implement (and / or may include circuitry configured to implement) the systems and methods of the present disclosure, which are described in more detail below. The system 100 may include an input interface for accessing training data 102 for the vision model. For example, as shown in FIG. 1 , the input interface may be comprised of a data storage interface 104 that can access the training data 102 from data storage 106. For example, the data storage interface 104 may be a memory interface or a persistent storage interface, such as a hard disk or SSD interface, or a personal, local, or wide area network interface, such as a Bluetooth, Zigbee, or Wi-Fi interface, or an Ethernet or fiber optic interface. The data storage 106 may be an internal data storage of the system 100, such as a hard drive or SSD, but may also be an external data storage, for example a network-accessible data storage.
[0015] In some embodiments, data storage 106 may further include a data representation 108 of an untrained version of the visual model, which may be accessed from data storage 106 by system 100. However, it will be appreciated that training data 102 and data representation 108 of the untrained visual model may each be accessed from different data stores, for example, via different subsystems of data storage interface 104. Each subsystem may be of the type described above for data storage interface 104.
[0016] In some embodiments, the data representation 108 of the untrained visual model may be generated internally by the system 100 based on design parameters for the visual model and, therefore, may not be explicitly stored in the data storage 106. The system 100 may further include a processor subsystem 110 that may be configured, during operation of the system 100, to provide an iterative function on behalf of the layer stack of the visual model to be trained, where each layer of the substituted layer stack may have mutually shared weights and may receive as input the output of the preceding layer, or, for the first layer of the layer stack, initial activations and a portion of the input of the layer stack.
[0017] The processor subsystem 110 may be further configured to iteratively train the visual model using the training data 102, where the training iterations by the processor subsystem 110 may include a forward propagation portion and a backward propagation portion. Among other operations that may be performed defining the forward propagation portion, the processor subsystem 110 may be configured to perform the forward propagation portion by determining an equilibrium point of the iterative function where the iterative function converges to a fixed point, where determining this equilibrium point includes using a numerical root-search algorithm to find a root solution of the iterative function minus its input, and providing the equilibrium point in place of the output of the layer stack of the visual model. The processor subsystem 110 is configured to train the visual model in accordance with the systems and methods of the present disclosure, as described in more detail below.
[0018] System 100 may further include an output interface for outputting a data representation 112 of the trained visual model. This data may also be referred to as trained model data 112. For example, as also shown in FIG. 1 , the output interface may be constituted by data storage interface 104, which in these embodiments is an input / output (“IO”) interface through which trained model data 112 may be stored in data storage 106. For example, data representation 108 defining an “untrained” visual model may be at least partially replaced during or after training by data representation 112 of a trained visual model, in which visual model parameters such as weights, hyperparameters, and other types of visual model parameters may be adjusted to reflect training on training data 102. This is also indicated in FIG. 1 by reference numerals 108 and 112, which refer to the same data record on data storage 106. In some embodiments, data representation 112 may be stored separately from data representation 108 defining the “untrained” visual model. In some embodiments, the output interface may be separate from the data storage interface 104, but may generally be of the type described above for the data storage interface 104.
[0019] 2 illustrates an exemplary content generation system 200 configured to implement (and / or include circuitry configured to implement) a system for data annotation, augmentation, and / or generation. The content generation system 200 may include at least one computing system 202 configured to implement all or a portion of the systems and methods of the present disclosure, which are described in more detail below. The computing system 202 may include at least one processor 204 operatively connected to a memory unit 208. The processor 204 may include one or more integrated circuits that implement the functionality of a central processing unit (CPU) 206. The CPU 206 may be a commercially available processing unit that implements an instruction set, such as one of the x86, ARM, Power, or MIPS instruction set families. Various components of the system 200 may be implemented in the same or different circuits.
[0020] During operation, CPU 206 may execute stored program instructions retrieved from memory unit 208. The stored program instructions may include software that controls the operation of CPU 206 to perform the operations described herein. In some embodiments, processor 204 may be a system-on-chip (SoC) that integrates the functionality of CPU 206, memory unit 208, network interface, and input / output interface into a single integrated device. Computing system 202 may implement an operating system to manage various aspects of its operation.
[0021] The memory unit 208 may include volatile and nonvolatile memory for storing instructions and data. Nonvolatile memory may include solid-state memory such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is deactivated or powered down. Volatile memory may include static and dynamic random access memory (RAM) for storing program instructions and data. For example, the memory unit 208 may store one or more machine learning models or algorithms (e.g., represented in FIG. 2 as machine learning model 210), training dataset 212 for the machine learning model 210, raw source dataset 216, etc.
[0022] The computing system 202 may include a network interface device 222 configured to provide communication with external systems and devices. For example, the network interface device 222 may include a wired and / or wireless Ethernet interface defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 may include a cellular communication interface for communicating with a cellular network (e.g., 3G, 4G, 5G). The network interface device 222 may be further configured to provide a communication interface to an external network 224 or the cloud.
[0023] External network 224 may be referred to as the World Wide Web or the Internet. External network 224 may establish standard communication protocols between computing devices. External network 224 may facilitate the exchange of information and data between computing devices and networks. One or more servers 230 may be in communication with external network 224.
[0024] Computing system 202 may include input / output (I / O) interface 220, which may be configured to provide digital and / or analog input and output. I / O interface 220 may include an additional serial interface (e.g., a Universal Serial Bus (USB) interface) for communicating with external devices.
[0025] Computing system 202 may include a human-machine interface (HMI) device 218, which may include any device that allows system 200 to receive control input. Examples of input devices may include human interface inputs such as a keyboard, a mouse, a touchscreen, a voice input device, and other similar devices. Computing system 202 may include a display device 232. Computing system 202 may include hardware and software for outputting graphics and text information to display device 232. Display device 232 may include an electronic display screen, a projector, a printer, or other suitable device for displaying information to a user or operator. Computing system 202 may be further configured to enable interaction with a remote HMI and remote display device via network interface device 222.
[0026] System 200 may be implemented using one or more computing systems. While this example shows a single computing system 202 implementing all of the described features, it is contemplated that various features and functionality may be separated and implemented by multiple computing units that communicate with each other. The particular system architecture selected may depend on various factors.
[0027] The system 200 can implement a machine learning model 210 that analyzes a raw source dataset 216. For example, the CPU 206 and / or other circuitry can implement the machine learning model 210. The raw source dataset 216 can include raw or unprocessed sensor data that can represent an input dataset for a machine learning system. The raw source dataset 216 can include images, videos, video segments, audio, text-based information, and raw or partially processed sensor data (e.g., radar maps of objects). In some embodiments, the machine learning model 210 can include a deep learning algorithm or a neural network algorithm designed to perform a predetermined function. For example, a neural network algorithm can be configured to identify events or objects in an image or video segment based on audio data.
[0028] The computer system 200 may store a training dataset 212 for the machine learning algorithm 210. This training dataset 212 may represent a set of data previously constructed for training the machine learning algorithm 210. The training dataset 212 may be used by the machine learning algorithm 210 to learn various terms and coefficients (e.g., weighting factors) associated with the ML algorithm. The training dataset 212 may include a set of source data having corresponding outcomes or results that the machine learning model 210 attempts to replicate through the learning process.
[0029] The machine learning model 210 may be operated in a learning mode using a training dataset 212 as input. The machine learning model 210 may be run over multiple iterations using data from the training dataset 212. With each iteration, the machine learning model 210 may update its internal weighting coefficients based on the results achieved. For example, the machine learning model 210 may compare its output results (e.g., generated content) to the results contained in the training dataset 212. Because the training dataset 212 contains expected results, the machine learning model 210 can determine when performance is acceptable. After the machine learning model 210 reaches a predetermined level of performance (e.g., 100% agreement with the results associated with the training dataset 212), the machine learning model 210 may be run using data not contained in the training dataset 212. The trained machine learning model 210 may be applied to a new dataset to generate content. The machine learning model 210 may include a visual model trained according to the systems and methods of the present disclosure.
[0030] The machine learning model 210 may be configured to identify specific features in the raw source data 216. The raw source data 216 may include multiple instances or input data sets (e.g., images, video streams, or segments containing audio data) for which an output result is desired. For example, the machine learning model 210 may be configured to identify objects or features in an image, identify objects or events in a video segment based on the audio data, or the like. In some examples, the machine learning model 210 may be configured to annotate the identified objects, features, or events. The machine learning model 210 may be configured to perform pose estimation in accordance with the principles of the present disclosure. The machine learning model 210 may be programmed to process the raw source data 216 to identify the presence of specific features. The machine learning model 210 may be configured to identify a feature in the raw source data 216 as a predetermined feature. The raw source data 216 may be derived from a variety of sources. For example, the raw source data 216 may be actual input data collected by a machine learning system. The raw source data 216 may be machine-generated for testing the system. By way of example, the raw source data 216 may include raw image data, raw video and / or audio data from a camera, audio data from a microphone, etc.
[0031] In one example, the machine learning model 210 can process the raw source data 216 and output video and / or audio data that includes one or more indicators of an identified event. The machine learning model 210 can generate a confidence level or coefficient for each generated output. For example, a confidence value above a predetermined high confidence threshold can indicate that the machine learning model 210 is confident that the identified event (or feature) corresponds to a particular event. A confidence value below a low confidence threshold can indicate that the machine learning model 210 has some uncertainty about the existence of a particular feature.
[0032] As generally shown in FIGS. 3A and 3B , an exemplary system 300 may include an image (e.g., image and / or video) capture device 302, an audio capture array 304, and a computing system 202. The system may receive video stream data related to a data capture environment from the image capture device 302. The system 202 may be configured to perform video object detection to identify one or more objects in corresponding images of the video stream data. The system 202 may receive audio stream data corresponding to at least a portion of the video stream data from the audio capture array 304. The audio capture array 304 may include one or more microphones 306 or other suitable audio capture devices. The systems and methods described herein may be configured to label at least some objects in the video stream data and / or audio stream data using output from at least a first machine learning model (e.g., machine learning model 210 or other suitable machine learning model configured to provide output including one or more object or event detection predictions).
[0033] The system 202 can calculate at least one offset value for at least a portion of the audio stream data corresponding to the at least one labeled object in the video stream data based on the at least one data capture characteristic. The system 202 can use the at least one offset value to synchronize the at least a portion of the video stream data with the portion of the audio stream data corresponding to the at least one labeled object in the video stream data. The at least one data capture characteristic can include one or more characteristics of the at least one image capture device, one or more characteristics of the at least one audio capture array, one or more characteristics corresponding to a location of the at least one image capture device relative to the at least one audio capture array, one or more characteristics corresponding to a movement of an object in the video stream data, one or more other suitable data capture characteristics, or a combination thereof.
[0034] The system 202 may label at least a portion of the audio stream data corresponding to at least one labeled object in the video stream data using one or more labels of the labeled object in the video stream data and at least one offset value. Each respective label may include an event type, an event start indicator, and an event end indicator. The system 202 may generate training data using at least some of the labeled portions of the audio stream data. The system 202 may use the training data to train a second machine learning model. The system 202 may use the second machine learning model to detect one or more sounds associated with the audio data provided as input to the second machine learning model. The second machine learning model may include any suitable machine learning model and may be configured to perform any suitable function, such as those described herein with respect to FIGS. 4 through 11.
[0035] In some embodiments, as generally shown in FIG. 3C , computing system 202 may be configured to label audio data based on sensor data received from one or more sensors, such as, for example, one or more sensors as described herein or any other suitable sensor or combination thereof. System 202 may receive audio stream data related to the data capture environment from audio capture array 354 or from any suitable audio capture device, such as, for example, one or more microphones 306 or other suitable audio capture devices. It should be understood that audio capture array 354 may include features similar to those of audio capture array 304 and may include any suitable number of audio capture devices. System 202 may receive sensor data related to the data capture environment from at least one sensor (e.g., sensor 352) that is asynchronous relative to audio capture array 354. Sensor 354 may include at least one of an inductive coil, a radar sensor, a LiDAR sensor, a sonar sensor, an image capture device, any other suitable sensor, or a combination thereof. The audio capture array 354 may be located remotely from the sensor 354, may be located proximate to the sensor 354, or may be located in any suitable relationship to the sensor 354.
[0036] The system 202 can identify at least some events in the sensor data using output from at least a first machine learning model, such as output from the machine learning model 210 or other suitable machine learning model. The machine learning model 210 can be configured to provide an output including one or more event detection predictions based on the sensor data. The system 202 can synchronize at least a portion of the sensor data associated with a portion of the audio stream data corresponding to at least one event in the sensor data. The system 202 can label at least a portion of the audio stream data corresponding to at least one event in the sensor data using one or more labels extracted for each event in the sensor data values. Each respective label can include an event type, an event start indicator, and an event end indicator. The system 202 can generate training data using at least some of the labeled portions of the audio stream data. The system 202 can use the training data to train a second machine learning model. The system 202 can use the second machine learning model to detect one or more sounds associated with the audio data provided as input to the second machine learning model. The second machine learning model may include any suitable machine learning model and may be configured to perform any suitable function, such as those described herein with respect to Figures 4-11.
[0037] The systems and methods of the present disclosure (e.g., any of systems 100, 200, etc.) are configured to perform pose estimation and use a visual model to train a visual model (e.g., model 210) for generating refined poses, as described in more detail below. The techniques of the present disclosure may be referred to as unified NeRF ("UPNeRF") techniques, which provide a unified solution that jointly predicts the pose, shape, and texture of observed objects from a single network. The visual models of the present disclosure can be trained using real scenes (e.g., real driving scenes) with inaccurate predicted labels.
[0038] 4A illustrates an example overall processing pipeline 400 for a visual model 402 (e.g., the UPNeRF visual model) configured to perform pose estimation according to the present disclosure. For example, one or more computing devices, processors, or processing devices may be configured to execute instructions that implement the functionality of pipeline 400, e.g., one or more processors of the systems (e.g., 100, 200, etc.) described herein.
[0039] The vision model 402 is trained using a training dataset 404. This training dataset 404 includes a plurality of input images 406 (e.g., 2D images of an object, such as a vehicle) and corresponding poses of the object, which in this example are represented as bounding boxes 408 with rotations. As used herein, "with rotations" refers to data / values indicating one or more rotation angles of the bounding box 408. For example, the rotation angle indicates the orientation of the bounding box 408 (and the object) relative to the camera, the ground, etc. The bounding box 408 may be defined by data / values identifying one or more corner coordinates (e.g., using X, Y, and Z), a width and / or length of the bounding box 408, etc. In some examples, the training dataset 404 further includes a shape 410 or shape data. The shape 410 provided along with the images 406 indicates the overall shape, contour, form, etc. of the object in the image 406.
[0040] During and / or after training, the visual model 402 is provided with test images 412 (e.g., a set of test images of objects extracted from a scene 416). As shown, the test images 412 may be provided to the visual model 402 as additional inputs, along with occlusion masks, random poses (e.g., bounding boxes representing the random poses), etc. Based on the test images 412, the visual model 402 is configured to generate and output features such as texture, shape, and refined pose of the objects in the test images, as shown in step 418.
[0041] 4B illustrates an example pose estimation module 424 according to the present disclosure. As used herein, pose estimation module 424 may correspond to a model implemented by visual model 402, a model separate from visual model 402, circuitry configured to perform pose estimation functions or techniques, etc. For example, one or more computing devices, processors, or processing devices are configured to execute instructions to implement the functionality of pose estimation module 424, such as one or more processors of the systems (e.g., 100, 200, etc.) described herein. Pose estimation module 424 is configured to provide reliable poses for target objects at multiple ranges and orientations and to perform robustly under various conditions (e.g., for occluded images in which at least a portion of the target object is occluded / hidden).
[0042] 4B, the pose estimation module 424 iteratively updates the input pose 426 based on the visual difference between the input pose 426 and the object observed in the input image 428. B ,W B ,L B Given the pose (t) and pose (t) (e.g., height, width, and length, respectively), the camera eigenvalues K, and the current pose (t), (t) (corresponding to rotation and translation, respectively), the pose estimation module 424 obtains the image projections of the 3D box corners (t) (e.g., the coordinates of the eight corners of the bounding box 430). The box corners (t) correspond to a visual representation of the current (i.e., input) pose 426. In one example, (t) is a 16-bit vector. The box encoder 432 encodes (t) to generate a box code 436. For example, the box code 436 corresponds to a high-dimensional code or vector based on (t). In one example, the box code 436 is a 255-bit vector or other representation of (t).
[0043] The input image 428 is provided to an image encoder 438. The image encoder 438 is configured to generate and output an estimated pose, such as a pose code 440, based on the input image 428. For example, the pose code 440 is a code value or vector corresponding to the estimated pose. In one example, the postcode 400 is a low-dimensional code or vector (e.g., compressed) representation of the estimated pose obtained using principal component analysis (PCA) or other techniques.
[0044] The box code 436 and the attitude code 440 are provided as inputs to an attitude refiner 444. The attitude refiner 444 is configured to predict attitude updates 446 or attitude changes Δ(t), ΔT(t), which represent the changes to R(t) and T(t), respectively, of the input attitude 426. The attitude updates 446 predict the next (refined or updated) attitude or attitude state 448 (R (t+1) ,T (t+1) ) is combined with the input pose 426 to obtain the pose update 446 and updated pose 448. The generation of the pose update 446 and updated pose 448 is repeated multiple times (e.g., by providing the pose update 446 to the box encoder 432), which updates the box code 436 based on the pose update 446. The pose estimation module 424 continues to generate the pose updates 446 and updated pose 448 until a final pose state is obtained.
[0045] 4C illustrates an exemplary integrated model or pipeline (e.g., a visual model pipeline) 450 according to the present disclosure. For example, one or more computing devices, processors, or processing devices, such as one or more of the processors of the systems (e.g., 100, 200, etc.) described herein, are configured to execute instructions to implement the functionality of integrated pipeline 450. Integrated pipeline 450 illustrates both the training of the visual model and the inference functions performed by the visual model. For example, as shown in FIG. 4C, the flows of the training and inference processes are indicated by respective dashed lines, and flows common to both the training and inference processes are indicated by solid lines.
[0046] The pipeline 450 includes an image encoder 438 (e.g., an image encoder based on a residual network (ResNet)), a pose estimation module 424, and a NeRF decoder 452. The image encoder 438 receives an input image 428 (and its associated occlusion mask, together referred to as a masked input image). The image encoder 438 converts the masked input image 428 into a shape code 454 and a texture code 456 (e.g., respective code values or vectors corresponding to the estimated shape and texture), and a pose code 440. The pose code 440 derives the object pose R along with the box code 436. o2c |T o2c After multiple iterations, the estimated pose is provided to the pose refiner 444 as described above to iteratively refine the camera pose R c2o |T c2o which can be input to a NeRF decoder 452 for inference tasks, or can be used to compute the pose loss (L) during training, as shown in step 458.
[0047] During the inference task, the NeRF decoder 452 generates a predicted output 460 based on the shape code 454, texture code 456, and updated pose 448, which identifies detected objects in the input image, such as corresponding bounding boxes. In one example, the NeRF decoder 452 performs volume rendering to generate an RGB image (e.g., rendered RGB values) and an occupancy image (e.g., aggregated occupancy values). The rendered RGB values are compared with the input image 428 to calculate the photometric loss Lγgb, and the aggregated occupancy values are compared with an occupancy mask received with the input image 428 to calculate the occupancy loss Lγgb. occ is obtained. The total loss L infer L infer =L γgb +w occ L occ can be obtained according to, where w occ is the loss term L γgb and L occ is a weighting factor configured to balance the loss L infer is used to update the optimizable variables of the NeRF decoder 542, which are defined differently for inference and training.
[0048] In one example, pipeline 450 may include one or more multi-layer perceptrons (MLPs) 462. For example, MLPs 462 may be configured during training to convert pose codes 440 into higher-dimensional codes or vectors, which may be used to generate a direct pose loss.
number
number
[0049] By integrating an object detector and an object-centric neural reconstruction as described above, the visual model according to the present disclosure significantly improves computational efficiency and generalization capabilities.
[0050] 4D illustrates steps of an exemplary method 470 for implementing (e.g., training and subsequently performing pose estimation) a visual model according to the principles of the present disclosure. For example, one or more processors or processing devices may be configured to execute instructions for implementing the method 470, e.g., instructions for implementing one or more processors of the systems described herein.
[0051] At step 472, the method 470 includes training a visual model to perform pose estimation using a training set of images, masks, and poses (e.g., pose information or data such as bounding boxes). Training the visual model includes training the pose estimation module 424, as described above with respect to Figures 4A-4C.
[0052] In step 474, the method 470 includes receiving an input image at a visual model. In step 476, the method 470 includes generating a pose code based on the input image. In step 478, the method 470 includes iteratively generating a refined pose based on the pose code and the box code. For example, the method 470 may repeat step 476 to refine the pose (e.g., until the pose loss is below a threshold). In step 480, the method 470 includes generating a predicted output (e.g., using a NeRF decoder as described herein) based on the input image (e.g., based on the shape code and texture code) and the refined pose.
[0053] In step 482, the method 470 includes controlling one or more functions of a system, device, machine, etc. based on the predicted output and refined pose. For example, the predicted output and refined pose can be used for various downstream object detection and image recognition tasks, such as controlling autonomous vehicles, robotics, AR / VR systems, etc. In some examples, the method 470 includes controlling functions of any of the systems described below in FIGS. 5-11.
[0054] 5-11 illustrate example systems and devices capable of implementing a visual model, such as a pose estimation model, a visual model configured to perform pose estimation, and the like, according to the present disclosure. FIG. 5 illustrates a schematic diagram of interaction between a computer-controlled device 500 and a control system 502. In one example, the control system 502 is configured to control the computer-controlled device 500 by executing a visual model according to the principles of the present disclosure. The computer-controlled device 500 includes an actuator 504 and a sensor 506. The actuator 504 may include one or more actuators, and the sensor 506 may include one or more sensors. The sensor 506 is configured to sense a state of the computer-controlled device 500. The sensor 506 may be configured to encode the sensed state into a sensor signal 508 and transmit the sensor signal 508 to the control system 502. Non-limiting examples of the sensor 506 include video, radar, LiDAR, ultrasonic, and motion sensors. In some embodiments, sensor 506 is an optical sensor configured to sense an optical image of an environment proximate to computer-controlled device 500. A visual model according to the present disclosure can perform pose estimation on the optical image as described herein.
[0055] The control system 502 is configured to receive sensor signals 508 from the computer-controlled device 500. As described below, the control system 502 may be further configured to calculate actuator control commands 510 dependent on the sensor signals and to transmit the actuator control commands 510 to the actuators 504 of the computer-controlled device 500.
[0056] 5, the control system 502 includes a receiving unit 512. The receiving unit 512 may be configured to receive sensor signals 508 from the sensors 506 and convert the sensor signals 508 into input signals x. In an alternative embodiment, the sensor signals 508 are received directly as input signals x without the receiving unit 512. Each input signal x may be part of each sensor signal 508. The receiving unit 512 may be configured to process each sensor signal 508 to generate a respective input signal x. The input signals x may include data corresponding to an image recorded by the sensors 506.
[0057] The control system 502 includes a classifier 514. The classifier 514 may be configured to classify input signals x into one or more labels using a machine learning (ML) algorithm, such as a neural network, as described above. For example, the classifier 514 corresponds to the classifier 408 described above. The classifier 514 is configured to be parameterized by parameters (e.g., parameters θ) as described above. The parameters θ may be stored in and provided by non-volatile storage 516. The classifier 514 is configured to determine output signals y from the input signals x. Each output signal y includes information that assigns one or more labels to each input signal x. The classifier 514 may transmit the output signals y to a conversion unit 518. The conversion unit 518 is configured to convert the output signals y into actuator control commands 510. The control system 502 is configured to transmit actuator control commands 510 to the actuator 504, which is configured to operate the computer-controlled device 500 in response to the actuator control commands 510. In some embodiments, the actuator 504 is configured to operate the computer-controlled device 500 directly based on the output signal y.
[0058] When an actuator control command 510 is received by an actuator 504, the actuator 504 is configured to perform an action corresponding to the associated actuator control command 510. The actuator 504 may include control logic configured to convert the actuator control command 510 into a second actuator control command that is used to control the actuator 504. In one or more embodiments, the actuator control command 510 may be used to control a display instead of or in addition to the actuator.
[0059] In some embodiments, the control system 502 includes a sensor 506 instead of or in addition to a computer-controlled device 500 including a sensor 506. The control system 502 can also include an actuator 504 instead of or in addition to a computer-controlled device 500 including an actuator 504.
[0060] 5, the control system 502 includes a processor 520 and a memory 522. The processor 520 may include one or more processors. The memory 522 may include one or more memory devices. One or more embodiments of the classifier 514 (e.g., an ML algorithm) may be implemented by the control system 502, which includes the non-volatile storage 516, the processor 520, and the memory 522.
[0061] The non-volatile storage 516 may include one or more permanent data storage devices such as a hard drive, optical drive, tape drive, non-volatile solid-state device, cloud storage, or any other device capable of permanently storing information. The processor 520 may include one or more devices selected from a high-performance computing (HPC) system including a high-performance core, microprocessor, microcontroller, digital signal processor, microcomputer, central processing unit, field programmable gate array, programmable logic device, state machine, logic circuit, analog circuit, digital circuit, or any other device that manipulates signals (analog or digital) based on computer-executable instructions resident in memory 522. The memory 522 may include a single memory device or multiple memory devices, including, but not limited to, random access memory (RAM), volatile memory, non-volatile memory, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, cache memory, or any other device capable of storing information.
[0062] The processor 520 may be configured to execute computer-executable instructions loaded into memory 522 and resident in non-volatile storage 516 that embody one or more anomaly detection methodologies of one or more embodiments. The non-volatile storage 516 may include one or more operating systems and applications. The non-volatile storage 516 may store compiled and / or interpreted computer programs written using various programming languages and / or technologies, alone or in combination, including, but not limited to, Java, C, C++, C#, Objective C, Fortran, Pascal, JavaScript, Python, Perl, and PL / SQL.
[0063] When executed by processor 520, the computer-executable instructions in non-volatile storage 516 cause control system 502 to implement one or more of the anomaly detection schemes disclosed herein. Non-volatile storage 516 may also include data that supports the functions, features, and processes of one or more embodiments described herein.
[0064] Program code embodying the algorithms and / or methodologies described herein may be distributed individually or collectively as a program product in a variety of different forms. The program code may be distributed using a computer-readable storage medium having computer-readable program instructions for causing a processor to execute aspects of one or more embodiments. Computer-readable storage media that are non-transitory in nature may include volatile and non-volatile, removable and non-removable tangible media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media may also include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technology, portable compact disc read-only memory (CD-ROM), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be read by a computer. The computer-readable program instructions may be downloaded from the computer-readable storage medium into a computer, other type of programmable data processing device, or other device, or may be downloaded over a network to an external computer or external storage device.
[0065] Computer-readable program instructions stored on a computer-readable medium may be used to direct a computer, other type of programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored on the computer-readable medium produce an article of manufacture including instructions that implement the functions, acts, and / or operations specified in the flowcharts or diagrams. In certain alternative embodiments, the functions, acts, and / or operations specified in the flowcharts and diagrams may be reordered, processed sequentially, and / or processed simultaneously, consistent with one or more embodiments. Furthermore, any of the flowcharts and / or diagrams may include more or fewer nodes or blocks than those illustrated, consistent with one or more embodiments.
[0066] The processes, methods, or algorithms may be implemented in whole or in part using suitable hardware components, such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), state machines, controllers, or other hardware components or devices, or a combination of hardware, software, and firmware components.
[0067] 6 shows a schematic diagram of a control system 502 configured to control a vehicle 600, which may be an at least partially autonomous vehicle or an at least partially autonomous robot. In one example, the control system 502 is configured to control the vehicle 600 by executing a vision model in accordance with the principles of the present disclosure. The vehicle 600 includes actuators 504 and sensors 506. The sensors 506 may include one or more video sensors, cameras, radar sensors, ultrasonic sensors, LiDAR sensors, and / or location sensors (e.g., GPS). One or more of the one or more specific sensors may be integrated into the vehicle 600. Alternatively or in addition to the one or more specific sensors identified above, the sensors 506 may include a software module configured, upon execution, to determine the state of the actuators 504. One non-limiting example of a software module includes a weather information software module configured to determine current or future weather conditions proximate the vehicle 600 or other location.
[0068] The classifier 514 of the control system 502 of the vehicle 600 may be configured to detect an object in the vicinity of the vehicle 600 depending on the input signal x. In such an embodiment, the output signal y may include information characterizing the proximity of the object to the vehicle 600, such as a pose estimate obtained by a vision model. The actuator control commands 510 may be determined according to this information. The actuator control commands 510 may be used to avoid a collision with the detected object.
[0069] In some embodiments, vehicle 600 is an at least partially autonomous vehicle, and actuators 504 may be implemented in the brakes, propulsion system, engine, drivetrain, or steering of vehicle 600. Actuator control commands 510 may be determined to control actuators 504 to cause vehicle 600 to avoid a collision with a detected object. The detected object may be classified according to what classifier 514 considers to be most likely, such as a pedestrian or a tree. Actuator control commands 510 may be determined depending on the classification. In scenarios where adversarial attacks may occur, the system may be further trained to better detect objects or identify changes in lighting conditions or angles for sensors or cameras on vehicle 600.
[0070] In some embodiments where the vehicle 600 is an at least partially autonomous robot, the vehicle 600 may be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and stepping. The mobile robot may also be an at least partially autonomous lawnmower or an at least partially autonomous vacuum cleaner. In such embodiments, the actuator control commands 510 may be determined such that the propulsion, steering, and / or braking units of the mobile robot may be controlled to enable the mobile robot to avoid collision with the identified object.
[0071] In some embodiments, vehicle 600 is an at least partially autonomous robot in the form of a gardening robot. In such embodiments, vehicle 600 may use optical sensors as sensors 506 to determine the condition of plants in the environment immediately surrounding vehicle 600. Actuator 504 may be a nozzle configured to spray a chemical. Depending on the identified species and / or the identified condition of the plant, actuator control commands 510 may be determined to cause actuator 504 to spray an appropriate amount of an appropriate chemical on the plant.
[0072] Vehicle 600 may be an at least partially autonomous robot in the form of a household appliance. Non-limiting examples of a household appliance include a washing machine, a stove, an oven, a microwave, or a dishwasher. In such a vehicle 600, sensor 506 may be an optical sensor configured to detect a condition of an object being treated by the household appliance. For example, if the household appliance is a washing machine, sensor 506 may detect a condition of laundry in the washing machine. Actuator control command 510 may be determined based on the detected condition of the laundry.
[0073] 7 shows a schematic diagram of a control system 502 configured to control a system 700 (e.g., a manufacturing machine), such as a punch cutter, cutter, or gun drill, of a manufacturing system 702, such as part of a production line. The control system 502 may be configured to control an actuator 504 configured to control the system 700 (e.g., the manufacturing machine). In one example, the control system 502 is configured to control the system 700 by executing a visual model in accordance with the principles of the present disclosure.
[0074] The sensor 506 of the system 700 (e.g., a manufacturing machine) may be an optical sensor configured to capture one or more characteristics of the manufactured product 704. The classifier 514 may be configured to determine a state of the manufactured product 704 from the one or more captured characteristics. The actuator 504 may be configured to control the system 700 (e.g., a manufacturing machine) for a subsequent manufacturing step of the manufactured product 704 depending on the determined state of the manufactured product 704. The actuator 504 may be configured to control a function of the system 700 (e.g., a manufacturing machine) on a subsequent manufactured product 706 of the system 700 depending on the determined state of the manufactured product 704.
[0075] 8 shows a schematic diagram of a control system 502 configured to control a power tool 800, such as a power drill or driver, having at least a partially autonomous mode. The control system 502 may be configured to control an actuator 504 configured to control the power tool 800. In one example, the control system 502 is configured to control the power tool 800 by executing a visual model in accordance with the principles of the present disclosure.
[0076] The sensor 506 of the power tool 800 may be an optical sensor configured to capture one or more characteristics of the work surface 802 and / or the fastener 804 being driven into the work surface 802, which may include using a visual model to perform pose estimation. The classifier 514 may be configured to determine a state of the work surface 802 and / or the fastener 804 relative to the work surface 802 from the one or more captured characteristics. The state may be that the fastener 804 is flush with the work surface 802. Alternatively, the state may be the hardness of the work surface 802. The actuator 504 may be configured to control the power tool 800 such that a drive function of the power tool 800 is adjusted depending on the determined state of the fastener 804 relative to the work surface 802 or depending on one or more captured characteristics of the work surface 802. For example, the actuator 504 may discontinue a drive function when the state of the fastener 804 is flush with the work surface 802. As another non-limiting example, the actuator 504 may apply additional or less torque depending on the hardness of the working surface 802 .
[0077] 9 shows a schematic diagram of a control system 502 configured to control an automated personal assistant 900 (e.g., a robot). The control system 502 may be configured to control an actuator 504 configured to control the automated personal assistant 900. The automated personal assistant 900 may be configured to control a household appliance such as a washing machine, stove, oven, microwave, or dishwasher. In one example, the control system 502 is configured to control the automated personal assistant 900 by executing a visual model according to the principles of the present disclosure.
[0078] The sensor 506 may be an optical sensor and / or an acoustic sensor. The optical sensor may be configured to receive a video image of the gestures 904 of the user 902. The acoustic sensor may be configured to receive voice commands of the user 902.
[0079] The control system 502 of the automated personal assistant 900 may be configured to determine actuator control commands 510 configured to control the system 502. The control system 502 may be configured to determine the actuator control commands 510 according to sensor signals 508 of the sensors 506, which may include performing pose estimation using a visual model. The automated personal assistant 900 is configured to transmit the sensor signals 508 to the control system 502. The classifier 514 of the control system 502 may be configured to execute a gesture recognition algorithm to identify a gesture 904 performed by the user 902, determine the actuator control commands 510, and transmit the actuator control commands 510 to the actuators 504. The classifier 514 may be configured to retrieve information from non-volatile storage in response to the gesture 904 and output the retrieved information in a form suitable for receipt by the user 902.
[0080] 10 shows a schematic diagram of a control system 502 configured to control a surveillance system 1000. The surveillance system 1000 may be configured to physically control access through a door 1002. The sensor 506 may be configured to detect a scene relevant to determining whether access is permitted. The sensor 506 may be an optical sensor configured to generate and transmit image and / or video data. Such data may be used by the control system 502 to detect human faces. In one example, the control system 502 is configured to control the surveillance system 1000 by executing a vision model in accordance with the principles of the present disclosure.
[0081] A classifier 514 of the control system 502 of the surveillance system 1000 may be configured to interpret the image and / or video data by matching it with known person identities stored in non-volatile storage 516, thereby determining the person's identity. The classifier 514 may be configured to generate an actuator control command 510 in response to interpreting the image and / or video data. The control system 502 is configured to transmit the actuator control command 510 to the actuator 504. In this embodiment, the actuator 504 may be configured to lock or unlock the door 1002 in response to the actuator control command 510. In some embodiments, non-physical, logical access control is also possible.
[0082] The surveillance system 1000 may be a surveillance system. In such an embodiment, the sensor 506 may be an optical sensor configured to detect a scene under surveillance, and the control system 502 is configured to control the display 1004. The classifier 514 is configured to classify the scene, e.g., determine whether the scene detected by the sensor 506 is suspicious. The control system 502 is configured to transmit actuator control commands 510 to the display 1004 in response to the classification. The display 1004 may be configured to adjust the displayed content in response to the actuator control commands 510. For example, the display 1004 may highlight objects deemed suspicious by the classifier 514. Utilizing embodiments of the disclosed system, the surveillance system may predict objects at predetermined points in the future.
[0083] 11 shows a schematic diagram of a control system 502 configured to control an imaging system 1100, such as an MRI device, an X-ray imaging device, or an ultrasound device. In one example, the control system 502 is configured to control the imaging system 1100 by implementing a vision model according to the principles of the present disclosure. The sensor 506 may be, for example, an imaging sensor. The classifier 514 may be configured to determine a classification of all or a portion of the sensed image. The classifier 514 may be configured to determine or select an actuator control command 510 depending on the classification obtained by the trained neural network. For example, the classifier 514 may interpret a region of the sensed image as potentially anomalous. In this case, the actuator control command 510 may be determined or selected to cause the display 1102 to display the image and highlight the potentially anomalous region.
[0084] While exemplary embodiments have been described above, these embodiments are not intended to describe every possible embodiment encompassed by the scope of the claims. The terms used herein are terms of description rather than limitation, and it is understood that various modifications are possible without departing from the spirit and scope of the present disclosure. As noted above, features of various embodiments can be combined to form additional embodiments of the present invention not explicitly described or shown. While various embodiments have been described as offering advantages or being more preferred over other embodiments or prior art implementations with respect to one or more desired characteristics, those skilled in the art will recognize that, depending on the particular application and implementation, one or more characteristics or characteristics may be compromised to achieve overall desirable system attributes. These attributes may include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, maintainability, weight, manufacturability, ease of assembly, etc. Thus, to the extent that any embodiment is described as less desirable with respect to one or more characteristics than other embodiments or prior art implementations, these embodiments are not outside the scope of the present disclosure and may be desirable for particular applications.
Claims
1. 1. A method for performing pose estimation for an image, comprising: The method includes, in one or more processing devices: receiving an input image; generating a pose code based on the input image, the pose code corresponding to an estimated pose of an object in the input image; generating a box code corresponding to a bounding box of the object in the input image; performing pose estimation for the input image by generating a refined pose of the object using the pose code and the box code; generating a predicted output of the object in the input image based on the input image and the refined pose; controlling one or more functions of a device based on the predicted output; A method comprising:
2. generating the predicted output comprises: (i) generating shape and texture codes using an image encoder; (ii) generating the predicted output using the shape code, the texture code and the refined pose; Including, The method of claim 1.
3. generating the predicted output includes generating the predicted output using a neural radiance field (NeRF) decoder. The method of claim 2.
4. generating the refined pose includes iteratively calculating the refined pose using the pose code and the box code. The method of claim 1.
5. The method further comprises: iteratively updating the box code using the refined pose. The method of claim 4.
6. The method further comprises: obtaining a posture loss based on the posture code using a multi-layer perceptron. The method of claim 1.
7. generating the predicted output comprises: converting the refined pose to a camera pose; generating the predicted output based on the camera pose; Including, The method of claim 1.
8. 1. A computing device configured to perform pose estimation for an image, the computing device comprising: receiving an input image; generating a pose code based on the input image, the pose code corresponding to an estimated pose of an object in the input image; generating a box code corresponding to a bounding box of the object in the input image; performing pose estimation for the input image by generating a refined pose of the object using the pose code and box code; generating a predicted output of the object in the input image based on the input image and the refined pose; controlling one or more functions of a device based on the predicted output; 1. A computing device comprising: a processing device configured to execute instructions stored in a memory to perform the steps of:
9. generating the predicted output comprises: (i) generating shape and texture codes using an image encoder; (ii) generating the predicted output using the shape code, the texture code and the refined pose; Including, The computing device of claim 8 .
10. generating the predicted output includes generating the predicted output using a neural radiance field (NeRF) decoder. The computing device of claim 9.
11. generating the refined pose includes iteratively calculating the refined pose using the pose code and the box code. The computing device of claim 8 .
12. the processing device is configured to iteratively update the box code using the refined pose. The computing device of claim 11.
13. the processing device is configured to use a multi-layer perceptron to obtain a posture loss based on the posture code. The computing device of claim 8 .
14. generating the predicted output comprises: converting the refined pose to a camera pose; generating the predicted output based on the camera pose; Including, The computing device of claim 8 .
15. 1. A computer-controlled device configured to operate according to a pose estimate that generated a visual model, comprising: The computer-controlled device comprises:
1. A control system comprising: receiving an input image captured by a camera; generating a pose code based on the input image, the pose code corresponding to an estimated pose of an object in the input image; generating a box code corresponding to a bounding box of the object in the input image; performing pose estimation for the input image by generating a refined pose of the object using the pose code and box code; generating a predicted output of the object in the input image based on the input image and the refined pose; outputting a control signal based on the predicted output; a control system configured to: an actuator configured to control operation of the computer-controlled device based on the control signal; A computer-controlled device comprising:
16. generating the predicted output comprises: (i) generating shape and texture codes using an image encoder; (ii) generating the predicted output using the shape code, the texture code and the refined pose; Including, 16. A computer-controlled apparatus according to claim 15.
17. generating the predicted output includes generating the predicted output using a neural radiance field (NeRF) decoder.
17. A computer-controlled apparatus according to claim 16.
18. generating the refined pose includes iteratively calculating the refined pose using the pose code and the box code.
16. A computer-controlled apparatus according to claim 15.
19. The control system further comprises: configured to iteratively update the box code using the refined pose.
19. The computer-controlled apparatus of claim 18.
20. The control system further comprises: and configured to use a multi-layer perceptron to obtain a posture loss based on the posture code.
16. A computer-controlled apparatus according to claim 15.