Unified model for 3D object detection and object-centric neural reconstruction
By using a unified object-centered neural reconstruction (UPNeRF) technique combined with a pose estimation module, the problem of high computational overhead in existing technologies is solved, achieving efficient and generalized 3D object reconstruction and pose estimation.
Patent Information
- Application Number
- CN202510992905.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-20
AI Technical Summary
Existing single-view 3D object reconstruction techniques rely on external 3D object detection during training and deployment, resulting in high computational overhead. Furthermore, the requirements for multi-view observation and accurate object pose estimation are stringent, making it difficult to achieve efficient and generalized reconstruction results.
We employ a unified object-centered neural reconstruction (UPNeRF) technique, combined with a pose estimation module, to improve reconstruction efficiency and generalization ability by iteratively updating the object pose. We also utilize a pose improver and NeRF decoder to generate more accurate 3D reconstruction results.
It enables more efficient and generalizable 3D object reconstruction, reduces computational overhead, and improves the accuracy and flexibility of object pose estimation.
Smart Images

Figure CN121366410A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to artificial intelligence (AI) techniques for image recognition and processing. BACKGROUND
[0002] Various systems are configured to perform tasks using machine learning (ML) or other artificial intelligence (AI) techniques. For example, systems configured to perform image recognition, object detection, and / or other automated tasks can implement AI techniques. As one example, image detection systems and methods use various detection models trained for object and feature detection. SUMMARY
[0003] A method of performing pose estimation for an image includes receiving, at one or more processing devices, an input image, generating, based on the input image, a pose code corresponding to an estimated pose of an object in the input image, generating a box code corresponding to a bounding box of the object in the input image, performing pose estimation for the input image by generating an improved pose of the object using the pose code and the box code, generating a predicted output for the object in the input image based on the input image and the improved pose, and controlling one or more functions of a device based on the predicted output.
[0004] Other embodiments include a non-transitory computer readable storage medium configured to store instructions that, when executed by a processor included in a computing device, cause the computing device to perform various steps of any of the preceding methods. Further embodiments include a computing device configured to perform various steps of any of the preceding methods. Further embodiments include a machine configured to perform various steps of any of the preceding methods.
[0005] Other aspects and advantages of the present application will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, illustrating by way of example the principles of the described embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0006] Figure 1 A system for training a machine learning model according to the principles of the present disclosure is generally shown.
[0007] Figure 2 A computer-implemented method for training and implementing a machine learning model according to the principles of the present disclosure is generally shown.
[0008] Figure 3A An audio data labeling system according to the principles of the present disclosure is generally shown.
[0009] Figure 3B A portion of a data capture system according to the principles of the present disclosure is generally shown.
[0010] Figure 3C An alternative audio data labeling system according to the principles of the present disclosure is generally shown.
[0011] Figure 4A An example overall processing pipeline for a visual model according to the principles of the present disclosure is shown.
[0012] Figure 4B An example pose estimation module according to the principles of the present disclosure is shown.
[0013] Figure 4C An example unified model or pipeline for both training and inference according to the principles of the present disclosure is shown.
[0014] Figure 4D Steps of an example method for implementing a visual model (e.g., with which to train and subsequently perform pose estimation) according to the principles of the present disclosure are shown.
[0015] Figure 5 A schematic diagram of an interaction between a computer-controlled machine and a control system according to the principles of the present disclosure is shown.
[0016] Figure 6 A schematic diagram of a control system of Figure 5 configured to control a vehicle, which can be a partially autonomous vehicle, a fully autonomous vehicle, a partially autonomous robot, or a fully autonomous robot, according to the principles of the present disclosure is shown.
[0017] Figure 7 A schematic diagram of a control system of Figure 5 configured to control a manufacturing machine (e.g., a punch cutter, a cutter, or a gun drill) of a manufacturing system (e.g., a portion of a production line) is shown.
[0018] Figure 8 A schematic diagram of a control system of Figure 5 configured to control a power tool, such as a power drill or driver having at least a partially autonomous mode is shown.
[0019] Figure 9 A schematic diagram of a control system of Figure 5 configured to control an automated personal assistant is shown.
[0020] Figure 10 A schematic diagram of a control system of Figure 5 configured to control a surveillance system, such as a control access system or a monitoring system is shown.
[0021] Figure 11 A schematic diagram of a control system of Figure 5Fig. 1 shows a schematic diagram of a control system configured to control an imaging system, such as an MRI device, an x-ray imaging device or an ultrasound device. DETAILED DESCRIPTION
[0022] Embodiments of the present disclosure are described herein. It is to be understood, however, that the disclosed embodiments are merely examples and other embodiments can take various and alternative forms. The figures are not necessarily to scale; some features can be exaggerated or minimized for the purpose of clarity and illustration. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the embodiments. As those skilled in the art will appreciate, the various features shown and described either with reference to any of the figures can be combined or otherwise rearranged within additional structures to produce yet further embodiments, which are also within the scope of the present disclosure. Combinations of the features shown and described with reference to any of the figures are contemplated and within the scope of the disclosure.
[0023] As used herein, the singular forms "a", "an" and "the" refer to both the singular as well as plural, unless the context clearly indicates otherwise. For example, a "processor" programmed to perform various functions refers to one processor programmed to perform each function, or more than one processor collectively programmed to perform each function.
[0024] As used herein, "content" can refer to raw content corresponding to input data (e.g., data representing captured images, video, sound, text, etc.) or synthetic content (e.g., synthetic images, video, sound, text, etc.). In some examples, "content" can include images, which can correspond to captured images, synthetic images, or a combination thereof. Images can be represented by image data. In some contexts herein, the terms "image" and "image data" can be used interchangeably and can refer to actual pixel values, color channels, vectors, and / or binary data corresponding to the visual content of an image. In examples, "image" and / or "image data" refer to a raw representation of an image, such as an array of numerical values representing pixel intensities, which in some examples can include pre-processed data originating from an image sensor. In contrast, "metadata" or "image metadata" can refer to contextual or supplemental details about an image, such as image size, format, date of creation, geolocation data, etc. In various examples, "image" and "image data" can but need not also include metadata.
[0025] Various systems are configured to perform tasks using machine learning (ML) or other artificial intelligence (AI) techniques (e.g., ML or other AI models). For example, systems configured to perform image recognition, object detection, and / or other automated tasks can implement AI techniques. As one example, image detection systems and methods use various detection (e.g., vision model) models trained for object and feature detection.
[0026] Some vision models are configured to generate images of 3-dimensional (3D) objects from 2-dimensional (2D) images (e.g., reconstruct 3D objects from a single 2D image), which can be referred to as single-view 3D object reconstruction. Single-view 3D object reconstruction is a key technology with a wide range of applications, including but not limited to autonomous driving, augmented / virtual reality (AV / VR) systems, robotics, and embodied AI. Single-view 3D object reconstruction techniques are limited by constraints of primary data sources, which can include sparse views and dynamic objects.
[0027] Vision models can implement neural radiance field (“NeRF”) techniques to perform 3D reconstruction, which offers particular advantages in terms of rendering scenes at fine resolutions and generating novel view images from reconstructed scenes. In some examples, object-centric NeRF techniques further enhance the flexibility of novel data synthesis. However, object-centric NeRF approaches place strict requirements on multi-view observations and accurate object poses, and / or rely heavily on third-party object detection to provide initial object poses. The reliance on external 3D object detection introduces computational overhead in both training and deployment.
[0028] Systems and methods according to the present disclosure implement visual models configured to perform unified object-centric reconstruction (e.g., object-centric NeRF 3D reconstruction) techniques. In particular, the systems and methods described herein combine object-centric neural reconstruction and pose estimation to obtain more efficient and generalizable reconstruction results. In one example, a visual module includes, implements, contains, and / or is in communication with a pose estimation module configured to generate updated pose data (“improved pose” or “improved pose data”) for an object in an input image from the input image and an input / current pose of the object (i.e., the pose of the object as shown in the input image). As used herein, “pose” refers to the position and orientation of an object in 3D space relative to a camera. As used herein, “camera” can refer to a camera that captures an image. Thus, an estimated or computed pose or pose data can include coordinates defining the pose of the object (e.g., X, Y, and Z coordinates in a 3D coordinate space), an angle heat map, a bounding box (with rotation / rotation angle), etc. In examples described herein, the pose or pose data includes a bounding box. The “improved” pose, which is different from the input or current pose, is a predicted or computed pose (and / or corresponding pose data) for a different range (i.e., distance from the object), angle, orientation, etc.
[0029] Figure 1 One example system 100 for training an ML or other AI model (e.g., a visual model according to the present disclosure) is shown. As used herein, for simplicity, a “visual” model can refer to a pose estimation model or module, a visual model configured to perform pose estimation according to the techniques of the present disclosure, etc. The system 100 can be configured to (and / or include circuitry configured to) implement the systems and methods of the present disclosure described in more detail below. The system 100 can include an input interface for accessing training data 102 for a visual model. For example, as shown, the input interface can be made up of a data storage interface 104 that can access the training data 102 from a data storage 106. For example, the data storage interface 104 can be a memory interface or a persistent storage interface, such as a hard disk or SSD interface, but can also be a personal area network, local area network, or wide area network interface, such as a Bluetooth, Zigbee, or Wi-Fi interface or an Ethernet or fiber optic interface. The data storage 106 can be an internal data storage of the system 100, such as a hard drive or SSD, but can also be an external data storage, such as a network-accessible data storage. Figure 1
[0030] In some embodiments, the data storage 106 can also include a data representation 108 of an untrained version of a visual model that can be accessed by the system 100 from the data storage 106. However, it should be appreciated that the data representation 108 of the untrained visual model and the training data 102 can each be accessed from different data storages, e.g., via different subsystems of the data storage interface 104. Each subsystem can be of the type as described above for the data storage interface 104.
[0031] In some embodiments, the data representation 108 of the untrained visual model can be generated internally by the system 100 based on design parameters of the visual model and, thus, can not be explicitly stored on the data storage 106. The system 100 can also include a processor subsystem 110 that can be configured to provide an iterative function as a substitute for a layer stack of the visual model to be trained during operation of the system 100. Herein, respective layers of the substituted layer stack can have weights shared with each other and can receive an output of a previous layer as an input or, for a first layer of the layer stack, an initial activation and a portion of an input of the layer stack.
[0032] The processor subsystem 110 can also be configured to iteratively train the visual model using the training data 102. Herein, an iteration of the training by the processor subsystem 110 can include a forward propagation portion and a backward propagation portion. In addition to defining other operations that can be performed for the forward propagation portion, the processor subsystem 110 can be configured to, among other things, perform the forward propagation portion by determining a fixed point at which an iterative function converges, and by providing the fixed point as a substitute for an output of the layer stack in the visual model, where determining the fixed point includes finding a root solution of the iterative function minus its input using a numerical root-finding algorithm. The processor subsystem 110 is configured to train the visual model in accordance with the systems and methods of the present disclosure as described in more detail below.
[0033] The system 100 can also include an output interface for outputting a data representation 112 of the trained visual model. This data can also be referred to as trained model data 112. For example, as also shown, the output interface can be made up of the data storage interface 104, where in these embodiments the interface is an input / output (“IO”) interface via which the trained model data 112 can be stored in the data storage 106. For example, the data representation 108 defining the “untrained” visual model can be at least partially replaced by the data representation 112 of the trained visual model during or after training, as parameters of the visual model (e.g., weights, hyperparameters, and other types of parameters of the visual model) can be adapted to reflect the training on the training data 102. This is also shown in Figure 1 Figure 1 The data representations 112 are shown in FIG. 1 by reference numerals 108, 112 that refer to the same data records on the data storage device 106. In some embodiments, the data representations 112 can be stored separately from the data representations 108 that define the“untrained” visual model. In some embodiments, the output interface can be separate from the data storage interface 104, but can generally be of the type described above for the data storage interface 104.
[0034] Figure 2 An example content generation system 200 configured to implement (and / or include circuitry configured to implement) a system for annotating, augmenting, and / or generating data is described. The content generation system 200 can include at least one computing system 202 configured to implement all or part of the systems and methods of the present disclosure explained in greater detail below. The computing system 202 can include at least one processor 204 operably connected to a memory unit 208. The processor 204 can include one or more integrated circuits that implement the functionality of a central processing unit (CPU) 206. The CPU 206 can be a commercially available processing unit that implements an instruction set (e.g., one of the x86, ARM, Power, or MIPS instruction set families). The various components of the system 200 can be implemented with the same or different circuitry.
[0035] During operation, the CPU 206 can execute stored program instructions retrieved from the memory unit 208. The stored program instructions can include software that controls the operation of the CPU 206 to implement operations described herein. In some embodiments, the processor 204 can be a system on a chip (SoC) that integrates the functionality of the CPU 206, the memory unit 208, network interfaces, and input / output interfaces into a single integrated device. The computing system 202 can implement an operating system for managing aspects of the operation.
[0036] The memory unit 208 can include volatile memory and non-volatile memory for storing instructions and data. The non-volatile memory can include solid state memory, such as NAND flash memory, magnetic and optical storage media, or any other suitable data storage device that retains data when the computing system 202 is deactivated or loses power. The volatile memory can include static and dynamic random access memory (RAM) that stores program instructions and data. For example, the memory unit 208 can store one or more machine learning models (e.g., as machine learning models 210) or algorithms, training data sets 212 for the machine learning models 210, raw source data sets 216, etc. Figure 2
[0037] The computing system 202 can include a network interface device 222 configured to provide communication with external systems and devices. For example, the network interface device 222 can include a wired and / or wireless Ethernet interface defined by the Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards. The network interface device 222 can include a cellular communication interface for communicating with a cellular network (e.g., 3G, 4G, 5G). The network interface device 222 can also be configured to provide a communication interface to an external network 224 or cloud.
[0038] The external network 224 can be referred to as the World Wide Web or the Internet. The external network 224 can establish standard communication protocols between computing devices. The external network 224 can allow for easy exchange of information and data between computing devices and networks. One or more servers 230 can be in communication with the external network 224.
[0039] The computing system 202 can include an input / output (I / O) interface 220 that can be configured to provide digital and / or analog input and output. The I / O interface 220 can include additional serial interfaces (e.g., Universal Serial Bus (USB) interfaces) for communicating with external devices.
[0040] The computing system 202 can include a human machine interface (HMI) device 218 that can include any device that enables the system 200 to receive control inputs. Examples of input devices can include human machine interface inputs such as a keyboard, mouse, touch screen, voice input device, and other similar devices. The computing system 202 can include a display device 232. The computing system 202 can include hardware and software for outputting graphical and textual information to the display device 232. The display device 232 can include an electronic display screen, a projector, a printer, or other suitable device for displaying information to a user or operator. The computing system 202 can also be configured to allow interaction with remote HMI and remote display devices via the network interface device 222.
[0041] The system 200 can be implemented using one or more computing systems. While the examples describe a single computing system 202 that implements all of the described features, it is intended that various features and functionality can be separated and implemented by multiple computing units in communication with one another. The particular system architecture chosen can depend on various factors.
[0042] The system 200 can implement a machine learning model 210 to analyze a raw source dataset 216. For example, the CPU 206 and / or other circuitry can implement the machine learning model 210. The raw source dataset 216 can include raw or unprocessed sensor data, which can represent an input dataset for a machine learning system. The raw source dataset 216 can include images, videos, video clips, audio, text-based information, and raw or partially processed sensor data (e.g., a radar map of an object). In some embodiments, the machine learning model 210 can include a deep learning or neural network algorithm designed to perform a predetermined function. For example, the neural network algorithm can be configured to identify events or objects in an image or video clip based on audio data.
[0043] The computer system 202 can store a training dataset 212 for the machine learning model 210. The training dataset 212 can represent a previously constructed dataset used to train the machine learning model 210. The machine learning model 210 can use the training dataset 212 to learn various conditions and other factors (e.g., weighting factors) associated with the ML algorithm. The training dataset 212 can include source datasets with corresponding outcomes or results that the machine learning model 210 attempts to replicate via a learning process.
[0044] The machine learning model 210 can operate in a learning mode using the training dataset 212 as input. The machine learning model 210 can be executed over multiple iterations using data from the training dataset 212. For each iteration, the machine learning model 210 can update internal weighting factors based on the results achieved. For example, the machine learning model 210 can compare the output results (e.g., generated content) with those included in the training dataset 212. Since the training dataset 212 includes expected results, the machine learning model 210 can determine when performance is acceptable. After the machine learning model 210 reaches a predetermined level of performance (e.g., 100% compliance with the results associated with the training dataset 212), the machine learning model 210 can be executed using data that is not in the training dataset 212. The trained machine learning model 210 can be applied to new datasets to generate content. The machine learning model 210 can include a visual model trained according to the systems and methods of the present disclosure.
[0045] Machine learning model 210 can be configured to identify specific features in raw source data 216. Raw source data 216 may include multiple instances or input datasets (e.g., images, video streams or clips including audio data, etc.) for its desired output. By way of example only, machine learning model 210 can be configured to identify objects or features in images, or objects or events in video clips, based on audio data, etc. In some examples, machine learning model 210 can be configured to annotate the identified objects, features, or events. Machine learning model 210 can be configured to perform pose estimation according to the principles of this disclosure. Machine learning model 210 can be programmed to process raw source data 216 to identify the presence of specific features. Machine learning model 210 can be configured to identify features in raw source data 216 as predetermined features. Raw source data 216 can be derived from various sources. For example, raw source data 216 can be actual input data collected by a machine learning system. Raw source data 216 can be machine-generated for use in testing the system. As an example, the raw source data 216 may include raw image data, raw video and / or audio data from a camera, audio data from a microphone, etc.
[0046] In the example, machine learning model 210 can process raw source data 216 and output video and / or audio data including one or more indications of the identified events. Machine learning model 210 can generate confidence levels or factors for each output. For example, a confidence value exceeding a predetermined high confidence threshold can indicate that machine learning model 210 is confident that the identified event (or feature) corresponds to a specific event. A confidence value below a low confidence threshold can indicate that machine learning model 210 has some uncertainty regarding the existence of a specific feature.
[0047] like Figure 3A and Figure 3B As generally illustrated, example system 300 may include an image (e.g., image and / or video) capture device 302, an audio capture array 304, and a computing system 202. The system may receive video stream data associated with a data capture environment from image capture device 302. System 202 may be configured to perform video object detection to identify one or more objects in the corresponding image of the video stream data. System 202 may receive audio stream data corresponding to at least a portion of the video stream data from audio capture array 304. Audio capture array 304 may include one or more microphones 306 or other suitable audio capture devices. The systems and methods described herein may be configured to label at least some objects in the video stream data and / or audio stream data using output from at least a first machine learning model (e.g., machine learning model 210 or other suitable machine learning models configured to provide outputs including predictions of one or more object or event detections).
[0048] The system 202 can compute (e.g., using at least one probability-based function or other suitable technique or function) at least one offset value for at least a portion of the audio stream data corresponding to the at least one labeled object of the video stream data based on at least one data capture characteristic. The system 202 can synchronize at least a portion of the video stream data with at least a portion of the audio stream data corresponding to the at least one labeled object of the video stream data using at least the at least one offset value. The at least one data capture characteristic can include one or more characteristics of the at least one image capture device, one or more characteristics of the at least one audio capture array, one or more characteristics corresponding to a position of the at least one image capture device relative to the at least one audio capture array, one or more characteristics corresponding to movement of an object in the video stream data, one or more other suitable data capture characteristics, or a combination thereof.
[0049] The system 202 can label at least a portion of the audio stream data corresponding to the at least one labeled object of the video stream data using one or more labels of the labeled object of the video stream data and the at least one offset value. Each respective label can include an event type, an event start indicator, and an event end indicator. The system 202 can generate training data using at least some of the labeled portion of the audio stream data. The system 202 can train a second machine learning model using the training data. The system 202 can use the second machine learning model to detect one or more sounds associated with audio data provided as input to the second machine learning model. The second machine learning model can include any suitable machine learning model and can be configured to perform any suitable function, such as those described herein with respect to FIGS. 4-11.
[0050] In some embodiments, as Figure 3CAs generally illustrated, computing system 202 can be configured to label audio data based on sensor data received from one or more sensors, such as those described herein or any other suitable sensor or combination of sensors. System 202 can receive audio stream data associated with a data capture environment from audio capture array 354 or any suitable audio capture device, such as one or more of microphones 306 or other suitable audio capture devices. It should be understood that audio capture array 354 can include features similar to those of audio capture array 304 and can include any suitable number of audio capture devices. System 202 can receive sensor data associated with the data capture environment from at least one sensor (e.g., sensor 352) that is asynchronous with respect to audio capture array 354. Sensor 354 can include at least one of an inductive coil, a radar sensor, a LiDAR sensor, a sonar sensor, an image capture device, any other suitable sensor, or a combination thereof. Audio capture array 354 can be positioned remotely from sensor 354, proximate to sensor 354, or in any suitable relationship to sensor 354.
[0051] System 202 can use output from at least a first machine learning model, such as machine learning model 210 or other suitable machine learning model, to identify at least some events in the sensor data. Machine learning model 210 can be configured to provide output including one or more event detection predictions based on the sensor data. System 202 can synchronize at least portions of the sensor data associated with portions of the audio stream data corresponding to at least one event of the sensor data. System 202 can label at least portions of the audio stream data corresponding to at least one event of the sensor data using one or more labels extracted for respective events of the sensor data values. Each respective label can include an event type, an event start indicator, and an event end indicator. System 202 can use at least some of the labeled portions of the audio stream data to generate training data. System 202 can use the training data to train a second machine learning model. System 202 can use the second machine learning model to detect one or more sounds associated with audio data provided as input to the second machine learning model. The second machine learning model can include any suitable machine learning model and can be configured to perform any suitable function, such as those described herein with respect to FIGS. 4-11.
[0052] The systems and methods of the present disclosure (e.g., any of systems 100, 200, etc.) are configured to train a vision model (e.g., model 210) to perform pose estimation and generate improved poses using the vision model, as described in greater detail below. The techniques of the present disclosure can be referred to as unified NeRF (“UPNeRF”) techniques, which provide a unified solution to jointly predict the pose, shape, and texture of observed objects from a single network. The vision models of the present disclosure can be trained using real-world scenes (e.g., real-world driving scenes) with inaccurate predicted labels.
[0053] Figure 4A An example overall processing pipeline 400 of a vision model 402 (e.g., a UPNeRF vision model) configured to perform pose estimation is shown in accordance with the present disclosure. For example, one or more computing devices, processors, or processing devices are configured to execute instructions to implement the functionality of pipeline 400, such as one or more processors of a system (e.g., 100, 200, etc.) described herein.
[0054] Vision model 402 is trained with a training dataset 404. Training dataset 404 includes a plurality of input images 406 (e.g., 2D images of objects, such as vehicles) and corresponding poses of the objects, represented in this example as bounding boxes 408 with rotations. As used herein, “with rotations” refers to data / values that indicate one or more rotation angles of bounding boxes 408. For example, the rotation angles indicate the orientation of bounding boxes 408 (and the objects) relative to a camera, ground, etc. Bounding boxes 408 can be defined by data / values that identify one or more corner (e.g., using X, Y, and Z) coordinates, width, and / or length, etc. of bounding boxes 408. In some examples, training dataset 404 also includes shapes 410 or shape data. Shapes 410 provided with images 406 indicate the overall shape, contour, form, etc. of the objects in images 406.
[0055] During and / or after training, vision model 402 is provided with test images 412 (e.g., a test set of images of objects extracted from scenes 416). As shown, test images 412 can be provided to vision model 402 as additional inputs along with occlusion masks, random poses (e.g., bounding boxes representing random poses), etc. Vision model 402 is configured to generate and output features based on test images 412, such as the texture, shape, and improved poses of the objects in the test images, as shown at 418.
[0056] Figure 4BAn example pose estimation module 424 according to this disclosure is shown. As used herein, pose estimation module 424 may correspond to a model executed by visual model 402, a model separate from visual model 402, circuitry configured to perform pose estimation functions or techniques, etc. For example, one or more computing devices, processors, or processing devices are configured to execute instructions to implement the functions of pose estimation module 424, such as one or more processors of the systems described herein (e.g., 100, 200, etc.). Pose estimation module 424 is configured to provide a reliable pose for a target object in multiple ranges and orientations, and to perform robustly under various conditions (e.g., for an occluded image where at least part of the target object is occluded / masked).
[0057] like Figure 4B As shown, the pose estimation module 424 iteratively updates the input pose 426 based on the visual difference between the input pose 426 and the observed object in the input image 428. Given the size of the object [H] B W B and L B Given the camera's intrinsic K and current pose (t) and (t) (corresponding to rotation and translation, respectively), the pose estimation module 424 obtains image projections of the 3D box corners (t) (e.g., the coordinates of the eight corners of the bounding box 430). The box corners (t) correspond to a visual representation of the current (i.e., input) pose 426. In this example, (t) is a 16-bit vector. The box encoder 432 encodes (t) to generate a box code 436. For example, the box code 436 corresponds to a higher-dimensional code or vector based on (t). In this example, the box code 436 is a 255-bit vector or another representation of (t).
[0058] Input image 428 is provided to image encoder 438. Image encoder 438 is configured to generate and output an estimated pose, such as pose code 440, based on input image 428. For example, pose code 440 is a code value or vector corresponding to the estimated pose. In the example, pose code 400 is a lower-dimensional code or vector (e.g., compressed) representation of the estimated pose obtained using principal component analysis (PCA) or other techniques.
[0059] Box code 436 and attitude code 440 are provided as input to attitude improver 444. Attitude improver 440 is configured to predict attitude update 446 or attitude change Δ(t), ΔT(t), which represent the corresponding changes in R(t) and T(t) for the input attitude 426. Attitude update 446 is combined with input attitude 426 to obtain the next (improved or updated) attitude or attitude state 448(R). (t+1) T (t+1)). The generation of pose updates 446 and updated poses 448 repeats over multiple iterations (e.g., by providing the pose updates 446 to the box encoder 432, which updates the box codes 436 based on the pose updates 446. The pose estimation module 424 continues to generate pose updates 446 and updated poses 448 until a final pose state is obtained.
[0060] Figure 4C An example unified model or pipeline 450 (e.g., pipeline of a vision model) is shown in accordance with the present disclosure. For example, one or more computing devices, processors, or processing devices are configured to execute instructions to implement the functionality of the unified pipeline 450, such as one or more processors of a system (e.g., 100, 200, etc.) described herein. The unified pipeline 450 shows both training of a vision model and inference functions performed by the vision model. For example, as shown, the training and inference process flows are indicated by respective dashed lines, and the flows common to both the training and inference processes are indicated by solid lines. Figure 4C
[0061] The pipeline 450 includes an image encoder 438 (e.g., a residual network (ResNet)-based image encoder), the pose estimation module 424, and a NeRF decoder 452. The image encoder 438 receives the input image 428 (and an associated occlusion mask, together referred to as a masked input image). The image encoder 438 converts the masked input image 428 into shape codes 454 and texture codes 456 (e.g., respective code values or vectors corresponding to estimated shapes and textures) and pose codes 440. The pose codes 440 are provided to the pose improver 444, as described above, along with the box codes 436 to iteratively improve the object poses R o2c |T o2c After multiple iterations, the estimated poses can be converted to camera poses R c2o |T c2o and input to the NeRF decoder 452 for inference tasks, or for computing pose loss (L) during training, as shown at 458.
[0062] During inference tasks, the NeRF decoder 452 generates a predicted output 460 based on the shape codes 454, the texture codes 456, and the updated poses 448, which identifies detected objects within the input image, corresponding bounding boxes, etc. In examples, the NeRF decoder 452 performs volume rendering to generate an RGB image (e.g., rendered RGB values) and an occupancy image (e.g., aggregated occupancy values). The rendered RGB values are compared to the input image 428 to compute a photometric loss L rgb and the aggregated occupancy values are compared to an occupancy mask received with the input image 428 to obtain an occupancy loss L occ total loss L infer may be obtained according to L infer = L rgb + w occ L occ where w occ is a weight coefficient configured to balance the two loss terms L rgb and L occ . The loss L infer is used to update the optimizable variables of the NeRF decoder 542, which are defined differently for inference and training.
[0063] In an example, the pipeline 450 can include one or more multi-layer perceptrons (MLPs) 462. For example, the MLPs 462 are configured to convert the pose codes 440 into higher dimensional codes or vectors during training, which can be used to obtain the direct pose loss In contrast, the output of the pose improver 444 can be used to obtain the pose loss
[0064] By unifying object detectors and object-centric neural reconstruction in the above-described manner, the visual model according to the present disclosure significantly improves computational efficiency and generalization capability.
[0065] Figure 4D Steps of an example method 470 for implementing a visual model (e.g., with which pose estimation is trained and subsequently performed) according to the principles of the present disclosure are shown. For example, one or more processors or processing devices are configured to execute instructions to implement the method 450, such as one or more processors of the systems described herein.
[0066] At 472, the method 470 includes training the visual model to perform pose estimation using a training set of images, masks, and poses (e.g., pose information or data, such as bounding boxes). Training the visual model includes training the pose estimation module 424, as described above in Figures 4A-4C .
[0067] At 474, the method 470 includes receiving an input image at the visual model. At 476, the method 470 includes generating a pose code based on the input image. At 478, the method 470 includes iteratively generating an improved pose based on the pose code and the box code. For example, the method 470 can repeat step 476 to improve the pose (e.g., until the pose loss is below a threshold). At 480, the method 470 includes generating (e.g., using a NeRF decoder as described herein) a predicted output based on the input image (e.g., based on the shape code and the texture code) and the improved pose.
[0068] At 482, method 470 includes one or more functions for controlling a system, device, machine, etc., based on predicted output and improved posture. For example, predicted output and improved posture can be used for various downstream object detection and image recognition tasks, such as the control of autonomous vehicles, robots, AR / VR systems, etc. In some examples, method 470 includes controlling the following... Figures 5-11 The functionality of any system described in the document.
[0069] Figures 5-11 Example systems and devices that can implement visual models (e.g., pose estimation models, visual models configured to perform pose estimation, etc.) according to this disclosure are described. Figure 5 A schematic diagram illustrating the interaction between a computer-controlled machine 500 and a control system 502 is provided. In one example, the control system 502 is configured to control the computer-controlled machine 500 by executing a visual model according to the principles of this disclosure. The computer-controlled machine 500 includes actuators 504 and sensors 506. Actuators 504 may include one or more actuators, and sensors 506 may include one or more sensors. Sensors 506 are configured to sense the condition of the computer-controlled machine 500. Sensors 506 may be configured to encode the sensed condition into a sensor signal 508 and transmit the sensor signal 508 to the control system 502. Non-limiting examples of sensors 506 include video, radar, LiDAR, ultrasonic, and motion sensors. In some embodiments, sensor 506 is an optical sensor configured to sense an optical image of the environment near the computer-controlled machine 500. A pose estimation for the optical image can be performed according to the visual model of this disclosure, as described herein.
[0070] The control system 502 is configured to receive sensor signals 508 from the computer-controlled machine 500. As described below, the control system 502 can also be configured to calculate actuator control commands 510 based on the sensor signals and transmit the actuator control commands 510 to the actuator 504 of the computer-controlled machine 500.
[0071] like Figure 5 As shown, the control system 502 includes a receiving unit 512. The receiving unit 512 can be configured to receive sensor signals 508 from sensor 506 and transform the sensor signals 508 into input signals x. In an alternative embodiment, sensor signals 508 are received directly as input signals x without the receiving unit 512. Each input signal x may be a portion of each sensor signal 508. The receiving unit 512 can be configured to process each sensor signal 508 to generate each input signal x. The input signal x may include data corresponding to an image recorded by sensor 506.
[0072] The control system 502 includes a classifier 514. The classifier 514 can be configured to classify an input signal x into one or more labels using a machine learning (ML) algorithm (e.g., a neural network). For example, classifier 514 corresponds to classifier 408 described above. The classifier 514 is configured to be parameterized by parameters (e.g., the parameters described above, such as parameter θ). Parameter θ can be stored in and provided by a non-volatile storage device 516. The classifier 514 is configured to determine an output signal y from the input signal x. Each output signal y includes information assigning one or more labels to each input signal x. The classifier 514 can transmit the output signal y to a conversion unit 518. The conversion unit 518 is configured to convert the output signal y into an actuator control command 510. The control system 502 is configured to transmit the actuator control command 510 to an actuator 504, which is configured to actuate a computer-controlled machine 500 in response to the actuator control command 510. In some embodiments, actuator 504 is configured to actuate computer-controlled machine 500 directly based on output signal y.
[0073] When actuator 504 receives actuator control command 510, actuator 504 is configured to perform an action corresponding to the relevant actuator control command 510. Actuator 504 may include control logic configured to transform actuator control command 510 into a second actuator control command for controlling actuator 504. In one or more embodiments, actuator control command 510 may be used to control a display instead of an actuator or in addition to an actuator.
[0074] In some embodiments, the control system 502 includes a sensor 506, and the computer-controlled machine 500 includes a sensor 506 as a replacement or addition. The control system 502 may also include an actuator 504, and the computer-controlled machine 500 includes an actuator 504 as a replacement or addition.
[0075] like Figure 5 As shown, the control system 502 also includes a processor 520 and a memory 522. The processor 520 may include one or more processors. The memory 522 may include one or more memory devices. A classifier 514 (e.g., an ML algorithm) of one or more embodiments may be implemented by the control system 502, which includes a non-volatile storage device 516, a processor 520, and a memory 522.
[0076] The non-volatile storage 516 can include one or more persistent data storage devices such as a hard disk drive, optical drive, tape drive, non-volatile solid-state device, cloud storage, or any other device capable of persistently storing information. The processor 520 can include one or more devices selected from a high-performance computing (HPC) system, including high-performance cores, microprocessors, microcontrollers, digital signal processors, microcomputers, central processing units, field-programmable gate arrays, programmable logic devices, state machines, logic circuits, analog circuits, digital circuits, or any other devices that manipulate signals (analog or digital) based on instructions resident in the memory 522. The memory 522 can include a single memory device or multiple memory devices, including but not limited to random-access memory (RAM), volatile memory, non-volatile memory, static random-access memory (SRAM), dynamic random-access memory (DRAM), flash memory, cache, or any other device capable of storing information.
[0077] The processor 520 can be configured to read into the memory 522 and execute computer-executable instructions that reside in the non-volatile storage 516 and embody one or more anomaly detection methods of one or more embodiments. The non-volatile storage 516 can include one or more operating systems and applications. The non-volatile storage 516 can store computer programs created using various programming languages and / or technologies, including but not limited to Java, C, C++, C#, Objective C, Fortran, Pascal, JavaScript, Python, Perl, and PL / SQL, individually or in combination.
[0078] The computer-executable instructions of the non-volatile storage 516, when executed by the processor 520, can cause the control system 502 to implement one or more of the anomaly detection methods as disclosed herein. The non-volatile storage 516 can also include data that supports the functions, features, and procedures of one or more embodiments described herein.
[0079] Program code embodying the algorithms and / or methods described herein can be distributed as a program product in a variety of forms. Program code can be distributed using computer-readable storage media, computer-readable storage devices, or computer networks, as exemplary forms of program product. Computer-readable storage media, computer-readable storage devices, or computer networks can be used to distribute program code in a variety of ways. Computer-readable storage media, computer-readable storage devices, or computer networks can include any tangible, non-transitory medium that is made available to a computer, other programmable data processing apparatus, or other device to enable the computer, other programmable data processing apparatus, or other device to read data, instructions, messages or any other information from the computer-readable storage media, computer-readable storage devices, or computer networks. Computer-readable storage media, computer-readable storage devices, or computer networks can also include any medium that can be used to store information for access by a computer, other programmable data processing apparatus, or other device, including a memory device or storage device that can be connected to a computer, other programmable data processing apparatus, or other device. Program code segments can be downloaded to a computer, other programmable data processing apparatus, or other device from the computer-readable storage media, computer-readable storage devices, or computer networks, or to an external computer or external storage device, via a network.
[0080] Computer readable program instructions stored in the computer-readable media can be used to direct a computer, other programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored in the computer-readable media produce an article of manufacture including instructions which implement the functions, acts, and / or operations specified in the flowcharts or diagrams. In some alternative embodiments, functions, acts, and / or operations specified in the flowcharts or diagrams can be reordered, processed serially, and / or processed in parallel, according to one or more embodiments. In addition, any one of the flowcharts or diagrams can include more or fewer nodes or blocks than shown, according to one or more embodiments.
[0081] The processes, methods, or algorithms can be embodied directly in hardware, software, firmware, or any combination thereof, and can be implemented indirectly as computer-readable storage media encoded with a computer program, which can be executed by a processor. Computer-readable storage media can include any medium for storing information, including volatile and non-volatile, removable and non-removable storage devices, and tangible and non-tangible media.
[0082] Figure 6An illustration of a control system 502 configured to control a vehicle 600, which can be an at least partially autonomous vehicle or an at least partially autonomous robot, is described. In an example, the control system 502 is configured to control the vehicle 600 by executing a vision model in accordance with the principles of the present disclosure. The vehicle 600 includes actuators 504 and sensors 506. The sensors 506 can include one or more video sensors, cameras, radar sensors, ultrasonic sensors, LiDAR sensors, and / or position sensors (e.g., GPS). One or more of the one or more particular sensors can be integrated into the vehicle 600. Alternatively or additionally to the one or more particular sensors described above, the sensors 506 can include a software module configured to determine a state of the actuators 504 when executed. One non-limiting example of a software module includes a weather information software module configured to determine a current or future state of the weather proximate to the vehicle 600 or other location.
[0083] The classifier 514 of the control system 502 of the vehicle 600 can be configured to detect objects proximate to the vehicle 600 from the input signal x. In such embodiments, the output signal y can include information characterizing the proximity of the objects to the vehicle 600, such as pose estimates obtained from the vision model. The actuator control commands 510 can be determined from this information. The actuator control commands 510 can be used to avoid collisions with the detected objects.
[0084] In some embodiments, the vehicle 600 is an at least partially autonomous vehicle, and the actuators 504 can be embodied in brakes, propulsion systems, engines, drivetrains, or steering devices of the vehicle 600. The actuator control commands 510 can be determined such that the actuators 504 are controlled such that the vehicle 600 avoids collisions with the detected objects. The detected objects can also be classified according to what the classifier 514 believes they most likely are, such as pedestrians or trees. The actuator control commands 510 can be determined from the classification. In scenarios where adversarial attacks can occur, the above-described system can be further trained to better detect objects or identify changes in lighting conditions or angles of sensors or cameras on the vehicle 600.
[0085] In some embodiments where the vehicle 600 is an at least partially autonomous robot, the vehicle 600 can be a mobile robot configured to perform one or more functions, such as flying, swimming, diving, and stepping. The mobile robot can be an at least partially autonomous lawnmower or an at least partially autonomous cleaning robot. In such embodiments, the actuator control commands 510 can be determined such that propulsion units, steering units, and / or braking units of the mobile robot can be controlled such that the mobile robot can avoid collisions with the identified objects.
[0086] In some embodiments, the vehicle 600 is an at least partially autonomous robot in the form of a gardening robot. In such embodiments, the vehicle 600 can use an optical sensor as the sensor 506 to determine a state of a plant in an environment near the vehicle 600. The actuator 504 can be a nozzle configured to spray a chemical. Depending on an identified species of the plant and / or an identified state, the actuator control command 510 can be determined to cause the actuator 504 to spray an appropriate amount of an appropriate chemical to the plant.
[0087] The vehicle 600 can be an at least partially autonomous robot in the form of a household appliance. Non-limiting examples of household appliances include a washing machine, a stove, an oven, a microwave, or a dishwasher. In such a vehicle 600, the sensor 506 can be an optical sensor configured to detect a state of an object to be processed by the household appliance. For example, in the case that the household appliance is a washing machine, the sensor 506 can detect a state of laundry within the washing machine. The actuator control command 510 can be determined based on the detected state of the laundry.
[0088] Figure 7 A schematic of the control system 502 is described, which is configured to control a system 700 (e.g., manufacturing machine) of a manufacturing system 702 (e.g., portion of a production line), such as a punch tool, a cutter, or a gun drill. The control system 502 can be configured to control an actuator 504, which is configured to control the system 700 (e.g., manufacturing machine). In an example, the control system 502 is configured to control the system 700 by performing a vision model in accordance with the principles of the present disclosure.
[0089] The sensor 506 of the system 700 (e.g., manufacturing machine) can be an optical sensor configured to capture one or more properties of a manufactured product 704. The classifier 514 can be configured to determine a state of the manufactured product 704 from the captured one or more properties. The actuator 504 can be configured to control the system 700 (e.g., manufacturing machine) in accordance with the determined state of the manufactured product 704 for a subsequent manufacturing step of the manufactured product 704. The actuator 504 can be configured to control a function of the system 700 (e.g., manufacturing machine) on a subsequent manufactured product 706 of the system 700 (e.g., manufacturing machine) in accordance with the determined state of the manufactured product 704.
[0090] Figure 8A schematic of a control system 502 configured to control a power tool 800, such as a power drill or driver, having at least a partially autonomous mode is described. The control system 502 can be configured to control an actuator 504 configured to control the power tool 800. In one example, the control system 502 is configured to control the power tool 800 by executing a vision model in accordance with the principles of the present disclosure.
[0091] The sensor 506 of the power tool 800 can be an optical sensor configured to capture one or more properties of the work surface 802 and / or the fastener 804 driven into the work surface 802, which can include performing pose estimation using the vision model. The classifier 514 can be configured to determine a state of the work surface 802 and / or the fastener 804 relative to the work surface 802 from the captured one or more properties. The state can be that the fastener 804 is flush with the work surface 802. The state can alternatively be a hardness of the work surface 802. The actuator 504 can be configured to control the power tool 800 such that a driving function of the power tool 800 is adjusted according to the determined state of the fastener 804 relative to the work surface 802 or the captured one or more properties of the work surface 802. For example, if the state of the fastener 804 is flush relative to the work surface 802, the actuator 504 can discontinue the driving function. As another non-limiting example, the actuator 504 can apply additional or less torque according to the hardness of the work surface 802.
[0092] Figure 9 A schematic of a control system 502 configured to control an automated personal assistant 900, such as a robot, is described. The control system 502 can be configured to control an actuator 504 configured to control the automated personal assistant 900. The automated personal assistant 900 can be configured to control a household appliance, such as a washing machine, a stove, an oven, a microwave, or a dishwasher. In an example, the control system 502 is configured to control the automated personal assistant 900 by executing a vision model in accordance with the principles of the present disclosure.
[0093] The sensor 506 can be an optical sensor and / or an audio sensor. The optical sensor can be configured to receive a video image of a gesture 904 of the user 902. The audio sensor can be configured to receive a voice command of the user 902.
[0094] The control system 502 of the automated personal assistant 900 can be configured to determine actuator control commands 510 configured to control the system 502. The control system 502 can be configured to determine the actuator control commands 510 from sensor signals 508 of the sensors 506, which can include performing pose estimation using a vision model. The automated personal assistant 900 is configured to transmit the sensor signals 508 to the control system 502. The classifier 514 of the control system 502 can be configured to execute gesture recognition algorithms to identify a gesture 904 made by the user 902, determine the actuator control commands 510, and transmit the actuator control commands 510 to the actuators 504. The classifier 514 can be configured to retrieve information from the non-volatile storage in response to the gesture 904 and output the retrieved information in a form suitable for receipt by the user 902.
[0095] Figure 10 An illustration of a control system 502 configured to control a surveillance system 1000 is described. The surveillance system 1000 can be configured to physically control access through a door 1002. The sensors 506 can be configured to detect a scene relevant to deciding whether to grant access. The sensors 506 can be optical sensors configured to generate and transmit image and / or video data. The control system 502 can use such data to detect a face of a person. In an example, the control system 502 is configured to control the surveillance system 1000 by executing a vision model according to the principles of the present disclosure.
[0096] The classifier 514 of the control system 502 of the surveillance system 1000 can be configured to interpret the image and / or video data by matching the identity of a known person stored in the non-volatile storage 516, thereby determining the identity of the person. The classifier 514 can be configured to generate actuator control commands 510 in response to the interpretation of the image and / or video data. The control system 502 is configured to transmit the actuator control commands 510 to the actuators 504. In this embodiment, the actuators 504 can be configured to lock or unlock the door 1002 in response to the actuator control commands 510. In some embodiments, logical access control that is not physical is also possible.
[0097] The monitoring system 1000 can also be a surveillance system. In such embodiments, the sensor 506 can be an optical sensor configured to detect a scene under surveillance, and the control system 502 is configured to control the display 1004. The classifier 514 is configured to determine a classification of the scene, e.g., whether the scene detected by the sensor 506 is suspicious. The control system 502 is configured to transmit actuator control commands 510 to the display 1004 in response to the classification. The display 1004 can be configured to adjust the displayed content in response to the actuator control commands 510. For example, the display 1004 can highlight objects deemed suspicious by the classifier 514. With embodiments of the disclosed system, a surveillance system can predict objects appearing at some time in the future.
[0098] Figure 11 A schematic of a control system 502 configured to control an imaging system 1100, e.g., an MRI device, an x-ray imaging device, or an ultrasound device, is described. In an example, the control system 502 is configured to control the imaging system 1100 by executing a vision model according to principles of the present disclosure. The sensor 506 can be, for example, an imaging sensor. The classifier 514 can be configured to determine a classification of all or part of a sensed image. The classifier 514 can be configured to determine or select actuator control commands 510 in response to a classification obtained by a trained neural network. For example, the classifier 514 can interpret a region of a sensed image as a potential anomaly. In this case, the actuator control commands 510 can be determined or selected to cause the display 1102 to display the imaging and highlight the potential anomaly region.
[0099] While the example embodiments have been described above, these embodiments are not intended to describe all possible forms of the claims. The words used in the specification are words of description rather than limitation, and it is understood that various changes can be made without departing from the spirit and scope of the disclosure. As previously described, features of various embodiments can be combined to form further embodiments of the present disclosure that can not be expressly described or illustrated. While various embodiments can have been described as providing advantages or being preferred over other embodiments or prior art implementations, one of ordinary skill in the art will recognize that one or more features or characteristics can be substituted for other features or characteristics with the same or similar effect. These attributes can include, but are not limited to, cost, strength, durability, life cycle cost, saleability, appearance, packaging, size, maintainability, weight, manufacturability, ease of assembly, and the like. Thus, any embodiment described as not having some features or characteristics that are present in other embodiments is not outside the scope of the present disclosure, and is contemplated as being within the scope of the present disclosure.
Claims
1. A method for performing pose estimation for an image, the method comprising at one or more processing devices: Receive input image; Based on the input image, a pose code is generated, wherein the pose code corresponds to the estimated pose of an object in the input image; Generate box codes corresponding to the bounding boxes of the objects in the input image; Pose estimation for the input image is performed by generating an improved pose of the object using the pose code and the box code; Generate a prediction output for the object in the input image based on the input image and the improved pose; as well as The predicted output is used to control one or more functions of the device.
2. The method according to claim 1, wherein, Generating the predicted output includes: (i) using an image encoder to generate shape code and texture code, and (ii) using the shape code, the texture code, and the improved pose to generate the predicted output.
3. The method according to claim 2, wherein, Generating the predicted output includes using a Neural Radiation Field (NeRF) decoder.
4. The method according to claim 1, wherein, Generating the improved pose involves iteratively calculating the improved pose using the pose code and the box code.
5. The method of claim 4, further comprising iteratively updating the box code using the improved pose.
6. The method of claim 1, further comprising using a multilayer perceptron to obtain a pose loss based on the pose code.
7. The method according to claim 1, wherein, Generating the predicted output includes converting the improved pose into a camera pose and generating the predicted output based on the camera pose.
8. A computing device configured to perform pose estimation for an image, the computing device including a processing device configured to execute instructions stored in a memory to perform the following operations: Receive input image; Based on the input image, a pose code is generated, wherein the pose code corresponds to the estimated pose of an object in the input image; Generate box codes corresponding to the bounding boxes of the objects in the input image; Pose estimation for the input image is performed by generating an improved pose of the object using the pose code and box code; Generate a prediction output for the object in the input image based on the input image and the improved pose; as well as The predicted output is used to control one or more functions of the device.
9. The computing device according to claim 8, wherein, Generating the predicted output includes: (i) using an image encoder to generate shape code and texture code, and (ii) using the shape code, the texture code, and the improved pose to generate the predicted output.
10. The computing device according to claim 9, wherein, Generating the predicted output includes using a Neural Radiation Field (NeRF) decoder.
11. The computing device according to claim 8, wherein, Generating the improved pose involves iteratively calculating the improved pose using the pose code and the box code.
12. The computing device according to claim 11, wherein, The processing device is configured to iteratively update the box code using the improved pose.
13. The computing device according to claim 8, wherein, The processing device is configured to use a multilayer perceptron to obtain the attitude loss based on the attitude code.
14. The computing device according to claim 8, wherein, Generating the predicted output includes converting the improved pose into a camera pose and generating the predicted output based on the camera pose.
15. A computer-controlled machine configured to operate based on a pose estimate generated from a visual model, the computer-controlled machine comprising: The control system is configured to: Receives input images captured by the camera. A pose code is generated based on the input image, wherein the pose code corresponds to the estimated pose of an object in the input image. Generate bounding box codes corresponding to the bounding boxes of the objects in the input image. Pose estimation for the input image is performed by generating an improved pose of the object using the pose code and box code. Based on the input image and the improved pose, generate a prediction output for the object in the input image, and The control signal is output based on the predicted output; as well as An actuator configured to control the operation of the computer-controlled machine based on the control signal.
16. The computer-controlled machine according to claim 15, wherein, Generating the predicted output includes: (i) using an image encoder to generate shape code and texture code, and (ii) using the shape code, the texture code, and the improved pose to generate the predicted output.
17. The computer-controlled machine according to claim 16, wherein, Generating the predicted output includes using a Neural Radiation Field (NeRF) decoder.
18. The computer-controlled machine according to claim 15, wherein, Generating the improved pose involves iteratively calculating the improved pose using the pose code and the box code.
19. The computer-controlled machine according to claim 18, wherein, The control system is also configured to iteratively update the box code using the improved posture.
20. The computer-controlled machine according to claim 15, wherein, The control system is also configured to use a multilayer perceptron to obtain attitude loss based on the attitude code.