Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1344 results about "2d images" patented technology

Building prefabricated part quality detection method based on multi-modal vision

The invention relates to a building prefabricated part quality detection method based on multi-modal vision. According to the method, multi-modal data including 2D image data and 3D point cloud data are obtained, an improved YOLOv8 model is used for performing defect coarse positioning on the 2D image data, defect parameters are calculated, defect areas such as cracks and exposed ribs can be quickly locked, and the parameters of the defect areas can be obtained. And by means of SIFT feature matching and ICP point cloud registration technologies, comparing with a two-dimensional template and three-dimensional geometric parameters of the BIM standard component model to obtain a two-dimensional registration difference chart and a three-dimensional deviation thermodynamic chart. And finally, according to the defect confidence coefficient, the two-dimensional registration difference chart and the three-dimensional deviation thermodynamic diagram, a preset dynamic weighting rule is adopted to carry out joint decision making, and a quality detection result is obtained. According to the method, through the multi-modal data, the improved YOLOv8 model, the point cloud registration technology and the preset dynamic weighting rule, the false detection problem can be effectively solved, the detection reliability and accuracy under the complex working condition are improved, and the quality management level of the building prefabricated part is improved.
Owner:SOUTHWEST JIAOTONG UNIV

System and method for reconstructing 3D scene data from 2D image data

A method and apparatus for reconstructing a three-dimensional (3D) scene from a two-dimensional (2D) input image of the scene using a fully-differentiable transformer-based encoder-decode. A 2D input image encoded into a set of image features using a pre-trained vision transformer model, wherein the vision transformer model is pre-trained with multi-view RGB image supervision and point cloud supervision. The set of image features is projected onto a 3D triplane representation using a transformer decoder to obtain output triplane tokens. A triplane representation is created from the tokens and queried. 3D point features of color and density for volumetric rendering re predicted using a multi-layer perceptron. The geometry of the generated 3D asset is represented with a surface mesh including vertices and triangular faces. A texture map by is created with a multichannel image in UV space. Multiple views of the 3D scene are simultaneously generated based on the surface mesh.
Owner:FUTUREVERSE IP LTD

Using image proccessing, machine learning and images of a human face for prompt generation related to beauty products for the human face

A method includes receiving 2D image data corresponding to a 2D image of a human face. The method further includes determining a textual identifier that describes a facial feature of the human face based on the 2D image data. The method further includes providing, to a generative machine learning model, a first prompt including information identifying the textual identifier that describes the facial feature of the human face. The method further includes obtaining, from the generative machine learning model, a first output identifying, among a plurality of beauty products, a subset of the plurality of beauty products, the subset of the plurality of beauty products related to the facial feature of the human face.
Owner:BRILLIANCE OF BEAUTY INC

Methods and processors for rendering a 3D object using multi-camera image inputs

Methods and processors for rendering a 3D object are disclosed. The method includes acquiring multi-camera image input including first image frames of the 3D object generated by a first camera and second image frames of the 3D object generated by a second camera, acquiring an initial 3D Gaussian Splatting (3DGS) model having a plurality of initial parameters including an initial frame-wise GS parameter and an initial camera-wise GS parameter, generating an adjusted 3DGS model by adjusting, based on the multi-camera image input, at least one of: the initial frame-wise GS parameter, the initial camera-wise GS parameter, generating, by the adjusted 3DGS model, a 3DGS output and rendering a 2D image of the 3D object using the 3DGS output.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Tunnel fracture identification method based on three-dimensional live-action reconstruction and orderly reacquisition of virtual camera

The invention discloses a tunnel fracture recognition method based on three-dimensional live-action reconstruction and ordered re-acquisition of a virtual camera. The method specifically comprises the following steps: acquiring a multi-view rock fracture image; acquiring three-dimensional point cloud data of a tunnel wall surface according to the acquired rock mass fracture image, reconstructing a geometrical shape of the tunnel, mapping acquired texture information to a three-dimensional grid vertex through an image processing algorithm, and generating a three-dimensional tunnel model; setting parameters of the virtual camera, orderly collecting to form a virtual image, and enabling the position of the virtual camera to correspond to the real position of the model; and extracting information of the crack of the virtual image, and projecting the coordinates of the two-dimensional image back to the three-dimensional space through the projection matrix to obtain the space coordinates of the crack. A virtual image is generated by combining a texture mapping technology of a three-dimensional model, so that the limitation of a visual angle, illumination and space in actual image acquisition is overcome, and high-precision identification of tunnel cracks is realized.
Owner:XIAN UNIV OF TECH

RMIP: fast tessellation-free GPU displacement ray tracing via inversion and oblong bounding simulation

A system generates, based on a displacement bounds data structure and a triangle mesh modeling a surface of a 3D virtual object within a 3D virtual scene, a displaced triangle mesh including one or more displaced surface bounding prisms, each of the one or more displaced surface bounding prisms displaced from a respective base triangle of a plurality of base triangles of the triangle mesh structure based on displacement bounds defined in a displacement bounds data structure for an area of a 2D texture space corresponding to a location of the respective base triangle defined by the 3D virtual scene. The system performs, using the displaced triangle mesh structure, a ray tracing process for a ray associated with a pixel of a 2D image of the virtual scene including determining, responsive to determining the ray intersects the particular displaced surface bounding prism, a location of an intersection of the ray.
Owner:ADOBE INC

Shield segment slab staggering detection method and system fusing visual and geometric features

The invention discloses a shield segment dislocation detection method and system fusing visual and geometric features, and the method comprises the steps: employing a mobile track three-dimensional laser scanning system, and rapidly obtaining the three-dimensional point cloud data of a tunnel segment; projecting the three-dimensional point cloud data of the tunnel segment to a two-dimensional image according to the scanning parameters and the image preset resolution, and mapping the laser point reflection intensity into a pixel gray value; shield segment inter-ring joints and shield segment in-ring joints are extracted respectively, and segment joint positioning is completed; and according to the seam positioning information, extracting local point clouds at the two sides of the seam, respectively carrying out circular model fitting, calculating the height difference of circular models at the two sides at the seam, and obtaining an in-ring slab staggering value. According to the method, the visual information of the point cloud reflection intensity and the geometric information of the spatial position are fused, rapid and accurate shield segment in-ring slab staggering detection is realized, and the efficiency and the accuracy are high.
Owner:CHINA RAILWAY DESIGN GRP CO LTD

Method and system for recovering a three-dimensional human mesh in camera space

A method for recovering a 3D mesh of N humans in a 3D scene comprises: encoding a 2D image from an image capturing device to extract embedded features for each of a plurality of regions; detecting N humans in N respective regions among the plurality of regions; processing the embedded features in the N respective regions and the embedded features for each of the plurality of regions to predict body model and depth parameters using a decoder comprising a cross-attention module; providing the predicted body model parameters to a 3D parametric model for generating 3D meshes; and placing the generated 3D meshes at respective 3D spatial locations based on the predicted depth parameters.
Owner:NAVER CORP

Laser welding line on-line quality detection system and method based on multi-mode visual fusion

The invention discloses a laser welding seam online quality detection system and method based on multi-mode visual fusion, and particularly relates to the technical field of laser welding seam quality detection. Comprising a data acquisition and synchronous control unit, a 2D image processing and defect detection unit, a 3D point cloud processing and size measurement unit, a multi-modal information fusion and collaborative analysis unit and a comprehensive quality judgment and output unit. According to the laser welding seam online quality detection system and method based on multi-modal visual fusion, the 2D and 3D multi-modal visual fusion technology is adopted, an improved defect detection algorithm and an accurate size measurement method are combined, welding seam surface defects are accurately recognized through 2D images, key size data are obtained by means of 3D point cloud, and the detection accuracy is improved. And the limitation of'heavy defects and light sizes' or'heavy sizes and light defects' in single-modal detection is avoided, and comprehensive and accurate evaluation of the welding seam quality is realized.
Owner:SUZHOU UNIV OF SCI & TECH

Single 2d image capture system, processing & display of 3D digital image

A system to capture a two dimensional digital source image of a scene by a user, including a smart device having a memory device for storing an instruction, a processor in communication with the memory and configured to execute the instruction, a digital image capture device in communication with the processor, said processor configured to capture a first two dimensional digital source image of the scene, said processor configured to execute an instruction to generate a second two dimensional digital image of the scene from said first two dimensional digital image of the scene via a camera angle rotation of between 1-180 degrees of said first two dimensional digital image of the scene, and a display in communication with the processor, the display configured to display a multidimensional digital image.
Owner:NIMS JERRY +2

Self-adaptive mechanical arm clamping control method and system based on image recognition

The invention is suitable for the technical field of mechanical arm control, and provides a self-adaptive mechanical arm clamping control method and system based on image recognition, the system comprises a mechanical arm, and the mechanical arm comprises a mechanical arm body and a mechanical clamping jaw; the image acquisition module is used for acquiring a 2D image and depth information of an object; the image processing module is used for carrying out preprocessing, contour extraction and pose coordinate determination on the acquired image; the data transmission module is used for transmitting the contour, size and pose coordinate information of the object processed by the image processing module to the controller; the controller is used for receiving information sent by the data transmission module and controlling movement and clamping force of the mechanical arm; according to the mechanical arm, the path can be automatically planned and proper force can be applied according to the characteristics of different objects, the flexibility, the accuracy and the safety of the clamping operation of the mechanical arm are improved, and the mechanical arm is suitable for various automatic scenes.
Owner:SHANGHAI SECOND POLYTECHNIC UNIVERSITY

Dynamic digital twinning system based on 3D Gaussian splashing

The invention discloses a dynamic digital twinning system based on 3D Gaussian splashing. The dynamic digital twinning system comprises a flexible electric power inspection robot, a 3DGS dynamic enhancement module and a fusion digital twinning background centralized control system. The flexible electric power inspection robot is used for collecting 2D images and environment point cloud data of substation equipment; the 3DGS dynamic enhancement module is used for carrying out 3DGS dynamic enhancement and generating a 3D Gaussian model; the standard digital twinborn centralized control platform is used for storing and calling data and digital twinborn bodies of the traditional static substation digital model; the model fusion port carries out space alignment and attribute interpolation fusion on the 3D Gaussian model and the static digital twin; and the visual terminal is used for displaying the fused enhanced digital twinborn body and the equipment state evaluation result. According to the invention, the dynamic enhancement of the digital twinborn model of the substation equipment can be realized, and the operation state and change condition of the equipment can be reflected more truly and intuitively.
Owner:ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD

Deburring method combining 2D image and 3D point cloud

The invention discloses a deburring method combining a 2D image and a 3D point cloud, and relates to the technical field of image processing, and the method comprises the steps: carrying out the point cloud registration under the verification of the surface orientation consistency of a standard point cloud and an actual point cloud; constructing a pixel mapping matrix between the workpiece image and the RGB image; semantic-level edge coarse extraction is carried out on the workpiece image, and extraction of a sub-pixel-level actual edge contour is carried out on the basis of coarse extraction in combination with a traditional edge detection algorithm; aligning the actual edge contour to the actual point cloud through the pixel mapping matrix; and the target polishing surfaces are aligned to the actual point cloud, to-be-polished edge contours are screened based on the distances between the contours of all the target polishing surfaces in the actual point cloud and all the actual edge contours, and a polishing track is obtained in the workpiece image. According to the method, the sub-pixel edge positioning capability of the 2D image and the spatial topology information of the 3D point cloud are fused, and the inherent defect of a single mode is overcome.
Owner:ZHEJIANG YIMU INTELLIGENT TECH CO LTD

Generating three-dimensional models using machine learning models

The present disclosure describes techniques for generating three-dimensional models using machine learning models. A two-dimensional (2D) image is input into a machine learning model. The machine learning model is configured to generate three-dimensional (3D) models with accurate geometry and detailed textures. A set of multi-view images is generated based at least in part on the 2D image by a first sub-model of the machine learning model. The first sub-model comprises a multi-level image prompt controller configured to implement hierarchical controls over generating multi-view images by the first sub-model based at least in part on an input image. A 3D model is generated based at least in part on the set of multi-view images by a second sub-model of the machine learning model. The second sub-model is configured to implement a background alignment and a camera alignment for improving quality and geometric accuracy of the generated 3D models.
Owner:LEMON INC(GB)

Imaging equipment positioning guiding method, system and equipment based on sparse reconstruction

The invention discloses an imaging equipment positioning guiding method, system and equipment based on sparse reconstruction, and the method comprises the steps: employing imaging equipment, and automatically collecting 2D images of a designated region of a simulation patient from a plurality of sparse angles; selecting a target tissue to be modeled on the 2D image; performing 3D modeling on the selected target tissue by adopting an extremely sparse reconstruction algorithm to form a 3D modeling image of the target tissue; registering the 3D modeling image and the 2D image, and calculating a first conversion relation between a visual coordinate system and a motion coordinate system of the imaging equipment; adjusting the initial pose of the 3D modeling image, and calculating the target position of the imaging device corresponding to the adjusted pose based on the first conversion relation and the adjusted pose of the 3D modeling image; and controlling the imaging equipment to reach the target position so as to obtain an expected imaging view angle of the target tissue.
Owner:JIANGSU FIRST-IMAGING MEDICAL EQUIPMENT CO LTD

Generating textures from text and models

A three-dimensional (3D) texture is generated for an input 3D model based on input text describing the desired texture. Example methods include rendering the 3D model and trainable 3D texture to generate a first two-dimensional (2D) image, adding noise to the first 2D image to generate a first 2D image with added noise, and inputting the first 2D image with added noise and input text into a trained neural network to generate a predicted noise of the first 2D image with added noise. The methods further include determining a loss between the first 2D image with added noise and the predicted noise and updating the trainable 3D texture based on the loss. The method is repeated for a number of time or until a loss between the first 2D image and the predicted noise transgresses a threshold.
Owner:SNAP INC

Point cloud compression with supplemental information messages

A system comprises an encoder configured to compress attribute information and / or spatial for a point cloud and / or a decoder configured to decompress compressed attribute and / or spatial information for the point cloud. To compress the attribute and / or spatial information, the encoder is configured to convert a point cloud into an image based representation. Also, the decoder is configured to generate a decompressed point cloud based on an image based representation of a point cloud. Additionally, an encoder is configured to signal and / or a decoder is configured to receive a supplementary message comprising volumetric tiling information that maps portions of 2D image representations to objects in the point. In some embodiments, characteristics of the object may additionally be signaled using the supplementary message or additional supplementary messages.
Owner:APPLE INC

3D model generation using multimodal generative ai

In various examples, systems and methods are disclosed relating to generating an output 3D latent representation by encoding, using a text encoder, a text prompt and encoding, using a 2D / 3D encoder, a 2D image of an object or a 3D representation of the object. A 3D output is generated by applying the output 3D latent representation to a decoder. A reconstruction loss and a SDS loss are determined for the 3D output. At least one of the text encoder, the 2D / 3D encoder, and the decoder is updated using the reconstruction loss and the SDS loss.
Owner:NVIDIA CORP

Semantic simultaneous localization and mapping method and system based on Gaussian splashing

The invention relates to a semantic simultaneous localization and mapping method and system based on Gaussian splashing. The method comprises the following steps: firstly, collecting a frame of RGB-D image, modeling a scene into a 3D semantic Gaussian field containing a plurality of 3D semantic gausses according to the RGB-D image, and rendering the 3D semantic gausses by using a tile rasterization technology to obtain 2D image plane gausses; rendering results of RGB color, depth and semantic features are extracted from the 2D image plane in a Gaussian mode, and the semantic features are decoded into semantic tags; constructing a mapping and tracking loss function by using an RGB color rendering result, a depth rendering result, a semantic tag and a truth value, and jointly optimizing a camera pose and a semantic Gaussian field based on a tracking stage and a mapping stage; and repeating the steps for each new frame of RGB-D image to complete the construction of the incremental semantic Gaussian map. Compared with the prior art, the method has the advantages of realizing robust camera tracking, real-time high-quality rendering, accurate 3D semantic reconstruction and the like.
Owner:TONGJI UNIV

Multimodal free space prediction by cross-modal deformable stixel predictor

Example systems and techniques are described for controlling operation of a vehicle. An example system includes one or more memories configured to store a machine learning model and one or more processors. The one or more processors are configured to obtain two-dimensional (2D) image data and three-dimensional (3D) point cloud data. The one or more processors are configured to generate one or more multimodal fused 3D stixels based on the 2D image data and the 3D point cloud data. As part of generating the one or more multimodal fused 3D stixels, the one or more processors are configured to execute a machine learning model, the machine learning model having been trained with a 3D stixel correction. The one or more processors are configured to control operation of a vehicle based on the one or more multimodal fused 3D stixels.
Owner:QUALCOMM INC

Spinal medical image registration method and system based on semantic reconstruction

The invention provides a spinal medical image registration method and system based on semantic reconstruction. Constructing spinal column 3D-CT image data and corresponding 2D image data of the spinal column 3D-CT image data; providing a 2D-3D reconstruction model, and reconstructing the 2D image data to obtain a 3D feature map b; providing a 3D-3D feature extraction model, and taking 3D-CT image data as input to obtain a feature map s with the same dimension as the 3D feature map; providing a parallel double-U-shaped network model based on cross attention, and fusing and registering the feature map b and the feature map s; a multi-weight loss function based on pixel and semantic information is adopted for optimization; and training to obtain a registration model, and performing spine medical image registration by using the registration model. According to the method, the information and precision loss caused by dimension reduction is avoided, the 2D / 3D information fusion capability is improved, the attention weight of the key semantic region is improved, and accurate spine image registration can be realized while the image shooting times can be reduced.
Owner:SHANGHAI JIAOTONG UNIV

Systems and methods for calculating refueling tanker boom 3D position for aerial refueling

Disclosed herein is methods, systems, and aircraft for performing image analysis for aiding refueling operations. A tanker aircraft includes a camera, a refueling boom, a camera configured to generate a two-dimensional (2D) image of the refueling boom, a processor, and non-transitory computer readable storage media storing code. The code is executable by the processor to perform operations including receiving the two-dimensional (2D) image from the camera, determining 2D keypoints of the refueling boom located within the 2D image based on a predefined point model of the refueling boom, determining a 6 degree-of-freedom (6DOF) pose using the 2D keypoints and the corresponding three-dimensional (3D) space 3D keypoints, optimizing 3D keypoints associated with moveable components of the refueling boom in response to a plurality of boom control parameters to produce optimized 3D keypoints, and estimating a position of a tip of the refueling boom based on the 6DOF pose.
Owner:THE BOEING CO

Machine-Learned Monocular Depth Estimation and Semantic Segmentation for 6-DOF Absolute Localization of a Delivery Drone

A method includes receiving a two-dimensional (2D) image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV. The method further includes applying a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment, where the semantic image comprises one or more semantic labels. The method additionally includes retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels. The method also includes aligning the depth image of the environment with the reference depth data representative of the environment to determine a location of the UAV in the environment, where the aligning associates the one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data.
Owner:WING AVIATION LLC

Method of estimating uncertainty in a vision-based tracking system and associated apparatus and system

Methods, apparatuses and systems of estimating uncertainty in a vision-based tracking system are disclosed. The method includes receiving a two-dimensional (2D) image of at least a portion of a first object via a camera on a second object. A set of keypoints are predicted on the first object in the 2D image by each one of a plurality of keypoint detectors (i.e., neural networks), organized into an ensemble. A three-dimensional (3D) pose is predicted for each one of the plurality of keypoint detectors from the corresponding set of keypoints. Additionally, the method includes deriving a measure of variation between each one of the 3D poses of the plurality of keypoint detectors and computing a Euclidean norm of the measure of variation to produce an uncertainty value. A process between the first object and the second object can be controlled in response to the calculated uncertainty value.
Owner:THE BOEING CO

Three-dimensional point clouds based on images and depth data

Techniques are discussed herein for generating three-dimensional (3D) representations of an environment based on two-dimensional (2D) image data, and using the 3D representations to perform 3D object detection and other 3D analyses of the environment. 2D image data may be received, along with depth estimation data associated with the 2D image data. Using the 2D image data and associated depth data, an image-based object detector may generate 3D representations, including point clouds and / or 3D pixel grids, for the 2D image or particular regions of interest. In some examples, a 3D point cloud may be generated by projecting pixels from the 2D image into 3D space followed by a trained 3D convolutional neural network (CNN) performing object detection. Additionally or alternatively, a top-down view of a 3D pixel grid representation may be used to perform object detection using 2D convolutions.
Owner:ZOOX INC

Method of training a neural network for vision-based tracking and associated apparatus and system

Methods, apparatuses and systems of training a neural network using in a vision-based tracking system are disclosed. A two-dimensional (2D) image of at least a portion of an object is received via a camera that is fixed, relative to the object. Subsequently, a keypoint detector predicts a set of keypoints on the object in the 2D image, generating predicted 2D keypoints. These predicted 2D keypoints are then projected into three-dimensional (3D) space, and keypoint depth information is added to generate predicted 3D keypoints. To enhance the training process, a 3D model of the object is utilized. Known rotational and translational information of the object in the 2D image is incorporated to known 3D model keypoints, resulting in transformed 3D model keypoints. Following this, a comparison between predicted 3D keypoints and transformed 3D model keypoints is made to calculate a loss value. The training process is further refined using an optimizer, minimizing the loss value during a training period.
Owner:THE BOEING CO

3D generative model training method and device based on hybrid training framework, equipment and storage medium

The invention discloses a 3D generative model training method and device based on a hybrid training framework, equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: pre-training a 3D generation model by utilizing a preset 3D data set; quickly generating initial 3D assets according to the text cue words through a 3D generation model; a strategy based on two-dimensional knowledge distillation is adopted, a text-to-2D image generation model is used as a teacher model to carry out iterative distillation optimization on the initial 3D assets, and the 3D assets with improved visual fidelity are obtained; in the distillation process, dynamically adjusting the teacher model to adapt to the distribution of the initial 3D assets by using a self-adaptive teacher model guide strategy; and taking the optimized high-fidelity 3D assets as enhanced training samples, and training the 3D generation model again to improve the generation quality. According to the method, the 3D generation model is pre-trained, so that the 3D generation model is initially converged; according to the distillation method, strong text-to-2D image generation model knowledge is transferred to initial 3D assets and serves as an enhanced sample to train a 3D generation model again for knowledge internalization. The adaptive teacher model guiding strategy reduces the distribution difference between the teacher model and the student model, and ensures the generation quality of the 3D generation model.
Owner:张家辉 +2

3DGS equipment monomer method based on dynamic segmentation of SAM2 large model

The invention relates to the technical field of computer vision and three-dimensional modeling, and discloses a 3DGS (three-dimensional ground structure) equipment monomer method based on dynamic segmentation of an SAM2 large model, which comprises the following steps of: calling a visual large model to segment a multi-view two-dimensional image to obtain a component contour mask bearing semantic information; the method comprises the following steps of: establishing a mask of a three-dimensional point cloud, establishing space mapping of the mask and the three-dimensional point cloud, executing structure segmentation on the point cloud under double dimensions of fusing geometry and semantics, and further generating a three-dimensional monomer model for each separated equipment component. The segmentation processing of the three-dimensional point cloud is no longer limited to fuzzy geometric inference, but is carried out under clear semantic affiliation, so that the accuracy and robustness of monomer separation in the component dense adjacent region are improved.
Owner:BEIJING HUAQING QIHANG TECH CO LTD

Systems and methods for processing 2d / 3d data for structures of interest in a scene and wireframes generated therefrom

Examples relate generally to improvements in generation of wireframe renderings derived from 2D and / or 3D data that includes at least one structure of interest in a scene. Such wireframe renderings and similar formats can be used in, among other things, 2D / 3D CAD drawings, designs, drafts, models, building information models, augmented reality or virtual reality, and the like. Measurements, dimensions, geometric information, and semantic information generated according to the inventive methods can be accurate in relation to the actual structures. The wireframe renderings can be generated from a combination of a plurality of 2D images and point clouds, processing of point clouds to generate virtual / synthetic views to be used with the point clouds, or from 2D image data that has been processed in a machine learning process to generate 3D data. In some aspects, the wireframe renderings are accurate in relation to the actual structure of interest, automatically generated, or both.
Owner:BENTLEY SYSTEMS CAPITAL LLC