Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

334 results about "Back projection" patented technology

What is Back Projection?¶. Back Projection is a way of recording how well the pixels of a given image fit the distribution of pixels in a histogram model. To make it simpler: For Back Projection, you calculate the histogram model of a feature and then use it to find this feature in an image.

Unmanned aerial vehicle pose visual angle optimization method and system for fracture refined shooting

The invention provides an unmanned aerial vehicle pose visual angle optimization method and system for fracture refined shooting. The method comprises the steps of performing fracture detection and boundary extraction on a coarse inspection image; recovering a camera track and sparse point cloud based on multi-view three-dimensional reconstruction, carrying out back projection and estimating a normal vector of a crack surface; constructing a shooting spherical shell with limited inner and outer radiuses by taking the crack point as a center, and generating a view cone and spherical shell intersection region allowed to be shot as a candidate set in combination with a normal vector; sampling in the candidate area to generate a plurality of candidate shooting points, and synchronously resolving the flight and holder integrated pose of the corresponding unmanned aerial vehicle; constructing a multi-target cost function including path length, attitude, pan-tilt angle and shooting error, and generating an optimal shooting point sequence and an inspection path through an optimization algorithm; the system realizes fine, efficient and automatic shooting of cracks in a complex structure environment through cooperation of multiple modules, and effectively improves the imaging quality and the detection precision.
Owner:SHANDONG XIEHE UNIV +1

Semantic aerial view visual relocation method and device in non-exposed scene, electronic equipment, storage medium and program product

The invention provides a semantic aerial view visual relocation method and device in a non-exposed scene, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a multi-view image sequence under a non-exposed scene (such as a tunnel, an underground pipe gallery or an underground parking lot); semantic recognition is carried out based on a pre-trained semantic target detection model, and spatial consistency semantic features are extracted through a semantic-geometric dual-channel fusion mechanism combining a semantic mask and geometric constraints; the method comprises the following steps of: realizing three-dimensional reconstruction by using a voxel micro-renderable modeling method (VGGT), and generating a dense three-dimensional semantic point cloud fusing semantics and a geometric structure; two-dimensional semantics are mapped to a three-dimensional space through a projection and back projection relation, and point cloud semantics are endowed; main structure planes such as the ground, the left wall surface and the right wall surface are extracted, and a two-dimensional semantic aerial view with semantic annotation is generated; and pose estimation is carried out based on a reciprocal matching strategy guided by a semantic mask, so that visual repositioning with high precision, high robustness and semantic interpretability is realized. The method breaks through the problems of low precision, sparse features and poor semantic consistency of traditional visual repositioning in a non-exposed environment, and can be widely applied to the fields of intelligent transportation, underground inspection and unmanned system positioning.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Multi-modal remote sensing target tracking positioning and intention discrimination method and device

The invention provides a multi-mode remote sensing target tracking and positioning and intention discrimination method and device. The method comprises the following steps: acquiring a plurality of visible light image frames and a plurality of infrared light image frames, and carrying out frame alignment operation on each visible light image frame and each infrared light image frame to obtain a plurality of groups of effective image frame pairs; for each group of effective image frame pairs, determining tracking identification information of each detection object in the effective image frame pairs based on the effective image frame pairs and a pre-trained multi-modal detection tracking model; for each detection object, determining longitude and latitude tracks of the detection object based on the tracking identification information and a back projection mapping function; and determining the behavior intention of each detection object based on a behavior recognition model and the longitude and latitude tracks of each detection object. The accuracy of target tracking and behavior intention recognition in the remote sensing video can be improved.
Owner:AEROSPACE INFORMATION RES INST CAS

Deep learning-based CT artifact removal method and system

The present invention relates to the technical field of medical images. Disclosed are a deep learning-based CT artifact removal method and system. The method comprises: acquiring a CT image, and separately performing sine transform and wavelet transform processing on the CT image; constructing an image enhancement model, and performing image optimization on the processed image separately by means of the image enhancement model and a random inversion layer which are connected in sequence; coupling the optimized image and the original image and then inputting the coupled image into the image enhancement model for reprocessing; and performing element-wise addition on the reprocessed image and the optimized image to obtain an artifact-removed CT image. In the present invention, wavelet transform is introduced to process the CT image to extract the context and spatial information of the CT image, effectively extracting feature information in an artifact removal process and improving the performance of image enhancement; and a CT image resolution enhancement model based on a VMamba model is established, enhancing the long-term dependencies in network training, effectively recognizing and removing radioactive artifacts, and improving the network training efficiency.
Owner:PEKING UNIV SCHOOL OF STOMATOLOGY

Method for optimizing micro-crack segmentation based on deep learning and super-resolution reconstruction

The invention discloses a method for optimizing micro-crack segmentation based on deep learning and super-resolution reconstruction, and the method comprises the steps: cutting an image in real time, obtaining an image block which takes a component as a target main body, and synchronously recording a homography matrix for geometric mapping; selecting an amplification strategy to improve the resolution, and recording a scale mapping relation; inputting the enhanced image block into a double-flow network; adaptive fusion and reconstruction are carried out on the two branch features, and a high-resolution texture image is output; generating a geometrically corrected ortho-image, fusing the geometrically corrected ortho-image with original illumination information, and outputting a corrected image with a known pixel size; identifying cracks, spalling and honeycomb diseases in parallel; generating a unified defect confidence map; calculating real geometric parameters of the BIM in a BIM global coordinate system through coordinate back projection; and generating quantitative defect reports and maintenance suggestions. The method has the advantage that seamless connection between the detection result and the BIM global coordinates is realized.
Owner:CHINA RAILWAY SHANGHAI DESIGN INST GRP CO LTD +1

Vision-driven multi-modal fusion lightweight semantic map construction method and system

The invention discloses a vision-driven multi-mode fusion lightweight semantic map construction method and system. The method comprises the following steps: synchronously acquiring a binocular image pair sequence, IMU data and GNSS data of a target area; based on the acquired multi-modal data, performing multi-sensor joint state estimation through a differential weighted fusion strategy, and outputting camera global pose and scene depth information; based on a current frame and a historical frame in the binocular image pair sequence, combining a camera global pose, extracting geometric prior auxiliary time sequence cross-frame semantic feature fusion through stereoscopic vision, and outputting a two-dimensional semantic segmentation result of the current frame; and back-projecting the two-dimensional semantic segmentation result into a lightweight global three-dimensional semantic map based on camera pose and scene depth information, and carrying out maintenance and updating through a voxelization statistical mechanism. Compared with a traditional dense point cloud map, the method has the light weight effect that the storage space is greatly reduced.
Owner:BEIHANG UNIV

Art and craft material detection method based on multi-modal deep learning

The invention discloses an industrial art material detection method based on multi-modal deep learning, particularly relates to the field of material analysis, and is used for solving the problems that cross-modal data alignment of curved surface utensils in highlight and multi-layer coating scenes is difficult, and material boundaries are easy to drift along with shooting batches and visual directions. The method comprises the following steps of: positioning a hyperspectral line scanning fragment; constructing a local reference grid to realize initial pairing; generating pixel-level residual image quantization dislocation; estimating luminosity mapping to form consistent bimodal data pairs after eliminating a high-reflectivity region; iteratively supplementing anchor points or adjusting grid rigidity, generating a stable material graph, updating an index table and a consistency record, ensuring cross-batch reusability and traceability, and improving the material detection precision.
Owner:WUXI GONGCHUN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Bridge disease spatial form quantitative characterization method based on fusion of three-dimensional laser point cloud and two-dimensional image

The invention discloses a bridge disease spatial form quantitative characterization method based on fusion of a three-dimensional laser point cloud and a two-dimensional image. The method comprises the following steps: synchronously acquiring image data and laser point cloud data of a bridge disease; the data sum is preprocessed; constructing a PSAG-Net model, and inputting the point cloud data into the PSAG-Net model for training; constructing a CM-FPN model, and inputting the image data into the CM-FPN model for training; segmenting the crack point cloud data by using a PSAG-Net model, carrying out post-processing on the completely segmented large-scale crack point cloud, complementing the missing crack region, calculating the depth value of the large crack, and counting the depth distribution; and for the small-scale crack point cloud which is not completely segmented, converting the small-scale crack point cloud into a depth map and carrying out binarization, carrying out crack pixel-level detection and segmentation by using a CM-FPN model, carrying out back projection to a three-dimensional coordinate system, and calculating a depth value. According to the invention, accurate quantitative characterization of depth information is realized.
Owner:SOUTHEAST UNIV

Plane CT fan-beam projection filtering back projection reconstruction method and CT detection equipment

The invention relates to the technical field of CT detection, and discloses a plane CT fan-beam projection filtering back projection reconstruction method and CT detection device.The method comprises the steps that S1, tilt geometric configuration parameters are determined according to the distance from a detector to a rotation center and a tilt angle, a sample rotating table is controlled to rotate, and the detector is triggered at each rotation angle to collect first projection data; s2, calculating a global three-dimensional coordinate of a detector pixel point under each rotation angle, and generating second projection data; s3, performing frequency domain filtering according to the thickness of the sample and the second projection data to obtain third projection data; and S4, calculating a ray intersection point coordinate from a reconstructed slice pixel point to a detector plane, and reconstructing the third projection data according to the ray intersection point coordinate to obtain a tomographic reconstruction image.According to the method, high-frequency detail features are reserved while noise is suppressed, the problem that a traditional fixed filtering kernel parameter is poor in adaptability to different samples is avoided, and the method is suitable for being applied to the field of image reconstruction. Therefore, high-precision three-dimensional imaging is realized.
Owner:SHENZHEN SANYING PRECISION INSTR CO LTD

Systems and methods for determining a location of a gross target volume of a patient

Provided herein are systems for determining a location of a gross target volume of a patient. In some examples, systems can include one or more processors that are configured to obtain image data associated with a plurality of images of a lesion of a patient. For each image, the one or more processors can be configured to backproject points representing the lesion into the 3D space to determine a plurality of distribution confidence values for a subset of voxels within the three-dimensional space. The one or more processors can be configured to determine a three-dimensional confidence distribution based on confidence values from the plurality of distribution confidence values corresponding to each voxel of the 3D space and determine a position of the lesion within the 3D space based on the 3D confidence distribution.
Owner:SIEMENS HEALTHINEERS INTERNATIONAL AG

Multi-fisheye camera aerial view construction method and system for automatic driving

The invention relates to the technical field of vehicle visual perception, and particularly discloses a multi-fisheye camera aerial view construction method and system for automatic driving, and the method comprises the steps: carrying out the offline calibration of each fisheye camera based on a polynomial fisheye imaging model, constructing a pixel back projection lookup table, and carrying out the offline calibration of each fisheye camera; performing table lookup type distortion removal processing on the acquired fisheye image; extracting multi-region checkerboard angular points, and establishing a cross-view-angle corresponding point set; solving an initial homography matrix from each fisheye camera to the aerial view reference plane, constructing a joint objective function including a checkerboard re-projection error, an overlapping region alignment error and a physical scale constraint, and performing global joint optimization on all the homography matrixes; generating each path of aerial view sub-graph by adopting a reverse mapping strategy; according to the method, the calibration process is visual, the geometric accuracy, the real-time performance and the engineering transportability are high, and the method is suitable for automatic parking, a vehicle looking-around system and automatic driving environment perception.
Owner:JILIN UNIVERSITY

DDPM-based CT image metal artifact elimination method

The invention discloses a DDPM-based CT image metal artifact elimination method, and aims to solve the artifact problem caused by a metal object in an existing CT image and improve image quality and diagnosis reliability. A diffusion model of unconditional training is adopted, step-by-step back diffusion repair of an artifact area is carried out in a sinogram domain, an unrepaired area is dynamically adjusted by combining with a metal mask, accurate repair of the artifact area is achieved, and original data of the area which is not affected by artifacts are kept. In the training stage of the system, artifact-free data are gradually converted into standard Gaussian noise through forward diffusion; in the inference stage, data are gradually recovered by utilizing back diffusion, block repair is carried out on an artifact region by combining with a metal mask, and a complete sinogram is generated through region merging. And finally, reconstructing a CT image by using a filtered back projection algorithm, and optimizing boundary transition through a smoothing algorithm to ensure seamless connection between the metal object and surrounding tissues.
Owner:SHANGHAI UNIV

Virtual camera projection method and system, electronic equipment and medium

The invention provides a virtual camera projection method and system, electronic equipment and a medium, and belongs to the technical field of target detection, and the method comprises the steps: constructing a training data set according to sampling points in a three-dimensional space; modeling projection and back projection between the real camera view and the virtual camera view to obtain a two-stage multi-layer perceptron model; training a two-stage multi-layer perceptron model according to the training data set to obtain a trained model; and obtaining a pixel mapping table between the virtual camera view and the real camera view based on the trained model. According to the method, the reversible pixel mapping relation between the real camera and the virtual camera is constructed by adopting the two-stage multi-layer perceptron model, and the pictures under any group of camera parameters are uniformly converted to be under the virtual camera parameters, so that high-quality label data are efficiently generated through the automatic labeling model; the data utilization rate and the generalization ability of the automatic labeling model are remarkably improved, so that the model can directly process multi-source heterogeneous data, and the model adaptation cost is reduced.
Owner:DONGFENG MOTOR GRP

Multi-view feature fusion operable component semantic segmentation method and system

The invention belongs to the field of robot control, and provides a multi-view feature fusion operable part semantic segmentation method and system, and the method comprises the steps: obtaining the point cloud data of an operable part, and generating a corresponding multi-view image; carrying out target detection, extracting features for each view angle, and obtaining a two-dimensional bounding box and a semantic tag; based on the extracted two-dimensional bounding box, processing by using SAM to obtain a foreground mask; on the basis of the extracted two-dimensional bounding box, the relevance of the same semantic target under different visual angles is captured on the global scale by using the constructed global visual angle interaction module, and the feature consistency is enhanced; and processing the fused bounding box features obtained by the global view angle interaction module by using a weight prediction network, predicting the response weight of each bounding box at each super point, combining the obtained foreground mask and the predicted response weight, and back-projecting to a 3D point cloud space through view point information to obtain a final 3D semantic segmentation result. According to the invention, the segmentation accuracy is improved.
Owner:UNIV OF JINAN +1

Multi-view three-dimensional reconstruction method based on full-dimensional dynamic convolution and content-guided attention mechanism

The invention provides a multi-view three-dimensional reconstruction method based on full-dimensional dynamic convolution and a content guidance attention mechanism. The method comprises a multi-scale feature extraction module, a cost body construction and aggregation module, a cost body regularization module and a depth map filtering fusion point cloud generation module. According to the method, feature extraction is carried out in a feature extraction network fusing full-dimensional dynamic convolution and a content-guided attention mechanism, and cross-layer feature transfer is realized by utilizing adaptive modulation convolution kernel capability of the full-dimensional dynamic convolution and content-guided attention, so that depth estimation in a weak texture region is more stable; the depth drift caused by insufficient feature information is reduced; and meanwhile, a weight map of the neighborhood map to the reference map is obtained by utilizing matching correlation, visible information enhancement is carried out on the weight map obtained by carrying out back projection on the neighborhood map in combination with the reference map, and the cost body is guided to aggregate more reliable information, so that error cost information of a shielding region or a visual angle range region is effectively inhibited when the cost body is aggregated. According to the image feature extraction method provided by the invention, the feature expression capability is enhanced, and visible information is further enhanced by using the weight map generated by bidirectional projection, so that a more accurate depth map is obtained, and finally, the accuracy and integrity of three-dimensional reconstruction are remarkably improved.
Owner:南宁桂电电子科技研究院有限公司 +1

Automatic surveying and mapping system and method for urban and rural planning complex terrain based on image analysis

The invention belongs to the technical field of surveying and mapping remote sensing, and discloses an automatic surveying and mapping system and method for urban and rural planning complex terrains based on image analysis. The method comprises the following steps of: acquiring multi-view heterogeneous image data of a target surveying and mapping area, performing layered decoupling processing and multi-view registration on an image, identifying spectral texture features and geometric three-dimensional features of a terrain, and performing complexity function partitioning on the surveying and mapping area according to the spectral texture features and the geometric three-dimensional features; the regional surveying and mapping demand degree is calculated by analyzing a terrain local complexity index and spatial distribution dispersion characteristics, so that a personalized sampling parameter range is screened out; constructing an optimized three-dimensional terrain model by fusing a terrain feature map, a point cloud reconstruction result and a geometric constraint attribute, predicting surveying and mapping precision, and performing comparative analysis on the predicted surveying and mapping precision and a back projection result to obtain a confidence evaluation result; the method is applied to a personalized urban and rural planning surveying and mapping scheme; according to the invention, accurate identification and personalized automatic surveying and mapping of the complex terrain are realized.
Owner:SHANDONG HUIYU AVIATION REMOTE SENSING TECH CO LTD

Small reservoir water level intelligent identification method and system based on general picture

The invention provides a small reservoir water level intelligent identification method and system based on a general picture, and the method comprises the following steps: firstly constructing a three-dimensional twin grid model of a reservoir dam, extracting the three-dimensional coordinates of structural feature points, and storing the three-dimensional coordinates in a database; the method comprises the following steps: acquiring a monitoring image to preliminarily identify an initial water level line, and matching two-dimensional pixel coordinates of corresponding feature points in the image; and solving a real-time pose matrix of the camera based on the matching pair of the three-dimensional and two-dimensional coordinates. And projecting the three-dimensional model to the image by using the pose matrix to generate a virtual measurement grid with three-dimensional information. Pixels on an initial water level line are back-projected to a physical space through a grid to form a three-dimensional point cloud, and an optimal water body plane is fitted through iterative optimization. And finally, calculating the real physical height difference between the water body plane and the dam crest, and subtracting the height difference from the known absolute elevation of the dam crest to obtain the final real-time absolute water level elevation. The method has the effect of improving the accuracy of water level recognition based on image processing.
Owner:HUBEI WATER CONSERVANCY & HYDROPOWER RES INST

Multi-modal three-dimensional point cloud semantic segmentation method for noise self-adaptive filtering

The invention discloses a multi-modal three-dimensional point cloud semantic segmentation method based on adaptive noise filtering, and belongs to the technical field of automatic driving environment perception. The method comprises the following steps: firstly, acquiring data by using a laser radar and a monocular camera, and constructing a multi-modal panoramic feature tensor containing a geometric structure and color textures through projection and mapping; extracting shallow geometric distribution features through a residual context module, and extracting multi-scale environment features through an expanded residual encoder; performing global context aggregation by using a self-attention mechanism of a Transform architecture, and establishing a full-image pixel dependency relationship to make up for a convolution locality defect; in the decoding stage, a channel cross fusion attention module (CCA) is adopted to process deep semantic features and shallow jump connection features, and a channel weight mask is dynamically generated to adaptively screen effective features and suppress high-frequency noise; and finally, outputting a two-dimensional semantic segmentation result and back-projecting the result to a three-dimensional space to obtain a semantic point cloud. According to the method, through global semantic integration and local detail screening, the problems of terrain misjudgment and noise interference of severe weather (such as rain, snow and dust) in a cross-country scene are effectively solved, and the robustness and precision of automatic driving perception are remarkably improved.
Owner:BEIHANG UNIV

Cleaning path planning method for photovoltaic cleaning robot

The invention discloses a method for planning a cleaning path of a photovoltaic cleaning robot. The method comprises the following steps: acquiring a visible light image and an infrared thermogram; performing noise reduction and image enhancement processing on the visible light image and the infrared thermal image, and determining a fault position and a fault type according to the visible light image and the infrared thermal image which are subjected to noise reduction and image enhancement processing; constructing a two-dimensional grid digital orthoimage map of each photovoltaic panel in the photovoltaic power station by using a motion recovery structure and a multi-view three-dimensional algorithm; pixel coordinates of each photovoltaic panel in the two-dimensional grid digital orthophoto map are converted into actual geographic coordinates through a homography-coplanar back projection formula, fault identification points and heavy dirty area identification points are arranged in the two-dimensional grid digital orthophoto map of the actual size, and a cleaning path is generated through a grid search method. And secondary damage possibly caused by the photovoltaic cleaning robot to an existing fault area on the photovoltaic panel in the cleaning process is avoided, the cleaning efficiency is improved, and the service life of the photovoltaic panel is prolonged.
Owner:INNER MONGOLIA UNIVERSITY

CT full-chain intelligent reconstruction method based on deep learning

A CT full-chain intelligent reconstruction method based on deep learning. Original measurement data is sequentially processed by a trained pixel-wise intelligent correction network, a trained angle-wise intelligent filtering network and a trained back-projection tensor intelligent reconstruction network to finally obtain a final reconstructed CT image. In the present invention, firstly, a pixel-wise intelligent correction network can perform pixel-wise adaptive Gaussian filtering on low-dose original measurement data by learning the variance of each pixel point. Moreover, an angle-wise intelligent filtering network can perform angle-wise adaptive filtering on sinogram data by learning a filtering kernel of each projection angle. Finally, by learning the mapping of a back-projection tensor to a target image, a back-projection tensor intelligent reconstruction network can ensure image quality while reducing the scanning dose.
Owner:SOUTHERN MEDICAL UNIVERSITY

Fish body detecting and counting method for passage behind fish pump of fishing boat

The invention relates to a fish body detection counting method for a channel behind a fish pump of a fishing boat, and provides an NMS-free fish body detection model MGI-RTDETR based on an improved RT-DETR for small targets, rapid movement, strong blur, highlight shielding and limited shipborne computing power. The model integrates multi-scale grouping interaction, dynamic context mixing, detail fidelity fusion, deployment period re-parameterization and back projection up-sampling, and real-time detection of one fish and one frame is achieved. And in combination with lightweight multi-target tracking, outputting paragraph counting by adopting a one-way cross-line combined de-duplication strategy. According to the method, missing detection and time delay can be remarkably reduced, edge deployment is facilitated, and the method is suitable for near-real-time fishing estimation, operation monitoring and resource evaluation.
Owner:EAST CHINA SEA FISHERIES RES INST CHINESE ACAD OF FISHERY SCI +1

Forging surface defect space positioning method

The invention relates to a forge piece surface defect space positioning method, and belongs to the technical field of forge piece surface defect intelligent detection.The method comprises the steps that a complete three-dimensional point cloud of a forge piece is obtained through three-dimensional scanning; the method comprises the following steps: performing image acquisition on the surface of a forge piece by using a binocular vision system, performing forge piece surface defect detection on an acquired two-dimensional image, if a defect is detected, performing defect segmentation to obtain a defect image, and performing three-dimensional reconstruction on the local part of the forge piece according to the two-dimensional image of the forge piece to obtain a local three-dimensional point cloud of the forge piece; converting the complete three-dimensional point cloud into a coordinate system of the local three-dimensional point cloud to obtain a target complete three-dimensional point cloud; projecting the target complete three-dimensional point cloud to the two-dimensional image to obtain a depth map; fusing the defect image with the depth map to obtain a fused image; and back-projecting the fused image back to a three-dimensional space to obtain a space coordinate of the forging defect. According to the invention, the positioning precision and the positioning efficiency of forging defects are improved.
Owner:WUHAN UNIV OF TECH

Single-frame unmanned aerial vehicle image pixel positioning method and system based on elevation map

The invention discloses a single-frame unmanned aerial vehicle image pixel positioning method and system based on an elevation map, and relates to the technical field of unmanned aerial vehicle image processing and geographic positioning, and the method comprises the steps: obtaining a single-frame original image shot by an unmanned aerial vehicle, unmanned aerial vehicle attitude information during shooting, camera internal and external parameters, and a digital elevation map of a shooting region; according to the internal reference of the camera, converting the pixel coordinate of the target pixel in the original image into a normalized plane coordinate system of the camera, and carrying out coordinate distortion correction to obtain the back projection point coordinate of the target pixel; converting a camera-unmanned aerial vehicle body coordinate system and a body-geographic inertial reference coordinate system according to camera external parameters and unmanned aerial vehicle attitude information to obtain normalized ray direction vectors corresponding to target pixels; an earth surface height constraint is constructed based on a digital elevation map, and an intersection point of a ray and the ground is determined by performing iterative approximation on a space point along a ray direction. According to the invention, rapid and accurate conversion from the single-frame unmanned aerial vehicle image pixel to the geographic coordinate can be realized.
Owner:SHANDONG SYNTHESIS ELECTRONICS TECH

New energy equipment intelligent inspection method based on laser radar and infrared imaging fusion

The invention relates to the technical field of new energy equipment intelligent detection, in particular to a new energy equipment intelligent inspection method based on laser radar and infrared imaging fusion, and the method comprises the steps: building a unified space-time coordinate system, and completing the external parameter calibration and radiometric calibration of two sensors. Through point cloud registration, voxel division and infrared data back projection, a hot geometric voxel graph is generated in combination with weighted average of multi-view projection weights. And dynamically correcting the voxel temperature by adopting a joint compensation model, and constructing a super voxel graph structure after extracting fusion features. And performing feature weighted fusion and abnormal probability calculation by using a graph attention network, and generating a time sequence change graph in combination with historical baseline data to identify an abnormal type. And finally, the severity is calculated through the anomaly probability, the temperature trend and the neighborhood difference, the anomaly position, the class, the grade, the reinspection track and the inspection period suggestion are output, and accurate anomaly detection and intelligent inspection decision making are achieved.
Owner:HUANENG BAOTOU NEW ENERGY POWER CO LTD +1

Three-dimensional editing method and system based on anchor point visual angle guidance and three-dimensional perception attention pooling

The invention discloses a three-dimensional editing method and system based on anchor point visual angle guidance and three-dimensional perception attention pooling. According to the method, a framework combining a diffusion model and three-dimensional Gaussian rendering is designed and is used for improving the consistency and efficiency of multi-view editing. According to the method, firstly, an anchor point view angle guiding strategy is put forward, editing is carried out through cross-view angle attention on a small number of key view angles, and an editing result is spread to other view angles by utilizing depth information, so that cross-view angle calculation overhead is reduced, and consistency is kept. Then, a three-dimensional perception attention pooling mechanism is constructed, text-image attention results of different visual angles are back-projected to unified three-dimensional Gaussian representation, and global consistent semantic attention distribution is realized through cross-visual-angle aggregation so as to enhance three-dimensional structure perception of a diffusion model. And finally, the three-dimensional model is directly updated by generating a multi-view editing result at one time, so that a redundant iterative optimization process in a traditional method is avoided.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

A 4D video generation method and system fusing monocular vision and terrain elevation data

The application discloses a 4D video generation method fusing monocular vision and terrain elevation data, comprising the following steps: acquiring video frame images and corresponding camera pose data collected by a monocular camera; acquiring terrain elevation data of a target area, and constructing a three-dimensional elevation network model; based on the camera pose data, establishing a ray projection model for pixels in the video frame images, calculating the intersection of the rays and the three-dimensional grid model, and obtaining the reference distance from the monocular camera to the ground surface corresponding to the pixels; generating an initial relative depth map of the video frame images by using a depth estimation model; selecting high-confidence ground pixels from the video frame images as reference points, taking the reference distance corresponding to the reference points as the true value, and performing absolute scale correction on the initial relative depth map to obtain an absolute depth map; performing back projection calculation on the absolute depth map, the camera pose data and pixel color information to generate a three-dimensional point cloud, mapping the three-dimensional point cloud to a real geographic coordinate system, and generating 4D video data in time sequence.
Owner:GUANGXI ACAD OF SCI +1

A multi-modal data fusion and ground feature boundary identification processing method for field verification

The present application relates to the field of natural resource investigation and geographic information processing, and discloses a multi-modal data fusion and ground object boundary identification processing method for field verification; current phase remote sensing images of a region to be verified, reference vector map patches, historical land attribute data, positioning trajectory data of a field verification terminal and a field photo collection correlation field template are acquired to construct a patch-level multi-modal basic data set; candidate ground object segmentation results and change area results are extracted to determine suspected problem patches; a boundary uncertainty band is generated for the boundary of the suspected problem patches and the to-be-verified boundary segments are divided; shooting guidance is triggered during the field verification process, field photos are collected, coarse registration and fine registration with local remote sensing images are completed; edge features of the field photos are back-projected to a map coordinate system to form a field confirmation boundary point set, the boundary of the suspected problem patches is corrected, verification confirmation patches are obtained, and patch-level verification evidence data is formed.
Owner:ANHUI COALFIELD GEOLOGICAL BUREAU EXPLORATION & RESEARCH INSTITUTE

Three-dimensional map construction and camera trajectory estimation method using edge information fusion

The application provides a three-dimensional map construction and camera trajectory estimation method based on edge information fusion, and belongs to the field of computer vision and robot perception. An image sequence in a scene is collected, and a two-dimensional edge intensity map is generated after edge detection; a three-dimensional Gaussian point cloud is initialized or updated: an initial point cloud is generated by back projection of the color, depth and edge intensity value of the first frame; a point cloud is added in the required area according to the rendering difference in subsequent frames; the camera pose of the current frame is initialized, a prediction map is generated based on the pose and the point cloud of the previous frame, a loss function is constructed to optimize the camera pose; the point cloud attributes are optimized again based on the optimized pose and the point cloud of the previous frame, and the point cloud of the current frame is generated; the sequence is processed frame by frame, the optimized point cloud is accumulated to form a global three-dimensional map, and the camera motion trajectory is output by accumulating the camera pose. The application improves the SLAM precision, stability and reconstruction quality in various complex scenes, especially in low-texture and strong-structure environments.
Owner:ZHEJIANG UNIV

An AI vision-based human motion posture correction method, device and medium

This invention discloses a method, device, and medium for correcting human motion posture based on AI vision, relating to the field of AI vision technology. The method includes: acquiring multi-channel image sequences of human posture using multiple cameras; performing temporal alignment and normalization on the multi-channel image sequences; obtaining a human probability map using a pre-trained human semantic segmentation network; performing morphological cleanup to obtain a human silhouette sequence; performing back-projection according to camera intrinsic and extrinsic parameters to obtain a 3D joint sequence from a 3D AI visual volume; extracting the human motion manifold to perform contact-sensing neural motion priors and outputting prior signals; and using posture feasibility scores and joint importance weights to weightedly fuse the geometric difference vector and the center of gravity correction component to obtain a comprehensive joint correction vector. This invention generates a 3D AI visual volume through back-projection and extracts the central skeleton, using constant bone segment lengths for constraint and correction, thus enhancing the reliability of posture stability analysis and alignment with standard action templates.
Owner:延安大学西安创新学院