Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

698 results about "Single image" patented technology

Edge-deployed semi-supervised anomaly detection method and system for railway track foreign object

Disclosed in the present invention are an edge-deployed semi-supervised anomaly detection method and system for a railway track foreign object. The method comprises the following steps: an edge device encoding and decoding a video stream captured by a camera to obtain an image frame sequence, and performing frame extraction; and using a semantic segmentation model to perform image segmentation on a certain image frame obtained by means of frame extraction, to obtain a railway track region segmentation image. The use of a single image as input may generate an expert model result having a high weight value; however, the determination based on a single image is not stable, multiple consecutive images of the task scene need to be inputted, the frequency of each expert model obtaining the highest weight is computed, and the expert model corresponding to the highest frequency is the final solution. The present invention supports scene-adaptive foreign object detection algorithm automatic selection, and a user can perform selection on the basis of prior knowledge, or selection may be performed by a scene-adaptive automatic algorithm selection method; the user only needs to provide a batch of image data of the current scene, and the optimal algorithm selection can be evaluated.
Owner:GUANGZHOU EMBEDDED MACHINE TECH CO LTD

Building defect detection method and intelligent imaging device

The invention relates to the technical field of building structure health monitoring, in particular to a building defect detection method based on infrared thermal imaging and visible light bimodal fusion and a matched intelligent imaging device. According to the method, infrared thermal imaging and visible light bimodal image fusion are combined, and an improved YOLOv8m-seg model is combined, so that accurate detection of the defects of the outer wall of the building is realized. Firstly, a bimodal data set containing multi-material temperature difference data is constructed, ORB feature matching is combined with a feature alignment module (FAM) to achieve accurate image alignment, the defect detection robustness is improved through an improved bimodal feature fusion network, and finally automatic recognition and quantitative analysis of cracks, fractures and other defects are achieved. The problems that single-mode detection is interfered by dirt and is poor in multi-material adaptability are solved, the detection accuracy rate reaches 96% or above, the recall rate exceeds 93%, the single image detection time is shorter than or equal to 0.3 s, and the method is suitable for efficient detection of various building outer wall defects.
Owner:CHANGSHA XINTAI INSTR CO LTD

Single-view large-scale outdoor scene three-dimensional reconstruction method based on three-dimensional Gaussian splashing

The invention discloses a single-view large-scale outdoor scene three-dimensional reconstruction method based on three-dimensional Gaussian splashing, and the method comprises the steps: collecting a pseudo aerial image, and constructing a panoramic multi-mode supervision end-to-end single-view three-dimensional reconstruction model; meanwhile, panoramic consistency supervision, semantic constraint depth regularization and a radial weighted luminosity loss and Gaussian cutting mechanism are introduced, so that the defect of insufficient geometric constraint of traditional single-view three-dimensional reconstruction is effectively overcome, and high-efficiency and high-fidelity three-dimensional modeling of a large-scale outdoor scene under single image input is realized; the method is suitable for various actual scenes such as smart city construction, automatic driving simulation, virtual reality / augmented reality, digital twinning and the like.
Owner:HANGZHOU MAQUAN INFORMATION TECH CO LTD

Three-dimensional model reconstruction and image generation method, device, storage medium, and program product

Embodiments of the present application provide a three-dimensional model reconstruction and image generation method, a device, a storage medium, and a program product. In the method, multi-stage three-dimensional reconstruction is performed on the basis of a single image of a target object; in a first stage, a plurality of view images are generated on the basis of an image generation model, and an initial three-dimensional model is reconstructed on the basis of the plurality of view images; and in a second stage, on the basis of the plurality of view images and an initial prompt containing set marker information, a text-to-image model is used to learn an association relationship between the target object and the set marker information, and a plurality of scene images are generated on this basis. Compared with the plurality of view images, the scene images generated in the second stage have higher resolution and richer image details, and then the initial three-dimensional model is optimized on the basis of the plurality of scene images, so that a target three-dimensional model having higher resolution and clearer model details can be obtained, thereby paving the way for practical application of a three-dimensional reconstruction solution based on a single image.
Owner:TAOBAO CHINA SOFTWARE

Single image super-resolution reconstruction method and system based on wavelet transform and cross-domain feature fusion

The invention discloses a single image super-resolution reconstruction method and system based on wavelet transform and cross-domain feature fusion. According to the method, firstly, a low-resolution RGB image is mapped to a high-dimensional feature space through a shallow feature extraction module; performing up-sampling and discrete wavelet decomposition on the features by using a wavelet feature mixing module to obtain multi-band features; low-frequency and high-frequency depth features are respectively extracted through a double-branch structure, cross-domain fusion is realized by means of a deformable cross attention mechanism, and the feature expression ability is enhanced in combination with residual connection; and finally, reconstructing a high-resolution image through convolution, up-sampling and regularization processing. In the training process, a pixel-level loss function is adopted to optimize network parameters, the multi-frequency-domain feature sensitivity is effectively improved, texture and structure information is balanced, the image contrast, definition and structural integrity are improved, and high-quality real-time super-resolution reconstruction can be achieved.
Owner:HUNAN UNIV

Systems and methods of determining changes in pose of an autonomous vehicle

A vehicle comprises a sensor configured to capture images and one or more processors. The one or more processors can be configured to receive a single image from the sensor, the single image captured by the sensor as the autonomous vehicle was moving; execute a machine learning model using the single image as input to generate a change in pose of the autonomous vehicle, the machine learning model trained to output changes in pose of autonomous vehicles based on blurring in individual images; determine a global position of the autonomous vehicle based on the generated change in pose of the autonomous vehicle; and transmit the global position to an autonomous vehicle controller configured to control the autonomous vehicle.
Owner:TORC ROBOTICS INC

Single-image dressed human body reconstruction method based on posture guide diffusion and view angle consistency constraint

The invention relates to a single-image dressed human body reconstruction method based on posture guide diffusion and view angle consistency constraint, and belongs to the field of computer vision and graphic images. The method comprises the following steps: firstly, carrying out multi-view image generation on a single image in an input dressed human body data set by adopting posture guide diffusion to obtain a two-dimensional diffusion image of an invisible view angle; secondly, establishing a visual angle consistency constraint module, taking the input image and the diffusion image as bidirectional visual angle features, performing implicit geometric guidance in the normal direction of the SMPLX human body template, and respectively extracting three-dimensional space features corresponding to cross visual angles; and finally, fusing voxelization features and three-dimensional space features in the SMPLX human body template to carry out human body implicit reconstruction to obtain a dressed human body model. According to the method, the problems of low reconstruction quality, clothing detail missing and the like caused by insufficient input information of a single image can be effectively solved, and high-quality three-dimensional reconstruction of a dressed human body is realized.
Owner:KUNMING UNIV OF SCI & TECH

Power transmission line inspection system based on autonomous correction long-endurance unmanned airship

The invention relates to the technical field of intelligent inspection of power equipment, and particularly discloses a power transmission line inspection system based on an autonomous correction long-endurance unmanned airship. The system aims to solve the problems of short endurance, small operation radius, single image angle, intermittent operation, leak detection, low precision, poor environmental adaptability and the like caused by dependence on manpower or GPS navigation in the conventional unmanned aerial vehicle inspection. The system comprises an energy supply module, a data acquisition module, an intelligent control module and a communication and data return module. Wherein the energy supply module is integrated with a solar energy collection, storage and intelligent scheduling unit to realize day-and-night continuous energy supply; the intelligent control module adopts an RT-DETR visual algorithm to identify a tower in real time, performs automatic deviation correction by fusing pose information, and accurately tracks a preset track; the data acquisition module supports multi-angle high-definition imaging, and the communication module guarantees real-time data return. And high-precision unmanned inspection with ultra-long endurance, full-automatic and wide-area continuous coverage is integrally realized.
Owner:SICHUAN SHUJU INTELLIGENT MFG TECH CO LTD

Fluorescence lifetime microscopic image large-view-field splicing method and system and electronic equipment

The invention relates to the technical field of image processing, and discloses a fluorescent lifetime microscopic image large-view-field splicing method and system and electronic equipment, and the method comprises the steps: constructing a tissue region mask for a to-be-spliced image, and constructing a vignetting model in a tissue region to complete vignetting correction; counting brightness indexes in an organization area to determine a global brightness reference, and realizing image group brightness unification through global zooming and single image brightness adaptive correction; further realizing alignment of adjacent images through feature point matching, and finishing image block splicing in combination with a minimum color difference suture fusion method to obtain an image band; an overlapping area is extracted from an image belt to generate an effective content mask, relative displacement is estimated by using a phase correlation method after low-frequency suppression, image belt splicing is completed by multiplexing a minimum color difference suture fusion method, a large-view-field spliced image is output, the automation degree and robustness of splicing are improved, the spliced image is geometrically consistent and visually seamless, and the splicing efficiency is improved. And the requirements of medical research on high-resolution and large-field-of-view fluorescence lifetime microscopic images are met.
Owner:SHENZHEN UNIV

Target detection method and system based on dynamic memory enhancement

The invention discloses a target detection method and system based on dynamic memory enhancement, relates to the technical field of computer vision and target detection, and aims to solve the problems that a traditional method depends on local features of a single image, is difficult to detect a target, is low in precision and is weak in generalization ability. The method comprises the steps that after a to-be-detected image is preprocessed, multi-stage multi-scale features are extracted through a backbone network; an encoder outputs single image global features for feature interaction fusion, and a global distribution extraction module constructs and updates a data set level global feature distribution library; the decoder adopts a hierarchical structure, and optimizes the target query token layer by layer in combination with self-attention, cross attention and global distribution fusion sub-modules; and finally outputting a target category and bounding box coordinates. Through global context injection, small target and shielding target detection precision is improved, model training stability and generalization ability are enhanced, and the method is suitable for diversified detection scenes.
Owner:CHONGQING UNIV

Sports equipment management personnel identification method, system and equipment

The invention discloses a sports equipment management personnel identification method, system and equipment, and the method comprises the steps: starting and initializing a camera through detecting a sports equipment use trigger signal, and stopping automatic focusing; and shooting multiple frames of images at different focal lengths according to a preset interval, determining a clear area by using a gradient magnitude algorithm, and performing image registration and fusion to form a composite image. And if the whole area is not covered, shooting and splicing are carried out again. And through composite image splicing, distortion cutting, noise reduction and normalization processing, a final image is output for determining the identity of a person. According to the invention, through multi-frame shooting and accurate image processing, a high-quality image is obtained, and information limitation and quality defects of a single image are effectively avoided. Various algorithms are fused to ensure that the image is clear, complete and standardized, the accuracy and reliability of personnel identity recognition are remarkably improved, an accurate image basis is provided for sports equipment management, the normalization and safety of equipment use are guaranteed, meanwhile, the image collecting and processing efficiency is improved, and the management cost is reduced.
Owner:SHENZHEN ONSAFE TECH DEV

Multi-modal data feature representation optimization method based on comparative learning and two-stage mask

The invention discloses a multi-modal data feature representation optimization method based on comparative learning and two-stage masks. The method comprises the following steps of: firstly, aiming at an input single image-text pair, generating two groups of heterogeneous image-text data with mode missing as two inputs of a model by applying two random mask strategies with different mask areas and proportions; wherein the first group of data applies a high-proportion mask to the image and applies a low-proportion mask to the text; the second group applies a low-scale mask to the image and a high-scale mask to the text. Then, the two groups of data respectively pass through an image encoder and a text encoder which share parameters, and two different multi-modal fusion feature vectors are generated through a cross-modal fusion encoder; according to the method, masked modal information is recovered through a decoder, and the reconstruction loss of the difference between a recovery result and original data is calculated. Meanwhile, two multi-modal fusion feature vectors generated twice are subjected to comparative learning, and the comparative loss of feature representation distances under different enhanced views for approaching the same image-text pair is calculated. And finally, performing weighted summation on the reconstruction loss and the comparison loss to form an overall loss function, and optimizing model parameters through back propagation. When the trained encoder is used for a downstream multi-modal classification task, the classification precision and generalization ability of the model can be effectively improved.
Owner:NORTHWEST UNIV

Peripheral out-of-focus lens assembly and glasses

The utility model discloses a peripheral out-of-focus lens assembly and glasses, and the assembly comprises a basic lens which generates perspective focal power for perspective light; the single image source is arranged at the edge position of the near-eye side of the basic lens and is far away from the basic lens; the plurality of light splitting surfaces are embedded into the basic lens and are distributed in the basic lens at different inclination angles, and the central normal lines of the plurality of light splitting surfaces are inclined to the side where the image source is located relative to the central normal line of the lens; the light irradiation range of a single image source covers all the light splitting surfaces; each light splitting surface and the shared image source form an out-of-focus unit, each light splitting surface is used for reflecting light rays emitted by the image source to form convergent light rays, the convergent light rays are emitted to an exit pupil position, and an out-of-focus image is formed after the convergent light rays enter human eyes; a plurality of out-of-focus images formed by reflection of the plurality of light splitting surfaces are respectively distributed in a peripheral view field. The peripheral out-of-focus lens assembly and the glasses are used for realizing out-of-focus stimulation, inhibiting eye axis growth and achieving the purpose of preventing and controlling myopia.
Owner:BEIJING NEDPLUSAR DISPLAY TECH CO LTD

Image watermark processing method and device, electronic equipment and storage medium

The invention discloses an image watermark processing method and device, electronic equipment and a storage medium, and relates to the field of financial science and technology. The method comprises the following steps: in response to an image sending request, generating a first chaotic sequence according to a sender identity, a sending timestamp and a chaotic sequence control parameter; segmenting the to-be-sent image into a plurality of image blocks, and determining basic watermark data according to channel values of pixels in the image blocks, a pre-specified private key, a sender identity identifier and a sending timestamp; determining the watermark length of a single image block according to the total length of the basic watermark data and the number of the image blocks, and distributing block-level watermark data for the image blocks from the basic watermark data according to the watermark length; according to the first chaos sequence, determining a first starting coordinate in the image block and a first channel in which a watermark needs to be embedded; and embedding a watermark into the image block according to the first initial coordinate, the first channel and the block-level watermark data. According to the technical scheme, the security and reliability of the image watermark are improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

System and methods for gait analysis and longitudinal health and aging assessments including musculoskeletal disorders using video-trained spatio-temporal graph neural networks

A method for pose and gait classification and motion prediction using spatio-temporal relationships between body joints includes capturing a sequence of images or video frames of a subject; applying a neural network-based pose estimation algorithm to the sequence of images or video frames to detect landmark positions of anatomical joints; constructing a spatio-temporal graph from the detected landmark positions of the one or more anatomical joints, wherein nodes of the spatio-temporal graph correspond to the anatomical joints and the landmark positions, spatial edges of the spatio-temporal graph represent anatomical connections between the anatomical joints within a single image or frame, and temporal edges connect the one or more anatomical joints across successive images or frames of the sequence of images or video frames; and inputting the constructed spatio-temporal graph into a spatio-temporal graph convolutional network (ST-GCN) to classify pose and gait patterns and predict motion or stability states.
Owner:IMAGINE DESIGN LLC

Apparatus and method for image conversion

An image conversion apparatus according to one embodiment includes: a memory that stores an image conversion program to compress a plurality of images into a single image or decompress the compressed single image into the plurality of images; and a processor that executes the image conversion program, and the image conversion program inputs the plurality of images into an encoder model and outputs the compressed single image in which the remaining images are inserted into one of the plurality of images, and the encoder model is machine-learned to compress a plurality of initially input images into a single image by hierarchically compressing the plurality of images into one according to a tree structure, ensuring the final compressed image is identical to one of the initially input images.
Owner:SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION

Single image defogging method based on block-by-block nonlinear brightness prior

The invention discloses a single image defogging method based on block-by-block nonlinear brightness prior, which belongs to the technical field of image processing, and comprises the following steps of: dividing a fog-containing image into local blocks, calculating the average brightness of the blocks, and constructing prior block-by-block monotone increasing nonlinear mapping to represent the corresponding relationship between the brightness of fog-containing blocks and the brightness of clear blocks; an atmospheric scattering model is combined, an atmospheric light vector is modeled into a vector, the vector and a PPWF form a parameterized recovery model, three scalar parameters are taken as a core, an optimal parameter is obtained through multi-target joint optimization, alternate optimization and golden section search, and finally a defogged image is generated. The method has the advantages of few parameters, low complexity and no need of training data, the definition and global contrast of far and near scenery can be remarkably improved while the image structure is maintained, and halo, excessive enhancement and color cast are effectively inhibited; the method has good robustness for different fog densities and illumination conditions, and is suitable for real-time and embedded defogging application of single-channel or multi-channel images.
Owner:NANJING UNIV OF POSTS & TELECOMM

Automatic medicine identification system based on computer vision

The invention discloses an automatic medicine identification system based on computer vision, and relates to the technical field of computer technology and artificial intelligence, and the system comprises a feature extraction module which is used for extracting an intrinsic identity vector and a degradation state vector from an obtained digital image of a to-be-identified medicine; the identity recognition module is used for generating an identity recognition result of the medicine according to the intrinsic identity vector and a preset classifier network; the quality evaluation module is used for generating a quality quantitative score of the medicine according to the degradation state vector and a preset degradation tolerance threshold value; and the risk discrimination module is used for generating a shoddy drug risk discrimination clue according to the degradation state vector, a preset certified product degradation mode library and a preset mode difference threshold. According to the invention, through a unique confrontation training mechanism, the system can respectively extract the characteristics representing the inherent identity of the medicine and the characteristics representing the appearance degradation state of the medicine from a single image, and the accuracy and robustness of recognition are greatly improved.
Owner:SUZHOU MUNICIPAL HOSPITAL

Text-to-three-dimensional object generation method based on dual color consistency matching

The invention discloses a text-to-three-dimensional object (3D) generation method based on dual color consistency matching, and aims to solve the problem of 3D asset quality reduction caused by multi-view color inconsistency in the prior art. According to the method, in each training iteration, a group of multi-view images are firstly rendered from a 3D model, and a high-fidelity forward view is specified as a color reference; subsequently, optimization is carried out through a double matching mechanism: 1, matching among images: aligning colors of other views to the reference view by using depth semantic features, thereby breaking circulation of error accumulation; 2, image internal matching: in a single image, color and geometric smoothness in the same semantic region are standardized by improving total variation loss; through the dual mechanism, accumulation and propagation of unnatural color distortion are effectively prevented, color inconsistency is reduced, and the visual fidelity and color consistency of the generated 3D model are remarkably improved.
Owner:SOUTHEAST UNIV

Medical image fusion method based on visual state space and gated attention mechanism

PendingCN121502717AImage enhancementImage analysisCross modalityFeature extraction
The invention relates to the technical field of multi-modal medical image fusion, provides a medical image fusion method based on a visual state space and a gated attention mechanism, and effectively captures cross-modal local details and a global dependency relationship through collaborative combination of a convolutional neural network and a state space model. Specifically, a multi-scale feature extraction module is introduced to extract hierarchical features of different scales of each mode, and the module further introduces an optimized channel-space attention module to enhance cross-scale feature expression and inter-channel interaction. In order to model long-distance spatial dependence and enhance cross-modal feature interaction, a multi-scale visual state space module is developed based on SSM, and the module can realize global context aggregation while keeping fine-grained structure details. And finally, integrating the information of the two modes together through a specific fusion strategy, and generating a single image with rich details.
Owner:CHANGCHUN UNIV

Bacterial cluster motion classification method and system based on single image and deep learning

The invention discloses a bacterial cluster motion single image detection method and system based on deep learning, and the method and system are used for quickly and automatically distinguishing the states of cluster motion, swimming and the like of bacteria. According to the method, high-precision classification is realized by acquiring a long-exposure single blurred image of bacteria in a circular limited space and extracting spatial-temporal characteristics by using a dense connected neural network (DenseNet) with an attention module. The kit is suitable for a high-throughput environment, can be integrated into portable equipment, and is used for early diagnosis and treatment evaluation of diseases such as urinary system infection (UTI) and inflammatory bowel disease (IBD).
Owner:KEYIN (SHANGHAI) MEDICAL TECHNOLOGY CO LTD

A glass detection method based on deep learning and ghost phenomenon

The application relates to a glass detection method based on deep learning and ghosting, and relates to the technical field of computer vision. The method comprises the following steps: performing glass detection based on a single original input image, the image being a single RGB image; extracting ghost features through a deep learning method of a backbone network based on the original input image to obtain a ghost area prediction map; connecting the ghost area prediction map with the original input image channel, extracting glass features based on the backbone network under the guidance of ghost clues, then performing glass feature decoding and glass area segmentation based on a convolutional neural network; and outputting a glass area prediction map. Compared with the prior art, the application performs glass detection based on a single image, and is more widely applicable. The backbone network can more accurately and efficiently extract ghost features and glass area features, and the ghosting phenomenon can be used to more accurately locate the glass area, obtain a high-quality glass area prediction map, and has good robustness.
Owner:JIANGNAN UNIV

Visual precise positioning method based on pixel fusion alignment

The invention discloses a visual precise positioning method based on pixel fusion alignment, and relates to the technical field of visual positioning, and the method comprises the steps: obtaining a target image shot by a to-be-positioned terminal and rough pose data, screening a candidate grid set in a pixel fusion grid template dictionary based on the rough pose, and obtaining a candidate grid set; performing multi-channel pixel fusion alignment on the target image in the candidate grid set, performing step-by-step detailed search from top to bottom according to grid levels, and calculating a target three-dimensional coordinate in a minimum grid according to a pixel fusion alignment result; multi-source three-dimensional scene modeling, multi-layer grid organization, multi-channel pixel-level alignment and an online self-learning updating mechanism are comprehensively utilized, and the problem that in the prior art, in the indoor and outdoor integrated complex environment where satellite signals are weak and illumination and scenes change frequently, the multi-source three-dimensional scene modeling, multi-layer grid organization, multi-channel pixel-level alignment and online self-learning updating are difficult to achieve is solved. The method solves the problem that high-precision and high-robustness three-dimensional visual positioning is realized only by means of a single image of the terminal, and is particularly suitable for navigation guidance and intelligent equipment positioning application of large comprehensive transportation hubs, airports and urban complexes.
Owner:BEIJING SILICON INTELLIGENCE TECH CO LTD

Glare mitigation techniques in symbologies

Methods, systems, and apparatus, including medium-encoded computer program products, for glare mitigation techniques include: obtaining images containing a representation of a mark, the images comprising multiple poses of the mark and generating a single image from the images that contain a representation of the mark with reduced glare when compared to the images comprising multiple poses of the mark. The single image is provided for processing of the representation of the mark to identify information associated with the mark.
Owner:SYS TECH SOLUTIONS INC

Camera condition guide viewpoint synthesis method based on video diffusion model

The invention relates to a video diffusion model-based camera condition-guided viewpoint synthesis method, which belongs to the technical field of computer vision, and comprises the following steps of: taking a stable video diffusion model SVD as a basic generator, and jointly inputting Gaussian noise and encoded image potential representation to obtain a video diffusion model SVD-based viewpoint synthesis model; meanwhile, a pose encoder based on a time attention mechanism is used for encoding camera parameters represented by Plcarbon coordinates, and then the camera parameters are used as pose conditions to be embedded into a time attention layer in a video diffusion model de-noising U-Net; pixel-level features of an input single image are extracted through an image coding embedding network and serve as image feature conditions to be embedded into a space attention layer in a de-noised U-Net of a video diffusion model, and a pre-trained video diffusion model is finely adjusted; the space consistency of the synthetic image and the input image and the track consistency of the synthetic image and the camera parameters are improved, and the generation diversity is increased while the generation quality is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method, device, storage medium and program product for stylized image guided three-dimensional model texture generation

The application discloses a style image guided three-dimensional model texture generation method, a computer device, a storage medium and a program product. The style image guided three-dimensional model texture generation method is a pipeline based on a diffusion model, which generates a style texture under the guidance of a single image, extracts style information from a reference image, and ignores content information. First, a mesh model of a three-dimensional object and a style image used for guidance are obtained, text embedding features and style features are obtained, and style injection is completed in a diffusion model cross-attention mechanism in layers. The texture parameters of a texture representation model are optimized by using interval score matching loss and style guidance loss, and finally, the style texture of the high-quality three-dimensional object is obtained by sampling the texture representation model.
Owner:ZHEJIANG UNIV

Nerve radiation field supervision-based single repositioning method for large-view-field light-changing scene

The invention discloses a single relocation method for a large-view-field light-changing scene based on nerve radiation field supervision. The method comprises the following steps: acquiring a large-view-field multi-view-angle image under the condition of luminosity change by using a camera; estimating the initial pose of the camera by using an SFM algorithm; taking the original image and the camera pose together as input to train a neural radiation field model of a large-view-field light-changing scene; training a pose estimation model by using the original image in combination with a CNN, and inputting the estimated pose into a neural radiation field model to query a corresponding rendered image; constraining the difference between the rendered image and the original image, and optimizing the pose estimation model; and obtaining a relative pose by using the pose estimation model, and obtaining an absolute pose in combination with an absolute pose obtaining strategy and a coordinate system conversion scheme. After training is completed, high-precision positioning can be achieved only through a single image, the reasoning speed is high, and the storage space needed by the model is small.
Owner:NANJING UNIV OF SCI & TECH

Optical fiber sound wave event identification method, equipment and medium

The invention discloses an optical fiber sound wave event identification method and device and a medium, and relates to the field of sound wave signal processing, and the method comprises the steps: obtaining a time sequence signal of an optical fiber sound wave; performing wavelet denoising processing on the time sequence signal; converting the denoised time sequence signal into a multi-modal time sequence image by adopting a multi-modal feature conversion method; performing fusion processing on the multi-modal time sequence image to obtain an RGB fusion image; a Transform-CNN (Convolutional Neural Network) hybrid network is constructed; and inputting the RGB fusion image into a Transform-CNN hybrid network to obtain an event identification result of the optical fiber sound wave. The method can significantly enhance the discrimination capability of the neural network model for the complex signal form, solves the problems that the image single-mode modeling capability is limited, a single image representation mode cannot comprehensively express the multi-dimensional features of the event, and improves the recognition precision.
Owner:JIMEI UNIV

Camera self-calibration method based on single picture and target

The invention relates to a camera self-calibration method based on a single picture and a target, and the method comprises the steps: determining an optimal homography matrix through an RANSAC algorithm and a radius distribution method, effectively eliminating the influence of low-quality feature points on a calibration result, and improving the calibration precision, robustness and stability; then, camera parameters are further optimized by adopting a maximum likelihood method and a free scale factor algorithm, a remapping matrix is calculated, and the camera calibration precision is improved; and finally, correcting image coordinates through the remapping matrix, recovering the image coordinates to world coordinates by using the optimal homography matrix, and checking a calibration effect. According to the method, manual intervention is not needed, automatic calibration of the target area image is achieved, reliable high-precision calibration images and stable tracking points are provided for structural health monitoring, test errors caused by image distortion are effectively reduced, and an important foundation is laid for application of a machine vision algorithm in the civil engineering field.
Owner:TIANJIN UNIV +1

Dynamic four-dimensional content generation method and device based on single image, equipment and medium

The invention relates to the technical field of computer vision, and discloses a dynamic four-dimensional content generation method and device based on a single image, equipment and a medium, and the method comprises the steps: generating a corresponding multi-view image set based on an input single static image; constructing a static three-dimensional scene representation model based on the multi-view image set; converting the static three-dimensional scene representation model into a dynamic four-dimensional scene representation model; performing time consistency optimization processing on the dynamic four-dimensional scene representation model to generate a dynamic frame sequence coherent in time dimension; and background illumination controllable editing processing is carried out on the dynamic frame sequence after time consistency optimization processing, and final editable dynamic four-dimensional content is generated. According to the method and the device, the three-dimensional model is constructed by utilizing the multi-view image set and is further converted into the dynamic four-dimensional model, so that the problems of insufficient multi-view image continuity and unstable dynamic details in time dimension in the prior art are solved, and the reality sense of dynamic contents and the user experience are improved.
Owner:GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)