Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

497 results about "Computer vision and image processing" patented technology

High-precision image processing method and system based on illumination adaptive compensation

The invention discloses a high-precision image processing method and system based on illumination adaptive compensation, and relates to the technical field of computer vision and image processing, and the method comprises the steps: inputting an original image, and dividing the image into a high-frequency edge layer, an intermediate-frequency texture layer and a low-frequency illumination layer through a multi-scale residual network; acquiring illumination intensity, color temperature and scene categories in real time by using an ambient light sensor and a scene semantic segmentation model, and generating dynamic compensation parameters; carrying out dynamic range expansion on a low-frequency illumination layer based on a physical illumination model, and adjusting the weight of highlight suppression and dark area enhancement through a self-adaptive S-shaped exposure curve; a double-branch generative adversarial network is adopted, noise suppression and super-resolution reconstruction are carried out on the high-frequency layer, and texture detail enhancement is carried out on the intermediate-frequency layer; aligning the data of the depth camera and the infrared sensor with the visible light image through a cross-modal fusion module; and performing tone mapping on the fused image based on human visual characteristics, and outputting an enhanced image with a high dynamic range and reserved details.
Owner:SHANXI UNIV

Unmanned aerial vehicle multi-modal feature fusion target tracking method and system based on natural language description

The invention discloses an unmanned aerial vehicle multi-modal feature fusion target tracking method and system based on natural language description, belongs to the technical field of computer vision and image processing, and solves the problem that in the prior art, when the quality of an image collected by an unmanned aerial vehicle is poor or image features are not obvious, the target tracking capability and the long-time tracking capability are poor. Natural language description is carried out on a traffic accident scene in an image of an unmanned aerial vehicle visual angle, and a language prompt is obtained; constructing a scene-context feature pyramid network to perform context information enhancement processing on the image of the view angle of the unmanned aerial vehicle to obtain a feature-enhanced image; respectively carrying out visual coding and language coding on the enhanced image and language prompt to obtain visual features and language feature vectors, and carrying out visual-language bimodal feature local alignment; and fully fusing the obtained aligned new language features with the visual features to obtain multi-modal features for target tracking. The method is used for multi-modal feature fusion target tracking of the unmanned aerial vehicle.
Owner:YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA

Blind guiding scene identification method based on multi-modal visual large model

The invention relates to the field of computer vision and the field of image processing, in particular to a blind guiding scene recognition method based on a multi-modal visual large model. Comprising the following steps: data acquisition: acquiring image or video data of a typical blind guiding scene, marking, constructing a data set, introducing an image prompting mechanism, and dividing a training set and a test set; constructing a multi-modal visual large model which comprises an image coding module, a semantic prompt coding module, a visual language fusion module and a scene analysis semantic decoding module; the model is tested and optimized, and a multi-task loss function is adopted for optimization; a blind guiding auxiliary system is constructed, image acquisition, visual understanding and voice broadcasting functions are integrated, and a closed-loop process is realized. According to the method, the problems of strong target dependence, poor universality and weak semantic comprehension ability of the existing blind guiding identification technology are solved, accurate identification and semantic feedback of complex scenes are supported, and the intelligent level and the environment adaptability of a blind guiding system are improved.
Owner:JIANGSU INST OF ECONOMIC & TRADE TECH +1

Panoramic image three-dimensional reconstruction method based on 3DGS

The invention provides a panoramic image three-dimensional reconstruction method based on 3DGS, and relates to the technical field of computer vision and image processing, and the method comprises the steps: 1, collecting a panoramic image sequence of a to-be-reconstructed scene, carrying out the unified coding through equidistant cylindrical projection, and obtaining a coding data set; and step 2, executing panoramic motion recovery structure processing on the coded data set, and generating six-degree-of-freedom camera pose parameters of each image and a corresponding sparse three-dimensional point cloud. According to the invention, high-precision and high-efficiency panoramic scene three-dimensional reconstruction is realized, a high-quality panoramic image or a dense three-dimensional model can be output, and the three-dimensional reconstruction quality and efficiency are improved.
Owner:LANJIAN (SUZHOU) TECH CO LTD

White vehicle body welding seam recognition and automatic welding method based on machine vision technology

The invention discloses a body-in-white welding seam recognition and automatic welding method based on a machine vision technology, particularly relates to the technical field of computer vision and image processing, and is used for solving the problem of welding seam track recognition accuracy caused by insufficient processing capability of an existing three-dimensional vision recognition method on incomplete and uncertain point cloud data. Through the steps of multi-view point cloud acquisition and registration, probabilistic confidence evaluation, region growth of track continuity constraint, multi-track fusion optimization and the like, accurate identification of a body-in-white welding seam track under a complex working condition is realized; firstly, multi-view point cloud data are obtained, probabilistic registration is carried out to generate a confidence evaluation result, then candidate tracks are generated based on confidence weighting and semantic constraint, finally, an optimal track is generated through intelligent optimization and converted into a welding instruction which can be executed by a robot, and the accuracy and robustness of weld joint recognition are effectively improved.
Owner:CHONGQING MULSTRONG INTELLIGENT TECH CO LTD

Image feature matching optimization method based on intra-class space consistency

The invention discloses an image feature matching optimization method based on intra-class space consistency in the technical field of computer vision and image processing. The method comprises the following steps: feature point extraction and preliminary matching: extracting feature points from a query image and a reference image and performing preliminary matching; initialization and transformation model estimation: initializing a matching point set and a residual error, and calculating an initial transformation model; error calculation and matching point set updating: calculating the error of the matching point pair, and updating the matching point set by adopting a dynamic screening method; performing residual optimization: judging whether the optimal condition is reached or not based on the residual, and deciding whether to continue iteration or not; performing intra-class space consistency clustering and isolated cluster elimination: performing clustering analysis after the optimal residual error is obtained, and eliminating isolated clusters based on an intra-class space consistency separation ion structure; and outputting a result: outputting a matching point set after the isolated clusters are removed. The method solves the problem that a traditional feature matching method is difficult to completely remove mismatching in a complex scene and is sensitive to noise.
Owner:CHANGCHUN UNIV OF SCI & TECH

Morphological gradient region replacement method based on SAM semantic segmentation and user guidance

The invention discloses a morphological gradient region replacement method based on SAM semantic segmentation and user guidance, and relates to the technical field of computer vision and image processing, and the method comprises the steps: carrying out the semantic segmentation of a to-be-processed image through an SAM model, extracting a multi-level semantic feature, carrying out the standardization and dimension reduction, extracting a causal factor based on independent component analysis, and carrying out the user guidance. A directed causal factor association graph is generated through Granger causal relationship test, and a causal attribution probability graph is generated through reverse mapping; constructing a structured causal graph, and generating a causal mask through a graph convolutional network; encoding the original interaction signal into a guide thermodynamic diagram; constructing a diffusion equation, forming a gradual change control equation by dynamically fusing and guiding the intensity distribution of the thermodynamic diagram and an image semantic diffusion item, and iteratively solving the gradual change control equation; generating an anisotropic morphological operation kernel according to the geometric curvature characteristics of each region in the replacement mask; and fusing the optimized replacement mask with the target content based on a gradient domain optimization algorithm to generate a gradient replacement image.
Owner:BEIJING YIBAIYISHIYI MEDICINE SCI & TECH CO LTD

Road video event rapid detection method based on edge calculation

The invention discloses a road video event rapid detection method based on edge calculation, particularly relates to the field of computer vision and image processing, and is used for solving the problem of motion evidence distortion caused by frame-level sequential disorder and content gaps in a road monitoring video. The method comprises the following steps of: reading a video clip, constructing a time sequence correction graph according to inter-frame content similarity and boundary continuity, rearranging to generate a corrected sequence, outputting a frame consistency index, estimating a stable motion field and a continuity score on the corrected sequence, and writing back an image stabilization amplitude correction time anchor point; generating an event evidence sequence according to the stable motion field, combining track breakpoints according to a time sequence correction graph, initiating backtracking revision to return candidate event segments, and if consistency verification is executed on candidate events, performing cross-level revision on the time sequence correction graph and image stabilization amplitude to trigger the stable motion field to re-estimate and output a confirmation event; and issuing a trigger result and establishing a mapping index as a new fragment to initialize priori acceleration time sequence correction and judgment convergence.
Owner:SHANDONG HUAREN INFORMATION TECH CO LTD

Fine-grained video behavior recognition method based on double-text prompt

The invention belongs to the field of computer vision and image processing, and relates to fine-grained action classification of a picture sequence after video framing by adopting a deep convolutional neural network, in particular to a fine-grained video behavior recognition method based on double-text prompt. The method is simple in program and easy to implement, actions capable of recognizing the fine grit of the human body can be obtained, text description can be divided into different fine grits through a large language model for the fine grit actions of the human body, and then the generated text feature vectors and video features of different time scales are subjected to cross attention mechanism response, so that the recognition accuracy of the human body is improved. Unique details of human body motion in a video can be better found, so that fine-grained motion is more accurately reasoned.
Owner:DALIAN UNIV OF TECH

Remote sensing small sample target detection method based on double-attention guided transfer learning

The invention belongs to the technical field of computer vision and image processing, and discloses a remote sensing small sample target detection method based on double-attention guided transfer learning, and the method comprises the steps: obtaining a remote sensing image data set, and carrying out the preprocessing; taking the preprocessed remote sensing image training set as input, constructing a basic detection model by using a ResNet-101 backbone network, a feature pyramid network and a content awareness upsampling and regional proposal network, and obtaining basic model parameters; basic model parameters are used as input, a DA-FSDET network is trained based on a content awareness strip pyramid and a deformable attention area proposal network, and the trained DA-FSDET network is used to acquire a category detection frame containing small sample categories and confidence. Through cascading and cooperative work of the content awareness stripe pyramid and the deformable attention area proposal network, the detection precision and robustness of the multi-scale target in the remote sensing image are effectively improved.
Owner:ZHONGYUAN ENGINEERING COLLEGE

Infrared polarization image super-resolution method based on cross attention double-branch network

The invention relates to the technical field of computer vision and image processing, in particular to an infrared polarization image super-resolution method based on a cross attention double-branch network, and the method comprises the steps: obtaining infrared polarization images with two different resolutions in the same scene through two infrared polarization cameras with different image resolutions; constructing a double-branch super-resolution network model, training the double-branch super-resolution network model according to the multiple pairs of image data sets until a robust fitting state is reached, and obtaining the trained double-branch super-resolution network model; and obtaining a high-resolution infrared polarization image from the low-resolution infrared polarization image to be subjected to super-resolution through the double-branch super-resolution network model. According to the method, the infrared polarization intensity image and the infrared polarization degree image are subjected to super-resolution at the same time through the deep learning model, the features are fused mutually, and high-quality super-resolution image reconstruction is achieved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Remote controller liquid crystal screen image detection method based on computer vision

The invention relates to the field of computer vision and image processing, and discloses a remote controller liquid crystal screen image detection method based on computer vision, which comprises the following steps: S1, acquiring an original image containing a remote controller, and carrying out gray level conversion and filtering denoising processing on the original image to obtain a preprocessed image; s2, calculating the local information entropy and the change rate of the image through a sliding window under multiple scales based on the preprocessed image, generating a multi-scale entropy gradient heat map, and screening out a candidate region with high entropy difference as a liquid crystal display region according to a set threshold value; and S3, carrying out vectorization partitioning on the candidate region, and constructing an image matrix. According to the invention, by introducing an image preprocessing mode of low-rank sparse decomposition and normalization processing, the expression ability of structural information in the character image is effectively enhanced, and the feature extraction accuracy under the conditions of complex background, low character contrast and the like is improved.
Owner:BEIJING HTDISPLAY ELECTRONICS CO LTD

Image rain removal method based on multi-scale state space model

The embodiment of the invention provides an image rain removal method based on a multi-scale state space model, and belongs to the field of computer vision and image processing. The method comprises the following steps: acquiring an image with rain stripes, and decomposing the image into a multi-scale pyramid image through down-sampling; respectively inputting the image with the rain stripes and the multi-scale pyramid image into a pre-constructed multi-scale state space model, and performing feature extraction and image reconstruction through a convolutional layer and an encoder-decoder network to obtain a reconstructed residual image; and adding the image with the rain stripes to obtain a target rain-removed image. According to the method, global information is modeled by using a state space model, multi-scale complementary information is effectively utilized in combination with a multi-scale framework, cross-scale complementary features are explicitly mined, rain stripes are more accurately removed, a high-quality clear image is gradually recovered from a rain image, and the definition and visual quality of the rain-removed image are ensured.
Owner:NAVAL AVIATION UNIV

Ancient mural restoration method based on progressive reconstruction and damage perception self-adaption

The invention discloses an ancient fresco restoration method based on progressive reconstruction and damage perception self-adaption, and relates to the technical field of computer vision and image processing, a damaged fresco image is input into a coarse restoration network for initial restoration, in the initial restoration process, a discriminator and the coarse restoration network are adopted for adversarial training, and a damaged fresco image is obtained; generating an initial restoration result of the mural image; inputting the initial repair result and the mask image into a mask-guided local information extraction network to obtain a local optimization result; and inputting the local optimization result into a global information extraction network to obtain an overall restoration result of the mural image. The problem that an existing network is low in efficiency in the aspect of capturing local details and global styles is solved. The method has good applicability to damaged wall paintings; the problem of fuzzy texture in the repairing result is effectively solved; local features are adaptively extracted and fused according to the damage degree of the mural, and the multi-stage residual information distillation module further refines details of the mural on different scales.
Owner:NORTHWEST UNIV

Hyperspectral image classification method and system based on double-branch lightweight algorithm

The invention discloses a hyperspectral image classification method and system based on a double-branch lightweight algorithm, belongs to the technical field of computer vision and image processing, and solves the problems that an existing image classification model is complex in structure, so that the calculated amount is large, the occupied memory is high, and balance between classification precision and model lightweight is difficult to achieve. The method comprises the steps of collecting a hyperspectral image, preprocessing the hyperspectral image, performing iterative training on a double-branch lightweight algorithm model based on a training set, executing the double-branch lightweight algorithm model, and classifying a test set by the double-branch lightweight algorithm model. According to the hyperspectral image classification method, the algorithm complexity is greatly reduced, the high precision of hyperspectral image classification is ensured, the lightweight space-spectrum Transformer is introduced into the double-branch lightweight algorithm model, and through the optimization design of the whole structure and the lightweight efficient Attention-Aware mechanism, the classification accuracy of the hyperspectral image is improved. The complexity of the model is reduced, and the characteristic relation between the modeling space and the spectrum sequence can be better established.
Owner:HENAN VOCATIONAL & TECHN COLLEGE OF COMM +1

Low-light image enhancement method based on conditional diffusion model and attention mechanism

The invention relates to the technical field of computer vision and image processing, and discloses a low-light image enhancement method based on a conditional diffusion model and an attention mechanism. The method comprises the following steps: converting an input low-light image and a normal-light image from an RGB color space into an HVI color space, and decomposing the HVI color space; obtaining a brightness component and a horizontal and vertical component; performing iterative brightness recovery on the obtained brightness component through a conditional diffusion model to obtain an enhanced brightness component; performing color preservation and local and global detail enhancement on the obtained horizontal and vertical components through a residual attention module to obtain enhanced horizontal and vertical components; and synthesizing the obtained enhanced components into an HVI image, converting the HVI image back to an RGB image, and finally obtaining an enhanced image. According to the method, the problem that an enhancement algorithm in the prior art is not easy to obtain a local detail which is fully reserved, the color is kept undistorted, and the generalization ability is improved to adapt to images under different conditions is solved.
Owner:ANHUI UNIV OF SCI & TECH

Underwater image enhancement method based on relation-driven dynamic state propagation

The invention discloses an underwater image enhancement method based on relation-driven state space modeling, belongs to the technical field of computer vision and image processing, and aims to solve the problems of color distortion, detail blurring and the like of an underwater image caused by water attenuation and scattering. Carrying out structure perception enhancement modeling; and image reconstruction and decoding output. Wherein the structure sensing module extracts spatial continuity information through an offset generation network, adaptively rearranges scanning paths, preferentially focuses on a semantic rich region and executes dynamic state propagation, so that the accuracy and interpretability of global modeling are improved; and meanwhile, a local convolution kernel is dynamically generated according to global channel statistical characteristics by inputting a dependent convolution branch, so that the adaptability to a background region is enhanced. In order to further improve the feature fusion effect, a cross-feature fusion bridge module is provided, multi-level features are guided and fused through bidirectional attention of a structural path and a semantic path, and detail information and context semantics are effectively integrated.
Owner:HARBIN INST OF TECH

RGBT target tracking network and method fusing multi-interaction feature enhancement mechanism

The invention discloses an RGBT target tracking network and method fused with a multi-interaction feature enhancement mechanism, relates to the technical field of computer vision and image processing, and aims to solve the problems that multi-modal fusion is insufficient and tracking is easy to drift due to fixed or blind updating of a template. The network adopts an end-to-end tracking framework, a backbone network of the network extracts visible light and infrared template images and searches feature tokens of the images through a convolution token embedding module, and performs intra-modal and inter-modal mixed attention interaction by using a multi-interaction Transform module to realize multi-level feature fusion. A target frame is output by adopting an angular point prediction head, a template updating module is introduced, and a template token is dynamically evaluated and updated through two Transform modules, so that the long-term tracking stability is improved. The network effectively deals with complex environment changes by fusing multi-modal information and a dynamic updating mechanism, and the tracking accuracy and robustness are remarkably improved.
Owner:HEFEI NORMAL UNIV

X corner detection method based on multilevel matching and feature optimization

The invention relates to an X corner detection method based on multilevel matching and feature optimization, and belongs to the technical field of computer vision and image processing. The method comprises the following steps: making a high-precision template corresponding to a predetermined X corner mark, so that the high-precision template can be accurately focused on a corner area; carrying out noise reduction and filtering operations on the image before identifying the mark so as to remove noise interference; carrying out template matching on the preprocessed image by adopting a self-adaptive threshold value and a multi-scale search strategy; and performing multi-level detection on the basis of template matching, including ORB feature extraction, Harris corner optimization and sub-pixel level positioning, and outputting the accurate position of the corner and a feature descriptor. The method can still maintain high precision and stability in a complex scene, is suitable for the fields of visual positioning, three-dimensional reconstruction and the like, and has the advantages of high detection efficiency, strong anti-interference capability, wide applicability and the like.
Owner:FUZHOU UNIV

Intelligent identification system for puffed corn

The invention discloses an intelligent identification system for puffed corn, and relates to the technical field of computer vision and image processing, and the system comprises an image collection module which integrates multispectral imaging and a near-infrared sensor and is used for obtaining surface morphology, internal structure and moisture content data of puffed corn; an environment light intensity sensor and an LED array are arranged in the self-adaptive optical compensation module, the wavelength and illumination intensity of a light source are adjusted in a closed-loop feedback mode, and color feature deviation caused by workshop environment light changes is compensated; the image processing module is used for deploying a lightweight hybrid neural network model and is used for adhesive particle segmentation and defect detection; the real-time sorting control module is used for driving a three-axis mechanical arm and a pneumatic spray valve according to an identification result so as to realize defective product rejection and grade subpackaging; and the digital twinning optimization platform is used for outputting a process adjustment suggestion to the production line PLC. According to the invention, through the holographic sensing network, the hybrid intelligent decision and the cross-domain cooperative control, the problems of inaccurate reading, unquick judgment and inaccurate control in puffed corn identification are solved.
Owner:JIANGSU CHENWEI BIOLOGICAL TECH CO LTD

Vegetation image recognition method, system and device based on OfficientNet and medium

The invention discloses a vegetation image recognition method, system, equipment and medium based on OfficientNet, and belongs to the technical field of computer vision and image processing, and the method comprises the steps: collecting data through image collection and synchronous utilization of a laser radar, and carrying out the preprocessing of the collected data; an improved OfficientNet network model is constructed, and training and optimization are carried out by using the improved OfficientNet network model; and carrying out vegetation classification and hidden danger detection on the input image, inputting the collected real-time data into a comprehensive risk scoring model for hidden danger detection for calculation, and carrying out risk early warning positioning. The hidden danger trees are accurately positioned based on the GIS technology, the omission ratio is reduced through a dynamic early warning system, and the manual inspection risk is reduced through automatic reporting; different climate monitoring requirements are adapted based on the expansibility of the model, the tree barrier hidden danger can be blocked and the maintenance efficiency can be improved in practical application, and an integrated solution integrating accurate identification and real-time positioning is formed.
Owner:GUIZHOU POWER GRID CO LTD

Method and System for Industrial Defect Detection and Location Based on Frequency Domain Enhancement

The present invention discloses a method and system for industrial defect detection and localization based on frequency domain enhancement, which relates to the fields of computer vision and image processing, can discover abnormal patterns in industrial images from multiple perspectives, and also has a relatively fast inference speed. The method includes: processing an input image through a high-pass filter to obtain a high-frequency image; inputting the original image and the high-frequency image into an improved U-Net network integrated with a dual-domain feature selection module; performing image reconstruction simultaneously in the spatial domain and the frequency domain; calculating an anomaly score and a localization map based on the reconstruction results. By combining spatial domain and frequency domain analysis, the present invention realizes a comprehensive characterization of local pixel changes and global frequency patterns in the image, improving the accuracy of anomaly detection and localization. The method has achieved optimal performance on multiple real industrial scenario datasets. The present invention is applicable to anomaly detection and anomaly area localization of industrial product images.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Remote sensing image change detection method

The invention discloses a remote sensing image change detection method, and relates to the technical field of computer vision and image processing, and the method comprises the steps: inputting a preprocessed dual-time image of a target region into a remote sensing image change detection model, and outputting a change probability graph of the target region; the remote sensing image change detection model is obtained by training a global network model by adopting a training set, and the global network model comprises an initial network model and a classifier; the initial network model is constructed on the basis of a transform model; performing thresholding processing on the change probability graph of the target area to obtain a binary change detection graph of the target area; the initial network model comprises a pyramid segmentation attention feature enhancement module, a cross-spatio-temporal feature interaction module, a cross-scale attention transformer module and a cross-level pyramid transformer module. According to the invention, the accuracy and reliability of remote sensing image change detection are improved.
Owner:NANKAI UNIV

Construction scene dynamic obstacle avoidance method and system based on multi-source image fusion

The invention discloses a construction scene dynamic obstacle avoidance method and system based on multi-source image fusion, and belongs to the technical field of computer vision and image processing, and the method comprises the steps: mapping multi-source image data to a symmetric positive definite matrix manifold space, carrying out the high-precision registration based on Riemannian geometric measurement, achieving the self-adaptive feature fusion through geometric flow optimization, and achieving the dynamic obstacle avoidance of a construction scene. According to the method, spatial topological features of obstacles are extracted through topological data analysis, and probability trajectory prediction is carried out through a variational inference method. Compared with the prior art, the method has the advantages that the obstacle avoidance success rate is increased by 35%-50%, the false alarm rate is reduced by 40%-60%, the similarity of the technical scheme is lower than 20%, and the method has the advantages that the method is suitable for large-scale popularization and application. And the accuracy, the reliability and the self-adaptive capability of dynamic obstacle avoidance in the construction scene are remarkably improved.
Owner:济南市莱芜区建筑业服务中心

Remote sensing and unmanned aerial vehicle image defogging method

The invention belongs to the technical field of computer vision and image processing, and particularly relates to a remote sensing and unmanned aerial vehicle image defogging method, a defogging network DWTMA-Net is adopted, and the DWTMA-Net is constructed based on a U-shaped architecture and comprises a discrete wavelet block DWB, a multi-dimensional attention module MAB and a wavelet down-sampling module WDM; downsampling of the encoder part is carried out by using Haar discrete wavelet transform through a WDM expansion downsampling method, and frequency information of wavelet transform is combined with spatial information of convolution downsampling; the DWB and the MAB are sequentially arranged between the encoder part and the decoder part, the DWB decomposes features into four frequency components by using Haar discrete wavelet transform (DWT), the low-frequency features are processed by a small AOD network to extract the features, and the high-frequency features are refined by using an expansion residual block; and then spatial information is reconstructed by applying inverse wavelet transform. According to the method, feature representation is improved, and the defogging performance of the network is effectively enhanced.
Owner:ANHUI UNIVERSITY OF TECHNOLOGY

Omnidirectional image super-resolution reconstruction method based on latitude perception potential diffusion model

The invention relates to the field of computer vision and image processing, and discloses an omnidirectional image super-resolution reconstruction method based on a latitude perception potential diffusion model, which comprises the following steps of: acquiring a low-resolution omnidirectional image and performing up-sampling through an interpolation method to obtain a guide image; constructing a latitude perception graph, fusing the latitude perception graph with image features, inputting the fused image features into a latitude perception network, and extracting multi-scale image features; encoding the high-resolution image and the guide image by using an image encoder to obtain a potential representation and a potential residual error; constructing a potential diffusion sequence according to the diffusion scheduling function; starting from the final diffusion representation, performing reverse sampling by using a de-noising network to recover potential representation; the input decoder decodes the image into a high-resolution image; and training optimization is carried out through a joint loss function. According to the technical scheme, the latitude graph is introduced to serve as an auxiliary regulation and control signal, and the spatial semantic features are fused through the multi-scale control module, so that the effect of adaptively adjusting the texture reduction intensity in different latitude regions is achieved.
Owner:COLLEGE OF SCI & TECH NINGBO UNIV

Image generation method and device and electronic equipment

The invention discloses an image generation method and device and electronic equipment, and belongs to the technical field of computer vision and image processing. The method comprises the steps of inputting feature information of a target building into a multi-modal big language model, and performing cross-modal reasoning based on the feature information through the multi-modal big language model to obtain key environment semantic information of the target building; constructing a space block model of the target building by utilizing the key environment semantic information; generating a target space condition control chart according to the space body block model and target observation configuration selected by a user; wherein the target space condition control chart is a visual angle screenshot, corresponding to the target observation configuration, of the space block model; inputting the key environment semantic information and the target space condition control chart into a condition diffusion model to obtain an initial view angle image, corresponding to the target observation configuration, of the target building; and determining a target view angle image of the target building according to the initial view angle image. According to the invention, the image generation effect can be improved.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Object motion trail prediction method and device

The invention provides an object motion trail prediction method and device, and belongs to the technical field of computer vision and image processing, and the method comprises the steps: reconstructing an obtained multi-view time sequence two-dimensional image sequence of a target object, obtaining an original particle point cloud sequence, and generating a sparse key point cloud sequence after sampling; constructing a particle map sequence based on the spatial position of the key point at each moment in the sequence, and performing spatial semantic completion; performing object-level dynamic space-time aggregation on the complemented key point feature sequence to obtain aggregated key point features, inputting the aggregated key point features into a particle diagram converter, and outputting key point updating features after long-range force propagation modeling through the converter; and decoding and predicting key point displacement at a future moment based on the updated feature, and generating a future movement track of the target object according to the key point displacement. Based on the method, the invention further provides a device for predicting the object motion trail. According to the method, the fine appearance and the stable and accurate motion trail of the object can be predicted at the same time.
Owner:HARBIN INST OF TECH AT WEIHAI

Feature fusion method and system for infrared and visible light images

The invention provides a feature fusion method and system for infrared and visible light images, and belongs to the technical field of computer vision and image processing. According to the scheme, important features are captured from the perspective of local reservation and global utilization by adopting a double-branch structure feature extraction encoder based on a convolutional neural network and Transform; and secondly, in consideration of richness and accuracy of enhanced feature representation, an iterative fusion network is constructed, cross-modal deep fusion is realized through multi-level feature interaction, and the accuracy of image feature fusion is effectively ensured.
Owner:XINJIANG UNIVERSITY

Cross-modal adaptive feature integrated RGB-D saliency target detection method

The invention relates to a computer vision and image processing technology, in particular to a cross-modal adaptive feature integrated RGB-D saliency target detection method, which comprises the following steps: S1, data preparation: acquiring an RGB-D data set of a task for training and testing; s2, constructing a network model, firstly constructing a backbone network through feature extraction, then constructing a cross-modal feature integration module TFM, then constructing a modal specific decoder network, and finally constructing an adaptive feature fusion module AFI and calculating a loss function; and S3, indexes are evaluated, and the effectiveness of the network model is evaluated by using four evaluation indexes. The method has the advantages that a polarized attention mechanism (PSA) is introduced to strengthen feature expression, a mid-term fusion method is adopted to realize multi-level cross-modal interaction, single-modal feature redundancy is effectively reduced, cross-modal complementary information is aggregated, and the feature expression capability is enhanced.
Owner:CHANGCHUN UNIV