Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

111 results about "Perceptual image" patented technology

A perceptual illusion differs from a strictly optical illusion, which is essentially an image that contains conflicting data that causes you to perceive the image in a way that differs from reality.

Fine-grained image recognition method and system based on shared visual backbone network

The invention relates to the technical field of image recognition and classification, in particular to a fine-grained image recognition method and system based on a shared visual backbone network. The method comprises the following steps: acquiring a fine-grained image data set; preprocessing the acquired image data; constructing a deep network model, including performing feature extraction by using the deep network model; object proposing and screening are carried out on the extracted feature map; performing feature pooling and information fusion on the screened features; performing global abstract extraction and gating context fusion on the fused features; projecting an object sequence in the gated and fused features to a feature space of a language model; training and optimizing the constructed deep network model; and performing fine-grained image classification by using the optimized model. According to the method, a parallel local feature extraction path is introduced, and a cross-attention fusion module is designed, so that the model can sense the global context and local discriminative details of the image at the same time.
Owner:YANTAI UNIV

Unmanned aerial vehicle full-time perception image reconstruction method based on multi-modal collaborative reinforcement learning and degeneration decoupling

The invention provides an unmanned aerial vehicle full-time perception image reconstruction method based on multi-mode cooperative reinforcement learning and degeneration decoupling. A frequency perception feature modulation model, a dual-mode dual-domain transformation module and a dynamic bidirectional guide mechanism are included. According to the system, firstly, feature information of different frequency bands is adaptively separated and modulated through a frequency sensing feature modulation model, and decoupling and compensation of composite unknown degradation are achieved; realizing cross-domain interaction and information fusion of visible light and infrared characteristics in a spatial domain and a channel domain by using a bimodal dual-domain transformation module; and finally, realizing collaborative enhancement of cross-modal degradation perception through a bidirectional dynamic guide mechanism, and generating an unmanned aerial vehicle visible light reconstruction image and an infrared super-resolution image with higher structural consistency and texture fidelity. According to the method, deep fusion and degeneration decoupling of multi-modal information can be realized in a complex degeneration environment, and the imaging quality and the environmental adaptability of an unmanned aerial vehicle full-time sensing system are remarkably improved.
Owner:HENAN UNIV OF SCI & TECH

Detecting Phishing Websites Using Perceptual Image Hashing

Systems and methods for detecting phishing using image hashing include obtaining a plurality of images from different sources, generating a hash for each image, comparing at least one hash associated with a first image to one or more hashes associated with a second image, calculating a similarity score based on the comparing, and classifying the first image based on the similarity score.
Owner:ZSCALER INC

Efficient image super-resolution method and device based on dual guidance semantic enhancement

The invention provides an efficient image super-resolution method and device based on dual guidance semantic enhancement, and relates to the technical field of image super-resolution, visual content generation and the like in computer image process.The method comprises the steps that an image super-resolution model containing a pre-training diffusion model backbone structure is constructed, and a pre-training diffusion model is constructed; a trainable LoRA module is inserted into a U-Net structure and a VAE encoder of the image super-resolution model; calculating the perception image loss of the generated image and the real image by adopting a dual-guidance quality enhancement training strategy; based on a semantic alignment method of score matching, performing semantic feature constraint on the generated image through a condition vector of a pre-training diffusion model to form semantic consistency loss aligned with real image distribution; and carrying out end-to-end training on the image super-resolution model under joint optimization of perceptual image loss and semantic consistency loss, and after the training is finished, realizing high-resolution generation of a low-resolution image only through one-time U-Net reasoning.
Owner:TSINGHUA UNIVERSITY

Extracting images and determining their meaning for semantic image retrieval and training a transformer-based multi-modal large language model to generate domain-aware images based on image meanings

The disclosure relates to systems and methods automatically extracting an image and related image components, computationally determining an understanding of the image, and generating mathematical vector embeddings via sentence encoders based on the computationally determined understanding. The mathematical vector embeddings may be used for semantic image retrieval that enables image searching based on a semantic understanding of input images and / or input text. The mathematical vector embeddings may be used for training and executing generative Artificial Intelligence (AI) models to create new content that includes retrieved images and / or generate new images.
Owner:ROHIRRIM INC

View-conditioned diffusion for real-world vehicle gaussian splatting

Systems and methods for view-conditioned diffusion for real-world vehicle gaussian splatting. A single perspective image can be transformed using image transformation techniques to generate a training dataset that addresses a domain gap between synthetic data and real-world data in a traffic scene. A pre-trained diffusion model can be finetuned with the training dataset to obtain a fine-tuned diffusion model. Perspective-aware images having different perspective views of an entity from the single perspective image can be generated using the fine-tuned diffusion model. A large generative model (LGM) can be trained using the perspective-aware images to generate a gaussian splatting model for the entity. View-conditioned simulations from the single perspective image can be generated by using the gaussian splatting model for downstream tasks.
Owner:NEC LABORATORIES AMERICA INC

Wire mesh information processing method and device based on real-time state perception image

The present application provides a method and device for processing rail network information based on real-time state perception images, relating to the field of information processing technology. The method comprises: determining an initial rail network image under normal conditions for all rail transit network information; the initial rail network image represents the attributes, state, and location of rail transit network equipment; and performing pixel-level updates on the initial rail network image based on the real-time state of the rail transit network to obtain a real-time rail network image. The method and device for processing rail network information based on real-time state perception images provided by the present application can provide a real-time and intuitive description of rail transit network information that is difficult to obtain.
Owner:CRSC URBAN RAIL TRANSIT TECH CO LTD

Utilizing a diffusion neural network for mask aware image and typography editing

The present disclosure relates to systems, methods, and non-transitory computer readable media for utilizing a diffusion neural network for mask aware image and typography editing. For example, in one or more embodiments the disclosed systems utilize a text-image encoder to generate a base image embedding from a base digital image. Moreover, the disclosed systems generate a mask-segmented image by combining a shape mask with the base digital image. In one or more implementations, the disclosed systems utilize noising steps of a diffusion noising model to generate a mask-segmented image noise map from the mask-segmented image. Furthermore, the disclosed systems utilize a diffusion neural network to create a stylized image corresponding to the shape mask from the base image embedding and the mask-segmented image noise map.
Owner:ADOBE INC

Equipment identification and quality grade joint evaluation method for substation inspection image

The invention discloses an equipment identification and quality grade joint evaluation method for a transformer substation inspection image, and the method comprises the steps: inputting a collected original transformer substation image into an image identification and quality grade joint evaluation model; the model comprises a detection perception image enhancement module, a target detection module, an ROI extraction module and an image quality evaluation module which are connected in sequence. Performing enhancement processing on the original transformer substation image by using a detection perception image enhancement module to obtain an enhanced image; performing target detection on the enhanced image by using a target detection module to obtain a power equipment category label and ROI frame coordinate information; an ROI extraction module is utilized to cut the original substation image according to the ROI frame coordinate information to obtain a corresponding equipment ROI image; performing no-reference image quality evaluation on the ROI image of the equipment by using an image quality evaluation module to obtain a quality grade label; according to the method, the equipment identification accuracy is improved, and the image quality grade is credibly output.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Semantic and structure perception image restoration method and device based on diffusion model

The invention provides a semantic and structure perception image restoration method and device based on a diffusion model. The method comprises the steps of obtaining a to-be-processed image and a mask used for identifying a to-be-restored area; performing panoramic segmentation on the to-be-processed image to obtain a semantic feature map, and constructing a peripheral semantic candidate set for the mask boundary to constrain semantic distribution of the to-be-restored region; extracting a structural feature map of the to-be-processed image, wherein the structural feature map comprises at least one of edge information, line segment information and angular point information; and inputting the to-be-processed image, the mask, the semantic feature map and the structural feature map into a multi-conditional diffusion generation network, and generating repair content in the to-be-repaired area by the multi-conditional diffusion generation network under the constraint of the semantic feature map, the structural feature map and the peripheral semantic candidate set, and outputting the restored image.
Owner:XIAMEN ZHENJING TECH CO LTD

Image forgery detection method based on multi-modal large language model

The invention relates to the technical field of image detection, and provides an image forgery detection method based on a multi-modal large language model, which comprises the following steps: inputting a to-be-detected image into the multi-modal large language model; the multi-modal large language model outputs an authenticity identification result of the to-be-detected image; wherein the multi-modal large language model has three parallel expert branches of physical consistency, semantic consistency and underlying structure clue analysis. According to the method, semantic information in the image is perceived and the context logic and the physical logic of the image are understood by fully utilizing the powerful priori knowledge, the deep semantic understanding capability and the chained reasoning capability of the multi-modal large language model, and meanwhile, complex reasoning and explanation can be performed on a forged image through a natural language.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

A data processing method and apparatus

According to embodiments of this disclosure, a method and apparatus for data processing are provided. The method includes: determining a reference region corresponding to a target traffic light indicated by map data in an image coordinate system corresponding to a perceived image; determining a set of candidate objects from a plurality of candidate objects identified based on the perceived image, based on the positional relationship between the reference region and a plurality of candidate regions, wherein the plurality of candidate regions correspond to the plurality of candidate objects identified as traffic lights; determining a first correlation between a first attribute of the target traffic light and a second attribute of the set of candidate objects, wherein the first attribute is determined based on the map data and the second attribute is determined based on the perceived image; and matching the target traffic light with the target objects in the set of candidate objects based on the first correlation. Based on this approach, embodiments of this disclosure can more accurately identify and match traffic lights.
Owner:BEIJING VOYAGER TECH CO LTD

Reconstruction of a mixed strategy three-dimensional medical image visual language model pre-training method

The application discloses a three-dimensional medical image visual language model pre-training method of reconstruction mixed strategy, belongs to the technical field of medical image computation, and comprises the following steps: constructing a medical image text pair data set, a language text mask reconstruction strategy, a visual image mask reconstruction strategy, a semantic perception fusion strategy and multi-task joint learning; the large language model is fine-tuned, the fine-tuned large language model is used to extract diagnosis and attribute information in a medical report and generate efficient prompts, the large language model has strong generalization ability, and the cost of manual labeling is greatly saved; the semantic perception fusion strategy is used for combining text features obtained by a text encoder and image features obtained by an image encoder to obtain new text features, so that the text perceives diagnosis and attribute information of the image in advance, the alignment of the image and the text in an embedding space is further optimized, and the pre-training efficiency is improved.
Owner:XIDIAN UNIV

Method and Display System for Operating a Projection Display System in a Mobile Device

A display system for a motor vehicle includes a projection device having a display device which is configured to output image-bearing light for projecting a display image in an eye area of a user via a display surface. A combiner surface reflects the image-bearing light into the eye area of the user. A device configured to obtain an eye pose of the user transforms a display image with the aid of a post-transformation matrix so that local flaws in a surface section of the combiner surface which result in a distortion or a fuzziness of a perception image perceptible by the user are compensated. A display image is displayed on the display surface using brightness values so that a homogeneous and high contrast representation results based on the eye pose of the user and the detail of the combiner surface. The transformed display image is output via the display device.
Owner:BAYERISCHE MOTOREN WERKE AG

Vehicle control method, device, equipment, medium, vehicle and program product

The invention relates to the field of vehicle energy management, and discloses a vehicle control method, device and equipment, a medium, a vehicle and a program product, and the method comprises the steps: determining the multi-dimensional features of the vehicle based on the driving data of the vehicle and a perception image; acquiring a torque control strategy matched with the multi-dimensional features; and controlling the output torque of the vehicle according to the torque control strategy. According to the method, the multi-dimensional features can be extracted from the driving data and the sensing image, then the torque control strategy is matched, optimal control over the output torque of the vehicle is achieved, and energy consumption of the vehicle is greatly reduced.
Owner:BYD CO LTD +1

Method, device and equipment for generating online map based on diffusion model and medium

The application relates to a remote sensing image generation online map method, device, equipment and medium, a plurality of different transformed sample images are obtained by performing geometric transformation on remote sensing sample images in each group of sample image pairs, the plurality of transformed sample images are added to the corresponding sample image pairs to obtain a training data set, each training data set is used to train an online map generation model to obtain a trained online map generation model, an encoder in a perception image compression network is used to map remote sensing sample images and corresponding transformed sample images from a pixel space to a feature space, a forward diffusion and reverse denoising are performed on the feature space by a denoising diffusion bridge network, and then an output of the denoising diffusion bridge network is mapped from the feature space to the pixel space by a decoder in the perception image compression network, and a real-time remote sensing image is input into the trained online map generation model to obtain a real-time network map. By adopting the method, a map with clearer boundaries and brighter colors can be generated in real time.
Owner:NAT UNIV OF DEFENSE TECH

Methods and systems for image co-registration of multi-modal temporal sensing

The disclosure generally relates to methods and systems for image co-registration of multi-modal temporal sensing. Conventional techniques for image co-registration of multi-modality images focus on either the spatial or temporal domain and thus are not of high accuracy and do not preserve both global and local characteristics for the matching. The present disclosure solves the technical problems in the art for image co-registration of multi-modal temporal sensing using a deep-learning based multi-input-output encoder-decoder network with a Gabor Jet Model. The deep-learning based multi-input-output encoder-decoder network is utilized for the feature extraction. A distinctive Gabor-jet layer of the Gabor Jet Model is utilized for the similarity matching. The Gabor-jet layer generates a Gabor jet graph which provides sparse feature points for matching between matching images and reference images.
Owner:TATA CONSULTANCY SERVICES LTD

Multi-modal CNN-Transform fused image tampering detection method

The invention belongs to the field of computer vision, and particularly relates to a multi-mode CNN-Transform fused image tampering detection method, which designs a CNN and Transform double-flow parallel feature extraction structure, effectively combines the advantage of CNN at capturing fine local features and the advantage of Transform in capturing long-distance dependency relationship and global semantic information, and improves the accuracy of image tampering detection. Local details and global tampering features in the image can be sensitively perceived at the same time, so that the detection capability and stability of image tampering are effectively improved. Processing the noise domain image filtered by the SRM through a CNN (Convolutional Neural Network), and capturing local texture and noise artifact features; meanwhile, a Transform branch processes an RGB spatial domain image, extracts rich global semantic information, and realizes deep interaction and advantage complementation of two types of modal information through a feature fusion module. The BAFM module provided by the invention can deeply mine feature information of different scales and different spatial directions, and more accurate spatial attention features are generated, so that the sensitivity and expression ability of the network to tampered regions are improved.
Owner:NANTONG UNIV

Visual SLAM method and system based on saliency prediction and storage medium

The invention discloses a visual SLAM method and system based on saliency prediction and a storage medium, and belongs to the technical field of synchronous positioning and mapping. Firstly, saliency prediction based on multi-modal fusion is carried out according to a current image frame and a depth frame, a grey-scale map of the current image frame and a saliency mask containing an effective structured region are obtained, and then feature extraction and matching are carried out to obtain matched feature points; then calculating the saliency entropy of the current image frame and judging a key frame, and creating map points for real-time grading to obtain a graded local map; and finally, performing global BA weighted optimization, and constructing a global map by continuously expanding and maintaining the graded local map. According to the method, a significance prediction technology based on multi-modal fusion is introduced, geometric information and depth information are fused to accurately perceive an effective structured region in an image, significant features are captured and accurately matched, the perception and association capability of the system is enhanced, and the stability and overall performance of the system are comprehensively improved.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Illumination perception adaptive method and system for night pedestrian re-identification

The invention discloses an illumination perception adaptive method and system for night pedestrian re-identification. The method comprises the following steps: S1, obtaining an input image; s2, processing the input image through an illumination perception image preprocessing network, and dynamically generating an enhancement coefficient graph and a noise confidence graph; s3, performing detail enhancement on the input image through an adaptive enhancement branch to obtain an enhanced image; performing noise suppression on the input image through a feature purification module to obtain a purified image; s4, performing feature extraction on the enhanced image and the purified image by using a shared backbone network to obtain enhanced features and purified features respectively; and S5, fusing the enhancement feature and the purification feature through an illumination perception controller to obtain a final identity representation. According to the invention, through end-to-end illumination perception adaptive processing, image enhancement and feature purification are effectively cooperated, and the discrimination capability and robustness of a night pedestrian re-identification model in a low-light environment are significantly improved.
Owner:SUN YAT SEN UNIV

Method and system for performing three-dimensional (3D)-aware image editing

PendingUS20260253316A1Pattern recognitionColor texture
A method and a system for performing image editing based on an attribute-specific text prompt includes acquiring a noise code (z), a textual instruction (Ai) specifying a target facial attribute to be edited, and a target camera pose (pt). Upon acquiring, mapping the noise code (z) to a latent code (w), via a mapping network. Once the mapping is done, editing the latent code (w) based on the textual instruction (Ai) to generate an edited latent code (ŵ), via a text-driven Latent Attribute Editor (LAE). Further, based on the edited latent code (ŵ), generating a color texture image and a set of alpha maps via a three-dimensional Generative Adversarial Network (3D GAN). Furthermore, based on the color texture image and the set of alpha maps, generating a 3D-aware and view-consistent image at the target camera pose (pt) via a differentiable renderer.
Owner:MOHAMED BIN ZAYED UNIV OF ARTIFICIAL INTELLIGENCE

Method, apparatus, device, storage medium and program product for generating trajectories

According to embodiments of this disclosure, a method, apparatus, device, storage medium, and program product for generating trajectories are provided. The method includes: determining visual features of a set of perceived images captured by an autonomous vehicle; acquiring a reference trajectory generated by a visual language model based on the visual features; and determining a target trajectory of the autonomous vehicle by a trajectory planning model based on the visual features and the reference trajectory, wherein the trajectory planning model is an end-to-end model. Based on this approach, embodiments of this disclosure can improve the quality of the generated trajectory.
Owner:BEIJING VOYAGER TECH CO LTD

A multispectral image fusion method, device and equipment based on multi-target segmentation

The application discloses a multispectral image fusion method and device based on multi-target segmentation, and equipment, and relates to the technical field of image processing. Different fusion methods are adopted for different targets in the feature domain according to the category of multi-target segmentation, so that the image conforming to the human eye visual perception is generated, and the significant target in the infrared image is effectively highlighted. The method comprises the following steps: collecting a visible light image and an infrared image, performing image registration processing on the visible light image and the infrared image to obtain a registered target visible light image and a target infrared image; performing multi-target semantic segmentation on the target visible light image and the target infrared image by using a multi-target segmentation network to generate a multi-target segmentation image; extracting a first deep feature corresponding to the target visible light image and a second deep feature corresponding to the target infrared image based on a multispectral image fusion network; and generating a target fusion image by fusing the first deep feature, the second deep feature and the multi-target segmentation image.
Owner:ZHEJIANG UNIV +1

Automatic driving perception data processing method, device and system and vehicle

The invention relates to an automatic driving perception data processing method, device and system and a vehicle, and the method comprises the steps: obtaining a data file of at least one automatic driving task, and recording perception annotation data and perception image data in the data file; based on the data file of the at least one automatic driving task, dictionary data of the at least one automatic driving task is established, and the dictionary data indicates summarized data feature values of the corresponding automatic driving tasks; based on the dictionary data and the data file, a report file is generated, the report file comprises at least one task record, and the task records are in one-to-one correspondence with the automatic driving tasks. The dictionary data is established for the data file, the summarized data type of each automatic driving task can be quickly obtained, the report file is generated by utilizing the dictionary data and the data file, the data of different files can be summarized into the report file in a unified format according to the automatic driving tasks, the data summarizing efficiency can be improved, and the data summarizing efficiency is improved. And the report error rate is reduced.
Owner:CHINA AUTOMOTIVE INNOVATION CORP

Display system having 1-dimensional pixel array with scanning mirror

Display systems are described including augmented 1-dimensional pixel arrays and scanning mirrors. In one example, a pixel array includes first and second columns of pixels, relay optics configured to receive incident light and to output the incident light to a viewer, and a scanning mirror disposed to receive the light from the first and second columns of pixels and to reflect the received light toward the relay optics. The scanning mirror may move between a plurality of positions while the first and second columns emit light in temporally spaced pulses so as to form a perceived image at the relay optics having a higher resolution relative to the pixel pitch of the individual columns. Foveated rendering may provide for more efficient use of power and processing resources.
Owner:MAGIC LEAP INC

Quantification method and device for moire patterns of screen shot image and storage medium

ActiveCN121860870AHigh compensation accuracyImage enhancementImage analysisVisual perceptionSpectral amplitude
The invention discloses a screen shooting image moire quantification method and device and a storage medium, which are used for improving the compensation precision of a display screen. Acquiring a sub-pixel unit image of the display screen to be tested; performing brightness compensation processing on the sub-pixel unit image to generate an expected brightness compensation image; generating a compensation graph frequency spectrum amplitude image according to the brightness compensation image; generating a stripe frequency band sensing image according to the visual distance and the pixel pitch corresponding to the target color channel; carrying out region division on the frequency spectrum amplitude image of the compensation graph; standard deviations of the elliptical region and the plurality of elliptical ring regions are calculated, and a region standard deviation set is generated; filling a blank image with the standard deviation in the regional standard deviation set to generate a first elliptical ring standard deviation image; performing frequency domain filtering on the first elliptical ring standard deviation image by using the stripe frequency band sensing image to generate a first frequency domain filtering image; and generating a first moire quantization value corresponding to the brightness compensation image according to the compensation image frequency spectrum amplitude image and the first frequency domain filtering image.
Owner:SHENZHEN SEICHITECH TECHN CO LTD

Vehicle-road-cloud integrated automatic driving vehicle test method and system

The invention relates to the related technical field of automatic driving vehicle testing, in particular to a vehicle-road-cloud integrated automatic driving vehicle testing method and system, and the method comprises the steps: a cloud obtains a background vehicle motion trail through analyzing an OpenScenario test case, and transmits the background vehicle motion trail to a tested automatic driving vehicle through road end equipment; inputting randomly generated Gaussian noise into a pre-training diffusion model by taking an image obtained by a vehicle end and a virtual background vehicle state as input, and performing fusion image generation through T times of noise reduction at a cloud end; and the generated vivid image is sent to a vehicle end measured algorithm, so that the self-driving vehicle test based on virtual-real fusion is realized. According to the invention, the test efficiency can be effectively improved by combining the real automobile, the real driving road and the virtual-real test scene of the tested automatic driving automobile, and the reliability of the test result can be effectively ensured by the generated high-fidelity perception image.
Owner:XIAN TECH UNIV +1

Perceptual image compression model evaluation method and system based on multi-index fusion

The invention discloses a perception image compression model evaluation method based on multi-index fusion, and belongs to the technical field of digital image compression. In order to solve the problem that an existing image quality evaluation method is difficult to comprehensively reflect subjective perception of human eyes, after multiple evaluation indexes are subjected to normalization processing, adaptive fusion is carried out through a three-layer full-connection neural network, and a comprehensive score is output based on a weighted summation mode. According to the method, the image compression quality can be measured more comprehensively and accurately, and the consistency of the model and human eye perception is improved.
Owner:PEKING UNIV

Contact lens with gradient optical system

The present invention relates to wearable optics, in particular to contact lenses containing an integrated information display device in the form of a micro-display with gradient lenses, and can be used to form augmented reality, virtual reality or augmented reality (AR / VR / XR). The claimed contact lens comprises a flat panel display having a maximum size d, the screen of which is located on the axis of symmetry of the contact lens and towards the eyes of the user; a power supply and a control unit; and a collimating optical system located between the display and the eye. The collimating optical system is implemented as a single converging lens having a diameter D and a radial refractive index gradient and the optical axis of which coincides with the axis of symmetry of the contact lens. The technical result consists in reducing the overall size of the contact lens while maintaining the clarity of the perceived image on the display screen.
Owner:XPANCEO RESEARCH ON NATURAL SCIENCE LLC