Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4633 results about "Image generation" patented technology

Image Generation is your video production partner because we care as much as you do, and we have the experience and ability to bring that passion to the screen on every project. Kevin McKeever, owner of Image Generation, produces and shoots every project personally, giving you a single point of contact.

Cross-modal joint source-channel coding and decoding method adaptable to changeable scenarios

PCT designated stageWO2026031415A1Internal combustion piston enginesBiological modelsChannel decoderImage signal
The present invention relates to the technical field of cross-modal image signal reconstruction. Disclosed is a cross-modal joint source-channel coding and decoding method adaptable to changeable scenarios. The method comprises: first designing a Transformer encoder-based cross-modal channel coding and decoding optimization solution, so as to achieve the performance improvement and robustness of a channel encoder and a channel decoder; then designing a cross-modal source coding and decoding optimization solution for haptic-to-image generation based on a latent diffusion model, so that under an image signal loss scenario, haptic information is used to guide image generation; and finally, incorporating transfer learning technology, so as to reduce additional training costs caused by a system facing changeable cross-modal communication scenarios such as a changeable channel signal-to-noise ratio and different transmission tasks. Under cross-modal changeable communication scenarios, the joint source-channel coding and decoding method provided in the present invention can solve the problems of the inability of a receiving end to well complete image reconstruction, and additional model training costs caused by changeable channel environments and scenarios.
Owner:NANJING UNIV OF POSTS & TELECOMM

Automobile part quality detection method and system based on artificial intelligence visual inspection

The invention discloses an automobile part quality detection method and system based on artificial intelligence visual inspection, and belongs to the field of artificial intelligence machine visual inspection, and the method comprises the steps: firstly, carrying out the registration of a collected RGB image and a depth image, and extracting a part region through a saliency detection network; two-dimensional key points are extracted based on an RGB region image and are matched with key points of a three-dimensional model, an initial three-dimensional attitude is obtained by adopting a PnP algorithm, iterative registration is performed with the three-dimensional model in combination with a point cloud generated by a depth image, and a fine three-dimensional attitude is obtained. And calculating a geometric transformation matrix from the part to a standard front view attitude according to the attitude, and performing attitude correction on the RGB and depth region image. And then matching the corrected image with a standard template image by using a feature detection and matching network so as to correct the position of the detection window. And finally, the three-dimensional size of the part is calculated in the corrected detection window in combination with the depth value, and tolerance judgment is carried out. And the precision, the robustness and the automation level of online detection of the automobile parts can be obviously improved.
Owner:XIANYANG VOCATIONAL TECHN COLLEGE

Location Search Based on Model-Generated Synthetic Images

Systems and methods for searching using machine-learned model-generated outputs can provide a user with a medium for generating synthetic images depicting synthetic environments that can then be matched to a real world example. The systems and methods can include obtaining a search query, which can be utilized to generate a prompt input that can be processed by an image generation model to generate a plurality of model-generated images. A selection can then be received that selects a particular model-generated image to utilize to query a database.
Owner:GOOGLE LLC

Multimodal sentiment analysis method based on diffusion model and self-paced learning

The invention provides a multi-modal sentiment analysis method based on a diffusion model and self-paced learning. The method comprises the following steps: firstly, dividing a data set into a missing image modal data set and a complete modal data set according to image modal integrity; thirdly, constructing a feature alignment diffusion model, and performing image generation; training the diffusion model by adopting a self-paced learning strategy and a missing image data set; and based on the trained diffusion model, guiding a reverse process through text features to generate feature representation of the missing image. And carrying out weighted fusion on the generated image features and text features by using an attention mechanism, and dynamically adjusting contribution weights of all modalities to generate a complete multi-modal feature representation. And finally, integrating a missing modal completion result and the complete modal features to form a unified multi-modal representation, inputting the unified multi-modal representation into a multi-modal sentiment classification module, and outputting a sentiment classification result. According to the method, the problem of multi-modal sentiment analysis under random missing of image modals is effectively solved, and the generation quality and semantic consistency are improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Single view reconstruction and rendering method

The invention discloses a single view reconstruction and rendering method, which belongs to the field of view reconstruction and rendering, and comprises the following steps of: firstly, generating multi-view feature representation with strong geometric consistency from a single input image by introducing an image diffusion module of a cross attention mechanism; a point cloud reconstruction module with self-attention and cross-attention is utilized, and multi-view information is fused to reconstruct an accurate three-dimensional point cloud; and finally, constructing differentiable three-dimensional Gaussian representation based on the point cloud, rendering the differentiable three-dimensional Gaussian representation, and outputting a new view angle image, a normal map and a depth map. According to the system, a staged training strategy is adopted, and a composite loss function including multi-scale bidirectional consistency smooth loss and feature consistency loss is innovatively used for optimization. According to the method, the problems of low geometric accuracy, multi-view inconsistency, detail missing and the like in single-view reconstruction are effectively solved, and the method can be widely applied to the fields of virtual reality, digital twinning, cultural heritage digitization and the like.
Owner:北京渲光科技有限公司

Retrieval augmented text-to-image generation

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output image using a text-to-image model and conditioned on both the input text and image and text pairs selected from a multi-modal knowledge base. In one aspect, a method includes, at each of multiple time steps: generating a first feature map for the time step; selecting one or more neighbor image and text pairs based on their similarities to the input text; for each of the one or more neighbor images and text pairs, generating a second feature map for the neighbor image and text pair; applying an attention mechanism over the one or more second feature maps to generate an attended feature map; and generating an updated intermediate representation of the output image for the time step.
Owner:GOOGLE LLC

Appearance defect detection system and method based on CCD image recognition

The invention provides an apparent defect detection system and method based on CCD image recognition. A specular reflection interference area meeting a preset difference condition is recognized through a first surface reflection image and a second surface reflection image of a target patch element; extracting corresponding orthographic gray scale data from a third surface orthographic image obtained by adopting coaxial light source irradiation based on the specular reflection interference area; according to the brightness distribution gradient of the specular reflection interference area in the first surface reflection image and the second surface reflection image, constructing a texture restoration guide map, and using the texture restoration guide map to carry out texture restoration on the orthographic gray scale data to generate a texture restoration image of the specular reflection interference area; and generating a defect identification reference image for eliminating specular reflection interference based on the texture repair image, and outputting an appearance defect identification result of the target patch element by the defect identification reference image. According to the technical scheme provided by the invention, the defect detection of the patch element can be realized under the interference of mirror reflection.
Owner:SHENZHEN SANYILIANGUANG INTELLIGENT EQUIP CO LTD

Generative object compositing by learning identity-preserving representation

A method, apparatus, non-transitory computer readable medium, and system for image generation include obtaining a foreground image and a background image. The foreground image depicts an object and the background image depicts a scene. The foreground image is encoded, using an image encoder of an image generation model, to obtain a foreground embedding. The foreground embedding preserves the identity of the object. A composite image is generated, using the image generation model, based on the background image and the foreground embedding. The composite image depicts the object from the foreground image within the scene from the background image.
Owner:ADOBE INC

Systems and methods for segmentation using retrieval augmentation

A system and a method are disclosed for classifying features from an input image. The method includes generating, by a processing circuit, a segment feature from a first input image, the segment feature corresponding to a group of pixels associated with an object represented in the first input image, and being an out-of-vocabulary segment feature; performing, by the processing circuit, a retrieval of a first feature vector, corresponding to the segment feature, from a database of feature vectors, the first feature vector representing an object-specific segmentation mask; and generating, by the processing circuit, an output segmentation mask based on the first feature vector.
Owner:SAMSUNG ELECTRONICS CO LTD

Blind person navigation path planning method combining visual SLAM and semantic segmentation

The invention relates to the technical field of visual navigation, in particular to a blind person navigation path planning method combining visual SLAM and semantic segmentation. The method comprises the following steps: acquiring an environment image and pose data; performing front-end tracking by using the pose data, and determining a key frame sequence; generating a depth map based on the environment image, and identifying a passable area according to the depth map; constructing a three-dimensional point cloud map according to the key frame sequence; semantic segmentation is carried out according to the three-dimensional point cloud map, and obstacle type labels are recorded; according to the obstacle type label, carrying out safety level layering on the passable area, and determining a path cost weight; performing global path planning based on the path cost weight, and generating a semantic enhancement navigation trajectory; and driving real-time voice guidance by using the semantic enhancement navigation trajectory, and calculating a path execution deviation corresponding to the real-time voice guidance. According to the method, obstacle types are distinguished through semantic segmentation based on a visual navigation technology, a semantic enhancement path is generated, semantic hierarchical navigation is realized, and the intelligence of path decision is improved.
Owner:SHANDONG SAIFEITE SAFETY ENG TECH DEV CO LTD

Three-dimensional target detection method based on double-point cloud enhancement and fusion

The invention discloses a three-dimensional target detection method based on double-point cloud enhancement and fusion, and relates to the technical field of automatic driving. The method comprises the following steps: firstly, respectively extracting local features of a laser radar point cloud and a pseudo point cloud generated by a front view image in a candidate region by using the laser radar point cloud and the pseudo point cloud; the features of the false point cloud are enhanced through residual geometric modeling, depth separable convolution and a space-channel attention mechanism, and a three-layer cascade adaptive unit is constructed for further optimization; the original point cloud RoI features and the pseudo point cloud RoI features are mapped to a unified space, a bidirectional attention mechanism is adopted to realize fine-grained cross-modal fusion, and complementary fusion features are generated; and finally, outputting a category confidence coefficient and a 3D bounding box through a multi-branch detection head and a multi-task loss function. According to the method, the problem of point cloud sparseness is effectively relieved, the detection precision is improved, the bottlenecks of roughness, alignment difficulty and the like of a traditional fusion method are overcome, and the perception robustness and accuracy of an automatic driving system in a complex scene are remarkably enhanced.
Owner:NORTH CHINA ELECTRIC POWER UNIV

Hardware tool defect online detection system based on AI image recognition

The invention discloses a hardware tool defect online detection system based on AI image recognition, and belongs to the technical field of image generation type adversarial networks. A YOLOv8 detection framework model is trained, a CBAM attention module is specifically added into a YOLOv8-S detection model, feature extraction of a hardware tool rare defect area is enhanced, feature distribution of hardware tool rare defect samples generated by a GAN is fused in a network bottleneck layer, the feature learning ability of the YOLOv8-S detection model for hardware tool rare defects is improved, and the hardware tool rare defect feature extraction method based on the CBAM attention module is obtained. The detection effect on the rare defects of the hardware tool is enhanced, and the problem of model performance bottleneck caused by scarcity of rare defect samples of the hardware tool is solved.
Owner:金华高格软件有限公司

Context-enhanced image generation method, and model training method and system

The present invention provides a context-enhanced image generation method, and a model training method and system, comprising: acquiring a training image data set; using a preset noise-adding mechanism to perform multi-timestep noise-adding on a training image, wherein a variance of Gaussian noise added at each timestep depends on a current timestep and progressively increases until a clean training image is transformed into standard Gaussian noise so as to obtain a noisy image at each timestep; using a random masking mechanism to generate a mask, and according to the mask, discarding a masked region of the noisy image and discarding an unmasked region of the clean image; and on the basis of reconstructing a discarded region using a non-discarded noisy image and the clean image, training a context-enhanced image generation model so as to obtain a trained image generation model. The present invention effectively improves the context comprehension capabilities of image generation methods, improves image generation quality, and achieves high-resolution diversified image generation.
Owner:SHANGHAI JIAOTONG UNIV

Electric power operation target detection method based on multi-mode large model knowledge distillation

The invention relates to the field of target detection, and particularly discloses an electric power work target detection method based on multi-modal large model knowledge distillation, which utilizes a vision-language multi-modal large model as a teacher model, and improves the target detection efficiency by expanding prompt word guidance. A high-quality pseudo label and a region-text pair are generated for an unlabeled electric power work image as a supervision signal, and on this basis, through joint optimization of detection loss, feature distillation loss, logic distillation loss and multi-modal contrast learning loss, a lightweight YOLO student model is guided to learn positioning and classification knowledge and to learn a multi-modal contrast learning loss. And deep alignment with the open vocabulary understanding ability of the teacher model is carried out on the feature space and semantic level, so that a semantic gap between closed category detection and open world perception is effectively bridged. Through the mode, the detection precision and generalization ability of the student model on common, rare and even unseen targets in the electric power work scene are remarkably improved.
Owner:MARKETING SERVICE CENT OF STATE GRID HENAN ELECTRIC POWER CO

Identifying Items in Images Using Embeddings Generated from the Images and Ranking Candidates Using a Language Model

An online system applies a visual language model and an optical character recognition model to a received image to generate descriptive information about unknown items in the image. The online system prompts a generative model with the descriptive information about unknown items in the image to separate the descriptive information into different bins each corresponding to a different unknown item in the image. For each unknown item detected in the image, the online system generates a target embedding from its descriptive information and performs a nearest neighbor search on an item catalog including embeddings for various items to find a set of candidate embeddings matching the target embedding. The online system retrieves item attributes of candidate items each corresponding to a candidate embedding of the set and prompts the generative model with this information to rank candidate items for the unknown item in the image.
Owner:MAPLEBEAR INC

Abnormal sample image generation method, electronic equipment and storage medium

The invention is suitable for the technical field of artificial intelligence, and provides an abnormal sample image generation method, electronic equipment and a storage medium, and the method comprises the steps: setting a material parameter library and a scene parameter library, and constructing a paired industrial defect sample data set in combination with three-dimensional geometric models of a plurality of sample workpieces; constructing category text description, defect text description and material category text description of various workpieces, and constructing and training an abnormal sample image generation model in combination with the paired industrial defect sample data set; obtaining a target normal image, a candidate defect area mask image, a target category text description, a target defect text description and a target material category text description of a target category workpiece, and inputting the target normal image, the candidate defect area mask image, the target category text description, the target defect text description and the target material category text description into an abnormal sample image generation model for processing to obtain a target defect image and a target defect area mask image; the training precision and flexibility of the abnormal sample image generation model are improved, and then the efficiency and precision of abnormal sample generation are improved.
Owner:SPEEDBOT ROBOTICS CO LTD

Orientation sensitive target detection method based on sub-aperture color image saturation characteristics

The invention discloses an SAR image orientation sensitive target detection method based on sub-aperture image saturation characteristics, and belongs to the technical field of synthetic aperture radar image target detection. The azimuth sensitive target detection is realized through the following steps: 1) sub-aperture image generation: dividing an azimuth frequency spectrum into a plurality of sub-bands, and generating a plurality of sub-aperture images through inverse Fourier transform; 2) color synthesis and feature extraction: allocating different hues to each sub-aperture image by using an HSV color space, synthesizing an RGB color image, converting the RGB color image to the HSV color space, and extracting a saturation channel in the RGB color space as an azimuth sensitivity feature map; and 3) target detection: carrying out threshold segmentation on the saturation feature map, and identifying a pixel region with a high saturation value, namely, an orientation sensitive target. According to the method, the color saturation change caused by the scattering difference of the target in different sub-apertures is utilized, effective detection of azimuth sensitive targets such as artificial buildings and vehicles is achieved, and the method has the advantages of being simple in calculation and high in robustness.
Owner:NANJING UNIV OF SCI & TECH

Image watermarking method based on space channel interactive attention mechanism

The invention discloses an image watermarking method based on a space channel interactive attention mechanism. The image watermarking method mainly comprises the following steps: acquiring an original image and constructing an image watermarking network based on the space channel attention mechanism; inputting the original image and the watermark into a watermark-containing image generation framework to obtain a watermark-containing image; inputting the image with the watermark into an image processing noise layer to obtain a noise image; inputting the noise image into a watermark extraction framework to obtain an extracted watermark; and constraining the training of the whole image watermark network based on the space channel attention mechanism by using the total loss. The image watermark network based on the space channel attention mechanism is used for copyright protection, and high-capacity embedding of watermarks, high-quality generation of images with watermarks and high-robustness extraction of the watermarks are achieved. According to the method, a depth feature map of the watermark is generated through watermark dimension extension, remodeling and diffusion by using a multi-layer sensor and the like. Meanwhile, the feature retention loss can promote the optimization of the total loss on the network training.
Owner:SHANDONG NORMAL UNIV

Esophageal endoscope image generation system and method based on improved StyleGAN2-ADA network

The invention discloses an esophageal endoscope image generation system and method based on an improved StyleGAN2-ADA network. The system comprises a generator and a discriminator. The generator comprises a mapping network and a synthesis network; the discriminator comprises a discrimination network, adopts an adaptive discriminator to enhance an ADA mechanism, and dynamically adjusts the image enhancement intensity according to the training accuracy of the discriminator; the method comprises the following steps: S1, fusing multi-vector features of a mapping network; s2, performing a multi-scale feature guided auxiliary classifier method on the discriminant network; s3, in the generative network, calculating a potential spatial distance loss and a comparison loss method; according to the method, the problems of non-uniform sample types and scarcity of data samples in medical images such as esophageal endoscopes are solved.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Backboard video generation method for real-time interactive digital human and related device

The invention provides a backplane video generation method for real-time interactive digital humans and a related device, and relates to the technical field of image generation, in particular to the technical field of artificial intelligence such as human-computer interaction, digital humans, end-cloud integration and large models. The method comprises the following steps: acquiring a plurality of original images presented by the same target person at different angles; generating a digital human image taking a pure color as a background based on the plurality of original images; generating prompt information based on a preset action expression demand and the digital human image; and inputting the prompt information into a preset video generation large model, and generating a digital person bottom plate video which enables a digital person corresponding to the target person to show a target action corresponding to the action expression demand. According to the method, efficient and low-cost generation of the digital human bottom plate video is realized.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Image processing method, training method of image evaluation model and electronic equipment

The invention provides an image processing method, a training method of an image evaluation model and electronic equipment, and relates to the field of computer technology and artificial intelligence. The method comprises the steps that image data and inquiry text data are acquired, the image data are generated by an image generation model, and the inquiry text data are used for describing the requirement for quality evaluation of the image data by adopting a natural language; respectively inputting the image data and the inquiry text data into a plurality of image evaluation models, and respectively carrying out quality evaluation on the image data from corresponding evaluation dimensions by utilizing the plurality of image evaluation models to obtain a plurality of evaluation results; and integrating the plurality of evaluation results to obtain a quality evaluation result of the image data. According to the method and the device, the technical problems of relatively low accuracy and meticulousness of image quality evaluation in related technologies are solved.
Owner:ZHEJIANG TMALL TECH CO LTD

Three-dimensional grounded video generation

Systems and methods are disclosed related to a 3D grounded video foundation model. A video generation method and system provide 3D conditioning information to a video diffusion model to improve generated video quality (object and temporal consistency) that is grounded in three dimensions (3D). The video generation method and system also enable precise camera control, cinematic effects, and scene editing. Video output corresponding to a set of camera specifications is generated for a scene from input image(s) including one or more images of a static scene or a sequence of images (video) for a dynamic scene. The input image(s) are used to calculate a 3D cache representing the scene. The 3D cache is rendered according to the set of camera specifications to produce a frame sequence and a mask sequence that identifies missing pixels in each frame. The frame sequence is encoded and masked to generate the output video.
Owner:NVIDIA CORP

Training a machine learning model to predict images representative of defects on a substrate

A method for training a prediction model to generate a high-resolution image representing defects on a substrate from a low-resolution image of the substrate. The method includes inputting a first image and a reference image of defects on a substrate, which are representative of images captured using different image capture conditions, to a neural network. The neural network is executed to generate a predicted image in response to the first image. A loss function that is indicative of a difference between a defect distribution in the predicted image and a defect distribution in the reference image is calculated and the neural network is modified based on the loss function. The neural network may be trained until the loss function is minimized.
Owner:ASML NETHERLANDS BV

Metalearning-based small sample cable protection layer overfire temperature-image generation method

The invention discloses a small sample cable protection layer overfire temperature-image generation method based on meta-learning, and the method comprises the steps: constructing a multi-task small sample data set through a small number of fire resistance test images, proposing a meta-learning type image generation model, and achieving the quick adaption to different protection systems through an internal and external circulation meta-learning strategy; inputting a target protection configuration, a fire condition and a temperature condition, wherein the meta-learning type image generation model can generate a fire appearance image of each protection layer at a corresponding temperature grade; further performing physical consistency judgment based on temperature-damage trend, cross-layer association and protection system configuration to ensure that a generation result conforms to a real damage rule, and performing quantitative index verification by calculating a structural similarity index, learning and sensing image block similarity and a Frechet Inception distance; according to the method, the multi-working-condition high-quality fire passing image can be generated under the condition of image scarcity, the fire resistance test cost is reduced, and data support is provided for temperature inversion and damage evaluation after a bridge cable fire disaster.
Owner:CHINA UNIV OF MINING & TECH +2

A Generalist Framework for Panoptic Segmentation of Images and Videos

Provided are systems and methods for performing panoptic segmentation of images and videos using a denoising diffusion model. The panoptic segmentation task is formulated as a conditional discrete data generation problem. This is achieved by learning a generative model for panoptic masks, for example treated as an array of discrete tokens, conditioned on an input image. The generative model can also be applied to video data by including predictions from past frames as an additional conditioning signal. This enables the model to learn to track and segment objects automatically across video frames.
Owner:GOOGLE LLC

Generative adversarial network architecture search method and system and image generation method

The invention discloses a generative adversarial network architecture search method and system and an image generation method, and belongs to the technical field of network architecture search. The searching method comprises the following steps: performing single-path sampling on a pre-constructed generator super network according to a parameter quantity constraint range to obtain an effective subnetwork; training the generator super-net by adopting a complexity adaptive learning rate optimization strategy to obtain a pre-trained generator super-net; adversarial training is carried out on the generator hypernet and the discriminator to obtain a pre-trained discriminator; generating a generator super-network candidate architecture through a genetic algorithm in the early stage of the evolution stage and through a covariance matrix self-adaptive evolution strategy in the later stage of the evolution stage; and performing multi-target non-dominated sorting on the generator super-network candidate architecture, updating an effective sub-network and keeping a Pareto optimal solution to obtain a searched optimal generator architecture and further obtain a searched optimal generative adversarial network architecture. The method not only ensures the search quality, but also improves the calculation efficiency.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Image generation method, object image generation method, image generation model training method, and cloud training platform

Embodiments of the present description provide an image generation method, an object image generation method, an image generation model training method, and a cloud training platform. The image generation model training method is applied to the field of deep learning, and comprises: acquiring a description text of target content; using a target image generation model to perform inference denoising on a random noise feature on the basis of a description text, and generate a plurality of target images corresponding to the description text of target content, wherein the target image generation model is obtained by training an initial image generation model on the basis of a plurality of prediction images and a plurality of sample original images of sample content, the plurality of prediction images are generated by performing inference denoising on a plurality of sample noise features on the basis of sample description texts of the plurality of sample original images, and the plurality of sample noise features are obtained by performing noise injection and feature encoding on the plurality of sample original images. The model learns to generate a plurality of content-related images, thereby improving generalization ability and training flexibility, reducing training complexity, and improving image generation efficiency and accuracy.
Owner:ALIBABA (CHINA) CO LTD

Real-time sesame seed candy forming defect detection method and device based on AI vision

The invention relates to the technical field of AI vision, in particular to a sesame seed candy forming defect real-time detection method and device based on AI vision. The method comprises the following steps: respectively collecting multimode images of qualified sesame seed candies, generating a qualified characteristic fingerprint set, and calculating sugar body light transmission uniformity and sesame adhesion density as a domain parameter set; constructing a probability distribution model, and setting an anomaly judgment threshold value and a domain parameter early warning threshold value; collecting a multi-mode image of the to-be-detected sesame seed candy in real time, extracting a to-be-detected feature fingerprint, and calculating the light-transmitting uniformity of a to-be-detected candy body and the sesame adhesion density; calculating a comprehensive abnormal score, and judging a defect; and capturing a low-confidence sample based on the comprehensive anomaly score, obtaining an artificial correction feedback sample, and updating a probability distribution model, an anomaly judgment threshold value and a domain parameter early warning threshold value by utilizing the feedback sample through online incremental learning. According to the invention, the detection cost of a high-yield production line can be reduced, and the model deployment period is shortened.
Owner:XIAOGAN HONGLONG MATANG RICE WINE CO LTD