Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

46 results about "Image translation" patented technology

Image translation refers to a technology where the user can translate the text on images or pictures taken of printed text (posters, banners, menu list, sign board, document, screenshot etc.).This is done by applying optical character recognition (OCR) technology to an image to extract any text contained in the image, and then have this text translated into a language of their choice, and the applying Digital image processing on the original image to get the translated image with a new language. Image traslation is related to machine translation.

Grayscale image translational motion blurring recovery method and system based on optical flow guidance

The invention discloses a grayscale image translation motion blur recovery method and system based on optical flow guidance, and belongs to the technical field of image processing. A target area is determined based on a video stream H; calculating a translational motion vector field matrix between every two adjacent frames in the video stream H by adopting an optical flow method, obtaining a translational motion vector located in a target area, further constructing a track representation matrix for reflecting a motion track of a motion target generating translational motion, normalizing the track representation matrix to serve as a blurring kernel, and obtaining a motion vector field matrix; and translational motion blur of the target area is removed in a targeted manner. According to the method, the high-density time sampling information provided by the video stream H with the high frame rate and the low signal-to-noise ratio is utilized, and the complex translational motion track of the blurred image B with the low frame rate and the high signal-to-noise ratio in the exposure period can be accurately captured. The fuzzy kernel synthesized based on the accurate translational motion trails can highly approach a real physical fuzzy process, and an image with a high signal-to-noise ratio and high definition can be reconstructed.
Owner:HUAZHONG UNIV OF SCI & TECH

Visual stereoscopic enhancement method, device, equipment and storage medium

The present disclosure provides a visual stereoscopic enhancement method, device, equipment and storage medium, relates to the technical field of image processing, in particular to the technical field of image mask, image translation, image fusion and the like, and can be applied to the scenes of portrait visual stereoscopic enhancement, virtual digital person, image or video space sense enhancement and the like. The specific implementation scheme comprises the following steps: according to the configured light direction and light angle, a mask is translated by a target distance in a first mask image corresponding to a target object in a target image to obtain a second mask image; after the first mask image and the second mask image are subtracted pixel by pixel, the regions with a translation increment less than 0 and greater than 1 are discarded to obtain a shadow image; after the shadow image and the first mask image are added pixel by pixel, the target image is multiplied pixel by pixel to obtain an image after stereoscopic enhancement of the target object. The present disclosure can simply and quickly realize visual stereoscopic enhancement and improve visual stereoscopic sense.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Adaptive machine learning-based lesion identification

An adaptable deep learning method is provided that delivers sound hepatic lesion identification in NETs, while significantly reducing human effort for data annotation and improving model generalizability for PET image quantification. A region-guided GAN (RGGAN) model conducts image-to-image translation between list-mode simulated PET images and real-world clinical data, while preserving semantic content of interest, e.g., lesions. The RG-GAN model is integrated with a lesion detection model into an end-to-end, unified framework for joint-task learning, such that the two models can benefit from each other. The RG-GAN translates the list-mode simulated data into real world-style images, which appear to be drawn from the real clinical PET image dataset, and feeds the translated images into the lesion detection model for training. In order to deal with the limited diversity of list mode-simulated PET image data, a specific data augmentation module is incorporated into the unified framework to improve model training.
Owner:THE REGENTS OF THE UNIVERSITY OF COLORADO

Oblique view angle rotator vibration displacement measurement method and system based on image translation

The invention discloses an image translation-based inclined view angle rotator vibration displacement measurement method. The method comprises the following steps of: performing cutting operation on an acquired rotator vibration image under a standard view angle and an acquired rotator vibration image under L random view angles; randomly extracting a preset number of random view angle cutting images and standard view angle cutting images of matched frames from the L random view angles, and taking the standard view angle cutting images as labels of the random view angle cutting images to construct a data set; training and verifying the image translation model through the training set and the verification set to obtain a trained image translation model; selecting one of the obtained rotating body vibration images of the preset time length under the L random view angles, inputting the selected rotating body vibration images into a trained image translation model, and obtaining a standard view angle image of a frame number corresponding to the preset time length; carrying out image splicing operation according to the standard view angle images with the frame number corresponding to the preset duration to obtain a spliced image; and performing edge detection on the spliced image, extracting an edge contour of the rotating body, obtaining sub-pixel-level displacement data, and converting the sub-pixel-level displacement data into a displacement value under a world coordinate system. According to the invention, the vibration displacement of the rotor can be effectively and accurately extracted from a plurality of shooting angles, and a more reliable and accurate scheme is provided for the measurement of the vibration displacement of the rotating body.
Owner:KUNMING UNIV OF SCI & TECH

Methods and systems for image-to-image translation of microscopy images

PCT designated stageWO2026055610A1Image enhancementImage analysisRadiologyFluorophore
Described herein are methods and systems for training image to image translation machine-learning models, wherein the image to image translation machine-learning models are trained using unpaired images. The trained image to image translation machine-learning models can be used to digitally identify markers of a fluorescent image depicting a plurality of markers in a single color channel from a single fluorophore and / or enhance a fluorescence image.
Owner:DIFFINE LLC

A privacy protection heterogeneous federated learning method and system based on image translation

The application discloses a privacy protection heterogeneous federated learning method and system based on image translation, which are corresponding solutions, in the solutions: a generated model with local user data semantic retention capability is optimized through an image translation server, a local user performs data enhancement on a local heterogeneous data set by using the generated model, so that the local data presents a state of uniform distribution of class labels, and a classification model of the local user is trained by using transformed data and enhanced data, so that the global model accuracy and the privacy protection capability are balanced. Moreover, there is no interaction between the image translation server and the data classification server, and the image translation server is in a trusted computing environment, so that the privacy protection capability is ensured.
Owner:HEFEI UNIV OF TECH

Systems and methods for preprocessing immunocytochemistry images for machine learning image-to-image translation

ActiveUS12597142B2Image enhancementImage analysisCytochemistryRadiology
Methods and systems for pre-processing immunocytochemistry images for machine learning are provided. The pre-processing method includes receiving paired positive and negative multi-protein images and segmenting stains of the paired multi-protein images. The method also includes labelling each image pixel corresponding to a cell of the paired multi-protein images and translating image pixel information of the labelled image pixels to tabular form to generate a table of cell coordinates and geometrical characteristics for each of the paired multi-protein images. The method further includes generating two cell-paired tables for each of the paired multi-protein images based on Euclidean distance-based pairing prior to input for the machine learning, where the Euclidean distance-based pairing is based on the cell coordinates in the table of cell coordinates and geometrical characteristics for each of the paired multi-protein images.
Owner:CITY UNIVERSITY OF HONG KONG

Data processing methods and related devices

PendingCN122311237AData scienceImage based
This application discloses a data processing method and related apparatus, which enables the trained translation model to accurately generate descriptive information for an image to be translated, and to accurately translate the text to be translated contained in the image based on the descriptive information and the image to be translated. The descriptive information can accurately describe the image content in the image to be translated. By combining the descriptive information, the translation model can more accurately analyze and understand the image content, and thus more accurately analyze the translation method that should be selected under the image content corresponding to the image to be translated when translating the text to be translated. This makes the final translation result more consistent with the image content of the image to be translated, and improves the accuracy and rationality of image translation.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and system for unsupervised deep representation learning based on image translation

A system for unsupervised deep representation learning based on image translation is provided. The system includes an image translation transformation module used for performing a random translation transformation on an image and generating an auxiliary label; an image mask module connected with the image translation transformation module and used for applying a mask to the image after translation transformation; a deep neural network connected with the image mask module and used for predicting an actual auxiliary label of the image after the mask is applied and learning the deep representation of the image; a regression loss function module connected with the deep neural network and used for updating parameters of the deep neural network based on a loss function; and a feature extraction module connected with the deep neural network and used for extracting the representation of the image.
Owner:ZHEJIANG NORMAL UNIV

Method for few-shot unsupervised image-to-image translation

ActiveUS12675701B2Data setObject Class
A few-shot, unsupervised image-to-image translation (“FUNIT”) algorithm is disclosed that accepts as input images of previously-unseen target classes. These target classes are specified at inference time by only a few images, such as a single image or a pair of images, of an object of the target type. A FUNIT network can be trained using a data set containing images of many different object classes, in order to translate images from one class to another class by leveraging few input images of the target class. By learning to extract appearance patterns from the few input images for the translation task, the network learns a generalizable appearance pattern extractor that can be applied to images of unseen classes at translation time for a few-shot image-to-image translation task.
Owner:NVIDIA CORP

Knowledge distillation for semantic relation preservation in image-to-image translation

GAN-based generators are useful for performing image-to-image translations. GAN models have large storage sizes and resource requirements, making them too large to be directly deployed on mobile devices. The system and method define a student GAN model with a student generator scaled down from the teacher GAN model (and the generator) using knowledge distillation. A semantic relation knowledge distillation loss is used to transfer semantic knowledge from the teacher's intermediate layer to the student's intermediate layer. The student generator, thus defined, is stored and executed by mobile devices such as smartphones and laptops to provide augmented reality experiences. Effects are simulated on images, including makeup, hair, nails, and age simulations.
Owner:LOREAL SA

Pedestrian re-identification method and system based on multi-modal fusion

The invention discloses a pedestrian re-identification method based on multi-modal fusion. The method comprises the following steps: acquiring a visible light image and an infrared image containing a target pedestrian and carrying out format preprocessing on the visible light image and the infrared image; performing feature extraction on the visible light image after format preprocessing by using a ResNet50 modal encoder to obtain RGB image features; the infrared image after format preprocessing is converted into a pseudo RGB image through an unsupervised image translation method, and pseudo RGB image features are extracted; fusing the RGB image features and the pseudo RGB image features by using an attention mechanism to obtain preliminary fusion features; performing feature enhancement on the preliminary fusion feature by using a Transform encoder to obtain an enhanced fusion feature; and comparing the enhanced fusion features with candidate image features in a candidate database to obtain an identification result. According to the invention, the accuracy of cross-modal recognition in pedestrian re-recognition is improved.
Owner:NO 15 INST OF CHINA ELECTRONICS TECH GRP

Method and apparatus for image translation using diffusion model

A method and apparatus for translating an image using a diffusion model are provided. According to an embodiment, the method of translating a synthetic aperture radar (SAR) image into an electro-optical (EO) image based on additional conditioning data including map data that provides spatial information about a target region is provided.
Owner:KOREA ADVANCED INST OF SCI & TECH

End-to-end document image translation method and device fusing layout information

The application provides a page information fusion end-to-end document image translation method and device, which comprises the following steps: obtaining a character recognition result of a document image to be translated, wherein the character recognition result comprises a plurality of words in the document image to be translated and two-dimensional coordinate information of each word, and the two-dimensional coordinate information is determined based on pixel values of the document image to be translated; obtaining a first feature vector based on corresponding text of each word, two-dimensional coordinate information of each word and one-dimensional position information of each word, wherein the one-dimensional position information is used to indicate a position of the word in a word sequence, and the word sequence is used to indicate a one-dimensional sequence formed by all words recognized from the document image to be translated; and decoding the first feature vector to obtain a translation text corresponding to the document image to be translated. The end-to-end document image translation method provided by the application effectively improves the document translation effect.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

System and method of image-to-image translation in diffusion seed space

A computer-implemented method of image-to-image translation that, when executed by data processing hardware, causes the data processing hardware to perform operations comprising applying an inversion technique to an input image to generate a source-domain seed, translating the source-domain seed to a target-domain seed using a translation module, and sampling the target-domain seed to generate a denoised code.
Owner:YISSUM RESEARCH DEVELOPMENT COMPANY OF THE HEBREW UNIVERSITY OF JERUSALEM LTD +1

Parameter-efficient and resolution-robust network architectures for image-to-image translation

One embodiment provides a method of using a computing device for image-to-image translation including accessing an image file containing a first amount of data. The computing device inputs the image file into a convolutional neural network (CNN). The CNN includes multiple Fourier layers. Each Fourier layer includes a Fourier transform, a linear feature transformation in a frequency domain and an inverse Fourier transform. Each linear feature transformation in the frequency domain is shared by different frequency components to reduce a number of parameters. The CNN outputs an output image file that includes contents that are translated from the input image file.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Interpreting machine learning models using image translation

This application relates to interpreting machine learning models using image transformation. A system and method for identifying visual features that influence predictive models are provided. This technique employs an image transformation function to introduce visual features into an image to create a modified image that can be fed into a predictive model. When the predictive model generates a prediction for a given image that differs from the prediction generated for a modified version of the image, a further modified version exaggerating the introduced visual features can then be created using the image transformation function. Therefore, this technique helps identify visual features that influence the predictive model, enabling the understanding of the model's conclusions and allowing for further study and testing of these visual features.
Owner:GOOGLE LLC

Methods and systems for image-to-image translation of microscopy images

Described herein are methods and systems for training image to image translation machine-learning models, wherein the image to image translation machine-learning models are trained using unpaired images. The trained image to image translation machine-learning models can be used to digitally identify markers of a fluorescent image depicting a plurality of markers in a single color channel from a single fluorophore and / or enhance a fluorescence image.
Owner:SARPEDA INC

Space target radar image data generation method based on low-light enhancement correction

ActiveCN121391692BImage enhancementBiological modelsHadamard productImaging data
The application relates to a space target radar image data generation method based on low-light enhancement correction. The method comprises the following steps: designing a content transfer decomposition network, decomposing low-light features into illumination-independent reflection components and adaptive illumination components through Hadamard product constraints in a latent space, proposing a learnable intensity compression function, dynamically generating a latent space mask through gradient consistency constraints and sparse regularization, embedding the mask into a reverse denoising process to realize directional repair of degradation areas, combining the generation ability of a diffusion model and the guidance of a physical prior, reconstructing a diffusion path based on a bidirectional diffusion principle, generating the physical consistency of the process through endpoint binding constraints, adaptively adjusting the diffusion step, introducing a cyclic implicit iteration mechanism, gradually refining the generation result through a recurrent neural network module, and realizing accurate optical-ISAR image translation.
Owner:NAT UNIV OF DEFENSE TECH

Pet behavior recognition and image translation method based on image features

The invention provides a pet behavior recognition and image translation method based on image features. The method comprises the following steps: preprocessing a video frame; key points, appearance and motion features are extracted through a multi-branch network, and fusion time sequence features are obtained through skeleton communication and cross-branch fusion; comparing the features at each moment with a pre-learned behavior code book to form a code string, and outputting a behavior label and confidence through time sequence decoding; and generating a mask according to the key points, inputting the original image, the condition vector obtained by the code string and the mask into a condition generator, and outputting a behavior image with an interpretation element. Low-power-consumption real-time reasoning is achieved through quantification and pruning, privacy security is guaranteed through local storage and selectable reporting feature indexes, and the method is suitable for family cameras and mobile phone terminals, easy and convenient to deploy and maintain and high in robustness.
Owner:GUANGZHOU YUECHUANGFU TECH CO LTD

Cross-border e-commerce live broadcast synchronous translation method and system based on image analysis

The invention belongs to the technical field of image processing, and particularly relates to a cross-border e-commerce live broadcast synchronous translation method and system based on image analysis, and the method effectively eliminates irrelevant background interference and provides high-quality input for subsequent processing through the rapid acquisition of a translation video stream and the precise cutting of commodity information characters; an end-to-end image translation model integrating a convolutional neural sub-network and a translation sub-network is constructed, the defect that convolutional feature extraction and semantic mapping are disjointed in the prior art is overcome, and through an end-to-end collaborative architecture, error conduction of a traditional cascade (OCR recognition first and translation second) method is avoided, and the accuracy of semantic mapping is improved; in the preprocessing step, effective features are screened through a space-time attention weight formula in the preprocessing step, the effective features are aggregated through a multi-frame fusion formula, and the features are optimized and complemented through a generator, so that the problems of a fuzzy region and inclination of a local image are solved, and the accuracy of subsequent translation is improved.
Owner:GUANGZHOU JINGCUI EDUCATION TECH CO LTD

A method, apparatus and related medium for video image jitter detection and stabilization.

This invention discloses a method, apparatus, and related medium for video image jitter detection and stabilization. The method includes: acquiring a target video image to be jitter detected; performing semantic segmentation on the target video image using deep learning-based image semantic segmentation technology to obtain the region of interest (ROI); extracting key points within the ROI using a corner detection algorithm; calculating the optical flow map of any two consecutive frames of the target video image using optical flow method, and obtaining the displacement information of the key points in the two consecutive frames based on the optical flow map; calculating the image translation matrix and image rotation matrix based on the displacement information to obtain jitter information; comparing the jitter information with a preset jitter threshold; and if jitter occurs in the current frame image, performing a reverse translation and / or rotation on the current frame image based on the jitter information to cancel the jitter. This invention combines deep learning semantic segmentation technology and optical flow method, which can improve the accuracy of jitter detection and achieve good jitter cancellation effect.
Owner:SHENYAN ARTIFICIAL INTELLIGENCE TECH (SHENZHEN) CO LTD

A dual-branch generative network heterogeneous image change detection method based on content and attribute decoupling

PendingCN122347740AManual annotationImage translation
The application provides a dual-branch generative network heterogeneous image change detection method based on content and attribute decoupling. The method constructs a joint optimization framework of an image translation network and a change detection network, generates an intermediate translation image through cross-domain reconstruction, constructs a change response map using cross-feature difference, and introduces a confidence discrimination mechanism based on information entropy to realize momentum update of a soft label, thereby forming a self-adaptive optimization closed loop. The method improves the change detection precision and stability under the condition of no accurate manual annotation, realizes collaborative enhancement of the translation and detection tasks, and is suitable for remote sensing image change analysis and other scenes.
Owner:HARBIN ENG UNIV

An optical-sar image translation method based on a scattering feature enhancement strategy

ActiveCN120634889BScattering characteristics are effectively restoredAdd nonlinearityImage enhancementImage analysisEdge extractionOptical image
This invention discloses an optical-SAR image translation method based on a scattering feature enhancement strategy. The method proposes an image translation framework, SFEG, which mainly includes an SCG module and a GFE module. The SCG module first extracts edges from the input image, generates an edge mask by combining target annotation information, and models the scattering field using prior knowledge of the SAR target scattering characteristics. Finally, the generated scattering field and the corresponding optical image are input into the generator. This module can effectively restore the scattering characteristics of the target in the input image. The GFE module includes a channel attention branch and a spatial attention branch, which can enhance the input features from the channel and spatial dimensions respectively, and combine nonlinear functions to achieve feature recombination, further enhancing the nonlinear and refined output capabilities of SFEG. This module can improve the actual effect of generating SAR images.
Owner:HARBIN INST OF TECH

Image translation method, image translation aparatus and electronic device

Embodiments of this application provide an image translation method, image translation apparatus and electronic device. The method includes: obtaining a source image, where the source image includes a source text block with one or multiple source text lines; generating a target image based on geometric properties of the source text block, where the target image comprises one or multiple target text lines translated from the one or multiple source text lines and the geometric properties of the source text block comprises: a shape of the source text block and length difference(s) corresponding to the one or multiple source text lines. According to the technical solution, the image translation method is more efficient to adapt to different conditions.
Owner:HUAWEI TECH CO LTD +1

Artificial intelligence-based foreign language bill image translation method, device and medium

ActiveCN120913229BData setLinguistic model
The application provides a foreign language bill image translation method, device and medium based on artificial intelligence, and belongs to the technical field of data information. The foreign language bill image translation method comprises the following steps: acquiring a foreign language bill image, inputting the foreign language bill image into an image recognition model, and outputting bill text information and position coordinate information; inputting the bill text information into a pre-trained bill translation special model to obtain Chinese bill information; inputting the information into a large language model, optimizing the Chinese bill information according to a preset bill question and answer prompt word; using the position coordinate information, arranging and combining the optimized Chinese bill information according to a Chinese bill format, and drawing a Chinese bill image; using a key information extraction model to extract bill key information from the Chinese bill information; and displaying the Chinese bill image and the bill key information. The application can solve the problem that the prior art cannot accurately translate bill data sets, thereby causing difficulty in training accurate and effective bill information extraction.
Owner:INSPUR GENERSOFT CO LTD

Contrast learning based remote sensing image cross-domain simulation translation method and device

The application discloses a remote sensing image cross-domain simulation translation method and device based on contrast learning, a medium and equipment. The method comprises the following steps: acquiring a source image to be translated and a target image to be translated; inputting the source image to be translated and the target image to be translated into an image translation model; wherein the image translation model comprises a generator and a discriminator, and the generator comprises an encoder and a decoder; inputting the source image to be translated and the target image to be translated into the encoder to obtain source encoding features and target encoding features respectively; inputting the source encoding features and the target encoding features into the decoder to obtain a source pseudo sample image and a target pseudo sample image respectively; inputting the source pseudo sample image into the discriminator for identification; wherein the image translation model is obtained by training based on a content contrast loss function, a style contrast loss function and an adversarial loss function constructed based on the source pseudo sample image and the target pseudo sample image. The application can avoid the problem of image detail distortion caused by excessive stylization of the image.
Owner:WUHAN UNIV +1

A subway station space layout division method based on multi-stage mechanical simulation and region growing algorithm

ActiveCN120974594BAlgorithmMechanical models
This invention proposes a spatial layout partitioning method for subway stations based on multi-stage mechanical simulation and region growing algorithms. First, the topological relationships and areas of the functional units of the subway station are transformed into mechanical model parameters. An adaptive mechanical simulation, comprising three stages—global layout, fine-tuning, and stable convergence—is then performed. Based on the mechanical simulation results, a multi-region competitive growing algorithm is used to generate an initial spatial layout that meets the area ratio requirements of each functional zone. Finally, image translation is used to obtain a simplified spatial layout diagram, ensuring the rationality and integrity of the subway station's functional zoning. This method, through the organic combination of multi-stage mechanical simulation and region growing algorithms, achieves the automatic generation of spatial layout schemes from subway station functional requirements. The final output layout scheme not only meets the area ratio requirements of each functional zone but also maintains the overall coordination and functional continuity of the subway station space, providing a scientific and efficient preliminary spatial partitioning scheme for subway station design.
Owner:HARBIN INST OF TECH