Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

656 results about "Image representation" patented technology

Image-level representation is a (numerical) way to represent an image without a direct pixel representation. For example, one could represent an image by its histograms of luminance and chroma values, or by its Fourier transform, or by any other statistical measure. Such a representation helps compare images or detect specific features.

Image aspect ratio enhancement using generative ai

A method includes adding an outpaint mask to an image to generate a masked image. The method also includes processing the image using an encoder neural network to generate an image representation of the image in a latent space. The method further includes processing the masked image using a convolution neural network and adding the image representation to generate an image embedding. The method also includes processing the image representation and the image embedding using at least one of a diffusion model and an interpolation process to generate a noisy latent image representation. The method further includes using a large language model to contextualize an outpainting prompt. The method also includes denoising the noisy latent image representation based on the contextualized outpainting prompt to generate a denoised latent image representation. In addition, the method includes processing the denoised latent image representation using a decoder neural network to generate an outpainted image.
Owner:SAMSUNG ELECTRONICS CO LTD

Cross-Modal Adapters for Machine-Learned Sequence Processing Models

A machine-learned system for aligning textual and image representations prior to input to a sequence processing model is described. The system includes a machine-learned image embedding model configured to receive image data and generate one or more image embeddings and a machine-learned text embedding model configured to receive text data and the one or more image embeddings and generate one or more text embeddings. The system includes a machine-learned cross-modal adapter configured to generate one or more text tokens aligned with one or more image tokens based at least in part on aligning data associated with the one or more text embeddings and the one or more image tokens. The system includes a machine-learned sequence processing model configured to generate an output based at least in part on the one or more text tokens and the one more image tokens.
Owner:GOOGLE LLC

System and method for multi-level signal cyclic loop image representations for measurements and machine learning

A system includes an input to receive a digital waveform signal, a memory, and one or more processors configured to execute code to cause the one or more processors to: generate a horizontal ramp sweep signal based on the digital waveform signal; receive a selection input to identify a segment of the digital waveform signal; gate the horizontal ramp sweep signal and the digital waveform signal based on the selection input to produce cyclic loop image data for the segment of the digital waveform; store the cyclic loop image data in the memory; and provide the cyclic loop image data as one or more inputs into a machine learning system. A method of waveform classification using a cyclic loop image includes receiving an input waveform, receiving a selection of a segment of the input waveform, transforming the segment of the input waveform into cyclic loop image data, the transforming comprising generating a horizontal ramp sweep signal based on edge transitions in the input waveform, and storing the cyclic loop image data in a memory; and sending the cyclic loop image data to a machine learning system to determine an attribute of the input waveform.
Owner:TEKTRONIX INC

Fault identification method combining enhanced convolution and multi-scale attention

The invention discloses a fault identification method combining enhanced convolution and multi-scale attention, and relates to the technical field of seismic interpretation and reservoir prediction. The method comprises the following steps: firstly, on the basis of seismic data high-resolution enhancement processing, generating a contrast-invariant four-channel image representation by using a local intensity sequence transformation filter; then, constructing an enhanced convolution encoder in combination with differential convolution to extract fault edge information, and meanwhile, constructing a convolution Mama attention module by fusing multi-scale separable convolution and a Mama framework to extract space and global information; and finally, constructing a feature fusion module to fuse the shallow features of the jump connection with the deep features of the decoder, and outputting a fault recognition result through full connection layer mapping. Compared with the prior art, the method combines the advantages of the local intensity sequence transformation, the attention mechanism, the convolutional network and the Mama architecture, can more effectively learn the fault feature information in the seismic data, and improves the fault recognition precision.
Owner:CHENGDU UNIVERSITY OF TECHNOLOGY

Point cloud compression with supplemental information messages

A system comprises an encoder configured to compress attribute information and / or spatial for a point cloud and / or a decoder configured to decompress compressed attribute and / or spatial information for the point cloud. To compress the attribute and / or spatial information, the encoder is configured to convert a point cloud into an image based representation. Also, the decoder is configured to generate a decompressed point cloud based on an image based representation of a point cloud. Additionally, an encoder is configured to signal and / or a decoder is configured to receive a supplementary message comprising volumetric tiling information that maps portions of 2D image representations to objects in the point. In some embodiments, characteristics of the object may additionally be signaled using the supplementary message or additional supplementary messages.
Owner:APPLE INC

Human body posture estimation method based on radar point cloud imaging and multi-dimensional feature fusion

The invention discloses a human body posture estimation method based on radar point cloud imaging and multi-dimensional feature fusion, and relates to the cross technical field of computer vision and radar signal processing, and the method comprises the following three key technical links: firstly, improving the target resolution through spatial energy distribution estimation; reconstructing target three-dimensional space distribution by using the positive correlation between radar signal energy and a target reflection area and adopting a least square estimation algorithm; secondly, constructing a structured multi-dimensional point cloud matrix, and converting sparse radar point cloud into high-information-density imaging representation through a distance-speed hierarchical sorting strategy; and finally, designing a multi-dimensional feature fusion attitude estimation network, integrating three-dimensional convolution, a multi-head attention mechanism and a gating circulation unit, and realizing collaborative extraction of spatio-temporal features. According to the method, the problems of sparse target features, noise sensitivity and poor universality in traditional millimeter wave radar attitude estimation are solved.
Owner:DALIAN MARITIME UNIVERSITY

Method and system for determining a treatment strategy for a vessel with multiple or diffuse lesions

PendingUS20250217985A1Image enhancementImage analysisRadiologyDiffuse Lesion
Non-invasively determining a treatment strategy for a blood vessel by obtaining a prediction of a total pressure drop value in the blood vessel based on features generated from a plurality of images of the blood vessel, the images representing all anatomic parts of the blood vessel, and calculating a contribution of one or more portion(s) of the blood vessel to the total pressure drop value. Based on the calculated contribution, a new simulated total pressure drop value of the blood vessel is calculated by neutralizing the contribution of the one or more portion(s) to the total pressure drop value to determine if the new simulated pressure drop improves, thereby indicating portions which should be treated to restore healthy pressure drop values.
Owner:MEDHUB LTD

Systems and methods for identifying brands utilized in website phishing campaigns

A computer-implemented method for identifying brands utilized in website phishing campaigns may include (i) capturing a website screenshot including visual elements representing a potential phishing vulnerability, (ii) transforming, utilizing a deep learning model, the website screenshot into an image representation including embeddings, (iii) determining whether the transformed website screenshot matches a dataset including reference transformed website screenshots representing previously identified brands utilized in phishing campaigns, (iv) clustering, upon determining a mismatch between the transformed website screenshot and the dataset, the transformed website screenshot with other transformed website screenshots sharing the visual elements representing the potential phishing vulnerability and one or more visual similarities, and (v) performing, based on the clustering, a security action that protects against potential phishing attacks by extracting brand information for adding to the dataset. Various other methods, systems, and computer-readable media are also disclosed.
Owner:GEN DIGITAL INC

Method, device, and product for retrieval

The present disclosure provides a method, a device, and a product for retrieval. The method includes acquiring context information related to an image and determining a representation of the image based on image data and the context information of the image, where the context information includes at least one of environment parameters, user behavior data, time elements, or field metadata. The method further includes encoding the representation as an image vector in a high-dimensional vector space and storing it into an image vector database. When retrieval is performed, a query that includes text information and that is for the image vector database is received, and an image associated with the text information is determined from the image vector database. The method according to the present disclosure can improve accuracy and efficiency for image retrieval.
Owner:DELL PROD LP

Multistage cross-modal alignment method based on comparative learning

The invention discloses a multi-level cross-modal alignment method based on comparative learning, which is used for improving the accuracy and efficiency of multi-modal sentiment analysis. According to the method, a RoBERTa model and a Vision Transform model are used for coding a text and an image respectively, and text representation and image representation are obtained. The global cross-modal alignment module aligns the representation of the text and the representation of the image by adopting a comparative learning technology to enhance the consistency between the text and the representation of the image. Furthermore, through a local cross-modal alignment module, a cross-attention mechanism is used to perform fine-grained alignment on the text and image representations to identify smaller, more specific semantic units in the associated image and text. According to the method, a multi-task learning framework is adopted to integrate cross-modal information from texts and images, sequence tag prediction is carried out through a conditional random field, and terms and emotions in the aspects are recognized and classified. Experimental results show that the performance of the method on a Twitter-2015 data set and a Twitter-2017 data set is superior to that of an existing single-mode model and an existing multi-mode model, and the performance of multi-mode sentiment analysis is effectively improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Bit stream structure for compressed point cloud data

A system comprises an encoder configured to compress attribute information and / or spatial information for a point cloud and / or a decoder configured to decompress compressed attribute and / or spatial information for the point cloud. To compress the attribute and / or spatial information, the encoder is configured to convert a point cloud into an image based representation. Also, the decoder is configured to generate a decompressed point cloud based on an image based representation of a point cloud. In some embodiments, a bit stream structure may be used to communicate compressed point cloud data. The bit stream structure may include point cloud compression network abstraction layer (PCCNAL) units that enable use of groups of frames (GOFs), frame, and sub-frame signaling of patch information. Such a bit stream structure may permit low delay streaming and random access reconstruction of point clouds amongst other applications.
Owner:APPLE INC

3D printing process defect monitoring method and system based on multi-modal large model

The invention discloses a 3D printing process defect monitoring method and system based on a multi-modal large model. The method comprises the steps that S1, before related operation is carried out, corresponding preposition work needs to be completed, wherein the preposition work comprises knowledge graph creation, retrieval enhancement generation, reasoning and action framework and fine adjustment of the large model; s2, collecting and preprocessing industrial camera cluster data; s3, according to the preprocessed image, segmenting the image, constructing a region adjacency graph, finding an optimal region combination by applying a random algorithm, and combining some regions to obtain image representation; s4, calling the multi-mode large model subjected to pre-training and specific field fine adjustment to realize real-time defect monitoring in the 3D printing process; s5, a detection report is given according to a defect detection result, user interaction intelligent consultation is opened, and the model is updated and adjusted according to printing data; according to the method, the problems of standard, efficiency, trust and supervision of defect monitoring in the 3D printing process can be effectively solved.
Owner:BEIJING HENGCHUANG ADVANCED MATERIALS & ADDITIVE MFG INST CO LTD +1

Soil plastic detection method and system based on multi-modal fusion and deep learning

The invention relates to the field of soil substance detection, and discloses a soil plastic detection method and system based on multi-modal fusion and deep learning, and the method comprises the following steps: obtaining one-dimensional hyperspectral data of a to-be-detected soil sample; converting the one-dimensional hyperspectral data based on an image conversion algorithm to obtain two-dimensional image representation; performing feature extraction on the one-dimensional hyperspectral data based on a first preset algorithm to obtain spectral features; performing feature extraction on the two-dimensional image representation based on a second preset algorithm to obtain image features; performing multi-view probability fusion on the spectral features and the image features to obtain fusion features; and inputting the fusion features into a pre-trained double-path attention residual convolutional network model for classification to obtain a detection result of the micro-plastics in the to-be-detected soil sample. According to the method, the detection accuracy is improved, the model robustness is enhanced, and low-concentration detection is realized.
Owner:SICHUAN AGRI UNIV

General traffic image generation method for solving unbalanced network traffic classification

The invention relates to the technical field of computers, in particular to a general traffic image generation method for solving unbalanced network traffic classification, which comprises the following steps of: arranging and combining original network traffic into a session stream according to a time sequence based on a session stream mode, and converting the session stream into an image by taking a data packet as a unit. Converting the effective load of each data packet into a grayscale image; the SCGAN is used for training; eliminating noise by adopting a convolution noise reduction auto-encoder, and performing high-definition reconstruction on the generated flow sample; and combining a minority of types of traffic image samples generated through high-definition reconstruction with the original real traffic samples. According to the method, when the data packets are converted into the traffic images, the time sequence dependency relationship of the network traffic is reserved, the structural features among the data packets in image representation are also reserved, a balanced new network traffic data set is constructed, the authenticity and diversity of the data set are kept, and the generalization ability and effect of the model are improved.
Owner:GUANGDONG UNIV OF SCI & TECH

Image retrieval method and device, computer equipment and storage medium

The invention discloses an image retrieval method and device, computer equipment and a storage medium, belongs to the technical field of artificial intelligence, and is provided with a retrieval system applied to an insurance marketing scene image. According to the method and the device, the to-be-processed image is segmented and coded, the content features of the image are combined with the position information to form high-quality image representation, and the corresponding image index information is generated by utilizing the preset image index generator, so that efficient indexing and retrieval of the image are realized. In the encoding process, local features of the image are reserved, spatial position information is fused, discrimination and uniqueness of image indexing are improved, generated image indexing information and an original image are stored in an image retrieval database in an associated mode, it is ensured that the image retrieval process has high matching precision and response speed, and the image retrieval efficiency is improved. The image retrieval accuracy and processing efficiency are effectively improved, and the method is suitable for large-scale image library management and rapid retrieval.
Owner:PING AN TECH (SHENZHEN) CO LTD

Methods and systems for multiple instance learning of tissue sample images

PendingUS20250356486A1Image enhancementImage analysisFeature vectorNeedle core biopsy
Methods for multiple instance learning of tissue sample images are described. The methods may comprise, for example, receiving a whole slide image from a needle core biopsy sample from a subject; identifying a tissue region in the whole slide image; selecting a set of image patches from the identified tissue region; resampling the set of image patches at a plurality of image scales to generate a plurality of resampled image patches; generating image representations for the plurality of resampled image patches; extracting feature vectors based on the image representations; providing the feature vectors as input to a trained machine learning model configured to predict a gene alteration state; and outputting the predicted gene alteration state for the needle core biopsy sample for the subject.
Owner:FOUNDATION MEDICINE INC

Dirt identifying and cleaning method, device and equipment

The invention is suitable for the field of intelligent jet cleaning, and provides a dirt recognition cleaning method, device and equipment, and the dirt recognition cleaning method comprises the steps: obtaining an original image corresponding to target dirt and a sound reflection signal corresponding to the target dirt; converting the sound reflection signal corresponding to the target dirt into an image representation; obtaining a fused image corresponding to the target dirt according to the original image corresponding to the target dirt, the image representation and a preset multi-channel image fusion algorithm; inputting a fusion image corresponding to the target dirt into the improved deep neural network to obtain segmentation mask data and size data of the target dirt; adjusting the pose of the dirt cleaning equipment according to the segmentation mask data and a preset coordinate conversion algorithm; according to the size data of the target dirt, dirt cleaning parameters are adjusted, and dirt cleaning equipment is controlled to execute cleaning operation. According to the method, the dirt recognition and cleaning precision and efficiency are remarkably improved, and various cleaning requirements and operation environments can be better met.
Owner:CHINA MERCHANTS DEEPSEA RES INST SANYA CO LTD +1

Multi-object tracking method based on global-local feature joint modeling

The invention provides a multi-object tracking method based on global-local feature joint modeling, and the method comprises the steps: carrying out the multi-scale image pyramid generation of a current frame of a large-scene high-resolution video, and obtaining a plurality of multi-scale image representations with different resolutions; target detection is carried out in a sliding window mode, a non-maximum suppression algorithm is used for fusing detection results under all scales to construct a joint query group containing global target query and local target query, the joint query group is input into a decoder and is associated with encoded image features through a cross attention mechanism, and a target query result is obtained. Outputting global-local joint feature representation; and in combination with a shielding state prediction result of the target, performing optimal matching on the current detection target and the trajectory set by adopting a shielding state perception matching strategy, and dynamically updating or discarding the trajectory. According to the method, the tracking precision and continuity of multiple objects in a large-scene high-resolution video in a dense shielding environment can be effectively improved, and collaborative optimization of global and local features is realized.
Owner:TSINGHUA UNIVERSITY

Cross-modal attention collaborative jail break attack method

The invention discloses a cross-modal attention collaborative jail break attack method, and belongs to the technical field of artificial intelligence security. The method comprises the following steps of: constructing an input sequence representation containing a system prompt, an adversarial image representation, a malicious query and an adversarial text suffix according to a causal self-attention mechanism; inputting the sequence into a visual language model to execute forward propagation; based on the designed attention-oriented loss collaborative function, optimizing adversarial image representation through a joint gradient optimization algorithm and updating an adversarial text suffix to optimize an attack target; iteratively circulating until convergence, and outputting the optimized unified multi-modal knowledge; and finally, utilizing the knowledge to construct an attack sequence to realize jailbreak. According to the method, accurate control on an internal attention mechanism of the visual language model is realized for the first time, and through visual-text dual-mode cooperative attack, the attack success rate is remarkably improved while high concealment is kept, and the important driving force for promoting the progress of a safe alignment technology is achieved.
Owner:NAT UNIV OF DEFENSE TECH

Zero sample image description method based on filtering hybrid representation and hallucination suppression

PendingCN120219879ANatural language data processingWord tokenEncoder
The invention provides a zero sample image description method based on filtering hybrid representation and hallucination suppression. The method comprises the following steps: acquiring a key entity in an image; the method comprises the following steps: acquiring sub-region features and category tokens of an entity through an image encoder, projecting the sub-region features and the category tokens to obtain projected sub-region features and global image representation, filtering the projected sub-region features, and mixing the filtered sub-region features with the global image representation; and obtaining the original generation probability of the lexical elements by adopting a hallucination suppression method to form image description.
Owner:XUZHOU ANCHUANG MINING INTELLIGENT TECHNOLOGY DEVELOPMENT CO LTD

Sample processing agnostic image representation learning for digital pathology

Described herein are systems, methods, and programming for analyzing and classifying digital pathology images agnostic to sample processing techniques used to prepare the digital pathology images. In some embodiments, image data including a first image set and a second image set may be obtained. The first and second image sets may be processed using a first and second slide preparation machine, respectively. A first augmented view set and a second augmented view set may be generated based on augmentations applied to the first and second image sets. For each image, a first vision transformer to may be trained to: generate a first representation of an augmented view of the first augmented view set, and enhance a similarity between the first representation and a second representation of an augmented view of the second augmented view set. The second representation may be generated via a second vision transformer.
Owner:GENENTECH INC +1

Multi-modal computer vision data fusion method

The invention provides a multi-modal computer vision data fusion method, and relates to the field of computer vision data fusion. The method comprises the following steps: 1, firstly, carrying out data alignment, obtaining a new image representation through pixel-level fusion, and then carrying out sensor fusion to integrate data into a uniform format for subsequent analysis; 2, feature fusion is carried out, firstly, feature splicing is carried out to serve as input of a model, then an attention mechanism is used to pay attention to more important modal information, and finally joint embedding is carried out to compare and analyze data; and step 3, finally, decision fusion is carried out, classification results are weighted through a voting mechanism, and then weighted averaging is carried out on data to obtain a final result. By processing heterogeneity, missing data and noise among different modals, data fusion is performed among different modals in advance, the fusion effect is optimized through a more efficient algorithm by secondary data processing, and the fusion efficiency is improved.
Owner:XIAN INST OF INTERPRETATION & TRANSLATION

A system and a method for 3D image processing, and a method for rendering a 3D image

A system and method of three-dimensional image processing. comprising the steps of pre-processing at least one set of raw images each including a plurality of two-dimensional source images. wherein each of the two-dimensional source image represents a cross-sectional view of a three-dimensional object at different positions along an axis in a three-dimensional space: and constructing a three-dimensional image representing the three-dimensional object by integrating 3D mesh points extracted from the at least one set of raw images being pre-processed: wherein the three-dimensional image is readable by a first image viewer arrange to render and to facilitate manipulation of the three-dimensional image.
Owner:SYNGULAR TECHNOLOGY LIMITED

Cancer survival prediction method and system based on pathological image

The invention provides a cancer survival prediction method based on a pathological image, and relates to the technical field of artificial intelligence and medical image analysis. The method comprises the following steps: firstly, dividing an acquired pathological image to obtain a plurality of image blocks; and obtaining the feature vector of each image block and identifying the tissue type to which each image block belongs. Calculating spatial proximity, feature similarity and tissue type compatibility between the image blocks according to the center coordinates of the image blocks, the feature vectors and the tissue type to which each image block belongs; and based on a calculation result, constructing a dynamic heterogeneous graph, and performing feature extraction to obtain a multi-scale feature graph of the dynamic heterogeneous graph. And performing multi-prototype learning on the basis of the belonging organization type of the image block and a cross-category attention mechanism to generate multiple prototypes. And finally, the multi-scale feature map and the multiple prototypes are aggregated to obtain final image representation, and cancer survival prediction is performed based on the final image representation, so that the prediction accuracy of the cancer survival state is improved.
Owner:GUANGDONG UNIV OF TECH

Evaluation method and device for space movement track

The invention provides a method and a device for evaluating a space moving track, and relates to the technical field of artificial intelligence. The invention discloses a spatial movement track assessment method, which comprises the following steps of: generating track-risk image data according to a spatial movement track and a corresponding risk map; inputting the trajectory-risk image data and a pre-constructed task cue word into a pre-constructed visual-language large model to obtain an evaluation result; and transmitting the evaluation result to the target end. According to the technical scheme provided by the embodiment of the invention, the space movement track and the risk map which need to be evaluated are uniformly converted into the image representation as the input of the vision-language large model, and the image representation and the task cue word are input into the vision-language large model; the general knowledge and the cross-modal reasoning ability of the vision-language large model are utilized to finish track evaluation in real time, and the expansibility and the interpretability are good.
Owner:LOW-ALTITUDE ECONOMIC BRANCH OF GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE

Text-guided multi-dimensional and multi-modal image clustering method and system

The invention discloses a text-guided multi-dimensional multi-modal image clustering method and system, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a target clustering dimension and a plurality of to-be-clustered images, and generating a description text and two answer texts for the target clustering dimension for each image; performing feature coding on each image and the corresponding description text and answer text to obtain image features, description text features and answer text features; calculating similarities between the image features and the latter two, and performing weighted fusion based on the similarities to obtain fused text features; performing cross attention calculation on the image features by taking the fused text features as query to obtain image representation focused on a target clustering dimension after text guidance, and clustering according to the image representation to obtain a clustering result; the method has the advantages that the problems of text and image content disjunction and semantic dimension conflict in image clustering can be relieved, and semantic interpretability and accuracy of clustering results are improved.
Owner:XIDIAN UNIV

Image processing method, image processing device, printing system, and image processing program

To provide a technique that enables a user to easily grasp each printing layer when a plurality of printing layers are formed.SOLUTION: An image processing method includes the steps of: (a) receiving the stacking order of printing medium and one or more printing layers; (b) displaying, on a display device, a preview image representing a state in which one or more virtual three-dimensional objects corresponding to one or more printing layers and the three-dimensional object corresponding to the printing medium are stacked and combined in the stacking order, the preview image corresponding to how the preview image appears in a three-dimensional virtual space; (c) displaying, on the display device, the preview image representing a state in which one or more three-dimensional objects corresponding to one or more printing layers and the three-dimensional object corresponding to the printing medium are stacked in the stacking order with intervals between them, the preview image corresponding to how the preview image appears in the virtual space; and (d) receiving an instruction to execute either the step (b) or the step (c), and executing the instructed step of the step (b) and the step (c).SELECTED DRAWING: Figure 19
Owner:SEIKO EPSON CORP

Vehicle camera system for view creation of viewing locations

Aspects of the subject disclosure relate to a vehicle camera system for view creation of view locations. A device may include a processor that obtains first data and second data from different cameras of a vehicle. The first data may include an image representation of a scene in a first field of view being observed from a first perspective view and the second data may include an image representation of the scene in a second field of view being observed from the first perspective view in a second field of view. The processor may modify the first data and the second data by a perspective transformation to adjust the image representation to a second perspective view. The processor may create a stitched image representing the scene in a combined field of view observed from the second perspective view by stitching the modified first data with the modified second data.
Owner:RIVIAN HOLDINGS LLC

Machine-learned text alignment prediction for providing an augmented-reality translation interface

Systems and methods for providing an augmented-reality translation interface can include obtaining an image, processing the image to generate an image representation and one or more paragraph bounding boxes, and processing the image representation and the one or more paragraph bounding boxes with a machine-learned alignment classification model to generate one or more alignment classifications. In parallel or in series, text from the image can be determined and translated. The image, the translated text, and the one or more alignment classifications can then be processed to generate and provide the translation in an augmented-reality interface.
Owner:GOOGLE LLC