Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

433 results about "Image representation" patented technology

Image-level representation is a (numerical) way to represent an image without a direct pixel representation. For example, one could represent an image by its histograms of luminance and chroma values, or by its Fourier transform, or by any other statistical measure. Such a representation helps compare images or detect specific features.

Human body posture estimation method based on radar point cloud imaging and multi-dimensional feature fusion

The invention discloses a human body posture estimation method based on radar point cloud imaging and multi-dimensional feature fusion, and relates to the cross technical field of computer vision and radar signal processing, and the method comprises the following three key technical links: firstly, improving the target resolution through spatial energy distribution estimation; reconstructing target three-dimensional space distribution by using the positive correlation between radar signal energy and a target reflection area and adopting a least square estimation algorithm; secondly, constructing a structured multi-dimensional point cloud matrix, and converting sparse radar point cloud into high-information-density imaging representation through a distance-speed hierarchical sorting strategy; and finally, designing a multi-dimensional feature fusion attitude estimation network, integrating three-dimensional convolution, a multi-head attention mechanism and a gating circulation unit, and realizing collaborative extraction of spatio-temporal features. According to the method, the problems of sparse target features, noise sensitivity and poor universality in traditional millimeter wave radar attitude estimation are solved.
Owner:DALIAN MARITIME UNIVERSITY

Method, device, and product for retrieval

The present disclosure provides a method, a device, and a product for retrieval. The method includes acquiring context information related to an image and determining a representation of the image based on image data and the context information of the image, where the context information includes at least one of environment parameters, user behavior data, time elements, or field metadata. The method further includes encoding the representation as an image vector in a high-dimensional vector space and storing it into an image vector database. When retrieval is performed, a query that includes text information and that is for the image vector database is received, and an image associated with the text information is determined from the image vector database. The method according to the present disclosure can improve accuracy and efficiency for image retrieval.
Owner:DELL PROD LP

Image retrieval method and device, computer equipment and storage medium

The invention discloses an image retrieval method and device, computer equipment and a storage medium, belongs to the technical field of artificial intelligence, and is provided with a retrieval system applied to an insurance marketing scene image. According to the method and the device, the to-be-processed image is segmented and coded, the content features of the image are combined with the position information to form high-quality image representation, and the corresponding image index information is generated by utilizing the preset image index generator, so that efficient indexing and retrieval of the image are realized. In the encoding process, local features of the image are reserved, spatial position information is fused, discrimination and uniqueness of image indexing are improved, generated image indexing information and an original image are stored in an image retrieval database in an associated mode, it is ensured that the image retrieval process has high matching precision and response speed, and the image retrieval efficiency is improved. The image retrieval accuracy and processing efficiency are effectively improved, and the method is suitable for large-scale image library management and rapid retrieval.
Owner:PING AN TECH (SHENZHEN) CO LTD

Methods and systems for multiple instance learning of tissue sample images

PendingUS20250356486A1Image enhancementImage analysisFeature vectorNeedle core biopsy
Methods for multiple instance learning of tissue sample images are described. The methods may comprise, for example, receiving a whole slide image from a needle core biopsy sample from a subject; identifying a tissue region in the whole slide image; selecting a set of image patches from the identified tissue region; resampling the set of image patches at a plurality of image scales to generate a plurality of resampled image patches; generating image representations for the plurality of resampled image patches; extracting feature vectors based on the image representations; providing the feature vectors as input to a trained machine learning model configured to predict a gene alteration state; and outputting the predicted gene alteration state for the needle core biopsy sample for the subject.
Owner:FOUNDATION MEDICINE INC

Dirt identifying and cleaning method, device and equipment

The invention is suitable for the field of intelligent jet cleaning, and provides a dirt recognition cleaning method, device and equipment, and the dirt recognition cleaning method comprises the steps: obtaining an original image corresponding to target dirt and a sound reflection signal corresponding to the target dirt; converting the sound reflection signal corresponding to the target dirt into an image representation; obtaining a fused image corresponding to the target dirt according to the original image corresponding to the target dirt, the image representation and a preset multi-channel image fusion algorithm; inputting a fusion image corresponding to the target dirt into the improved deep neural network to obtain segmentation mask data and size data of the target dirt; adjusting the pose of the dirt cleaning equipment according to the segmentation mask data and a preset coordinate conversion algorithm; according to the size data of the target dirt, dirt cleaning parameters are adjusted, and dirt cleaning equipment is controlled to execute cleaning operation. According to the method, the dirt recognition and cleaning precision and efficiency are remarkably improved, and various cleaning requirements and operation environments can be better met.
Owner:CHINA MERCHANTS DEEPSEA RES INST SANYA CO LTD +1

Multi-object tracking method based on global-local feature joint modeling

The invention provides a multi-object tracking method based on global-local feature joint modeling, and the method comprises the steps: carrying out the multi-scale image pyramid generation of a current frame of a large-scene high-resolution video, and obtaining a plurality of multi-scale image representations with different resolutions; target detection is carried out in a sliding window mode, a non-maximum suppression algorithm is used for fusing detection results under all scales to construct a joint query group containing global target query and local target query, the joint query group is input into a decoder and is associated with encoded image features through a cross attention mechanism, and a target query result is obtained. Outputting global-local joint feature representation; and in combination with a shielding state prediction result of the target, performing optimal matching on the current detection target and the trajectory set by adopting a shielding state perception matching strategy, and dynamically updating or discarding the trajectory. According to the method, the tracking precision and continuity of multiple objects in a large-scene high-resolution video in a dense shielding environment can be effectively improved, and collaborative optimization of global and local features is realized.
Owner:TSINGHUA UNIVERSITY

Cross-modal attention collaborative jail break attack method

The invention discloses a cross-modal attention collaborative jail break attack method, and belongs to the technical field of artificial intelligence security. The method comprises the following steps of: constructing an input sequence representation containing a system prompt, an adversarial image representation, a malicious query and an adversarial text suffix according to a causal self-attention mechanism; inputting the sequence into a visual language model to execute forward propagation; based on the designed attention-oriented loss collaborative function, optimizing adversarial image representation through a joint gradient optimization algorithm and updating an adversarial text suffix to optimize an attack target; iteratively circulating until convergence, and outputting the optimized unified multi-modal knowledge; and finally, utilizing the knowledge to construct an attack sequence to realize jailbreak. According to the method, accurate control on an internal attention mechanism of the visual language model is realized for the first time, and through visual-text dual-mode cooperative attack, the attack success rate is remarkably improved while high concealment is kept, and the important driving force for promoting the progress of a safe alignment technology is achieved.
Owner:NAT UNIV OF DEFENSE TECH

Multi-modal computer vision data fusion method

The invention provides a multi-modal computer vision data fusion method, and relates to the field of computer vision data fusion. The method comprises the following steps: 1, firstly, carrying out data alignment, obtaining a new image representation through pixel-level fusion, and then carrying out sensor fusion to integrate data into a uniform format for subsequent analysis; 2, feature fusion is carried out, firstly, feature splicing is carried out to serve as input of a model, then an attention mechanism is used to pay attention to more important modal information, and finally joint embedding is carried out to compare and analyze data; and step 3, finally, decision fusion is carried out, classification results are weighted through a voting mechanism, and then weighted averaging is carried out on data to obtain a final result. By processing heterogeneity, missing data and noise among different modals, data fusion is performed among different modals in advance, the fusion effect is optimized through a more efficient algorithm by secondary data processing, and the fusion efficiency is improved.
Owner:XIAN INST OF INTERPRETATION & TRANSLATION

Evaluation method and device for space movement track

The invention provides a method and a device for evaluating a space moving track, and relates to the technical field of artificial intelligence. The invention discloses a spatial movement track assessment method, which comprises the following steps of: generating track-risk image data according to a spatial movement track and a corresponding risk map; inputting the trajectory-risk image data and a pre-constructed task cue word into a pre-constructed visual-language large model to obtain an evaluation result; and transmitting the evaluation result to the target end. According to the technical scheme provided by the embodiment of the invention, the space movement track and the risk map which need to be evaluated are uniformly converted into the image representation as the input of the vision-language large model, and the image representation and the task cue word are input into the vision-language large model; the general knowledge and the cross-modal reasoning ability of the vision-language large model are utilized to finish track evaluation in real time, and the expansibility and the interpretability are good.
Owner:LOW-ALTITUDE ECONOMIC BRANCH OF GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE

Text-guided multi-dimensional and multi-modal image clustering method and system

The invention discloses a text-guided multi-dimensional multi-modal image clustering method and system, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a target clustering dimension and a plurality of to-be-clustered images, and generating a description text and two answer texts for the target clustering dimension for each image; performing feature coding on each image and the corresponding description text and answer text to obtain image features, description text features and answer text features; calculating similarities between the image features and the latter two, and performing weighted fusion based on the similarities to obtain fused text features; performing cross attention calculation on the image features by taking the fused text features as query to obtain image representation focused on a target clustering dimension after text guidance, and clustering according to the image representation to obtain a clustering result; the method has the advantages that the problems of text and image content disjunction and semantic dimension conflict in image clustering can be relieved, and semantic interpretability and accuracy of clustering results are improved.
Owner:XIDIAN UNIV

Machine-learned text alignment prediction for providing an augmented-reality translation interface

Systems and methods for providing an augmented-reality translation interface can include obtaining an image, processing the image to generate an image representation and one or more paragraph bounding boxes, and processing the image representation and the one or more paragraph bounding boxes with a machine-learned alignment classification model to generate one or more alignment classifications. In parallel or in series, text from the image can be determined and translated. The image, the translated text, and the one or more alignment classifications can then be processed to generate and provide the translation in an augmented-reality interface.
Owner:GOOGLE LLC

An image representation model pruning method based on multi-granularity importance measurement

The application relates to a multi-granularity importance metric-based image representation model pruning method, and belongs to the technical field of model compression. The method combines information loss and redundancy criteria to construct a multi-granularity importance metric-based model channel pruning method, compensates for the single importance metric and unstable performance of an existing algorithm, and improves the fine-tuning performance of a pruned model by keeping the balance and diversity of pruned channel distribution, thereby expanding the model pruning idea based on importance metric.
Owner:BEIHANG UNIV

Importance sampling environment maps for real-time path tracing

PendingUS20260080608A13D-image renderingComputer graphics (images)Importance map
Approaches presented herein provide systems and methods for path tracing using a set of textured spherical surfaces obtained from an importance map for an image. An image representation may be generated using the importance map and evaluated to identify a first set of nodes. The nodes may have associated values, such as luminance values, that may be used to subdivide the nodes into bins to maintain a weighed distribution for the associated values. An array of nodes may be generated for sampling and conversion to a three-dimensional direction that may be applied to one or more lighting effects.
Owner:NVIDIA CORP

Synthetic generation of training data

The present application relates to image processing. A computer-implemented method is provided for generating synthetic training data that is usable for training a data-driven model for analysing a surface image of a physical product that comprises at least one object, the method comprising:a) providing image data that comprises:an object image dataset comprising a plurality of object images of the at least one object, at least one object image being associated with a label usable for annotating a content of the object image; anda background image representing a background of a surface image of the physical product;b) generating a synthetic object image dataset from the object image dataset, wherein the synthetic object image dataset comprises a plurality of synthetic object images of the at least one object, at least one synthetic object image being associated with a label; andc) generating a plurality of first synthetic training data samples, wherein each first synthetic training data sample is generated by selecting one or more object images from the synthetic object image dataset and by plotting the selected one or more object images at one or more locations on the background image. The computer-implemented method may be used to improve the computer vision technique for the application in the technical field of agriculture and in production environment.
Owner:BASF SE

Three-dimensional Gaussian splash reconstruction method for underwater scene

The invention discloses a three-dimensional Gaussian splash reconstruction method for an underwater scene, and belongs to the technical field of computer vision and three-dimensional reconstruction. Comprising the following steps: acquiring a monocular video frame sequence of a target underwater scene, a corresponding camera pose sequence, an initial sparse point cloud, an initial three-dimensional Gaussian point set and learnable physical parameters of an underwater imaging model; in the training process, performing weighted evaluation on a reconstruction error based on a multi-view consistency mechanism of opacity weighting, calculating an importance score of each Gaussian point, and performing densification operation on a three-dimensional Gaussian point set; adopting a staged freezing strategy to cooperatively optimize the three-dimensional Gaussian point set and underwater imaging model parameters; and performing rendering and underwater image synthesis on any new view angle camera pose based on the optimized three-dimensional Gaussian point set and underwater imaging model parameters, and outputting a new view angle synthesized image to represent a reconstruction result. According to the method, the geometric compactness, the visual fidelity and the physical interpretability of an underwater three-dimensional reconstruction result are improved.
Owner:ZHEJIANG UNIV

Method and system for enhancing images using machine learning

A method for enhancing images for metrological applications. A first set of first images is provided, wherein the first set comprises at least one first image, the first set representing a scene of interest, and a second set of second images is provided, wherein the second set comprises at least one second image, the second set representing at least a part of the scene of interest. The method comprises the following steps: 1) jointly processing the first set and the second set by a neural network, and 2) outputting a third set of images by the neural network, the third set comprising at least one processed image, from which is determinable the position and / or orientation of the at least one object with a processed precision, wherein the processed precision is equal to or higher than the initial precision.
Owner:HEXAGON INNOVATION HUB GMBH

System and Method for Object Segmentation for Task Performance

Embodiments disclosing a controller for controlling a robot to perform a task are provided. The task is performed in an environment that is represented by an input image. The controller causes segmenting of an object in the input image. A confidence level of segmentation is updated by comparing the segmented object with constrained affined transformations of a template of the object. The constrained affine transformations are based on constraints indicative of a property of the object. The property of the object and the updated confidence level of segmentation are then used for performing the task.
Owner:MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC

Multi-capture for remote deposit via scanner

Systems and techniques may be used to perform multi-capture for remote deposit via scanner. For example, a technique may include scanning a first side and a second side of a plurality of checks, and comparing geometrical features of the first side of each check of the plurality of checks to geometrical features of the second side of each check of the plurality of checks. Based on the comparison, the technique may include selecting an image from a first set of respective individual images that corresponds to an image from a second set of respective individual images to form a pair of images, the pair of images representing a single check of the plurality of checks, and outputting the pair of images.
Owner:WELLS FARGO BANK NA

System

A system is provided.SOLUTION: A system comprising: means for registering a face image of a user; means for extracting feature points from the registered face image; means for generating a 3D model of a face based on the feature points; means for automatically generating an angle, a pose, and an expression to be combined with a background picture; means for representing the face image by the 3D model according to the generated angle, pose, and expression and naturally combining the face image with the background picture; and means for presenting the combined image to the user and storing the combined image after checking.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Knowledge distillation-based lightweight insect identification method and customs real-time universal equipment

The invention relates to the technical field of insect image recognition, and discloses a lightweight insect recognition method based on knowledge distillation, which comprises the following steps: carrying out local preprocessing on an insect image to be recognized to obtain a preprocessed image; performing feature extraction and channel expansion on the preprocessed image to obtain a convolution feature map; dividing the convolutional feature map into a plurality of image blocks with fixed sizes, and mapping each image block into a corresponding Token to obtain a Token sequence for a local self-attention mechanism; applying self-attention to local regions of the Token sequence for the local self-attention mechanism to model a global relationship of the local regions in the image to obtain a Token sequence for classification; the Token sequences used for classification are converted and fused into global image representation, and softmax probability vectors corresponding to candidate insect species are output; the softmax probability vectors are processed in a descending order, the first k categories are selected as candidate results, a top-k candidate list is output, and local lightweight insect recognition processing can be achieved.
Owner:中国电子口岸数据中心黄埔分中心

Crowd counting method and system based on WiFi and video modal cross-level attention

The invention relates to the technical field of crowd counting, in particular to a crowd counting method and system based on WiFi and video modal cross-level attention, and the method comprises the steps: constructing a WiFi density map at a WiFi sensing side, and converting an irregular detection record into a fixed-size image representation; on a video sensing side, marking a region of interest for video frames collected by cameras with different visual angles, and cutting the region of interest to serve as a video side image; respectively carrying out feature coding on the WiFi density map and the video side image by adopting a convolutional neural network and self-attention combined mode; gradually aligning the WiFi modal feature embedded representation and the video modal feature embedded representation through multi-layer stacked cross-modal attention to obtain a cross-modal fusion feature; and inputting the cross-modal fusion features into a lightweight multilayer perceptron, and outputting crowd count. According to the invention, through hierarchical alignment and fusion of WiFi signals and video features, accurate estimation of the number of crowds in a large-scale complex scene can be realized.
Owner:INNER MONGOLIA ZHIXING HUILIAN TECHNOLOGY CO LTD

Image generation

Computer implemented methods and associated systems are described, which have particular application to image generation by machine learning models. A method of generating a composite image is described that is based on two images using a controlled machine learning model. A method of processing a composite image is also described which includes determining that a transition region of the composite image is similar to one of the images on which the composite image was based and using in the transition region visual elements from the basic image. A method for providing a user interface is also described. The method includes displaying representations of images generated using common input and different hyperparameters.
Owner:CANVA PTY LTD

Horizontal panoramic unmanned aerial vehicle detection method and system based on rotation event camera

The invention discloses a horizontal panoramic unmanned aerial vehicle detection method and system based on a rotation event camera. The method comprises the following steps: acquiring an event stream which is shot by the event camera installed on a rotation platform and contains unmanned aerial vehicle information; performing event stream preprocessing to obtain event groups in a plurality of continuous fixed time intervals; performing image generation for each event group to form a class image representation; performing input construction on the class image representations of all the event groups to obtain input features; inputting the input features into a pre-trained space-time fusion detection network to obtain a detection result, wherein the detection result comprises a detection frame and a target confidence coefficient of an unmanned aerial vehicle target; and performing orientation estimation according to the detection frame of the unmanned aerial vehicle target and the attitude of the event camera on the rotating platform to obtain the relative azimuth angle of the unmanned aerial vehicle. The objective of the invention is to provide an all-directional 360-degree horizontal field angle for the event camera, realize real-time and reliable detection and orientation estimation of the unmanned aerial vehicle, and meet actual requirements in a dynamic deployment scene.
Owner:HUNAN UNIV

Risk assessment method and device based on artificial intelligence, computer equipment and medium

The invention belongs to the technical field of artificial intelligence, and relates to an artificial intelligence-based risk assessment method and device, computer equipment and a storage medium, and the method comprises the steps: receiving business data inputted by a user; performing data conversion processing related to noise representation on the business data to obtain target data; calling a diffusion generation model; wherein the diffusion generation model comprises a text representation generator, an image representation generator and an iterative optimization module; performing iterative optimization processing on the target data based on an iterative optimization module and a text representation generator to obtain a risk assessment result; performing image visualization on the risk assessment result based on an image representation generator to obtain a target risk assessment result in an image form; generating decision interpretation information of the risk assessment result based on an interpretation module; and displaying the target risk assessment result and the decision interpretation information. The method can be applied to data risk assessment scenes in the financial field and the medical field, and the accuracy of risk assessment is effectively improved through the method.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Training methods, devices, electronic equipment, and media for image and text generation models

This application discloses a training method, apparatus, electronic device, and medium for a text-image generation model, belonging to the field of artificial intelligence. The method includes: inputting a first training sample pair from a first training sample pair set into a first text-image generation model, and outputting a second training sample pair. The first training sample pair includes a first image and first text describing the image content of the first image. The second training sample pair includes a second image and second text, where the second image is an image obtained by converting the first text into a text-image representation, and the second text is the text obtained by converting the first image into a text-image representation. Based on the first and second training sample pairs, M training sample pairs are generated. The first training sample pair in the first training sample pair set is replaced with a third training sample pair to obtain a second training sample pair set, where the third training sample pair is the training sample pair with the highest text-image similarity among the M training sample pairs. The first text-image generation model is trained based on the second training sample pair set to obtain a target text-image generation model.
Owner:VIVO MOBILE COMM CO LTD

Unsupervised pre-training of geometric vision models

A method includes: performing unsupervised pre-training of a model, the model including and a decoder including: obtaining a first image and a second image under different conditions or from different viewpoints; encoding, by the encoder, the first image into a representation of the first image and the second image into a representation of the second image; transforming the representation of the first image into a transformed representation; decoding, by the decoder, the transformed representation into a reconstructed image, where the transforming of the representation of the first image and the decoding of the transformed representation is based on the representation of the first image and the representation of the second image; and adjusting one or more parameters of at least one of the encoder and the decoder based on minimizing a loss; and fine-tuning the model, initialized with a set of task specific encoder parameters, for a geometric vision task.
Owner:NAVER CORP

Highway long tunnel real-time fire source detection method and system based on multi-modal fusion

The application discloses a kind of based on multimodal fusion expressway long tunnel real-time fire source detection method and system, it is estimated by introducing smoke temperature airflow dynamic propagation delay under wind speed constraint, visible light, infrared, smoke concentration and environmental temperature and so on heterogeneous data are phase shift compensation and space-time unification, different modal can correspond to the same fire physical time, then respectively extract visual side flame profile, thermal gradient feature and sensing side smoke temperature mutation, time evolution characteristics, and utilize cross-modal interactive fusion to form the fusion feature expression with image representation ability and environmental perception ability, finally, combine target detection and inverse perspective positioning Output fire confidence and fire source coordinates, to solve the problem that multi-source data is difficult to effectively align, early fire source feature weak leads to inaccurate identification, false alarm under complex disturbance is higher and fire source position is difficult to determine in real time accurately, improve the real-time, accuracy and engineering applicability of tunnel fire source detection.
Owner:LUOYANG INST OF SCI & TECH

Analysis support device and analysis support system

An analysis support device includes a display; a first display control section configured to display a first screen including a device scheduled image and a device actual result image on the display when a device actual result time at which a mounting device actually performs a mounting step of mounting a component on a substrate is delayed; and a second display control section configured to display a second screen including a preparation scheduled image and a preparation actual result image on the display when a predetermined first instruction is given by a user after the first screen is displayed on the display, the preparation scheduled image indicating a preparation scheduled time at which a preparation step, which is preparation for the mounting step, is to be performed, and the preparation actual result image indicating a preparation actual result time at which the preparation step is actually performed.
Owner:FUJI CORP