Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

499 results about "Candidate image" patented technology

Artificial intelligence-based image search refinement

Systems and methods for image search result filtering can include obtaining a search query, determining a plurality of candidate image search results, processing the search query with a generative model to determine a plurality of search result criteria, and refining the plurality of candidate image search results based on determining whether the candidate results satisfy the plurality of search results criteria. The systems and methods can perform a plurality of determinations based on the output of the generative model.
Owner:GDM HOLDING LLC

Method and device for detecting surface defects of injection molded part based on double-model collaboration

The invention provides an injection molding part surface flaw detection method and device based on double-model collaboration, and the method comprises the steps: carrying out the local abnormal reflection feature analysis of a multi-view optical image collected on the surface of an injection molding part based on a first detection model, and obtaining a candidate flaw region; performing spatial positioning in the multi-view optical image based on the candidate flaw area to obtain a multi-view candidate image segment; performing surface geometric continuity analysis on the multi-view candidate image segments based on a second detection model to obtain a three-dimensional surface geometric consistency result; carrying out authenticity discrimination on the candidate flaw area based on a three-dimensional surface geometric consistency result to obtain a real flaw area; wherein the first detection model is used for capturing an optical response model of local abnormal reflection characteristics; the second detection model is a stereoscopic vision model for analyzing the geometric continuity of the multi-view lower surface. According to the invention, the false alarm rate of surface defect detection of the injection molded part under a complex surface condition is reduced.
Owner:SHENZHEN SUCCESS RAIN TECH CO LTD

Enzyme freeze-dried powder quality detection method based on machine vision

The invention discloses a machine vision-based enzyme freeze-dried powder quality detection method, and belongs to the technical field of machine vision, and the method specifically comprises the following steps: in a closed optical cavity, synchronously collecting a multispectral polarization sequence of a sample through annular multispectral polarization illumination and microscopic imaging; performing short-time airflow pulse disturbance and continuous imaging on the sample, and constructing a polarization time sequence texture map by means of inter-frame polarization retention difference; according to Stokes parameters, calculating a deskewness graph and a specular reflection component graph, and carrying out joint normalization on the deskewness graph and the specular reflection component graph and the polarization time sequence texture graph to form a multi-channel feature stack; constructing a double-branch encoder, and performing comparative learning on the feature stack to generate an impurity prototype dictionary; using the dictionary to match the feature stack pixel by pixel, and tracking and inhibiting powder agglomeration false detection through a topological constraint connected domain; and recovering a height map by using an oblique incidence fringe phase shift method, performing pixel-level fusion with the candidate image, retaining targets with height abrupt change and consistent polarization characteristics, and outputting an impurity mask.
Owner:LIAONING INST OF SCI & TECH

Scheme detection method and device for complex image-text mixed file

The invention provides a scheme detection method and device for a complex image-text mixed file, and the method comprises the steps: extracting standard review items and standard review contents in a standard manual, and carrying out the structural processing, and obtaining a standard review file; converting a to-be-detected file into a to-be-detected image, extracting a to-be-examined multi-scale candidate region image from the to-be-detected image by using an initial recognition model and a matching model obtained by OCR and pre-training, and extracting character information of the multi-scale candidate image to obtain character information of the candidate region; inputting the multi-scale candidate region image and the standard review file into a lightweight hybrid twin network obtained by pre-training for image feature matching, and outputting a matched feature image; inputting the matched feature image, the standard review file and the candidate area text information into a multi-modal large model obtained by pre-training for compliance analysis, and outputting a detection result; according to the method, the automation and intelligence degree of review of the complex image-text mixed file can be remarkably improved.
Owner:ZHEJIANG SHUANGYUAN TECH CO LTD

Picture display method, electronic equipment and computer readable storage medium

The embodiment of the invention discloses a picture display method, electronic equipment and a computer readable storage medium, in the method, the electronic equipment responds to a target operation used for uploading a picture, and target information related to interface content of an application program is obtained; according to the first picture number and the target information, a target candidate picture set of each target element is obtained from the picture set, the target elements are determined according to the target information, the target candidate picture set comprises at least one target candidate picture, and the total number of the target candidate pictures of each target element is smaller than or equal to the first picture number; the first picture number is the number of pictures capable of being displayed on a first screen interface of the picture selection interface; and displaying the target candidate picture of each target element on the first screen interface. Thus, the user can obtain the pictures of the multiple target elements on the first screen interface, the pictures of the target elements do not need to be obtained through page turning, and the convenience of selecting the pictures by the user is improved.
Owner:HUAWEI TECH CO LTD

Method and electronic device for automatic generation of high quality data for image editing applications

A method for generating an image editing dataset is provided. The method may include obtaining a candidate image for inputting an AI model from among at least one candidate image. The method may include determining an editing operation to be performed by the AI model on the candidate image from among at least one editing operation. The method may include determining a base prompt and a control prompt based on inputting of the candidate image and the editing operation into the AI model, wherein the base prompt comprises base instructions for detection of at least one object within the candidate image, and the control prompt comprises control instructions relevant to the editing operation to be performed by the AI model on the candidate image. The method may include generating the image editing dataset based on the base prompt and the control prompt.
Owner:SAMSUNG ELECTRONICS CO LTD

Zero-sample multi-modal combined image retrieval method and device based on context awareness

The invention provides a zero-sample multi-modal combined image retrieval method and device based on context awareness, and relates to the technical field of image retrieval, and the method comprises the following steps: obtaining a training sample set; the training sample set is input into a target CLIP model, target image embedding and target text embedding of the training sample set are obtained, and learnable language prompt embedding and visual prompt embedding are introduced into a text encoder and an image encoder of the target CLIP model respectively; training the mapping network model according to the target image embedding, the target text embedding and the target loss function to obtain a target mapping network model; obtaining a target image retrieval model according to the target CLIP model and the target mapping network model, and inputting the reference image and the text description into the target image retrieval model to obtain a target query feature; and outputting a retrieval result according to the similarity between the target query feature and the image embedding of each candidate image. According to the image retrieval method, a multi-modal combined image retrieval task can be completed.
Owner:HUBEI UNIV

Automobile part detection method

The invention relates to the technical field of automobile part detection, in particular to an automobile part detection method, which comprises the following steps of: dividing a region range of a to-be-detected surface of an automobile part to obtain a plurality of image acquisition regions, and setting acquisition points for the plurality of image acquisition regions; acquiring basic feature data of a plurality of image acquisition areas, setting corresponding acquisition environment light sources, and performing image acquisition by an image acquisition camera corresponding to the acquisition environment light sources to obtain a plurality of detection images; performing feature analysis comprehensive processing on the plurality of detection images to obtain a defect area candidate image, and performing defect feature enhancement processing on the defect area candidate image to obtain a defect positioning image; and inputting the defect positioning image into the pre-training model, and outputting a detection result according to the defect parameters. According to the invention, the acquisition effect of the detection image is guaranteed, the visibility of the detection image is improved, and the accuracy of defect identification is guaranteed while the detection efficiency is improved.
Owner:ANHUI YONGMAOTAI AUTO PARTS CO LTD

Multi-modal visual position identification reordering method and system based on guidance

The invention relates to the technical field of visual position recognition, and particularly discloses a multi-modal visual position recognition reordering method and system based on guidance, and the method comprises the steps: obtaining a query image, and retrieving a plurality of candidate images based on a pre-trained visual basic model and the query image; constructing a composite multi-modal prompt object, wherein the composite multi-modal prompt object comprises an image pair formed by the query image and the current candidate image, and an instruction text used for guiding a multi-modal large language model to perform visual comparison; outputting a structured similarity judgment result, wherein the result comprises a quantitative similarity score; and sorting based on the similarity scores corresponding to all the candidate images, and determining the candidate image with the highest score as an optimal matching result. Through combination of guiding type prompt engineering and structured output, an intermediate text generation link is avoided fundamentally, and the calculation efficiency is improved while the fidelity of all original visual information is reserved.
Owner:SHENZHEN 1024 ROBOT TECHNOLOGY CO LTD

Task execution method and device, equipment and medium

The invention discloses a task execution method and device, equipment and a medium, and relates to the technical field of artificial intelligence. The method comprises the steps that a task prompt word and an interface image are obtained, the task prompt word is used for describing a to-be-executed interaction task, and the interface image is divided into at least two image areas; acquiring image feature representations corresponding to the at least two image areas and text feature representations corresponding to the task cue words; identifying the relevance between the image feature representation and the text feature representation corresponding to the at least two image regions, and determining at least one candidate image region of which the relevance meets the requirement; and generating a task instruction based on the at least one candidate image area, wherein the task instruction is used for instructing execution of an interaction task on the first interface. In combination with the text feature representation corresponding to the task prompt word, redundant feature input is reduced, and the correlation between the candidate image area and the interaction task is improved, so that the accuracy of the task instruction and the execution efficiency of the interaction task are improved.
Owner:MOORE THREADS TECH CO LTD

Vehicle violation manned intelligent detection method and system based on CLIP and D-Fine cascade framework

The invention discloses a vehicle violation manned intelligent detection method and system based on a CLIP and D-Fine cascade framework, and the method comprises the steps: collecting a real-time video frame image of a preset traffic monitoring network, processing the real-time video frame image, inputting a monitoring image into a CLIP-ILP model, employing the CLIP-ILP model as a filter, and combining with a Top-K screening strategy, and screening out candidate images; and inputting the candidate image into a constructed context enhanced D-FINE detection model, carrying out positioning and multi-class detection on a preset target, outputting a multi-class detection result, and executing verification of a spatial co-occurrence rule, a license plate position heuristic rule, a size consistency rule and a context consistency rule. And marking the result which does not pass the verification as suspicious or rejecting the result, and outputting the result which passes the verification as an illegal manned detection result to the target terminal. According to the method, the problem of balance between the recall rate and the precision in manned detection can be solved.
Owner:YUNNAN MINZU UNIV +1

Embedding based retrieval for image search

Methods, systems, and apparatus including computer programs encoded on a computer storage medium, for retrieving image search results using embedding neural network models. In one aspect, an image search query is received. A respective pair numeric embedding for each of a plurality of image-landing page pairs is determined. Each pair numeric embedding is a numeric representation in an embedding space. An image search query embedding neural network processes features of the image search query and generates a query numeric embedding. The query numeric embedding is a numeric representation of the image search query in the same embedding space. A subset of the image-landing page pairs having pair numeric embeddings that are closest to the query numeric embedding of the image search query in the embedding space are identified as first candidate image search results.
Owner:GOOGLE LLC

Industrial finished product microdefect detection method based on generative adversarial network

The invention discloses an industrial finished product microdefect detection method based on a generative adversarial network, and the method comprises the following steps: S1, collecting a high-resolution image of the surface of an industrial finished product, and constructing a data set of a normal image and a defect image; s2, extracting an image feature vector, splicing the image feature vector with a defect type vector, and inputting the spliced image feature vector and the defect type vector into a generator network containing a geometric structure transformation module to generate a defect candidate image; s3, inputting the defect candidate image into a discriminator network containing a topological structure analysis module, and outputting an authenticity discrimination result, a structure connectivity characteristic index and a defect type prediction label; s4, jointly training a generator and discriminator network; s5, inputting a to-be-detected image, and obtaining a defect candidate image and a prediction result; s6, generating a defect probability graph, a saliency graph and a binary mask, and positioning a defect area; and S7, outputting a microdefect detection and classification result in combination with the prediction label, the saliency map and the mask. According to the method, the detection precision and the classification capability of the micro-defects under the complex background are improved.
Owner:CHANGCHUN GUANGHUA UNIV

Image retrieval method, electronic device, and computer-readable storage medium

An image retrieval method includes acquiring an image retrieval condition, the image retrieval condition comprising a reference image and modification text, the modification text being configured to indicate a modification expectation for the reference image; composing the reference image and the modification text, to obtain an image-text composition; acquiring a plurality of candidate images, and determining, for a candidate image, a first similarity between the candidate image and the image-text composition, and a second similarity between the candidate image and the modification text; and determining at least one target image satisfying the image retrieval condition from the plurality of candidate images with reference to the first similarity and the second similarity.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Inplanatory visual question-answering method and system based on question perception and confidence constraint

The invention discloses an explanatory visual question-answering method and system based on question perception and confidence constraint. The system comprises a selection-enhancement module and a fusion enhancement confidence constraint module, the selection-enhancement module is used for selecting candidate image areas based on question semantics and enhancing salient areas related to questions through a learnable enhancement mechanism to realize accurate visual localization; and the fusion enhancement confidence constraint module is used for fusing the enhanced regional features and the multi-modal features, and ensuring that the confidence of a prediction answer is improved when explanation information is introduced through a double-branch prediction and confidence constraint mechanism, so that the reliability of a prediction result is enhanced. According to the method, the defects of problem insensitive positioning and positioning-reasoning disjunction in the prior art can be effectively overcome, the superiority of the method is verified on a public data set, the answer prediction accuracy and the explanation generation quality are remarkably improved, and the method has wide application prospects.
Owner:NANJING UNIV OF POSTS & TELECOMM

Gaming machine

To provide a gaming machine that enhances the enjoyment of the game by utilizing presentations that suggest reliability based on the player's choices. [Solution] In the pre-performance, one of several candidate images is selected as the pre-selected image, and in the post-performance, one of several candidate images is selected as the post-selected image. When the combined performance is executed in state A, the reliability is higher when the pre-selected image and post-selected image are of the same type than when they are of different types. When the performance is executed in state B, which is different from state A, the reliability is higher when the pre-selected image and post-selected image are of different types than when they are of the same type.
Owner:SANSEI R&D KK

Image staged matching and positioning method

The invention discloses an image staged matching and positioning method, and relates to the technical field of image analysis. Comprising the following specific steps: constructing a model architecture, and inputting an unmanned aerial vehicle image u; candidate images {u, Si} are screened through global semantic feature matching; the method comprises the following steps: extracting image key points and description words, establishing an initial matching point pair set, introducing a mismatching screening mechanism, executing homography transformation on geotagging information in a satellite image, and determining ground target positioning under the view angle of an unmanned aerial vehicle. According to the method, through a mode of combining global feature optimization and local mismatching screening, accurate positioning of a ground target under the view angle of the unmanned aerial vehicle is realized, and an optimal matching candidate is rapidly screened out; in the local matching stage, mismatching point pairs are eliminated, and the geotagging information in the satellite image is accurately mapped to the unmanned aerial vehicle image coordinate system through homography transformation, so that the image matching problem under the large view angle change is effectively solved, and the accuracy and efficiency of image matching and positioning are improved.
Owner:BEIJING INFORMATION SCI & TECH UNIV

High-precision localization of a moving object on a trajectory

Techniques for generating high-precision localization of a moving object on a trajectory are provided. In one technique, a particular image that is associated with a moving object is identified. A set of candidate images is selected from a plurality of images that were used to train a neural network. For each candidate image in the set of candidate images: (1) output from the neural network is generated based on inputting the particular image and said each candidate image to the neural network; (2) a predicted position of the particular image is determined based on the output and a position that is associated with said each candidate image; and (3) the predicted position is added to a set of predicted positions. The set of predicted positions is aggregated to generate an aggregated position for the particular image.
Owner:ORACLE INT CORP

Construction waste intelligent classification method based on deep learning

The invention discloses an intelligent construction waste classification method based on deep learning. The method comprises the steps of obtaining a standardized construction waste operation site image data set; generating a global candidate image region set based on the standardized construction waste operation site image data set; forming a node feature vector set; obtaining a construction waste heterogeneous object-relation graph structure containing node type labels and edge type labels; generating a gating coefficient and fused multi-scale feature representation; generating a sub-graph level representation; generating an edge consistency constraint signal based on a boundary clue output by the local feature extraction branch and a pixel mask of the global candidate image region set; and according to the updated node representation, the sub-graph-level representation and the edge consistency constraint signal, defining a construction waste category label, an instance contour and a segmentation mask. According to the method, the instance-level and material-level segmentation precision of the construction waste target under the scenes of stacking, shielding, similar inter-class appearance and large intra-class difference is remarkably improved.
Owner:THE ARCHITECTURAL DESIGN & RES INST OF ZHEJIANG UNIV CO LTD +1

Street lamp visual coding method and system for non-directional dynamic patrol

The invention discloses a non-directional dynamic patrol street lamp visual coding method and system. The method comprises the following steps: acquiring continuously acquired road images and geographic position and attitude data thereof, and grouping after primary processing and spatial index construction; screening candidate image pairs by calculating image similarity, extracting multi-dimensional features of street lamps and environments thereof, constructing local feature descriptors including street lamp bodies and environment context features, and forming small-range environment fingerprints to determine successfully matched street lamp entities; hybrid codes are generated for the entities, and the codes are combined with geographic coordinates, the fused local feature descriptors, the fused fixed object list and the fused identifier of the nearest fixed object distance, and a final code value is obtained through Hash calculation. By implementing the method provided by the invention, visual features, small-range environment fingerprints and spatial priori knowledge can be effectively fused, so that the problems are solved, and intelligent management of urban lighting facilities is promoted.
Owner:WINTOO INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Artificial intelligence-generated backgrounds

Generating images of items with artificial intelligence (AI)-generated backgrounds is described. An example process includes capturing an image(s) of an item, obtaining first image data that identifies a real-world background in the image(s), and presenting a user interface for a user of an electronic device to indicate a descriptor(s). The process may further include receiving, via the user interface, an indication of the descriptor(s), generating prompt data representing a prompt(s) based at least in part on the descriptor(s), sending the first image data and the prompt data to a server computer(s), receiving, from the server computer(s), second image data output by a trained AI model(s), the second image data representing an AI-generated image(s) with an AI-generated background(s), and presenting, based at least in part on the second image data, a candidate image(s) with the AI-generated background(s) for selection by the user.
Owner:BLOCK INC

Literature image-text pair quality control method and system, storage medium and equipment

The invention relates to the technical field of data processing, and discloses a literature image-text pair quality control method, which comprises the following steps: inputting academic literatures, identifying and extracting related descriptive characters of an image in a text, and generating a preliminary image-text pair; extracting a picture title text, calculating a semantic matching degree between the picture title text and the related descriptive text, filtering the preliminary image-text pair based on a first matching degree threshold, and generating a first candidate image-text pair; generating a semantic vector of the academic literature theme, mapping the title text to generate a picture title vector, calculating a correlation degree between the picture title vector and the semantic vector, filtering the first candidate image-text pair based on a second correlation degree threshold, and generating a second candidate image-text pair; processing the pictures in the second candidate image-text pair, taking the full text as background information, generating a corresponding picture description text, performing semantic similarity comparison on the corresponding picture description text and the related descriptive text, and outputting image-text pair data higher than a third similarity threshold value through the preset third similarity threshold value. The method can improve the control quality of literature image-text pairs.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Electronic device and method of controlling same

An electronic device includes memory storing instructions; and at least one processor, wherein the instructions, when executed, cause the electronic device to receive an input prompt; identify feature information in a first portion of a first generated image based on the input prompt; obtain a modified prompt by modifying the input prompt based on the feature information; obtain candidate images of a first image quality corresponding to the modified prompt using a first generative artificial intelligence (AI) model; display a user interface (UI) including the candidate images via a display; and based on a selected candidate image being identified from among the candidate images, obtain a second generated image of a second image quality corresponding to the selected candidate image, wherein a second image quality parameter of the second generated image is higher than a first image quality parameter of the first generated image.
Owner:SAMSUNG ELECTRONICS CO LTD

Gaming machine

To provide a gaming machine that enhances the enjoyment of the game by utilizing presentations that suggest reliability based on the player's choices. [Solution] In the pre-performance, one of several candidate images is selected as the pre-selected image, and in the post-performance, one of several candidate images is selected as the post-selected image. When the combined performance is executed in state A, the reliability is higher when the pre-selected image and post-selected image are of the same type than when they are of different types. When the performance is executed in state B, which is different from state A, the reliability is higher when the pre-selected image and post-selected image are of different types than when they are of the same type.
Owner:SANSEI R&D KK

Motif-based image classification

A method for displaying images similar to a selected image includes receiving, from a user, a selection of an anchor image, generating, using a machine learning model, an anchor embeddings set for the anchor image and respective candidate embeddings sets for a plurality of candidate images. The method also includes calculating a distance between the anchor embeddings set and each of the plurality of candidate embeddings sets and displaying at least one of the plurality of candidate images based on the calculated distance.
Owner:HOME DEPOT PRODUCT AUTHORITY LLC

Reference image providing apparatus, reference image providing method, reference image providing program, and reference image providing system

A reference image providing apparatus capable of reducing user's labor required for registration of a reference image, wherein the reference image providing apparatus includes an extractor configured to extract a reference image or a candidate image of the reference image from diagnostic images used for diagnosis, the reference image serving as an imaging reference for determining whether or not an image can be used for diagnosis; and a provider configured to provide an extraction result. For example, when the reference image is not set, the extractor extracts the reference image or the candidate image.
Owner:KONICA MINOLTA INC

Underground safety production simulation modeling method based on VR technology

The invention relates to the technical field of environment modeling, in particular to an underground safety production simulation modeling method based on the VR technology, and the method comprises the steps: obtaining all RGB images of an underground scene and corresponding depth images, and the attitude information and position coordinates of a camera when each frame of RGB image is shot, so as to generate a dense point cloud, and constructing a triangular mesh model; calculating the screening value of each point cloud, and obtaining the feature point cloud of each triangular patch; and obtaining all candidate images of each triangular patch, determining a mapping deviation value, a texture distortion degree and a quality evaluation value of each triangular patch in each candidate image, screening the candidate images of each triangular patch, performing texture mapping on the triangular mesh model, and importing the texture mapping result into a VR engine to construct a virtual underground environment. According to the method, the overall effect of texture mapping of the model can be improved, so that a highly vivid virtual underground scene is constructed, and the immersion and reality of underground scene safety production virtual practical training are enhanced.
Owner:MINGCHUANG HUIYUAN (GUIZHOU) MINE DESIGN & RES INST CO LTD

Gps-based visual positioning method, system, computer and storage medium

ActiveCN120765753BImage enhancementImage analysisComputer graphics (images)Geographical distance
The application provides a GPS-based visual positioning method, system, computer and storage medium, which comprises the following steps: acquiring the GPS coordinates of a query image, and screening reference images with close geographical positions from a database based on the GPS coordinates; screening reference images with a geographical distance less than or equal to a preset radius as a candidate set; determining several similar images based on k-nearest neighbor search; clustering according to the co-visibility relationship of each similar image to generate a plurality of locations corresponding to the candidate image, each location containing commonly observed 3D points; performing 2D-3D matching on each location, and estimating and outputting a six-degree-of-freedom camera pose through a perspective n-point algorithm and a random sample consensus algorithm to realize visual positioning. The GPS coordinates are used to dynamically screen reference images, the search time is reduced, the quick response requirement of real-time scene application is met, and the cross-environment high robustness is guaranteed.
Owner:JIANGXI QIUSHI INST OF ADVANCED STUDIES

Instruction processing method and device, electronic equipment and storage medium

The invention provides an instruction processing method and device, electronic equipment and a storage medium, and the method comprises the steps: responding to received input instruction information, and determining target image data from at least one piece of candidate image data according to the instruction information; wherein the at least one piece of candidate image data is image data shot at different visual angles; and based on the instruction information, performing analysis processing on the three-dimensional representation corresponding to the candidate image data and the target image data to obtain a reasoning result corresponding to the instruction information. Analysis is carried out by combining the three-dimensional representation of the candidate image data and the target image data, the two-dimensional and three-dimensional information of the image is fully utilized, and compared with a traditional processing mode which only depends on the two-dimensional image information, the image content can be understood more comprehensively and deeply, and the accuracy and reliability of a reasoning result are improved.
Owner:XIAOMI EV TECH CO LTD

Image detection method and device, computer device, storage medium and program product

This application discloses an image detection method, apparatus, computer device, storage medium, and program product, relating to the field of artificial intelligence. The method includes: extracting visual features from an input image; obtaining category text features corresponding to at least two candidate image categories, where the category text features are features used to characterize the candidate image categories from a textual perspective; obtaining category visual features corresponding to at least two candidate image categories, where the category visual features are features used to characterize the candidate image categories from an image perspective; and determining the image category to which the input image belongs from the at least two candidate image categories based on the visual features, the category text features corresponding to the at least two candidate image categories, and the category visual features. Using the method of this application can improve the accuracy of determining the image category to which the input image belongs.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD