Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

36 results about "Visual search tasks" patented technology

Visual search is a type of perceptual task requiring attention that typically involves an active scan of the visual environment for a particular object or feature (the target) among other objects or features (the distractors). Visual search can take place with or without eye movements.

Open vocabulary target detection method and system based on visual retrieval enhancement prompt

The invention relates to the technical field of computer vision, discloses a visual retrieval enhancement prompt-based open vocabulary target detection method and system, and aims to solve the problem of poor rare category detection performance caused by training data long tail distribution in open vocabulary target detection. The method comprises two stages of off-line construction of a visual concept knowledge base and on-line retrieval enhanced reasoning, wherein a large language model is utilized to generate discriminative visual descriptors for each category, and the discriminative visual descriptors are encoded by a CLIP text encoder and then stored in an FAISS vector database; during reasoning, global visual features of an input image are extracted through a CLIP image encoder, the most relevant descriptor in the knowledge base is retrieved and added to an original category name prompt, and an enhanced prompt is formed and input to an open vocabulary detector to complete detection. Experiments show that the method obviously improves the detection performance in COCO and LVIS benchmark tests, and especially obviously improves the detection effect on rare categories.
Owner:CHONGQING UNIV

Display method and system of electronic rearview mirror based on multi-mode risk field coupling

The invention discloses a display method and system of an electronic rearview mirror based on multi-mode risk field coupling. The display method comprises the steps that driving state data of a current vehicle and one or more target vehicles and observation point data of a current vehicle driver are obtained; constructing a coupling dynamic risk field model based on the driving state data, and generating a comprehensive risk value; generating an attention entropy and a visual search efficiency value based on the observation point data; generating a display control signal by the comprehensive risk value, the attention entropy and the visual search efficiency value through a visual information flow optimizer based on an optimal control theory; and according to the display control signal, simplifying the background region by adopting an anisotropic diffusion equation, enhancing the foreground target by utilizing a nonlinear rendering pipeline, generating an electronic rearview mirror synthetic picture, and displaying the electronic rearview mirror synthetic picture on an in-vehicle display screen. According to the invention, the prominent presentation of key risk information and the effective compression of redundant information are realized, the cognitive load of a driver is reduced, and the driving decision-making efficiency and safety are improved.
Owner:SAIC GM WULING AUTOMOBILE CO LTD

Visual search refinement

Described is a system and method for enabling visual search for information. With each selection of a search term, additional search terms are dynamically selected and presented to the user in conjunction with results matching the currently selected search terms. Likewise, a selected search term may be tokenized and a graphical token presented to the user to represent the selected search term.
Owner:PINTEREST INC

Digital supplement association and retrieval for visual search

ActiveUS12579194B2Multimedia data indexingWeb data retrievalData miningVisual search tasks
Systems and methods for identification and retrieval of content for visual search are provided. An example method includes receiving data specifying a digital supplement. The data may identify a digital supplement and a supplement anchor for associating the digital supplement with visual content. The method may also include generating a data structure instance that specifies the digital supplement and the supplement anchor and, after generating the data structure instance, enabling triggering of the digital supplement by an image based at least on storing the data structure instance in a database that includes a plurality of other data structure instances. The other data structure instances may each specify a digital supplement and one or more supplement anchors.
Owner:GOOGLE LLC

Online re-judgment method and system based on large model and visual retrieval screening method

The invention provides an online re-judgment method and system based on a large model and a visual retrieval screening method, and belongs to the related technical field of artificial intelligence and industrial quality inspection.The online re-judgment method comprises the steps that an efficient visual model is constructed through pre-training and generation technologies, and the efficient visual model is deployed to a production line for trial operation; carrying out high-value data collection and labeling by utilizing an active learning and automatic labeling technology, carrying out efficient visual model iteration, and starting mass production operation; the difficult cases detected by the model are re-judged through an industrial quality inspection large model and a retrieval enhancement technology, and data which are not confidence in re-judgment are handled manually. According to the method and the system, more than 80% of investment of manual re-judgment can be reduced, and misjudgment caused by subjectivity of manual re-judgment can be remarkably reduced, so that remarkable cost reduction and benefit increase are brought to enterprises; and popularization and application in the fields of quality inspection, screening and the like are facilitated.
Owner:TZTEK TECHNOLOGY CO LTD

An online fiber visualized searching system

The application discloses an online fiber visual searching system, which comprises luminous electronic tags, a first tag reader and a second tag reader. The luminous electronic tags are provided with inductive coils and luminous LED lamps. The luminous electronic tags are arranged at two ends of optical fibers, and the luminous electronic tags at the two ends of the same optical fiber have corresponding marks. The first tag reader is used for reading a single luminous electronic tag at a time. The second tag reader is used for reading a single or multiple luminous electronic tags at a time. The first tag reader and the second tag reader are in communication connection, and the first tag reader transmits tag information to the second tag reader after reading the tag information. The application has the advantages of realizing online optical fiber identification and being beneficial to fiber information management.
Owner:NANJING HUAMAI TECH

Attention state detection device

PendingCN122056594APsychotechnic devicesSensorsVisual technologyAttention Concentration
The invention discloses an attention state detection device, and relates to the technical field of computer vision. In the device, eye movement characteristics are acquired through a first acquisition module, then behavior characteristics are determined based on the eye movement characteristics in a first determination module, and finally, the attention state is determined based on a pre-established attention stage, a pre-established mapping relation between the behavior characteristics and a threshold value and the current behavior characteristics in a second determination module. In the attention detection device, the attention is divided into a plurality of stages in advance, such as a visual search characterization stage, a sight orientation characterization stage, an attention maintenance characterization stage and an attention concentration characterization stage, and a mapping relationship among the attention stages, behavior characteristics and threshold values is pre-established, so that the attention detection accuracy is improved. According to the invention, the existing attention stage and the missing attention stage can be determined based on the mapping relation and the current behavior characteristics, so that the user can know the specific stage where the attention is missing, and the accurate judgment of the attention state is realized.
Owner:ANYANG XIANGYU MEDICAL EQUIP

Visual search method with feedback loop based on interactive sketch

The disclosure relates to the fields of electronic processing of images and e-commerce, specifically—to the visual search with feedback loop based on interactive sketch for transparency and user control. Presented method adds transparency and user control to the traditionally “black-box” visual search process by introducing a feedback loop, based on interactive sketch. The method comprises of following main steps: submission of an information about the object being searched for by the user (in a form of image, written text, spoken text or selection of product category from menu in GUI), identification of product attributes and verification of identified attributes against the ontology rules, selection of building blocks from database for filtered attributes, connection of selected blocks to form a sketch of object being searched for, application of projection rules to produce a form vector, comparison of vector from previous step with database, presenting user with interactive sketch of product alongside the list of matching products from database, enabling user to select a part of the object in sketch, presenting user with possible modifications for the selected part of the object, enabling user to select needed modification and re-iterate the search process as many times as needed, until the wanted results are achieved.
Owner:GASEFIS UAB

Automatic conference spokesman positioning method, system and device and medium

The invention relates to the technical field of intelligent positioning, in particular to a conference spokesman automatic positioning method, system and device and a medium. The method comprises the following steps: driving a microphone array to collect training sound data, and controlling a pan-tilt camera to spirally scan a conference room speaking area to collect training image data; when training human voice and lip micro-motion data are synchronously detected, extracting three-dimensional coordinates of a sound source and recording current shooting parameters (at least one of a horizontal rotation angle, a pitch angle and a focal length) of a camera to form a group of training samples; not less than 30 groups of samples are collected and input into a neural network for training, a coordinate transformation matrix of the camera relative to the microphone array is obtained, the microphone array obtains real-time sound source coordinates, visual search center coordinates are generated through the transformation matrix, and accordingly the camera is driven to complete spokesman close-up shooting. The method can be completed without manual intervention.
Owner:SHENZHEN MINRRAY IND CORP LTD

Application prediction based on a visual search determination

Visual search in an operating system of a computing device can process and provide additional information on the content being provided for display. The computing device can include an operating system that includes a visual search interface that obtains and processes display data associated with content currently being provided for display. The visual search interface can generate display data based on the current content provided for display, process the display data with one or more on-device machine-learned models, and provide additional information to the user. The visual search interface may transmit data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.
Owner:GOOGLE LLC

Visual search and discovery via generative model inversion

Solutions for visual search and discovery include performing unsupervised training of a generative adversarial network that has a generator and an assessor. Training the generative adversarial network involves alternating training the assessor with the generator and a plurality of catalog images with training the generator with the assessor. The catalog images are inverted into catalog vectors by leveraging the trained generator. A query image is inverted into a query vector, and image similarity is determined by calculating a distance between the query vector and a catalog vector. In some examples, inversion is performed by training an encoder with the trained generator and inverting the catalog images with the encoder. In some examples, the trained generator is used to perform a search in a vector space. A weighting vector may be used to weight elements of the vectors, effectively prioritizing image features for image similarity determination.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Dynamic search control invocation and visual search

Described is a system and method for enabling dynamic selection of a search input. For example, rather than having a static search input box, the search input may be dynamically positioned such that it encompasses a portion of displayed information. For example, a user may touch a touch-based display using two fingers to invoke the dynamic search input and then determine a size and a position of the dynamic search input by moving their fingers on the display. An image segment that includes a representation of the encompassed portion of the displayed information is generated and processed to determine an object represented in the portion of the displayed information. Additional images with visually similar representations of objects are then determined and presented to the user.
Owner:PINTEREST INC

Visual search interface in operating system

A visual search in an operating system of a computing device may process providing content for display and providing additional information about the content. The computing device may include an operating system including a visual search interface that retrieves display data associated with content currently provided for display and processes the display data. The visual search interface may generate display data based on current content provided for display, process the display data using one or more on-device machine learning models, and provide additional information to the user. The visual search interface may send data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.
Owner:GOOGLE LLC

Visual search response generation based on intent determination

Systems and methods for unimodal visual search can include obtaining an image query, processing the image query to generate an intent determination, determining a particular prompt based on the intent determination, determining a plurality of search results, and processing the particular prompt, the image query, and at least a subset of the plurality of search results to generate a model-generated response. A generation model can render different responses based on different prompts related to different intents.
Owner:GOOGLE LLC

Depth-based visual search area

In one implementation, a method of extracting information from a physical environment is performed at a device including an image sensor, one or more processors, and non-transitory memory. The method includes determining a gaze location and a distance to an object in a physical environment at the gaze location. The method includes selecting a field-of-view of the physical environment based on the gaze location and the distance to the object. The method includes obtaining, using the image sensor, an image corresponding to the field-of-view of the physical environment. The method includes extracting information from the image.
Owner:APPLE INC

Visual search for electronic devices

The subject technology provides visual search systems and methods that can be used to efficiently perform visual searches on an electronic device. The subject technology provides systems and methods for presenting one or more visual indicators corresponding to the searchable portions of digital content. A visual search may include identifying, at an electronic device, an element of interest in an image. A visual indicator for the element of interest may be overlaid on the image at a location corresponding to the element of interest. The visual indicator may be selectable to cause a display of information associated with the element of interest responsive to the selection.
Owner:APPLE INC

Multi-scale visual language search method for open vocabulary target detection

The invention discloses a multi-scale visual language search method for open vocabulary target detection. The method comprises the following steps: firstly, designing a multi-scale visual search mechanism based on a class-independent candidate box, performing multi-scale visual search on an input image by taking the candidate box as a main body, and simultaneously sensing a target main body and local and surrounding contexts thereof on multiple scales; secondly, a vision-language soft alignment method based on a pre-trained vision language model is provided, multi-scale vision features and an attribute description code table of open vocabularies are jointly mapped into a unified cross-modal feature space, soft alignment is achieved through similarity calculation, and therefore key semantic clues of a target body and contexts of the target body are generated; and the visual feedback semantic information is used for subsequent decision making as interpretable visual feedback semantic information.
Owner:BEIJING INST OF TECH

Photo operation display section of character visual search graphical user interface of electronic device

1. Name of the product in this design: Photo operation display unit of a graphical user interface for visual search of people in electronic devices. 2. Purpose of this design: An electronic device. 3. The key design features of this product are the parts of the graphical user interface that require protection. 4. The image or photograph that best illustrates the design's key features: the front view. 5. Purpose of the graphical user interface: The overall interface is a graphical user interface used for visual search of people, while the part of the interface is the photo operation display area. The top of the interface displays a circular area and a rounded rectangular area surrounding the circular area. The circular area displays thumbnails, which are usually the user's thumbnails. When a user touches a portion of the screen, the user can select an image thumbnail from a pre-stored collection of photo images, and the corresponding photo image will be displayed as a slideshow in a rectangular area. 6. Other situations requiring explanation: The dotted lines in the diagram represent content for which protection is not sought.
Owner:SAMSUNG ELECTRONICS CO LTD

Intelligent system and method for visual search queries

A user can submit a visual query that includes one or more images. Within or in connection with the visual query, various processing techniques, such as optical character recognition (OCR) techniques, can be used to recognize text (e.g., in the image(s), around the image(s), etc.) and / or various object detection techniques (e.g., machine learning object detection models, etc.) can be used to detect objects (e.g., products, landmarks, animals, humans, etc.). Content related to the detected text or object(s) can be identified and potentially provided to the user as search results or proactive content feeds. Thus, various aspects of the present disclosure enable visual search systems to more intelligently process visual queries to provide improved search results and content feeds, including those that are more personalized and / or take into account implicit features of the visual query and / or user search intent.
Owner:GOOGLE LLC

Method, system and non-transitory computer-readable recording medium for searching similar products using a multi task learning model

A method of searching for similar products using a multi-task learning (MTL) model is provided. The method includes converting, by using a multi-task learning model utilizing a unified backbone network, each of a plurality of original images including a fashion item into a single vector including at least two multi-task attributes, and generating a visual search database by storing the original image and the single vector of the original image together.
Owner:MUSINSA CO LTD

Digital Supplement Association and Retrieval for Visual Search

PendingUS20260187154A1Data miningVisual search tasks
Systems and methods for identification and retrieval of content for visual search are provided. An example method includes receiving data specifying a digital supplement. The data may identify a digital supplement and a supplement anchor for associating the digital supplement with visual content. The method may also include generating a data structure instance that specifies the digital supplement and the supplement anchor and, after generating the data structure instance, enabling triggering of the digital supplement by an image based at least on storing the data structure instance in a database that includes a plurality of other data structure instances. The other data structure instances may each specify a digital supplement and one or more supplement anchors.
Owner:GOOGLE LLC

Visual field rehabilitation training method, device, system, storage medium and program product

The application relates to the technical field of healthcare information, and provides a visual field rehabilitation training method, device and system, a storage medium and a program product, which can improve visual neglect of a target object. According to an eye movement heat map of the target object, a visual low-attention area of the target object is determined; a virtual scene is created, and a virtual object matching the virtual scene is generated in the visual low-attention area; the virtual object is used for visual search training of the target object; according to an upper limit value of a direction integration angle corresponding to the visual low-attention area, an evaluation baseline corresponding to the visual low-attention area is determined; the higher the upper limit value of the direction integration angle, the higher the evaluation baseline; according to the evaluation baseline, visual search training of the target object on the virtual object in the visual low-attention area is evaluated to obtain an evaluation result; and the virtual object is adjusted according to the evaluation result to perform the next visual search training.
Owner:SHANGHAI UNITED IMAGING RES INST OF INTELLIGENT IMAGING +1

Skier video snapshot method and system based on RFID and vision cooperation

The invention relates to the technical field of computer vision and radio frequency identification, particularly discloses a skier video snapshot method and system based on RFID and vision cooperation, and aims to solve the technical problems of missing shooting, excessive shooting, high false detection rate and insufficient accuracy of skier video automatic snapshot in a complex environment of a snowfield. The method comprises the following steps: acquiring motion parameters such as instantaneous speed and acceleration of a skier through an RFID card reader array deployed along a slideway, predicting a time node when the skier arrives at an optimal ROI (Region of Interest) region of a camera view, starting visual identification based on a YOLO V5 model before and after the node, capturing a video, and binding the video with an RFID tag code; the system comprises an RFID card reader array, an electronic tag, a high-definition camera device and a plurality of functional modules, can dynamically adjust an ROI (Region of Interest) and a visual search window, and is equipped with a standby monitoring mechanism. According to the invention, space-time cooperation of RFID and visual technology is realized, snapshot accuracy is significantly improved, missed shooting rate and multi-shooting rate are reduced, congestion analysis can be provided for scenic spots, and value-added services are provided for skiers.
Owner:HUIPAI INTELLIGENT COMPUTING (HANGZHOU) TECHNOLOGY CO LTD

Pet-friendly hair drier

The pet-friendly hair drier comprises a hair drier main body which is provided with a visual friendly surface; wherein the visual friendly surface is arranged at an air outlet of the hair dryer main body. And the visual friendly surface is provided with a specular reflection area. The visual friendly surface has a main view direction; and the main view direction is parallel to the airflow direction of the air outlet of the blower. The hair dryer has the beneficial effects that pet-friendly elements such as the mirror surface are designed on the hair dryer, so that the unfamiliar feeling of a pet to the hair dryer can be reduced. The pet-friendly element is arranged at the air outlet of the hair drier, particularly, the hair drier usually has sound and easily frightens the pet, the pet-friendly element is arranged at the air outlet of the hair drier, the pet-friendly element has special significance, and the pet-friendly element at least has sound and wind feelings given to the pet by the hair drier, and when the pet perceives the existence of the hair drier through the sound or the wind feelings, the pet can be frightened. The air outlet position is the large probability of the area seen by the instinctive visual search. And the mirror surface arranged at the position enables the pet to see himself / herself, so that the pet mirror has relatively strong familiarity and is pet-friendly.
Owner:谢城玲

Visual Search Pivot Generation

In accordance with techniques for visual search pivot generation, a visual search request is received to trigger a visual search for items that are visually similar to a seed item. Using a machine learning model, one or more pivots representing visual attribute values for refining the visual search are generated based on information associated with the seed item. The one or more pivots are communicated for display in a user interface, and a user selection of a pivot is received. In response, at least one item is communicated for display in the user interface that is visually similar to the seed item and has a visual attribute value corresponding to the pivot.
Owner:EBAY INC

Spatial visual search method for long video multi-mode perception

The invention discloses a spatial visual search method for long video multi-mode perception. The spatial visual search method comprises the following steps: a space-time projection step: projecting an input long video sequence into a two-dimensional spatial grid image; a spatial reasoning step: performing semantic matching on the two-dimensional spatial grid image and an input natural language query by using a pre-trained video-language model to predict a grid unit index which is most matched with the natural language query, and converting the grid unit index into a corresponding time sequence window through an inverse mapping function; an iteration refining step: adopting an iteration layering refining mechanism to carry out recursion processing on the time sequence window; and a double-path switching step: introducing a duration sensing double-path mechanism. According to the method, a time sequence positioning task is converted into a spatial visual search problem through space-time projection, and a duration sensing double-path adaptive mechanism is introduced, so that efficient, low-overhead, zero-sample and robust long video event positioning is realized.
Owner:LANZHOU UNIV

Systems and methods for masking query input driven visual search

An image masking system (100) is disclosed. The system (100) includes a segmentation module (130) configured to extract one or more regions of interest by processing an input query image using a segmentation network. The segmentation module (130) is further configured to select one or more segments from the one or more extracted regions of interest based on a user indicator. The system (100) includes a mask generation module (140) configured to generate a binary mask for the selected segments on the input query image. The system (100) further includes an image transformer module (150) configured to generate a masked image by performing an arithmetic operation on the binary mask and the input query image. The image transformer module (150) is further configured to extract an embedding of the masked image using one of a convolutional neural network architecture or a Transformer architecture.
Owner:ROBERT BOSCH GMBH +1