Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

66 results about "Visual search tasks" patented technology

Visual search is a type of perceptual task requiring attention that typically involves an active scan of the visual environment for a particular object or feature (the target) among other objects or features (the distractors). Visual search can take place with or without eye movements.

Visual search experiment method for head fixing brain imaging

The invention relates to the technical field of visual search, in particular to a visual search experiment method for head fixing brain imaging. Comprising the following steps: fixing a target object on an experiment table, and training the target object through a display screen to obtain a trained target object; obtaining a target stimulus for an experiment and a target display strategy of the target stimulus in a display screen, the target stimulus being an icon for identification; and executing the target display strategy in the display screen, and obtaining behavioral statistical data and brain region microscopic imaging data of the trained target object based on the execution action of the trained target object. According to the method, the brain region microscopic imaging data of the trained target object can be obtained by executing the target display strategy corresponding to the target stimulation in the display screen, the display strategy of the target stimulation in the display screen is optimized, the motion noise interference of the target object is reduced, the brain region microscopic imaging data with higher quality is obtained, and the accuracy of the brain region microscopic imaging data is improved. And the comprehensiveness and the reliability of an experimental result are improved.
Owner:TSINGHUA UNIVERSITY

Visual guidance robot material box grabbing module

The invention relates to the technical field of robot visual guidance, and discloses a visual guidance robot material box grabbing module, which comprises a mechanical coarse positioning and state sensing unit for physically clamping a material box and collecting physical state data; a prediction type pre-positioning and dynamic reference generation unit receives the data, calls a historical deviation database, and generates a prediction target position and a temporary standard reference in parallel; the nested visual search and calibration unit moves to a prediction position and searches to obtain a complete feature image; the accurate deviation calculation and fine adjustment unit compares the image with a reference, calculates a final residual deviation and generates an accurate grabbing position; and the closed-loop learning and data feedback unit executes capturing and feeds back the final residual deviation to the historical deviation database. According to the method, a predicted target position and a dynamic temporary reference are established by using a physical state data vector and a historical deviation database, and parallel compensation of system errors and mechanical errors is realized.
Owner:WUXI JIANGLAN INTELLIGENT EQUIP CO LTD

Open vocabulary target detection method and system based on visual retrieval enhancement prompt

The invention relates to the technical field of computer vision, discloses a visual retrieval enhancement prompt-based open vocabulary target detection method and system, and aims to solve the problem of poor rare category detection performance caused by training data long tail distribution in open vocabulary target detection. The method comprises two stages of off-line construction of a visual concept knowledge base and on-line retrieval enhanced reasoning, wherein a large language model is utilized to generate discriminative visual descriptors for each category, and the discriminative visual descriptors are encoded by a CLIP text encoder and then stored in an FAISS vector database; during reasoning, global visual features of an input image are extracted through a CLIP image encoder, the most relevant descriptor in the knowledge base is retrieved and added to an original category name prompt, and an enhanced prompt is formed and input to an open vocabulary detector to complete detection. Experiments show that the method obviously improves the detection performance in COCO and LVIS benchmark tests, and especially obviously improves the detection effect on rare categories.
Owner:CHONGQING UNIV

Display method and system of electronic rearview mirror based on multi-mode risk field coupling

The invention discloses a display method and system of an electronic rearview mirror based on multi-mode risk field coupling. The display method comprises the steps that driving state data of a current vehicle and one or more target vehicles and observation point data of a current vehicle driver are obtained; constructing a coupling dynamic risk field model based on the driving state data, and generating a comprehensive risk value; generating an attention entropy and a visual search efficiency value based on the observation point data; generating a display control signal by the comprehensive risk value, the attention entropy and the visual search efficiency value through a visual information flow optimizer based on an optimal control theory; and according to the display control signal, simplifying the background region by adopting an anisotropic diffusion equation, enhancing the foreground target by utilizing a nonlinear rendering pipeline, generating an electronic rearview mirror synthetic picture, and displaying the electronic rearview mirror synthetic picture on an in-vehicle display screen. According to the invention, the prominent presentation of key risk information and the effective compression of redundant information are realized, the cognitive load of a driver is reduced, and the driving decision-making efficiency and safety are improved.
Owner:SAIC GM WULING AUTOMOBILE CO LTD

Identifying and localizing editorial changes to images utilizing deep learning

The present disclosure relates to systems, methods, and non-transitory computer readable media that utilize deep learning to identify regions of an image that have been editorially modified. For example, the image comparison system includes a deep image comparator model that compares a pair of images and localizes regions that have been editorially manipulated relative to an original or trusted image. More specifically, the deep image comparator model generates and surfaces visual indications of the location of such editorial changes on the modified image. The deep image comparator model is robust and ignores discrepancies due to benign image transformations that commonly occur during electronic image distribution. The image comparison system optionally includes an image retrieval model utilizes a visual search embedding that is robust to minor manipulations or benign modifications of images. The image retrieval model utilizes a visual search embedding for an image to robustly identify near duplicate images.
Owner:ADOBE INC +1

Visual search refinement

Described is a system and method for enabling visual search for information. With each selection of a search term, additional search terms are dynamically selected and presented to the user in conjunction with results matching the currently selected search terms. Likewise, a selected search term may be tokenized and a graphical token presented to the user to represent the selected search term.
Owner:PINTEREST INC

Digital supplement association and retrieval for visual search

Systems and methods for identification and retrieval of content for visual search are provided. An example method includes receiving data specifying a digital supplement. The data may identify a digital supplement and a supplement anchor for associating the digital supplement with visual content. The method may also include generating a data structure instance that specifies the digital supplement and the supplement anchor and, after generating the data structure instance, enabling triggering of the digital supplement by an image based at least on storing the data structure instance in a database that includes a plurality of other data structure instances. The other data structure instances may each specify a digital supplement and one or more supplement anchors.
Owner:GOOGLE LLC

Visual search refinement

Described is a system and method for enabling visual search for information. With each selection of a search term, additional search terms are dynamically selected and presented to the user in conjunction with results matching the currently selected search terms. Likewise, a selected search term may be tokenized and a graphical token presented to the user to represent the selected search term.
Owner:PINTEREST INC

Old people subthreshold depression cognitive deviation correction training system

The invention relates to the technical field of cognitive deviation correction training, in particular to an old people subthreshold depression cognitive deviation correction training system which comprises an image processing module, an image storage module, a point probe output module, a visual search output module, a response data statistics module, an analysis calibration module, a detection implantation module and a unified processing module. Determining an attention change value for each user based on the response delay time length, dividing an attention category of each user based on the attention change value, and determining whether to correct the doping degree of the probe stimulation identifier output by the identifier generation unit based on the attention category; when it is determined that the attention category of the single user is the weak attention category, determining whether the emotional picture output for the single user is qualified or not based on the attention change representation value; personalized adjustment is carried out according to the specific condition of the user, the emotion picture is adjusted according to the response condition of the user, and the processing efficiency of the correction training data is improved.
Owner:JILIN UNIVERSITY

Online re-judgment method and system based on large model and visual retrieval screening method

The invention provides an online re-judgment method and system based on a large model and a visual retrieval screening method, and belongs to the related technical field of artificial intelligence and industrial quality inspection.The online re-judgment method comprises the steps that an efficient visual model is constructed through pre-training and generation technologies, and the efficient visual model is deployed to a production line for trial operation; carrying out high-value data collection and labeling by utilizing an active learning and automatic labeling technology, carrying out efficient visual model iteration, and starting mass production operation; the difficult cases detected by the model are re-judged through an industrial quality inspection large model and a retrieval enhancement technology, and data which are not confidence in re-judgment are handled manually. According to the method and the system, more than 80% of investment of manual re-judgment can be reduced, and misjudgment caused by subjectivity of manual re-judgment can be remarkably reduced, so that remarkable cost reduction and benefit increase are brought to enterprises; and popularization and application in the fields of quality inspection, screening and the like are facilitated.
Owner:TZTEK TECHNOLOGY CO LTD

An online fiber visualized searching system

The application discloses an online fiber visual searching system, which comprises luminous electronic tags, a first tag reader and a second tag reader. The luminous electronic tags are provided with inductive coils and luminous LED lamps. The luminous electronic tags are arranged at two ends of optical fibers, and the luminous electronic tags at the two ends of the same optical fiber have corresponding marks. The first tag reader is used for reading a single luminous electronic tag at a time. The second tag reader is used for reading a single or multiple luminous electronic tags at a time. The first tag reader and the second tag reader are in communication connection, and the first tag reader transmits tag information to the second tag reader after reading the tag information. The application has the advantages of realizing online optical fiber identification and being beneficial to fiber information management.
Owner:NANJING HUAMAI TECH

Attention state detection device

PendingCN122056594APsychotechnic devicesSensorsVisual technologyAttention Concentration
The invention discloses an attention state detection device, and relates to the technical field of computer vision. In the device, eye movement characteristics are acquired through a first acquisition module, then behavior characteristics are determined based on the eye movement characteristics in a first determination module, and finally, the attention state is determined based on a pre-established attention stage, a pre-established mapping relation between the behavior characteristics and a threshold value and the current behavior characteristics in a second determination module. In the attention detection device, the attention is divided into a plurality of stages in advance, such as a visual search characterization stage, a sight orientation characterization stage, an attention maintenance characterization stage and an attention concentration characterization stage, and a mapping relationship among the attention stages, behavior characteristics and threshold values is pre-established, so that the attention detection accuracy is improved. According to the invention, the existing attention stage and the missing attention stage can be determined based on the mapping relation and the current behavior characteristics, so that the user can know the specific stage where the attention is missing, and the accurate judgment of the attention state is realized.
Owner:ANYANG XIANGYU MEDICAL EQUIP

Visual search method with feedback loop based on interactive sketch

The disclosure relates to the fields of electronic processing of images and e-commerce, specifically—to the visual search with feedback loop based on interactive sketch for transparency and user control. Presented method adds transparency and user control to the traditionally “black-box” visual search process by introducing a feedback loop, based on interactive sketch. The method comprises of following main steps: submission of an information about the object being searched for by the user (in a form of image, written text, spoken text or selection of product category from menu in GUI), identification of product attributes and verification of identified attributes against the ontology rules, selection of building blocks from database for filtered attributes, connection of selected blocks to form a sketch of object being searched for, application of projection rules to produce a form vector, comparison of vector from previous step with database, presenting user with interactive sketch of product alongside the list of matching products from database, enabling user to select a part of the object in sketch, presenting user with possible modifications for the selected part of the object, enabling user to select needed modification and re-iterate the search process as many times as needed, until the wanted results are achieved.
Owner:GASEFIS UAB

Long document understanding method and device, electronic equipment and storage medium

The invention discloses a long document understanding method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the layout analysis of a target document through a layout analysis model, and obtaining a layout analysis result; wherein the layout analysis result at least comprises coordinates and types of the element areas; cutting an element sub-graph based on the coordinates of each element region; generating a text query instruction by using the current query input and the layout analysis result every time the query input is received; inputting the text query instruction and all the element sub-graphs into a visual searcher for visual search to obtain a plurality of key element sub-graphs associated with the current query input and the current search task type; performing document understanding on each key element sub-graph based on the text cue word through a multi-model large model to obtain a current document understanding result; wherein the text cue word is generated at least based on the current query input and the current search task type; and utilizing the current document to understand and output a feedback result of the current query input.
Owner:ASIAINFO TECH CHINA INC

Automatic conference spokesman positioning method, system and device and medium

The invention relates to the technical field of intelligent positioning, in particular to a conference spokesman automatic positioning method, system and device and a medium. The method comprises the following steps: driving a microphone array to collect training sound data, and controlling a pan-tilt camera to spirally scan a conference room speaking area to collect training image data; when training human voice and lip micro-motion data are synchronously detected, extracting three-dimensional coordinates of a sound source and recording current shooting parameters (at least one of a horizontal rotation angle, a pitch angle and a focal length) of a camera to form a group of training samples; not less than 30 groups of samples are collected and input into a neural network for training, a coordinate transformation matrix of the camera relative to the microphone array is obtained, the microphone array obtains real-time sound source coordinates, visual search center coordinates are generated through the transformation matrix, and accordingly the camera is driven to complete spokesman close-up shooting. The method can be completed without manual intervention.
Owner:SHENZHEN MINRRAY IND CORP LTD

A method and apparatus for supporting multi-time axis visual search and playback

The present application relates to the technical field of video retrieval and playback, and provides a method and device supporting multi-time axis visual retrieval and playback. Each terminal of the present application respectively collects, encodes and stores the fragmented data files from multi-channel audio and video signals, and sends the fragmented data files to a client through key frame mapping; the client performs multi-source time axis aggregation indexing on the fragmented data files from each terminal, matches the playback time of each channel of audio and video signals specified by a user, obtains the closest key frame information, and sends the corresponding terminal; the terminal outputs the file stream of the fragmented data file to the client from the specified key frame according to the key frame information, and the client receives each channel of file stream and decodes and plays back. The present application solves the problems that the prior art cannot intuitively obtain the audio and video recording situation of a channel or multiple channels at a certain time point, and cannot select and directly start single-channel or multi-channel fast positioning and playing from a certain time point.
Owner:709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD

Application prediction based on a visual search determination

Visual search in an operating system of a computing device can process and provide additional information on the content being provided for display. The computing device can include an operating system that includes a visual search interface that obtains and processes display data associated with content currently being provided for display. The visual search interface can generate display data based on the current content provided for display, process the display data with one or more on-device machine-learned models, and provide additional information to the user. The visual search interface may transmit data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.
Owner:GOOGLE LLC

Visual search and discovery via generative model inversion

Solutions for visual search and discovery include performing unsupervised training of a generative adversarial network that has a generator and an assessor. Training the generative adversarial network involves alternating training the assessor with the generator and a plurality of catalog images with training the generator with the assessor. The catalog images are inverted into catalog vectors by leveraging the trained generator. A query image is inverted into a query vector, and image similarity is determined by calculating a distance between the query vector and a catalog vector. In some examples, inversion is performed by training an encoder with the trained generator and inverting the catalog images with the encoder. In some examples, the trained generator is used to perform a search in a vector space. A weighting vector may be used to weight elements of the vectors, effectively prioritizing image features for image similarity determination.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Display board for visual training and use method thereof

The invention relates to the technical field of medical instrument engineering, and discloses a display board for visual training and a using method thereof.The display board comprises a board body, the board body is provided with a board containing area and a game area separately, a board containing groove is formed in the board body, a loudspeaker module is arranged on the outer wall of the board body, LED lamp beads are arranged in the board body, and the LED lamp beads are arranged in the board body. A supporting plate is fixedly connected to the interior of the board, a first motor, a limiting seat and a vibration assembly are fixedly connected to the upper surface of the supporting plate, a first gear is arranged at the output end of the first motor, a second gear is connected to the tooth end of the first gear in a meshed mode, and a connecting shaft is fixedly connected to the interior of the second gear in a penetrating mode. The game area is matched with the card placing frame to rotate to adjust the direction of the card body, so that the visual search training difficulty is effectively improved; the pressure sensor monitors the pick-and-place state of the board body in real time, the LED lamp beads provide traffic light feedback according to the matching result, multi-mode interactive training is achieved jointly, and the guidance and feedback timeliness of the training process are enhanced.
Owner:烟台畅源健康科技有限公司

Dynamic search control invocation and visual search

Described is a system and method for enabling dynamic selection of a search input. For example, rather than having a static search input box, the search input may be dynamically positioned such that it encompasses a portion of displayed information. For example, a user may touch a touch-based display using two fingers to invoke the dynamic search input and then determine a size and a position of the dynamic search input by moving their fingers on the display. An image segment that includes a representation of the encompassed portion of the displayed information is generated and processed to determine an object represented in the portion of the displayed information. Additional images with visually similar representations of objects are then determined and presented to the user.
Owner:PINTEREST INC

Visual search request initiating method, visual search method and electronic equipment

The embodiment of the invention discloses a visual search request initiating method, a visual search method and electronic equipment, and the method comprises the steps: starting the perception of an element dragging event after the loading of a visual search function of a search component in a commodity information page is completed; after sensing an element dragging event, if a dragged target element is related to visual data, activating an operation panel corresponding to the visual search function; and after the target element is dragged to the target area in the operation panel, submitting a search request for commodity search based on the visual data content corresponding to the target element. Through the embodiment of the invention, the operation complexity of initiating the visual search request can be reduced, and the probability of intention interruption caused by operation failure is reduced.
Owner:HANGZHOU ALIBABA INT INTERNET IND CO LTD

Visual search interface in operating system

A visual search in an operating system of a computing device may process providing content for display and providing additional information about the content. The computing device may include an operating system including a visual search interface that retrieves display data associated with content currently provided for display and processes the display data. The visual search interface may generate display data based on current content provided for display, process the display data using one or more on-device machine learning models, and provide additional information to the user. The visual search interface may send data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.
Owner:GOOGLE LLC

Visual search response generation based on intent determination

Systems and methods for unimodal visual search can include obtaining an image query, processing the image query to generate an intent determination, determining a particular prompt based on the intent determination, determining a plurality of search results, and processing the particular prompt, the image query, and at least a subset of the plurality of search results to generate a model-generated response. A generation model can render different responses based on different prompts related to different intents.
Owner:GOOGLE LLC

Depth-based visual search area

In one implementation, a method of extracting information from a physical environment is performed at a device including an image sensor, one or more processors, and non-transitory memory. The method includes determining a gaze location and a distance to an object in a physical environment at the gaze location. The method includes selecting a field-of-view of the physical environment based on the gaze location and the distance to the object. The method includes obtaining, using the image sensor, an image corresponding to the field-of-view of the physical environment. The method includes extracting information from the image.
Owner:APPLE INC

Multi-section low-carbon vehicle routing problem optimization method based on improved fruit fly optimization algorithm

The embodiment of the invention relates to the technical field of logistics distribution management, and particularly discloses a multi-section low-carbon vehicle routing problem optimization method based on an improved fruit fly optimization algorithm. According to the embodiment of the invention, the method comprises the steps: generating an initialized fruit fly population through obtaining logistics distribution data; carrying out natural number coding; performing olfaction search simulation on the initialized fruit fly population, and selecting a plurality of first candidate individuals; performing visual search simulation on the initialized fruit fly population, and selecting a plurality of second candidate individuals; dividing a plurality of fruit fly sub-populations, and performing information interaction, cooperation and competition; and when a termination condition is satisfied, outputting an optimal distribution route scheme and performing visual display. According to the method, the fruit fly optimization algorithm can be improved, and the multi-section low-carbon vehicle routing problem is optimized, so that the distribution vehicles are reasonably dispatched, the distribution route is optimized, the fuel consumption can be effectively reduced, the fuel cost and carbon emission in logistics distribution are reduced, the logistics cost of a distribution company is finally reduced, and the social benefit is improved.
Owner:JIAN COLLEGE +1

Visual search for electronic devices

The subject technology provides visual search systems and methods that can be used to efficiently perform visual searches on an electronic device. The subject technology provides systems and methods for presenting one or more visual indicators corresponding to the searchable portions of digital content. A visual search may include identifying, at an electronic device, an element of interest in an image. A visual indicator for the element of interest may be overlaid on the image at a location corresponding to the element of interest. The visual indicator may be selectable to cause a display of information associated with the element of interest responsive to the selection.
Owner:APPLE INC

Visual search refinement in computer-generated rendering environments

This disclosure relates to the refinement of visual search in computer-generated rendering environments (CGR). Various disclosed embodiments include devices, systems, and methods capable of faster and more efficient real-time physical object identification, information retrieval, and CGR environment updates. In some embodiments, providing the CGR environment at a first device based on the classification of the physical objects, image, or video data includes: transmitting the physical objects from the first device to a second device, and updating the CGR environment by the first device based on responses associated with the physical objects received from the second device.
Owner:APPLE INC

Visual search-oriented biological brain neuron activation mechanism analysis method and device

The invention provides a biological brain neuron activation mechanism analysis method and device oriented to visual search. A mouse is taken as a model animal, (1) a neuron activation analysis method based on target stimulation correlation is provided, and the representation problem of a sparse activation mechanism of visual cortical neurons is solved; (2) a neuron activity pattern analysis method based on mouse visual brain region function hierarchy division is provided, and the problem of representation of a hierarchical activation mechanism of visual cortical neurons is solved; and (3) a contrast experiment of background noise is introduced into visual stimulation, so that the representation problem of a visual cortical neuron activation distribution rule under the influence of the background stimulation noise is solved.
Owner:TSINGHUA UNIVERSITY