Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

89 results about "Visual search tasks" patented technology

Visual search is a type of perceptual task requiring attention that typically involves an active scan of the visual environment for a particular object or feature (the target) among other objects or features (the distractors). Visual search can take place with or without eye movements.

Visual search interface in an operating system

Visual search in an operating system of a computing device can process and provide additional information on the content being provided for display. The computing device can include an operating system that includes a visual search interface that obtains and processes display data associated with content currently being provided for display. The visual search interface can generate display data based on the current content provided for display, process the display data with one or more on-device machine-learned models, and provide additional information to the user. The visual search interface may transmit data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.
Owner:GOOGLE LLC

Visual and text search interface for text-based video editing

Embodiments of the present invention provide systems, methods, and computer storage media for a visual and text search interface used to navigate a video transcript. In an example embodiment, a freeform text query triggers a visual search for frames of a loaded video that match the freeform text query (e.g., frame embeddings that match a corresponding embedding of the freeform query), and triggers a text search for matching words from a corresponding transcript or from tags of detected features from the loaded video. Visual search results are displayed (e.g., in a row of tiles that can be scrolled to the left and right), and textual search results are displayed (e.g., in a row of tiles that can be scrolled up and down). Selecting (e.g., clicking or tapping on) a search result tile navigates a transcript interface to a corresponding portion of the transcript.
Owner:ADOBE INC

Multifunctional mobile platform for detecting, interfering, trapping, striking and picking up unmanned aerial vehicle

The utility model discloses a multifunctional mobile platform for detecting, interfering, trapping, striking and picking up an unmanned aerial vehicle. The multifunctional mobile platform comprises a mobile platform and portable reconnaissance and striking equipment which is arranged on the mobile platform and is used for striking the unmanned aerial vehicle, the mobile platform is provided with a positioning navigation antenna for autonomous navigation; the mobile platform is provided with a navigation laser radar for ground environment detection. An unmanned aerial vehicle countering antenna and an unmanned aerial vehicle countering controller are arranged on the mobile platform, and the unmanned aerial vehicle countering controller is in communication connection with the unmanned aerial vehicle countering antenna; a pickup unit is arranged on the moving platform, the pickup unit comprises a mechanical arm mounted on the moving platform, a pickup clamping jaw assembled at the end of the mechanical arm and a 3D camera, and the pickup clamping jaw, the 3D camera and the mechanical arm are in communication connection with the unmanned aerial vehicle countering controller. According to the utility model, the unmanned aerial vehicle countering controller sends a control instruction to the 3D camera and carries out visual search on the forced landing unmanned aerial vehicle; and a control instruction is sent to the mechanical arm and the picking clamping jaw to pick the unmanned aerial vehicle.
Owner:RES INST OF PHYSICAL & CHEM ENG OF NUCLEAR IND

Visual search experiment method for head fixing brain imaging

The invention relates to the technical field of visual search, in particular to a visual search experiment method for head fixing brain imaging. Comprising the following steps: fixing a target object on an experiment table, and training the target object through a display screen to obtain a trained target object; obtaining a target stimulus for an experiment and a target display strategy of the target stimulus in a display screen, the target stimulus being an icon for identification; and executing the target display strategy in the display screen, and obtaining behavioral statistical data and brain region microscopic imaging data of the trained target object based on the execution action of the trained target object. According to the method, the brain region microscopic imaging data of the trained target object can be obtained by executing the target display strategy corresponding to the target stimulation in the display screen, the display strategy of the target stimulation in the display screen is optimized, the motion noise interference of the target object is reduced, the brain region microscopic imaging data with higher quality is obtained, and the accuracy of the brain region microscopic imaging data is improved. And the comprehensiveness and the reliability of an experimental result are improved.
Owner:TSINGHUA UNIVERSITY

Visual guidance robot material box grabbing module

The invention relates to the technical field of robot visual guidance, and discloses a visual guidance robot material box grabbing module, which comprises a mechanical coarse positioning and state sensing unit for physically clamping a material box and collecting physical state data; a prediction type pre-positioning and dynamic reference generation unit receives the data, calls a historical deviation database, and generates a prediction target position and a temporary standard reference in parallel; the nested visual search and calibration unit moves to a prediction position and searches to obtain a complete feature image; the accurate deviation calculation and fine adjustment unit compares the image with a reference, calculates a final residual deviation and generates an accurate grabbing position; and the closed-loop learning and data feedback unit executes capturing and feeds back the final residual deviation to the historical deviation database. According to the method, a predicted target position and a dynamic temporary reference are established by using a physical state data vector and a historical deviation database, and parallel compensation of system errors and mechanical errors is realized.
Owner:WUXI JIANGLAN INTELLIGENT EQUIP CO LTD

Visual Search Query Intent Extraction and Search Refinement

Visual search query intent extraction and search refinement is described. In one or more implementations, a visual search query system receives a search query for items listed on an online marketplace, the search query including an image. Using one or more machine learning models, the visual search query system analyzes the image to determine characteristics of an object in the image. Based on the characteristics of the object in the image, the visual search query system automatically generates one or more search terms and searches the online marketplace to locate items matching the one or more search terms. The visual search query system then displays visual indications of the located items matching the one or more search terms in a user interface of the online marketplace.
Owner:EBAY INC

Open vocabulary target detection method and system based on visual retrieval enhancement prompt

The invention relates to the technical field of computer vision, discloses a visual retrieval enhancement prompt-based open vocabulary target detection method and system, and aims to solve the problem of poor rare category detection performance caused by training data long tail distribution in open vocabulary target detection. The method comprises two stages of off-line construction of a visual concept knowledge base and on-line retrieval enhanced reasoning, wherein a large language model is utilized to generate discriminative visual descriptors for each category, and the discriminative visual descriptors are encoded by a CLIP text encoder and then stored in an FAISS vector database; during reasoning, global visual features of an input image are extracted through a CLIP image encoder, the most relevant descriptor in the knowledge base is retrieved and added to an original category name prompt, and an enhanced prompt is formed and input to an open vocabulary detector to complete detection. Experiments show that the method obviously improves the detection performance in COCO and LVIS benchmark tests, and especially obviously improves the detection effect on rare categories.
Owner:CHONGQING UNIV

Display method and system of electronic rearview mirror based on multi-mode risk field coupling

The invention discloses a display method and system of an electronic rearview mirror based on multi-mode risk field coupling. The display method comprises the steps that driving state data of a current vehicle and one or more target vehicles and observation point data of a current vehicle driver are obtained; constructing a coupling dynamic risk field model based on the driving state data, and generating a comprehensive risk value; generating an attention entropy and a visual search efficiency value based on the observation point data; generating a display control signal by the comprehensive risk value, the attention entropy and the visual search efficiency value through a visual information flow optimizer based on an optimal control theory; and according to the display control signal, simplifying the background region by adopting an anisotropic diffusion equation, enhancing the foreground target by utilizing a nonlinear rendering pipeline, generating an electronic rearview mirror synthetic picture, and displaying the electronic rearview mirror synthetic picture on an in-vehicle display screen. According to the invention, the prominent presentation of key risk information and the effective compression of redundant information are realized, the cognitive load of a driver is reduced, and the driving decision-making efficiency and safety are improved.
Owner:SAIC GM WULING AUTOMOBILE CO LTD

Identifying and localizing editorial changes to images utilizing deep learning

The present disclosure relates to systems, methods, and non-transitory computer readable media that utilize deep learning to identify regions of an image that have been editorially modified. For example, the image comparison system includes a deep image comparator model that compares a pair of images and localizes regions that have been editorially manipulated relative to an original or trusted image. More specifically, the deep image comparator model generates and surfaces visual indications of the location of such editorial changes on the modified image. The deep image comparator model is robust and ignores discrepancies due to benign image transformations that commonly occur during electronic image distribution. The image comparison system optionally includes an image retrieval model utilizes a visual search embedding that is robust to minor manipulations or benign modifications of images. The image retrieval model utilizes a visual search embedding for an image to robustly identify near duplicate images.
Owner:ADOBE INC +1

External human-computer interface position determination method based on pedestrian visual search features

The invention discloses an external human-computer interface position determination method based on pedestrian visual search features, and particularly relates to the technical field of traffic safety, comprising the following steps: dividing a plurality of sight focusing parts according to a vehicle profile, and performing a pedestrian visual search test to collect test data; performing difference analysis on the data, and determining factors which have significant influence on the watching behavior of the testee; using a Markov chain to analyze the visual transfer characteristics of the pedestrian under different influence factor conditions, and quantifying the transfer probability of the pedestrian between different sight focusing parts under different influence factor conditions; through quantification of influence of factors on visual search, powerful support is provided for arrangement of the external human-computer interface position of the vehicle, and the arranged external human-computer interface position can be quickly noticed by pedestrians.
Owner:ANHUI KASIPU INTELLIGENT TECH CO LTD

Visual search refinement

Described is a system and method for enabling visual search for information. With each selection of a search term, additional search terms are dynamically selected and presented to the user in conjunction with results matching the currently selected search terms. Likewise, a selected search term may be tokenized and a graphical token presented to the user to represent the selected search term.
Owner:PINTEREST INC

User interface sketch retrieval method based on modal fusion

The invention discloses a sketch retrieval method based on a modal fusion network, belongs to the field of computer vision, and develops a sketch retrieval visual search framework based on modal fusion by taking a user interface sketch as an efficient information representation in a software development process in mobile application program development. Making a user interface sketch data set by using a data synthesis technology, and preprocessing the data set; a network of a feature extraction encoder and a modal fusion operator is utilized to perform feature extraction on a user interface sketch, and a feature encoder module is responsible for extracting meaningful visual feature representation from an original image and a sketch of a UI example so as to accurately capture information such as design elements and layout. After the modal fusion network extracts visual feature representations from natural images and sketches, the visual feature representations are mapped to a unified modal fusion embedding space. By learning a representation space fusing modal data points and learning and calculating cross-modal attention mapping between the modal data points, the research can reserve semantically related and modal specific features so as to unify information of two modals. Through the proposed multi-modal embedding framework, not only can the structure of the original image and the sketch of the UI example and the joint feature of the associated content be learned, but also the multi-modal embedding framework can be applied to guide the retrieval process of the UI example sketch.
Owner:NORTHWEST UNIV

Digital supplement association and retrieval for visual search

Systems and methods for identification and retrieval of content for visual search are provided. An example method includes receiving data specifying a digital supplement. The data may identify a digital supplement and a supplement anchor for associating the digital supplement with visual content. The method may also include generating a data structure instance that specifies the digital supplement and the supplement anchor and, after generating the data structure instance, enabling triggering of the digital supplement by an image based at least on storing the data structure instance in a database that includes a plurality of other data structure instances. The other data structure instances may each specify a digital supplement and one or more supplement anchors.
Owner:GOOGLE LLC

Visual search refinement

Described is a system and method for enabling visual search for information. With each selection of a search term, additional search terms are dynamically selected and presented to the user in conjunction with results matching the currently selected search terms. Likewise, a selected search term may be tokenized and a graphical token presented to the user to represent the selected search term.
Owner:PINTEREST INC

Driver target attention prediction method, system, device, medium and product

The invention discloses a driver target attention method, system and device, a medium and a product, and relates to the field of target detection, and the method comprises the steps: constructing a driver target attention prediction model; the driver target attention prediction model comprises a target recognition model, a driving event classification model and an attention prediction model; the attention prediction model comprises an environment model fusing multi-layer visual memory and a target-level visual attention prediction model; obtaining visual behavior data and driving data of a driver in real time; and inputting the driver visual behavior data and the driving data into the driver target attention prediction model to obtain a visual attention prediction target. According to the method, the adaptive capacity of the driver visual target attention prediction model can be improved, the accuracy and robustness of model prediction are improved, the prediction result better fits the visual search mechanism of the driver, and the practicability of the driver visual target attention prediction model is improved.
Owner:BEIJING INST OF TECH

Visual search based real time e-commerce system and method with computer vision

InactiveUS20250272732A1CommerceEngineeringVisual search engine
A visual search based real time e-commerce system and method with computer vision is disclosed. A visual search engine (150) of the said system draws item detection and classification (110) sub module to retrieve information from a remote database (118). The said item detection sub module (110) on the server having instructions, that when executed by a processor, cause operations to be performed, wherein the operations includes receiving at least one image (102) and / or a video uploaded (104) by a user using a mobile application to the said visual search engine (150). The visual search (150) engine requests vendors registered in the system to quote availability and a cost associated, payment method, and shipping options (124) to a preferred location of the said user. The real-time e-commerce system trained using deep learning models accurately identifies and classifies the objects, maps the object label set with its respective product category which in turn, creates product attributes to search vendors in the specific category in real-time.
Owner:PABBATHIREDDY RAMAKRISHNA REDDY +1

Old people subthreshold depression cognitive deviation correction training system

The invention relates to the technical field of cognitive deviation correction training, in particular to an old people subthreshold depression cognitive deviation correction training system which comprises an image processing module, an image storage module, a point probe output module, a visual search output module, a response data statistics module, an analysis calibration module, a detection implantation module and a unified processing module. Determining an attention change value for each user based on the response delay time length, dividing an attention category of each user based on the attention change value, and determining whether to correct the doping degree of the probe stimulation identifier output by the identifier generation unit based on the attention category; when it is determined that the attention category of the single user is the weak attention category, determining whether the emotional picture output for the single user is qualified or not based on the attention change representation value; personalized adjustment is carried out according to the specific condition of the user, the emotion picture is adjusted according to the response condition of the user, and the processing efficiency of the correction training data is improved.
Owner:JILIN UNIVERSITY

Online re-judgment method and system based on large model and visual retrieval screening method

The invention provides an online re-judgment method and system based on a large model and a visual retrieval screening method, and belongs to the related technical field of artificial intelligence and industrial quality inspection.The online re-judgment method comprises the steps that an efficient visual model is constructed through pre-training and generation technologies, and the efficient visual model is deployed to a production line for trial operation; carrying out high-value data collection and labeling by utilizing an active learning and automatic labeling technology, carrying out efficient visual model iteration, and starting mass production operation; the difficult cases detected by the model are re-judged through an industrial quality inspection large model and a retrieval enhancement technology, and data which are not confidence in re-judgment are handled manually. According to the method and the system, more than 80% of investment of manual re-judgment can be reduced, and misjudgment caused by subjectivity of manual re-judgment can be remarkably reduced, so that remarkable cost reduction and benefit increase are brought to enterprises; and popularization and application in the fields of quality inspection, screening and the like are facilitated.
Owner:TZTEK TECHNOLOGY CO LTD

An online fiber visualized searching system

The application discloses an online fiber visual searching system, which comprises luminous electronic tags, a first tag reader and a second tag reader. The luminous electronic tags are provided with inductive coils and luminous LED lamps. The luminous electronic tags are arranged at two ends of optical fibers, and the luminous electronic tags at the two ends of the same optical fiber have corresponding marks. The first tag reader is used for reading a single luminous electronic tag at a time. The second tag reader is used for reading a single or multiple luminous electronic tags at a time. The first tag reader and the second tag reader are in communication connection, and the first tag reader transmits tag information to the second tag reader after reading the tag information. The application has the advantages of realizing online optical fiber identification and being beneficial to fiber information management.
Owner:NANJING HUAMAI TECH

Attention state detection device

PendingCN122056594APsychotechnic devicesSensorsVisual technologyAttention Concentration
The invention discloses an attention state detection device, and relates to the technical field of computer vision. In the device, eye movement characteristics are acquired through a first acquisition module, then behavior characteristics are determined based on the eye movement characteristics in a first determination module, and finally, the attention state is determined based on a pre-established attention stage, a pre-established mapping relation between the behavior characteristics and a threshold value and the current behavior characteristics in a second determination module. In the attention detection device, the attention is divided into a plurality of stages in advance, such as a visual search characterization stage, a sight orientation characterization stage, an attention maintenance characterization stage and an attention concentration characterization stage, and a mapping relationship among the attention stages, behavior characteristics and threshold values is pre-established, so that the attention detection accuracy is improved. According to the invention, the existing attention stage and the missing attention stage can be determined based on the mapping relation and the current behavior characteristics, so that the user can know the specific stage where the attention is missing, and the accurate judgment of the attention state is realized.
Owner:ANYANG XIANGYU MEDICAL EQUIP

Visual search method with feedback loop based on interactive sketch

The disclosure relates to the fields of electronic processing of images and e-commerce, specifically—to the visual search with feedback loop based on interactive sketch for transparency and user control. Presented method adds transparency and user control to the traditionally “black-box” visual search process by introducing a feedback loop, based on interactive sketch. The method comprises of following main steps: submission of an information about the object being searched for by the user (in a form of image, written text, spoken text or selection of product category from menu in GUI), identification of product attributes and verification of identified attributes against the ontology rules, selection of building blocks from database for filtered attributes, connection of selected blocks to form a sketch of object being searched for, application of projection rules to produce a form vector, comparison of vector from previous step with database, presenting user with interactive sketch of product alongside the list of matching products from database, enabling user to select a part of the object in sketch, presenting user with possible modifications for the selected part of the object, enabling user to select needed modification and re-iterate the search process as many times as needed, until the wanted results are achieved.
Owner:GASEFIS UAB

Long document understanding method and device, electronic equipment and storage medium

The invention discloses a long document understanding method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the layout analysis of a target document through a layout analysis model, and obtaining a layout analysis result; wherein the layout analysis result at least comprises coordinates and types of the element areas; cutting an element sub-graph based on the coordinates of each element region; generating a text query instruction by using the current query input and the layout analysis result every time the query input is received; inputting the text query instruction and all the element sub-graphs into a visual searcher for visual search to obtain a plurality of key element sub-graphs associated with the current query input and the current search task type; performing document understanding on each key element sub-graph based on the text cue word through a multi-model large model to obtain a current document understanding result; wherein the text cue word is generated at least based on the current query input and the current search task type; and utilizing the current document to understand and output a feedback result of the current query input.
Owner:ASIAINFO TECH CHINA INC

Automatic conference spokesman positioning method, system and device and medium

The invention relates to the technical field of intelligent positioning, in particular to a conference spokesman automatic positioning method, system and device and a medium. The method comprises the following steps: driving a microphone array to collect training sound data, and controlling a pan-tilt camera to spirally scan a conference room speaking area to collect training image data; when training human voice and lip micro-motion data are synchronously detected, extracting three-dimensional coordinates of a sound source and recording current shooting parameters (at least one of a horizontal rotation angle, a pitch angle and a focal length) of a camera to form a group of training samples; not less than 30 groups of samples are collected and input into a neural network for training, a coordinate transformation matrix of the camera relative to the microphone array is obtained, the microphone array obtains real-time sound source coordinates, visual search center coordinates are generated through the transformation matrix, and accordingly the camera is driven to complete spokesman close-up shooting. The method can be completed without manual intervention.
Owner:SHENZHEN MINRRAY IND CORP LTD

External human-machine interface position determination method based on pedestrian visual search features

The present invention discloses a method for determining the position of an external human-machine interface based on pedestrian visual search characteristics, which specifically relates to the field of traffic safety technology, including: dividing a plurality of sight-focusing areas according to the vehicle shape, conducting a pedestrian visual search test and collecting test data; performing a difference analysis on the data to determine the factors that have a significant impact on the tester's gaze behavior; using a Markov chain to analyze the visual transfer characteristics of pedestrians under different influencing factors, and quantifying the probability of pedestrians transferring between different sight-focusing areas under different influencing factors; by quantifying the impact of factors on visual search, providing strong support for the arrangement of the vehicle's external human-machine interface position, and the arranged external human-machine interface position can be quickly noticed by pedestrians.
Owner:ANHUI KASIPU INTELLIGENT TECH CO LTD

Marine boundary monitoring system and method based on image analysis and electronic chart

The invention relates to a system and a method for marine boundary monitoring. In the prior art, video images are generally obtained through a fixed camera installed on the land, monitoring personnel conduct monitoring for 24 hours all day long, the obtaining and utilization efficiency of image information is low and limited, a large number of ocean boundary monitoring personnel are needed, and meanwhile the risk of a monitoring blind area is possibly caused due to fatigue accumulation of each monitoring personnel. In addition, when an unidentified ship appears, a person needs to be dispatched to go out on site for face-to-face confirmation, and the problem of low efficiency in the aspects of manpower and cost is caused. In order to solve the problem, the invention provides the ocean boundary monitoring system and method based on the image analysis and the electronic chart, by automatically analyzing camera and radar images, the intervention of ocean monitoring personnel is minimized, when an unidentified ship appears, an unmanned ship and an unmanned aerial vehicle are used for obtaining a video image of the ship, and the video image of the ship is displayed. And identification is carried out through a visual search method, so that on-site operation and face-to-face confirmation are reduced. On the basis, efficient operation can be achieved with a simpler structure and lower cost.
Owner:KOREA INSTITUTE OF OCEAN SCIENCE & TECHNOLOGY

A method and apparatus for supporting multi-time axis visual search and playback

The present application relates to the technical field of video retrieval and playback, and provides a method and device supporting multi-time axis visual retrieval and playback. Each terminal of the present application respectively collects, encodes and stores the fragmented data files from multi-channel audio and video signals, and sends the fragmented data files to a client through key frame mapping; the client performs multi-source time axis aggregation indexing on the fragmented data files from each terminal, matches the playback time of each channel of audio and video signals specified by a user, obtains the closest key frame information, and sends the corresponding terminal; the terminal outputs the file stream of the fragmented data file to the client from the specified key frame according to the key frame information, and the client receives each channel of file stream and decodes and plays back. The present application solves the problems that the prior art cannot intuitively obtain the audio and video recording situation of a channel or multiple channels at a certain time point, and cannot select and directly start single-channel or multi-channel fast positioning and playing from a certain time point.
Owner:709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD

Application prediction based on a visual search determination

Visual search in an operating system of a computing device can process and provide additional information on the content being provided for display. The computing device can include an operating system that includes a visual search interface that obtains and processes display data associated with content currently being provided for display. The visual search interface can generate display data based on the current content provided for display, process the display data with one or more on-device machine-learned models, and provide additional information to the user. The visual search interface may transmit data associated with the display data to perform additional data processing tasks. Application suggestions may be determined and provided based on the visual search data.
Owner:GOOGLE LLC