Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Visual matching" patented technology

Low-altitude visual matching navigation methods, devices, systems, and storage media

PendingCN122083949ANavigational calculation instrumentsVisual matchingFlight vehicle
This invention discloses a low-altitude visual matching navigation method, device, system, and storage medium, comprising: in an offline phase, optimizing an aerial image sequence into a high-fidelity 3DGS map model; in an online phase, rapidly fusing coarse poses for rendering using inertial pre-integration and global descriptor retrieval; based on the coarse pose, using the 3DGS differentiable rasterization pipeline to synthesize a high-fidelity new perspective reference image in real time by jointly optimizing the pose increment and reference view fusion weights; obtaining a 2D-2D correspondence between the real-time image and the new perspective reference image through deep learning matching, and converting the depth map synchronously generated by the 3DGS model into a 2D-3D association; and calculating the high-precision visual pose of the aircraft through the PnP algorithm and nonlinear optimization. This invention can improve the problems of low quality of new perspective reference images, insufficient multi-sensor information fusion, and limited positioning accuracy in complex low-altitude scenarios.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Leveraging audio matches to improve visual matching recall between video content items

ActiveUS12677030B1Visual matchingAudio frequency
Audio matching is performed between a first video content item and a second video content item to identify a matching audio segment. First temporal boundaries within the first video content item and second temporal boundaries within the second video content item corresponding to the identified matching audio segment are identified. A visual matching between the first video content item within the first temporal boundaries and the second video content item within the second temporal boundaries is performed using a modified visual similarity threshold that is lower than a baseline visual similarity threshold. Whether a match exists between the first and second video content items is determined based on the visual matching.
Owner:GOOGLE LLC

Image automatic labeling method and system based on landscape element knowledge graph and visual matching

A method and system for automated image annotation based on a landscape element knowledge graph and visual matching is disclosed. The method includes: constructing a landscape element knowledge graph; performing structured association modeling of landscape nodes, landscape element nodes, image nodes, and cultural knowledge nodes; extracting relevant landscape nodes from the graph based on user-inputted images and shooting locations to limit the range of elements to be identified; extracting features from the input image using a zero-shot visual matching model and comparing them with image samples of candidate landscape elements to obtain identification results; extracting attributes and associated cultural knowledge content from the graph based on the identified landscape element nodes; semantically fusing the identification results with cultural knowledge according to preset rules to generate structured or natural language annotation text; and outputting the annotation text to a terminal for display. This invention combines geographic location, knowledge graph, and visual matching to achieve accurate, information-rich, and culturally profound automated annotation of landscape images.
Owner:ZHEJIANG UNIV OF TECH

Webpage data processing method, apparatus, device, and medium

The present disclosure provides a webpage data processing method and device, equipment and medium, relates to the technical field of artificial intelligence, in particular to the technical field of webpage development and deep learning. The method comprises: determining an interactive operation for a target webpage; obtaining a screenshot of the target webpage and attributes of a plurality of webpage elements in the target webpage, the attributes indicating functions possessed by the webpage elements; based on the attributes of the plurality of webpage elements, screening a plurality of candidate webpage elements related to the interactive operation from the plurality of webpage elements; determining the positions and sizes of the plurality of candidate webpage elements, and cutting the screenshot to obtain visual segments of the plurality of candidate webpage elements; determining the intention matching degrees of the attributes of the plurality of candidate webpage elements and the interactive operation; determining the visual matching degrees of the visual segments of the plurality of candidate webpage elements and the interactive operation; and based on the intention matching degrees and the visual matching degrees, determining a target webpage element from the plurality of candidate webpage elements and executing the interactive operation.
Owner:BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD

Robot task planning method based on multi-modal large model

PendingCN122347174AVisual matchingData set
The application relates to the technical field of robot control, and discloses a robot task planning method based on a multimodal large model. The application collects multimodal data of a robot motion scene, including voice information and visual information, pre-processes the information to obtain a task instruction sequence required to be executed by the robot and an entity data set in a scene where the robot is located, inputs the pre-processed data into a multimodal large model based on a large language model, calculates, based on the pre-processed data, a semantic evaluation coefficient, a visual matching coefficient, a feasibility coefficient and a priority coefficient corresponding to each task instruction in the task instruction sequence, and obtains a task comprehensive execution coefficient corresponding to each task instruction by weighted summation, reorders the task instructions of the task instruction sequence based on the task comprehensive execution coefficient, and obtains a final task planning sequence, thereby improving the rationality and efficiency of task execution of the robot.
Owner:KEYI COLLEGE OF ZHEJIANG SCI TECH UNIV

Data transmission system and method for operation guidance in remote interventional surgery

This application discloses a data transmission system and method for operation guidance in remote interventional surgery, relating to the field of intelligent sensor technology. The method includes: acquiring a control dataset; calculating the video end-to-end latency, visual matching ambiguity, and operator hand tremor components based on the control dataset; processing the optimal pairing set corresponding to the control dataset and the visual matching ambiguity using a preset filter to obtain motion state estimates; performing temporal cross-correlation analysis on the video end-to-end latency and operator hand tremor components to calculate the delay-tremor coupling strength; constructing a future position prediction distribution model based on the motion state estimates and the delay-tremor coupling strength; superimposing the future position prediction distribution model with a preset anatomical structure model, and displaying the superimposed result on the current surgical video screen to guide the surgeon in remotely controlling the surgical operation from the control terminal. This application improves the safety of remote precision operations.
Owner:THE SECOND HOSPITAL AFFILIATED TO WENZHOU MEDICAL COLLEGE

Fast relocalization method based on sparse semantic anchor points and inertial sensor tight coupling

This invention discloses a fast relocalization method based on tight coupling between sparse semantic anchors and inertial sensing, belonging to the fields of augmented reality and computer vision. It constructs a lightweight sparse semantic map by extracting sparse feature points with stable geometric and semantic attributes from the environment as sparse semantic anchors. When device tracking is unstable or lost, inertial navigation and visual matching threads based on sparse anchors are activated in parallel. The core PnP algorithm is used to quickly recover the visual pose, and the method is tightly coupled with inertial data for optimization, ultimately outputting a high-precision, smooth six-DOF pose. This solves the problem of excessively long relocalization time and experience interruption in traditional visual SLAM scenarios with fast motion and weak textures, achieving millisecond-level, user-unnoticed tracking recovery, significantly improving the robustness and user experience of augmented reality systems.
Owner:CHONGQING AEROSPACE POLYTECHNIC COLLEGE

A script visual asset structured generation method based on a large model

PendingCN122262323ASemantic analysisBiological modelsSemantic vectorVisual matching
The present application relates to the technical field of natural language processing, and more particularly to a script visual asset structured generation method based on a large model, which comprises the following steps: obtaining a script original text and dividing it into scene text units; extracting deep semantic vectors by using a large-scale pre-training language model; mapping literary descriptions to a primary visual feature set based on a visual semantic feature library; encapsulating each component by using a structured mapping engine, and constructing a structured visual asset description model including global scene parameters, entity object parameters and dynamic interaction parameters; and detecting and outputting consistent structured data through a cross-scene logic checking mechanism. The present application realizes the automatic conversion of script text into parameterized data, and improves the semantic analysis depth and visual matching accuracy.
Owner:SHANGHAI CHENGRONG NETWORK TECHNOLOGY CO LTD

A multi-source fusion reliable navigation positioning method for urban night complex scenes

PendingCN122306056AVisual matchingData acquisition
This invention provides a multi-source fusion reliable navigation and positioning method for complex urban nighttime scenarios, including: GNSS, INS, and visual data acquisition and preprocessing to achieve quality control of input data; visual feature extraction and initial matching using deep learning technology; identification and removal of dynamic object interference by combining optical flow residuals and IMU motion parameters; estimation of the confidence of visual feature matching point pairs based on LSTM; construction of a GNSS / INS / Vision tightly coupled navigation and positioning model and a filter optimizer by fusing the matching confidence; and output of the final navigation result. This invention improves the robustness of visual matching in complex environments with changing urban nighttime lighting and dynamic object interference, significantly enhancing the continuity and reliability of autonomous perception and navigation and positioning of the vehicle.
Owner:LIAONING TECHNICAL UNIVERSITY

A method, system, device and storage medium for industrial safety scenario matching

The application discloses a kind of industrial safety scene matching method, system, equipment and storage medium.The method first constructs the standardized triad data including hidden danger image, description text and rectification image;Through image feature matching module, feature matching proportion, internal point proportion and structure similarity index are extracted and fused, and quantitative result including comprehensive confidence is generated;Further, large language model analysis module according to quantitative result and hidden danger description, through structured prompt word guide carries out multi-step semantic reasoning, and outputs matching analysis result with grade;Finally, fusion visual and semantic evidence completes safety hidden danger rectification verification and decision.The application solves the problem of semantic gap and high false alarm rate of pure visual method through "visual matching-semantic reasoning" collaborative architecture, realizes intelligent and accurate identification of compliance rectification and real hidden danger.
Owner:YUANZHIFU (HANGZHOU) TECH CO LTD

A parking inspection method of an unmanned aerial vehicle based on image path planning

The application discloses a parking inspection method based on image path planning of a UAV, and comprises the following steps: acquiring an inspection image shot by the UAV on an inspection path; dividing the inspection image into a plurality of inspection unit images corresponding to preset sub-regions; matching each inspection unit image with a reference image of the corresponding sub-region to obtain a similarity; when the similarity is lower than a preset threshold, performing visual matching on the inspection unit image and a pre-stored positioning map to calculate a current position coordinate of the UAV; matching the calculated current position coordinate of the UAV with a preset inspection parking point coordinate to determine a nearest target parking point; and controlling the UAV to fly to the determined nearest target parking point.
Owner:广东科陆智泊信息科技有限公司 +1

An accessory identification method and apparatus

This application discloses a method for parts identification. The method first acquires image information of the part to be identified and extracts corresponding part information. Then, it uses this information to query a preset electronic parts catalog. Based on the query results, it flexibly uses an object storage service image library to match candidate part images. Finally, it determines the target part based on the degree of matching between the candidate part information and the part information to be identified. This method integrates the advantages of accurate catalog querying and image visual matching, overcoming the shortcomings of single identification schemes. It comprehensively mines key part information to reduce misjudgments, uses layered processing to balance recognition efficiency and adaptability to complex scenarios, reduces the mismatch rate of similar parts, and quantifies matching judgments to improve the accuracy of results. This achieves efficient and accurate parts identification, meeting the high reliability and efficiency requirements of the automotive aftermarket, reducing user time costs, and improving business transaction efficiency and data update timeliness.
Owner:SHENZHEN CASSTIME TECH CO LTD

Three-dimensional recognition vision accuracy matching optimization method and system for industrial robot

PendingCN122289308APattern recognitionVisual matching
This invention discloses a method and system for optimizing visual accuracy matching in 3D recognition for industrial robots, relating to the field of visual guidance technology. The method includes the following steps: during the normal movement of the industrial robot, acquiring motor parameters, IMU data, and visual images of the work area; constructing and selecting 3D probability models for various working conditions; predicting the 3D pose distribution for the next sampling period; and outputting a pose prediction sequence composed of multiple candidate pose points. Under the condition of triggering a pose mutation, a preset rule engine is invoked to reconstruct the 3D visual image, plan the motion trajectory of the industrial robot, perform a matching between the motion trajectory and the reconnection and alignment pose, and select the optimal reconnection and alignment pose. This invention improves the visual matching accuracy in weak texture and smooth scenes through texture determination and the construction of full disparity cost analysis, effectively improving the adaptability and efficiency of industrial robot reconnection and alignment scenarios.
Owner:陕西天捷创科智能科技有限公司

A multi-unmanned aerial vehicle cooperative interception method based on a gimbal visual guide

PendingCN122431399AVisual matchingObservation data
The application discloses a multi-unmanned aerial vehicle cooperative interception method based on a gimbal visual guide, relates to the technical field of unmanned aerial vehicle cooperative control and intelligent visual perception, and comprises the following steps: a distributed cooperative network of multiple unmanned aerial vehicles is constructed to realize sharing of identification information, position information and resource states; a target multi-view image is acquired by using a gimbal camera of each unmanned aerial vehicle; a three-dimensional position and a velocity vector of the target are calculated through stereo visual matching and multi-frame difference calculation; multi-source observation data are weighted and fused based on an attention mechanism to obtain a target motion state; an optimal interception path is planned by using a distributed predictive control algorithm in combination with resource states and relative positions of the unmanned aerial vehicles, and an elastic formation is formed; gimbal parameters are dynamically adjusted, and a main tracking unmanned aerial vehicle is adaptively switched to realize continuous and stable tracking. The application improves target positioning accuracy and cooperative interception efficiency, and enhances system robustness and environmental adaptability.
Owner:CHANGZHOU YUNSHAO LOW ALTITUDE INTELLIGENT TECHNOLOGY CO LTD

A method and system for intelligent inspection of security robots for multiple scenarios

ActiveCN121541646BPattern recognitionVisual matching
This invention discloses an intelligent inspection method and system for security robots in multiple scenarios, relating to the field of intelligent security. The method includes: in environments with good visibility, collecting vibration data generated by the robot's interaction with the ground to establish a standard fingerprint strip and historical images; in low-visibility environments, the robot collects blurred visual images and vibration data in real time, generating the current inspection fingerprint and matching it with the standard fingerprint strip; when a deviation is detected, the system calls upon a visual memory library and uses a deep learning visual matching network to determine the closest reference landmark, thereby adjusting the robot's turning angle to bring it back to the center path of the fingerprint strip. This invention can be applied to industrial control software and deployed on various models of security robots to achieve process control during inspections. It solves the problem of unstable positioning of security robots during inspections in low-visibility environments, realizing path self-correction and intelligent inspection control based on vibration fingerprint and deep visual matching.
Owner:ANLIZHI INTELLIGENT ROBOT TECH (BEIJING) CO LTD

An unmanned aerial vehicle image matching method based on rotation isovist visual features

This invention belongs to the field of computer vision and UAV autonomous localization technology. It discloses a UAV image matching method based on rotationally equivariant visual features. By introducing a group-equivariant convolutional structure, a rotationally equivariant feature extraction network is constructed for target matching, ensuring that the feature representation maintains mathematical structural consistency as the input image rotates. The method for generating feature point detection heatmaps is improved, and an offset correction network and a global offset loss function are designed to achieve automatic coordinate alignment of matching points, eliminating projection deviations caused by rotation and improving pose calculation accuracy. The global offset loss function constrains the spatial consistency of feature points under different viewpoints, significantly improving matching accuracy and stability. This invention achieves rotational equivariance and self-correction of matching projection errors from both network structure and training mechanism levels, enabling UAV visual matching capabilities that adapt to continuous rotation. This invention operates quickly and maintains stable performance in complex environments, closely reflecting real-world UAV flight scenarios.
Owner:BEIHANG UNIV

End-to-end visual matching method and device based on dynamic computation graph, equipment and medium

PendingCN122391678AVisual matchingConfidence map
The application relates to an end-to-end visual matching method, device and equipment based on a dynamic calculation graph and a medium. The method extracts multi-scale features by inputting two visual images into a shared parameter neural network trunk. A normalized similarity matrix of corresponding pixels / feature blocks is calculated based on the features, a confidence map and an attention map are generated from the matrix, and the confidence map and the attention map are fed back to the trunk to enhance the features, so that enhanced multi-scale features are obtained. The matching probability is calculated by using a Softmax algorithm based on the similarity matrix, and the matching uncertainty is calculated by using an entropy index. If the uncertainty is greater than a threshold, a feature recalculation branch is triggered, refined features are obtained by performing deep calculation on the enhanced features through a feature recalculation network, and the enhanced features and the refined features are fused through a dynamic calculation mask to obtain the final features of the two images. The feature corresponding relationship is calculated based on the final features, and a visual feature matching result is output. The method can effectively improve the visual matching precision and optimize the calculation efficiency.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI