Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

75 results about "Visual localization" patented technology

Visual perception method and system based on multi-modal thinking tree

The invention relates to the technical field of artificial intelligence and computer vision, and provides a visual perception method and system based on a multi-modal thinking tree in order to solve the problem that a traditional expansion strategy of purely increasing the parameter scale cannot effectively break through the semantic refinement bottleneck. The visual perception method based on the multi-modal thinking tree comprises the steps of obtaining a to-be-processed original image and a target anaphora text, and constructing the multi-modal thinking tree; defining a reasoning action set for driving node extension; executing a multi-mode Monte Carlo tree search process; iteratively generating a reasoning path until a preset search depth is reached or a termination condition is triggered; all effective leaf nodes generated in the searching process are obtained at the same time; and carrying out aggregation optimization on all effective leaf nodes by adopting a regional feature weighted voting mechanism, and screening out a candidate scheme with the highest comprehensive weight as a final visual perception positioning result. According to the method, the perception performance can be effectively improved on the basis of not changing the original parameter scale of the model, and high-precision visual positioning is realized.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Inplanatory visual question-answering method and system based on question perception and confidence constraint

The invention discloses an explanatory visual question-answering method and system based on question perception and confidence constraint. The system comprises a selection-enhancement module and a fusion enhancement confidence constraint module, the selection-enhancement module is used for selecting candidate image areas based on question semantics and enhancing salient areas related to questions through a learnable enhancement mechanism to realize accurate visual localization; and the fusion enhancement confidence constraint module is used for fusing the enhanced regional features and the multi-modal features, and ensuring that the confidence of a prediction answer is improved when explanation information is introduced through a double-branch prediction and confidence constraint mechanism, so that the reliability of a prediction result is enhanced. According to the method, the defects of problem insensitive positioning and positioning-reasoning disjunction in the prior art can be effectively overcome, the superiority of the method is verified on a public data set, the answer prediction accuracy and the explanation generation quality are remarkably improved, and the method has wide application prospects.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-modal large model illusion detection method based on reverse visual localization

The invention belongs to the technical field of artificial intelligence, and particularly relates to a multi-modal large model illusion detection method based on reverse visual positioning. The method comprises the following steps: constructing a visual instruction fine tuning data set rich in context; training a visual positioning large model with pixel-level positioning and rejection capability based on the data set; performing sentence-by-sentence verification on a response generated by a to-be-detected multi-modal large model by using the trained model, and judging whether illusion exists or not by judging whether text description can be reversely positioned back to image pixels or not; according to the method, illusion rich in details can be effectively detected, pixel-level masks and natural language interpretation are provided, and the accuracy and transparency of evaluation are remarkably improved.
Owner:FUDAN UNIV YIWU RES INST +1

Geometrically assisted visual positioning method and system

A geometric structure aided visual localization method and a system implementing the method are provided. The method includes retrieving a 3D map of initialization location data, the 3D map being constrained by visual structures and geometric structures and modeled by a Gaussian mixture model of a set of Gaussian distributions with mapped landmarks; acquiring a series of real-time image frames by a camera of a mobile system; and for each real-time image frame: extracting local features from the real-time image frame; predicting a camera pose corresponding to the real-time image frame; creating a key frame by tracking the predicted camera pose in the 3D map; identifying temporally visible landmarks with respect to the created key frame; acquiring 2D-3D correspondences between the local features and the temporally visible landmarks; and localizing the mobile system by estimating a state of the real-time image frame based on the 2D-3D correspondences.
Owner:HONG KONG UNIV OF SCI & TECH R & D CORP LTD

A multi-modal visual understanding method based on consistent learning and mixed feature extraction

This invention relates to a multimodal visual understanding method based on consistency learning and hybrid feature extraction. It includes constructing an end-to-end fine-grained consistency learning framework, introducing a hybrid region extractor, fusing local details and global semantics to generate high-quality hybrid visual cue embeddings, combining self-reconstruction loss and latent spatial consistency loss to force the model to establish explicit alignment between the input visual cue and the output segmentation label, utilizing the geometric boundary constraints of the localization task for description generation, and simultaneously optimizing localization accuracy using the semantic depth of the description task. Furthermore, it constructs a detailed localization index expression and segmentation task to enhance the model's reasoning ability for complex long text instructions. The aim is to address the problems of feature fragmentation and insufficient accuracy in existing large models for fine-grained visual localization and description tasks. Compared with existing technologies, this invention has advantages such as high accuracy and strong generalization ability in pixel-level localization and fine-grained description.
Owner:TONGJI UNIV

High-efficiency visual positioning method under conservation of multi-modal large model dialogue capability

The invention relates to an efficient visual positioning method under conservation of a multi-modal large model dialogue capability, which is characterized in that an adopted positioning network D-LMM comprises a multi-modal large model LMM image encoder CLIP, a multi-modal large model LMM text encoder, a frozen LLM and an instance-level text feature extractor. A dynamic instance feature modulation module DIFM and a fusion segmentation mask head are adopted, and for given images and texts, the process of positioning by using the D-LMM comprises the following steps: inputting the images into an image encoder CLIP to obtain multilayer features; splicing the input visual features and the input text features together; inputting the output text features of the LLM into an instance-perceived text feature extractor ITE to obtain instance-perceived text features; and inputting the average text feature of each instance and the multi-level visual features obtained from the image encoder into a dynamic instance feature modulation module DIFM, and converting the multi-level visual features into a multi-level feature map perceived by the instances.
Owner:TIANJIN UNIV

Multimodal visual positioning method, apparatus, and electronic device

This application provides a multimodal visual localization method, apparatus, and electronic device, which can be applied to the field of computer vision technology. The method includes: responding to image data acquired by an image acquisition device mounted on a vehicle at a target time, extracting features from the image data to obtain image features; predicting the depth features at the target time based on the expected pose transformation of the image acquisition device and historical depth features to obtain predicted depth features; if the image data also includes a depth image, performing noise reduction processing on the predicted depth features according to the mapping relationship between the depth features and the predicted depth features to obtain optimized depth features; fusing the optimized depth features and visual features to obtain multimodal visual features; and localizing the image acquisition device based on the multimodal visual features to obtain the localization result at the target time.
Owner:TIANJIN UNIV

A map-free visual positioning method, device, equipment and medium

PendingCN122636718AVoxelVisual localization
This application discloses a map-free visual localization method, apparatus, device, and medium. The method includes acquiring scene images collected in real time by a robot and extracting an initial feature map from the scene images; enhancing the initial feature map into a voxel feature map using a sparse kernel self-attention mechanism; aggregating the voxel feature map from several dimensions and multiplying the aggregated features element-wise to obtain image feature identifiers; searching for at least one anchor point feature identifier based on the image feature identifiers; and determining a six-degree-of-freedom pose based on the set of anchor point feature identifiers. This application utilizes a sparse kernel self-attention mechanism to enhance the initial feature map using voxels and aggregates the voxel feature map from several dimensions, achieving cross-domain interactive sparse quantization. This overcomes feature interference caused by motion blur or illumination changes, improves the robustness of scene image feature extraction, thereby reducing the probability of matching failure or inaccurate matching results and improving the accuracy of map-free visual localization.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Visual localization and attitude estimation method based on prior search in rocket recovery section

The invention discloses a visual localization and attitude estimation method and device based on prior search in a rocket recovery section and a medium, and belongs to the technical field of spacecraft guidance, navigation and control, and the method comprises the steps: carrying out the systematic sampling of a rocket recovery tail end operation space containing a three-dimensional position and a three-dimensional camera attitude in advance; constructing a matching data set in which the rocket recovery state is matched with the image; learning mapping from an image to a low-dimensional feature vector by using a deep learning model, and enabling the structure of a feature space to be consistent with the structure of a physical state space through a designed loss function; and in the rocket recovery stage, pre-stored samples closest to real-time image features are retrieved, and state labels of the pre-stored samples are fused, so that pose estimation is realized. The method can adapt to complex three-dimensional attitude changes, is high in robustness, and can meet the real-time requirement.
Owner:ORIENTAL SPACE TECH (SHANDONG) CO LTD

Visual positioning method, storage medium and electronic device

Provided are a visual positioning method, a non-transitory computer-readable storage medium and an electronic device. Surface normal vectors of a current image frame is obtained. A first transformation parameter between the current image frame and a reference image frame is determined, by projecting the surface normal vectors to a Manhattan coordinate system. A matching operation between feature points of the current image frame and feature points of the reference image frame is performed, and a second transformation parameter between the current image frame and the reference image frame is determined based on a matching result. A target transformation parameter is obtained, based on the first transformation parameter and the second transformation parameter. A visual localization result corresponding to the current image frame is output, based on the target transformation parameter.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Large-scale visual positioning optimization method based on OCP theory

The invention discloses a large-scale visual positioning optimization method based on an OCP theory, and relates to the technical field of deep learning. The method comprises the following steps: defining a target function of a visual positioning model; initializing model parameters, momentum vectors and related hyper-parameters; performing iterative optimization training, calculating a small-batch stochastic gradient in each iteration, performing exponential moving average and deviation correction on the gradient by using a diagonal element of an approximate Hessian matrix of element-by-element square of the gradient, performing weight attenuation, and finally calculating a parameter update quantity and updating a model by using an optimization method based on an OCP theory; and outputting the model with the optimal performance on the verification set after iteration is finished. According to the method, the OCP theory and approximate second-order information are combined, a new large-scale visual positioning method is provided, the convergence speed, stability and final test precision of visual positioning model training are effectively improved on the premise that linear complexity is kept, and the method is particularly suitable for large-scale non-convex optimization scenes.
Owner:SHANDONG UNIV OF SCI & TECH

AGV and turnover equipment docking method and system based on visual positioning

The invention belongs to the technical field of image analysis, and particularly relates to an AGV and turnover equipment docking method and system based on visual localization, and the method comprises the steps: calculating a specular reflection inhibition factor based on the gray value of a pixel point in an image and a neighborhood gray variance, a comprehensive matching cost function is obtained according to a geometric distance between a model projection point and a pixel point of a 3D model of the turnover device, an included angle cosine between an expected 2D normal direction and a gradient unit normal vector and a specular reflection suppression factor, a pixel point with the minimum comprehensive matching cost is searched in a local search neighborhood of the model projection point to serve as a matching point, and the matching point is matched with the model projection point. And updating the pose, and when the number of the effective matching points is smaller than a preset threshold value, switching to a dead reckoning mode based on the last effective pose. According to the method, the problem of unstable positioning caused by environmental interference in a complex industrial environment is solved, and the docking precision and robustness are improved.
Owner:XIAN CUMMINS ENGINE COMPANY

Visual map data processing method and device, computer equipment and storage medium

The invention relates to the technical field of visual SLAM positioning and mapping, and discloses a visual map data processing method and device, computer equipment and a storage medium. Firstly, panoramic semantic segmentation is carried out on an original image, a panoramic segmented image is generated, the panoramic segmented image comprises semantic information of each pixel, and the semantic information comprises prior dynamic semantic information. And performing optical flow estimation processing on the original image to obtain an optical flow image. Then, based on the panoramic segmented image, the optical flow image and the original image, feature points of a dynamic area are removed in real time, and a static feature image with higher precision is obtained; and finally, under the condition of performing motion tracking based on the static feature image and determining that the static feature image is a key frame, generating a dense point cloud map endowed with semantic information based on the original image, the static feature image and the panoramic segmentation image, and removing a point cloud with prior dynamic semantic information from the dense point cloud map. And a more accurate static environment map is constructed to update the point cloud map.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

Scene library based vehicle visual-only localization method

This application provides a vehicle pure vision localization method based on a scene library, relating to the field of vision localization. The method includes: performing early vision localization of the vehicle using a high-precision map to obtain an optimized image pose; performing quality detection on the image pose based on a deep learning model, and storing qualified image poses into a scene library; recalling matching historical images from the scene library based on the image content and rough position information of the current onboard camera image; and using the recalled historical images to assist or replace the high-precision map for pose estimation to obtain the vision localization result. The technical solution of this application forms a data closed loop in which image data accumulation and localization accuracy mutually promote each other, enabling the sustainable operation of the vision localization process.
Owner:JISHU TECHNOLOGY (WUHAN) CO LTD

Fusion localization methods, devices and electronic equipment for autonomous vehicles

This application discloses a fusion localization method, apparatus, and electronic device for autonomous vehicles. The method includes: acquiring localization data from multiple sensors of the autonomous vehicle, including satellite localization data, laser localization data, and visual localization data; determining the confidence level type corresponding to the satellite localization data and the laser localization data using pre-set confidence threshold conditions; determining a fusion localization strategy for the autonomous vehicle based on the confidence level types of the satellite localization data, the laser localization data, and the visual localization data; and performing fusion localization according to the fusion localization strategy to obtain the fusion localization result of the autonomous vehicle. This application performs mutual verification of confidence levels based on localization data from multiple sensors and adopts different fusion localization strategies based on the verification results, ensuring that the fusion localization algorithm has reliable observation input and improving the stability and accuracy of fusion localization.
Owner:ZHIDAO NETWORK TECH (BEIJING) CO LTD

A visual positioning method based on pose space texture continuity mapping

PendingCN122289381APattern recognitionRobotics
This invention discloses a visual localization method based on pose space texture continuity mapping, belonging to the fields of robotics and computer vision. In the mapping stage, image intrinsic parameters, pose, and sparse point cloud are first obtained through sparse reconstruction. Image features are extracted using a pre-trained model, and pose encoding and decoding models are trained to ensure consistency between pose encoding and image features in spatial metrics. Simultaneously, the decoding model outputs a multi-peak distribution to model pose ambiguity. Subsequently, a mapping model from image features to pose encoding is trained and optimized end-to-end to construct a scene map. In the localization stage, features of the image to be localized are extracted, pose encoding is obtained through the mapping model, and the pose distribution is output by the decoding model. A unique location or multiple candidate poses are determined based on the number of peaks in the distribution. This invention effectively improves the robustness and accuracy of robot localization in repetitive texture scenes by maintaining the continuity alignment between the pose space and texture space.
Owner:BEIJING UNIV OF TECH

A method, system, device, and storage medium for visual localization of pathological images.

This application provides a method, system, device, and storage medium for visual localization of pathological images, belonging to the field of image recognition technology. The method includes: extracting visual features based on a target pathological image; determining semantic feature vectors and knowledge feature vectors based on a first text description; the target pathological image is the pathological image for which target region localization is to be performed; the knowledge feature vectors are used to represent knowledge information associated with the content of the target pathological image; fusing the semantic feature vectors and knowledge feature vectors to obtain fused text features; performing cross-modal fusion of the fused text features and visual features to obtain fused multimodal features; obtaining a fused representation based on the fused multimodal features; and, based on the fused representation, locating the target region in the target pathological image using a multilayer perceptron to obtain the position information of the bounding box of the target region. This application can improve the ability to accurately and flexibly locate regions at the pathological image level.
Owner:XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV

A sparse-to-dense visual localization method and system based on feature gaussian splats

PendingCN122115572AReduce storage requirementsPreserve geometric richnessImage analysis3D modellingPattern recognitionHeat map
The application provides a sparse-to-dense visual positioning method and system based on feature Gaussian splash, and the method comprises the following steps: initializing a color-decoupled feature Gaussian field based on a training image set, optimizing the color-decoupled feature Gaussian field based on a query feature map set in combination with feature rendering and feature alignment loss cyclic optimization, and outputting a compact feature Gaussian scene model; screening a Gaussian landmark set in the compact feature Gaussian scene model by using a matching-oriented sampling strategy; training a scene-specific detector; extracting sparse local features of a query landmark heat map corresponding to a query image and performing sparse feature matching with the Gaussian landmark set to obtain an initial pose of a query perspective camera; based on 3D Gaussian splash, rendering a dense feature map and a depth map of the query perspective in the compact feature Gaussian scene model by using the initial pose of the query perspective camera, performing cluster-based proxy matching-based sparse-to-dense accelerated pose optimization, and obtaining accurate positioning of the query perspective camera.
Owner:WUHAN UNIV

Dynamic screen area monitoring and self-adaptive character extraction system based on visual positioning

The invention discloses a dynamic screen area monitoring and self-adaptive character extraction system based on visual localization, which relates to the technical field of dynamic screen character perception, and comprises the following steps: acquiring an original image frame of an electronic screen, carrying out resolution adaptation, automatically identifying a target area and outputting coordinates; executing character recognition; detecting a target area content change of the target area image; and driving the extraction system to re-execute the positioning-extraction process. According to the method, dynamic screen area monitoring and adaptive character extraction without manual intervention are realized, the manual operation cost and the system maintenance burden are reduced, and high-reliability and high-adaptive intelligent screen sensing capability is provided for application scenes such as software automatic testing, user behavior analysis, content security auditing and barrier-free assistance.
Owner:SHANGHAI SHANHAO INTELLIGENT TECH DEV CO LTD

Robot vision positioning method and device based on target detection

This invention provides a robot visual localization method based on target detection, comprising: an image acquisition module that records in real time IR and depth images of the forklift robot's picking direction; an edge computing module that detects the pallet's border information based on the IR image; a false detection filtering module that filters the pallet's border information based on prior knowledge and establishes internal relationships within the borders; and a distance determination module that performs visual localization of the pallet based on the internal relationships within the borders and the depth image. This invention also provides an apparatus for this method, using a more easily deployable single-stage image target detection model to detect the pallet's support legs, enabling the robot to locate the pallet. Then, by combining the depth image with the detection results, the distance between the pallet and the forklift robot is evaluated, and a route is planned to pick up the pallet along with the goods on it.
Owner:BEIJING GRAY TIANZE TECHNOLOGY CO LTD

Tray planting machine, interaction unit, picking method and inserting method

This invention discloses a tray loading machine, an interaction unit, a picking method, and an insertion method. The tray loading machine consists of a feeding mechanism, a first transport mechanism, a tray loading workstation, a second transport mechanism, a vision positioning system, a first barcode reader, and a second barcode reader. The interaction unit includes the tray loading machine and a tray turnover box. The tray turnover box vertically stores SMT material trays with an internal array of storage positions, ensuring the top of the trays protrudes from the upper surface of the storage positions. The picking method uses the second transport mechanism to directly grasp the protruding part of the material for retrieval without extending into the storage positions. The insertion method uses the second transport mechanism to align the material with the storage position and allows it to slide in using its own weight for retrieval. This invention, by converging tray-level operations within the tray loading workstation, combined with the vertical storage and top protrusion features of the tray turnover box in the interaction unit, and the optimization of the gripping position and insertion method in the picking and insertion methods, enables the warehousing system to reduce storage position gaps and ineffective handling strokes, thereby improving warehousing efficiency.

Nuclear-targeting multi-modal imaging diagnosis and treatment probe with bnct efficacy and preparation and application thereof

PendingCN122251581ACapable of imagingSolve the problem of blind treatmentEnergy modified materialsEchographic/ultrasound-imaging preparationsNeutron irradiationFluorescence
This invention relates to a nuclear-targeted multimodal imaging diagnostic probe with BNCT therapeutic efficacy, its preparation, and its application, belonging to the fields of biomedicine and nuclear medicine technology. This boron drug contains... 10 Using acetylacetone difluoroborate as the parent nucleus, this invention endows the tumor with nuclear targeting ability by introducing N- or S-containing heterocyclic structures, while utilizing its conjugated structure to provide fluorescence / photoacoustic dual-modal imaging functionality. Based on this, the invention achieves three core functions: first, through fluorescence / photoacoustic dual-modal imaging, it enables the visual localization and real-time monitoring of tumor sites; second, by leveraging the binding of nuclear targeting groups to nuclear DNA / RNA, it... 10 B is precisely delivered into the nucleus, significantly improving the killing efficiency of BNCT; thirdly, thermal neutron irradiation is performed under imaging guidance, triggering... 10 The B(n,α)⁷Li nuclear fission reaction ultimately enables high-precision integrated diagnosis and treatment. This invention also discloses a method for preparing this probe, which has promising clinical application prospects.
Owner:NANTONG UNIV

A Cross-View Visual Localization Method and System Based on Spectrum-Aware Adaptive Convolution

This application discloses a cross-view visual localization method and system based on spectrum-aware adaptive convolution, belonging to the field of computer vision and localization technology. The method includes acquiring ground images and aerial images, and using a pre-trained large visual model for feature extraction, obtaining a ground image feature map Fg and an aerial image feature map Fa, respectively. Based on a 3D-2D projection model, the ground image feature map Fg is upscaled into a 3D point cloud, and a deformable attention mechanism is used to generate a bird's-eye view feature map. Multi-scale features of the bird's-eye view feature map and the aerial image feature map Fa are extracted based on spectrum-aware adaptive convolution, and a dense correspondence between the bird's-eye view feature map and the aerial image feature map Fa is established. Based on the dense correspondence, the three-degree-of-freedom pose of the ground image relative to the aerial image is calculated by regression. This invention solves the problems of viewpoint differences, geometric inaccuracies, and multi-scale matching, and has excellent generalization, practicality, economy, and high application value.
Owner:CHINA TOWER CO LTD +1

Automatic report generation method based on cross-modal memory network

The invention discloses an automatic report generation method based on a cross-modal memory network, and relates to the technical field of automatic generation of pipe network reports, comprising the following steps: step 1, constructing a defect image-report data set of an urban underground drainage pipe network; step 2, extracting image features by adopting an aggregation discriminant attention mechanism and generating a defect area attention mask; step 3, reinforcing the key area by utilizing the defect area attention mask; step 4, constructing a cross-modal memory network to realize feature alignment; and 5, inputting the enhanced features output by the cross-modal memory network into an encoder-decoder architecture to generate a defect report. According to the method, a collaborative architecture of a fusion aggregation discriminant attention mechanism and a cross-modal memory network is adopted, multi-defect gradient fusion is used for enhancing visual localization and sharing a memory matrix to realize image-text feature explicit alignment, the attention precision of a key area is improved, semantic consistency of terminologies and image features is ensured, and the accuracy of image recognition is improved. And high-precision and structured report output is realized.
Owner:SUZHOU HUIYIKANG DATA TECH CO LTD

Computer mainboard surface mounting real-time control method based on visual positioning

The invention discloses a computer mainboard surface mounting real-time control method based on visual localization, which relates to the technical field of automatic mounting, and comprises the following steps: respectively acquiring the actual position information of a to-be-mounted area and the actual position information of a to-be-mounted element; dynamically adjusting the image acquisition resolution and the processing strategy; calculating a first position deviation of the to-be-mounted area relative to the theoretical position; calculating a second positional deviation of the element relative to the pick-up center of the mounting head; generating a real-time motion trail compensation amount of the mounting head; and according to the real-time motion trail compensation amount, dynamically controlling a mounting head to finish accurate mounting of the element. The method has the advantages that by introducing a visual strategy dynamic adjustment mechanism based on the predefined mounting area risk level, the optimal balance of high-precision mounting and system efficiency is realized, a complete closed-loop quality control chain is formed, data support is provided for process optimization, and the quality of the system is improved. And the quality, the reliability and the intelligent level of surface mounting of the computer mainboard are comprehensively improved.
Owner:SHENZHEN TUOJUNCHENG TECH CO LTD

A 3DGS visual relocalization method, system, terminal, and storage medium based on intrinsic image decomposition.

This invention relates to the field of image localization technology, and discloses a 3DGS visual relocalization method, system, terminal, and storage medium based on intrinsic image decomposition. The method includes: acquiring multi-view scene image data; inputting the multi-view scene image data into a 3D Gaussian sputtering model for training to obtain an intrinsic 3DGS map model; acquiring a query image to be localized and a set of landmarks; inputting the query image into the intrinsic 3DGS map model for decomposition to obtain a target albedo map; performing feature matching based on the target albedo and the set of landmarks to obtain an initial pose; rendering the initial pose based on the intrinsic 3DGS map model to obtain a dense candidate albedo map and a depth map; and adjusting the initial pose based on the dense candidate albedo map, the depth map, and the target albedo map to obtain the localization camera pose. This invention achieves accurate visual localization by visually relocalizing a query image using an intrinsic 3DGS map model.
Owner:GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)

A sensor-computer integrated visual localization method and system for robot target localization

This invention discloses a sensor-computer integrated visual positioning method and system for robot target localization. Image data is acquired by a sensor board equipped with sensors and sent to a core processor. Image processing and target edge extraction are performed at the programmable logic unit (PL), and the data is transmitted to the processing system unit (PS). The PS reads edge points, performs arc segment extraction to obtain elliptical arc segments, and writes them to the PL. Based on this, the PL combines this with a symmetric matrix to optimize the pulsation array and achieve ellipse fitting to complete target localization. An interface board microcontroller controls the power supply timing and manages external interfaces. The core processor (PS) and peripheral interfaces are responsible for overall system control and Ethernet communication. Integrating image sensing and edge computing into a single hardware component, through collaborative processing, hardware acceleration of key algorithms, optimized data transmission, and integrated management, this method achieves low latency, high accuracy, and high robustness in robot target localization, and possesses advantages such as compact structure, low power consumption, high adaptability, and rapid deployment.
Owner:HUNAN UNIV

Multimodal pixel-level detection method, device, system, and storage medium for road surface cracks

This invention discloses a method, device, system, and storage medium for multimodal pixel-level detection of road cracks, comprising: constructing an automated data synthesis pipeline; constructing a multimodal road crack dataset; deeply fusing the general visual language model Qwen2.5-VL with a visual segmentation model through a cue injection mechanism to construct a multimodal pixel-level detection model for road cracks; designing a coordinate transformation module and a text projection module to encode and inject the explicit bounding box cues and implicit semantic cues generated by Qwen2.5-VL into the visual segmentation model; fine-tuning Qwen2.5-VL using a hierarchical symmetric low-rank adaptation strategy to enable it to learn professional knowledge in the crack detection field; and finally, outputting image-to-text understanding, visual location coordinates, and inference segmentation masks through the multimodal pixel-level detection model for road cracks. This invention addresses the shortcomings of visual detection methods lacking semantic understanding and large models lacking pixel-level detection accuracy.
Owner:CHENGDU UNIV OF INFORMATION TECH