Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

38 results about "Visual computing" patented technology

Visual computing is a generic term for all computer science disciplines handling with images and 3D models, i.e. computer graphics, image processing, visualization, computer vision, virtual and augmented reality, video processing, but also includes aspects of pattern recognition, human computer interaction, machine learning and digital libraries. The core challenges are the acquisition, processing, analysis and rendering of visual information (mainly images and video). Application areas include industrial quality control, medical image processing and visualization, surveying, robotics, multimedia systems, virtual heritage, special effects in movies and television, and computer games.

Concrete wall full-section deformation monitoring device and method

The invention discloses a concrete wall full-section deformation monitoring device and method, and relates to the technical field of engineering structure health monitoring, and the device comprises an integrated measurement unit, a passive stable characteristic target group and a processing unit. The integrated measuring unit is composed of an area array solid-state laser radar and a high-resolution binocular vision camera which are rigidly and fixedly connected, and is used for synchronously collecting three-dimensional point cloud and two-dimensional images of a wall and an environment. The target group is arranged on an independent stable object near the wall body; the processing unit identifies coordinates of the target group based on laser radar data, establishes a world coordinate system, calculates self pose drift of the measuring unit by periodically detecting position change of the target, and performs coordinate correction on a wall surface relative displacement field obtained by visual calculation by using the drift distance; and finally, outputting a full-section absolute deformation field under a world coordinate system. According to the invention, high-precision monitoring of full-field continuous deformation of the concrete wall is realized, and the reliability and comprehensiveness of structural health monitoring data are improved.
Owner:TAIXING ENG CONSTR SUPERVISION CO LTD

Method and system for quickly classifying human body scars through visual calculation

The invention discloses a quick classification method and system for human scars through visual calculation, and relates to the technical field of artificial intelligence, and the method comprises the following steps: S1, scar image collection and preprocessing: obtaining a human scar region image and shooting parameters, and carrying out the noise removal, illumination correction and scar region segmentation of an original image; s2, multi-modal scar feature extraction: based on the segmented scar region, extracting color features, texture features, morphological features and depth features; s3, feature fusion and dimension reduction; s4, constructing and training a lightweight scar classification model; s5, outputting a classification result and dynamically optimizing the model; according to the rapid classification method and system for the human body scars through visual calculation, through multi-modal feature fusion and improvement of a lightweight model, the classification accuracy is better than that of a traditional manual classification and single feature machine learning method, subjective differences of doctors are eliminated, and the classification objectivity is guaranteed.
Owner:HANGZHOU PLASTIC SURGERY HOSPITAL CO LTD

Rapid dust removal system and method for engineering machinery welding

The invention provides a rapid dust removal system and method for engineering machinery welding, and the system comprises a rack, an air suction cover which is movably arranged on the rack and is used for capturing welding smoke dust; the dust removal purifier is arranged on the rack along with the suction hood and is used for purifying welding fume; the visual detection device comprises a camera and a visual calculation unit in communication connection with the camera, and the camera is arranged on the air suction cover and used for recognizing the welding point position in real time; the visual calculation unit is used for calculating the relative position of the welding point position and the suction hood and outputting a control signal; the servo driving mechanism is used for driving the dust removal purifier and the air suction cover to move along the rack, the servo driving mechanism is in communication connection with the visual inspection device, and the air suction cover is driven to move to the position over the welding point position according to the control signal; the dust removal efficiency can be improved, the energy consumption can be reduced, and the welding gun can be accurately tracked to collect smoke dust.
Owner:AEROSPACE KAITIAN ENVIRONMENTAL TECH CO LTD

Multi-modal data real-time analysis and feedback method and system

The invention provides a multi-modal data real-time analysis and feedback method and system, and the method comprises the steps: collecting a video frame and an audio frame, reading a count value of the same monotonic timer when the collection of the video frame and the audio frame is completed, and generating an audio and video sequence; calculating the behavior popularity of each region in each time slice based on the sequence, and generating a region set for multi-modal event analysis in combination with a preset threshold and a quantity upper limit; scheduling the corresponding video sub-blocks to a visual computing power unit for target positioning and action classification, and executing voice activity detection and keyword category judgment on the time-aligned audio clips at the same time; visual and audio results are fused, and a structured event sequence is generated according to time slice and region compression; and constructing a classroom teaching chain through the sequence, identifying a key event, and finally generating a teaching event description. According to the invention, under the condition of domestic chip combination, cost, time delay, multi-modal consistency and data security controllability are considered, and real-time perception and feedback of teaching behaviors are effectively supported.
Owner:GUANGZHOU KINDLINK INTELLIGENT TECHNOLOGY CO LTD

Three-dimensional digital garment stylization method and device

The invention relates to the technical field of crossing of computer vision, computer graphics and digital fashion, and discloses a three-dimensional digital clothing stylization method, which comprises the following steps of: firstly, constructing a double-encoder color mapping network to realize color alignment of a clothing image and a style image; the problem that training data is difficult to obtain is solved through a self-supervision strategy; then extracting depth features of the multi-view garment image and the style image by using a pre-trained VGG network, and embedding the style features into a garment feature space by means of an edge enhancement nearest neighbor feature matching algorithm; secondly, introducing an attention-guided edge enhancement feature extraction network, extracting edge features, and performing alignment optimization; finally, a stylized three-dimensional digital clothing image which has high style fidelity and clear structure details and supports rendering at any viewing angle is generated, the problems of color distortion, fuzzy details, inconsistent styles and the like in the prior art are effectively solved, and the method is suitable for the fields of virtual fitting, digital fashion design and the like.
Owner:ZHEJIANG SCI-TECH UNIV +1

Dynamic scene image domain adaptation system and method based on four-dimensional Gaussian sputtering, and computer storage medium

The invention discloses a dynamic scene image domain adaptation system and method based on four-dimensional Gaussian sputtering and a computer storage medium, and relates to the field of computer vision and computer graphics, and the method comprises the steps: obtaining dynamic images captured at different time points from a plurality of visual angles; generating a four-dimensional Gaussian model based on the dynamic image; decomposing the four-dimensional Gaussian model into a conditional three-dimensional Gaussian model and an edge one-dimensional time component; extracting a target domain embedding vector; based on the extracted target domain embedding vector and the embedding vector of the three-dimensional Gaussian model, performing affine transformation on each Gaussian point, and mapping the Gaussian representation to the distribution of the target domain; and the multi-view consistency is maintained by predicting the corresponding relationship between different training views. According to the dynamic scene image domain adaptation method based on four-dimensional Gaussian sputtering provided by the embodiment of the invention, the four-dimensional Gaussian is decomposed into the conditional three-dimensional Gaussian and the one-dimensional Gaussian based on time distribution by utilizing extension, so that the multi-view visual consistency can be ensured.
Owner:HARBIN INST OF TECH AT WEIHAI +1

Soil cation exchange capacity detection method based on color developing solution image recognition

The invention discloses a soil cation exchange capacity detection method based on color developing solution image recognition, and relates to the technical field of soil physicochemical property detection. Comprising the following steps: acquiring a supernatant image after a soil sample of to-be-detected soil reacts with a methylene blue solution under a multi-light-source condition; zero-sample semantic segmentation is carried out on color blocks in the supernatant image to extract position information of different color developing areas in the color blocks, and color partitions are obtained; performing color correction on the color subareas to eliminate color deviation of the color subareas under different light source conditions to obtain RGB parameters of the methylene blue circular color development area; and inputting the RGB parameters into an extreme gradient lifting algorithm so as to fit a numerical relationship between the color concentration of the color developing area and the soil cation exchange capacity, thereby obtaining the soil cation exchange capacity of the soil to be detected. According to the method, a complex chemical detection process is converted into a visual calculation task, and an artificial chemical process depending on multiple ion exchange, washing and titration is simplified.
Owner:SICHUAN AGRI UNIV

A Method and System for Synchronous Monitoring of Bridge Structural Displacement Based on Wireless Intelligent Vision

PendingCN122360296APhase correlationLinear drift
This invention discloses a method and system for synchronous monitoring of bridge structural displacement based on wireless intelligent vision, belonging to the field of bridge health monitoring technology. The method involves acquiring images of the bridge monitoring area and performing format standardization processing; synchronizing the clock and frame acquisition timing of multiple monitoring nodes, achieving microsecond-level synchronization through linear drift compensation, hardware-triggered synchronization, and frame buffering technology; calculating sub-pixel displacement data using either target-independent or improved marker tracking modes, the former based on frequency domain phase correlation analysis, and the latter using a constrained attitude decomposition algorithm to improve accuracy; performing localized visual calculations on the displacement data and outputting displacement time-series data; and wirelessly transmitting the data to achieve real-time visualization and data interaction with downstream systems. The system includes modules for image acquisition, node synchronization, displacement calculation, local calculation, wireless transmission, and data interaction. This invention achieves non-contact, high-precision displacement monitoring, improves multi-node synchronization and data real-time performance, and is adaptable to various inspection scenarios.
Owner:GUANGXI UNIV

A blockchain-based complex fusion network training system and method

ActiveCN118865254BData setSimulation
The application relates to a kind of complex fusion network training system and method based on blockchain, including target visual computing network and target feature chain;Target visual computing network is based on public data set and adaptive loss function, and pre-training is carried out for different task types to the target frame in the target monitoring video collected, and target feature information related to the current specific task is calculated;Target feature chain is used to store the target feature information after encryption.The application can process the identification of multiple tasks of pedestrians and vehicles in parallel, while introducing blockchain technology to effectively guarantee the security and non-tamperability of data, greatly improving the usability and practicality of the system, achieving the effect of enhancing system robustness, promoting multi-source data security fusion and improving complex scene perception ability.
Owner:BEIHANG UNIV

A garment three-dimensional reconstruction method based on geometric prior and generated image assistance

PendingCN122391482APattern recognitionData set
The present application belongs to the field of augmented reality, computer vision, computer graphics, and relates to a kind of garment three-dimensional reconstruction method based on geometric prior and generated image auxiliary, comprising: obtaining image, inputting image into trained garment three-dimensional reconstruction model, and obtaining detailed three-dimensional garment grid;Garment three-dimensional reconstruction model includes: multi-modal prior information extraction module, garment generation network and geometry enhancement module;The present application extracts prior information that can reflect geometric structure from synthetic image, i.e.semantic segmentation map and surface normal map, which is jointly trained with real data set, and combined with real data set to construct garment statistical model, and according to garment statistical model, principal component coefficient is restored to garment grid, to make up for the problem of insufficient three-dimensional supervision signal, without relying on additional three-dimensional scanning grid data, improve the representation ability of model to complex garment shape, thereby improve the precision and generalization performance of three-dimensional garment reconstruction.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Active safety intelligent prevention and control system for berthing and unberthing in ship-end sailing

The invention discloses a ship-end sailing berthing and departing active safety intelligent prevention and control system, which comprises a ship visual perception module, navigation equipment, a video storage module, a visual computing server, a terminal host, a network switch, a serial port server and a wireless module, and is characterized in that the ship visual perception module comprises a berthing visual perception module and a sailing visual perception module; the acquisition module is used for acquiring berthing environment video data and navigation video data of a ship; the wireless module is used for transmitting the berthing environment video data and the navigation video data to the management platform in real time; the system can provide the functions of situation awareness, sailing collision prevention and bridge collision prevention early warning in the sailing process, provide ship side images, berthing and leaving data and early warning information for a driver in the ship berthing process, reduce the berthing blind area of the driver, provide berthing safety guarantee and reduce the collision risk of the ship in the berthing and leaving period to the maximum extent.
Owner:JIANGLONG BOAT TECH

Unmanned aerial vehicle navigation method and system based on vision and reinforcement learning

ActiveCN121740052BEnsure strict spatial and temporal consistencySuppress step jumpsSensing dataVision based
The present application relates to navigation control technical field, especially to a kind of unmanned aerial vehicle navigation method and system based on vision and reinforcement learning, method includes: obtaining lagged first time visual image, real-time second time depth image and the continuous sensing data between them;Through the preset distillation vision network extraction global semantic features and with positive shadow image matching, the visual positioning anchor point of first time is solved;Inertial compensation positioning point of current time is obtained by using sensing data to calculate displacement and attitude change in time window, and mapping visual positioning anchor point to second time;Tensor splicing is carried out to inertial compensation positioning point and depth image to build reinforcement learning state space, and control instruction is generated by inputting strategy network;Through inertial compensation mechanism, the time-space deviation caused by visual calculation delay is eliminated, the flight trajectory oscillation problem caused by different step of multi-source heterogeneous data is solved, and the high robustness autonomous navigation and smooth control of unmanned aerial vehicle in dynamic environment are realized.
Owner:GHOSTCLOUD

A construction safety visual calculation and resource scheduling method, system and electronic equipment

The application provides a construction safety visual calculation and resource scheduling method, system and electronic equipment, which comprises the following steps: deploying and initializing the calculation and resource scheduling system; the edge end collects video streams in real time and performs preprocessing, light-weight model inference and local alarm; the edge end performs three-level screening on the video data, intelligent compression of the region of interest (ROI) and end-to-end encryption before uploading to the cloud; a multi-objective optimization decision model is constructed based on FAHP, and the allocation strategy of the task in the edge end and the cloud is dynamically adjusted according to the real-time collected network state parameters, computing power state parameters and event characteristic parameters; the cloud adopts a high-precision model for inference, links a knowledge base to generate a rectification scheme, and realizes full-process closed-loop management; an incremental data set is constructed to fine-tune the light-weight model and the high-precision model. The scheme provided by the application can balance the contradiction between the real-time alarm demand of the construction site and the high-precision calculation demand of the complex model.
Owner:CCCC SHANGHAI DREDGING CO LTD

Diffusion model-based three-dimensional digital human and object interactive motion synthesis method and system

The invention belongs to the field of computer vision, computer graphics and robots, and relates to a three-dimensional digital human and object interactive motion synthesis method and system based on a diffusion model. The method comprises the following steps: acquiring three types of feature vectors: a three-dimensional grid shape of an object, a frame-by-frame motion sequence of the object and a body type feature vector of a digital human; performing conditional feature coding on the obtained three types of feature vectors to obtain conditional vectors; on the basis of the diffusion model, predicting noiseless estimation of the current time step by using a denoising device in a condition vector iteration manner, and optimizing the generated interactive motion sequence of the three-dimensional digital human and the object according to a generation guide strategy; and obtaining a three-dimensional grid of the human body based on the interactive motion sequence of the three-dimensional digital human and the object, and obtaining an interactive motion synthesis result of the three-dimensional digital human and the object through three-dimensional modeling software. According to the invention, a three-dimensional digital human and object interactive motion sequence with strong sense of reality can be synthesized, and the synthesized three-dimensional digital human has the motion of the trunk and the two hands at the same time.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

A scene gray image recovery system and method based on a dynamic vision sensor

The application discloses a scene gray image recovery system and method based on a dynamic visual sensor, and belongs to the field of computer vision and computational imaging. The application gradually opens the light of an image acquisition system by an incident light intensity control component, acquires the brightness change events on the sensor plane in the process by the dynamic visual sensor, obtains a time mapping gray image based on the brightness change events and the change relationship of brightness with the acquisition time, obtains an event number mapping gray image according to the mapping relationship of the brightness change events, the gray scale and the event number, and obtains the scene gray image according to the time mapping gray image and the event number mapping gray image. The application enables the dynamic visual sensor which can only output the brightness change discrete information to also output the gray image. Compared with the prior art, the application is compatible with static scenes and has better recovery effect.
Owner:ZHEJIANG UNIV

Multi-split control method and system for logistics packet supply station based on vision and cloud collaboration

The invention discloses a logistics package supply station multi-split control method and system based on vision and cloud collaboration, and the method comprises the steps: independently collecting package bar codes and image information by each package supply station, and transmitting task data to a server after local verification; the server performs priority ranking and resource allocation on the tasks according to a dynamic scoring mechanism; reasoning the batch processing tensor by using a detection model, and outputting information of packages in all the images; performing multi-dimensional rule filtering on an original result output by the detection model, and starting a processing flow for a triggered abnormal condition; and the packet supply station verifies and executes the control instruction, and feeds back confirmation information to the server after completing the packet loading action. Through a multi-split intensive processing mode, the visual computing units originally needing to be independently deployed on each packet supply platform are concentrated to the server for unified processing, the hardware cost is greatly reduced, meanwhile, only the packet supply platforms need to be added during expansion, and computing resources do not need to be repeatedly configured.
Owner:ZHEJIANG KANGLI AUTOMATIC CONTROL TECH

A method and system for adaptive alignment of table data rows and columns based on edge visual computing

The application discloses a kind of based on edge visual computing's table data row and column self-adapting alignment method and system, the method includes creating table video stream self-adapting extractor, real-time intercepts video frame containing table data, identifies target table position information, extracts table image to be detected;For target table image, construct the self-adapting alignment model based on visual computing, splice multiple sets of table profile of video frame, align table row and column data;Based on table row and column data, construct deep reinforcement learning model, fine adjustment table structured information;Based on table structure information, customization identification and processing cell text data;According to table structure and text data, carry out structured encoding and storage.The application realizes the automatic identification, storage and alignment of table data in video stream, improves the acquisition efficiency and accuracy of table data, greatly saves manpower cost.
Owner:CHENGDU JINFA EDGE INTELLIGENT TECHNOLOGY CO LTD

Image classification method and device based on lateral inhibition attention mechanism, and electronic equipment

The invention relates to an image classification method and device based on a side suppression attention mechanism and electronic equipment, and the method comprises the steps: constructing an image classification model based on the side suppression attention mechanism, and carrying out the training through employing a neuromorphic data set, and obtaining a trained image classification model based on the side suppression attention mechanism; inputting a to-be-processed neuromorphic image into the trained image classification model based on the lateral inhibition attention mechanism to obtain an image classification result; according to the method, secondary features and background information are actively inhibited, so that the network is more focused on a key visual mode, and the feature identification capability and robustness of the model are improved under the condition that the parameter quantity is not excessively increased; according to the method, the effectiveness and generalization ability of a side suppression attention mechanism in an image processing task provide a new thought and method support for a low-power-consumption and high-efficiency bionic vision calculation model.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Event-driven asynchronous graph neural network FPGA accelerator for real-time edge vision

The invention belongs to the technical field of artificial intelligence chips and reconfigurable computing, and particularly relates to an event-driven asynchronous graph neural network FPGA accelerator for real-time edge vision. Aiming at the problems of low storage utilization rate, obvious data access bottleneck, limited degree of parallelism, too high calculation redundancy and the like of a traditional GNN accelerator in an event vision scene, the invention provides a novel architecture combining efficient graph feature storage, hierarchical graph construction and redundancy elimination convolution calculation. By introducing parallel read-write optimization, low-dependence graph structure generation and a reusable computing cache mechanism, the computing delay is remarkably reduced while resource consumption is kept controllable, the overall throughput rate is increased, and therefore the sub-microsecond real-time reasoning capacity is achieved. The method is suitable for unmanned driving, intelligent monitoring, robot navigation and other application scenes needing low-delay visual calculation on edge equipment.
Owner:SHANGHAI TECH UNIV

Low-power always-on face detection, tracking, recognition and / or analysis using event-based vision sensors

Relates to low-power always-on face detection, tracking, recognition and / or analysis using event-based visual sensors. The techniques disclosed herein utilize a vision sensor that integrates a dedicated camera with dedicated computer vision (CV) computing hardware and a dedicated low-power microprocessor for the purpose of detecting, tracking, recognizing, and / or analyzing subjects, objects, and scenes in the field of view of the camera. The vision sensor processes information retrieved from the camera using the included low-power microprocessor and sends an "event" (or an indication that one or more reference conditions have occurred, and possibly associated data) for a host processor only when needed or as defined and configured by an application. This allows a general purpose microprocessor, which is typically relatively high speed and high power to support various applications, to remain in low power (e.g., sleep mode) as conventionally for most of the time, while becoming active only upon receiving an event from the vision sensor.
Owner:QUALCOMM INC

SYSTEM INSIDE A VEHICLE FOR COMMUNICATION WITH A PERSON NEAR THE VEHICLE

A system (11) within a vehicle (10) for communication with a person (50) in the vicinity of the vehicle (10), comprising the following: a perception system (38) that is connected to a control unit (34A) of the system (11), wherein the perception system (38) is capable of detecting the presence of a person (50) in the vicinity of the vehicle (10); a monitoring system (52) that is connected to the control unit (34A) of the system (11) and can monitor the position of the head and eyes of the person (50); a generator (55) for holographic images comprising the following: a visual computing machine (56) that communicates with the control unit (34A) of the system (11) and the monitoring system (52) and can compute a holographic image (58) and encode the holographic image (58) for a display (60) of an image generation unit (PGU) (62); and a beam guidance device (68), wherein: the beam guidance device (68) is designed to receive information regarding the position of the head and eyes of the person (50) from the monitoring system (52) via the visual computing machine (56); and The display (60) is designed to project the holographic image (58) onto the beam guidance device (68), and the beam guidance device (68) is designed to redirect the projected holographic image (58) to the eyes of the person (50) on the basis of the information received from the monitoring system (52). wherein the visual computing machine (56) is further configured to encode a lens function into the holographic image (58) on the basis of the information received from the monitoring system (52), wherein the beam guiding device (68) is a MEMS mirror, microelectromechanical systems, a two-galvanometer mirror, a rotating Risley prism pair or a one-dimensional illusion mirror in combination with a rotating polygon mirror.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Method and device for calculating pipeline flow based on computer vision and storage medium

The invention relates to the field of computer vision, and provides a method and device for calculating pipeline flow based on computer vision and a storage medium, and the method comprises the steps: obtaining a target video stream at a pipeline opening of a target pipeline, the target video stream comprising a flowing fluid and at least one drifting object; acquiring a pipe wall mask image and a flowing fluid mask image from the target video stream based on a computer vision algorithm; determining the cross sectional area of the target pipeline according to the pipe wall mask image and the flowing fluid mask image; obtaining the flow velocity of the drifting object according to the time point of the drifting object and the pixel point position change of the drifting object in the plurality of video frames; taking the flow velocity of the drifting object as the flow velocity of the flowing fluid, and obtaining the target flow of the flowing fluid according to the flow velocity and the area of the flowing fluid passing through the cross section of the target pipeline. According to the scheme, the flow of various types of fluid can be detected without damaging the pipeline structure, the implementation is high, the pipeline flow detection equipment does not need to be in direct contact with the fluid in the pipeline, and the installation environment is simple.
Owner:WUHAN DAOXIAOFEI TECH CO LTD

Directional visual and audio communication

A system within a vehicle for communicating with a person in proximity to the vehicle includes a perception system in communication with a system controller and adapted to detect the presence of a person, a monitoring system adapted to monitor the position of the person's head and eyes, a holographic image generator including a visual compute engine adapted to calculate a holographic image and encode the holographic image onto a spatial light modulator of a picture generating unit, and a beam steering device, wherein the beam steering device is adapted to receive information related to a position of the person's head and eyes from the monitoring system, and the display is adapted to project the holographic image to the beam steering device and the beam steering device is adapted to re-direct the projected holographic image to the eyes of the person.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

A polarization-sensitive all-optical modulation memristor device based on a directional single-walled carbon nanotube array and titanium dioxide heterojunction and a preparation method thereof

PendingCN122396215AHeterojunctionExcitatory synapse
The application discloses a polarization-sensitive all-optical modulation memristor device based on a directional single-wall carbon nanotube array and a titanium dioxide heterojunction and a preparation method thereof, and relates to the technical field of intelligent neuro-morphological visual computing and microelectronic devices. The application constructs a heterojunction by magnetron growth of a titanium dioxide film on a directional arrangement single-wall carbon nanotube array film prepared by a DLSA method to obtain a polarization-sensitive all-optical modulation memristor device. The memristor device under the structure has enhanced conductance under 980 nm near-infrared light stimulation and reduced conductance under 320 nm ultraviolet light stimulation. Based on the non-volatile characteristics of all-optical modulation, the device simulates basic synaptic plasticity functions, including excitatory postsynaptic current (EPSC) and inhibitory postsynaptic current (IPSC). Meanwhile, due to the photoelectric anisotropy of the directional arrangement single-wall carbon nanotube array, the optical response of the device also exhibits polarization angle-dependent characteristics, and the anisotropy ratios of the optical response amplitudes under 980 nm and 320 nm excitation are 1.85 and 1.52, respectively.
Owner:NORTHEAST NORMAL UNIVERSITY

Cross-modal target tracking method based on natural language and visual model fusion of unmanned aerial vehicle

The invention belongs to the technical field of cross-modal target tracking. The invention provides a cross-modal target tracking method based on natural language and visual model fusion of an unmanned aerial vehicle. According to the embodiment of the invention, the natural language, the search image and the template image are respectively encoded by adopting the pre-trained language encoder and the visual encoder, and the calculation overhead is reduced by freezing a part of network layers. In order to reduce redundancy of visual calculation, an image block selection strategy based on content awareness is designed, and similar image blocks are dynamically shielded by using cosine similarity, so that storage and calculation efficiency is optimized. In the feature fusion stage, a visual guidance channel attention mechanism is introduced, and a multi-head self-attention network is combined to realize deep interaction of cross-modal features. And the spatial position and size of the target are accurately captured by using the modal self-adaptive frame head module.
Owner:XIAN TECH UNIV

Methods and systems for multi-hypothesis-based fusion of sensor data

ActiveCN114942430BImage enhancementImage analysisMultiple hypothesisRadar
This document describes a multi-hypothesis-based data fusion tracker. Each hypothesis corresponds to a different pseudo-measurement type. The fusion tracker automatically determines which pseudo-measurement type is more likely to be accurate in the current situation using a predefined error covariance associated with radar. The fusion tracker can rely on either a combination of radar and vision computations, or it can ignore vision-based pseudo-measurements and rely solely on radar pseudo-measurements. By selecting between three different bounding boxes (view-angle-based, view-lateral-position-based, or radar-based only), the fusion tracker can balance accuracy and speed when drawing, repositioning, or resizing bounding boxes, even in congested traffic or other high-volume conditions.
Owner:APTIV TECHNOLOGIES AG

A robust event-driven gait recognition method, system, device and storage medium based on event stream

This invention discloses a robust event-driven gait recognition method, system, device, and storage medium based on event flow, relating to the fields of event vision, computer vision, pattern recognition, and intelligent security technologies. The method includes renormalizing spatial displacement, temporal displacement, and edge length according to a unified scale and robustness scale, recalculating edge attributes after each pooling, introducing a motion intensity index to assist in determining edge reliability, employing continuous reweighting instead of direct edge deletion, and using motion consistency, radius validity, orientation validity, and entropy constraints to weakly supervise edge confidence. A graph convolutional backbone network is used to extract spatial graph features for each time slice, and temporal relationships are jointly modeled through difference and similarity branches. The additive angular interval loss and dynamic center loss are jointly optimized. This invention improves message passing stability, enhances anti-disturbance robustness, overcomes intra-class fluctuations such as cross-viewpoint and low-light conditions, and possesses excellent potential for edge device deployment.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Unmanned aerial vehicle navigation method and system based on vision and reinforcement learning

The invention relates to the technical field of navigation control, in particular to an unmanned aerial vehicle navigation method and system based on vision and reinforcement learning, and the method comprises the following steps: obtaining a lagged first moment visual image, a real-time second moment depth image and continuous sensing data between the lagged first moment visual image and the real-time second moment depth image; extracting global semantic features through a preset distillation visual network, matching the global semantic features with the positive image, and resolving a visual positioning anchor point at a first moment; calculating displacement and posture change in the time window by using the sensing data, mapping the visual positioning anchor point to a second moment, and obtaining an inertia compensation positioning point at the current moment; performing tensor splicing on the inertial compensation positioning point and the depth image to construct a reinforcement learning state space, and inputting a strategy network to generate a control instruction; the space-time deviation caused by visual calculation delay is eliminated through an inertia compensation mechanism, the problem of flight path oscillation caused by asynchronism of multi-source heterogeneous data is solved, and high-robustness autonomous navigation and smooth control of the unmanned aerial vehicle in a dynamic environment are realized.
Owner:GHOSTCLOUD