Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

491 results about "Vision processing" patented technology

Panoramic photographing and machine vision deformation monitoring integrated Beidou monitoring machine system

The invention discloses an integrated Beidou monitoring machine system integrating panoramic photographing and machine vision deformation monitoring, and relates to the technical field of safety monitoring. The system comprises a hardware structure unit and a software function unit, the hardware structure unit comprises a Beidou positioning module, an inertial navigation compensation module, a panoramic vision acquisition module, a machine vision processing module, a multi-mode communication transmission module and a power management module; the software function unit comprises a multi-source data fusion engine, a panoramic three-dimensional modeling sub-module, an intelligent early warning decision module and a cloud collaborative resolving interface. According to the invention, through real-time fusion of Beidou positioning data and panoramic vision imaging, an improved dense optical flow method is adopted to track local crack propagation, and global displacement monitoring of Beidou is combined, so that cross-scale synchronous sensing from macroscopic displacement to microscopic deformation is realized, and the comprehensive monitoring precision is improved.
Owner:HUASI (GUANGZHOU) MEASUREMENT & CONTROL TECH CO LTD

Self-adaptive three-dimensional scene reconstruction method and system based on single panorama

The invention discloses a self-adaptive three-dimensional scene reconstruction method and system based on a single panorama, and belongs to the field of computer vision processing. The method comprises the following steps: firstly, generating a depth map through an indoor scene panorama and constructing an initial three-dimensional grid; then generating a multi-view image and a shielding mask thereof based on view conversion, forming a training sample pair for fine tuning of the diffusion completion model, and injecting scene prior information for the model; then extracting a grid boundary contour and calculating a central axis, and adaptively constructing a camera pose set; after the integrity of the grid is optimized through iteration completion, the grid is converted into a 3D Gaussian sputtering field, and Gaussian point parameters are optimized through multi-view data; in the optimization process, an up-sampling strategy of rendering error feedback and edge Gaussian points is adopted, and finally a scene model with complete geometric textures is output. According to the method, a high-quality three-dimensional scene is reconstructed from a single panorama through a structure self-adaptive completion and Gaussian refining technology, and the method is particularly suitable for immersive roaming reconstruction of indoor scenes.
Owner:ZHEJIANG UNIV

Resource and task aware visual processing edge adaptive decision-making method

The invention belongs to the technical field of artificial intelligence and computer vision, particularly relates to a visual processing edge adaptive decision-making method for resource and task perception, and aims to solve the problem of scheduling mismatch caused by resource dynamic change and task demand diversity in visual task processing in an edge computing environment. The method comprises the following steps: collecting multi-dimensional resource state data of edge nodes in real time to form a resource state vector with high time resolution; analyzing the visual task request, and constructing a quantifiable task feature vector; and establishing a resource-task association mapping model based on a dynamic weight distribution mechanism. The method also supports cross-edge domain collaborative decision, and processes a pipeline dynamic reconstruction and security isolation mechanism. According to the technical scheme, the fluctuation of the resource utilization rate is reduced to 15% or below, the average task processing delay is reduced to 60%, the scheduling satisfaction degree is improved by 40% or above, and the self-adaptability and the service quality guarantee capability of the edge vision system are remarkably enhanced.
Owner:SHENZHEN IBD INTELLIGENT TECH CO LTD

Unattended crucible loading and unloading operation control system and method based on visual processing

The invention relates to the technical field of operation control, in particular to an unattended operation crucible loading and unloading operation control system and method based on visual processing, and the method comprises the following steps: monitoring a static accumulation value in a graphite powder fluidization conveying process in real time through a distributed static sensor; when the static accumulation value exceeds a safety threshold value, pulse type ion wind neutralization treatment is triggered, and a static balance material flow is generated; the dynamic alignment deviation of the stock bin and the crucible opening is calculated; the stepping angle and the rotating speed of a spiral auger in the stock bin are controlled based on the dynamic alignment deviation, and an axial vibration field is synchronously applied in the discharging process; the position of an air gap in a material is predicted based on a convolutional neural network, and a vacuum negative pressure value is dynamically adjusted to form a gradient suction mode. According to the device, the risks of powder agglomeration, conveying pipeline blockage or electrostatic discharge caused by electrostatic accumulation are effectively prevented, the safety of the conveying process is ensured, and meanwhile, the flowing stability of graphite powder is optimized.
Owner:青岛雷英智能科技有限公司

Obstacle avoidance navigation system based on binocular vision

The invention discloses an obstacle avoidance navigation system based on binocular vision. The obstacle avoidance navigation system comprises a binocular camera and a GPS module. A visual processing and depth estimation algorithm module; a map generation module; and a path planning and dynamic obstacle avoidance module. The binocular camera and the GPS module are combined, so that visual-geographic information fusion is realized, and the environmental perception precision is improved; the three-dimensional point cloud generated by depth calculation provides three-dimensional space data support for obstacle avoidance, and is safer and more reliable than traditional two-dimensional recognition. The navigation map constructed by the semantic segmentation and target detection technology can dynamically distinguish complex obstacle types; the path combination strategy of the shortest path algorithm and the map API not only ensures the real-time performance of local obstacle avoidance, but also gives consideration to the global path optimality, so that the navigation efficiency of the system in complex terrains is higher.
Owner:LINKER

Improved YOLOv5-based unmanned aerial vehicle target detection method and system, equipment and medium

The invention relates to the technical field of visual processing, and provides an improved YOLOv5-based unmanned aerial vehicle target detection method, which comprises the steps of S1, acquiring a training image set; s2, performing target labeling on each training image to form a labeling file; s3, training the training image set and the annotation file based on the improved YOLOv5 to obtain an improved YOLOv5 model; the improved YOLOv5 model comprises an input end, a backbone network, a neck network and a detection head which are connected in sequence; the backbone network is used for extracting features; the backbone network comprises a multi-stage depth separable convolution module, a multi-stage C3 module and a dynamic window attention module; each stage of C3 module is located between the two stages of depth separable convolution modules; the neck network is used for carrying out feature fusion; the detection head is used for outputting detection information; and S4, inputting a to-be-detected image set to the improved YOLOv5 model to obtain a detection target. According to the scheme, the scale adaptability can be improved, dense target leak detection is prevented, and the calculation efficiency and the positioning precision are improved.
Owner:BLUE SKY LABORATORY +1

Self-supervised end-to-end visual reconstruction method and system

The invention provides a self-supervised end-to-end vision reconstruction method and system, and relates to the technical field of computer vision processing. The method comprises the following steps: firstly, acquiring multi-camera parameters and different-view-angle image data, including focal length, lens distortion and main coordinate point data of each camera, and pixel point coordinates of a current frame of a first camera and a reference frame of a second camera, then constructing an end-to-end training model, and calculating a re-projection error of pixel points of the current frame and the reference frame to obtain a multi-view-angle image; and solving the parameter update quantity by using a Gaussian Newton iteration method, iteratively optimizing the camera pose, the pixel corresponding relation and the depth data, and reconstructing a three-dimensional coordinate by combining the obtained data set after the re-projection error is converged, and converting and splicing to realize three-dimensional scene reconstruction. By implementing the scheme, end-to-end self-supervised training can be realized under the condition of not depending on manual annotation, so that three-dimensional visual reproduction is realized.
Owner:DOMINANT INTELLIGENT TECH (SUZHOU) CO LTD

Three-dimensional scene target blanking method and system based on 3DGS

The invention discloses a three-dimensional scene target blanking method and system based on 3DGS, and belongs to the field of computer vision processing. The method comprises the following steps: firstly, performing three-dimensional reconstruction through a multi-view image to generate an original Gaussian point cloud scene; after a scene is divided into foreground and background point clouds based on target text information, spatial expansion is performed on the foreground to generate a three-dimensional soft mask. And projecting the mask to a two-dimensional image space to generate a two-dimensional mask image of each view angle, and inputting the two-dimensional mask image and an original image into an image blanking model to generate a candidate blanking image set. After a user selects a reference image from an initial view angle, a candidate image most similar to a preorder view angle is selected as a final blanking result from a non-initial view angle, and fidelity enhancement is carried out. And finally, optimizing the original Gaussian point cloud parameters by using the blanking result of each view angle to obtain a target three-dimensional scene after blanking. Based on the invention, a user can carry out texture remodeling on the structure of the original three-dimensional scene, and the multi-view consistency of the three-dimensional scene is maintained.
Owner:ZHEJIANG UNIV

Time sequence NDVI crop distribution extraction method based on mask auto-encoder

The invention belongs to the technical field of visual processing, and particularly relates to a time sequence NDVI crop distribution extraction method based on a mask auto-encoder, which mainly comprises four steps. Firstly, time sequence NDVI data are prepared, a multi-temporal remote sensing image in a complete growth cycle of target crops in a target area is obtained and processed, and the recognition precision is improved by using NDVI feature changes in the growth stage of the crops. And then performing time sequence transformation on the ViT model, dividing data space dimensions, completing Patch flattening and embedding, enabling the Patch to be adaptive to time sequence data, and simultaneously performing image processing advantages. Then, model training is carried out, a mask auto-encoder is pre-trained in a self-supervised mode through a large amount of unlabeled data, and then supervised fine tuning is carried out through a small amount of labeled data; and finally, processing new data by using the fine-tuned model, obtaining pixel-level crop category prediction, and generating a complete distribution map. According to the method, by means of time sequence data and model transformation, crop types and growth stage differences are effectively distinguished, and accurate extraction is achieved.
Owner:HUANTIAN SMART TECH CO LTD

Unmanned aerial vehicle multi-scale crop detection method based on neurodynamics model

The invention discloses an unmanned aerial vehicle multi-scale crop detection method based on a neurodynamics model, and belongs to the technical field of crop detection, and the method comprises the steps: obtaining a visible light image, a long-wave infrared image and hyperspectral data of a crop under the multi-scale condition of an unmanned aerial vehicle; performing registration fusion processing on the visible light image, the long-wave infrared image and the hyperspectral data to obtain fused hyperspectral cube data; constructing a neurodynamic model for crop detection, inputting the fused hyperspectral cube data into the neurodynamic model for crop detection for training, and outputting a crop detection result; and calculating a loss function of the neurodynamic model for crop detection based on a crop detection result, and feeding back the loss function to the model for constraint. According to the invention, visual processing, deep learning and agricultural scene perception technologies are combined, and efficient and accurate crop monitoring is realized by using the parallel computing capability and anti-interference characteristics of the neurodynamics model.
Owner:SOUTHWEAT UNIV OF SCI & TECH +3

Vision processing and model training method, device, storage medium and program product

The present disclosure provides a vision processing and model training method, device, storage medium and program product. A specific implementation solution is as follows: establishing an image classification network with the same backbone network as the vision model, performing a self-monitoring training on the image classification network by using an unlabeled first data set; initializing a weight of a backbone network of the vision model according to a weight of a backbone network of the trained image classification network to obtain a pre-training model, the structure of the pre-training model being consistent with that of the vision model, and optimize the weight of the backbone network by using real data set in a current computer vision task scenario, so as to be more suitable for the current computer vision task; then, training the pre-training model by using a labeled second data set to obtain a trained vision model.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Method to Use a Single Camera for Barcoding and Vision

Systems and methods for performing barcoding and machine vision with a single camera are disclosed herein. An example system includes an image sensor configured to capture low-resolution image data of a large field of view and high-resolution image data of a small field of view. A first data pipeline is configured to transmit the low-resolution image data to a first module configured to perform image processing on the low-resolution image data. A second data pipeline is configured to transmit the high-resolution image data to a second module configured to perform image processing on the high-resolution image data. Machine readable instructions cause the system to capture image data of the large field of view or the small field of view and the processor transmits either the low-resolution image data via the first data pipeline or the high-resolution image data via the second pipeline.
Owner:ZEBRA TECHNOLOGIES CORP

Multi-modal fusion bank receipt intelligent processing method and system based on vision and NLP

The invention discloses a multi-modal fusion bank receipt intelligent processing system and method based on vision and NLP, and the method comprises the steps: receiving a bank receipt picture or a PDF document, and completing the text detection, direction correction and character recognition through a visual processing engine; a three-level receipt independent segmentation mechanism is applied, and independent receipt records are divided through spatial clustering analysis, semantic analysis, visual verification and cross-page association processing; jointly extracting text features and layout features of each receipt through a double-flow multi-modal fusion model, and fusing the text features and the layout features; field-level data extraction is executed through a field extraction engine, and financial data verification including account number, amount, date and abnormity quadruple verification is carried out; and generating structured JSON output to obtain a bank receipt processing result. According to the invention, the innovative five-layer processing architecture realizes high-precision analysis of the bank receipts through an independently researched and developed receipt independent segmentation engine, a vision-semantic fusion model and a financial data verification system.
Owner:QINGDAO WHALE ABACUS TECHNOLOGY CO LTD

Tomato stem picking point attitude estimation method

The invention discloses a tomato stem picking point attitude estimation method, and relates to a computer vision processing technology. The method comprises the following steps: detecting a long-range directional frame target, acquiring a color flow by a depth camera, and detecting position information and shielding conditions of fruits and fruit stems through a directional frame; visual guidance: matching fruit stems based on tomato growth constraint conditions, and visually guiding the mobile platform and the mechanical arm to move in a coupling manner until the fruit stems are observed when the fruit stems are not matched; detecting close-range key points, regulating and controlling a mechanical arm to a close-range observation position, identifying two points of a main stem and three points of a fruit stem through a key point detection model, and extracting point clouds; performing point cloud data processing, removing point cloud noise through point cloud filtering, and performing adaptive compensation on key points with depth anomaly by using a neighborhood mean value; calculating picking poses, and establishing the picking poses of the fruit stem shear points according to the three-dimensional information of the five key points; according to the method, the problem of tomato shielding in a greenhouse complex scene can be effectively solved, and the fruit stem shearing point pose is obtained.
Owner:CHINA AGRI UNIV

Construction safety real-time high-precision detection method and system based on environmental characteristics

The invention relates to the technical field of video recognition, in particular to a construction safety real-time high-precision detection method and system based on environmental characteristics. Arranging a plurality of sensors in a to-be-detected area to collect field environment data in real time; analyzing the illumination change of the to-be-detected area through the light and shadow analysis model, calculating a light-safety misjudgment index, and identifying a potential light misjudgment area; a visual processing model is utilized to analyze the worker video, human skeleton key points are recognized in real time, personnel behavior safety indexes are obtained, and potential dangerous behaviors are detected; detecting local wind speed and turbulence conditions and vibration and resonance phenomena of a to-be-detected area in real time through a local breeze and vibration detection model, and evaluating a local environment stability index; and the safety level of the to-be-detected area is calculated in real time by using the light-safety misjudgment index, the personnel behavior safety index and the local environment stability index, so that multi-dimensional and real-time construction safety risk assessment is realized.
Owner:ZHEJIANG INST OF COMM CO LTD +2

Image classification with modality dropout

Systems and methods are provided for classifying images associated with an item, and generating an image set for that item which includes image classifications determined to be helpful for the item type of the item. To classify images, an image classification model is generated and trained using two phases. The first phase uses intermediate model with text and visual processing to teach the model to recognize patterns created by text without requiring OCR at inference. The second phase uses visual processing to refine the model for use at inference. To generate an image set, image classifications helpful to an item type are identified, items are associated with item types, images are obtained for an item, the images are classified using the image classification model, missing image classifications set out in the preferred image set are identified, and a request or requests is generated for the missing image classifications.
Owner:AMAZON TECH INC

Computer vision processing method and system for industrial defect real-time detection

The invention relates to a computer vision processing method and system for industrial defect real-time detection. The method comprises the following steps: extracting geometric features and textural features of predefined defect types, and generating a structured descriptor set; generating a synthetic defect image set based on the defect-free image set and the structured descriptor set; inputting the synthesized defect image set into a double-flow feature extraction network to obtain a fusion feature vector; generating a defect category threshold set based on the vector and the structured descriptor set; and inputting the to-be-detected image and the corresponding defect-free reference image into the double-flow feature extraction network, calculating defect probability distribution in combination with the structured descriptor set and the dynamic classifier, and outputting a defect category decision result based on the defect category threshold set. According to the method, the precision, robustness and adaptability of defect detection are improved by means of fusing the global features of the defect-free reference image and the local features of the defect image and expanding training samples by using the synthetic defect image set.
Owner:周骏

Visual processing

According to embodiments of the disclosure, a method, an apparatus, a device, and a storage medium for visual processing are provided. A method includes: converting a plurality of image blocks divided from visual data into a plurality of embedding representations respectively, where the visual data includes an image or a video; extracting, by using a first processing block in a trained visual encoder, first feature information from the plurality of embedding representations according to a first attention mechanism; extracting, by using a second processing block in the visual encoder, second feature information from the first feature information according to a second attention mechanism; and generating, by using a tokenizer in the visual encoder, an encoding representation corresponding to the visual data based on the second feature information. In this manner, the encoding efficiency can be improved, and better universality and scalability can be achieved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Driver fatigue state detection system based on computer vision processing

The invention discloses a driver fatigue state detection system based on computer vision processing. The system comprises a facial feature extraction unit which selects a high-definition camera to capture a driver facial image; the upper body feature monitoring unit collects upper body image data through a high-definition camera to accurately extract the action frequency of the upper body; the hand-held steering wheel characteristic monitoring unit is used for sensing pressure applied by hands of a driver by arranging a grip strength sensor on a steering wheel through a protective sleeve; the data processing unit receives data information collected by the facial feature extraction unit, the upper body feature monitoring unit and the handheld steering wheel feature monitoring unit for feature fusion; the fatigue state judgment unit judges the fatigue state of the driver according to a preset fatigue judgment rule; the face blinking frequency is combined with multi-dimensional information such as the upper body action frequency and the grip strength, the state of a driver can be reflected from different angles, richer and more comprehensive data support is provided for fatigue judgment, and the fatigue state can be judged more accurately.
Owner:SHANDONG JIANZHU UNIV

Three-dimensional Gaussian splash multi-view video joint semantic coding method

The invention provides a three-dimensional Gaussian splash multi-view video joint semantic coding method, and relates to the technical field of computer vision processing. According to the method, the multi-view video image coding and the feedforward 3DGS technology are combined, the architecture of double-view coding and Gaussian parameter prediction parallel operation is adopted, left and right view branch model parameters are shared, the processing real-time performance is improved, and the complexity is reduced. Meanwhile, a cross-view-angle parallel semantic feature auto-encoder is designed, and information interaction between left and right view angle branches is realized by transmitting cross-view-angle context information, so that redundancy between view angles is reduced. And finally, realizing 3DGS lightweight prediction on the decoding features at a receiving end, and rendering to obtain a video image of a user view angle so as to meet the real-time requirement of immersive semantic communication.
Owner:TSINGHUA UNIVERSITY

AOI detection data processing system and method for PCB

The invention relates to the field of PCB detection, in particular to an AOI detection data processing system and method for a PCB, and the system comprises a visual processing module, a contour recognition module, a PCB alignment module, a region aggregation module and a defect detection module, the visual processing module is used for rasterizing each image pixel point, the contour recognition module is used for detecting a communication contour of the PCB, and the PCB alignment module is used for aligning the PCB alignment module. The PCB alignment module is used for aligning a PCB image with a standard image, the area aggregation module is used for identifying a defect view field and judging a qualified state, and the defect detection module is used for establishing a reinspection path and outputting a block defect rate. Production defects can be found in time, the probability of missing detection and false detection is reduced, the production quality of the PCB is improved, stable operation of an AOI system in a production environment is guaranteed, and the detection efficiency and accuracy are improved.
Owner:KAIPING ELEC & ELTEK CO LTD

Garment color fastness gradient detection method based on visual processing

The invention belongs to the field of garment processing, and particularly relates to a garment color fastness gradient detection method based on visual processing, which comprises the following steps: carrying out standardized color fastness test on a garment sample, and collecting corresponding picture samples and color fastness parameters according to a test time sequence; preprocessing the collected picture samples according to a test time sequence, wherein the preprocessing comprises denoising, enhancement and color correction of the picture samples; extracting color features and texture features of the preprocessed picture samples, performing gradient classification on various picture samples according to a test time sequence, and establishing a color feature data set, a texture feature data set and a color fastness parameter data set which are associated; through standardized testing and automatic image acquisition, hundreds of samples can be processed at a time, predicted values are directly output based on a trained model, physical testing steps are reduced, key areas such as necklines and cuffs are positioned through area segmentation, local area testing is rapidly carried out, and compared with global testing, the method has the advantages of being high in pertinence and accuracy, and missing detection can be avoided.
Owner:BEIJING KECE INFORMATION TECHNOLOGY CO LTD

Machine-learning algorithms for low-power applications

Systems, computer programs, devices, and methods that enable ML-based vision processing for low-power, embedded, and / or real-time applications. In one exemplary embodiment, smart glasses use classifiers that are based on machine-learned (ML) patch relationships. The ML patch features are determined during an offline training process. The ML patch features are grouped into weak classifiers, strong classifiers, and detectors to progressively improve prediction accuracy. An object detection architecture uses triggering logic, search management, and a classification neural network to enable event-based searching, interest-based searching, and / or dynamic search control. In some cases, pre-processing may also be used to minimize the neural network complexity (e.g., pre-processing for scaling, rotations, translations, etc.).
Owner:SOFTEYE INC

Dynamic self-adaptive edge server visual task processing method and system

The invention discloses a dynamic self-adaptive edge server visual task processing method and system, and belongs to the technical field of visual task processing, and the method comprises the steps: constructing an edge load model and an energy risk model, outputting an edge load rate and an energy risk coefficient, and generating a local execution confidence coefficient based on the edge load rate and the energy risk coefficient; generating a cloud execution confidence coefficient in combination with the network index; quantizing the data quality and the environment state through the data evaluation model and the data acquisition environment model; fusing the data timeliness factors to construct a local-data adaptation degree and a cloud-data adaptation degree; local / cloud execution is decided according to the decision model, and the resolution is optimized based on the execution adaptation degree. According to the method, resource dynamic adaptation and task quality optimization are realized, and the edge visual processing efficiency is remarkably improved.
Owner:SHENZHEN IBD INTELLIGENT TECH CO LTD

Aviation transient electromagnetic pod coil attitude correction method and system

The invention discloses an aviation transient electromagnetic pod coil attitude correction method and system, and belongs to the field of aviation geophysical exploration, and the system comprises a surface scanning industrial camera module which is used for obtaining the image information of a suspension coil in real time; the industrial-grade integrated navigation module is used for outputting inertial navigation data; the time synchronization and data acquisition module is used for establishing a time unification mechanism; the visual processing module is used for extracting the image processed by the time unification mechanism to obtain visual features; the inertial navigation resolving module is used for obtaining an inertial navigation state of a coil attitude angle through angular velocity integration; the data fusion and attitude estimation module is used for performing fusion calculation by combining the visual features and the inertial navigation state to obtain a fusion result; and the projection area and magnetic moment calculation module is used for calculating the ground projection area and the effective emission magnetic moment component of the coil according to the fusion result. According to the method, the system complexity and the electromagnetic interference risk caused by multi-inertial navigation layout are reduced, and the detection precision and stability of the aviation transient electromagnetic data are improved.
Owner:INSTITUTE OF GEOLOGY AND GEOPHYSICS CHINESE ACADEMY OF SCIENCES

Visual intelligent reconstruction evaluation system for three-dimensional wave liquid level

The invention discloses a three-dimensional wave liquid level visual intelligent reconstruction evaluation system, and belongs to the field of ocean engineering monitoring. The system comprises an image acquisition module, an image stereoscopic vision processing module, an attention-enhanced reconstruction neural network module, a camera attitude evaluation module, a visualization and output module and a hydrodynamic parameter analysis module. An image is collected through a fixed baseline binocular camera system, after preprocessing, a neural network fused with a multi-scale attention mechanism is utilized to reconstruct a three-dimensional wave structure, coordinate system conversion is achieved in combination with self-supervised attitude evaluation, and finally hydrodynamic parameters such as significant wave height and a three-dimensional velocity field are extracted and visualized. The system does not need explicit calibration and large-scale data, has high automation, real-time performance and strong environmental adaptability, can be deployed on various platforms, and significantly improves the precision and engineering applicability of non-contact wave observation.
Owner:HARBIN INST OF TECH

Real-time visual processing method and system based on ESN-CV cooperative processing

The invention discloses a real-time visual processing method and system based on ESN-CV coprocessing, and relates to the technical field of visual processing, and the method comprises the steps: firstly, synchronously collecting a camera image and road surface humidity data of a physical sensor, and extracting an image reflection intensity distribution matrix as a visual feature; then humidity time sequence data and visual features are input into an echo state network, a dynamic weight coefficient matrix is generated through spatio-temporal feature fusion, the dynamic weight coefficient matrix is injected into a predefined convolutional layer of a target detection network, kernel weight parameters are adjusted in an element-by-element superposition mode, and the feature extraction capacity of a high-sensitivity area is enhanced. And after a detection result is output, the system triggers closed-loop feedback through a confidence coefficient deviation value: when the deviation exceeds a limit, the actual offset is calculated by combining a high-precision map, an error correction vector is generated, the state of the ESN reserve pool is updated by utilizing a Hadamard product, and the weight generation logic of the next frame is optimized in real time.
Owner:UNIV OF JINAN

Augmented reality display system for evaluation and modification of neurological conditions, including visual processing and perception conditions

In some embodiments, a display system comprising a head-mountable, augmented reality display is configured to perform a neurological analysis and to provide a perception aid based on an environmental trigger associated with the neurological condition. Performing the neurological analysis may include determining a reaction to a stimulus by receiving data from the one or more inwardly-directed sensors; and identifying a neurological condition associated with the reaction. In some embodiments, the perception aid may include a reminder, an alert, or virtual content that changes a property, e.g. a color, of a real object. The augmented reality display may be configured to display virtual content by outputting light with variable wavefront divergence, and to provide an accommodation-vergence mismatch of less than 0.5 diopters, including less than 0.25 diopters.
Owner:MAGIC LEAP INC

Scalable coding of video and associated features

The present disclosure relates to scalable encoding and decoding of pictures. In particular, a picture is processed by one or more network layers of a trained module to obtain base layer features. Then, enhancement layer features are obtained, e.g. by a trained network processing in sample domain. The base layer features are for use in computer vision processing. The base layer features together with enhancement layer features are for use in picture reconstruction, e.g. for human vision. The base layer features and the enhancement layer features are coded in a respective base layer bitstream and an enhancement layer bitstream. Accordingly, a scalable coding is provided which supports computer vision processing and / or picture reconstruction.
Owner:HUAWEI TECH CO LTD

Defect automatic positioning method based on BIM virtual image

The invention relates to the technical field of computer vision processing, in particular to an automatic defect positioning method based on a BIM virtual image. The method comprises the following steps: establishing a high-rise building model, and making a data set integrating an illumination condition and a full view angle; calculating camera parameters; performing cross-modal image retrieval: performing targeted optimization on the basis of a classical ResNet architecture, and constructing a backbone network structure suitable for a building image retrieval task; initializing position attitude estimation based on matching; and correcting the camera position posture. According to the method, a defect fixed frame based on a building BIM virtual image is provided, a building image data set BIM-Vision based on Revit is constructed, rich visual angles and illumination condition setting are achieved, accurate camera position postures and 3D labels between beam columns are provided, high-quality basic data support is provided for building visual research, and the method has the advantages of being high in practicability and high in practicability. And during inspection, accurate positioning of defect positions and component association can be completed only by shooting a field image, so that the field operation process is greatly simplified.
Owner:DALIAN NATIONALITIES UNIVERSITY