Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2349 results about "Visual technology" patented technology

Visual technology is the engineering discipline dealing with visual representation.

Real-time video analysis method based on deep learning

The invention relates to the technical field of computer vision, and discloses a real-time video analysis method based on deep learning. The method comprises the following steps: acquiring a real-time video stream through image acquisition equipment, and performing frame segmentation processing to generate a continuous video frame sequence; and extracting features of the video frame sequence by using a pre-trained convolutional neural network to obtain a multi-dimensional feature vector, inputting the multi-dimensional feature vector into the time sequence analysis model to calculate dynamic relevance, and outputting an inter-frame movement track and object behavior features. And constructing a scene understanding map containing a spatial position and a time evolution relationship according to the above-mentioned data, and carrying out abnormal event detection and generating event marking data based on the map. And performing semantic analysis on the event marking data, determining an abnormal event type and a confidence score, triggering a real-time alarm signal according to a result, and updating a historical event database. In the analysis process, the resource occupancy rate of the system is continuously monitored, the calculation precision is dynamically adjusted, a degradation processing mechanism is started when a preset threshold value is exceeded, and key area analysis is preferentially guaranteed.
Owner:HANGZHOU SIYUAN INFORMATION TECH CO LTD

Defect segmentation positioning method and system for inorganic mineral casting image

The invention relates to the technical field of computer vision, in particular to a defect segmentation positioning method and system for an inorganic mineral casting image, and the method comprises the following steps: calling an illumination image to analyze brightness, matching exposure parameters, splicing the image, analyzing a gradient, recognizing a defect, screening an effective region, calculating a gray variance, and constructing roughness weight recognition texture features. According to the method, the high-reflection area identification, the brightness gradient analysis, the pixel-level roughness weight and the structure tensor analysis are combined, the exposure interval can be dynamically adjusted when the casting image is processed, the defect type information is output in the direction, and the positioning information is generated by correcting the recognition position in combination with the actual coordinate of the target spot. The method has the advantages that the high-reflection area identification, the brightness gradient analysis, the pixel-level roughness weight and the structure tensor analysis are combined; the method has the advantages that the method is simple and easy to implement, detail loss of overexposure areas is reduced, the recognition precision of defect areas is improved, accurate area segmentation and classification processing are achieved, roughness weight calculation combining gray variance and pixel density is combined, the sensitivity to surface fine defects is enhanced, and the precision and reliability of defect positioning are improved.
Owner:SHANDONG CLAREMONT NEW MATERIAL TECH CO LTD

Image super-resolution method and system based on semantic perception token

The invention discloses an image super-resolution method and system based on semantic perception tokens, and relates to the technical field of computer vision, and the method comprises the steps: generating semantic confidence and grouping information through the aggregation of content perception tokens, and decoupling a basic residual error into a texture enhancement and degradation inhibition guidance graph; in combination with a static semantic constraint mask and a sparse matrix multiplication mechanism, progressive focusing of attention is realized; a diffusion time step embedding and cooperative modulator is introduced, semantic guidance information is dynamically injected into a multi-step denoising process, adaptive attention features and diffusion reconstruction features are fused, and finally a high-fidelity and high-resolution image is output. According to the method, content-adaptive high-resolution image reconstruction is realized through collaborative modulation of a sparse attention mechanism guided by semantic grouping and diffusion denoising guided by semantic decoupling.
Owner:HUAQIAO UNIVERSITY

Resource and task aware visual processing edge adaptive decision-making method

The invention belongs to the technical field of artificial intelligence and computer vision, particularly relates to a visual processing edge adaptive decision-making method for resource and task perception, and aims to solve the problem of scheduling mismatch caused by resource dynamic change and task demand diversity in visual task processing in an edge computing environment. The method comprises the following steps: collecting multi-dimensional resource state data of edge nodes in real time to form a resource state vector with high time resolution; analyzing the visual task request, and constructing a quantifiable task feature vector; and establishing a resource-task association mapping model based on a dynamic weight distribution mechanism. The method also supports cross-edge domain collaborative decision, and processes a pipeline dynamic reconstruction and security isolation mechanism. According to the technical scheme, the fluctuation of the resource utilization rate is reduced to 15% or below, the average task processing delay is reduced to 60%, the scheduling satisfaction degree is improved by 40% or above, and the self-adaptability and the service quality guarantee capability of the edge vision system are remarkably enhanced.
Owner:SHENZHEN IBD INTELLIGENT TECH CO LTD

Photovoltaic power station intelligent inspection system based on AI vision

The invention relates to the technical field of photovoltaic power station operation and maintenance, in particular to a photovoltaic power station intelligent inspection system based on AI vision, which comprises a video acquisition module, a geometric reference construction module, a tremor offset resolving module, a coordinate inverse correction module and a defect fine calibration module, the video acquisition module is used for acquiring a real-time video stream of an unmanned aerial vehicle polling photovoltaic array, and performing time-space synchronization calibration on the video stream to generate an original image sequence. According to the invention, the computer vision technology is utilized to calculate a current frame blanking point in a video picture as a tiny offset of a reference object relative to a reference position in real time, the displacement is deducted from a GPS coordinate, and a coordinate inverse correction module is utilized to inversely calculate a visual axis offset into a ground projection error. The shake amount and the shake direction of the camera at each moment can be accurately calculated, and then the GPS coordinates are corrected in turn, so that the geographic accuracy of defect positioning is improved, and the operation and maintenance personnel can accurately find a fault component.
Owner:ATLAS POWER TECHNOLOGY (XUZHOU) CO LTD

Non-contact physiological signal extraction method and system based on frequency self-adaption and illumination noise perception

The invention relates to the technical field of biomedical engineering and computer vision, in particular to a non-contact physiological signal extraction method and system based on frequency self-adaption and illumination noise perception.The method comprises the following steps of multi-mode video stream collection and spatio-temporal data preprocessing, illumination-noise perception mask generation and feature filtering, multi-mode video stream collection and spatio-temporal data preprocessing, illumination-noise perception mask generation and feature filtering, and non-contact physiological signal extraction. Frequency adaptive gating and frequency domain feature enhancement, depth time attention feature re-calibration, physiological signal regression and closed loop optimization; the method has the beneficial effects that a lightweight end-to-end deep learning network architecture is constructed by systematically fusing three core modules of illumination-noise perception mask, frequency adaptive gating and depth time attention, and the defects that a traditional physical model depends on artificial prior and is poor in anti-interference performance and high in reliability are overcome. And the one-sidedness caused by high calculation complexity and difficulty in distinguishing the signal and noise of the existing deep learning model is avoided, and the weak physiological signal can be recovered from the face video more accurately and robustly.
Owner:CENT SOUTH UNIV

General multi-modal target tracking method based on space-time propagation and modal cooperation

The invention discloses a universal multi-modal target tracking method based on space-time propagation and modal cooperation, and belongs to the technical field of computer vision. The method comprises the following steps: converting RGB and X modal images into a token form, and constructing initial features in combination with modal specific time tokens; the method comprises the following steps of: extracting a multi-level enhanced feature through a Transform encoder and a Mama collaborative prompt block; generating discriminative fusion features by using a gating fusion and context sensing module; a time-guided attention mechanism is adopted to strengthen search area features, and a result is output through a tracking prediction head; and transmitting the fusion time token as historical information to the next frame, and dynamically updating the template by combining a long-short time template updating strategy. According to the method, complementarity and space-time dependence between modes are effectively mined, tracking robustness and generalization ability in a complex scene are improved, and the method is suitable for various mode combination tasks such as RGB-D, RGB-T and RGB-E.
Owner:INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI

Medical image segmentation method and system based on deep learning

The invention relates to the technical field of medical image processing and computer vision, in particular to a medical image segmentation method and system based on deep learning, the method is based on a U-shaped encoder-decoder architecture, a DSAB module is introduced into an encoder, and context perception of a directional anatomical structure is enhanced through complementary directional space shift and CSA mechanism weighting; an MGCF module is designed in a decoder, and a parallel multi-scale convolution path and an AGCA mechanism are combined, so that multi-level features are efficiently fused to recover boundary details. Meanwhile, links of data preprocessing, Transform structure details, segmentation result post-processing and the like are supplemented, the model performance is improved through a mixed loss function and an optimization training strategy, and the method has remarkable advantages in segmentation precision and boundary definition and provides powerful support for clinical auxiliary diagnosis.
Owner:ANHUI POLYTECHNIC UNIV

Industrial equipment defect detection method based on multi-modal feature fusion and dynamic optimization

The invention discloses an industrial equipment defect detection method based on multi-modal feature fusion and dynamic optimization, and belongs to the technical field of computer vision, and the method comprises the steps: 1, multi-modal industrial data collection: deploying multiple sensors in a production line, collecting data in multiple periods, constructing a defect-free and multi-type defect sample library, and carrying out multi-modal industrial data collection; a time-space aligned multi-modal label is marked; step 2, data enhancement and defect synthesis; step 3, multi-modal hybrid model training: constructing a hybrid network, and performing pre-training and fine tuning by using a dynamic loss function; step 4, edge end dynamic optimization and deployment: edge end reasoning is realized through dynamic knowledge distillation, and model fine tuning is automatically triggered when false detection and missing detection are found; and step 5, intelligent labeling and result visualization: a front-end interface displays a detection result in real time. The problem that a current target detection framework is not high in small target recognition accuracy and low in efficiency is solved, and the reliability of industrial equipment defect detection is improved.
Owner:NANJING CHENGUANG GRP

Subway key component fault detection method and system based on AI visual large model

The invention relates to the technical field of artificial intelligence and computer vision, discloses a subway key component fault detection method and system based on an AI visual large model, and aims to solve the problems of low detection precision, weak generalization ability, insufficient multi-mode understanding, poor real-time performance and lack of state evolution modeling in the prior art. The method comprises the following steps: acquiring images of key components through a multi-view industrial camera array, and performing distortion correction, illumination normalization and noise suppression; a pre-trained visual large model is utilized to extract deep space features, and modeling is carried out on a continuous frame feature sequence through bidirectional LSTM to capture a time sequence change trend. By introducing the large-scale visual large model and spatio-temporal joint modeling, the identification capability of tiny defects is improved, the discrimination stability is enhanced, high-precision and low-delay automatic detection is realized, the false alarm rate and the omission ratio are remarkably reduced, and the detection efficiency and the system maintainability are improved.
Owner:GUANGDONG HUANENG ELECTROMECHANICAL GRP CO LTD

Visual obstacle avoidance method for flight of low-altitude inspection unmanned aerial vehicle

The invention discloses a flight vision obstacle avoidance method for a low-altitude inspection unmanned aerial vehicle, particularly relates to the technical field of image processing and computer vision, and is used for solving the problem of how to stably detect and reliably estimate the distance and contact time of such obstacles on the premise of only depending on airborne vision and being limited in calculation power and time delay. A space-time skeleton of a linear obstacle is constructed and continuously tracked in the strip-shaped candidates, and a grading distance and contact time are given in combination with an imaging time sequence difference and projection drift; uncertainties generated by the detection side and the self-motion are unified to the same coordinate and time mark, and risk density and safety envelope are formed through recursion; and then a path is updated according to a fixed priority of'deceleration-bypassing-still hovering 'by taking the risk as a hard constraint, so that high recall and stable distance evaluation and executable real-time avoidance can still be realized for a slender and swinging real obstacle under the conditions of only dependence on airborne vision and limited computing power and time delay. And the near loss event occurrence rate and the meaningless braking ratio are obviously reduced.
Owner:四川吉利学院

CNN and Transform fused self-supervised monocular depth estimation system and method

The invention relates to the technical field of computer vision, and particularly discloses a CNN and Transform fused self-supervised monocular depth estimation system and method. According to the system, local representation is enhanced through a multi-scale feature fusion mechanism, a cross-regional attention network is constructed to realize global context association, and fine reduction of a fine structure of a complex scene is realized. Firstly, based on DCB, multilayer expansion convolution is adopted to expand a receptive field, multi-scale pixel features are fused, and local details of a key area are enhanced; secondly, capturing fine-grained local information of the image by using parallel local convolution of ELGF, and acquiring long-distance dependency by using a self-attention mechanism, thereby realizing collaborative modeling of local information and global dependency, and remarkably enhancing feature expression ability; and finally, estimating a relative pose between adjacent images through a pre-trained ResNet18-based lightweight encoder, and constructing reprojection loss to optimize depth prediction. Experiments show that the model constructed by the method provided by the invention reaches 0.102 and 4.430 in AbsRel and RMSE indexes respectively, and is obviously superior to the existing mainstream method.
Owner:SHANGHAI DIANJI UNIV

Intelligent conference video frame dynamic coding method based on multi-mode semantic understanding

The invention relates to the technical field of computer vision, in particular to an intelligent conference video frame dynamic coding method based on multi-modal semantic understanding, which comprises the following steps: acquiring a video stream sequence and a synchronous audio stream in a conference scene in real time; performing semantic analysis and decoupling on the video stream sequence, and extracting key frames and subsequent frames; extracting a sparse motion field from a subsequent frame, and segmenting a video frame into candidate visual areas including a face, a mouth shape and a background; extracting audio semantic features, executing cross-modal semantic correlation analysis, calculating semantic correlation between the sparse motion field distribution features and the audio semantic features, and positioning a pronunciation area highly related to the voice content; and calculating a quantization offset value of each candidate visual area according to the semantic relevancy, applying the quantization offset values in different areas, and packaging the quantization offset values into a variable-code-rate video code stream. According to the invention, the multi-mode semantic understanding model is constructed to carry out deep semantic analysis on the video frame content so as to realize the dynamic coding of the conference video frame.
Owner:SHENZHEN JIKEYUAN ELECTRONIC TECH CO LTD

Pavement crack accurate segmentation method based on histogram interaction attention

The invention relates to the technical field of deep learning and computer vision, and discloses a histogram interactive attention-based pavement crack segmentation network processing method and system, so as to enhance the edge detail fidelity and improve the crack segmentation precision. The method comprises the steps of image preprocessing, up-sampling, down-sampling, feature fusion and image reconstruction processing. Wherein global feature modeling in the intensity sub-boxes and among the sub-boxes is realized by constructing a histogram interactive attention module (HIA); a double-branch detail enhancement feedforward module (DDEF) is introduced to enhance spatial detail and high-frequency edge information expression; meanwhile, a Fourier jump enhancement module (FFSM) is adopted to jointly refine jump connection features in a spatial domain and a frequency domain. Through the synergistic effect of the modules, the network can realize continuous recovery and structural consistency modeling of a crack boundary in a complex pavement environment, so that the accuracy and the stability of a segmentation result are remarkably improved.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Multi-modal dynamic compensation road disease intelligent detection and risk assessment system

The invention discloses a multi-modal dynamic compensation road disease intelligent detection and risk assessment system, and relates to the technical field of artificial intelligence and computer vision, and the system comprises an image collection module which is used for obtaining a road surface image in real time through a camera device, and transmitting the image to a preprocessing module; the preprocessing module is electrically connected with the image acquisition module and is used for carrying out graying, noise reduction, contrast enhancement and geometric correction operation on the image and outputting a standardized image; the feature extraction module is electrically connected with the preprocessing module. According to the road disease detection system provided by the invention, by integrating a plurality of modules, high efficiency and intelligence of road disease detection are realized, compared with traditional manual inspection, the system not only improves the detection efficiency, but also remarkably enhances the objectivity and accuracy of detection, and is particularly suitable for real-time monitoring requirements of a large-scale road network; the image acquisition quality is effectively improved, and the effectiveness of feature extraction can be ensured.
Owner:ZHEJIANG NORMAL UNIV

Spindle coating defect online detection method and system based on machine vision

The invention relates to the technical field of image processing, in particular to a picking ingot coating defect online detection method and system based on machine vision, and the method comprises the steps: obtaining a multi-view image, carrying out the feature extraction of an image data stream after preprocessing, and carrying out the matching or clustering with a preset defect dictionary; the method comprises the following steps: preliminarily identifying a potential defect area, triggering refined multi-angle image acquisition, applying a multi-task CNN model fused with a polarization perception convolution kernel, jointly optimizing pixel-level segmentation loss and boundary prediction loss, outputting pixel-level semantic segmentation and an accurate boundary of the defect area, and calculating a reconstruction error to carry out defect classification. And finally, the system also has the functions of defect root cause analysis and process flow adjustment, so that defect tracing and production optimization are realized. According to the system, advanced visual technology, deep learning and multi-modal information fusion are integrated, and online, accurate and automatic detection of picking ingot coating defects is realized.
Owner:NANJING YISEN IND TECHNOLOGY CO LTD

Indoor scene three-dimensional reconstruction method based on deep fusion and confidence modeling

The invention discloses an indoor scene three-dimensional reconstruction method based on deep fusion and confidence modeling, and belongs to the technical field of computer vision. According to the method, composite data frames such as a color image, a depth map, an IMU (Inertial Measurement Unit) and a camera attitude are comprehensively utilized to carry out regional three-dimensional reconstruction on an indoor scene: firstly, the scene is divided into a smooth region (such as a wall, a ground, a ceiling, a glass plane, a mirror surface, a blackboard and other planes) and a complex curved surface region; aiming at the smooth area, adopting geometric prior guide plane fitting provided by a visual large model, and combining sensor attitude information to quickly reconstruct a regular plane model; for a complex curved surface area, a multi-frame point cloud fusion strategy is adopted to accumulate different view angle information, and a deep residual error refining network is utilized to recover curved surface details, so that a high-precision curved surface model is obtained. According to the method, the three-dimensional structure of the indoor scene can be efficiently reconstructed, the global framework of the smooth area is reserved, and the details of the surface of a complex object are depicted in detail.
Owner:CHONGQING UNIV OF EDUCATION +1

Method for generating blind person navigation obstacle avoidance map

The invention relates to the technical field of computer vision, in particular to a method for generating a blind person navigation obstacle avoidance map. The method comprises the following steps: driving a blind guiding robot to a preset inspection point, and collecting a surrounding environment image in real time; generating a three-dimensional point cloud according to the environment image, and constructing an environment map; extracting feature points based on the surrounding environment image, performing feature matching, and outputting a feature matching result; evaluating an inter-frame motion state by using a feature matching result; determining pose features of the blind guiding robot according to the inter-frame motion state; semantic segmentation is carried out by using the pose features, and an obstacle area is determined; acquiring an obstacle image of the obstacle area, and fusing the obstacle image with the surrounding environment image to generate an obstacle map; and constructing a dynamic occupancy grid based on the obstacle map. According to the invention, real-time updating and intelligent obstacle avoidance functions of the blind navigation map are realized based on the computer vision technology, and the navigation safety and the path planning accuracy are improved.
Owner:SHANDONG SAIFEITE SAFETY ENG TECH DEV CO LTD

Multi-modal large model hidden danger identification method and system based on feature retrieval enhancement

The invention discloses a multi-modal large model hidden danger identification method and system based on feature retrieval enhancement, and relates to the technical field of computer vision. The method comprises the following steps: quickly scanning an input image through a lightweight target detection model, and positioning a hidden danger candidate area; performing feature retrieval in a pre-constructed standard hidden danger knowledge base based on the candidate region to obtain related standard hidden danger reference information; and the candidate region and the reference information are coded and fused and then input into a multi-modal large model, and a fine-grained recognition result containing hidden danger categories, position coordinates, confidence coefficients and professional text description is output. According to the method, a professional knowledge base retrieval mechanism and a multi-modal feature fusion technology are introduced, so that the problems of lack of professional knowledge and insufficient fine-grained discrimination ability in an existing hidden danger recognition method are effectively solved, and the recognition accuracy and the result specialty are improved while the detection efficiency is ensured.
Owner:浙江省应急管理科学研究院(浙江省安全生产技术检测检验中心浙江省危险化学品登记中心) +1

Visual perception method and system based on multi-modal thinking tree

The invention relates to the technical field of artificial intelligence and computer vision, and provides a visual perception method and system based on a multi-modal thinking tree in order to solve the problem that a traditional expansion strategy of purely increasing the parameter scale cannot effectively break through the semantic refinement bottleneck. The visual perception method based on the multi-modal thinking tree comprises the steps of obtaining a to-be-processed original image and a target anaphora text, and constructing the multi-modal thinking tree; defining a reasoning action set for driving node extension; executing a multi-mode Monte Carlo tree search process; iteratively generating a reasoning path until a preset search depth is reached or a termination condition is triggered; all effective leaf nodes generated in the searching process are obtained at the same time; and carrying out aggregation optimization on all effective leaf nodes by adopting a regional feature weighted voting mechanism, and screening out a candidate scheme with the highest comprehensive weight as a final visual perception positioning result. According to the method, the perception performance can be effectively improved on the basis of not changing the original parameter scale of the model, and high-precision visual positioning is realized.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Face forgery detection method and system based on multi-view collaborative fusion

The invention discloses a face forgery detection method and system based on multi-view collaborative fusion, and relates to the technical field of computer vision, and the method comprises the steps: obtaining to-be-detected video frame data; preprocessing the video frame data to be detected to obtain a standardized input image tensor; and inputting the standardized input image tensor into a pre-trained face counterfeiting detection model, and processing the standardized input image tensor by the face counterfeiting detection model to generate a face counterfeiting detection result. According to the method, the technical problem of poor detection performance of the model in cross-library and complex environments is solved, the detection robustness of the model in complex scenes such as fuzzy and compressed scenes is remarkably improved, and the detection precision is remarkably improved.
Owner:XUZHOU UNIV OF TECH +1

Physical attribute inversion and three-dimensional reconstruction method based on Gaussian splashing and micro rendering

The invention discloses a physical attribute inversion and three-dimensional reconstruction method based on Gaussian splashing and micro-rendering, and belongs to the technical field of computer vision, and the method comprises the steps: generating a sparse point cloud based on a multi-view image to initialize a Gaussian ball; determining a Gaussian pivot point based on a Gaussian ball, constructing a self-adaptive tetrahedral mesh in combination with gradient optimization, extracting an explicit surface from the self-adaptive tetrahedral mesh, and sampling to obtain geometric information; predicting the physical attribute of each sampling point through an implicit texture network, and providing ambient light through an HDR ambient light map; geometric information, physical attributes and ambient light are used as input, a differentiable PBR renderer is adopted to generate a physical rendering image, multi-term loss function reverse optimization is constructed, and collaborative optimization and inversion of three-dimensional geometry and surface physical attributes are realized. According to the method, the problems of texture dislocation, thin-wall structure disappearance and high rendering calculation cost under dynamic topology are effectively solved, and a digital model with fine geometry and a realistic material can be efficiently recovered from a multi-view image.
Owner:ZHEJIANG UNIV

Space-time consistency data generation method for visual target tracking

The invention relates to the technical field of computer vision, in particular to a space-time consistency data generation method for visual target tracking. The method comprises the following steps: firstly, training a path generator on a target tracking training set, learning a motion law of a target in a time sequence by using optical flow estimation and conditional variation coding technologies, and generating a target motion track conforming to physical constraints; and then, based on the generated target trajectory, introducing a space-time consistency attention mechanism to guide a text-video generation model, and under the condition of keeping basic model parameter freezing, constraining the position, scale and continuity of a target in a generation frame through an attention network, thereby synthesizing a video frame sequence with real motion features. According to the method, target tracking video data with real motion characteristics and high time sequence consistency is generated, and the robustness of the model to complex motion, illumination change and shielding conditions can be improved in different scenes.
Owner:QINGDAO UNIV OF TECH

Causal perception sentiment analysis method and system based on thinking chain reasoning

The invention discloses a causal perception sentiment analysis method and system based on thinking chain reasoning, and belongs to the technical field of computer vision. The method comprises the following steps: acquiring and reading a multi-modal sentiment analysis data set; extracting video features from the video data in the multi-modal sentiment analysis data set, including voice, text and visual modal features; using the training set and the test set to train and verify the causal perception emotion polarity alignment model; inputting the test set into the trained causal perception emotion polarity alignment model to obtain an emotion state prediction result; video features are input into a causal perception emotion polarity alignment model, and emotion clues are extracted through thinking chain prompt and a self-supervision verification mechanism; then performing causal intervention and anti-factual reasoning on each modal feature by using an emotion clue to obtain a causal-related single-modal feature; and finally, obtaining joint feature representation from the causal-related single-mode features through cross-mode interaction by using a multi-mode representation learning method, and predicting an emotional state.
Owner:NANJING UNIV OF POSTS & TELECOMM

Aviation part crack detection and repair method

The invention relates to the technical field of industrial vision, in particular to an aviation part crack detection and repair method which comprises the following steps: acquiring an aviation part optical image at a reference time point and an aviation part optical image at a to-be-detected time point; according to the method, logarithmic polar coordinate transformation is carried out on different time point images, rotation, scaling and translation parameters are extracted, an image registration relation is established, the structural consistency of time sequence images in a local area is enhanced, a structural tensor is constructed for each pixel neighborhood, the change characteristics of the dual-time-phase tensor are compared, a structural change saliency map is generated, and the structural change saliency map is obtained. Sensitive capture of a tiny deformation area is achieved, the responsiveness to an initial crack is improved, then a crack propagation interval is further refined into a main crack path in a self-adaptive threshold segmentation and skeleton extraction mode, the tip acutance of the crack is calculated in combination with tip contour information of a geometric boundary of the crack, and the initial crack is obtained. And quantitative support is provided for the crack danger degree.
Owner:SHENYANG AEROSPACE UNIVERSITY

Drainage pipe network defect detection method based on data optimization and multi-model fusion

The invention relates to the technical field of drainage pipe network defect detection and computer vision, in particular to a drainage pipe network defect detection method based on data optimization and multi-model fusion, and the method comprises the steps: collecting a video in a pipeline through a detection robot, and carrying out the frame extraction to obtain an original image; carrying out the quality optimization of the image through employing a lens fouling restoration algorithm, an image defogging algorithm and an illumination enhancement algorithm; positioning a defect area by using the target detection model, and performing pixel-level segmentation through the U-Net semantic segmentation model to obtain a defect contour; visual features and morphological features are extracted, feature fusion is carried out in combination with an attention mechanism, and finally a classifier is input to realize accurate identification of defect types. Targeted solutions are provided for multiple common imaging defects in the drainage pipe network, the method can adapt to complex and changeable actual detection environments, good generalization ability and practical value are achieved, and reliable technical support can be provided for intelligent operation and maintenance of the urban drainage pipe network.
Owner:NANJING TECH UNIV

Image target detection system and method based on deep learning

The invention relates to the technical field of computer vision, in particular to an image target detection system and method based on deep learning, and the system comprises a dynamic feature alignment unit, a motion blur compensation unit and a feature fusion control unit. A dynamic feature alignment unit generates spatial deformation parameters through a deformable convolutional layer and an offset prediction sub-network, resamples a shallow high-resolution feature map, and realizes deep and shallow feature space alignment, and a motion blur compensation unit generates a motion vector based on brightness gradient field difference, constructs a mask and weights a suppression blur region, so as to realize deep and shallow feature space alignment. The feature fusion control unit analyzes local entropy and target size distribution, dynamically distributes feature weights and feeds back and optimizes offset parameters, a closed-loop learning loop is formed by the method, and the problems of inaccurate feature alignment, fuzzy interference and poor scene adaptation are solved.
Owner:ZHEJIANG KANGXU TECH CO LTD

Jade defect intelligent detection method and system based on machine vision and deep learning

The invention relates to the technical field of computer vision, and discloses a jade defect intelligent detection method and system based on machine vision and deep learning, and the method comprises the following steps: S1, based on a high-resolution industrial camera and a laser three-dimensional scanner, adopting a multi-mode synchronous collection strategy, and rotating a jade sample through a precise motion control system, a jade surface high-resolution two-dimensional color image and high-precision three-dimensional point cloud data are respectively obtained, and a jade multi-mode original data set is generated. A high-resolution two-dimensional color image and high-precision three-dimensional point cloud data are integrated through a multi-modal synchronous acquisition strategy, and multi-dimensional feature expression under unified coordinates is constructed, so that the limitation of a single data source is effectively overcome; an image registration algorithm and a feature pyramid network are combined with a point cloud network to perform multi-modal feature fusion, complementarity of color texture and geometric morphology information is enhanced, and image quality is optimized based on adaptive histogram equalization and non-local mean filtering.
Owner:SHENZHEN BAIHAI DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Large model navigation method guided by historical topological graph based on manifold perception

The invention discloses a manifold perception-based large model navigation method guided by a historical topological graph, and relates to a computer vision technology. The method aims at solving the challenges that in the navigation process, long-distance reasoning experience is insufficient, instruction fragments and dynamic visual observation are difficult to align, and large model reasoning is prone to illusion interference. Firstly, a large model based on an encoder-decoder structure is used for supplementing historical information coding for a visual observation sequence, and therefore global topological information guidance is provided for long-distance reasoning. And secondly, in order to effectively solve the problem that large model reasoning is subjected to illusion interference, significant space-time differences in a visual observation sequence are mined by using a multi-curvature manifold, so that the large model can accurately describe the current environment and make a decision according to a visual reference object in thinking. Besides, in order to strengthen the perception capability of the large model to the space structure and establish a graph self-attention mechanism, the node distance embedded in the constructed historical topological graph is combined with the visual similarity so as to model the space relationship between the nodes.
Owner:WENZHOU TAIYI INTELLIGENT TECHNOLOGY CO LTD +2

Railway scene target detection and behavior identification method based on space-time double-flow characteristics

The invention belongs to the technical field of railway engineering safety monitoring and computer vision, and discloses a railway scene target detection and behavior recognition method based on space-time double-flow features, and the method comprises the steps: obtaining and preprocessing video data into an image frame sequence; a sequence is input to an improved target detection network (Mamba-Yov11), the network integrates local spatial features and global time sequence context information by setting a convolutional neural network path and a state space model path in parallel, and adopts an improved C3k2UIB module to realize dynamic path selection and improve parameter efficiency; the network training adopts a Focaler-IoU loss function to solve the problem of unbalanced training of small samples and difficult samples; and for complex behaviors, the detected target area is sent to the MILA-SF behavior recognition network, and the space-time behaviors are efficiently recognized through fast and slow dual-path design. According to the method, the detection precision of targets such as wearable equipment and operation tools in a railway scene can be remarkably improved.
Owner:EAST CHINA JIAOTONG UNIVERSITY