Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41108 results about "Pattern recognition" patented technology

Pattern recognition is the automated recognition of patterns and regularities in data. Pattern recognition is closely related to artificial intelligence and machine learning, together with applications such as data mining and knowledge discovery in databases (KDD), and is often used interchangeably with these terms. However, these are distinguished: machine learning is one approach to pattern recognition, while other approaches include hand-crafted (not learned) rules or heuristics; and pattern recognition is one approach to artificial intelligence, while other approaches include symbolic artificial intelligence. A modern definition of pattern recognition is...

System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance

A system and method are disclosed for generating hyperspectral images from multi-modal sensor data including RGB, LiDAR, thermal, and near-infrared inputs. Training data includes hyperspectral images and corresponding multi-modal measurements. Spectral band grouping is performed based on correlation coefficients. A multi-modal decomposition network with cross-modal attention mechanisms generate reconstructed hyperspectral images by fusing complementary sensor information. A fine-tuning network creates reconstructed RGB images. A comprehensive quality assurance system analyzes spectral consistency, cross-modal coherence, and fusion artifacts to generate quality metrics. Missing data compensation strategies handle corrupted sensor inputs using information from other modalities. The system includes temporal integration for video sequences and multi-resolution processing for different sensor resolutions. Quality metrics guide network weight adjustments to improve reconstruction accuracy while maintaining robustness to sensor failures and environmental variations.
Owner:ATOMBEAM TECH INC

Mapping locating systems and methods

Locator systems are disclosed for locating buried utilities such as conduits, cables, pipes. or wires. The locator system may include image-capturing and position and orientation measuring devices such as magnetic compass, accelerometers, gyros, GPS, and DGPS. The system may associate images captured during the locate process with buried utility position data and the other sensor data. These data and images may be further associated with terrain images from satellites or aerial photography to provide a highly precision map and database of buried utility locations and visual images of the burial site. Information so obtained may be transferred to a hand-held personal communication device, such as a smart phone to show the location of buried utilities in combination with photo-images and / or terrain maps.
Owner:SEESCAN INC

System and method for efficient scene continuity in visual and multimedia using generative artificial intelligence

ActiveUS20250378537A1Image enhancementPattern recognitionGenerative process
A system and method for generating multimedia artifacts with managed scene continuity in visual and multimedia using an AI-based and scene continuity aware media generation platform. The system receives a user or AI agent specification or simulation result(s), selects or trains generative models based on the specification, preprocesses relevant data, and generates scene narrative or frame-specific, sequence specific or broader continuity aware content using the selected or trained model(s). The generated content may be further enhanced using frame interpolation and view synthesis techniques to create smooth transitions or novel viewpoints or to aid in more efficient transmission or viewing or persistence of resultant content. The system enables efficient and customizable generation of high-quality scene continuity aware content for various applications in visual and multimedia production using neuro-symbolic and simulation enhanced compression, representation and generation processes.
Owner:QOMPLX INC

Optimization method and device for sparse view angle three-dimensional Gaussian splashing

The invention relates to an optimization method and device for sparse view angle three-dimensional Gaussian splash, and belongs to the technical field of three-dimensional reconstruction in computer vision, and the method comprises the steps: collecting a sparse view angle image; a multi-view stereoscopic vision model based on deep learning generates a geometrically consistent depth map for the sparse view image, converts the depth map into point clouds and fuses the point clouds to obtain dense point clouds; sampling dense point clouds by adopting voxel-guided farthest point sampling to obtain initialized point clouds, and constructing a three-dimensional Gaussian field; rendering the three-dimensional Gaussian field through an enhanced geometric renderer to obtain a rendering depth and a rendering normal; constructing a multi-level geometric regularization loss function, and optimizing the three-dimensional Gaussian field; and performing optimization adjustment on the three-dimensional Gaussian field based on a shape-scale constraint criterion and a two-stage adaptive opacity constraint strategy to obtain an optimized three-dimensional Gaussian field. According to the method, the problems of initialization failure, insufficient geometric supervision and element out-of-control of 3D Gaussian splashing under the sparse view angle are solved.
Owner:CHINESE ACAD OF SURVEYING & MAPPING

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Knowledge graph link prediction method

The present invention relates to the technical field of knowledge graph completion tasks, and particularly relates to a knowledge graph link prediction method. The method comprises: using a precoding model to obtain an embedding layer vector, and constructing a corresponding masked triple; adding a corresponding position code to each element in the masked triple, so as to obtain a corresponding input sequence, inputting the input sequence into a trained main masking model, and outputting an entity classification probability; and on the basis of the entity classification probability, predicting potential candidate entities. The method further comprises: concatenating semantic information corresponding to the embedding layer vector and structural information obtained by an embedding model, so as to obtain fused head entity and relation representations, and constructing a corresponding fused masked triple; and adding a corresponding position code to each element in the fused masked triple, so as to obtain a corresponding fused input sequence. The present invention uses a precoding method, thereby effectively reducing the training burden on a model, and improving the inference speed of a model; and a fusion module is used before inputs are fed into a main masked model, thereby ensuring the integrity of textual description information and improving prediction accuracy.
Owner:JIANGNAN UNIV

Injection product defect detection method based on machine vision

The invention relates to an injection molding product defect detection method based on machine vision, which comprises the following steps: collecting material information of a to-be-detected injection molding product in real time, and dynamically matching and adjusting light source parameters according to spectral reflection characteristics of materials to ensure image collection quality; secondly, the collected images are preprocessed, edge features and texture features are extracted, a three-dimensional model is constructed through multi-view image splicing, and three-dimensional defect features are extracted; thirdly, the multi-dimensional features are input into a deep learning model, the defect probability is calculated through feature fusion and forward propagation, and whether the product has defects or not is judged; if the defect exists, further identifying the defect category, and calculating the number and size of the defect; and generating a standardized detection report based on the defect information. According to the method, the image adaptability of products made of different materials is improved through dynamic light source adjustment, the two-dimensional and three-dimensional features are fused, the defect recognition accuracy is improved, and full-process automation from qualitative judgment to quantitative analysis of the defects is achieved.
Owner:SICHUAN YUJIA MOLDS&PLASTICS CO LTD

Road crack detection method and system based on fused image

The invention relates to the technical field of road crack detection, in particular to a road crack detection method and system based on a fused image. The method comprises the following steps: acquiring road multi-source monitoring data including a visible light image, infrared thermal imaging data and laser radar point cloud data, and performing multi-modal image fusion and road three-dimensional point cloud reconstruction to generate a fused road image and road three-dimensional modeling data; performing crack curvature analysis based on the fused road image to generate crack curvature data; performing reflection crack contour recognition and positioning on the fused road image through the crack curvature data to generate reflection crack initial positioning data; obtaining road base material data; and performing reflection crack stress field reconstruction on the road area according to the reflection crack initial positioning data to obtain a reflection crack stress field. According to the invention, through multi-modal fusion, curvature identification, stress field modeling and crack channel analysis, the accuracy and strain of road reflection crack detection are improved.
Owner:BINHAI BAY BRANCH OF DONGGUAN CITY URBAN MANAGEMENT & COMPREHENSIVE LAW ENFORCEMENT BUREAU

Microblog emotion analysis method based on standard dictionaries and semantic rules

The invention discloses a microblog emotion analysis method based on standard dictionaries and semantic rules. The microblog emotion analysis method comprises the following steps: collecting microblog data and manually labeling and marking the emotion value of each microblog; proposing corresponding standard micrblog emotion dictionaries, and establishing an emotion dictionary database; based on the standard emotion dictionaries, adding the semantic rules for assistance, and performing parameter adjustment and optimization on parameters of the semantic rules; based on a real dataset experiment, acquiring the final classification accuracy and precision. The technical scheme provided by the invention is adopted to well analyze the emotion tendency of each microblog user by introducing the standard emotion dictionaries, microblog expression dictionaries and the semantic rules, therefore, higher classification accuracy and precision are achieved.
Owner:BEIJING UNIV OF TECH

Image segmentation and dynamic target identification method based on artificial intelligence

The invention relates to the technical field of artificial intelligence, in particular to an artificial intelligence-based image segmentation and dynamic target recognition method, which comprises the following steps of: accurately positioning a candidate region through multi-modal space-time fusion and dynamic confidence coefficient screening; strengthening spatial-temporal feature expression in a layering manner through a multi-level feature decoupler, and generating a multi-dimensional feature enhanced spatial-temporal candidate region; through a deformable segmentation network, a deformation convolution kernel and edge motion matching loss are combined, joint optimization of a geometric boundary and motion continuity is realized, and the segmentation robustness of a flexible target is improved; through optical flow back propagation dynamic correction and confidence coefficient propagation, high-precision segmentation masks with consistent time and space are output; and through a target trajectory re-identification and completion mechanism driven by a graph attention network, and in combination with optical flow deformation prediction, stable tracking in a shielding scene is realized.
Owner:CHANGSHA INSTITUTE OF TECHNOLOGY

Weak supervision target detection method guided by cross-modal pseudo tag

The invention relates to the technical field of computer vision and multi-modal learning, in particular to a weak supervision target detection method guided by cross-modal pseudo labels. According to the method, a labeled source domain data set is constructed to train an image classification teacher model, and a teacher-student network structure is constructed; clustering the regional features of the target domain image, allocating pseudo tags to each cluster by optimizing the allocation cost between the source domain category and the target domain cluster, and constructing a pseudo tag pool; and training a student model on the pseudo label pool for region feature detection of the target domain image. According to the method, a cross-modal attention mechanism is introduced, so that more accurate semantic alignment between a source category label and a target domain feature is realized; the stability of label distribution is improved by a structure keeping regular term; the generalization ability of the model is further enhanced by multiple rounds of pseudo-label confidence learning. The method can be widely applied to tasks such as target detection, cross-domain transfer learning and open world recognition, and efficient and accurate weak supervision target detection is realized.
Owner:DATA SPACE RES INST

Multi-target detection and tracking method

The invention discloses a multi-target detection and tracking method, and relates to the technical field of computer vision and intelligent monitoring. The method comprises the following steps: acquiring multi-source video data of an unmanned aerial vehicle and a middle-high point fixed camera, and after scene adaptation preprocessing, outputting a target detection frame by using a multi-scale detection model fused with scene context; block enhanced appearance features and geometrical relationship features of the target are extracted to construct a dynamic feature library, and an initial track is generated based on a multi-stage adaptive association mechanism; through a child-mother type multi-machine collaborative optimization track, linkage control is triggered in combination with abnormal behavior analysis, and close-range evidence obtaining of the unmanned aerial vehicle and linkage of fixed equipment recording are controlled. According to the method, the multi-source data fusion capability, the multi-scale target detection precision and the trajectory association robustness in a complex scene are improved, and intelligent management and control requirements in the fields of traffic, forestry and the like can be efficiently supported.
Owner:CHINA TOWER CO LTD XIANGTAN BRANCH +1

Cutting workpiece defect detection method and system based on image feature feedback

The invention discloses a cut workpiece defect detection method and system based on image feature feedback, and relates to the technical field of image processing.The method comprises the steps that cut workpiece technological characteristics are obtained, a preset defect type library is constructed, a hardware system is built, and parameters are initialized; synchronously acquiring a multi-view original image, and storing and associating annotation information; de-noising the original image, enhancing the contrast, and extracting a region of interest ROI; extracting texture, shape, edge and gray features from the ROI, and screening through a Relief-F algorithm to obtain an optimal feature subset; inputting into an SVM (Support Vector Machine) model for reasoning, and screening to obtain an effective defect detection result; and calculating an evaluation index and generating a feedback signal, and performing iterative optimization after adjusting parameters. The system comprises an acquisition module, a master control module, a data processing module and a display module. Through the precise design and closed-loop feedback of the whole process, the precision, efficiency and long-term adaptability of defect detection of the complex cutting workpiece are improved, and the industrial quality management and control requirements are met.
Owner:苏州艾克夫电子有限公司

Three-dimensional reconstructions based on gaussian primitives

In implementation of techniques for three-dimensional reconstructions based on Gaussian primitives, a computing device implements a reconstruction system to receive a first digital image depicting an object from a first angle and a second digital image depicting the object from a second angle. The reconstruction system segments the first digital image and the second digital image into patches. The reconstruction system then generates, using a machine learning model, three-dimensional Gaussian primitives that predict parameters of points of the object in a three-dimensional space that correspond on a per-pixel basis to pixels of the patches. The reconstruction system then forms a three-dimensional reconstruction of the object for display in a user interface by merging the three-dimensional Gaussian primitives.
Owner:ADOBE INC

Dynamic Latent Space Adaptation Based on Spatiotemporal Kernal Context for Multiscale Rendering

A system for dynamic latent space adaptation using spatiotemporal kernel context for multiscale rendering with hierarchical and Lorentzian autoencoders. The Spatiotemporal Kernel Estimator (SKE) analyzes media through motion field, temporal recurrence, frequency band, and scene semantics analyzers to generate adaptive kernel parameters encoding content-specific importance distributions. The system dynamically adapts latent manifold geometry by modifying metric tensor properties according to kernel context, enabling content-aware compression that allocates representational capacity based on visual significance. A multiscale cache implements kernel-adaptive retention policies prioritizing important regions. An adaptive renderer provides intelligent level-of-detail selection based on zoom level and kernel-estimated importance, optimizing processing allocation. The self-optimizing architecture continuously refines kernel context and geometric adaptation based on user interaction and performance feedback, achieving superior compression ratios and perceptual quality. Applications include bandwidth-efficient video streaming, virtual reality, scientific visualization, and cognitive video analytics requiring intelligent context-aware visual processing.
Owner:ATOMBEAM TECH INC

Three-dimensional dynamic scene reconstruction method and apparatus, and storage medium

The present disclosure relates to the field of computer vision and discloses a three-dimensional dynamic scene reconstruction method and apparatus, and a storage medium. The three-dimensional dynamic scene reconstruction method comprises: acquiring synchronized videos of a plurality of viewpoints of a dynamic scene; computing matching points between video images of different viewpoints, and estimating intrinsic and extrinsic parameters of each camera; obtaining a Gaussian splatting point set {p0} on the basis of a sparse point cloud constructed according to the depth of each matching point; for the first image frame of each video, using {p0} to perform static training thereon, to obtain a Gaussian splatting point set {p}; for the remaining image frames, dividing {p} into a static point set {S} and a dynamic point set {D}, performing dynamic training on {D}, and constructing a dynamic Gaussian splatting point set {P} from {p}, {S}, and the final {D}; and, in view of the intrinsic and extrinsic parameters of each camera, rendering {P} using a Gaussian splatting rendering pipeline, to obtain rendered images at different moments from new viewpoints.
Owner:TSINGHUA UNIVERSITY

Unmanned aerial vehicle image small target detection method based on dynamic filtering and adaptive sparse Transform

The invention discloses an unmanned aerial vehicle image small target detection method based on dynamic filtering and an adaptive sparse Transform. According to the method, an end-to-end target detection framework is adopted, a dynamic filtering module is introduced into a backbone network, global feature interaction is achieved through data-dependent frequency domain operation, and linear calculation complexity is maintained. For feature interaction in a scale, an adaptive sparse Transform module is introduced to enhance the capability of focusing key information on high semantic hierarchy features of a model, and noise interference and feature redundancy are effectively suppressed at the same time. Through the combination of dynamic filtering and adaptive sparse Transform, the model can extract image foreground information more effectively on the premise of not significantly increasing the calculation burden, and the problem that a traditional target detection model is susceptible to complex background interference is significantly relieved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

PCB welding spot defect detection system and method based on image recognition

The invention relates to the technical field of PCB welding spot defect detection, and discloses a PCB welding spot defect detection system and method based on image recognition, and the system comprises an image preprocessing module which is used for obtaining an original image flow and dividing an interested detection area; the feature fusion module is used for extracting multi-modal features to form a fusion set; the defect judgment module is used for establishing a mapping index and obtaining a judgment result; the parameter calibration module is used for verifying the detection parameters and adjusting the mapping index; and the report output module is used for generating a defect detection report. The method comprises the steps of image preprocessing, feature fusion, defect discrimination, parameter calibration, report generation and the like. According to the system and the method, the PCB welding spot defects can be efficiently and accurately detected, the detection precision and stability are improved, a structured report is generated, and an effective solution is provided for PCB quality detection.
Owner:GUILIN SHIYU ELECTRONIC TECH CO LTD

Membrane structure weld defect detection method based on image recognition

The invention relates to the technical field of material nondestructive testing, and discloses a membrane structure welding seam defect detection method based on image recognition, which is used for solving the problems of low defect segmentation accuracy and incapability of effectively recognizing internal defects caused by continuous image gray change of lap welding seams of unequal-thickness flexible materials in a traditional method. The method comprises the following steps: firstly, collecting an original image of the lap weld of the unequal-thickness flexible material, carrying out gray conversion, analyzing thickness gradient distribution, adjusting a gray value, carrying out region segmentation to lock a weld range, extracting potential defect edge features to form a candidate region, and carrying out classified verification to confirm internal defects. Aiming at the problem of low defect segmentation accuracy caused by continuous change of image gray in the prior art, the method improves the defect identification precision through segmented mapping and boundary tracking logic, and is suitable for membrane structure engineering quality control.
Owner:HUNAN ZHONGHUAN HI TECH MATERIALS CO LTD

System and method for reconstructing 3D scene data from 2D image data

A method and apparatus for reconstructing a three-dimensional (3D) scene from a two-dimensional (2D) input image of the scene using a fully-differentiable transformer-based encoder-decode. A 2D input image encoded into a set of image features using a pre-trained vision transformer model, wherein the vision transformer model is pre-trained with multi-view RGB image supervision and point cloud supervision. The set of image features is projected onto a 3D triplane representation using a transformer decoder to obtain output triplane tokens. A triplane representation is created from the tokens and queried. 3D point features of color and density for volumetric rendering re predicted using a multi-layer perceptron. The geometry of the generated 3D asset is represented with a surface mesh including vertices and triangular faces. A texture map by is created with a multichannel image in UV space. Multiple views of the 3D scene are simultaneously generated based on the surface mesh.
Owner:FUTUREVERSE IP LTD

Systems and methods for updating large language models

Techniques for updating a large language model (LLM) to correct generation of undesired responses, such as incorrect outputs, toxic outputs, etc. are described. Typical methods of retraining and fine-tuning are inefficient and computationally expensive for LLMs. Some embodiments of the present disclosure involve identifying a salient layer of the LLM that is responsible for the undesired response and editing only the salient layer. This layer is identified by computing a saliency value for the layer using a mean of gradient values for the layer, and the layer with the greatest saliency value is selected for editing. For editing, a small network is used to update the weights of the selected layer. The LLM is updated to include the edited layer, and the updated LLM is used for future processing.
Owner:AMAZON TECH INC

An AR home experience method in a large scene

The invention discloses an AR home experience method in a large scene. On the basis of combination of a natural feature identification-based three-dimensional registration method and a binocular tracking positioning and local mapping method, the camera attitude is estimated by using feature points of a real-time scene and corresponding three-dimensional points thereof under the binocular trackingpositioning and local map construction technology. According to the mode, on-site environment features shot in real time are used as recognition tracking objects, a virtual home model can still be normally positioned and tracked under the condition that no identification graph exists, the problems that an existing AR home experience application is small in use range and poor in stability are solved, and therefore the AR home experience of virtual and real fusion can be met in a wider range and more truly.
Owner:MAANSHAN JUMEI YOUPIN DECORATION ENGINEERING CO LTD

Traditional picture repairing method fusing low-resolution prior and efficient visual selection

The invention belongs to the technical field of digital restoration of computer vision and cultural heritage, and particularly relates to a traditional picture restoration method fusing low-resolution prior and efficient visual selection, which comprises the following steps: constructing a multi-source image data set, taking images in the multi-source image data set as high-resolution images, preprocessing the high-resolution images to obtain low-resolution images, and carrying out high-resolution priori and high-efficiency visual selection on the low-resolution images. The high-resolution image and the low-resolution image are respectively masked to generate simulated damage mask images, and the simulated damage mask images comprise a regular damage mask image and an irregular damage mask image; taking the multi-source image data set and the preprocessed multi-source image data set as training data, and training a multi-source image model; the dual-stage repair network comprises a coarse repair network and a fine repair network; according to the method, the problems of structural semantic loss, high priori information dependency and insufficient global and local coordination when an existing image restoration method is used for processing a complex scene and a large-range missing region are solved.
Owner:NORTHWEST UNIV

Respiratory system risk prediction method and system based on graph neural network

The invention relates to the technical field of respiratory system risk prediction, and provides a respiratory system risk prediction method and system based on a graph neural network, and the method comprises the steps: collecting the multi-modal medical data of a patient, and constructing a multilayer heterogeneous graph based on the multi-modal medical data; constructing a weighted adjacency matrix and a node feature vector through the multi-layer heterogeneous graph; matrix product operation and convolution operation are carried out based on the weighted adjacent matrix and the node feature vector, splicing combination with historical moment state information is carried out, graph state representation is obtained, weighted aggregation of time dimensions is carried out, and time sequence attention features are obtained; performing coding processing based on the clinical examination data to obtain multi-modal fusion features; and inputting the multi-modal fusion features into a risk classifier for classification calculation to obtain a respiratory system risk level prediction result, generating a risk assessment report, and outputting respiratory risk early warning information. The accuracy and clinical practicability of respiratory system risk prediction are improved.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Visual inertial positioning method based on dynamic target detection and semantic information constraint

The invention discloses a visual inertial positioning method based on dynamic target detection and semantic information constraint, and belongs to the field of motion estimation and dynamic environment processing. According to the method, a dynamic target detection mechanism is introduced, the dynamic target is effectively detected based on the target detection network, inertial navigation information and geometric constraints, the dynamic target is effectively recognized in the image processing process, the corresponding dynamic feature points are screened out, the mismatching rate of the dynamic features is remarkably reduced, and high-quality observation input is provided for back-end optimization. In the back-end sliding window optimization stage, a semantic information consistency constraint method is constructed, and the estimation stability of the system in a weak texture area or a repeated texture area is enhanced by utilizing the consistency of feature points in the same semantic area on a geometric structure. Visual inertia pose estimation is realized based on dynamic target detection and semantic information constraint, and high-robustness and high-precision pose estimation can still be realized in a complex environment with dynamic interference of pedestrians, vehicles and the like and severe scene change.
Owner:BEIJING INST OF TECH

Fabric defect intelligent detection method and system based on AI visual identification

The invention relates to the technical field of fabric detection, and discloses a fabric defect intelligent detection method and system based on AI visual identification. According to the method, motion blur is quantized through motion state data, optical blur caused by fabric motion is eliminated through deconvolution solution, so that motion interference in the fabric transmission process is processed in a targeted mode, self-adaptive balance of the deblurring capacity and the feature retention capacity is achieved, and then based on the optical interference principle, the deblurring capacity and the feature retention capacity are improved. Through a dynamic calibration system combining hardware-level real-time compensation and multi-dimensional optical parameter calibration, dynamic optical parameter calibration of primary correction data is realized, then fabric defect characterization data is extracted to accurately obtain defect features, and finally, a detection-production line control closed loop is constructed through a quality quantitative index and a comprehensive risk value, so that fabric defect detection is realized. The fabric defect detection precision can be improved, so that the problem of high defect missing detection and false detection rate caused by optical data distortion due to movement and environment interference in a traditional method is effectively solved.
Owner:HANGZHOU HANGSIYUE TEXTILE TECH CO LTD

Two-way visual saliency detection method and device combining difference guidance and texture enhancement

The invention discloses a two-way visual saliency detection method and device combining difference guidance and texture enhancement, and the method comprises the following steps: 1, obtaining an image to be subjected to saliency detection, and carrying out the marking and preprocessing of a data set; 2, constructing a visual saliency detection model which comprises a dual-path encoder (a saliency detection path and an image reconstruction path), an adaptive interaction network, a decoder network and an output network; a significance detection path in the dual-path encoder network uses a pre-trained ConvNeXt encoder, and an image reconstruction path uses VQ-VAE as a backbone network; the adaptive interactive network comprises a multi-scale convolution module and a gating fusion module; the decoder network comprises a mutual conversion attention module and a double-gating fusion module; the output network comprises a multi-level feature fusion module; 3, training the saliency detection model to obtain a trained saliency detection model; and 4, carrying out saliency detection on the image data by adopting the trained saliency detection model.
Owner:SICHUAN UNIV

Video content semantic understanding and text description generation method based on deep learning

The invention discloses a video content semantic understanding and text description generation method based on deep learning, and relates to the technical field of multimedia information processing.The method comprises the steps that the semantic similarity of a text and a video frame is calculated through a CLIP model, related key frames are selected, and features are aggregated; respectively extracting audio, visual and semantic features; aligning different modal features by using self-attention, unifying dimensions of the LSTM, and then splicing and fusing; attention weights are calculated at a video level, a frame level and a channel level, and key information expression is enhanced; swin Transform encodes fusion features, and LSTM (Long Short Term Memory) decodes step by step to generate natural language description; and a text-video index database is constructed, and rapid retrieval is realized based on semantic similarity. According to the method, the mapping relation between the video features and the natural language is learned end to end through the deep learning model, dependence on a fixed template can be eliminated, and semantic description with various sentence patterns and coherent logic is generated.
Owner:CHINA UNIV OF MINING & TECH YINCHUAN COLLEGE

Anti-unmanned aerial vehicle intelligent identification and tracking system based on multi-source data fusion

The invention provides an anti-unmanned aerial vehicle intelligent identification and tracking system based on multi-source data fusion, and relates to the technical field of anti-unmanned aerial vehicle detection, and the system comprises a multi-source data preprocessing module which is used for outputting preprocessed multi-source data; the target detection module is used for carrying out unmanned aerial vehicle target detection on visual data in the preprocessed multi-source data and outputting a detection result containing a bounding box position, confidence and morphological characteristics; the target tracking module is used for performing unmanned aerial vehicle target tracking based on the target detection result and outputting a tracking result; and the fusion decision module is used for confirming the target identity based on the tracking result and the preprocessed multi-source data and outputting a final recognition result. The technical problems of low detection precision of small targets, difficulty in distinguishing similar targets, inaccurate 3D motion prediction, difficulty in re-identification after long-time shielding and the like in the prior art can be solved, and accurate identification, stable tracking and intelligent decision making of the unmanned aerial vehicle target are realized.
Owner:ERDOS SHIDA TECH CO LTD

CT image segmentation and classification system based on segmentation feature guidance

The invention belongs to the technical field of medical image processing, and discloses a CT image segmentation and classification system based on segmentation feature guidance, and the specific technical scheme is as follows: the system adopts a shared encoder to extract general features, realizes collaborative optimization of segmentation and classification through a double decoding path, adopts a partial decoder in a segmentation path, and adopts a partial decoder in the segmentation path; in combination with a local feature attention module, through multi-scale feature fusion and a boundary perception mechanism, the region consistency of a global segmentation map is gradually optimized, edge detail information is supplemented, and a classification path generates a space attention weight through a segmentation feature guide module by utilizing segmentation prediction; the classification network is guided to focus on a focus area and suppress background interference, a self-adaptive loss weighting strategy based on multi-task learning is adopted, the double-task gradient flow is dynamically balanced, and the gradient competition problem in the multi-task learning is effectively relieved.
Owner:SHANXI MEDICAL UNIV +1