Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

45171 results about "Visual perception" patented technology

Visual perception is the ability to interpret the surrounding environment using light in the visible spectrum reflected by the objects in the environment. This is different from visual acuity, which refers to how clearly a person sees (for example "20/20 vision"). A person can have problems with visual perceptual processing even if they have 20/20 vision.

High-precision image processing method and system based on illumination adaptive compensation

The invention discloses a high-precision image processing method and system based on illumination adaptive compensation, and relates to the technical field of computer vision and image processing, and the method comprises the steps: inputting an original image, and dividing the image into a high-frequency edge layer, an intermediate-frequency texture layer and a low-frequency illumination layer through a multi-scale residual network; acquiring illumination intensity, color temperature and scene categories in real time by using an ambient light sensor and a scene semantic segmentation model, and generating dynamic compensation parameters; carrying out dynamic range expansion on a low-frequency illumination layer based on a physical illumination model, and adjusting the weight of highlight suppression and dark area enhancement through a self-adaptive S-shaped exposure curve; a double-branch generative adversarial network is adopted, noise suppression and super-resolution reconstruction are carried out on the high-frequency layer, and texture detail enhancement is carried out on the intermediate-frequency layer; aligning the data of the depth camera and the infrared sensor with the visible light image through a cross-modal fusion module; and performing tone mapping on the fused image based on human visual characteristics, and outputting an enhanced image with a high dynamic range and reserved details.
Owner:SHANXI UNIV

Adaptive Real Time Image and Video Processing Using PCM-Enhanced Visual Strategy Caching and Multi-Stage Cognitive Routing

A system and method for adaptive image and video processing using a Persistent Cognitive Machine (PCM) architecture with visual strategy caching. The system receives degraded input media and extracts degradation fingerprints to query a PCM-based visual strategy cache containing previously successful processing strategies. When matching cached strategies are found above a relevance threshold, they are retrieved and applied directly. When no match exists, the input is processed through transform-domain networks to generate new strategies. A pattern synthesizer combines multiple strategies for complex degradation types. The system evaluates processing effectiveness using a feedback controller and stores successful strategies in the hierarchical cache. This cognitive approach enables real-time processing with continuously improving performance as the cache learns from successful patterns. The adaptive architecture eliminates redundant processing while maintaining high-quality output, making it suitable for diverse imaging and video applications requiring efficient enhancement capabilities with superior performance over traditional methods.
Owner:ATOMBEAM TECH INC

Multi-modal data processing method and apparatus, electronic device, computer-readable storage medium, and computer program product

Disclosed in the present application are a multi-modal data processing method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a reference image and a reference text; extracting a reference visual feature of the reference image; by means of a multi-modal large language model, determining an embedding of the reference text, an embedding of a start mark of the reference visual feature, an embedding of the reference visual feature, and an embedding of an end mark of the reference visual feature; on the basis of the multi-modal large language model, splicing the embedding of the reference text, the embedding of the start mark, the embedding of the reference visual feature, and the embedding of the end mark into a target embedding sequence, performing attention processing on the basis of the embedding of the start mark, the embedding of the end mark, and an embedding selected by a sliding window in the target embedding sequence, and outputting a predicted sequence; and generating a predicted image and a predicted text on the basis of the predicted sequence.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

System and method for efficient scene continuity in visual and multimedia using generative artificial intelligence

ActiveUS20250378537A1Image enhancementPattern recognitionGenerative process
A system and method for generating multimedia artifacts with managed scene continuity in visual and multimedia using an AI-based and scene continuity aware media generation platform. The system receives a user or AI agent specification or simulation result(s), selects or trains generative models based on the specification, preprocesses relevant data, and generates scene narrative or frame-specific, sequence specific or broader continuity aware content using the selected or trained model(s). The generated content may be further enhanced using frame interpolation and view synthesis techniques to create smooth transitions or novel viewpoints or to aid in more efficient transmission or viewing or persistence of resultant content. The system enables efficient and customizable generation of high-quality scene continuity aware content for various applications in visual and multimedia production using neuro-symbolic and simulation enhanced compression, representation and generation processes.
Owner:QOMPLX INC

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Virtual stylist

An example operation may include at least one of receiving, via a user interface of a device, an activation input from a user to initiate a session, capturing, by a camera of the device, a scan of a body of the user, wherein the capturing comprises recording at least one image and / or at least one video of the user, processing the at least one image and / or video to generate a three- dimensional model of the user comprising measurements and contours of the body, retrieving, from a database, at least one clothing item associated with the user, the at least one clothing item comprising dimensional attributes and texture attributes, rendering, by a graphics processing unit, the at least one clothing item onto the three-dimensional model to generate a visual representation, wherein the rendering simulates draping behavior, movement, and light interaction of the at least one clothing item relative to the three-dimensional model, and displaying, on the user interface, an interactive visualization comprising the visual representation of the three-dimensional model with the at least one clothing item from multiple viewing angles.
Owner:ELGORT PENELOPE

Weldment welding seam automatic detection method and device based on machine vision

The invention discloses a weldment welding seam automatic detection method and device based on machine vision, and relates to the technical field of machine vision intelligent detection. The weldment welding seam automatic detection method and device based on machine vision comprises the steps that S1, surface images and forming feature data of a weldment are collected and preprocessed to construct a standardized image feature data set; s2, the boundary clearness of the weld joint is evaluated by combining the edge strength and the contour coherence, and the main contour extraction range is dynamically adjusted; s3, analyzing abnormal focusing characteristics of the candidate area, and adjusting a defect labeling range and a detection priority; and S4, integrating the boundary definition and the abnormal focusing features, analyzing the structure abnormality, and dynamically controlling and verifying a resource allocation strategy. The problems that in the weldment detection process, obvious light reflection and texture blurring phenomena exist in a heat affected area at a weld joint, a traditional image enhancement and edge extraction algorithm is difficult to stably recognize microdefects, and the credibility of a detection result is reduced are solved.
Owner:WUXI TIENENG PRECISION MASCH CO LTD

Unmanned aerial vehicle target detection method based on frequency-space joint attention and dynamic fusion

The invention relates to the technical field of computer vision detection, in particular to an unmanned aerial vehicle target detection method based on frequency-space joint attention and dynamic fusion, and the method comprises the steps: obtaining an unmanned aerial vehicle image data set, carrying out the preprocessing, and dividing a training set and a test set; constructing a target detection model, inputting the training set into the target detection model to extract image features, sequentially performing frequency domain detail enhancement, spatial domain salient region extraction and multi-scale feature adaptive fusion based on the image features, and establishing a feature sequence; screening the feature sequence to obtain an initial target query, and finishing target classification and positioning on the initial target query through a decoder; training a target detection model by using the training set, and inputting the test set into the trained target detection model to generate a detection result; on the premise that the real-time reasoning advantage of RT-DETR is kept as much as possible, the problems that in an unmanned aerial vehicle scene, a target is prone to missing detection, the scale change is large, the background is complex, and the target is fuzzy are effectively solved, and the detection precision is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Multi-modal visual fusion complex scene small target detection tracking method and system

The invention discloses a multi-modal visual fusion complex scene small target detection tracking method and system, and relates to the technical field of unmanned aerial vehicle target tracking, and the method comprises the steps: employing a visible light camera, an infrared thermal imager and a laser radar sensor which are carried on an unmanned aerial vehicle platform, and synchronously collecting RGB images, thermal infrared images and point cloud data; the consistency of the multi-modal data is ensured through data preprocessing and space-time alignment; constructing a lightweight double-branch network to extract multi-scale features, generating a fusion feature map by adopting adaptive weighted fusion, and generating depth information by utilizing point cloud to assist in scale estimation; a small target detection head is designed based on the fusion feature map, and precise detection is realized in combination with a feature pyramid network, adaptive scale prediction and a context awareness suppression mechanism; furthermore, through multi-mode cooperative tracking, including target association, spatio-temporal context modeling, trajectory prediction and a re-detection mechanism, tracking continuity is ensured.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

Uncoupling robot control system and method based on multi-source visual fusion

The embodiment of the invention provides an unhooking robot control method based on multi-source visual fusion, which is applied to the technical field of robot control and comprises the following steps: acquiring an RGB image, a depth image, an infrared image and IMU data through a multi-source sensing system mounted at the tail end of a robot; carrying out feature fusion identification by adopting a double-branch neural network, and outputting the boundary contour of the lifting hook and the three-dimensional coordinates of the optimal grabbing point; the visual coordinates are unified to a robot base coordinate system through a registration correction mechanism; a Transform prediction model is constructed based on the visual and inertial signals, and future pose changes of the lifting hook are estimated; a feedforward control track is generated to counteract swing of the lifting hook, and track correction is carried out in combination with visual servo feedback; and a joint instruction is generated through path planning and inverse kinematics solution, and the mechanical arm is driven to complete precise unhooking operation. According to the method, the recognition precision, the anti-interference capability and the operation success rate of unhooking operation in complex illumination and dynamic environments are effectively improved.
Owner:ANHUI HUADIAN SUZHOU POWER GENERATION

Human-guided vision-force fused impedance iterative learning control method for robotic arm

A human-guided vision-force fused impedance iterative learning control method for a robotic arm, comprising: analyzing a robot-environment interaction dynamics equation, solving a visual servo acceleration model, and making use of the equation to establish a human-robotic arm-environment interaction dynamics model in an image feature space; acquiring an image feature position and speed curve of a human-guided robot completing an assembly task, and using dynamic movement primitives for coding and generalization; and designing an impedance iterative learning controller which uses image feature tracking errors as control input, learning impedance characteristics when the human-guided robot performs a contact operation, identifying unknown contact dynamics under the interaction between the robot and the environment, and counteracting identified contact interference in the feature space, so as to implement a flexible assembly operation. The control method solves the problems in existing assembly operations that human-robotic arm-environment coupling nonlinear dynamics, unknown contact dynamics of intensive contact assembly tasks and poor generalization of assembly scenarios require relearning for different scenarios, etc.
Owner:HUNAN UNIV

Industrial part alignment method and system based on visual analysis and storage medium

The invention relates to the technical field of image processing, and discloses an industrial part alignment method and system based on visual analysis and a storage medium. The method comprises the steps that a three-view camera collects an industrial part image, and preprocessing is carried out through gradient magnitude local contrast enhancement to obtain an enhanced image; performing hierarchical feature extraction to identify edge contours and key control points to form a multi-dimensional feature set; and establishing a dynamic reference coordinate system based on the feature set to obtain a part space attitude matrix. And the attitude deviation is compensated through Z-axis offset and rotation coupling error analysis. Posture adjustment is decomposed into a plurality of sub-stages, an alignment track is optimized by adopting a variable speed planning strategy, and accurate alignment of the parts is achieved. The problems that multi-view visual information fusion is insufficient, a special recognition algorithm for geometrical characteristics of the industrial parts is lacked, and Z-axis offset and rotation coupling error compensation is inaccurate in the posture adjustment process are solved, and the precision and stability of alignment of the industrial parts are improved.
Owner:BEIJING TIANYUAN 3D TECH CO LTD

Dynamic Latent Space Adaptation Based on Spatiotemporal Kernal Context for Multiscale Rendering

A system for dynamic latent space adaptation using spatiotemporal kernel context for multiscale rendering with hierarchical and Lorentzian autoencoders. The Spatiotemporal Kernel Estimator (SKE) analyzes media through motion field, temporal recurrence, frequency band, and scene semantics analyzers to generate adaptive kernel parameters encoding content-specific importance distributions. The system dynamically adapts latent manifold geometry by modifying metric tensor properties according to kernel context, enabling content-aware compression that allocates representational capacity based on visual significance. A multiscale cache implements kernel-adaptive retention policies prioritizing important regions. An adaptive renderer provides intelligent level-of-detail selection based on zoom level and kernel-estimated importance, optimizing processing allocation. The self-optimizing architecture continuously refines kernel context and geometric adaptation based on user interaction and performance feedback, achieving superior compression ratios and perceptual quality. Applications include bandwidth-efficient video streaming, virtual reality, scientific visualization, and cognitive video analytics requiring intelligent context-aware visual processing.
Owner:ATOMBEAM TECH INC

Three-dimensional dynamic scene reconstruction method and apparatus, and storage medium

The present disclosure relates to the field of computer vision and discloses a three-dimensional dynamic scene reconstruction method and apparatus, and a storage medium. The three-dimensional dynamic scene reconstruction method comprises: acquiring synchronized videos of a plurality of viewpoints of a dynamic scene; computing matching points between video images of different viewpoints, and estimating intrinsic and extrinsic parameters of each camera; obtaining a Gaussian splatting point set {p0} on the basis of a sparse point cloud constructed according to the depth of each matching point; for the first image frame of each video, using {p0} to perform static training thereon, to obtain a Gaussian splatting point set {p}; for the remaining image frames, dividing {p} into a static point set {S} and a dynamic point set {D}, performing dynamic training on {D}, and constructing a dynamic Gaussian splatting point set {P} from {p}, {S}, and the final {D}; and, in view of the intrinsic and extrinsic parameters of each camera, rendering {P} using a Gaussian splatting rendering pipeline, to obtain rendered images at different moments from new viewpoints.
Owner:TSINGHUA UNIVERSITY

Cross-modal image-text analysis method for machine vision

The invention relates to the technical field of machine vision, and discloses a machine vision-oriented cross-modal image-text analysis method, which comprises the following steps of: partitioning an input image to generate an image block sequence; inputting the image block sequence into a visual converter for multi-scale feature extraction, and generating target visual features; encoding the input text to generate a target text feature; inputting the target visual features and the target text features into a deep reconstruction bottleneck network for compression alignment, and generating a cross-modal compression vector; and inputting the cross-modal compression vector into a large language model to generate cross-modal decoding information, so that cross-modal redundant information can be effectively filtered, compact shared semantic representation can be learned, the information integrity of the compression process is ensured through bidirectional reconstruction verification, cross-modal semantic alignment is realized, and the method has the advantages of high efficiency and high reliability. Omnibearing cross-modal content generation from the whole to details is achieved, and the requirements of different application scenes are met.
Owner:SHENZHEN YOULIANCHUANG WISDOM TECH CO LTD

Unmanned aerial vehicle laser and vision fusion inspection method and system for bridge bottom disease detection

The invention discloses an unmanned aerial vehicle laser and vision fusion inspection method and system for bridge bottom disease detection, and the method comprises the steps: carrying out the synchronous data collection through employing a calibrated laser radar, a camera and an IMU, and obtaining a three-dimensional laser point cloud and a two-dimensional visual image of the appearance of a bridge; sharpening the image containing the motion blur and completing brightness self-adaption of the image; stable feature points are extracted, multi-frame matching is carried out, the corresponding poses of the images are estimated, and bridge dense point cloud reconstruction is completed; performing geometric component segmentation on the point cloud to generate a geometric prior region; component segmentation is carried out on a support area in the image, and a continuous and accurate component segmentation result is obtained in combination with a geometric prior area; screening the image, calling a targeted disease detection model in a corresponding component area, and generating a segmentation mask for the disease; obtaining a real disease three-dimensional point cloud, and carrying out quantitative calculation on the physical size of the disease; and displaying the real disease three-dimensional point cloud data and the physical size of the disease. The method is high in efficiency and precision.
Owner:SOUTHEAST UNIV

Cable surface defect detection method and system based on machine vision

The invention relates to the technical field of image data processing, in particular to a cable surface defect detection method and system based on machine vision. The method comprises the following steps: acquiring a surface image of a cable, and converting the surface image into a grayscale image; determining the local complexity of each pixel point; obtaining a plurality of areas of the grayscale image, performing complexity determination, and dividing each area to obtain a plurality of windows of each area; determining a contrast limit threshold value of each window; performing image enhancement by using a CLAHE algorithm to obtain an enhanced grayscale image; and carrying out cable surface defect detection on the enhanced grayscale image by using a defect detection algorithm. According to the method, the CLAHE parameters are adaptively adjusted based on local complexity, the window size and the contrast threshold are dynamically determined in combination with gray and gradient information, discontinuity is corrected and eliminated through boundary similarity, and the quality and reliability of a cable defect detection image are improved.
Owner:CHUNHUA KUNLUN YOUJIA CABLE CO LTD

Multi-sensory autonomous multimodal emotion-synchronized environmental control architecture and regulation system (amesecar)

An autonomous environmental regulation and behavioral monitoring system is disclosed, configured to adapt temperature, lighting, and acoustic conditions based on real-time emotional and physiological data. The system includes a dual-redundant central processor, hierarchical communication networks, multi-angle visual acquisition units, infrared thermometers, and modular environmental subsystems. It detects posture, gestures, facial expressions, and thermal signals to classify user states and apply individualized airflow, light, and sound modulation without relying on external internet connectivity. The system also monitors connected appliances using voltage-based pressure analysis to forecast device degradation. With integrated gesture recognition, privacy-preserving data handling, and predictive adaptation, the invention enables multi-user personalization, long-term learning, and uninterrupted operation within residential, administrative, or healthcare infrastructures.
Owner:SEYEDKHAMOUSHI FAEZEHALSADAT +1

Multi-mode body-equipped intelligent robot control method and device

The invention relates to the technical field of body-equipped intelligent robots, in particular to a multi-mode body-equipped intelligent robot control method and device, and the method comprises the steps: synchronously collecting visual, auditory, tactile, force sense and body perception information, and unifying the information to the same time-space reference through a cross-mode time-space stamp alignment mechanism; hierarchical feature extraction and fusion are carried out on the multi-modal information, and unified multi-modal scene state representation is generated; reasoning a decision based on the representation by using a body agent framework, and outputting a control instruction; motion planning and control, visual servo tracking in a non-contact stage and dynamic parameter correction in a contact stage are executed according to instructions; optimizing the multi-modal strategy network through an incremental strategy distillation mechanism based on the interactive data flow; the problem of space-time asynchronization of multi-modal sensing information is solved through a cross-modal space-time stamp alignment mechanism.
Owner:CHONGQING IND INTELLIGENCE TECHNOLOGY RESEARCH INSTITUTE

Robot multi-modal fusion autonomous decision-making method and system based on large language model

The invention relates to the technical field of robot decision making, and provides a robot multi-modal fusion autonomous decision making method and system based on a large language model.The method comprises the steps that a robot obtains multi-modal environment information through a visual sensor, a touch sensor, an auditory sensor and a laser radar which are carried by the robot; performing preliminary filtering and noise reduction processing on the original sensor data, and synchronously recording all the sensor data by timestamps; performing space-time semantic alignment on the preprocessed multi-modal data, mapping pixel coordinates of a target in a visual target coordinate quantization original image to a robot coordinate system, performing uncertainty evaluation on a multi-modal signal through a dynamic Bayesian network, and taking entropy or variance as an uncertainty quantitative evaluation index. According to the method, the information quality is improved from a data fusion source, accurate and reliable basic support is provided for subsequent decision making, and decision making errors caused by data deviation are greatly reduced.
Owner:ANHUI UNIV +1

Tooth three-dimensional modeling system based on computer vision, computer equipment and readable storage medium

The invention relates to the technical field of tooth modeling, and discloses a three-dimensional tooth modeling system based on computer vision, computer equipment and a readable storage medium. According to the method, mirror reflection, diffuse reflection and subsurface scattering components in an original image are separated, mirror reflection intensity is normalized in combination with a dynamic truncation algorithm, pixel saturation is eliminated, groove and nest textures are reserved, a complete point cloud is obtained based on a two-dimensional texture image and cubic spline repair, and a multi-exposure point cloud sequence is obtained through bimodal calibration. The method comprises the following steps: solving the problem of data dislocation, carrying out weight assignment and data fusion on three-dimensional points in a plurality of exposure point cloud sequences to obtain three-dimensional fusion feature data, combining layered optical modeling and photon tracking compensation deviation, fusing clinical constraints, finally dynamically adjusting parameters, feeding back and optimizing, and outputting a micron-sized precision model. The modeling defect caused by difficulty in effectively coordinating feature contribution degrees under different exposure conditions is overcome, and high-precision modeling is realized.
Owner:SHENZHEN JINSHI LIMEI MEDICAL TECH CO LTD

Generative ai models for image rendering and inverse rendering

Embodiments of the present disclosure relate to rendering and inverse rendering using one or more generative models. “Rendering” refers to the process of generating a final visual image, video frame, or animation from a 2D or 3D model. “Inverse rendering” is a process that involves deducing or estimating the properties (e.g., material maps or other properties such as geometry, lighting, and textures) of a scene from observed images or visual data. Essentially, it aims to reverse the traditional rendering process. Various aspects of the present disclosure introduce editable light and material controls into generative models to allow for artistic creation. Various embodiments integrate generative models as a renderer for classic rendering pipelines to upcycle and enhance the style of rendered content.
Owner:NVIDIA CORP

Light guide plate defect detection method and system based on neural network

The invention discloses a light guide plate defect detection method and system based on a neural network, and particularly relates to the technical field of machine vision detection, and the method comprises the following steps: aiming at the problem of image instability of a light guide plate in a dynamic transmission or rotation process, continuously collecting an image sequence and extracting time domain features; and performing interference judgment in combination with the inter-frame consistency prediction coefficient and a first threshold to realize accurate identification of the abnormal image frame. For an abnormal image frame, further correcting the recognition credibility of the abnormal image frame by adopting a confidence adjustment and fusion mode, and meanwhile, introducing a frequency domain transformation and image enhancement strategy to compensate detail loss caused by motion blur; according to the method, inter-frame consistency analysis, confidence fusion regulation and control and frequency domain fuzzy recognition and compensation mechanisms are introduced, abnormal judgment and image quality restoration of the light guide plate image in the dynamic scene are realized, the recognition accuracy and stability of the neural network model on the defect type, position and confidence are improved, and the false detection and omission ratio is effectively reduced.
Owner:深圳市鸿卓电子有限公司

Waste metal classification and identification method and system based on image identification

The invention relates to the technical field of industrial visual inspection, and particularly discloses a waste metal classification and recognition method and system based on image recognition, and the method comprises the steps: obtaining metal surface visual information through an image collection system, extracting multi-level depth features after preprocessing, and generating a preliminary classification result and confidence evaluation; when the confidence coefficient is insufficient, starting a multi-mode verification mechanism, acquiring element composition data by adopting a laser-induced breakdown spectroscopy technology, and acquiring surface topological characteristics by adopting a structured light three-dimensional scanning technology; matching the element data with a component database to generate a component verification result, and comparing the morphology features with a morphology database to generate a morphology verification result; and finally, three types of results are integrated based on a weighted fusion algorithm to generate a final classification decision, and a sorting mechanism is controlled to complete accurate sorting.
Owner:JIANGXI JIANGLING NON-FERROUS METAL DIE-CASTING CO LTD

Visual language navigation method for cross-modal alignment in dynamic shielding environment

The invention discloses a visual language navigation method for cross-modal alignment in a dynamic shielding environment, and the method comprises the steps: collecting multi-modal data through a visual sensor, an inertial measurement unit, a laser radar and the like, and carrying out the preprocessing and time synchronization; sensing the dynamic shielding object through a model composed of a convolutional neural network and a long-short-term memory network, and estimating the future change of the dynamic shielding object in combination with a space-time sequence prediction algorithm; a double-branch convolutional neural network and a Transform based on a dynamic attention mechanism are adopted to respectively extract visual and semantic features and fuse the visual and semantic features; on the basis of occlusion prediction, potential occlusion region features are extracted in advance from a time dimension, an occluded image is repaired by using a generative adversarial network and geometric constraints in a space dimension, and cross-modal feature alignment is optimized through an attention mechanism; planning a path by using a hybrid reinforcement learning algorithm based on a deep Q network-space and a fast exploration random tree, and dynamically adjusting according to real-time shielding; according to the method, the accuracy, adaptability and reliability of visual language navigation in a dynamic shielding environment are improved.
Owner:SHANGHAI JIAOTONG UNIV

Laser engraving method and system for automatically correcting coordinates of galvanometer and camera

The invention relates to the technical field of laser engraving, and discloses a laser engraving method and system capable of automatically correcting coordinates of a galvanometer and a camera, and the method comprises the following steps: calculating a target conversion coefficient based on a calibration size and a pixel size of an image collected by the camera, and controlling the galvanometer to perform laser etching on a cross mark on the surface of a PCB (Printed Circuit Board), recording origin data of a galvanometer coordinate system; adjusting the position of the camera to enable the cross center of the view of the camera to coincide with the cross mark, and calculating the coordinate offset; coordinate space transformation is carried out on all the points to be machined, and galvanometer marking coordinate data are obtained; after laser engraving is carried out on the galvanometer marking coordinate data, an engraving area image is collected, the position residual error and the size proportion deviation are calculated, laser engraving is carried out again, and a laser engraving result is obtained. The technical problem that the visual positioning coordinate and the galvanometer marking coordinate in the laser engraving system are not consistent is effectively solved.
Owner:SHENZHEN ZHENHUAXING INTELLIGENT TECH CO LTD

Knowledge graph completion method based on multi-mode visual angle perception and deep neural network

The invention relates to the field of knowledge graph completion, provides a knowledge graph completion method based on multi-modal visual angle perception and a deep neural network, and aims to solve the problems of weak multi-modal information expression ability, rough fusion mode and insufficient structural reasoning ability in the prior art. According to the method, structure information, text description and visual image information of an entity in a knowledge graph are obtained, structure, text and image modal input is constructed respectively, and a graph neural network, a pre-training language model and a visual encoder are adopted for feature coding; weighted fusion and semantic enhancement of multi-modal features are realized through a visual angle fusion mechanism and hierarchical attention processing; cross-modal contrast learning is introduced to improve modal consistency; and carrying out triple reasoning by using a uniform Transform encoder, and verifying a completion result by scores. According to the method, multi-modal semantics are effectively integrated, the entity representation capability and the triple prediction accuracy are improved, the model robustness is enhanced, and the method is suitable for application scenes such as intelligent question answering and recommendation systems and has remarkable practical value and popularization prospects.
Owner:DALIAN NATIONALITIES UNIVERSITY

YOLOv8 algorithm improvement method based on unmanned aerial vehicle aerial image small target detection model

The invention belongs to the technical field of computer vision and artificial intelligence, belongs to the cross technical field of target detection, deep learning and image processing, and particularly relates to a YOLOv8 algorithm improvement method based on an unmanned aerial vehicle aerial image small target detection model, which comprises the following steps of: introducing a user-defined feature enhancement module into a YOLOv8 backbone network, a neck part and a detection head part; the self-defined feature enhancement module comprises a context guide self-adaptive fusion module introduced into a backbone network so as to replace part of traditional convolution operation; a space edge sensing feature up-sampling module and a space sensing enhanced convolution module are adopted in the neck fusion network; a fine-grained dynamic pruning detection head is introduced into a detection head detection network. According to the method, the performance of the model in a small target detection scene is effectively enhanced, and the accuracy, robustness and real-time response capability of a detection system are remarkably improved.
Owner:YANCHENG INST OF TECH

Multi-view three-dimensional point cloud reconstruction method and device based on DPE-SE depth estimation

The invention provides a multi-view three-dimensional point cloud reconstruction method and device based on DPE-SE depth estimation, and relates to the technical field of computer vision and three-dimensional reconstruction. The method comprises the following steps: acquiring multi-view image data; preprocessing the image; inputting the preprocessed image into a DPE-SE-based depth estimation model, carrying out key constraint on an edge region through a semantic edge guiding mechanism, carrying out adaptive propagation updating on a weak texture region, realizing accurate depth estimation, and generating a multi-view depth result; then geometric consistency check and multi-scale depth fusion are performed on a multi-view depth result, and a dense depth map is constructed; and finally, performing three-dimensional back projection reconstruction and point cloud optimization processing, and outputting high-quality point cloud data containing three-dimensional coordinates and confidence information. According to the method, the problems of edge mismatching and depth voids are remarkably improved in complex illumination, weak texture and shielding environments, the continuity and structural integrity of the point cloud are improved, and technical support is provided for unmanned aerial vehicle surveying and mapping, building detection and digital twin modeling.
Owner:HUAQIAO UNIVERSITY +1