Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

5100 results about "Object detection" patented technology

Object detection is a computer technology related to computer vision and image processing that deals with detecting instances of semantic objects of a certain class (such as humans, buildings, or cars) in digital images and videos. Well-researched domains of object detection include face detection and pedestrian detection. Object detection has applications in many areas of computer vision, including image retrieval and video surveillance.

Construction scene prediction method and device fusing image, text and BIM mode

The invention provides a construction scene prediction method and device fusing an image, a text and a BIM modal, and relates to the technical field of intelligent construction prediction management. According to the method, the BIM semantic graph is constructed by extracting the BIM semantic information of the BIM model, and the target detection recognition of the construction site video is carried out in combination with the YOLO model to obtain the target detection result; performing cross-modal alignment with the CLIP to realize deep fusion of the multi-modal data of the image, the text and the BIM to obtain a multi-modal heterogeneous graph; and inputting the multi-modal heterogeneous graph into a space-time sequence model for prediction, outputting prediction results of construction scenes at a plurality of moments in the future, and dynamically mapping the prediction results to a digital twinborn platform to realize risk early warning and visual display. The construction dynamic change can be captured in real time, the construction progress and risk can be accurately predicted, and the intelligent level of construction management is improved.
Owner:XIAMEN UNIV OF TECH

Improved YOLOv11s safety helmet wearing detection model and optimization method thereof

The invention provides an improved YOLOv11s safety helmet wearing detection model and an optimization method thereof, and relates to the technical field of computer vision target detection. According to the improved YOLOv11s safety helmet wearing detection model and the optimization method thereof, the improved YOLOv11s safety helmet wearing detection model comprises the following modules: a multi-modal fusion module, a space-time analysis module, a domain adaptation module, a topological optimization module and a dynamic architecture module; and the multi-modal fusion module is used for realizing feature decoupling by adopting channel separation convolution based on input RGB and near-infrared images, fusing visible light and thermal radiation features through a dynamic weight distribution algorithm, implementing affine transformation alignment on multi-scale features by utilizing a spatial transformation network, and generating a multi-modal feature graph. Through fusion of visible light and near infrared spectrum features and implementation of dynamic weight distribution, complementarity of target texture and thermal radiation features under a complex illumination condition is enhanced, and the problem of feature distortion of single-mode data in a strong backlight or low-illumination scene is solved.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Manipulator grabbing method based on deep learning target detection and image segmentation

The invention discloses a manipulator grabbing method based on deep learning target detection and image segmentation, and relates to the technical field of artificial intelligence and robotics.The manipulator grabbing method comprises the following steps that a scene image to be processed is collected, the image quality is improved through the multi-light-source fusion image enhancement technology, and recognition errors caused by uneven illumination are reduced; and inputting the enhanced image to a pre-trained deep learning model, executing a target detection task, and outputting an initial bounding box and a category label of the target object. According to the method, through multi-light-source image enhancement and high-precision image segmentation, the accuracy of target recognition and contour extraction is remarkably improved, and the capture failure rate caused by image misjudgment is reduced. And meanwhile, geometric consistency verification and a multi-factor grabbing scoring mechanism are introduced, dynamic screening and collision pre-detection are conducted on the paths, the grabbing stability and safety of the mechanical arm in the complex environment are effectively guaranteed, and the intelligence and robustness of the whole system are remarkably improved.
Owner:SHENZHEN BOCHUANG ROBOT TECH

Multi-view panoramic point cloud splicing method

The invention discloses a multi-view panoramic point cloud splicing method, which comprises the steps of performing joint calibration on a binocular camera and a laser radar through a calibration algorithm, and aligning a coordinate system; a binocular camera is used for collecting a two-dimensional image and carrying out distortion correction, a depth map is generated based on parallax calculation, and three-dimensional point cloud reconstruction is carried out in combination with a triangulation principle; performing deep learning target detection on the two-dimensional image, and screening a line rod assembly area through non-maximum suppression; multi-view point cloud data are collected through a laser radar, point cloud segmentation processing is carried out, and a telegraph pole assembly point cloud subset is reserved; matching the two-dimensional pixel area of the telegraph pole assembly with a laser radar point cloud projection result, and screening point cloud data belonging to the telegraph pole assembly; and carrying out alignment and fusion on the multi-view point clouds through initial registration and fine registration by adopting an improved point cloud splicing algorithm to generate a panoramic point cloud image. According to the method, the point cloud splicing processing time is shortened, and the robustness, precision and integrity of point cloud splicing are improved.
Owner:NANJING SIWEI VECTOR TECH CO LTD

Construction progress monitoring method and system based on big data

The invention relates to the technical field of construction progress monitoring, and discloses a construction progress monitoring method and system based on big data. The method comprises the following steps: forming a space-time alignment data set through multi-source data acquisition, filtering and quality evaluation; performing feature extraction and registration to generate a digital model; target detection classification is performed to form a completion state table; progress evaluation is achieved through component-task mapping; trend analysis and risk identification are performed to generate a prediction result; decision reference is provided for personalized information screening and augmented reality display. Through multi-source data acquisition, fusion and intelligent analysis, accurate perception, objective evaluation, scientific prediction and visual presentation of the actual state of the construction site are realized, so that a comprehensive, accurate and prospective construction progress monitoring method is provided, the construction period delay risk is effectively reduced, and the construction management efficiency is improved.
Owner:ZHEJIANG ENERGY CONSTR CO LTD

Power operator behavior identification early warning system and method based on video analysis

The invention discloses a video analysis-based electric power operation personnel behavior identification and early warning system and method, which realize accurate identification and real-time early warning of electric power operation personnel behaviors by combining a video analysis technology with multi-modal data fusion, and effectively improve the safety management level of an operation site. Compared with a traditional safety supervision mode, the method employs a mode of combining deep learning target detection and time sequence behavior analysis, improves the recognition accuracy of operators and safety equipment, fuses the data of equipment worn by the operators with video data, improves the detection precision, and improves the safety supervision accuracy. And misjudgment caused by illumination change, shielding or complex environment is reduced. Besides, high-risk violation behaviors such as no safety helmet wearing, no safety belt fastening, violation climbing, high-altitude object throwing and the like are accurately recognized through the violation detection module, and different levels of alarm measures are adopted according to the severity of the violation behaviors in combination with an early warning feedback mechanism, so that the pertinence and response efficiency of early warning are improved.
Owner:PENGLAI WIND POWER BRANCH OF HUANENG SHANDONG POWER GENERATION CO LTD +1

Robot obstacle avoidance and navigation method based on multi-modal fusion and visual language model

The invention provides a robot obstacle avoidance and navigation method based on multi-modal fusion and a visual language model, and the method comprises the steps: collecting different modal data in real time through a plurality of sensors, and carrying out time synchronization processing and normalization processing; extracting features of different modal data and fusing the features through an intermediate layer; and inputting the fused multi-modal data into a visual language model, generating a semantic map of the environment by utilizing semantic segmentation of the model and a target detection result, generating an action strategy in combination with a natural language instruction and a visual analysis result, and converting the generated action strategy into a control signal which can be executed by the robot to realize closed-loop control. According to the method, the visual language model and the multi-modal sensor fusion technology are combined, and the perception, decision making and real-time response capabilities of the mobile robot in a complex dynamic environment are improved.
Owner:ZHUHAI MAKERWIT TECH CO LTD

Data center inspection robot monitoring analysis method and system based on machine vision

The invention provides a data center inspection robot monitoring analysis method and system based on machine vision, and relates to the technical field of inspection robots, and the method comprises the steps: obtaining a video stream and environment parameter data collected by an inspection robot, and carrying out the detection and recognition of an abnormal state through a deep learning target; and constructing a multi-modal data fusion analysis model to form a knowledge graph to generate a root cause analysis result, and planning an inspection path based on priority scores. According to the invention, intelligentization and precision of data center monitoring are realized, and inspection efficiency and fault diagnosis accuracy are improved.
Owner:BEIJING AMPLI INFORMATION TECHNOLOGY CO LTD

Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Object detection and tracking using machine learning transformer models with attention

Object detection and tracking systems may use machine-learned transformer models with self-attention for detecting, classifying, and / or tracking objects in an environment. Techniques described herein may include receiving sensor data generated by different sensor modalities of a vehicle, determining different bounding shapes based on the different sensor modalities, and using a machine-learned transformer model to determine associated and / or combined bounding shapes. The machine-learned transformer model may receive a variable number of input bounding shapes representing any number of objects and various sensor modalities. Multiple stages of the transformer may be used to determine associated bounding shapes and to assign attributes for the associated bounding shapes, based on the individual bounding shapes of the different sensor modalities and / or previous bounding shapes for objects detected and tracked in a previous scene in the environment.
Owner:ZOOX INC

Industrial PCB defect identification method based on sample generation model

The invention belongs to the field of defect detection, and particularly relates to an industrial PCB defect identification method based on a sample generation model. Comprising the following steps: constructing an original PCB defect image data set with labels, and preprocessing original PCB defect images to obtain preprocessed images; inputting the preprocessed image into a sample generation model for processing to generate a defect sample; combining the original PCB defect image data set and the defect sample, and performing adaptive size filling and high-frequency noise injection processing to obtain an enhanced defect data set; training the YOLO-ADF target detection model by adopting the enhanced defect data set to obtain a trained YOLO-ADF target detection model; sampling the trained target detection model to carry out industrial PCB defect identification; according to the invention, the accuracy of small target defect detection is improved, the omission ratio is reduced, and the detection accuracy and the model robustness are improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Target detection method based on Mama feature fusion

The invention discloses a target detection method based on Mama feature fusion, and relates to the technical field of image target detection. According to the method, the innovative implementation of the VSSA module is utilized, a selective scanning mechanism of the state space model is applied to 2D visual data processing, the long-distance dependency relationship in the image is effectively captured through state space modeling in four directions, the limitation of a traditional state space model in the two-dimensional visual data processing process is solved through the multi-direction processing strategy, and the processing precision of the 2D visual data is improved. The model can comprehensively perceive spatial dependency relationships in different directions in an image, the VSSA adopts learnable state space parameters to dynamically model a feature sequence, the ability of the network to understand a complex space structure is enhanced, the method is particularly suitable for processing scenes needing long-distance context information, and in addition, the method is combined with MTMHSA, so that the complexity of the network is reduced. And the fusion capability of different levels of features in target detection is further enhanced. Through the innovation, the model can better understand the target in the image, and the positioning and classification precision of the target is improved.
Owner:CHONGQING UNIV OF TECH

Unmanned aerial vehicle image-based small object detection method for target areas

The present invention relates to the technical field of deep learning and computer vision. Disclosed is an unmanned aerial vehicle image-based small object detection method for target areas. The present invention crops images of obvious small objects in certain target areas, and annotates the small objects of different categories to form a raw training and testing dataset, so as to ensure the accuracy of data required in the early stage of the algorithm and further ensure the scientificity of the algorithm; uses the computing capability of an improved YOLOv7 detection model to collect image features of different degrees in the dataset, the improved YOLOv7 detection model using YOLOv7 as a basic model and adding to a neck network an MS-CET module, which is constituted by an improved self-attention mechanism and convolution module SPPCSP, and a BHC-FB module, which is constituted by bidirectional mixed convolution modules NConv and RPConv connected in parallel; and finally fuses different feature layers as a final judgment basis of an unmanned aerial vehicle for small object detection in the target areas, to further check the accuracy of the algorithm and criteria for dataset selection, thereby improving recognition accuracy.
Owner:CHONGQING UNIV OF TECH

Interaction action detection method and device based on multi-level features

The invention discloses an interactive action detection method based on multi-level features, and the method comprises the steps: S1, fusing the information of a low-level feature map, a middle-level feature map and a high-level feature map through a cascading fusion mode, and obtaining a fused feature map; s2, acquiring local detail features from the fused feature map, performing global modeling to obtain global features, fusing the global features with the local detail features to generate global context features, performing character detection, object detection and interaction detection in parallel through multi-task branches, outputting character features, object features and interaction features, and outputting the character features, the object features and the interaction features. S3, gradually fusing multi-level contexts through attention interaction of a unitary relationship, a pairwise relationship and a ternary relationship, generating text embedding for an interaction category by utilizing a pre-training text encoder, and aligning the text embedding with visual features to improve the accuracy of fine-grained interaction classification; and S4, explicitly modeling a complex interaction relationship, and finally predicting and outputting a final prediction result by using FFN.
Owner:NORTH CHINA UNIVERSITY OF TECHNOLOGY

Railway foreign-object intrusion detection method and system based on deep learning

The present invention relates to the technical field of railway inspection, and relates in particular to a railway foreign-object intrusion detection method and system based on deep learning. The method comprises: S10, acquiring image data to undergo detection; and S20, inputting the image data into a trained attention semantic segmentation network to obtain a foreign-object detection result. The attention semantic segmentation network is obtained by first training a preset attention semantic segmentation network architecture using a preconfigured data set, and then re-training the attention semantic segmentation network on specified image data using an adaptive correction algorithm. A backbone network architecture is obtained by inserting a specified attention mechanism at a specified position within a pre-selected residual neural network, and using three parallel dilated convolutions as an initial convolutional layer in the residual neural network. A dual-branch decoder combines an edge recognition branch and a semantic segmentation branch. The method is applicable to various scenarios, achieves improved detection accuracy and maintains a lightweight design.
Owner:BEIJING JIAOTONG UNIV

Multi-domain object detection method and apparatus

A multi-domain object detection method includes generating a teacher model and a student model from a pre-trained model, inputting an image with weak augmentation applied to a target image, for which an object is to be detected, to the teacher model, determining whether a pseudo label generated by the teacher model is below a preset threshold, performing negative learning for a class corresponding to the pseudo label when the pseudo label is determined to be below the threshold, inputting an image with strong augmentation applied to the target image to the student model, calculating an unsupervised loss by comparing a first prediction generated by the student model with the pseudo label, updating the teacher model using an exponential moving average (EMA) predetermined in the student model, and detecting an object in an image from another domain using the teacher model.
Owner:HYUNDAI MOTOR CO LTD +2

Traffic scene target detection method based on feature fusion and attention mechanism

The invention relates to the technical field of traffic scene target detection, in particular to a traffic scene target detection method based on feature fusion and an attention mechanism, and the method comprises the steps: obtaining a public traffic scene target image training data set; training a traffic target detection model by using the traffic scene target image training data set; and inputting a to-be-detected traffic scene image into the trained traffic scene target detection model to obtain an output detection result graph. A hierarchical receptive field is constructed through a feature fusion module, a fine structure of a target edge is reserved in a deep convolution stage, cross-channel semantic information is dynamically aggregated through point convolution, an attention module is put forward to dynamically adjust a convolutional receptive field and a space attention weight, interference of background noise is suppressed, and a target is obtained. The local detail feature expression of the target is better enhanced, a PIoU loss function is adopted, and the positioning precision and robustness are further improved by combining a target size adaptive penalty factor and an anchor frame quality-based gradient adjustment strategy.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Synchronizing camera, lidar and radar for object detection using radar-guided scene flow estimation and adaptive attention

This disclosure provides systems, methods, and devices for vehicle driving assistance systems that support enhanced sensor fusion techniques. In a first aspect, a method of includes receiving point cloud data for two or more frames from a radar device and generating scene flow parameter data based on the point cloud data. The method also includes generating voxel position adjustment data based on the scene flow parameter data, and generating feature concatenation information associated with two or more sensors based on the voxel position adjustment data and feature information associated with the two or more sensors. The method further includes performing feature detection and tracking based on the feature concatenation information to generate tracking information for one or more objects, and outputting the tracking information. Other aspects and features are also claimed and described.
Owner:QUALCOMM INC

Scene perception method and device based on multiple modes, electronic equipment and storage medium

The invention provides a scene perception method and device based on multiple modes, electronic equipment and a storage medium. According to the method provided by the embodiment of the invention, image features of a multi-view image sequence and radar features obtained by a 4D radar data sequence are firstly extracted, and spatial-temporal evolution of a dynamic scene and a static scene in a BEV space and a voxel space is modeled by using historical radar features and current radar features to obtain dynamic scene features and static scene features; and performing cross-modal interaction fusion on the image features, the dynamic scene features and the static scene features to obtain multi-modal fusion features, wherein the multi-modal fusion features can be directly used for 3D target detection, semantic occupancy prediction and / or motion state estimation. According to the invention, high-precision and high-efficiency scene understanding can be realized in a complex environment.
Owner:张家港港务集团有限公司 +1

Multi-stage filtering road thrown object detection method based on dynamic difference analysis

The invention relates to a multi-stage filtering road spilled object detection method based on dynamic difference analysis, which is suitable for automatic identification of unstructured foreign matters in video monitoring. The method comprises the following steps: firstly, extracting a reference image road mask, eliminating vehicle and pedestrian interference by using YOLOv8 detection, and extracting a motion candidate area through a frame difference method and background modeling; and then context expansion and super-resolution reconstruction are carried out on the candidate region, the candidate region is converted into an HSV space, multi-dimensional features such as color similarity, structural similarity and shadow determination are synthesized for screening, false detection is further removed in combination with inter-frame time sequence consistency, and finally a stable detection result is output. The method provided by the invention has the advantages of strong anti-interference capability, high adaptability, high detection precision and the like, and is suitable for the intelligent recognition task of the expressway thrown objects in a complex environment.
Owner:CCCC HUAKONG (TIANJIN) CONSTR GRP CO LTD

Physical education evaluation system based on artificial intelligence

The invention relates to the technical field of physical education, and discloses an artificial intelligence-based physical education evaluation system, which comprises a data acquisition module, a multi-target detection module, a key point modeling module, an action trajectory analysis module, a multi-modal fusion module and a dynamic feedback module. The data acquisition module receives a video stream of the camera device and motion physiological data of the wearable sensor; the multi-target detection module locates a target based on a YOLOv8 algorithm; the key point modeling module constructs a three-dimensional action attitude model by using an OpenPose algorithm; the action trajectory analysis module evaluates action standard; the multi-modal fusion module integrates the data to construct a multi-dimensional feature matrix; the dynamic feedback module generates a real-time score and a personalized correction instruction through the optimized LSTM network. In addition, the system further comprises a model self-adaption module, an abnormal action recognition module and a distributed calculation module. According to the system, precise evaluation of physical education is realized, the sports safety is guaranteed, personalized teaching is promoted, and the physical education quality is effectively improved.
Owner:HUNAN EDUCATION AUDIO-VISUAL ELECTRONIC PUBLISHING HOUSE CO LTD

Segmentation-assisted detection and tracking of objects or features

Disclosed are apparatuses, systems, and techniques for segmentation-assisted detection and tracking of objects or features in videos, across images, and / or in other 2D and / or 3D visual content. The techniques include processing a plurality of frames of a video to obtain a plurality of representations of an object depicted in the video. A first subset of the plurality of representations is obtained by processing, using an object detection model, a first subset of the plurality of frames. A second subset of the plurality of representations is obtained using visual similarity of an appearance of the object in a second subset of the plurality of frames to the appearance of the object in at least one other frame of the plurality of frames. The techniques further include obtaining, using the plurality of representations, segmentation masks for the plurality of frames and performing one or more operations based on the segmentation masks.
Owner:NVIDIA CORP

Sea surface target tracking method and system based on infrared and visible light image fusion

The invention provides a sea surface target tracking method and system based on infrared and visible light image fusion, and relates to the technical field of ocean monitoring, and the method comprises the steps: collecting visible light and infrared image data of a sea surface target; performing feature extraction and fusion through wavelet transform fusion to generate a comprehensive feature map, and performing target recognition on the comprehensive feature map by using a deep learning target detection model; after target recognition, the system calculates the position of a target based on image data and radar data, performs multi-target matching and association through a Hungary algorithm combined with multi-modal features, predicts the position of the target and updates trajectory information in combination with a Kalman filtering or particle filtering algorithm. The infrared camera and the visible light camera carry out dynamic angle adjustment according to the position and the movement track of the target; whether light information correction is carried out or not is judged based on the light correction threshold value, when light information correction is carried out, light information correction features are constructed based on the visible light compensation model and the infrared compensation model through the image data, and information errors caused by the illumination angle are eliminated.
Owner:HARBIN INST OF TECH AT WEIHAI +1

Cross-modal remote sensing target detection method and system based on space consistency constraint and deep feature alignment

The invention discloses a cross-modal remote sensing target detection method and system based on spatial consistency constraint and deep feature alignment, belongs to the field of computer vision and remote sensing science and technology and the technical field of machine learning and deep learning, and solves the problem that a conventional cross-modal method does not fully consider feature hierarchy difference. According to the invention, an improved teacher-student network model is constructed and comprises a student branch network, a teacher branch network and an optimization module for performing pseudo-label optimization, non-monitoring learning and supervised learning on the student branch network and the teacher branch network; training the improved teacher-student network model by adopting the target domain data set and the source domain data set to obtain a trained improved teacher-student network model; and carrying out cross-modal remote sensing target detection on a to-be-detected target domain image by adopting the trained improved teacher-student network model. The method is used for cross-modal remote sensing target detection.
Owner:SOUTHWEST JIAOTONG UNIV

Anti-falling identification early warning method based on image identification and related equipment thereof

The invention relates to the technical field of image recognition, and provides an anti-falling recognition early warning method based on image recognition and related equipment thereof. The method comprises the following steps: carrying out multi-modal preprocessing on a real-time image acquisition data set to obtain an RGB-D data stream, carrying out target detection and three-dimensional attitude modeling on the RGB-D data stream to obtain a target personnel label set and a personnel attitude parameter set, and carrying out trajectory prediction on the personnel attitude parameter set through a physical kinematics model to obtain predicted motion trajectory data. Acquiring an inertial monitoring data set in real time according to the target person label set, performing multi-modal fusion in combination with the predicted motion trajectory data to obtain a confidence evaluation value, and performing risk quantification on the confidence evaluation value and the predicted motion trajectory data to obtain a graded early warning instruction. And performing protocol coding and signal conversion on the graded early warning instruction to obtain a control signal. According to the invention, through image identification, motion prediction, multi-modal data fusion and fine processing, the accuracy of anti-falling monitoring in a complex scene is improved.
Owner:FOSHAN CHANCHENG DISTRICT GLOBAL ELECTRICAL PORCELAIN ELECTRICAL MATERIALS CO LTD

Unmanned aerial vehicle lightweight target detection method based on improved YOLOv8n model

The invention discloses an unmanned aerial vehicle lightweight target detection method based on an improved YOLOv8n model, a new PCE module is proposed for the first time, a new feature extraction network is obtained, the feature extraction capability is improved by using partial convolution and an efficient multi-scale attention module, and the target detection accuracy is improved. In addition, a lightweight detection head based on a shared parameter strategy and a diversified branch block is designed, the model parameter quantity is greatly reduced while the high detection precision is maintained by sharing the first two layers of reparameterized convolution parameters of a classification head and a regression head of feature maps with different sizes, a WIoU loss function is used to balance high and low quality samples, and the detection accuracy is improved. The model prediction deviation is reduced, and the generalization ability is improved.
Owner:NANJING UNIV OF SCI & TECH

3D object detection using temporal inputs

Apparatuses, systems, and techniques of using one or more machine learning processes (e.g., neural network(s)) to detect objects from a plurality of image frames. In at least one embodiment, a plurality of image frames are fused into a feature map using one or more neural networks. In at least one embodiment, a plurality of image frames are processed using one or more neural networks to detect objects in a 3D space.
Owner:NVIDIA CORP

Laser - based targeting and object detection system

A pest control system is disclosed comprising an optical, computational, and monitoring subsystem, optionally mounted on a mobile platform. The optical system may include a neutralizing laser or multi-wavelength light source, discovery and detail cameras (optionally stereo), a beam-steering mechanism, tunable focus, and optional thermal or depth sensors. The processor, such as a GPU or FPGA, identifies insect or biological targets, adjusts laser focus by depth, and controls beam activation. A monitoring system verifies safety by detecting humans or other non-target entities using environmental and thermal cameras; if detected, laser firing is inhibited. The mobile platform may use wheels, propellers, tracks, or cables, with GPS and data links for remote control. A visible light pre-flash may induce a blink reflex before firing. In some embodiments, a scouting drone transmits target coordinates to the neutralization unit, enabling coordinated, efficient, and safe laser-based pest control.
Owner:REYNTJENS NICK

Medical image fuzzy boundary segmentation method based on edge perception Mama network

The invention discloses a medical image fuzzy boundary segmentation method based on an edge perception Mama network. The method aims at solving the problems that a camouflage pathological structure in a medical image is visually similar to surrounding healthy tissues, so that segmentation is difficult, and clinical deployment is difficult due to secondary calculation complexity of an attention mechanism in a traditional camouflage target detection (COD) method. According to the invention, a boundary guiding module inspired by COD is creatively combined with linear complexity state space modeling, and an E-Mama edge guiding framework is provided. The framework adopts an encoder-edge guide-decoder structure, and accurate boundary description is realized while the computational efficiency is kept. The system comprises five core components: a Mama encoder; an adaptive fusion processing module (AFM); an edge detection module (EDM); an edge guidance module (EGM); the invention relates to a Mama decoder. Experiments show that advanced segmentation precision is realized on a plurality of medical data sets, and meanwhile, the calculation complexity is reduced from O (n2) to O (n), so that clinical deployment becomes possible.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Multi-modal sensor-based detection and tracking of objects using bounding boxes

A perception system may be used to generate bounding boxes for objects in a vehicle scene. The perception system may receive images and feature maps corresponding to the received images. The perception system may correlate object queries from previous time steps with object queries from the current time step.
Owner:MOTIONAL AD LLC