Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1708 results about "Goal recognition" patented technology

Target identification tracking method and system based on multi-source fusion imaging

The invention discloses a target identification tracking method and system based on multi-source fusion imaging, and the method comprises the steps: obtaining and preprocessing multi-modal environment monitoring data, carrying out the image enhancement based on the preprocessed multi-modal environment monitoring data, and obtaining the image-enhanced multi-modal environment monitoring data; performing multi-modal feature extraction on the image enhancement multi-modal environment monitoring data, and introducing a cross attention mechanism to fuse the extracted multi-modal features to construct cross-modal fusion features; constructing a cascade target recognition model, and inputting cross-modal fusion features to perform target recognition on the current monitoring scene; tracking a scene recognition target in the current monitoring scene to generate scene target tracking information, and storing the scene target tracking information in a preset track memory pool; and when a tracking target disconnection condition occurs during target tracking, obtaining a mismatched tracking trajectory, and performing tracking trajectory disconnection repair in combination with the trajectory memory pool. Therefore, the identification accuracy and tracking reliability of the target in the monitoring scene are improved.
Owner:SHENZHEN PARD TECH CO LTD

Manipulator grabbing method based on deep learning target detection and image segmentation

The invention discloses a manipulator grabbing method based on deep learning target detection and image segmentation, and relates to the technical field of artificial intelligence and robotics.The manipulator grabbing method comprises the following steps that a scene image to be processed is collected, the image quality is improved through the multi-light-source fusion image enhancement technology, and recognition errors caused by uneven illumination are reduced; and inputting the enhanced image to a pre-trained deep learning model, executing a target detection task, and outputting an initial bounding box and a category label of the target object. According to the method, through multi-light-source image enhancement and high-precision image segmentation, the accuracy of target recognition and contour extraction is remarkably improved, and the capture failure rate caused by image misjudgment is reduced. And meanwhile, geometric consistency verification and a multi-factor grabbing scoring mechanism are introduced, dynamic screening and collision pre-detection are conducted on the paths, the grabbing stability and safety of the mechanical arm in the complex environment are effectively guaranteed, and the intelligence and robustness of the whole system are remarkably improved.
Owner:SHENZHEN BOCHUANG ROBOT TECH

Image segmentation and dynamic target identification method based on artificial intelligence

The invention relates to the technical field of artificial intelligence, in particular to an artificial intelligence-based image segmentation and dynamic target recognition method, which comprises the following steps of: accurately positioning a candidate region through multi-modal space-time fusion and dynamic confidence coefficient screening; strengthening spatial-temporal feature expression in a layering manner through a multi-level feature decoupler, and generating a multi-dimensional feature enhanced spatial-temporal candidate region; through a deformable segmentation network, a deformation convolution kernel and edge motion matching loss are combined, joint optimization of a geometric boundary and motion continuity is realized, and the segmentation robustness of a flexible target is improved; through optical flow back propagation dynamic correction and confidence coefficient propagation, high-precision segmentation masks with consistent time and space are output; and through a target trajectory re-identification and completion mechanism driven by a graph attention network, and in combination with optical flow deformation prediction, stable tracking in a shielding scene is realized.
Owner:CHANGSHA INSTITUTE OF TECHNOLOGY

Grabbing attitude generation method and system based on multi-modal large model

The invention discloses a grabbing posture generation method and system based on a multi-modal large model, and the method comprises the steps: carrying out the cross-modal matching of visual features and semantic features in the multi-modal large model when a voice instruction and an RGB image are inputted, and obtaining the position information of a control function code and a target object; when an RGB image with a hand drawing instruction is input, obtaining position information of a control function code, a target object and a path point; calculating the point cloud data of the target object according to the position information of the target object and the depth information, inputting the ideal point cloud of the target object into a target recognition network model after preprocessing, carrying out the grabbing region recognition of the point cloud of the target object region, outputting a region with high grabbing confidence, and mapping a real coordinate system; constructing a point cloud bounding box, and generating a grabbing posture candidate set; the grabbing posture with the highest quality is selected as the grabbing posture of the robot by calculating the grabbing posture candidate score; and executing a target grabbing task in combination with the control function code and the grabbing path.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

Multi-modal coupled perception method for target recognition and region segmentation in confined space

The present invention relates to the technical field of environmental perception in confined spaces, and discloses a multi-modal coupled perception method for target recognition and region segmentation in a confined space. The method comprises: firstly, constructing a 3D point cloud data processing network, a 2D point cloud data processing network, and an image data processing network; next, constructing a feature fusion device, and outputting a bird's-eye view incorporating multi-sensor information; then constructing a multi-scale feature fusion extraction network and a network output head, outputting a target recognition and region segmentation result, and designing a loss function to train a network weight; and finally, inputting 3D point cloud data, 2D point cloud data, and camera image data into a network model, performing inference to obtain a target recognition and region segmentation prediction result, and performing visual rendering on the prediction result. In the present invention, in view of the characteristics of a millimeter-wave radar, a multi-modal coupled perception network is constructed, so that effective information in mass point cloud data of the millimeter-wave radar can be deeply mined, thereby effectively improving the detection accuracy of target recognition.
Owner:CHINA UNIV OF MINING & TECH

SAM2 multi-task perception binary segmentation method based on hybrid expert adapter

The invention discloses an SAM2 multi-task perception binary segmentation method based on a hybrid expert adapter. The method is specifically implemented according to the following steps: step 1, constructing a data set and an encoder; step 2, constructing a hybrid expert adapter module; and step 3, constructing a task awareness gating module. According to the method, a pre-trained SAM2 is taken as a main network, on the basis of freezing the main body weight, a standard adapter and a MoE-Adapter are respectively deployed on odd and even layers of an encoder, and a lightweight expert sub-network and a dynamic gating strategy are combined, so that unified processing of multiple tasks such as salient target detection, camouflage target identification, marine animal segmentation and the like is realized.
Owner:XIAN UNIV OF TECH

Unmanned aerial vehicle cruising method and system based on deep learning artificial intelligence image recognition algorithm

The invention discloses an unmanned aerial vehicle cruising method and system based on a deep learning artificial intelligence image recognition algorithm. According to the method, an unmanned aerial vehicle carrying an improved YOLOv7-SwinT target recognition model collects real-time image data of an inspection area, and the model fuses a single-stage target detection architecture of YOLOv7 and a visual feature extraction network of Swin Transform. Progressive target detection is realized by adopting a three-level recognition architecture, wherein the progressive target detection comprises primary anomaly detection based on lightweight CNN, intermediate accurate positioning in combination with an attention mechanism and advanced target classification of multi-sensor data fusion. And the system combines the electric quantity of the unmanned aerial vehicle, the environmental condition and the task priority according to the identification result, generates a dynamic inspection path through an adaptive path planning algorithm, and realizes multi-vehicle collaborative operation by using an intelligent task allocation algorithm. In the inspection process, sensor data are processed in real time through edge computing equipment, and charging scheduling is optimized by adopting an intelligent energy management system. The target recognition precision and the cruising efficiency of the unmanned aerial vehicle in a complex environment are remarkably improved, and the method is suitable for application scenes such as electric power inspection and security monitoring.
Owner:NAT ENERGY GRP DONGTAI OFFSHORE WIND POWER CO LTD

Low-altitude monitoring system based on multi-source heterogeneous sensing fusion and rule engine driving

The invention relates to the field of low-altitude airspace monitoring, and relates to a low-altitude monitoring system based on multi-source heterogeneous perception fusion and rule engine driving, which comprises a perception access layer used for completing data standardization and protocol adaptation and realizing data stream release through message middleware; the processing calculation layer is used for generating a structured alarm event; the application support layer is used for three-dimensional visual interactive display of target situation data and alarm events, unified access of multiple types of sensors such as radar, photoelectricity, radio frequency, ADS-B and meteorological sensors is realized through a plug-in protocol adapter, and target state estimation and track smoothing are realized by combining space-time registration and extended Kalman filtering. The limitation that an existing system can only process a single data source or simple superposition is overcome, and the target recognition precision and the track continuity are remarkably improved.
Owner:ZHONGKE BRILLIANT ROBOT (CHENGDU) CO LTD

Family service method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence, in particular to a family service method and system based on artificial intelligence, and the method comprises the following steps: obtaining a statement analysis action sequence and a task direction, collecting the feedback of a device to generate a state structure, judging an instruction trend, recombining a statement to generate a behavior chain, and disassembling an action recognition conflict to extract a main control path. Analyzing a behavior time sequence generation prominent trend, and adjusting a display sequence to generate a function scheme. According to the method, through semantic component extraction and equipment state mapping, the accuracy of instruction target recognition is improved, word order rearrangement and logic connection enhance the continuity of a behavior chain, a control path is optimized and sorted according to verb density and object association, the task scheduling precision is improved, and behavior time sequence analysis recognizes a high-frequency operation forward trend; according to the method, priority dynamic adjustment and rearrangement of recommended content in combination with instruction frequency and time sequence characteristics are realized, the response initiative and the service matching degree are enhanced, and collaborative optimization of semantic understanding, behavior prediction and content recommendation is integrally realized.
Owner:HUNAN UNIV

Aliasing image intelligent deformation small target identification algorithm based on detector multiplexing

The invention relates to an aliasing image intelligent deformation small target identification algorithm based on detector multiplexing, and relates to the technical field of aliasing image target identification, and the algorithm comprises the steps: designing a detector multiplexing optical system; designing a view field coding element, and splicing the whole system; simulating and generating a motion track of the early warning target under the space-based background; the target motion trail is imaged through a detector multiplexing optical system; aliasing video simulation of moving target imaging is carried out; an intelligent deformation small target recognition algorithm is applied to the aliasing image, and a moving target is recognized; target resolution and field-of-view positioning are realized by using the relevance between the inter-frame track of the moving target and the light spot shape. According to the aliasing image intelligent deformation small target recognition algorithm based on detector multiplexing, light originally imaged on a large-area-array detector is folded and imaged on a small-area-array detector, and therefore the purpose that the small-area-array detector receives large-view-field imaging is achieved.
Owner:CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Multi-modal ship target individual identification method and system

The invention discloses a multi-modal ship target individual identification method and system, and relates to the technical field of ship target identification, and the method comprises the steps: obtaining multi-modal data of a target region, and constructing a training set and a test set; each training sample in the training set is input into a target individual prediction model, a multi-granularity adaptive loss function is adopted as a loss function to carry out model training, and the target individual prediction model adopts an image feature extraction method based on cross-modal guide enhancement to carry out feature extraction and splicing fusion on each training sample; and a target individual identification result of each training sample is obtained by adopting a visual angle self-adaptive fusion method of embedded type label information. And after training is completed, a trained target individual prediction model is obtained, each test sample in the test set is input into the trained target individual prediction model, and a final target individual identification result is obtained. According to the invention, the generalization ability of model learning individual features can be improved, and the accuracy of ship target individual identification is improved.
Owner:NAVAL AVIATION UNIV

Material tracking multi-target identification method and system based on 3D +2D high-dimensional feature fusion

The invention provides a 3D + 2D high-dimensional feature fusion-based material tracking multi-target identification method and system. The method comprises the steps of performing target identification on a to-be-tracked material by using a YOLO algorithm; performing multi-target real-time tracking on the material based on a Deep SORT algorithm; generating three-dimensional point cloud data of a target by adopting a multi-view data calibration and alignment algorithm, and carrying out noise reduction, sampling and data enhancement preprocessing operation; performing three-dimensional reconstruction on the preprocessed three-dimensional point cloud data by using a Gaussian Splitting method, and introducing time sequence information at the same time; segmenting a dynamic three-dimensional target and a static background through a Mask R-CNN technology; aligning the position of the target in the two-dimensional image with the coordinates in the three-dimensional space by using a feature fusion method; a PNP algorithm is adopted to estimate attitude and depth information of a target in a three-dimensional space, accurate positioning and state restoration of materials are realized, and efficient and reliable technical support is provided for material management and tracking in an industrial environment.
Owner:UNIV OF SCI & TECH BEIJING

Anti-collision monitoring system based on depth estimation and instance segmentation fused three-dimensional model

The invention discloses an anti-collision monitoring system based on a depth estimation and instance segmentation fused three-dimensional model, particularly relates to the field of intelligent driving security and protection, is used for solving the problem of target recognition and anti-collision in an environment, realizes efficient feature extraction of global semantics and local textures through a multi-branch network structure, and ensures the accuracy and comprehensiveness of target segmentation. In combination with a cross-frame identity association and conflict detection mechanism, the consistency of target identities is effectively maintained, and the mismatching problem caused by shielding or similar appearances is reduced; multi-scale space-time shielding mode analysis and boundary motion consistency verification are adopted, a fusion algorithm is utilized to comprehensively evaluate shielding complexity and segmentation robustness, a high-risk shielding area is identified, secondary difference correction is executed, and space positioning information of a target is remarkably optimized; and finally, integrating the optimized segmentation and depth data into a global coordinate system, constructing a high-precision and coherent three-dimensional semantic scene, and providing stable and high-quality input data for collision detection and trajectory prediction.
Owner:CHINA AVIATION PLANNING AND DESIGN INSTITUTE (GROUP) CO LTD

Target intelligent collaborative identification method based on unmanned aerial vehicle cluster

The invention discloses a target intelligent cooperative identification method based on an unmanned aerial vehicle cluster, and belongs to the field of unmanned aerial vehicle cluster control and computer vision. According to the method, cluster networking and model initialization are realized through a dynamic heterogeneous federated learning architecture; a space-time attention mechanism is adopted to optimize task allocation, and a deformable network is utilized to extract multi-view target features; a cascade characteristic distillation fusion strategy is provided, and modal compression and cross-modal gating fusion are carried out on multi-source data such as multispectral data and laser radar data; an anti-interference elastic communication mechanism based on meta-learning is designed, and the system robustness is enhanced by combining space-time confrontation detection and a dynamic spectrum sensing technology; an unsupervised federal incremental learning system is established, and online evolution of the model is realized through momentum weighted aggregation. According to the method, the target identification accuracy is improved by 35% in a complex environment, the time delay is reduced to 200 ms, and high-precision real-time identification support is provided for military reconnaissance, disaster rescue and other scenes.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Intelligent driving target identification method based on polarization visual image fusion

The invention discloses an intelligent driving target identification method based on polarization visual image fusion. The method comprises the following steps: step 1, preprocessing an acquired polarization image; 2, constructing and training a polarization image fusion network based on deep learning; step 3, inputting the polarization image into the trained fusion network to generate a high-quality fusion image; and step 4, inputting the fused image into a target identification model to realize identification of targets such as vehicles and pedestrians. The method gives consideration to detail enhancement and color fidelity of the polarization image, can improve the accuracy and robustness of target recognition in a complex scene, and has a wide application prospect in the fields of intelligent driving and the like.
Owner:NORTHWEST A & F UNIV

Video target identification method and device based on artificial intelligence, and storage medium

The invention relates to the technical field of image recognition, and provides a video target recognition method and device based on artificial intelligence and a storage medium, and the method comprises the steps: obtaining a video frame data sequence of a target video, carrying out the multi-dimensional analysis of the video frame data sequence, obtaining global video frame information and local region-of-interest information, and generating a target feature map based on the global video frame information and the local region-of-interest information, then carrying out time sequence mode analysis to obtain time sequence evolution features, combining the generated semantic representation vector, inputting the semantic representation vector into a preset adaptive Transform model to carry out target recognition, and obtaining a target recognition result. A semantic representation vector is generated through feature fusion, and a self-adaptive Transform model is used for target recognition, so that the target recognition precision in a complex scene is improved, and the problems of low detection precision and insufficient time sequence information utilization during complex scene processing, dynamic change and long-time sequence analysis are solved.
Owner:HOHEM TECHNOLOGY CO LTD

Environment detection method and system based on multi-modal data fusion and deep learning

The invention provides an environment detection method and system based on a sample target detection model. The method comprises the following steps: synchronously acquiring an environment image, a video stream and physical parameters by using a multi-mode sensor; decomposing the data into image features and environmental parameter components through a dual-time sequence control signal, and realizing space-time alignment by adopting a linear phase filter; constructing a foreground region template based on the depth information, and generating target recognition feature representation containing an abnormal blurred target; adversarial training is carried out on the lightweight target detection network in combination with a transfer learning strategy, the network integrates convolutional features and a Transform attention mechanism, and the weight is dynamically adjusted through environmental parameters; fusing a target result and sensor data in real-time detection, and inputting a decision tree model for risk grading; and after the early warning is triggered, reconstructing a false detection sample through an online learning mechanism and iteratively optimizing the model. The system correspondingly comprises a multi-modal data acquisition module, a data enhancement and annotation module, a model training module, a real-time detection and fusion module and an early warning and optimization module. According to the invention, through multi-source data fusion, dynamic data enhancement and an adaptive compensation mechanism, the small target detection precision, the environmental adaptability and the real-time early warning capability are significantly improved.
Owner:SHANDONG HUANFA INSPECTION & TESTING CO LTD

Interaction method and device based on image recognition

The invention provides an interaction method and device based on image recognition, and relates to the technical field of man-machine interaction, and the technical scheme is characterized in that the method comprises the steps: analyzing the moving speed and change frequency of target and non-target elements in a training scene, and calculating an environment dynamic factor; adjusting a high threshold value and a low threshold value of target identification confidence; comparing the target recognition confidence output by multi-modal image recognition with the adjusted high threshold and low threshold, and determining a visual feedback type; setting a visual feedback intensity parameter according to the determined visual feedback type and the target type; and displaying the set visual feedback in the visual field of the technical object in an overlapping manner. The interaction method and device based on image recognition provided by the invention have the advantages that the cognitive load of the technical object is reduced, and the rapid target recognition capability of the technical object in a complex and uncertain battlefield environment is improved.
Owner:BEIJING ZHIFENG TECH CO LTD

Character recognition method and device, computer equipment and storage medium

The invention provides a character recognition method and device, computer equipment and a storage medium. According to the scheme, the slice scanning image is firstly acquired, then the label area is identified according to the slice scanning image, and the character information is close to the label, so that the determined target area is larger than the area of the label so as to cover the character information. And then the slice scanning image is intercepted according to the target area, and a target image only containing key information is obtained. And character recognition is carried out on the target image to obtain a character recognition result. And finally, the character recognition results are screened by using a rule base, and a target recognition result is screened out. According to the scheme, through tag region identification and target region interception, the character identification range is narrowed, the identification efficiency is improved, and the waste of computing resources caused by processing a large amount of irrelevant information is reduced. The rule base can better adapt to complex slice images and constantly changing requirements, information needed by a user is accurately screened out, and the character recognition efficiency is greatly improved.
Owner:MOTIC CHINA GROUP CO LTD

Three-dimensional shielded target tracking method based on multi-modal space-time interaction

The invention discloses a three-dimensional shielding target tracking method based on multi-modal space-time interaction, and relates to the technical field of target tracking. The method comprises the following steps: acquiring a point cloud and an image and preprocessing to obtain global fusion features; obtaining an initial detection frame and region-of-interest features through region proposal network processing; projecting the non-empty voxel point cloud to the image features, and reconstructing shielded target features; convolution and neural network processing are utilized to obtain a refined detection frame; screening legal detection frames through distance calculation and legality judgment; the bipartite graph and the self-adaptive channel graph are adopted for convolution, and appearance correlation scores are calculated; and matching the detection frame and the trajectory based on a Hungary algorithm to realize whole-course tracking. The target identification accuracy and robustness are improved, the shielding problem is solved, the accuracy of the detection frame is ensured, the correlation accuracy is improved by using the bipartite graph and the adaptive convolution, the nodes are matched in combination with the geometric cost matrix, and whole-course tracking and error calibration are realized.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Object monitoring method and device based on multi-source data and storage medium

The invention relates to the field of image processing, and discloses an object monitoring method and device based on multi-source data and a storage medium, and the method comprises the steps: synchronously collecting multi-channel video streams, environment parameters and equipment position information, and carrying out the time alignment and data association processing, and forming associated data; performing feature matching on the multiple paths of video streams, fusing position information and environment parameters, and establishing a mapping relation between video feature points and a unified space coordinate system; splicing the multiple paths of video streams in real time according to the mapping relation to generate a panoramic video stream; performing dynamic scene analysis based on the panoramic video stream and the environmental parameters, and identifying a target type and a state to obtain an analysis result; and generating an equipment regulation and control strategy based on the analysis result, and generating and issuing a regulation and control instruction for adjusting the working parameters of the multi-view image acquisition device based on the equipment regulation and control strategy. According to the invention, the splicing precision, the real-time performance and the target identification accuracy of panoramic monitoring can be improved, and the dynamic adaptive regulation and control of the equipment can be realized.
Owner:SHENZHEN STARCAM TECH

Multi-task parallel processing method and system based on AI target identification

The invention provides a multi-task parallel processing method and system based on AI target recognition, and the method comprises the steps: carrying out the target recognition analysis of input multi-mode data, obtaining an initial task list, and generating a parallel recognition task sequence through combining the context information of a system; performing adaptive feature extraction and optimization on the multi-modal data according to the parallel recognition task sequence to obtain a task specific feature set; based on the task specificity feature set and the system context information, optimizing a task execution sequence and resource allocation by using a deep reinforcement learning model to obtain a preliminary optimization scheme; performing task preheating and dynamic task merging on the preliminary optimization scheme to obtain and execute an optimization task execution scheme. According to the method, the processing strategy can be dynamically adjusted according to the multi-modal data features and the task requirements, the task execution sequence and resource allocation are effectively optimized, and the adaptability and efficiency of a target recognition system in a complex industrial environment are improved.
Owner:SHENZHEN LEKE INTELLIGENT CONTROL TECH CO LTD

Target detection method and system based on millimeter wave radar

The invention discloses a target detection method and system based on a millimeter wave radar. The target detection method and system are used for realizing accurate detection, positioning and dynamic and static recognition of multiple targets under a complex background. According to the method, a distance-Doppler spectrogram is generated through the technical means of sliding window construction, spectral analysis, clutter suppression and the like, and candidate target points are detected by adopting an SO-CFAR algorithm. Then, determining a target position through high-resolution direction estimation and coordinate transformation, performing spatial clustering in combination with a density-based DBSCAN algorithm, and extracting a target geometric center and a bounding box; in the aspect of target tracking, Kalman filtering is used for predicting and updating the position and speed of the target, and a beam forming technology is used for enhancing a target signal, so that the target recognition stability is improved. And finally, the system performs robust dynamic and static state recognition on the target through a dynamic and static judgment module, so that high precision and robustness of the target detection process are ensured. The method can effectively cope with static background interference and dynamic target changes, and is suitable for target detection and tracking in a complex environment.
Owner:HANGZHOU DIANZI UNIV

Target tracking method fusing visual perception and motion control

The invention discloses a target tracking method fusing visual perception and motion control. The method comprises the following steps: improving a YOLOv8 model and pre-training the improved YOLOv8 model; s2, preprocessing an image acquired by a camera, inputting the preprocessed image into the YOLOv8 model trained in the step S1, and outputting a target bounding box and a target center pixel coordinate; target center pixel coordinates are transmitted to a development board of a pre-burning algorithm program through a serial port, the development board is designed according to a stepping motor mathematical model, sine acceleration and deceleration serve as the basis, PID control, sliding mode control and fuzzy control algorithms are fused, and code burning is generated after in-loop simulation optimization; and the development board drives a stepping motor to adjust the posture of the holder according to the target center pixel coordinate control instruction, so that the target stops moving after being kept in the view center. The problems that a traditional YOLOv8 model is large in target recognition error under the complex background, PID control is prone to oscillation under the nonlinear working condition, buffeting exists in sliding mode control, and response lags when fuzzy control is independently used are solved.
Owner:SHAANXI UNIV OF SCI & TECH

Interaction method and device for upper limb cooperative control of humanoid robot

The embodiment of the invention provides an interaction method for upper limb cooperative control of a humanoid robot. The interaction method comprises the following steps: performing task analysis on a received voice request; if the target object needing to be executed exists, target object recognition processing is carried out according to the task analysis result; based on target identification information and target pose information output by identification processing and a task analysis result, performing task planning processing by utilizing a large language model according to a double-arm division cooperation principle, and generating a task sequence and a character instruction corresponding to each task; sequentially executing the following operations on each target task in the task sequence: based on the target pose information and the target identification information, controlling upper limbs of the humanoid robot to execute an operation action corresponding to the target task on the target object; therefore, on the basis of the large language model, the visual module and the voice module are matched, so that a user can communicate with the robot more visually; and the operation capability of the upper limbs of the humanoid robot for executing complex tasks can be improved.
Owner:江淮前沿技术协同创新中心

Photoelectric pod target identification and tracking system based on multi-scale attention mechanism

The invention provides a photoelectric pod target identification and tracking system based on a multi-scale attention mechanism, and belongs to the technical field of intelligent vision. Through combination of a multi-scale convolution module and a multi-head self-attention mechanism, accurate detection and tracking of a target in a photoelectric pod video image are realized. The multi-scale convolution module adopts convolution kernels of different sizes, and can extract local features of different scales to adapt to the change of the size of a target; the multi-head self-attention mechanism is used for capturing global features, especially long-distance dependency relationships between targets and backgrounds and between targets. Through fusion of local features and global features, the system improves the precision and robustness of target recognition in a complex scene. Meanwhile, by optimizing the structural design and the feature aggregation method, the calculation complexity of the system is remarkably reduced, and the requirements of the photoelectric pod for real-time performance and high efficiency are met.
Owner:GUANGDONG UNIV OF TECH

Visible light, infrared and IQ signal fusion individual identification method based on cross-modal cross attention

The invention discloses a visible light, infrared and IQ signal fusion individual identification method based on cross-modal cross attention, and belongs to the technical field of artificial intelligence and multi-modal image processing. Aiming at the problems of insufficient multi-modal heterogeneous feature fusion capability, unbalanced modal semantic expression and unstable classification precision in the prior art, a visible light image, an infrared image and an original IQ signal are acquired, and after preprocessing, a convolutional neural network is combined with a space attention module to extract image features; using a convolutional hybrid network to extract signal spectrum features; three groups of modal pairs are constructed by adopting a cross-modal bidirectional cross attention mechanism to carry out bidirectional semantic interaction and feature fusion; and finally, inputting the fusion features into a classifier to obtain an identification result. According to the method, deep semantic fusion can be realized, modal quality changes can be dynamically adapted, and the precision and robustness of target recognition in a complex environment are improved.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Target identification method, system and device for complex scene and medium

The invention discloses a complex scene-oriented target identification method, system and device and a medium, and relates to the technical field of computer vision. The method comprises the following steps: inputting a to-be-detected image into a target recognition model for processing to obtain a target recognition result; the target identification model is constructed based on a dual-channel coding network and a weighted optimization loss function; the dual-channel coding network comprises a MobileNet channel and a wavelet channel; wherein the MobileNet channel is used for extracting multi-scale spatial features and is composed of a layered encoder and a multi-feature prediction decoder which are connected in sequence; the wavelet path is used for extracting frequency domain features and is composed of a global semantic wavelet coding module and a local wavelet fusion decoding module which are connected in sequence; and the multi-feature prediction decoder is also respectively connected with the global semantic wavelet coding module and the local wavelet fusion decoding module. According to the invention, the target identification precision and accuracy for complex scenes can be improved.
Owner:CHINA CRIMINAL POLICE UNIV

Robot dynamic grabbing method and system based on multi-modal data fusion

The invention provides a robot dynamic grabbing method and system based on multi-modal data fusion, and the method comprises the steps: carrying out the multi-modal fusion of the collected visual data and tactile data of a to-be-grabbed dynamic object through a robot, and obtaining the multi-modal fusion data; according to the multi-modal fusion data, target recognition is conducted on the dynamic object to be grabbed through a target detection algorithm based on regional proposal, and position and posture information is obtained; based on the position and attitude information, state prediction is conducted on the dynamic object to be grabbed, and state prediction information is obtained; according to the state prediction information, utilizing a reinforcement learning algorithm to generate a grabbing path of the robot; according to the method, the collected visual data and tactile data are subjected to multi-modal fusion, so that the robot can obtain comprehensive environment information, and the adaptability of the robot to environment change is improved; the state of the to-be-grabbed object is predicted, so that the robot can pre-judge the motion trail of the object, and the grabbing success rate of the robot can be improved.
Owner:BEIJING HANXINSHENG TECH CO LTD

Underwater sonar target identification system based on multi-domain feature fusion and lightweight modeling

The invention discloses an underwater sonar target recognition system based on multi-domain feature fusion and lightweight modeling, and the system comprises a Trifusion block, a novel lightweight attention residual network, a long and short time attention LSTM and a Mamba module which are connected in sequence, and achieves target recognition through multi-domain parallel extraction of fusion features, lightweight convolution and attention optimization, long and short time dependence capture and long sequence modeling. The method has the advantages that complex noise is comprehensively represented, the model efficiency and stability are improved, the bottleneck of time sequence processing is broken through, and the method has high performance, high robustness and wide adaptability.
Owner:GUILIN UNIV OF ELECTRONIC TECH