Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

908 results about "Scene graph" patented technology

A scene graph is a general data structure commonly used by vector-based graphics editing applications and modern computer games, which arranges the logical and often spatial representation of a graphical scene.

Edge-deployed semi-supervised anomaly detection method and system for railway track foreign object

Disclosed in the present invention are an edge-deployed semi-supervised anomaly detection method and system for a railway track foreign object. The method comprises the following steps: an edge device encoding and decoding a video stream captured by a camera to obtain an image frame sequence, and performing frame extraction; and using a semantic segmentation model to perform image segmentation on a certain image frame obtained by means of frame extraction, to obtain a railway track region segmentation image. The use of a single image as input may generate an expert model result having a high weight value; however, the determination based on a single image is not stable, multiple consecutive images of the task scene need to be inputted, the frequency of each expert model obtaining the highest weight is computed, and the expert model corresponding to the highest frequency is the final solution. The present invention supports scene-adaptive foreign object detection algorithm automatic selection, and a user can perform selection on the basis of prior knowledge, or selection may be performed by a scene-adaptive automatic algorithm selection method; the user only needs to provide a batch of image data of the current scene, and the optimal algorithm selection can be evaluated.
Owner:GUANGZHOU EMBEDDED MACHINE TECH CO LTD

Systems and methods for multi-modal visual reasoning using multiple scene graphs

A system for processing multi-modal data representing an environment to generate scene graphs of the environment is described. The system can obtain sensor data associated with a vehicle operating in the environment. In examples, the system can determine a set of features from the sensor data, including one or more objects and one or more agents present in the environment, and can generate a scene graph that represents the poses and velocities of these objects and agents relative to the environment. In some examples, based on generating the scene graph, the system can generate a knowledge graph by encoding the relationships among the identified objects and agents. In some examples, the system can generate a control signal, using attributes that represent the states of objects and agents in the knowledge graph, and provide this control signal to the vehicle in order to adjust or cause the operation of the vehicle.
Owner:QPIAI INDIA PTE LTD

Diffusion model-guided training of generative models for rendering novel views of 3D scenes

Approaches presented herein provide for the training and use of generative models to generate high quality image data for novel reconstruction views. A generative model such as a neural radiance field (NeRF) can be trained to generate such content. In order to train the NeRF to represent a specific scene with high accuracy, the NeRF can be trained using a diffusion model for score distillation guidance. The diffusion model can be trained using a large set of environment data from a variety of different views, then fine-tuned for a specific domain and / or scene. A parameter-efficient training process can be used to avoid overfitting of the diffusion model to the domain- or scene-specific training data. Once fine-tuned, the “expert” diffusion model can be used with the NeRF during training to effectively transfer the expert knowledge to the NeRF, enabling the NeRF to generate high quality image data for the scene from viewpoints corresponding to extreme novel views.
Owner:NVIDIA CORP

High-place operation risk real-time judgment method based on image and point cloud bimodal data

The invention discloses a high-altitude operation risk real-time judgment method based on image and point cloud bimodal data, and belongs to the crossing field of computer vision, artificial intelligence and engineering construction safety, and the method comprises the following steps: S1, obtaining a bimodal data stream through a laser radar and an RGB camera; s2, fitting an external reference registration matrix by acquiring an internal reference matrix of the camera to obtain a calibration parameter set; s3, optimizing the cue word, and driving the SAM-2 segmentation model to obtain a two-dimensional mask; s4, performing reverse projection on the two-dimensional mask to obtain a color point cloud instance set; s5, calculating the three-dimensional geometric centroid to obtain the three-dimensional geometric centroid and a three-dimensional semantic topology scene graph; and S6, obtaining a plurality of security rule compliance results by executing the interpretable rule engine. By the adoption of the method, the defects of an existing high-place operation risk judgment technology in the aspects of complex environment adaptability, multi-modal fusion efficiency, three-dimensional space perception and risk judgment interpretability are overcome.
Owner:SOUTH CHINA UNIV OF TECH

Instrument monitoring method, system and equipment for inspection robot and medium

The invention relates to the technical field of robots, in particular to an instrument monitoring method, system and device for an inspection robot and a medium, the method is used for searching an instrument to be inspected by using the inspection robot and monitoring the instrument to be inspected, and the method comprises the following steps: obtaining a view scene image collected by the inspection robot, performing image segmentation and feature extraction on the view scene image; obtaining the exploration value of the instrument area; obtaining candidate paths from the inspection robot to the instrument areas; evaluating a cost function of each candidate path; carrying out adaptive visual angle adjustment and data acquisition on the inspection instrument; the monitoring method is an exploration type sensing strategy based on exploration value driving, and an instrument target can be found in the environment. And in combination with a historical task map, instrument semantic priori and current perception data, the information value of the candidate region is evaluated, and optimal decision making is performed on the path by fusing environmental factors such as illumination, shielding and equipment states, so that priority selection and dynamic reconstruction of the target region are realized.
Owner:SICHUAN ENVIRONMENTAL PROTECTION ENG CO LTD CNNC

High-precision map reconstruction method and system based on monocular vision, medium and equipment

The invention belongs to the technical field of robot positioning and three-dimensional mapping, and discloses a high-precision map reconstruction method and system based on monocular vision, a medium and equipment. Scene image data are acquired through a monocular image acquisition module, after feature extraction and matching are performed on each frame of image, attitude information acquired by an inertial measurement unit is fused, and pose calculation of a robot is completed through a sparse vision SLAM system; meanwhile, an image dense depth map is generated by a monocular depth estimation model based on an attention mechanism, and the image dense depth map is converted into a single-frame color dense point cloud in combination with image RGB color information. According to the invention, a loose coupling fusion strategy is adopted to carry out spatial registration on a single-frame colored dense point cloud and a robot pose at a corresponding moment, multi-frame data fusion is completed through point cloud splicing and optimization, and a globally consistent three-dimensional dense point cloud map is constructed to realize scene modeling. The method has the characteristics of high robustness and high reconstruction precision, and can be effectively applied to three-dimensional map construction in an outdoor complex environment.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Rendering Video Of A Scene Using Three-Dimensional Gaussians

A set of images of a scene re received. Each image includes temporal data and spatial data relating to the scene. Based on the spatial data of each image, three-dimensional (3D) Gaussian splatting data is generated. The temporal data of each image and the 3D Gaussian splatting data are inputted to a neural network to generate spatial-temporal 3D Gaussian embeddings. Offset data based on the spatial-temporal 3D Gaussian embeddings is generated. The video of the scene is rendered based on the 3D Gaussian splatting data and the offset data, allowing for improved rendering of video of the scene.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Dynamic visual target motion tracking control method and system based on deep learning

The invention relates to the technical field of dynamic visual target motion tracking control, in particular to a dynamic visual target motion tracking control method and system based on deep learning, and the method comprises the steps: synchronously collecting continuous multi-frame target scene image data through a visual multi-frame collection module; and performing time sequence association and memory fusion on target features in continuous multi-frame target scene image data through a cross-frame feature memory fusion module, and constructing a target feature model. According to the invention, the current and historical stable features are dynamically fused through the cross-frame feature memory fusion module, time sequence association is realized in combination with the long and short-term memory network, and the problem of slow feature model updating under target deformation and shielding is solved; the deformation-shielding bimodal recognition module accurately recognizes a scene state, provides a basis for the multi-branch Kalman filtering prediction module, enables the multi-branch Kalman filtering prediction module to call a corresponding branch, corrects a prediction equation through a compensation factor, and improves the position prediction accuracy.
Owner:FUZHOU UNIV

Target detection method and device in wide-area complex scene, equipment and storage medium

The invention relates to a target detection method and device in a wide-area complex scene, equipment and a storage medium, and the method comprises the steps: carrying out the multi-level feature extraction of a wide-area complex scene image through employing a backbone network, and obtaining a low-level feature, a middle-level feature and a high-level feature; encoding the high-level features by adopting an EDFPT module to obtain encoded features; inputting the low-level features, the middle-level features and the coding features into a feature fusion module to obtain fusion features; screening a fixed number of image features from the fusion features by adopting an IoU-perceived query selection strategy to obtain an initial query vector; and processing the initial query vector by adopting a decoder with an auxiliary prediction head to obtain a wide-area complex scene image target detection result. The method has higher feature sensitivity to dense small targets and special-shaped targets in a wide-area complex scene image, the detection precision is effectively improved while the light weight of the model is ensured, and the omission ratio is reduced compared with RT-DETR.
Owner:NANCHANG UNIV +1

Personal question and answer method based on observation-recording-decision-making mechanism

The invention provides a personal question and answer method based on an observation-recording-decision mechanism, which comprises the following steps: an intelligent agent captures RGB images and depth images from all directions through multi-view perception, is used for constructing a 3D scene graph and mapping the 3D scene graph to a 2D semantic map, and meanwhile, the intelligent agent marks each passing position on the 2D semantic map, so that the 3D scene graph is mapped to the 2D semantic map; and dynamically updating the weight of a non-visited position, reducing the selection probability of a boundary point in a marked passing region, based on a 2D semantic map in an observation stage, distinguishing whether a question can be answered or not by an intelligent agent according to an observed RGB picture, if so, directly generating a response, otherwise, navigating to a new region, and repeating the steps until an available or maximum step number is reached, and completing the question answering. According to the method, through construction of a semantic map, weight regulation and control navigation, and fusion design of special VLM analysis and double-criterion decision making, decision making of agent non-redundancy exploration and accurate question and answer is achieved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-degree-of-freedom mechanical arm obstacle avoidance path planning method based on industrial vision

The invention discloses a multi-degree-of-freedom mechanical arm obstacle avoidance path planning method based on industrial vision, and relates to the technical field of intelligent manufacturing, and the method comprises the steps: collecting and preprocessing an initial environment image of a current working scene of a mechanical arm, obtaining a standardized working scene image data set, carrying out dynamic object detection, and obtaining an environment perception evaluation report; identifying an unobserved area of the working scene based on the environment perception evaluation report, performing collision path calculation, and generating a visual angle adjustment action instruction; the safety path sequence is issued to a mechanical arm joint and executed, images in front of the mechanical arm are continuously collected during execution, and if an unpredicted sudden obstacle is detected, the four-dimensional risk map is dynamically updated, and path re-planning is conducted; and when the tail end of the mechanical arm successfully completes the task, obstacle avoidance path planning data of the mechanical arm are recorded and stored persistently, and an obstacle avoidance path planning record is generated. According to the method, the real-time problem of path planning of the mechanical arm is solved through construction of the four-dimensional dynamic risk map.
Owner:KUNSHAN GANYUAN KANGSHENG TECHNOLOGY CO LTD

Indoor scene image generation method and device, equipment and storage medium

The invention discloses an indoor scene image generation method and device, equipment and a storage medium. The method comprises the following steps: firstly, acquiring a to-be-processed image with furniture and information input by a user, and performing furniture elimination processing on the image to obtain an empty room image; and then generating a cue word for image generation based on a preset cue word template in combination with user input information. Meanwhile, extracting a depth map and a semantic segmentation map of the empty room image, and inputting the depth map and the semantic segmentation map into a pre-trained depth map control model and a pre-trained semantic segmentation control model for feature coding to obtain corresponding depth control features and semantic segmentation control features. And performing parameter adjustment on the large-scale graph generation model subjected to parameter fine adjustment by using the two types of control features to obtain a target graph generation model. And finally, inputting the empty room image and the cue word into the target image generation model to obtain a target image. According to the method, the spatial structure stability and the personalized design requirement of the user on the indoor layout can be considered.
Owner:BEIJING CALF INTERNET TECH CO LTD

Robotic task completion from natural language requests

A computer-implemented method, apparatus and system is provided for robotic task completion from natural language requests. The method may include: receiving a natural language command, processing the natural language command with a generative large language model to extract an intent and associated context, creating a three-dimensional (3D) open-vocabulary semantic scene graph of the environment, associating the scene graph with the intent and associated context, creating, based at least in part on the scene graph, an execution plan comprising a sequence of actions to complete the natural language command, and generating executable code or tool calls corresponding to one or more actions in the sequence of actions, and controlling one or more robotic manipulators and / or actuators to perform one or more actions of the sequence of actions based on the execution plan.
Owner:JOHNS HOPKINS UNIVERSITY

Semantic-based unmanned aerial vehicle autonomous navigation method and device, equipment and medium

The invention belongs to the technical field of unmanned aerial vehicle navigation, and relates to an unmanned aerial vehicle autonomous navigation method and device based on semantics, equipment and a medium. The method comprises the following steps: acquiring an image sequence and IMU data of an unmanned aerial vehicle, and constructing a three-dimensional geometric skeleton of an environment; forming a two-dimensional semantic map according to the image sequence of the unmanned aerial vehicle; marking corresponding semantic tags according to the three-dimensional geometric skeleton and the two-dimensional semantic map, and taking the semantic tags and the geometric information as observation results; fusing observation results from different visual angles to obtain a semantic-geometric coupling map; according to the semantic-geometric coupling map, different nodes are generated, and a dynamic three-dimensional scene graph is obtained by taking a relationship between the nodes as an edge for connecting the nodes; and obtaining an unmanned aerial vehicle instruction, and generating a flight path according to the dynamic three-dimensional scene graph to realize autonomous navigation of the unmanned aerial vehicle. According to the invention, instructions can be understood in a semantic level, so that autonomous navigation of the unmanned aerial vehicle is realized.
Owner:NAT UNIV OF DEFENSE TECH

Story-driven role and scene image generation method

The invention discloses a story-driven role and scene image generation method. The method comprises the steps that a natural language story text input by a user is received and preprocessed; through predefined role description structure constraints, enabling the language understanding and generation model to output a structured role description information structure under template constraints; generating a role image according to the structured role description information text, and extracting image features for consistency control; establishing a mapping table of structured role description information and image feature representation, and realizing the consistency of the appearance of roles in multiple scenes; automatically disassembling the complete story text into a plurality of scene nodes, and generating structured scene description information for each scene; and generating a complete story picture in combination with the scene description information and the role reference diagram. The invention provides a story-driven role and scene image generation method, which is used for automatically generating a story text to a role image and a scene image through semantic understanding, information description structured generation and image consistency management.
Owner:DEEP EXTENDED REALITY RES INC

Industrial quality inspection data cooperative transmission method driven by image recognition

The invention relates to an industrial quality inspection data cooperative transmission method driven by image recognition, in particular to the field of artificial intelligence, which analyzes video frames in real time through a lightweight neural network, accurately locates and extracts key areas and features in industrial product images, constructs a dynamic semantic scene graph to understand contents and divide priorities, and improves the quality inspection efficiency. The core of the method is that non-uniform intelligent coding is carried out on video streams according to semantic importance of contents, high-fidelity transmission of key information such as defects is ensured, non-key areas are greatly compressed to save bandwidth, scheduling is carried out in the transmission process according to priorities, confidence feedback based on a cloud recognition result is introduced, closed-loop control is formed, and high-fidelity transmission is realized. According to the method, the front-end analysis model, the coding strategy and the network path are adaptively optimized, so that low-delay and high-reliability visual quality inspection data transmission and accurate identification can be continuously and stably realized in a complex industrial network environment, and the overall efficiency and the intelligent level of online quality inspection are remarkably improved.
Owner:XIAN UNIV OF TECH

High-fidelity three-dimensional reconstruction method and system based on multi-source data fusion

The invention relates to a high-fidelity three-dimensional reconstruction method and system based on multi-source data fusion, and the method comprises the steps: obtaining multi-source heterogeneous data of a power grid target region, carrying out the time-space alignment and standardization preprocessing, and generating a standardized data set; performing multi-modal feature extraction and fusion coding on the standardized data set to generate uniform multi-dimensional feature representation; performing entity matching and semantic enhancement on the multi-dimensional feature representation and a pre-constructed power grid semantic knowledge graph to generate a semantic enhanced three-dimensional scene graph; carrying out three-dimensional geometric reconstruction and dynamic detail enhancement on the three-dimensional scene graph based on semantic enhancement, and constructing a high-fidelity power grid model; performing multi-level-of-detail optimization processing and format packaging on the high-fidelity power grid model, and outputting a high-fidelity three-dimensional power grid model suitable for the digital twin platform; according to the scheme, multi-source heterogeneous data can be fully fused, a high-precision and high-fidelity three-dimensional power grid model is constructed, and the requirements of fine management and real-time monitoring of a power grid are met.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO TANCHENG COUNTY POWER SUPPLY CO

Optical remote sensing ground feature relationship semantic understanding system and method in localization environment

The invention discloses an optical remote sensing ground feature relationship semantic understanding system and method in a localized environment, and the system constructs a remote sensing scene relationship analysis module, a multi-modal scene knowledge base module, a remote sensing scene graph construction module, a remote sensing scene representation module and a multi-modal sample pair construction module. Carrying out programmed modeling on the ground feature relationship by utilizing a code big language model adaptive to a domestic platform, and generating a scene graph triple; through a double-branch double-time phase comparison learning network, time sequence vision and structure knowledge are fused, and cross-modal joint representation learning is realized. According to the method, the problems of incomplete surface feature relationship expression, inaccurate modeling and lack of real-time semantics in the prior art are solved, deep optimization is carried out in the aspects of data annotation, heterogeneous computing power management and the like aiming at the localized environment, and the accuracy and integrity of semantic understanding of the surface feature relationship and the operation efficiency of the semantic understanding on a domestic software and hardware platform are remarkably improved.
Owner:SUZHOU AEROSPACE INFORMATION RES INST

Battery pack special-shaped part accurate positioning method and system fusing disassembly topological relation

The invention discloses a battery pack special-shaped part accurate positioning method and system fusing a disassembly topological relation, and belongs to the technical field of power battery disassembly positioning. The method comprises the following steps: acquiring a global point cloud of a battery pack by using a global camera, and registering with a preset CAD model to obtain a transformation matrix; a graph neural network assembly relation scene graph is constructed based on the matrix and a CAD model analysis result, and semantic association recognition of the parts is achieved; performing difference detection by comparing the registered CAD model with the global point cloud, and marking the state of the part; and planning an optimal recognition view angle of the local camera according to a sheltered or uncertain state, controlling the local camera to obtain a local point cloud and performing fine positioning, and updating pose information in the scene graph. The method is mainly used for accurate positioning of special-shaped parts in the power battery disassembling process, data support is provided for robot disassembling operation, the positioning robustness under the severe working condition is improved, and meanwhile the positioning problem caused by shielding is effectively solved.
Owner:INST OF INTELLIGENT MFG GUANGDONG ACAD OF SCI

Scene graph generation method and system based on dual-dependence joint learning

The invention belongs to the field of computer vision and artificial intelligence, and provides a scene graph generation method and system based on dual-dependence joint learning, and the method achieves the feature alignment through the cross attention operation of the flattening features of an input image and the query vectors of a subject and an object. Then, subject and object features are analyzed, high-confidence pairs are screened out, and the high-confidence pairs and predicate query vectors are processed in a decoder to output triple semantic features. And constructing a global association graph based on the features, updating node features by using an attention graph convolutional network, and finally predicting the category and bounding box of each triple by using a multi-layer perceptron to complete scene graph generation. According to the method, the generation accuracy and the relation context consistency are improved, the method is suitable for the fields of image understanding, intelligent monitoring and the like, and a complete system architecture solution is provided.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Layered three-dimensional scene generation method and system based on spatial super-division

The invention discloses a hierarchical three-dimensional scene generation method and system based on spatial super-division, and belongs to the technical field of computer graphics, and the method comprises the steps: carrying out the preprocessing of a scene image, and obtaining a high-resolution object image; generating initial rough scene voxels for the scene image, and screening rough voxels and structural latent variables aligned with the high-resolution object image from the initial rough scene voxels to construct a hierarchical scene tree; inputting the high-resolution object image and the rough voxel into a voxel super-resolution model, and generating a fine voxel which keeps geometric consistency with the rough voxel; performing scale alignment and attitude registration based on the rough voxels and the fine voxels; and generating fine voxels of the sub-components recursively by taking the rough voxels of the current node as conditions based on the hierarchical scene tree, and finally assembling to generate a high-resolution three-dimensional scene. According to the method, a high-quality three-dimensional scene with high visual fidelity, fine geometric details and global structure consistency can be efficiently and automatically reconstructed from a single RGB image.
Owner:ZHEJIANG UNIV +1

Bridge side falling behavior sensing method and system suitable for low-quality visual conditions

The invention discloses a bridge side falling behavior sensing method and system suitable for low-quality visual conditions, and belongs to the technical field of behavior sensing. The method comprises the following steps: judging whether weak light and / or blur exists in a bridge side scene image according to a calculated TBV index of the bridge side scene image; the bridge side scene image with the weak light is input into the D2A-DCE model, the bridge side scene image with the blurring is input into the improved NAFNet model, the bridge side scene image with the weak light and the blurring is sequentially input into the D2A-DCE model and the improved NAFNet model, and a bridge side scene enhanced image is output; inputting the bridge side scene enhanced image into the improved CMFF model, and outputting a fusion feature map; and inputting the fusion feature map into a pre-constructed Bridge-SIDE Net model, and outputting a bridge side falling behavior sensing result. According to the invention, the problems of poor bridge side behavior identification accuracy, false report and missing report in an image quality degradation environment can be solved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-modal semantic guided three-dimensional target positioning method, medium and equipment

The invention provides a multi-modal semantic guided three-dimensional target positioning method, a medium and equipment. The method is realized based on a multi-modal semantic guided three-dimensional target positioning model, and comprises a multi-view semantic prior module, a text coding module, a double-branch point cloud coding and multi-source comparison supervision module, a sparse scene graph construction and graph relationship learning module and a positioning decoding module. The multi-view semantic prior module is used for segmenting the 3D scene point cloud into a 3D object point cloud, generating a multi-view 2D visual representation and encoding the semantics of the multi-view 2D visual representation; the double-branch point cloud coding and multi-source comparison supervision module is used for injecting semantic features into a 3D object point cloud to obtain 3D fusion features and realizing multi-source feature alignment through multi-source comparison supervision; the sparse scene graph construction and graph relation learning module is used for constructing a sparse scene graph and optimizing the sparse scene graph through a graph attention network; and the positioning decoding module is used for decoding and outputting a positioning result. The method can improve the target positioning capability of the model in a complex scene.
Owner:SOUTH CHINA UNIV OF TECH

Cross-scene image target detection method based on DFIR-DETR architecture

The invention discloses a cross-scene image target detection method based on a DFIR-DETR framework, and the method comprises the following steps: collecting an image of the surface of industrial steel, and carrying out the preprocessing of an original image; inputting the preprocessed image into a backbone network, and extracting multi-scale depth features represented by defects through the backbone network; carrying out bidirectional fusion and enhancement processing on the extracted multi-scale depth features, and constructing a feature pyramid with global context information and local detail information; and based on the enhanced feature pyramid, performing defect classification and bounding box regression, and outputting the category, confidence and position information of the defect. A technical architecture of dynamic focusing, multi-scale cooperation and frequency domain enhancement is established, and high-precision, strong-robustness and high-efficiency industrial steel surface defect detection is realized. Not only is the significant improvement of the core index reflected, but also the inherent problems of the traditional method in the aspects of real-time performance, small target detection and cross-scene generalization are solved, and the method has high industrial application value.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

Robot operation method and system combined with scene graph

The invention relates to the technical field of robots, in particular to a robot operation method and system combined with a scene graph. According to the method, a monitoring camera can be utilized to perform semantic segmentation on a shot image (scene graph) of a target ground to identify a stain area and a stain type, the cleaning complexity is distinguished, and the stain cleaning sequence and the cleaning parameters are determined from low to high based on the complexity, so that low-complexity stains can be processed firstly, and the cleaning efficiency is improved. The method has the advantages that high-complexity stains are processed, the sequence is reasonable, secondary pollution is avoided, different cleaning parameters are determined according to different complexities, one-step operation is avoided, cleaning is more reasonable, in addition, a scene graph acquired by the robot is acquired by monitoring, scanning by the robot is not needed, the efficiency is high, and omission is avoided.
Owner:SHENZHEN AOWAN TECH CO LTD

Visual servo intelligent robust control method of mobile mechanical arm for complex operation tasks

The invention belongs to the technical field of robot control, and discloses a visual servo intelligent robust control method for a mobile mechanical arm for a complex operation task, which comprises the following steps of: 1, acquiring a working scene image, and calculating a space coordinate of a target feature point in real time through a projection transformation model; step 2, dynamically inhibiting visual measurement noise by adopting adaptive Kalman filtering, generating a feedforward compensation signal through an integral sliding mode observer, and decoupling chassis slippage disturbance; step 3, inputting the pose signal into a radial RBF neural network, designing an RBF gain scheduler, and adjusting an output proportion-integral gain in real time to suppress external time-varying disturbance; 4, adopting a hybrid visual servo mode switching mechanism, and generating a mobile platform control instruction through an integral sliding mode surface in a position servo mode; and 5, designing a three-order composite controller to drive the mechanical arm so as to realize accurate grabbing. According to the method, the grabbing deviation caused by kinematics uncertainty of the mobile platform and dynamic environment disturbance is eliminated.
Owner:NANTONG UNIV

Machine learning model for task and motion planning

Apparatuses, systems, and techniques are described that solve task and motion planning problems. In at least one embodiment, a task and motion planning problem is modeled using a geometric scene graph that records positions and orientations of objects within a playfield, and a symbolic scene graph that represents states of objects within context of a task to be solved. In at least one embodiment, task planning is performed using symbolic scene graph, and motion planning is performed using a geometric scene graph.
Owner:NVIDIA CORP

Multi-modal perception fusion and dynamic feedback-based intelligent task decision and execution optimization method

PendingCN121956516AAddress adjustment lagAdaptive controlData synchronizationClosed loop
The invention relates to the technical field of reinforcement learning, in particular to an intelligent task decision-making and execution optimization method based on multi-modal perception fusion and dynamic feedback, which comprises the following steps of: analyzing space coordinates and brightness difference to construct a dynamic scene graph by collecting and fusing multi-modal data such as visual laser radar torque and completing filtering normalization; the method comprises the following steps: acquiring a task target extraction target area and a reachable area, generating an action sequence in combination with a mechanical arm pose decomposition task, weighting according to joint angle deviation to form an initial trajectory, and performing prediction control rolling optimization of an execution trajectory in combination with a real-time joint state sequence. Vision, laser radar and force sense data synchronous representation is enhanced through multi-modal fusion, a dynamic scene graph is constructed in combination with spatial difference analysis and category division, a task action sequence is generated, a track is corrected, a continuous closed loop of perception, decision and execution is formed, and the problems of insufficient environmental representation, task decomposition stiffness and track adjustment lag are solved.
Owner:CRRC IND INST CO LTD

Visual duplicate removal method, device and equipment for overlapped targets

The invention discloses a visual duplicate removal method, device and equipment for overlapped targets, and relates to the technical field of automation control, and the method comprises the steps: collecting a scene image of a to-be-captured target, and recognizing a matching template and pose information of each target in the scene image based on template matching; constructing a reference polygon enveloping a target boundary according to the matching template; performing external expansion or internal shrinkage processing on the reference polygon based on a preset safe distance, and generating a safe region polygon in combination with the matching template; transforming the associated security area polygon into the same image coordinate system through the pose information of each target; and in the image coordinate system, judging whether any two targets have an overlapping region or not so as to carry out overlapping removal processing. Through accurate geometric description, preset safety distance and target matching and overlapping judgment, misjudgment and missing detection of target overlapping are remarkably reduced, and the grabbing collision risk caused by neglecting of the physical size of an actuator is avoided.
Owner:SHENZHEN ZMOTION TECH CO LTD

Three-dimensional image modeling method and device based on plane data

The invention provides a three-dimensional image modeling method and device based on plane data, and the method comprises the steps: carrying out the full-color / multispectral band fusion, bit depth adjustment, thin cloud removal, image enhancement, stripe removal and light and color uniformity of a satellite surveying and mapping image, and achieving the real color recovery of the image; performing splicing preprocessing on the plurality of small orthographic satellite surveying and mapping images after real color recovery, performing image registration on the to-be-registered image and a reference image, performing image geometric correction on the registered image, and performing image mosaic processing on the plurality of images after image geometric correction to obtain a single large-scene image; and three-dimensional image matching, building mask extraction, three-dimensional building modeling and three-dimensional building post-processing are carried out on a single large scene image to obtain a three-dimensional digital surface model, and ground three-dimensional real image modeling is completed. By applying the technical scheme of the invention, the technical problem that the existing two-dimensional surveying and mapping data is difficult to meet the increasing application requirements of three-dimensional scenes is solved.
Owner:BEIJING AEROSPACE TECH INST