Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1560 results about "Scene graph" patented technology

A scene graph is a general data structure commonly used by vector-based graphics editing applications and modern computer games, which arranges the logical and often spatial representation of a graphical scene.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Human abnormal behavior monitoring method based on large-model multi-agent

The invention discloses a human abnormal behavior monitoring method based on a large-model multi-agent, which is executed by a modular multi-agent system deployed on a back-end server, obtains information through a monitoring camera, and comprises the following steps: obtaining a video stream from the monitoring camera by a sensing agent and extracting human body posture features; analyzing the key frame by a scene understanding agent by using a visual large model, and constructing a time sequence dynamic scene graph; the core reasoning agent evaluates the scene semantic conformity based on the pre-trained large model and performs abnormal preliminary judgment; performing fine-grained classification, interpretation generation and risk assessment on the abnormal behaviors; and the report and action agent generates an alarm and records event data. According to the invention, through multi-agent cooperative work and a large model technology, efficient and accurate monitoring of human abnormal behaviors is realized, and the intelligent level of the monitoring system and the abnormal behavior identification accuracy are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Multi-modal large language model fine tuning method, system, equipment and medium

The invention relates to a multi-mode large language model fine tuning method, system and device and a medium, and belongs to the technical field of artificial intelligence and computer vision crossing. The fine tuning method comprises the steps that an original business scene image is acquired and preprocessed, and a preprocessed image is obtained; performing bounding box coordinate labeling and semantic label definition on the entity target in the preprocessed image through a labeling tool, and outputting a structured labeling file; based on the preprocessed image and the structured annotation file, constructing a training sample set comprising multiple rounds of image-text dialogues; loading the pre-trained multi-modal large language model, configuring low-rank matrix decomposition parameters, and generating a fine tuning instruction set; and inputting the training sample set into a pre-trained multi-modal large language model, carrying out joint training operation based on the fine tuning instruction set, and outputting the fine-tuned multi-modal large language model. According to the method, the identification accuracy, the interaction capability and the system availability of the visual question-answering system in an actual application scene are improved.
Owner:GOLDEN TIMES CULTURE COMM

VLM model intelligent decision-making-based driving method and device, and storage medium

PendingCN121291416AAlgorithmControl signal
The invention discloses a VLM model intelligent decision-making-based driving method and device and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: extracting key visual features from continuous multi-frame driving scene images based on a preset visual feature extraction algorithm; inputting a user instruction and a time sequence visual Token corresponding to the key visual features into a preset VLM model for multi-modal alignment, and generating a target planning Token; inputting the target planning Token and the time sequence vision Token into a preset trajectory generation model to obtain a predicted trajectory; and converting the predicted trajectory into a control signal, and controlling the mobile device to complete a moving action based on the control signal. The problem of modal difference between a semantic space and an action space is solved.
Owner:YOUDI ROBOT (WUXI) CO LTD

Multi-device cooperative control method and system under autonomous cooperation algorithm

The invention relates to the technical field of equipment cooperative control, and discloses a multi-equipment cooperative control method and system under an autonomous cooperation algorithm, and the method comprises the steps: collecting the physical attribute data, real-time operation data, production task data and a smart factory scene graph of operation equipment under a multi-equipment cooperative task in a smart factory application; based on the physical attribute data and the real-time operation data, identifying a space conflict mode and a time conflict mode of the operation equipment to obtain collaborative conflict information; according to the production task data and the collaboration conflict information, carrying out collaboration conflict decoupling on the operation equipment to obtain a feasible collaboration mode; performing obstacle avoidance reinforcement learning on the operation equipment by using the smart factory scene graph to construct an obstacle avoidance global path of the operation equipment in the smart factory; and based on the feasible coordination mode and the obstacle avoidance global path, constructing a target coordination control scheme of the operation equipment so as to perform coordination control on the operation equipment. According to the invention, the reliability of multi-device cooperative control under the autonomous cooperation algorithm can be improved.
Owner:HARBIN SAISI TECH CO LTD

Multi-view-angle-oriented three-dimensional scene image reconstruction registration and optimization method and system

The invention provides a multi-view-oriented three-dimensional scene image reconstruction registration and optimization method and system, and relates to the technical field of image processing, and the method comprises the steps: obtaining multi-frame three-dimensional scene image data, extracting a multi-level feature set, constructing a cross-view-angle semantic association graph, building a feature corresponding relation, and calculating a multi-view-angle spatial transformation relation parameter. Performing coordinate system alignment on the image data to generate an initial three-dimensional reconstruction result, and performing optimization in combination with a multi-target joint optimization function and a dynamic adaptive weight regulation and control mechanism. According to the method, the precision and robustness of three-dimensional scene reconstruction are improved, and the problem of registration errors caused by large view angle difference in a complex scene is solved.
Owner:BEIJING SETTALL TECH DEV CO LTD

Edge-deployed semi-supervised anomaly detection method and system for railway track foreign object

Disclosed in the present invention are an edge-deployed semi-supervised anomaly detection method and system for a railway track foreign object. The method comprises the following steps: an edge device encoding and decoding a video stream captured by a camera to obtain an image frame sequence, and performing frame extraction; and using a semantic segmentation model to perform image segmentation on a certain image frame obtained by means of frame extraction, to obtain a railway track region segmentation image. The use of a single image as input may generate an expert model result having a high weight value; however, the determination based on a single image is not stable, multiple consecutive images of the task scene need to be inputted, the frequency of each expert model obtaining the highest weight is computed, and the expert model corresponding to the highest frequency is the final solution. The present invention supports scene-adaptive foreign object detection algorithm automatic selection, and a user can perform selection on the basis of prior knowledge, or selection may be performed by a scene-adaptive automatic algorithm selection method; the user only needs to provide a batch of image data of the current scene, and the optimal algorithm selection can be evaluated.
Owner:GUANGZHOU EMBEDDED MACHINE TECH CO LTD

Systems and methods for multi-modal visual reasoning using multiple scene graphs

A system for processing multi-modal data representing an environment to generate scene graphs of the environment is described. The system can obtain sensor data associated with a vehicle operating in the environment. In examples, the system can determine a set of features from the sensor data, including one or more objects and one or more agents present in the environment, and can generate a scene graph that represents the poses and velocities of these objects and agents relative to the environment. In some examples, based on generating the scene graph, the system can generate a knowledge graph by encoding the relationships among the identified objects and agents. In some examples, the system can generate a control signal, using attributes that represent the states of objects and agents in the knowledge graph, and provide this control signal to the vehicle in order to adjust or cause the operation of the vehicle.
Owner:QPIAI INDIA PTE LTD

Intelligent accident liability judgment method and system fused with video tracking

The invention provides an accident liability intelligent determination method and system fused with video tracking, and the method comprises the steps: S1, carrying out the alignment evidence collection of a video, an IMU, a GPS, and an OBD, and forming an evidence chain through framing Hash, chain Hash, and a timestamp; s2, completing detection segmentation on the shared trunk multi-branch network; s3, in combination with appearance re-identification and a motion model, realizing cross-frame association, and outputting track, speed and shielding recovery; s4, mapping the trajectory to BEV according to depth and calibration, and constructing a space-time scene graph by fusing lanes, signals, speed limit and the like; s5, extracting events such as lane changing, line merging and signal rushing, and positioning a conflict point and a way giving relation; s6, generating error factors according to a rule base and priority judgment; s7, quantifying the responsibility proportion according to the weight, the collision participation time and the reaction time; and S8, generating a report evidence of the key frame and the rule list. The system is composed of an acquisition evidence obtaining module, a video understanding tracking module, a space-time reconstruction module, an event extraction module, a rule reasoning module, a responsibility quantification module and a report evidence storage module.
Owner:国任财产保险股份有限公司

Construction method of scene graph question and answer inference model based on reinforcement learning

The invention relates to the technical field of visual questions and answers, in particular to a scene graph question and answer inference model construction method based on reinforcement learning, which comprises the following steps: calling an existing large model interface, performing deep processing on an initial scene data set, and constructing a first batch of training sets; carrying out first-stage reinforcement learning training on a pre-constructed multi-modal large model by utilizing the first batch of training sets; obtaining an LLaVA-CoT data set, screening the LLaVA-CoT data set, and inputting the screened LLaVA-CoT data set into the multi-modal large model trained at the first stage to obtain a reasoning result; calling an existing large model interface to carry out accuracy evaluation and correction on the reasoning result of the multi-modal large model to obtain a second batch of training sets; and performing second-stage reinforcement learning training on the multi-modal large model by using the second batch of training sets to obtain a final scene graph question and answer inference model. According to the method, a staged and differentiated training strategy is adopted, the characteristic advantages of two batches of data are brought into full play, staged training is carried out on the large model, and the model performance is gradually improved.
Owner:DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD

Determining lighting and composition parameters using machine learning models for synthetic data generation

Approaches presented herein provide for the determination of realistic lighting parameters for a scene represented in an image. Realistic lighting parameters can allow for the insertion of one or more virtual objects into a scene image, where the lighting or shading applied to the virtual object(s) can be consistent with those for other objects in the scene. A machine learning model such as a discriminator or diffusion model can be used to analyze a composed image generated by a differential renderer, for example, in which at least one virtual object has been inserted into a scene image and had lighting effects applied in accordance with a set of lighting parameters. A loss value can be determined based on the results of this machine learning model, which can be used to optimize the lighting parameters and / or adjust the weights or parameters of a model used to generate the lighting parameters. Once fine-tuned or optimized, the lighting parameters can represent an accurate light map for the scene or environment that can be used to generate composed images.
Owner:NVIDIA CORP

Three-dimensional scene reconstruction method and apparatus, device, medium, and program product

Embodiments of the present disclosure provide a three-dimensional scene reconstruction method and apparatus, a device, a medium, and a program product. The three-dimensional scene reconstruction method comprises: acquiring a scene image collected for a three-dimensional scene; on the basis of the scene image, determining sparse point cloud data and camera parameter information corresponding to the scene image; performing model initialization on the basis of the sparse point cloud data to obtain a current three-dimensional data model; performing rendering on the basis of the camera parameter information and the current three-dimensional data model to obtain a current color map, a current depth map and a current normal map from the perspective of the scene image, and determining a pseudo normal map from the perspective of the scene image on the basis of the current depth map; and on the basis of the current color map, the current normal map, the pseudo normal map, and an actual color map, training the current three-dimensional data model to obtain a trained target three-dimensional data model. By means of the technical solution provided by the embodiments of the present disclosure, higher-quality automatic reconstruction of three-dimensional scenes can be achieved, reducing reconstruction costs, and improving the level of detail and rendering quality of scene models.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Three-dimensional model reconstruction and image generation method, device, storage medium, and program product

Embodiments of the present application provide a three-dimensional model reconstruction and image generation method, a device, a storage medium, and a program product. In the method, multi-stage three-dimensional reconstruction is performed on the basis of a single image of a target object; in a first stage, a plurality of view images are generated on the basis of an image generation model, and an initial three-dimensional model is reconstructed on the basis of the plurality of view images; and in a second stage, on the basis of the plurality of view images and an initial prompt containing set marker information, a text-to-image model is used to learn an association relationship between the target object and the set marker information, and a plurality of scene images are generated on this basis. Compared with the plurality of view images, the scene images generated in the second stage have higher resolution and richer image details, and then the initial three-dimensional model is optimized on the basis of the plurality of scene images, so that a target three-dimensional model having higher resolution and clearer model details can be obtained, thereby paving the way for practical application of a three-dimensional reconstruction solution based on a single image.
Owner:TAOBAO CHINA SOFTWARE

Image scene understanding method and system applied to Internet of Vehicles road condition analysis

The invention provides an image scene understanding method and system applied to vehicle networking road condition analysis, and the method comprises the steps: firstly obtaining a multi-source road condition image set which is collected by a vehicle networking roadside sensing node in a continuous time period and comprises road scene image sequences in different shooting angles and different exposure modes, and carrying out the time-space consistency calibration processing of the multi-source road condition image set; the method comprises the following steps: generating a standardized image sequence, performing hierarchical feature analysis on the standardized image sequence to form a feature hierarchical chain from a bottom layer to a high layer, inputting the feature hierarchical chain into a pre-training scene semantic understanding model for cross-layer association reasoning, and generating scene semantic description containing road element type labels and dynamic relationships among elements; and finally, generating road condition analysis data containing road condition element positioning information and interaction trend prediction based on the scene semantic description, and transmitting the road condition analysis data to an Internet of Vehicles communication terminal to support driving decision optimization, thereby effectively improving the accuracy and decision support capability of Internet of Vehicles road condition analysis.
Owner:BEIJING CHEXIAO TECH CO LTD

Diffusion model-guided training of generative models for rendering novel views of 3D scenes

Approaches presented herein provide for the training and use of generative models to generate high quality image data for novel reconstruction views. A generative model such as a neural radiance field (NeRF) can be trained to generate such content. In order to train the NeRF to represent a specific scene with high accuracy, the NeRF can be trained using a diffusion model for score distillation guidance. The diffusion model can be trained using a large set of environment data from a variety of different views, then fine-tuned for a specific domain and / or scene. A parameter-efficient training process can be used to avoid overfitting of the diffusion model to the domain- or scene-specific training data. Once fine-tuned, the “expert” diffusion model can be used with the NeRF during training to effectively transfer the expert knowledge to the NeRF, enabling the NeRF to generate high quality image data for the scene from viewpoints corresponding to extreme novel views.
Owner:NVIDIA CORP

High-place operation risk real-time judgment method based on image and point cloud bimodal data

The invention discloses a high-altitude operation risk real-time judgment method based on image and point cloud bimodal data, and belongs to the crossing field of computer vision, artificial intelligence and engineering construction safety, and the method comprises the following steps: S1, obtaining a bimodal data stream through a laser radar and an RGB camera; s2, fitting an external reference registration matrix by acquiring an internal reference matrix of the camera to obtain a calibration parameter set; s3, optimizing the cue word, and driving the SAM-2 segmentation model to obtain a two-dimensional mask; s4, performing reverse projection on the two-dimensional mask to obtain a color point cloud instance set; s5, calculating the three-dimensional geometric centroid to obtain the three-dimensional geometric centroid and a three-dimensional semantic topology scene graph; and S6, obtaining a plurality of security rule compliance results by executing the interpretable rule engine. By the adoption of the method, the defects of an existing high-place operation risk judgment technology in the aspects of complex environment adaptability, multi-modal fusion efficiency, three-dimensional space perception and risk judgment interpretability are overcome.
Owner:SOUTH CHINA UNIV OF TECH

Defenses for attacks against non-max suppression (NMS) for object detection

Systems and techniques are described for object detection. For example, a computing device can apply a transformation to an image of a scene to generate a transformed image. The computing device can determine a plurality of candidate bounding regions for the transformed image. Each candidate bounding region is associated with an object in the scene. The computing device can determine a subset of candidate bounding regions for the transformed image by removing, using a non-max suppression model, at least one candidate bounding region of the plurality of candidate bounding regions. The computing device can generate an output bounding box for the object based on the subset of candidate bounding regions. The computing device can output the output bounding box. In some cases, the computing device can use an output of an image processing operation on the image to reduce a number of the plurality of candidate bounding regions.
Owner:QUALCOMM INC

Instrument monitoring method, system and equipment for inspection robot and medium

The invention relates to the technical field of robots, in particular to an instrument monitoring method, system and device for an inspection robot and a medium, the method is used for searching an instrument to be inspected by using the inspection robot and monitoring the instrument to be inspected, and the method comprises the following steps: obtaining a view scene image collected by the inspection robot, performing image segmentation and feature extraction on the view scene image; obtaining the exploration value of the instrument area; obtaining candidate paths from the inspection robot to the instrument areas; evaluating a cost function of each candidate path; carrying out adaptive visual angle adjustment and data acquisition on the inspection instrument; the monitoring method is an exploration type sensing strategy based on exploration value driving, and an instrument target can be found in the environment. And in combination with a historical task map, instrument semantic priori and current perception data, the information value of the candidate region is evaluated, and optimal decision making is performed on the path by fusing environmental factors such as illumination, shielding and equipment states, so that priority selection and dynamic reconstruction of the target region are realized.
Owner:SICHUAN ENVIRONMENTAL PROTECTION ENG CO LTD CNNC

YOLOv8 traffic sign real-time detection method and system based on edge calculation optimization

The invention provides a YOLOv8 traffic sign real-time detection method based on edge calculation optimization, and the method comprises the steps: collecting traffic scene image data in real time, carrying out the preprocessing of the collected image data, inputting an optimized YOLOv8 network model, carrying out the feature extraction, carrying out the processing of feature maps of different scales through an SE module and a bidirectional feature fusion strategy, and carrying out the detection of a traffic sign. Entering a detection head for target detection to obtain a detection result; an obtained detection result is transmitted to a central server or an automatic driving system in a structured data format through a low-delay communication protocol; in a central server or an automatic driving system, real-time statistics and trend analysis are carried out on detection results, and through deep optimization and deployment strategy improvement on a YOLOv8 model, many limitations in an edge calculation scene in the prior art are overcome. Specifically, a lightweight optimization strategy combining model pruning, quantitative processing and an efficient inference engine is provided, and the detection precision and robustness are improved in combination with an environment adaptive image preprocessing method.
Owner:HARBIN INST OF TECH

Image acquisition test method and device, equipment and storage medium

The invention relates to the technical field of image acquisition and testing, and discloses an image acquisition and testing method, device and equipment and a storage medium, and the method comprises the steps: collecting a standard test card image through an image acquisition card in a multi-stage standard illumination environment, and calculating sensitivity benchmark test data; executing time domain and space domain combined sampling processing to obtain super-sensitivity image data; performing brightness analysis and dynamic adjustment on the currently shot first scene image to obtain a first image acquisition parameter; acquiring a second scene image, and performing image grid division and contrast compensation on the second scene image to obtain an enhanced image after contrast compensation; according to the super-sensitivity image acquisition method and the super-sensitivity image acquisition device, the sensitivity limitation of the position depth of a traditional image acquisition card is broken through, and super-sensitivity image acquisition is realized.
Owner:SHENZHEN LIANRUI ELECTRONICS CO LTD

Visual language action large model design method for enhancing spatial perception ability

The invention discloses a visual language action large model design method for enhancing spatial perception ability, which comprises the following steps of: acquiring a plurality of frames of task scene images, inputting the images into a two-dimensional image encoder and a visual geometry guidance Transformer VGGT encoder, extracting image features and performing time sequence spatial embedding, encoding a natural language instruction into language embedding representation, and extracting a visual language action large model. Image features and language embedding are fused through a cross attention mechanism, a pre-training large model is input to generate cross-modal representation, the fusion features are input into an action expert module to output an action control sequence in combination with robot body state information, and a robot is driven to execute an operation task. Therefore, damage of additional information to an original pre-training model is avoided, and compared with an original combined structure of a pre-training visual language model and an action expert, utilization of multi-view picture information is enhanced, so that higher understanding ability on space depth is achieved in the task execution process, the task success rate is increased, and the task execution efficiency is improved. And a more efficient and more robust robot sensing and decision-making integrated system is realized.
Owner:SHANGHAI WUZHI EVOLUTION TECHNOLOGY CO LTD

Identification method for matching scene behaviors by using multi-modal features

The invention relates to an identification method for matching scene behaviors by using multi-modal features, and belongs to the technical field of scene behavior matching. The method comprises the following steps: acquiring multi-modal scene behavior data, and carrying out noise self-adaptive purification processing on the multi-modal scene behavior data to obtain a preprocessed scene image, scene audio data and a scene label text; secondly, performing feature collaborative extraction on the preprocessed data to obtain visual features, audio features and text features, inputting the cooperatively extracted features into a scene behavior matching network, and performing scene feature fusion and behavior feature fusion respectively to obtain a scene feature vector and a behavior feature vector; and constructing a bipartite graph according to the scene feature vector and the behavior feature vector, calculating the semantic similarity between nodes of the bipartite graph, dynamically updating the edge weight of the bipartite graph according to the semantic similarity between the nodes, and normalizing the updated bipartite graph to obtain a scene behavior matching result. According to the method, the association degree of the scene and the behavior can be accurately quantified, and the accuracy of a matching result is greatly improved.
Owner:LUZHOU VOCATIONAL & TECHN COLLEGE

Mountain road vision field obstacle image segmentation and image enhancement processing method

The invention relates to the technical field of data processing, in particular to a mountainous area road vision field obstacle image segmentation and image enhancement processing method, which specifically comprises the following steps: collecting a mountainous area road scene image, and carrying out labeling and data set division; performing multi-modal illumination correction and adaptive perspective correction based on the acquired image, and then extracting amplitude images of the frequency domain gradient and the space gradient of the corrected image for adaptive fusion to generate a fusion enhanced gradient image; a mountainous area road visual impairment detection model is constructed based on the improved U-Net, a mixed cavity convolution module and channel importance re-calibration mechanism, a space-channel double-path attention gating mechanism and an iterative enhancement strategy are adopted in the improved U-Net, sub-pixel convolution operation is adopted at each level of a decoder, and a detection result is calibrated through an example transformation weighting mechanism. Optimizing a detection result by calculating a loss function; and finally, through training, verification and testing, outputting a most important detection result. According to the method, the detection accuracy can be improved, and the sensitivity of the model to obstacle degree detection is reduced.
Owner:SHANDONG LUQIAO GROUP CO LTD

Endoscope monocular dynamic scene reconstruction method based on dynamic Gaussian splashing and motion tracking

The invention discloses an endoscope monocular dynamic scene reconstruction method based on dynamic Gaussian splashing and motion tracking, and the method comprises the steps: firstly obtaining a dynamic scene image, obtaining an initial three-dimensional track through introducing auxiliary prior information, and then obtaining three-dimensional Gaussian; based on the auxiliary prior information, generating a group of compact Sim3 motion bases, and obtaining each three-dimensional Gaussian transformation function in the dynamic scene image by using the weighted combination of the Sim3 motion bases, the parameters of the three-dimensional Gaussian transformation function including the transformation of the position and the shape along with the time; and performing three-dimensional Gaussian transformation to each frame of image of the dynamic scene according to the three-dimensional Gaussian transformation function, and performing image rendering on each frame to obtain a rendered image. According to the method, monocular depth and dense displacement field prior are combined, and the problem of monocular scale drift is solved by utilizing depth relative sequence stability. According to the method, real-time rendering is ensured, and meanwhile, the dynamic scene reconstruction precision is remarkably improved.
Owner:HANGZHOU DIANZI UNIV

High-precision map reconstruction method and system based on monocular vision, medium and equipment

The invention belongs to the technical field of robot positioning and three-dimensional mapping, and discloses a high-precision map reconstruction method and system based on monocular vision, a medium and equipment. Scene image data are acquired through a monocular image acquisition module, after feature extraction and matching are performed on each frame of image, attitude information acquired by an inertial measurement unit is fused, and pose calculation of a robot is completed through a sparse vision SLAM system; meanwhile, an image dense depth map is generated by a monocular depth estimation model based on an attention mechanism, and the image dense depth map is converted into a single-frame color dense point cloud in combination with image RGB color information. According to the invention, a loose coupling fusion strategy is adopted to carry out spatial registration on a single-frame colored dense point cloud and a robot pose at a corresponding moment, multi-frame data fusion is completed through point cloud splicing and optimization, and a globally consistent three-dimensional dense point cloud map is constructed to realize scene modeling. The method has the characteristics of high robustness and high reconstruction precision, and can be effectively applied to three-dimensional map construction in an outdoor complex environment.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Blind sidewalk identification method based on image processing

The invention relates to the technical field of image recognition, in particular to a blind sidewalk recognition method based on image processing. The method comprises the following steps: collecting a ground scene image, and carrying out noise suppression and illumination normalization processing to obtain a ground scene image to be processed; inputting the to-be-processed ground scene image into a pre-constructed convolutional neural network model to identify feature textures of the blind sidewalk bricks; dividing a blind sidewalk area in the ground scene image to be processed by using the blind sidewalk brick feature texture, and determining a direction gradient feature of the blind sidewalk area; executing blind sidewalk direction consistency constraint based on the direction gradient features to collect blind sidewalk structured path segments; recognizing a blind sidewalk fracture area based on directional gradient features; and performing cross-frame target tracking according to the structured path segment of the blind sidewalk, and outputting a continuous blind sidewalk trajectory. The automatic recognition rate of the blind sidewalk area is improved based on the image recognition technology, and the continuous recognition capacity of the blind sidewalk path in the complex illumination and shielding environment is enhanced.
Owner:SHANDONG SAIFEITE SAFETY ENG TECH DEV CO LTD

Rendering Video Of A Scene Using Three-Dimensional Gaussians

A set of images of a scene re received. Each image includes temporal data and spatial data relating to the scene. Based on the spatial data of each image, three-dimensional (3D) Gaussian splatting data is generated. The temporal data of each image and the 3D Gaussian splatting data are inputted to a neural network to generate spatial-temporal 3D Gaussian embeddings. Offset data based on the spatial-temporal 3D Gaussian embeddings is generated. The video of the scene is rendered based on the 3D Gaussian splatting data and the offset data, allowing for improved rendering of video of the scene.
Owner:YINWANG INTELLIGENT TECHNOLOGIES CO LTD

Multi-modal interaction method and interaction system applied to intelligent robot

The invention discloses a multi-modal interaction method and interaction system applied to an intelligent robot. The method comprises the following steps: collecting a scene image and processing the scene image into a three-dimensional point cloud and a two-dimensional texture feature; the Gemini Robotics-ER model is used for extracting features, and the vision-language-action model is used for analyzing a language instruction into a sequence capable of being recognized by a machine; and fusing the features to generate an interactive decision matrix, planning a trajectory, calculating kinetic parameters, driving the robot to execute actions and feeding back in real time. The system comprises a multispectral visual information acquisition and preprocessing unit, a Gemini Robotics-ER model processing unit, a natural language instruction analysis unit, a vision-language-action cooperative processing unit, a trajectory planning and dynamics calculation unit and a motion control and feedback unit, and all the units work cooperatively. According to the method and the system, through multi-modal fusion and closed-loop control, interaction accuracy and real-time performance are improved, and industrial scene requirements are met.
Owner:ZHENGXIN (SUZHOU) TECHNOLOGY CO LTD

AR scene furniture identification and dynamic removal system based on artificial intelligence

The invention belongs to the technical field of augmented reality, and discloses an AR scenario furniture identification and dynamic removal system based on artificial intelligence, and the system comprises the steps: carrying out the multi-view image registration, furniture entity segmentation and shielding completion analysis based on the obtained hotel scene AR image sequence data, and forming a complete furniture contour feature; the style similarity of the furniture multi-dimensional feature vectors is analyzed, and a furniture style recognition result is obtained; the method comprises the following steps: acquiring AR scanning data of a home environment, performing furniture removal priority ranking through space conflict analysis and style coordination evaluation, generating a dynamic removal strategy, establishing furniture three-dimensional model data based on the dynamic removal strategy, and performing material texture mapping and shadow fusion to obtain a realistic rendering effect; an interactive optimization adjustment scheme is generated, and scene parameter real-time adjustment and optimization and visual effect improvement are realized through user feedback; the reality sense of virtual home display and the user experience effect are remarkably improved.
Owner:SHANGHAI XIANGYUE JIANGFENG DIGITAL TECHNOLOGY CO LTD

Automatic commodity grabbing method and system based on multi-mode sensing and six-axis mechanical arm

The invention relates to an automatic commodity grabbing method and system based on multi-mode perception and a six-axis mechanical arm. The method comprises the steps that 1, commodity detection and zero sample fine granularity recognition are carried out; collecting a current shelf scene image, and carrying out target detection; 2, grabbing posture generation and screening optimization are carried out; 3, grabbing action execution and obstacle avoidance motion planning are carried out; after the optimal grabbing posture is obtained, the grabbing action is executed through the six-axis mechanical arm; the method comprises the following steps: firstly, carrying out voxelization processing on modeling of a scene by BundleFusion; and then, a path planning algorithm is called to generate an obstacle avoidance grabbing path in the current scene, a six-axis mechanical arm end effector is closed to complete grabbing after reaching a target position, and the commodities are carried to a designated position according to task setting. According to the invention, automatic identification and grabbing of any specified commodity are realized, the generalization ability of the system is greatly improved, and the application range of the system is greatly expanded.
Owner:SHANDONG UNIV