Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2146 results about "Scene graph" patented technology

A scene graph is a general data structure commonly used by vector-based graphics editing applications and modern computer games, which arranges the logical and often spatial representation of a graphical scene.

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Community intelligent monitoring and emergency linkage method and system fusing BIM spatial semantics

The invention discloses a community intelligent monitoring and emergency linkage method and system fusing BIM spatial semantics, and the method comprises the steps: constructing a BIM scene map, and obtaining the attributes and mutual relationships of components and spatial regions in a BIM model; mapping a dynamic target detected in video monitoring into the BIM model, and obtaining spatial semantic information of the dynamic target; based on BIM spatial semantic information of a dynamic target, a target-environment interaction graph is constructed, a graph neural network model is used for training and reasoning, and specific complex events related to spatial contexts are recognized; taking the BIM model as a space-time reference, fusing multi-source heterogeneous data, and reconstructing by adopting a graph-based event association algorithm to form a complete event chain containing an atomic event sequence and an association relationship; and when an emergency event or an event chain is detected to indicate an emergency state, combining BIM preset information and real-time sensor data, dynamically generating an optimal emergency plan, and performing visual commanding and dispatching through a BIM three-dimensional scene and augmented reality.
Owner:ZHEJIANG LEISHENG CONSTRUCTION ENGINEERING CO LTD

Intelligent agent strategy generation and online optimization method based on dynamic scene perception

The invention provides an agent strategy generation and online optimization method based on dynamic scene perception, and the method comprises the steps: obtaining a scene demand description of a user, and carrying out the analysis of the scene demand description, so as to generate a target task sequence which can be executed by an agent disposed in a target scene; dynamically sensing the current environment characteristics of the target scene to generate a dynamic semantic topology network and generate a dynamic scene graph according to the dynamic semantic topology network; on the basis of the dynamic scene graph, hierarchical modeling of the incidence relation is carried out on the behavior space corresponding to the intelligent agent, and behavior semantic features containing scene perception are generated; the behavior semantic features containing scene perception are mapped to an intelligent agent strategy representation space, and a behavior feature strategy for controlling an intelligent agent to execute a target task sequence is obtained; and according to the determined scene value representation, decomposing the behavior feature strategy to obtain an advantage estimation value adapted to the scene so as to carry out online optimization on the behavior feature strategy. According to the invention, the intelligent agent strategy can accurately adapt to the requirements in the business process of an enterprise.
Owner:BEIJING ZHONGSHURUIZHI TECH CO LTD

Mechanical arm motion control method based on multi-agent cooperation

The invention discloses a mechanical arm motion control method based on multi-agent cooperation, and the method comprises the steps: firstly, receiving an RGB image through a sub-task generation agent, and generating a structured sub-task sequence according to a natural language task instruction of the RGB image; secondly, performing joint modeling on a task text and a scene image through a 3D sensing intelligent body, positioning specific coordinates of a target object in a three-dimensional space, reasoning dynamic characteristics of a current environment based on historical state information of a robot by combining an environment sensor, and generating an environment sensing vector; and finally, the action generation agent performs fusion modeling according to the subtask text, the subtask target coordinates, the current state of the robot and the environment perception vector, generates a continuous action vector, drives a mechanical arm to complete each subtask action, and constructs closed-loop feedback by a controller and a discriminator to realize task execution state judgment and automatic circulation. The precise action control instruction can be effectively generated, and the task execution fineness of the mechanical arm is remarkably improved.
Owner:CHINA JILIANG UNIV +1

Three-dimensional Gaussian sputtering method for sparse visual angle semantic priori

The invention discloses a three-dimensional Gaussian sputtering method for sparse visual angle semantic priori. The method comprises the following steps of: 1, acquiring a target scene image, constructing a real image set as a data set, and manually selecting an interested object in a target scene to perform semantic three-dimensional reconstruction to obtain an initial semantic image of a sparse view angle; a multi-view image is collected, camera external parameters and scene sparse point clouds are obtained, a plurality of views are selected and input into the SAM2 segmentation model, and a view set Is with semantic images is obtained; 2, using a pre-trained SAM2 segmentation model as an interactive image sequence segmentation model, and initializing a three-dimensional Gaussian primitive according to the scene sparse point cloud; and step 3, training parameters of semantic three-dimensional Gaussian sputtering based on the trained three-dimensional Gaussian sputtering model and the interactive image sequence segmentation model, and reconstructing a target scene. According to the method, the high-quality semantic model is efficiently reconstructed. And the reconstruction result can be easily corrected through the interactive graphical interface in the training process.
Owner:XIDIAN UNIV

Scene self-adaptive adjustment method and system for virtual-real fusion of intelligent internet of things and element universe

The invention provides a scene adaptive adjustment method and system based on intelligent Internet of Things and meta-universe virtual-real fusion, and relates to the technical field of artificial intelligence and Internet of Things, and the method comprises the steps: collecting a scene image and environment parameter data, carrying out the decomposition and partitioning of the image, and recognizing the type of the scene, environment parameter tensors are constructed to form a digital twinborn model, the digital twinborn model is mapped to a virtual space, a proper feature vector is selected according to an identification result to generate a target feature vector, a projection position of a virtual object in an entity space is obtained, and the feature vector is optimized and then combined with the projection position to generate an enhanced feature and a mapping matrix; and generating rendering parameters through the parameter mapping network, rendering the virtual object, and displaying the virtual object on the entity space interface in an overlapping manner.
Owner:HANGZHOU MOXI TECH DEV CO LTD

Human abnormal behavior monitoring method based on large-model multi-agent

The invention discloses a human abnormal behavior monitoring method based on a large-model multi-agent, which is executed by a modular multi-agent system deployed on a back-end server, obtains information through a monitoring camera, and comprises the following steps: obtaining a video stream from the monitoring camera by a sensing agent and extracting human body posture features; analyzing the key frame by a scene understanding agent by using a visual large model, and constructing a time sequence dynamic scene graph; the core reasoning agent evaluates the scene semantic conformity based on the pre-trained large model and performs abnormal preliminary judgment; performing fine-grained classification, interpretation generation and risk assessment on the abnormal behaviors; and the report and action agent generates an alarm and records event data. According to the invention, through multi-agent cooperative work and a large model technology, efficient and accurate monitoring of human abnormal behaviors is realized, and the intelligent level of the monitoring system and the abnormal behavior identification accuracy are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Video plot generation and scene synthesis method and system based on natural language processing

The invention provides a video plot generation and scene synthesis method and system based on natural language processing, and relates to the technical field of video generation, and the method comprises the steps: receiving a script text to construct a multilayer scene map, extracting keywords to calculate semantic relevancy, executing feature decomposition and reconstruction to obtain a scene synthesis vector, and generating a video scene based on distance measurement. And extracting spatio-temporal features by using the feature pyramid, executing adaptive feature fusion, segmenting the video by applying a self-attention mechanism, and after a transition effect is inserted, performing style migration to output a finished video.
Owner:SHANDONG FOREIGN LANGUAGES VOCATIONAL AND TECH UNIV +1

Multi-modal large language model fine tuning method, system, equipment and medium

The invention relates to a multi-mode large language model fine tuning method, system and device and a medium, and belongs to the technical field of artificial intelligence and computer vision crossing. The fine tuning method comprises the steps that an original business scene image is acquired and preprocessed, and a preprocessed image is obtained; performing bounding box coordinate labeling and semantic label definition on the entity target in the preprocessed image through a labeling tool, and outputting a structured labeling file; based on the preprocessed image and the structured annotation file, constructing a training sample set comprising multiple rounds of image-text dialogues; loading the pre-trained multi-modal large language model, configuring low-rank matrix decomposition parameters, and generating a fine tuning instruction set; and inputting the training sample set into a pre-trained multi-modal large language model, carrying out joint training operation based on the fine tuning instruction set, and outputting the fine-tuned multi-modal large language model. According to the method, the identification accuracy, the interaction capability and the system availability of the visual question-answering system in an actual application scene are improved.
Owner:GOLDEN TIMES CULTURE COMM

Accident scene generation method based on scene knowledge graph and considering accident causes

The invention belongs to the technical field of automatic driving testing, and particularly relates to an accident scene generation method based on a scene knowledge graph and considering accident causes. Comprising the following steps: step 1, modeling a scene knowledge graph; 2, modeling the scene graph time sequence prediction model; 3, modeling the time sequence causal inference model; step 4, modeling the scene graph time sequence decision generation model so as to generate an accident scene; according to the method, the accident scene database with high authenticity, high diversity and accident cause consistency can be constructed under the condition that the accident scene sample data size is limited, the number of accident scenes of the same type in the automatic driving algorithm closed-loop self-evolution cloud database is efficiently expanded, the constructed scene database is used for training the automatic driving algorithm, and the accuracy of the automatic driving algorithm is improved. The adaptive capacity of the self-driving automobile to scenes with the same type of accidents can be effectively enhanced, and safe and reliable operation of the self-driving automobile in the real world is guaranteed.
Owner:JILIN UNIVERSITY

Lightweight multi-target instance segmentation method and system for inspection robot

The invention relates to a lightweight multi-target instance segmentation method and system for an inspection robot, and belongs to the technical field of computer vision and deep learning. The system is used for executing the method and comprises the steps of constructing training data based on a public data set; an improved lightweight YOLOv11 network is constructed, the weight is adaptively adjusted at least through a dynamic convolution module according to the input features, and collaborative optimization is carried out on feature fusion through cascade grouping attention and channel-space attention; the method comprises the following steps: training an improved lightweight YOLOv11 network to obtain an optimal weight; and deploying the trained lightweight instance segmentation model to an inspection robot, processing campus scene image data in real time, and outputting target pixel-level contour and category information. According to the method, the model parameter quantity and the calculation cost are remarkably reduced, and meanwhile, high precision and real-time performance of multi-target instance segmentation in a campus scene are achieved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

VLM model intelligent decision-making-based driving method and device, and storage medium

PendingCN121291416AAlgorithmControl signal
The invention discloses a VLM model intelligent decision-making-based driving method and device and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: extracting key visual features from continuous multi-frame driving scene images based on a preset visual feature extraction algorithm; inputting a user instruction and a time sequence visual Token corresponding to the key visual features into a preset VLM model for multi-modal alignment, and generating a target planning Token; inputting the target planning Token and the time sequence vision Token into a preset trajectory generation model to obtain a predicted trajectory; and converting the predicted trajectory into a control signal, and controlling the mobile device to complete a moving action based on the control signal. The problem of modal difference between a semantic space and an action space is solved.
Owner:YOUDI ROBOT (WUXI) CO LTD

Multi-device cooperative control method and system under autonomous cooperation algorithm

The invention relates to the technical field of equipment cooperative control, and discloses a multi-equipment cooperative control method and system under an autonomous cooperation algorithm, and the method comprises the steps: collecting the physical attribute data, real-time operation data, production task data and a smart factory scene graph of operation equipment under a multi-equipment cooperative task in a smart factory application; based on the physical attribute data and the real-time operation data, identifying a space conflict mode and a time conflict mode of the operation equipment to obtain collaborative conflict information; according to the production task data and the collaboration conflict information, carrying out collaboration conflict decoupling on the operation equipment to obtain a feasible collaboration mode; performing obstacle avoidance reinforcement learning on the operation equipment by using the smart factory scene graph to construct an obstacle avoidance global path of the operation equipment in the smart factory; and based on the feasible coordination mode and the obstacle avoidance global path, constructing a target coordination control scheme of the operation equipment so as to perform coordination control on the operation equipment. According to the invention, the reliability of multi-device cooperative control under the autonomous cooperation algorithm can be improved.
Owner:HARBIN SAISI TECH CO LTD

Multi-view-angle-oriented three-dimensional scene image reconstruction registration and optimization method and system

The invention provides a multi-view-oriented three-dimensional scene image reconstruction registration and optimization method and system, and relates to the technical field of image processing, and the method comprises the steps: obtaining multi-frame three-dimensional scene image data, extracting a multi-level feature set, constructing a cross-view-angle semantic association graph, building a feature corresponding relation, and calculating a multi-view-angle spatial transformation relation parameter. Performing coordinate system alignment on the image data to generate an initial three-dimensional reconstruction result, and performing optimization in combination with a multi-target joint optimization function and a dynamic adaptive weight regulation and control mechanism. According to the method, the precision and robustness of three-dimensional scene reconstruction are improved, and the problem of registration errors caused by large view angle difference in a complex scene is solved.
Owner:BEIJING SETTALL TECH DEV CO LTD

Map dynamic construction method and system fused with laser vision

The invention relates to the technical field of robot navigation, and discloses a laser vision fused map dynamic construction method and system, and the system comprises a data perception module, a feature analysis module, an information fusion module and a dynamic optimization module. The data sensing module collects an environment point cloud sequence and a scene image sequence through laser ranging equipment and a visual imaging device; the feature analysis module extracts spatial structure features of the point cloud and visual semantic features of the image; the information fusion module adopts a weight distribution strategy to fuse the multi-modal features to generate an initial map; and the dynamic optimization module adjusts the map position and attribute parameters in real time by using the incremental data. Through multi-modal fusion and dynamic optimization, the problems of insufficient precision, semantic deficiency and low dynamic updating efficiency of a traditional single sensor map are solved, and the method is suitable for scenes of robot navigation, automatic driving and the like.
Owner:XIAN DASHENG TECH CO LTD

Intelligent environment analogue simulation method and system based on artificial intelligence

The invention discloses an artificial intelligence-based intelligent environment simulation method and system, and relates to the technical field of artificial intelligence environment simulation, and the method comprises the steps: forming a scene tensor based on an updated initial scene and a POMDP belief state set, using an optimized VAE model to generate an initial image sequence, using a U-Net diffusion model to optimize the initial image sequence, and using a U-Net diffusion model to optimize the initial image sequence; the method comprises the steps of calculating reconstruction loss of a VAE encoder and diffusion denoising loss of a U-Net diffusion model, calculating a joint loss function of the VAE encoder and the U-Net diffusion model, generating a high-quality image sequence by using the joint loss function through a Monte Carlo simulation method, and optimizing the high-quality image sequence through a conditional diffusion model to generate a final scene image sequence. According to the method, efficient fusion of multi-modal data is achieved, the defect that in the prior art, vision and semantics are inconsistent is overcome, task specific parameters are rapidly generated through inner circulation, meta-model generalization ability is optimized through outer circulation, a model can be rapidly matched with an edge scene, and the generation efficiency and fidelity of a rare scene are remarkably improved.
Owner:BEIJING JINGSI XINCHUANG TECHNOLOGY CO LTD

Edge-deployed semi-supervised anomaly detection method and system for railway track foreign object

Disclosed in the present invention are an edge-deployed semi-supervised anomaly detection method and system for a railway track foreign object. The method comprises the following steps: an edge device encoding and decoding a video stream captured by a camera to obtain an image frame sequence, and performing frame extraction; and using a semantic segmentation model to perform image segmentation on a certain image frame obtained by means of frame extraction, to obtain a railway track region segmentation image. The use of a single image as input may generate an expert model result having a high weight value; however, the determination based on a single image is not stable, multiple consecutive images of the task scene need to be inputted, the frequency of each expert model obtaining the highest weight is computed, and the expert model corresponding to the highest frequency is the final solution. The present invention supports scene-adaptive foreign object detection algorithm automatic selection, and a user can perform selection on the basis of prior knowledge, or selection may be performed by a scene-adaptive automatic algorithm selection method; the user only needs to provide a batch of image data of the current scene, and the optimal algorithm selection can be evaluated.
Owner:GUANGZHOU EMBEDDED MACHINE TECH CO LTD

Systems and methods for multi-modal visual reasoning using multiple scene graphs

A system for processing multi-modal data representing an environment to generate scene graphs of the environment is described. The system can obtain sensor data associated with a vehicle operating in the environment. In examples, the system can determine a set of features from the sensor data, including one or more objects and one or more agents present in the environment, and can generate a scene graph that represents the poses and velocities of these objects and agents relative to the environment. In some examples, based on generating the scene graph, the system can generate a knowledge graph by encoding the relationships among the identified objects and agents. In some examples, the system can generate a control signal, using attributes that represent the states of objects and agents in the knowledge graph, and provide this control signal to the vehicle in order to adjust or cause the operation of the vehicle.
Owner:QPIAI INDIA PTE LTD

Electronic device and method for restoring scene image of target view

A device and method for performing scene restoration, including: obtaining an input image of an object; based on an input viewpoint corresponding to the input image, determining a plurality of augmented viewpoints surrounding the object in a three-dimensional (3D) space including the object; generating a plurality of augmented images at the plurality of augmented viewpoints, wherein each augmented image from among the plurality of augmented images corresponds to a view of the object from a corresponding augmented viewpoint from among the plurality of augmented viewpoints, and wherein each augmented image is generated based on an image at a different viewpoint using a view change model; generating a scene restoration model based on the input image at the input viewpoint and the plurality of augmented images at the plurality of augmented viewpoints; and restoring a scene image of a target view of the object using the scene restoration model.
Owner:SAMSUNG ELECTRONICS CO LTD

Intelligent accident liability judgment method and system fused with video tracking

The invention provides an accident liability intelligent determination method and system fused with video tracking, and the method comprises the steps: S1, carrying out the alignment evidence collection of a video, an IMU, a GPS, and an OBD, and forming an evidence chain through framing Hash, chain Hash, and a timestamp; s2, completing detection segmentation on the shared trunk multi-branch network; s3, in combination with appearance re-identification and a motion model, realizing cross-frame association, and outputting track, speed and shielding recovery; s4, mapping the trajectory to BEV according to depth and calibration, and constructing a space-time scene graph by fusing lanes, signals, speed limit and the like; s5, extracting events such as lane changing, line merging and signal rushing, and positioning a conflict point and a way giving relation; s6, generating error factors according to a rule base and priority judgment; s7, quantifying the responsibility proportion according to the weight, the collision participation time and the reaction time; and S8, generating a report evidence of the key frame and the rule list. The system is composed of an acquisition evidence obtaining module, a video understanding tracking module, a space-time reconstruction module, an event extraction module, a rule reasoning module, a responsibility quantification module and a report evidence storage module.
Owner:国任财产保险股份有限公司

Autonomous vision semantic navigation system and method based on interactive semantic mapping

The invention discloses an autonomous visual semantic navigation system and method based on interactive semantic mapping, a semantic scene graph is constructed by sensing the environment through a first visual angle of a robot and combining top-down environment semantic priori, and the method comprises regional spatial distribution identification, object attribute and layout analysis and room type clustering. And in combination with priori knowledge, complete scene representation is formed, the scene graph can be dynamically adapted and updated according to real-time observation of the robot, and the accuracy and timeliness of scene representation are ensured. Besides, a high-level and low-level AI agent module cooperation mechanism is adopted, and the high-level AI agent module is responsible for understanding and decomposing navigation tasks and distributing subtasks and supervising the execution process; the low-layer AI agent module is responsible for further dynamic decomposition and execution of the subtasks and feeds back execution results to the high-layer Agent, a complete task closed-loop monitoring mechanism is formed, and the execution efficiency and robustness of the navigation tasks are improved.
Owner:PAZHOU LAB (HUANGPU) +1

Robot operation object positioning method and system

The invention discloses a robot operation object positioning method and system, and relates to the image processing related technical field, and the method comprises the steps: building a three-dimensional point cloud data set based on laser scanning and design data, and enabling the three-dimensional point cloud data set to be marked with key features; activating a positioning camera of the robot, executing operation scene image acquisition, and establishing a scene image; carrying out contour matching identification on the RGB image according to the workpiece features, establishing a matching identification result, and calling a mapping depth image; executing search matching of the key features, and establishing a search matching result; local point cloud region fusion matching is executed, and a spatial registration result is established; and resolving attitude parameters of the target workpiece in the three-dimensional point cloud data set, and completing positioning based on a resolving result. The technical problems of poor adaptability and insufficient positioning precision of robot operation object positioning in a complex environment in the prior art are solved, and the technical effects of realizing high-precision positioning of the operation object and improving the positioning real-time performance and reliability are achieved.
Owner:ANHUI RONGZHOU INTELLIGENT TECH CO LTD

Target segmentation method and device for multi-view image, electronic equipment and storage medium

The embodiment of the invention provides a target segmentation method and device for a multi-view image, electronic equipment and a storage medium, and the method comprises the steps: obtaining scene images of different views obtained by a plurality of image collection devices which are deployed in a target scene in advance, and carrying out the three-dimensional Gaussian reconstruction, and obtaining a three-dimensional Gaussian point cloud corresponding to the target scene; in response to a segmentation prompt instruction corresponding to a target segmentation article in the scene articles, performing two-dimensional article segmentation on the scene image, and generating a corresponding binary segmentation mask; carrying out projection processing on the three-dimensional Gaussian point cloud to obtain projection planes of a plurality of visual angles, and carrying out Gaussian point filtering on each projection plane based on the binarization segmentation mask to obtain a filtered three-dimensional Gaussian point set; and based on the pose point cloud estimation information, the filtered three-dimensional Gaussian point set is sequentially splashed to device planes corresponding to a plurality of image acquisition devices to obtain a corresponding target segmentation result, so that the accuracy of object two-dimensional image segmentation in a target scene is greatly improved.
Owner:PENG CHENG LAB

Dynamic three-dimensional scene reconstruction and real-time rendering method and device based on free space-time Gaussian sputtering

The invention discloses a dynamic three-dimensional scene reconstruction and real-time rendering method based on free space-time Gaussian sputtering, and the method comprises the steps: firstly obtaining a multi-view-angle dynamic scene video, extracting cross-view-angle feature points for three-dimensional reconstruction, initializing free space-time Gaussian primitives, and carrying out the real-time rendering of the free space-time Gaussian primitives; an optimizable explicit motion function and a time opacity function are set for each Gaussian primitive and are used for representing the geometry and appearance of the dynamic three-dimensional scene; then sputtering and rendering the Gaussian primitive at the current moment based on the observation visual angle, and outputting a high-fidelity dynamic scene image of the corresponding visual angle; finally, Gaussian primitive parameters are optimized in a combined mode through a rendering loss function and a four-dimensional regularization strategy, and meanwhile low-influence Gaussian primitives are relocated periodically. According to the method, the movable Gaussian primitive is introduced at any position in space and time, the explicit motion function and the time opacity function are combined, a complex dynamic scene is effectively modeled, and efficient reconstruction and real-time rendering of the dynamic three-dimensional scene are achieved.
Owner:ZHEJIANG UNIV

Three-dimensional dynamic scene graph construction method based on 3D Gaussian representation

The invention discloses a three-dimensional dynamic scene graph construction method based on 3D Gauss, and belongs to the field of computer body intelligence. The implementation method comprises the following steps of: realizing object perception and semantic feature extraction of open vocabularies by utilizing a visual basic model; a 3D Gaussian scene with high fidelity and continuous object semantics is constructed through multi-view multi-dimension optimization of 3D Gaussian representation; constructing a multi-level three-dimensional scene graph, extracting spatial levels and semantic relationships among objects by using 3D spatial positions and semantic tags of instance objects existing in a semantic Gaussian graph, and constructing a multi-level spatial semantic topology to accurately represent an environment layout; according to the method for realizing local updating for the Gaussian scene graph based on the environmental structural similarity, environmental change detection is carried out through real-time RGB-D observation and the structural similarity between high-quality rendering views of the Gaussian scene graph, and corresponding local updating is carried out by using rapid training and differentiable rendering of 3D Gaussian representation. And the capability of adapting to a complex dynamic environment of the 3D Gaussian scene graph is improved.
Owner:BEIJING INST OF TECH

Planar splatting

Techniques are described for image processing. For example, a computing device can segment, using a first neural network, image(s) of a scene to determine respective segments for each of the image(s). The computing device can determine, using a second neural network, normal vectors for each of the image(s). The computing device can generate a graph based on each respective segment for each image, each respective normal vectors for each image, and estimated planar distances. The computing device can partition, based on the normal vectors and the estimated planar distances, the graph to determine indexes associated with Gaussian primitives. The computing device can assign, using linear regression, each descriptor of a plurality of descriptors to an index of the plurality of indexes based on a respective weight. The computing device can merge, using a Gaussian tree, Gaussian primitives of the Gaussian primitives with associated indexes that are similar to each other.
Owner:QUALCOMM INC

Multi-modality reinforcement learning in logic-rich scene generation

Generating high-quality images of logic-rich three-dimensional (3D) scenes from natural language text prompts is challenging, because the task involves complex reasoning and spatial understanding. A reinforcement learning framework utilizing a ground truth data set can be implemented to train a policy network. The policy network can learn optimal parameters to refine a text prompt to obtain a modified text prompt. The modified text prompt can be used to obtain a three-dimensional scene, and the three-dimensional scene can be rendered and projected to obtain a rendered image. The framework involves an action agent for text modification, a generation agent to produce rendered images, and a reward agent to evaluate the rendered images. The loss function used in training the policy network optimizes visual accuracy and quality of the rendered images and semantic alignment between the rendered images and the text prompt.
Owner:INTEL CORP

Remote sensing scene graph guided semantic information reasoning method and device, equipment and medium

The invention provides a semantic information reasoning method guided by a remote sensing scene graph, which can be applied to the technical field of remote sensing image processing. The method comprises the following steps: segmenting a remote sensing image to generate a scene segmentation image; executing target detection to generate an image block set; performing fine-grained analysis on the image global features and the image block set to generate a scene description text; performing grammar analysis, extracting a triple of objects, object attributes and relation information among the objects, and generating a remote sensing scene graph related to the problem; encoding the problem text into an embedded vector, projecting the embedded vector to a visual feature space, processing the embedded vector through a frozen self-attention layer and a learnable gating layer, and outputting an image-text interaction feature; carrying out attention fusion and double gating balance on the question coding features, and outputting a text code guided by a scene graph; and fusing the interaction features and the text codes, and reasoning local semantic information of the target area. The invention further provides a semantic information reasoning device and equipment guided by the remote sensing scene graph and a medium.
Owner:AEROSPACE INFORMATION RES INST CAS

Construction method of scene graph question and answer inference model based on reinforcement learning

The invention relates to the technical field of visual questions and answers, in particular to a scene graph question and answer inference model construction method based on reinforcement learning, which comprises the following steps: calling an existing large model interface, performing deep processing on an initial scene data set, and constructing a first batch of training sets; carrying out first-stage reinforcement learning training on a pre-constructed multi-modal large model by utilizing the first batch of training sets; obtaining an LLaVA-CoT data set, screening the LLaVA-CoT data set, and inputting the screened LLaVA-CoT data set into the multi-modal large model trained at the first stage to obtain a reasoning result; calling an existing large model interface to carry out accuracy evaluation and correction on the reasoning result of the multi-modal large model to obtain a second batch of training sets; and performing second-stage reinforcement learning training on the multi-modal large model by using the second batch of training sets to obtain a final scene graph question and answer inference model. According to the method, a staged and differentiated training strategy is adopted, the characteristic advantages of two batches of data are brought into full play, staged training is carried out on the large model, and the model performance is gradually improved.
Owner:DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD

Text-to-dynamic three-dimensional scene generation method and device

The invention provides a text-to-dynamic three-dimensional scene generation method and device. The method comprises the following steps: acquiring a first natural language description input by a user and a plurality of scene images shot at different angles; generating a static three-dimensional scene according to the first natural language description and the plurality of scene images; a second natural language description input by the user to the large language model is obtained, the second natural language description comprises relative scale information and constraint information of the multiple objects, and the relative scale information is used for describing actual physical sizes of the multiple objects in the static three-dimensional scene; the constraint information is used for describing physical conflicts existing in the moving process of the multiple objects; and according to the second natural language description and the static three-dimensional scene, generating reasoning motion tracks and reasoning sizes of the plurality of objects in the static three-dimensional scene so as to obtain a dynamic three-dimensional scene. Through the scheme, the authenticity and richness of the automatic driving simulation test are improved.
Owner:CHANGAN UNIV