Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11688results about "3D-image rendering" patented technology

Three-dimensional environment reconstruction optimization method based on multi-sensor fusion data

The invention discloses a three-dimensional environment reconstruction optimization method based on multi-sensor fusion data, and relates to the field of three-dimensional environment reconstruction optimization, and the three-dimensional environment reconstruction optimization method based on the multi-sensor fusion data comprises the following steps: S1, collecting multi-source sensor data, and constructing a data set under a unified coordinate system; s2, generating dense visual point cloud, and extracting laser point cloud features to construct a model; s3, establishing a local three-dimensional model, and generating a local environment image; s4, shadow parameters are extracted through shadow geometric analysis, and time sequence optimization is carried out; s5, consistency verification and correction are carried out, and three-dimensional reconstruction data are output; and S6, comparing the reconstruction data with the navigation map database, and carrying out map optimization updating. According to the method, time synchronization and space calibration are carried out on data acquired by the depth camera and the laser radar, complete and accurate three-dimensional information modeling of the target environment is realized, and the geometric precision of environment reconstruction and the image detail reduction capability are improved.
Owner:NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Virtual stylist

An example operation may include at least one of receiving, via a user interface of a device, an activation input from a user to initiate a session, capturing, by a camera of the device, a scan of a body of the user, wherein the capturing comprises recording at least one image and / or at least one video of the user, processing the at least one image and / or video to generate a three- dimensional model of the user comprising measurements and contours of the body, retrieving, from a database, at least one clothing item associated with the user, the at least one clothing item comprising dimensional attributes and texture attributes, rendering, by a graphics processing unit, the at least one clothing item onto the three-dimensional model to generate a visual representation, wherein the rendering simulates draping behavior, movement, and light interaction of the at least one clothing item relative to the three-dimensional model, and displaying, on the user interface, an interactive visualization comprising the visual representation of the three-dimensional model with the at least one clothing item from multiple viewing angles.
Owner:ELGORT PENELOPE

Three-dimensional dynamic scene reconstruction method and apparatus, and storage medium

The present disclosure relates to the field of computer vision and discloses a three-dimensional dynamic scene reconstruction method and apparatus, and a storage medium. The three-dimensional dynamic scene reconstruction method comprises: acquiring synchronized videos of a plurality of viewpoints of a dynamic scene; computing matching points between video images of different viewpoints, and estimating intrinsic and extrinsic parameters of each camera; obtaining a Gaussian splatting point set {p0} on the basis of a sparse point cloud constructed according to the depth of each matching point; for the first image frame of each video, using {p0} to perform static training thereon, to obtain a Gaussian splatting point set {p}; for the remaining image frames, dividing {p} into a static point set {S} and a dynamic point set {D}, performing dynamic training on {D}, and constructing a dynamic Gaussian splatting point set {P} from {p}, {S}, and the final {D}; and, in view of the intrinsic and extrinsic parameters of each camera, rendering {P} using a Gaussian splatting rendering pipeline, to obtain rendered images at different moments from new viewpoints.
Owner:TSINGHUA UNIVERSITY

Diamond high-strength micro-powder quality detection method and system based on artificial intelligence

The invention relates to the technical field of quality monitoring, and discloses a diamond high-strength micro-powder quality detection method and system based on artificial intelligence. The method comprises the steps of obtaining a two-dimensional projection image sequence of diamond micro-powder particles, calculating a projection matrix based on camera calibration parameters and geometric constraints, obtaining a multi-view image data set of the particles, establishing a pixel-level corresponding relation, extracting three-dimensional space coordinates of the surfaces of the particles, and reconstructing dense point cloud data of the particles. Establishing a local coordinate system based on the dense point cloud data, determining attitude parameters of particles in a three-dimensional space, if the attitude parameters deviate from a normal range, performing attitude compensation processing to obtain standardized point cloud data, and performing three-dimensional grid model construction on the standardized point cloud data; and calculating geometrical characteristic parameters of the particles based on the three-dimensional grid model, performing defect detection on the surfaces of the particles, and generating a crystal integrity evaluation report of the particles. The quality detection accuracy of the diamond high-strength micro-powder particles is improved.
Owner:ZHECHENG HAOXIN SUPERHARD PROD CO LTD

Insurance claim settlement-oriented multi-modal image video evidence analysis method and system

The invention discloses an insurance claim settlement-oriented multi-modal image video evidence analysis method and system. The method comprises the following steps of: acquiring video / image and multi-source data such as metadata, audio, IMU (Inertial Measurement Unit), GPS (Global Positioning System), OBD (On-Board Diagnostic) and the like; calculating content Hash of the video and the audio according to frames, connecting the content Hash with time information in series to form chained Hash, and adding a verification digital signature and a credible timestamp; realizing cross-modal time sequence alignment based on self-adaptive time anchor-attitude coupling; tampering detection is carried out in combination with PRNU fingerprints, noise field consistency, dual compression, copy-movement and the like; multi-view geometry and monocular depth are fused, IMU scale constraint and micro rendering are introduced, three-dimensional reconstruction and re-projection optimization are completed, and collision dynamics verification is carried out; and constructing an event cause and effect graph, judging responsibility in combination with traffic rules, outputting a confidence coefficient vector and a structured report, and generating a verifiable evidence packet. The scheme has the advantages of high efficiency and traceability in the aspects of space-time restoration and interpretable responsibility judgment.
Owner:国任财产保险股份有限公司

Tooth three-dimensional modeling system based on computer vision, computer equipment and readable storage medium

The invention relates to the technical field of tooth modeling, and discloses a three-dimensional tooth modeling system based on computer vision, computer equipment and a readable storage medium. According to the method, mirror reflection, diffuse reflection and subsurface scattering components in an original image are separated, mirror reflection intensity is normalized in combination with a dynamic truncation algorithm, pixel saturation is eliminated, groove and nest textures are reserved, a complete point cloud is obtained based on a two-dimensional texture image and cubic spline repair, and a multi-exposure point cloud sequence is obtained through bimodal calibration. The method comprises the following steps: solving the problem of data dislocation, carrying out weight assignment and data fusion on three-dimensional points in a plurality of exposure point cloud sequences to obtain three-dimensional fusion feature data, combining layered optical modeling and photon tracking compensation deviation, fusing clinical constraints, finally dynamically adjusting parameters, feeding back and optimizing, and outputting a micron-sized precision model. The modeling defect caused by difficulty in effectively coordinating feature contribution degrees under different exposure conditions is overcome, and high-precision modeling is realized.
Owner:SHENZHEN JINSHI LIMEI MEDICAL TECH CO LTD

Generative ai models for image rendering and inverse rendering

Embodiments of the present disclosure relate to rendering and inverse rendering using one or more generative models. “Rendering” refers to the process of generating a final visual image, video frame, or animation from a 2D or 3D model. “Inverse rendering” is a process that involves deducing or estimating the properties (e.g., material maps or other properties such as geometry, lighting, and textures) of a scene from observed images or visual data. Essentially, it aims to reverse the traditional rendering process. Various aspects of the present disclosure introduce editable light and material controls into generative models to allow for artistic creation. Various embodiments integrate generative models as a renderer for classic rendering pipelines to upcycle and enhance the style of rendered content.
Owner:NVIDIA CORP

Urban building three-dimensional automatic modeling and visualization method

The invention discloses an urban building three-dimensional automatic modeling and visualization method, and belongs to the technical field of building three-dimensional modeling. The method comprises the steps that point cloud data, high-resolution images and geographic information system data of urban buildings are acquired, data cleaning, registration and alignment are carried out, and preliminary building digital representation is formed; accurately segmenting each building, and identifying the contour and main structural features of the building; based on the data integrity and the building complexity, adaptively selecting a proper reconstruction strategy to carry out three-dimensional reconstruction; in the reconstruction process, the geometric structure is analyzed and optimized in real time, and potential topological problems are repaired; automatically generating missing details based on a predefined architectural style library and a component library, and performing material inference and texture mapping; a graph structure is used for representing the relation between the buildings, and the positions and orientations of the buildings are adjusted through a global optimization algorithm; a rendering engine supporting multi-level detail switching is developed, and smooth visualization and interaction of a large-scale city scene are achieved.
Owner:CHANGZHOU JINTAN DISTRICT LUOSUI TECHNOLOGY CO LTD

Three-dimensional scene reconstruction method and device based on large model geometric prior, and medium

The invention discloses a three-dimensional scene reconstruction method and device based on large model geometric prior, and a medium, and aims to solve the problems that a conventional 3DGS is liable to have artifacts and detail loss in geometric discontinuity, data redundancy and illumination variation scenes, and predicts a dense depth map and a normal map from a monocular image by using a pre-trained large model. The position and form of the Gaussian kernel are constrained as additional geometric priori; a primitive adjustment strategy based on kernel density estimation is introduced in the training stage, small Gaussian primitives with similar structures and adjacent spaces are combined into a large Gaussian primitive, the rendering quality is kept, redundancy is reduced, and the volume of the model is reduced; an exposure coefficient is adaptively estimated for each input image, an exposure compensation image loss function is constructed, and floating artifacts caused by illumination differences at shooting moments are eliminated. Experiments show that compared with the prior art, the method improves the three-dimensional reconstruction precision and real-time rendering quality of complex illumination and less-texture areas in a public data set and an unmanned aerial vehicle aerial photography scene.
Owner:NARI INFORMATION & COMM TECH

Commodity display interaction visualization method and device

The invention relates to the field of commodity visualization, in particular to a commodity display interaction visualization method and device. The method comprises the following steps: collecting a multi-azimuth image of a commodity, carrying out three-dimensional texture modeling, and constructing a three-dimensional texture mapping model; performing material light rendering on the three-dimensional texture mapping model to generate a material rendering result; collecting an environment detection image of a commodity display environment, and performing environment illumination adaptation compensation on a material rendering result to obtain an illumination compensation rendering commodity; carrying out attribute information visual layout on the illumination compensation rendering commodity to obtain a commodity visual space; and carrying out interaction response animation analysis according to the commodity visualization space, carrying out multi-target parallel rendering, and executing commodity interaction visualization operation. The form and surface details of the commodity in the real world are accurately restored, the visual reality sense is improved, and the interactive experience feeling of browsing the commodity by a user is enhanced.
Owner:SHENZHEN XIAOYI SHUZHI TECH CO LTD

Text-driven CAD modeling method and system based on diffusion and visual language model

The invention relates to the technical field of computer aided design, in particular to a text-driven CAD modeling method and system based on a diffusion and visual language model.The method comprises the steps that natural language text description is obtained, and CAD semantic features of the natural language text description are extracted; carrying out geometric standardization on the CAD semantic features by adopting a fine-tuning diffusion model, and generating a CAD view image conforming to engineering specifications; carrying out fusion by adopting a fine-tuned visual language model to generate a parameterized CAD construction sequence; a three-mode alignment mechanism is adopted, and the semantic consistency of the CAD semantic features, the CAD view images and the CAD construction sequences is checked; performing verification and post-processing on the CAD construction sequence, and outputting an executable Python code or STEP file; the CAD modeling method disclosed by the invention performs explicit modeling based on flexible modal description, and has the characteristics of high geometric constraint and high usability.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Interactive simulation teaching processing method and device for marine electromechanical safeguard system

The invention provides an interactive simulation teaching processing method and device for ship electromechanical guarantee, and relates to the technical field of interactive simulation teaching, and the method comprises the steps: building a ship electromechanical full-structure model in a virtual simulation platform, digitally mapping real equipment operation parameters into virtual parts with real-time response characteristics, and carrying out the real-time simulation of the virtual parts; synchronous acquisition and normalized calibration of operation instructions, equipment feedback and environment disturbance in a teaching scene are realized, and a unified teaching simulation data set is generated. State recognition, fault simulation and operation evaluation calculation are completed through multi-layer task scheduling, and real-time teaching feedback is generated and visually displayed on a terminal. Based on difference analysis of dynamic evaluation data and a standard operation template, operation deviation and abnormity are automatically identified, and a targeted guide prompt is generated. According to the invention, unified modeling and data linkage can be carried out on student operation behaviors, equipment dynamic responses and teaching feedback results, and the problems of teaching feedback lagging and lack of pertinence in operation correction are solved.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719 +1

High-resolution three-dimensional reconstruction method of fusion diffusion model

The invention discloses a high-resolution three-dimensional reconstruction method of a fusion diffusion model, which belongs to the technical field of image data processing, and comprises the following steps: constructing an original data set D; constructing an enhanced training set; constructing a three-dimensional reconstruction network which comprises a text encoder, a renderer, a VAE encoder, a conditional diffusion model, a VAE decoder and an MVS module; training and fine-tuning the conditional diffusion model in three stages to obtain a three-dimensional reconstruction model, acquiring an image sequence and a text instruction of a scene to be reconstructed, and performing reconstruction by using the three-dimensional reconstruction model. According to the method, highly consistent geometric and color reduction can be kept under the multi-view condition, and splicing artifacts are remarkably reduced. Through semantic guidance optimization, texture details and structural consistency of the reconstruction model are greatly improved. Conditional diffusion sampling enables the model to accurately restore local details in a complex scene, and the stability of real-time rendering is improved.
Owner:SHENZHEN SENSING DATA TECH CO LTD +1

Cross-source data three-dimensional reconstruction method and system based on improved Gaussian sputtering

The invention discloses a cross-source data three-dimensional reconstruction method and system based on improved Gaussian sputtering, and the method comprises the steps: collecting an unmanned plane inclined image and a ground panoramic image of a target region, and constructing a time-space correlation data set; based on multi-view geometric constraints, space-time coding matching point pairs are established through an adaptive feature pyramid, intelligent incremental cross-source data sparse reconstruction is carried out, and point cloud and camera parameters are output; adopting improved Gaussian sputtering, compressing a three-dimensional Gaussian kernel into a two-dimensional Gaussian primitive through double tangent vector constraint, and fitting surface geometry to realize multi-scale reconstruction; and optimizing primitive parameters by using a differentiatable renderer, completing multi-scale fine reconstruction through gradient back propagation, and generating a high-precision three-dimensional model. According to the method, multi-scale accurate geometric prior input and accurate camera poses are provided for three-dimensional reconstruction, the dependence on professional manual operation in a traditional three-dimensional reconstruction method is greatly reduced, and meanwhile, the geometric accuracy and visual fidelity of a reconstruction result are remarkably improved.
Owner:HANGZHOU INST FOR ADVANCED STUDY UCAS

Patient registration for total hip arthroplasty procedure using pre-operative computed tomography (CT), intra-operative fluoroscopy, and / or point cloud data

ActiveUS12507972B2Image enhancementImage analysisPelvic regionPatient registration
A system for computer assisted navigation during surgery includes a computer platform that operates to register a target surgical area of a patient. In certain cases, a process includes: obtaining a pre-op CT image of a pelvic region of a patient and intra-operatively obtaining a point cloud data about the pelvic region with a navigated instrument, generating a 3D bone model which excludes non-targeted area such as a femur, and then merging the 3D bone model to the point cloud to register the target surgical area.
Owner:GLOBUS MEDICAL INC

Scene rendering method and system based on three-dimensional Gaussian splashing

The invention discloses a scene rendering method and system based on three-dimensional Gaussian splashing. The method comprises the following steps: organizing and constructing an original 3D Gaussian set to obtain a spatial hierarchical structure; generating corresponding level details for the primitives in the spatial hierarchical structure to obtain a spatial hierarchical structure associated with the level details; traversing the spatial hierarchical structure associated with level details, and executing hierarchical view cone cutting and shielding elimination to obtain a visible node list; traversing the visible node list, calculating a level detail selection standard, and generating a level detail activity Gaussian set; performing optimization sorting on the level detail activity Gaussian set to obtain an activity primitive list; and performing tile-based rasterization on the active primitive list, and performing adaptive processing according to level details of the primitives to obtain a final color value of each pixel. According to the method, the rendering performance can be greatly improved, the occupation of a memory and a video memory is remarkably reduced, the rendering quality is improved, visual flaws are reduced, and the expandability of the 3DGS rendering method is enhanced.
Owner:CHINA ORDNANCE SCI INST

Gaussian representation SLAM method based on dense matching prior and factor graph constraint

The invention discloses a Gaussian representation SLAM method based on dense matching priori and factor graph constraint, which comprises the steps of inputting a current image and a key frame image, outputting a point graph corresponding to the image through a pre-trained model, returning a matching condition of two frame image points and respective point cloud information, and obtaining a point-level matching result based on a point-level matching result. The method comprises the following steps: constructing a joint optimization problem of a current frame and a key frame by taking a luminosity consistency error and a geometric projection error as targets, performing joint estimation on a camera pose and a point cloud of the current frame, realizing high-precision pose solution, generating point diagram data after Gaussian scene representation and rasterized rendering processing, and transmitting the point diagram data to a rear end for global optimization. And the rear end receives the pose and point cloud data, executes loopback detection to identify repeated key frames, and performs Gaussian rendering through an optimized key frame image to complete global dense three-dimensional reconstruction. The method effectively solves the problem of track drift and scene inconsistency caused by lack of pose priori and global geometric constraints in an existing system.
Owner:HANGZHOU DIANZI UNIV

Three-dimensional gaussian splatting optimization method for unposed input

Disclosed in the present invention is a three-dimensional Gaussian splatting optimization method for an unposed input. The method comprises: for an input image, predicting a ray bundle distribution by using a ray prediction model, to obtain distribution features in the form of ray bundles; on the basis of the distribution features in the form of ray bundles, calculating a camera pose; for the ray bundle distribution, performing sampling on the basis of the volumetric density of light rays, to obtain an initial spatial distribution of a three-dimensional Gaussian point cloud focused on a visual center area; for the input image, obtaining a visible shell by means of view frustum projection and object-mask computation; on the basis of the initial spatial distribution of the three-dimensional Gaussian point cloud and the visible shell, performing three-dimensional Gaussian splatting scene training, to obtain a three-dimensional scene reconstruction model satisfying a preset loss function criterion, the loss function comprising a training regularization term for a camera pose parameter. The present invention provides important scene initialization information for three-dimensional Gaussian splatting training, and significantly improves the quality and the richness of detail of the final three-dimensional structure.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Method and system for generating 3D (three-dimensional) human motion under text driving by using 2D (two-dimensional) video

The invention discloses a method and a system for generating 3D (three-dimensional) human motion under text driving by utilizing a 2D (two-dimensional) video. The method comprises the following steps of: acquiring the video and preprocessing to obtain a two-dimensional key point sequence and text description; the two-dimensional key point sequence passes through a spatiotemporal feature adapter to obtain a potential spatiotemporal feature sequence, a residual vector quantizer quantizes and outputs a three-dimensional SMP L parameter sequence, and meanwhile, potential spatiotemporal features and a discrete Token sequence are mapped; preprocessing a text to extract a semantic vector, partially covering a Token sequence of a basic quantization layer, reconstructing a prediction sequence through a predictor in combination with the semantic vector, and obtaining a complete sequence through a refiner; constructing a total loss function and a text-to-action loss function to train the module; and inputting the text description and the basic quantization layer Token to a trained module, outputting a three-dimensional SMPL parameter sequence, and rendering to generate a three-dimensional human body grid and animation. According to the method, the end-to-end generation from the text to the three-dimensional SMPL action is realized only by two-dimensional key points and text description.
Owner:ZHEJIANG UNIV

Patient Registration For Total Hip Arthroplasty Procedure Using Pre-Operative Computed Tomography (CT), Intra-Operative Fluoroscopy, and / Or Point Cloud Data

PendingUS20250384569A1Image enhancementImage analysisPelvic regionPatient registration
A system for computer assisted navigation during surgery includes a computer platform that operates to register a target surgical area of a patient. In certain cases, a process includes: obtaining a pre-op CT image of a pelvic region of a patient and intra-operatively obtaining a point cloud data about the pelvic region with a navigated instrument, generating a 3D bone model which excludes non-targeted area such as a femur, and then merging the 3D bone model to the point cloud to register the target surgical area.
Owner:GLOBUS MEDICAL INC

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Model-free six-dimensional object pose estimation

A composite pose-estimation algorithm includes a video-object segmentation sub-algorithm (311) configured to determine a mask of a visual object in an image, and an object-pose tracking sub-algorithm (312) configured to track a pose of a visual object over multiple depth-video frames, wherein the pose-estimation algorithm is configured to input a depth video, from which frames are extracted and fed to the video-object segmentation sub-algorithm, which determines respective object masks to be used by the object-pose tracking sub-algorithm alongside the depth video. A method of tracking a pose of a physical object comprises: obtaining a depth video depicting a physical object in a plurality of poses from an input interface (330); forming a storable data item representing the physical object by applying the pose-estimation algorithm to the depth video; and tracking the physical object or a copy thereof using an instance of the pose-estimation algorithm which has been initialized by means of the storable data item.
Owner:ABB (SCHWEIZ) AG

Three-dimensional scene reconstruction method based on intelligent LED street lamp multi-mode sensor

The invention discloses a three-dimensional scene reconstruction method based on an intelligent LED street lamp multi-mode sensor. The three-dimensional scene reconstruction method comprises the following steps that RGB images are obtained and preprocessed; the information is input to a visual feature coding module, two-dimensional bounding box information is extracted, and an object segmentation module is guided to output a two-dimensional segmentation mask; obtaining point cloud data, projecting the point cloud data to the standardized RGB image, and screening target points in combination with the two-dimensional segmentation mask; complementing the preliminary segmentation result of the point cloud, mapping the result to a standardized RGB image, and extracting a pixel region; carrying out joint coding, implicit representation and neural decoding processing on the object-level RGB image and the point cloud complete segmentation result; and fusing into an original three-dimensional scene, and completing spatial restoration through point cloud registration, attitude optimization and semantic constraint. The invention provides an efficient three-dimensional scene reconstruction method in combination with a multi-mode sensor of an intelligent LED street lamp, and the method has high precision, real-time performance and dynamic target processing capability.
Owner:ZHEJIANG UNIV +1

Coal mine operation and maintenance monitoring method and device, electronic equipment and storage medium

The invention provides a coal mine operation and maintenance monitoring method and device, electronic equipment and a storage medium. Environmental parameters and equipment state data are collected in real time through the Internet of all things deployed in an underground coal mine; preprocessing the environment parameters and the equipment state data by utilizing an edge computing node; the method comprises the following steps: acquiring point cloud data and a multi-view image of an underground coal mine, and splicing the point cloud data and the multi-view image to construct a three-dimensional scene model; the preprocessed data and the three-dimensional scene model are dynamically fused, and the equipment state features, the environment features and the risk early warning parameters are displayed in the three-dimensional scene model in a three-dimensional visualization mode in an overlapping mode; under the abnormal condition, abnormal data are received, maintenance steps or fault points are marked in a virtual interface of the monitoring center, data islands are broken, and multi-source data are deeply fused; and real-time guidance can be provided through interaction in modes of gestures, voice and the like, so that the maintenance efficiency is improved.
Owner:BEIJING TIANMA INTELLIGENT CONTROL TECHNOLOGY CO LTD +1

Three-dimensional scene optimization method and system for collaborative rendering of dynamic LOD and view cone elimination based on space-time prediction

The invention relates to the field of scene rendering, and particularly discloses a three-dimensional scene optimization method and system for collaborative rendering of dynamic LOD and view cone rejection based on space-time prediction, and the method comprises the steps: dynamically calculating the visibility frequency of an object, and dynamically adjusting the loading and rendering modes of objects with different priorities; predicting a view cone range of multiple frames in the future by using an LSTM network architecture; dividing the scene into uniform grid blocks, and then performing grading elimination; optimizing a heterogeneous computing pipeline; rendering the execution process; the system comprises a dynamic LOD and visual cone rejection collaborative optimization module, an LSTM visual cone prediction module, a block-level mixed rejection module, a heterogeneous calculation pipeline module, an optimization CPU-GPU task allocation and data transmission module and a rendering execution control module. According to the method, the technologies of collaborative optimization of dynamic LOD and view cone removal, block-level mixed removal, heterogeneous calculation assembly line optimization and the like are adopted, so that unnecessary rendering calculation and resource loading are reduced, and the rendering efficiency is improved.
Owner:YANTAI JIERUI NETWORK TRADING

Garden landscape design method based on digital twinning

The invention discloses a garden landscape design method based on digital twinning, and belongs to the technical field of landscape garden design. According to the method, a target site environment data set is collected, and a site thermodynamic distribution cloud picture is generated through a geographic space grid processing unit; extracting a site space pattern feature matrix on the basis; constructing a landscape effect distribution model, decomposing the feature matrix into topographic relief, seasonal phase color and illumination reflection parameters, driving the model to generate a three-dimensional scene rendering sequence, embedding the model into a three-dimensional digital twin scene of the garden site, and associating the parameters with physical attributes of the twin scene in real time; dynamic optimization of the landscape scheme is executed, an element replacement instruction is generated, and the spatial topological relation network is updated; and outputting a site design scheme map, adjusting facility layout coordinates, and generating a final scheme data packet in combination with the path connectivity, so that the precision and adaptability of garden landscape design can be improved.
Owner:JINAN TINGYING INTELLIGENT EQUIP TECH CO LTD +2

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Cross-platform virtual-real fusion scene construction method and system based on AI space calculation

The invention discloses a cross-platform virtual-real fusion scene construction method and system based on AI space calculation, and relates to the technical field of artificial intelligence and space calculation, and the method comprises the steps: carrying out the multi-scale feature fusion based on a received cross-modal conversion instruction set, and generating an initial image sequence; carrying out implicit field coding on a target object by combining a three-dimensional reconstruction algorithm to obtain an initial parameterized model; performing space-time alignment on the multi-view video stream, loading a digital scene asset package in combination with physical sensing data and a preset spatial index structure, and establishing a bidirectional data channel between a virtual scene and a physical sensor; performing rendering and illumination parameter adjustment on the initial parameterized model to obtain an optimized parameter model; performing differential coding processing on the optimization parameter model to obtain a target virtual-real scene fusion model; and distributing the target virtual-real scene fusion model to a preset terminal. The invention provides a virtual-real fusion construction method for end-to-end collaborative optimization, which is suitable for cross-platform live broadcast or dynamic interaction scenes.
Owner:ZHONGJING TECH (GUANGZHOU) CO LTD