Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

384 results about "Morphing" patented technology

Morphing is a special effect in motion pictures and animations that changes (or morphs) one image or shape into another through a seamless transition. Traditionally such a depiction would be achieved through cross-fading techniques on film. Since the early 1990s, this has been replaced by computer software to create more realistic transitions.

Rendering model training method and device, illumination rendering method and device and equipment

The invention relates to the technical field of computers, and relates to a rendering model training method and device, an illumination rendering method and device, a computer program product and electronic equipment. The training method comprises the following steps: extracting inherent attribute information and global light and shadow information of a to-be-rendered object from an image frame sample based on a teacher model; performing feature decoding on the Gaussian primitives by using an initial student model to obtain a basic physical attribute, and determining a first loss according to the basic physical attribute and the inherent attribute information, the initial student model comprising a plurality of Gaussian primitives attached to a deformable grid; interpolating light and shadow data of the probe model through the initial student model to obtain current light and shadow information, constructing second loss according to the current light and shadow information and global light and shadow information, and arranging the probe model on the surface of the deformable grid; and adjusting parameters in the initial student model according to the first loss and the second loss, and adjusting shadow data of the probe model to obtain a target student model and a target probe model.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

Texture mapping method and device, storage medium and electronic equipment

The invention provides a texture mapping method, a texture mapping device, a computer storage medium and electronic equipment, and relates to the technical field of computer graphic processing. The method comprises the following steps: acquiring a world space coordinate and a world space normal of a vertex of a target model; performing texture projection on the world space coordinates on a plurality of world coordinate axes to obtain target texture coordinates of the vertexes of the target model on the world coordinate axes; determining a texture weight value based on the included angles between the world space normal and the plurality of world coordinate axes; and carrying out texture sampling according to the target texture coordinate to obtain an initial map sampling result, and carrying out mixed sampling on the initial map sampling result and the texture weight value on the corresponding world coordinate axis to carry out texture mapping on the to-be-processed model based on the target map sampling result. The method can get rid of dependence on a model topological structure and UV expansion of a three-dimensional model, and realizes a more real, natural, accurate and flexible texture fitting effect without deformation stretching.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

6D pose estimation method and device fusing attention mechanism, equipment and medium

The invention discloses a 6D pose estimation method and device fusing an attention mechanism, equipment and a medium, and relates to the technical field of object space poses, and the method comprises the steps: processing an RGB image and a depth image of a target object through a preset multi-modal feature fusion model, the model comprises a semantic segmentation module, a feature extraction module, a feature fusion module, an attitude estimation module and an attitude iterative optimization module. A target mask point cloud is obtained through semantic segmentation, feature fusion is performed by using a cross attention mechanism and deformable convolution, and a final pose is obtained through pose estimation and iterative optimization, so that the problems of insufficient feature extraction and weak multi-modal feature association in complex scenes such as weak texture and shielding are solved, and the accuracy and robustness of pose estimation are improved.
Owner:湖南工商大学

Long-tail data generation method based on feed-forward stylized reconstruction

The invention discloses a long-tail data generation method based on feed-forward stylized reconstruction. The method comprises the following steps: constructing multi-modal geometric perception input and feature codes; reconstructing a semantic guidance scene based on a feature-level linear modulation mechanism; the invention relates to 3D consistency controllable style migration based on a feedforward reconstruction model. According to the invention, a feedforward reconstruction network based on DINOv2 improvement is constructed. Multi-modal adaptation is performed on a backbone network input layer, and decoupled high-level semantic features are explicitly injected by using a feature-level linear modulation mechanism, so that the geometric accuracy of automatic driving scene reconstruction is improved. On the basis of a three-dimensional scene of feedforward reconstruction, a distillation strategy is adopted, and the generation capability of a 2D diffusion model is migrated to a 3D Gaussian field. According to the method, the phenomena of object deformation, disappearance and the like in the generated data are effectively avoided, and long-tail weather data with geometric consistency of rain, snow and the like can be generated through the text instruction.
Owner:DALIAN UNIV OF TECH

Deforming real-world object using image warping

Methods and systems are disclosed for performing real-time deforming operations. The system receives an image that includes a depiction of a real-world object. The system applies a machine learning model to the image to generate a warping field and segmentation mask, the machine learning model trained to establish a relationship between a plurality of training images depicting real-world objects and corresponding ground-truth warping fields and segmentation masks associated with a target shape. The system applies the generated warping field and segmentation mask to the image to warp the real-world object depicted in the image to the target shape.
Owner:SNAP INC

Deforming real-world object using image warping

Methods and systems are disclosed for performing real-time deforming operations. The system receives an image that includes a depiction of a real-world object. The system applies a machine learning model to the image to generate a warping field and segmentation mask, the machine learning model trained to establish a relationship between a plurality of training images depicting real-world objects and corresponding ground-truth warping fields and segmentation masks associated with a target shape. The system applies the generated warping field and segmentation mask to the image to warp the real-world object depicted in the image to the target shape.
Owner:SNAP INC

Method for reducing motion blur of erf model by using event and frame

PendingCN120912819AImage enhancementImage analysisCamera response functionMorphing
The invention relates to a method for reducing motion blur of an erf model by using events and frames, and belongs to the technical field of computer vision. The method comprises the following steps: converting an event stream into triple representation; generating a bidirectional optical flow field based on the event flow, and aligning the fuzzy frame and the event flow through event-guided deformation convolution; constructing a double-branch optimization framework; the double-branch features are dynamically fused through frequency domain attention gating, and the weight of the frequency domain attention gating is generated by event frequency spectrums through MLP; generating a pseudo-true value based on an event double integral model, aligning an event spectrum and rendering a high-frequency component, and carrying out chain constraint on the continuity of a cross-frame radiation field through event brightness change; event-frame non-linear response differences are modeled by a learnable camera response function. Clear reconstruction and rendering of the moving target in a complex dynamic scene are realized.
Owner:SHANGHAI DECHENG DATA TECHNOLOGY CO LTD

Semantic segmentation method based on phase contour constraint

The invention discloses a semantic segmentation method based on phase contour constraint, and belongs to the technical field of computer vision. The method comprises the following steps of: acquiring a two-dimensional texture image and a deformed stripe image of a scene to be detected in the same view field by using a calibrated camera-projector system; resolving the deformed fringe image through a multi-step phase shift and Gray code technology to obtain absolute phase distribution, and calculating a gradient field; according to a preset sudden change condition, extracting a physical edge contour reflecting the depth change of the object from the gradient field; and mapping the physical edge contour into the two-dimensional texture image based on a pixel coordinate corresponding relation to accurately define a semantic object region, endowing a corresponding category label, and finally generating a semantic label mask image. According to the method, the uncertainty of the visual texture edge is corrected by utilizing the physical phase information, the problem of edge extraction in a complex illumination scene is effectively solved, automatic generation of semantic segmentation data is realized, and the training precision and robustness of a semantic segmentation model are improved.
Owner:NANJING NANXUAN HEYA TECH CO LTD

Real-time software deformation interaction method based on physical engine in environment

PendingCN121960046AReduce the amount of synced dataGuaranteed deformation accuracyDesign optimisation/simulationCollision detectionInstruction stream
The invention discloses a real-time software deformation interaction method based on a physical engine in an environment. The real-time software deformation interaction method comprises the following steps: S1, establishing a virtual environment and a software model; s2, constructing an interactive perception and collision detection network; s3, receiving and analyzing a multi-modal interaction instruction; s4, real-time deformation calculation based on a physical engine; s5, performing multi-user collaborative state synchronization and rendering feedback; and S6, carrying out fault state simulation on the software object. According to the method, real-time and high-precision simulation of software deformation is realized through deep fusion of software dynamics calculation of a physical engine and an interaction instruction stream, and through hierarchical physical attribute binding and client prediction, the network synchronization data volume is greatly reduced while the millimeter-level deformation precision is ensured, so that the network synchronization efficiency is improved. And low-delay and high-consistency experience under multi-person cooperative operation is ensured.
Owner:BEIJING JUNHE CHUANGXIANG TECH DEV CO LTD

Photorealistic 4d scene generation using video diffusion models

A method for generating photorealistic 4D scenes from text inputs is disclosed. The method utilizes a text-to-video diffusion model to generate a reference video and a freeze-time video. A canonical 3D representation is reconstructed using deformable 3D Gaussian Splats (D-3DGS) based on the freeze-time video. Temporal deformations are learned to capture dynamic interactions in the reference video. The method employs a novel Score Distillation Sampling strategy combining multi-view and temporal aspects to enhance consistency and robustness. The resulting 4D scenes feature multiple objects interacting with detailed background environments, viewable from different angles and times. The method enables flexible camera control and integration with augmented and virtual reality applications. Some examples include features such as image-to-4D generation.
Owner:SNAP INC

Systems and methods for diffusion-based facial performance relighting

The disclosed computer-implemented method may include receiving, by a computing device, multi-view flat-lit performance data of a subject. Additionally, the method may include rendering, by the computing device, a dynamic sequence of novel-view flat-lit images of the subject based on a deformable three-dimensional Gaussian splatting (3DGS) model. The method may also include providing the rendered dynamic sequence of flat-lit images as input to a diffusion-based relighting model trained on the multi-view flat-lit performance data of the subject. Furthermore, the method may include generating, by the computing device using the diffusion-based relighting model, a relit sequence of the subject under a specified lighting condition. Various other methods, systems, and computer-readable media are also disclosed.
Owner:NETFLIX INC

Text-guided creative and time consistency-oriented video coloring method and system

The invention relates to the technical field of computer vision, in particular to a text-guided creativity and time consistency-oriented video coloring method and system. By introducing a time deformable attention block, according to the position and shape change dynamic state of an object in a video frame, dynamic features of an instance are captured in the time dimension, so that feature representation consistency is kept, and color flickering and shifting in the coloring process are effectively prevented; semantic representation of a text noun concept is adjusted through a cross-modal pre-fusion module, mask cross attention in the module is used for enhancing understanding of the model on the noun concept and instance perception, and color distribution accuracy is improved; the global structure of the colored video is maintained through gray level video information introduced in the video decoding process; a cross-fragment fusion mechanism is introduced during reasoning, so that the long-term consistency of video coloring is maintained; through text guidance, the user can color the video according to own demands and creativity.
Owner:PEKING UNIV

Efficient warping-based neural video codec

An example computing device may include memory and one or more processors. The one or more processors may be configured to parallel entropy decode encoded video data from a received bitstream to generate entropy decoded data. The one or more processors may be configured to predict a motion vector based on the entropy decoded data. The one or more processors may be configured to decode a motion vector residual from the entropy decoded data. The one or more processors may be configured to add the motion vector residual and motion vector. The one or more processors may be configured to warp previous reconstructed video data with an overlapped block-based warp function using the motion vector to generate predicted current video data. The one or more processors may be configured to sum the predicted current video data with a residual block to generate current reconstructed video data.
Owner:QUALCOMM INC

Texture coordinate generation method and device of curved surface three-dimensional model, electronic equipment and storage medium

The embodiment of the invention provides a texture coordinate generation method and device of a curved surface three-dimensional model, electronic equipment and a storage medium. The method comprises the following steps: acquiring three-dimensional coordinate information of a model vertex; determining curved surface geometric feature parameters corresponding to the vertexes based on the three-dimensional coordinate information; determining a plane mapping conversion relation based on the curved surface geometric feature parameters; and generating texture coordinates of the model by applying the plane mapping conversion relation. According to the technical scheme, the accurate plane mapping conversion relation is established based on the geometric feature parameters of the curved surface, the problem of pole deformation in traditional spherical UV mapping is effectively solved, seamless mapping of the square continuous mapping on the curved surface model is achieved, meanwhile, the calculation load of a graphics processor is reduced, and the processing efficiency of texture mapping is improved.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

Personalized replacing, cutting and deformation rendering method and system for design file

ActiveCN121392102AFile system administration3D-image renderingPersonalizationTransparency (graphic)
The invention relates to the technical field of graphic image processing, and discloses a personalized replacement cutting deformation rendering method and system for a design file, and the method comprises the steps: 1, analyzing the design file, binding an external material with a target layer according to a replacement mapping table, and obtaining an adaptive material texture; 2, performing normalization deformation to form a single inverse mapping rule; step 3, sampling according to a single inverse mapping rule and applying a mask to obtain a clipped layer; 4, setting a synthesis area for the cutting group, and pre-multiplying transparency accumulation to obtain local synthesis textures of the cutting group; step 5, pasting the texture back to the main canvas and combining the texture with the non-group layer according to a mixed mode; step 6, determining a dirty rectangular region according to propagation change of the layer dependency graph, and generating a multiplexing identifier and a texture fingerprint; and step 7, outputting the main canvas and exporting the texture and metadata. According to the invention, stable, efficient and traceable reuse of integrated online rendering of personalized replacement, cutting and deformation of design files is realized.
Owner:XIAMEN FINGERPRINT TECH CO LTD

Coarse prediction driven context enhancement for joint multi-modal sensor representation learning

Certain aspects of the present disclosure provide techniques for coarse-to-fine attention-based sensor fusion. The method includes obtaining a 3D voxel image space of an environment, the 3D voxel image space comprising first features of the environment extracted from a plurality of images; obtaining a point cloud corresponding to the environment, the point cloud comprising points, wherein the points are labeled with second features; generating a coarse representation of the environment, the coarse representation comprising a projection of the first features onto the points of the point cloud, wherein the projection is based on combining a respective set of the first features within a first radius from a point with a respective set of the second features corresponding to the point; applying a deformable attention module to predict fine sampling locations and extract fine features from the first features; and generating, with an attention module, a fine representation of the environment.
Owner:QUALCOMM INC

Three-dimensional model dynamic rendering method and storage medium

The invention discloses a three-dimensional model dynamic rendering method and a storage medium, and the method comprises the steps: obtaining three-dimensional data of a target scene, obtaining a plurality of voxels, and enabling each voxel to correspond to a plurality of voxel attributes and an initial comprehensive displacement vector; in any voxel in any time step, obtaining a comprehensive displacement vector of the voxel in the current time step according to a preset physical deformation rule and / or animation curve deformation rule, and performing coordinate transformation on the voxel; after the coordinate transformation, updating the voxel attribute of the voxel corresponding to the current time step length; and performing real-time three-dimensional rendering on the target scene according to the vertex normal of each voxel at the current time step length and the voxel attribute. The physical consistency is ensured in the dynamic deformation process by synthesizing the displacement vector; by updating the voxel attributes, the voxel attributes are kept in smooth transition after deformation; and high-quality real-time rendering of the target scene is realized by recalculating the deformed vertex normal and combining the illumination model.
Owner:HUANENG LANCANG RIVER HYDROPOWER CO LTD

Deformed character typesetting rendering method, device and equipment based on path

The invention discloses a path-based deformed character typesetting rendering method, device and equipment. The method comprises the following steps: receiving character content input by a user and a corresponding deformation strength parameter; the text content is transversely typeset by adopting non-automatic line folding; according to the deformation strength parameters, the curvature radius of the target path is dynamically calculated through a Newton iteration method; the origin coordinates of the character content characters are mapped to the target path, and the rotation angle of each character is calculated; when the total width of the characters exceeds the perimeter of the target path, automatically generating a multi-circle path and distributing the characters; a final graph is output through a cross-platform rendering engine, and the display consistency of all platforms is ensured. According to the method and the device, the problems of inconvenience in path adjustment, cross-platform inconsistency and incapability of processing super-long characters in path-based character typesetting in the prior art can be solved. By means of the method, a designer only needs to adjust the strength value, the bending effect of the characters along the dynamic path can be previewed in real time, all devices and browsers are compatible, and the design efficiency is greatly improved.
Owner:BEIJING YIYUANKU TECH CO LTD

Gaussian splatting with neural spline deformation

One embodiment of the present invention sets forth a technique for determining a time-varying deformation associated with a scene. The technique includes matching a query time to a time interval associated with the scene and generating, via execution of a machine learning model, (i) a first set of attributes associated with a set of canonical coordinates in the scene at a starting time of the time interval and (ii) a second set of attributes associated with the set of canonical coordinates at an ending time of the time interval. The technique also includes computing a third set of attributes associated with the set of canonical coordinates at the query time based on a spline interpolation associated with the first and second sets of attributes. The technique further includes generating a representation of the scene at the query time based on the third set of attributes.
Owner:ETH ZURICH +1

Model training method, video encoding method and decoding method

Embodiments of the present application provide a model training method, a video encoding method and a decoding method. The model training method comprises: obtaining a reference sample frame and a plurality of continuous to-be-encoded sample frames; performing morphing processing on the reference sample frame by a generator in an initial generation model to generate a reconstructed sample frame; inputting each reconstructed sample frame and the corresponding to-be-encoded sample frame into a first discriminator in the initial generation model to obtain a first discrimination result; splicing the to-be-encoded sample frames in chronological order to obtain a spliced to-be-encoded sample frame, and splicing the reconstructed sample frames to obtain a spliced reconstructed sample frame; inputting the spliced to-be-encoded sample frame and the spliced reconstructed sample frame into a second discriminator in the initial generation model to obtain a second discrimination result; obtaining an adversarial loss value based on the first discrimination result and the second discrimination result, and training the initial generation model based on the adversarial loss value. The present application maintains the consistency of the reconstructed video frame sequence and the to-be-encoded video frame sequence in the time domain, and improves the reconstruction quality.
Owner:ALIBABA (CHINA) CO LTD

Generative animatable gaussian avatar

Animation systems including an expressive deformation model configured to transform expression settings, pose settings, and a template mesh into an animatable mesh, a first generator branch configured to transform identity controls for the animatable mesh into base Gaussian attributes, a second generator branch configured to transform detail controls for the animatable mesh into residual Gaussian attributes, the system configured to embed the base Gaussian attributes and residual Gaussian attributes in UV maps of the animatable mesh and to combine the UV maps and the animatable mesh to form an animatable Gaussian representation of an object to animate.
Owner:NVIDIA CORP

Morphing of Watertight Spline Models Using As-Executed Manufacturing Data

Methods, computer systems, and computer-readable memory media for determining a warp function. An as-designed watertight spline model of an object is received. A point cloud and the as-designed watertight spline model are used to construct a model of the object. The point cloud is obtained from a physical or virtual (simulated) inspection and / or manufacturing process. A warp function is determined based on a difference between the as-designed watertight spline model and the constructed model. The warp function is a continuous function quantifying differences between the as-designed model and the constructed model. As-preprocessed instructions for a simulation or analysis process of the object are determined based on metadata of the as-designed watertight spline model and the warp function. The simulation or analysis process is performed on the object according to the as-preprocessed instructions to produce as-simulated data, and the as-simulated data is stored in a non-transitory computer-readable memory medium.
Owner:NVARIATE INC

Monocular video dynamic human body reconstruction method and system based on three-dimensional gaussian splashing

This invention relates to the fields of computer vision and computer graphics, and provides a method and system for dynamic human body reconstruction from monocular video based on 3D Gaussian splashing. The method includes the following steps: data preprocessing; initialization of the normalized space 3D Gaussian; deformation of the normalized space 3D Gaussian to an intermediate pose space to obtain a non-rigidly deformable 3D Gaussian and pose-related features; transformation of the non-rigidly deformable 3D Gaussian to the observation space using linear blending skinning to obtain the observation space 3D Gaussian; decoding the viewpoint-related color based on Gaussian color features, pose-related features, and viewpoint direction; constructing a total loss function including a normal consistency regularization term to optimize the 3D Gaussian attributes and network parameters; and rendering the target human body image using a differentiable Gaussian splash rasterizer. This invention enables rapid and fully automatic reconstruction from monocular video to a high-fidelity, animable human body model, applicable to fields such as virtual reality and film production.
Owner:CHANGCHUN UNIV

Electrocardiogram technology synthesized by auto-encoder

The invention discloses an electrocardiogram technology synthesized by an auto-encoder, relates to the technical field of crossing of medical image processing and computer vision, in particular to the electrocardiogram technology synthesized by the auto-encoder, and aims to solve the problem that an existing ECG image generation method is difficult to simulate real paper wrinkles and geometric deformation. According to the method, ECG images are input in a partitioned mode to serve as sub-blocks, pixel offset is extracted from style images for geometric deformation, a diffusion model is adopted for style transfer, content feature anchoring and a style mapping enhancement mechanism are combined, and the vivid wrinkle visual effect is injected while it is ensured that ECG waveform features are reserved. The image generated by the method has high structural similarity and visual authenticity, is suitable for medical data enhancement and diagnosis model training, and improves the generalization ability and clinical practicability of downstream tasks.
Owner:TIANJIN UNIVERSITY OF TECHNOLOGY

Spatial prior calibration and feature adaptation method and system for sparse perception architecture

PendingCN122637386AMorphingFeature adaptation
The application relates to the technical field of automatic driving perception, in particular to a space prior calibration and feature adaptation method and system for a sparse perception architecture, which freezes the interpolated position embedding matrix into a non-trainable state, so that the space layout sensitive representation (such as object relative position and scale relationship) learned in the pre-training stage is completely preserved, the space prior destruction caused by the re-learning of a downstream task is avoided, and the overfitting risk is effectively reduced. A learnable 1x1 one-dimensional convolution is used for calibration along the channel dimension, so that the trainable parameter quantity and the input resolution are completely decoupled. The image features output after the space prior protection and the channel response calibration have accurate geometric structures, so that the Deformable Attention (deformable attention) mechanism in a sparse perception method such as Sparse4D is accurate in positioning and rich in feature semantics when four-dimensional key point sampling is performed, and the overall performance of three-dimensional target detection is improved.
Owner:HONEYCOMB (WUHAN) MICROSYSTEM TECH CO LTD

A deformation field prediction method based on image quality evaluation and adaptive pre-training

The application relates to a deformation field prediction method based on image quality evaluation and adaptive pre-training. Wavelet packet decomposition is performed on a reference image and a deformation image of a collected component to obtain coefficients of each subband of each layer of the images, and Pearson correlation coefficients of each subband are calculated based on all the coefficients corresponding to the subband. The subbands are screened based on L2 norms of coefficient vectors of the subbands and the Pearson correlation coefficients, and the screened coefficient vectors of the subbands are reconstructed to obtain reconstructed reference images and reconstructed deformation images. Quality evaluation indexes of each reconstructed sample pair are calculated. Training branches to which the corresponding reconstructed sample pairs belong are determined based on the quality evaluation indexes, a U-Net model is preliminarily trained by using a preset supervised training strategy based on the reconstructed sample pairs of different training branches, the U-Net model is unsupervisedly trained, and displacement fields and strain fields are predicted based on the trained U-Net model.
Owner:HUNAN UNIV

Anti-parallax image non-rigid real-time interactive deformation method

The invention relates to the technical field of image processing, and discloses an anti-parallax image non-rigid real-time interactive deformation method, which comprises the following steps: S1, acquiring a first image and a second image; s2, extracting a semantic primitive corresponding to the initial control point pair; s3, receiving an operation instruction of a user for the semantic primitive; s4, updating the position of a control point associated with the operated semantic primitive; s5, calculating a non-rigid deformation field by adopting a moving least square deformation algorithm; and S6, applying the non-rigid deformation field to the second image to generate the second image. According to the method, a structure regular term is introduced, and non-similar components in local affine transformation are quantized and punished, so that a finally calculated non-rigid deformation field keeps the local shape and angle of image content unchanged to the greatest extent while meeting control point constraints, unnatural distortion is avoided, and the accuracy of a deformation result is ensured.
Owner:SHENZHEN ZHENCHENG TECH CO LTD

Multi-language display method and system and wearable intelligent device

The invention relates to the technical field of computer data processing, in particular to a multi-language display method and system and wearable intelligent equipment. The method comprises the following steps: segmenting the content of a display text based on language types of various font sub-files cached in advance; all the segments are reset based on a typesetting engine, font contour information of all the characters is output, and the font contour information comprises deformation and connected characters; typesetting information is determined on the basis of line width limitation of the lightweight general graphics library, and the typesetting information comprises line tail broken lines, word spacing and self-adaptive zooming; and based on the typesetting information and the font contour information of all the characters, selecting a dot matrix font drawing mode / vector font drawing mode for rendering. By means of the method, multi-language compatibility can be achieved, and meanwhile the problems of line breaking errors, character pattern dislocation or extremely low rendering speed and the like occurring when mixed text rendering is processed are solved.
Owner:ZHENSHI INFORMATION TECH SHANGHAI CO LTD

Lightweight deformable strip feature extraction method for remote sensing image building extraction

The invention discloses a lightweight deformable strip feature extraction method for remote sensing image building extraction, and belongs to the technical field of remote sensing. The method comprises the steps that a deformable strip attention backbone network is constructed, and ImageNet1K is used for pre-training for 300 rounds; the method comprises the following steps: cutting a building-containing data set into 1024 * 1024 image blocks and dividing a training / verification / test set; embedding the backbone network into a remote sensing task framework; loading a pre-training weight training network; and the test model outputs a result. The backbone network is divided into four stages, down-sampling is carried out in each stage, then deformable strip attention modules are stacked, the number of the deformable strip attention modules in the first second stage and the fourth stage is two, the number of the deformable strip attention modules in the third stage is four, the modules are subjected to 1 * 1 convolution dimensionality reduction and GELU activation, and branches are subjected to depth separable convolution and strip convolution K value 7 / 9 / 11 / 13 processing and then multiplied with original features to be output. The method gives consideration to precision and efficiency, and is suitable for high-resolution remote sensing image building extraction.
Owner:CHINA THREE GORGES UNIV

Human-centered video scene reconstruction and separation method, system, medium and device

This application provides a method, system, medium, and device for human-centered video scene reconstruction and separation. The method includes: for a first-person perspective video sequence, initializing a 3D Gaussian set covering the background, hands, and objects based on a priori knowledge of structure recovery from motion; assigning a learnable dynamic category probability vector to each Gaussian point in the 3D Gaussian set; constructing a dedicated deformation branch; according to the learnable dynamic category probability vector of each Gaussian point, assigning each Gaussian point to the dedicated deformation branch for processing through a preset soft-hard two-stage routing mechanism, determining the Gaussian points processed by the dedicated deformation branch; rendering the Gaussian points processed by the dedicated deformation branch to determine the 4D scene reconstruction image and the decomposed reconstruction images of the background, hands, and objects. This application achieves 4D scene reconstruction of human-centered video and explicit, fine-grained separation of the background, hands, and objects.
Owner:SHANGHAI JIAOTONG UNIV