Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

261 results about "Latent vector" patented technology

The latent vector is a a lower dimensional representation of the features of an input image. The space of all latent vectors is called the latent space. The latent vector denoted by the symbol $z$, represents an intermediate feature space in the generator network.

System and method for causality-augmented generative intelligence to discover non-obvious insights from heterogeneous data sources

The present invention provides a system and method for causality-augmented generative intelligence capable of autonomously discovering non-obvious actionable insights from heterogeneous and multimodal data sources. The system integrates a data ingestion unit for semantic and temporal harmonization of structured and unstructured datasets, a causal inference processor for constructing a dynamically evolving directed causal knowledge representation using perturbation-based validation, a latent representation processor that combines multimodal semantic embeddings with causal parameters to generate fused latent vectors, and a generative insight processor utilizing causally constrained generative reasoning to synthesize hypotheses anchored to verified cause-effect dependencies. A validation processor performs counterfactual assessment and observational verification to ensure retention of only those insights that remain consistent with causal ground truth.
Owner:MIA MD TOFAYEL GONEE MANIK

Spatial intelligent three-dimensional modeling method for designing sketch image based on two-dimensional structure

The invention provides a spatial intelligent three-dimensional modeling method based on a two-dimensional structure design sketch image, and the method comprises the steps: carrying out the preprocessing of an input two-dimensional structure design sketch image, and obtaining a preprocessed design sketch image, extracting two-dimensional structure design features from the preprocessed design sketch image, and encoding the two-dimensional structure design features to generate structured tensor representation; deducing space components, function partitions and constraint logic based on structured tensor representation to construct a sketch semantic graph, nodes of the sketch semantic graph representing physical structure units, and edges of the sketch semantic graph representing connection relations and stress constraints among the physical structure units; determining a node embedding vector corresponding to each node in the sketch semantic graph so as to perform three-dimensional space coding on the node embedding vector to obtain a node potential vector, and generating a geometric reasoning sequence during three-dimensional structure geometric reasoning according to a mechanical logic definition on the basis of edges in the sketch semantic graph, and performing three-dimensional structure geometric reasoning based on the node potential vectors and the geometric reasoning sequence to generate a three-dimensional structure model.
Owner:BEIJING FEIDU TECH CO LTD

Robot motion control model training method, device and equipment based on deep reinforcement learning, robot and medium

The invention provides a robot motion control model training method, device and equipment based on deep reinforcement learning, a robot and a medium, and relates to the technical field of robots. The method comprises the following steps: acquiring a first potential vector obtained after a student encoder encodes robot body observation data, and a second potential vector obtained after a teacher encoder encodes privilege observation data; based on the current training step number and a preset probability function, calculating a sampling probability for controlling a fusion proportion of the first potential vector and the second potential vector; fusing the first potential vector and the second potential vector based on the sampling probability to generate a third potential vector, and inputting the third potential vector into a strategy network; and updating the parameters of the policy network based on the value estimation of the current state output by the value network and the action policy output by the policy network. According to the method, updating oscillation caused by sudden change of input distribution in the training process of the strategy network can be avoided, the training efficiency is improved, and the training cost is reduced.
Owner:SHENZHEN ZHUJI POWER TECH CO LTD

Latent Transformer Architecture with Attention Mechanisms and Expert Systems for Federated Deep Learning with Homomorphic Encryption

ActiveUS20260039311A1Code conversionMachine learningMixture of expertsEngineering
A latent transformer architecture with latent attention mechanisms and expert processing systems for federated deep learning is disclosed. The system operates entirely within latent space, eliminating traditional embedding and positional encoding layers while maintaining full attention capabilities. Input data is compressed into latent vectors via variational autoencoder encoding, then processed by a latent attention module that computes query, key, and value matrices directly from latent representations. The architecture incorporates expert processing systems including gated latent expert networks for sparse computation and latent mixture of experts for collaborative processing. In the gated approach, a routing network selectively activates specialized expert modules based on latent vector characteristics. The mixture approach enables all experts to contribute through weighted combination, facilitating distributed computation and enhanced model expressiveness.
Owner:ATOMBEAM TECH INC

Method and system for motion generation from input text

A method for training a model for generating a representation of long-term motion from a text input comprises: training a motion encoder of an autoencoder to compress and map an input motion into a latent representation comprising a sequence of latent vectors in a discrete latent space, each latent vector representing a fixed length of motion; training a quantization module to quantize the latent vectors to a sequence of quantized latent vectors in quantized latent space; and training a motion decoder to reconstruct the quantized sequence as a sequence of single-frame pose representations. A text encoder is trained to predict a latent sequence conditioned on a text input and a duration using the mapped latent representation as a target.
Owner:NAVER CORP

Causal discovery using knowledge graph link prediction

Causal discovery is performed using knowledge graph link prediction. Information from a causal network is transformed into a causal knowledge graph according to a mapping, the causal knowledge graph including a plurality of causal links, wherein each causal link includes a cause entity, a causal relation, an effect entity, and a causal weight indicating a relative strength of causal influence of the cause entity on the effect entity. The causal knowledge graph is converted into embeddings, where the embeddings include a latent vector space representation of the causal knowledge graph. The embeddings are trained using a subset of the causal links of the causal knowledge graph. The embeddings are used for causal discovery to predict additional causal links of the causal knowledge graph.
Owner:ROBERT BOSCH GMBH

Method for predicting load displacement curve of CFRP reinforced concrete filled steel tubular column

The invention relates to the technical field of civil engineering structure calculation and artificial intelligence crossing, in particular to a load displacement curve prediction method for a CFRP reinforced concrete filled steel tubular column, which comprises the following steps: acquiring load displacement curves and corresponding input parameters of the CFRP reinforced concrete filled steel tubular column under different parameter combinations, and constructing a training database; constructing a load displacement curve prediction model by taking a conditional variation auto-encoder as a framework; for the real-time state parameters of the target CFRP reinforced concrete filled steel tubular column, N times of sampling is carried out from prior distribution of a potential space to obtain N potential vectors, and the N potential vectors and the corresponding real-time state parameters are input into a decoder for prediction to obtain a curve cluster; and averaging the curve clusters to obtain a final prediction curve. According to the method, the variational auto-encoder is combined with the deep neural operator, so that a continuous and smooth whole-process load displacement curve can be directly predicted and generated, and the prediction precision and integrity are greatly improved.
Owner:XIHUA UNIV

Latent transformer core for a large codeword model

A Large Codeword Model (LCM) with a latent transformer core is a deep learning architecture that operates on discrete, compressed representations of data called codewords. The latent transformer core incorporates a Variational Autoencoder (VAE) which allows for the removal of the embedding and positional encoding layers from the Transformer. Input data is compressed into a latent space representation using the VAE encoder, which is then processed by the Transformer. The VAE decoder generates outputs based on the processed latent vectors. This approach enables efficient handling of diverse data types beyond language, including time series, images, and audio.
Owner:ATOMBEAM TECH INC

Generated image detection method and system for face privacy protection

The invention discloses a face privacy protection-oriented generated image detection method and a face privacy protection-oriented generated image detection system. The method comprises the following steps of: firstly, preparing face and text pairing data, and finely adjusting a diffusion model; secondly, on the basis of the diffusion model after fine tuning, potential vectors are extracted and clustered, and text prompts and center vectors obtained through clustering form a dictionary; and finally, based on the obtained dictionary, obtaining a pseudo image and a label through the fine-tuned diffusion model, and outputting a detection result through a classifier. According to the method, potential spatial clustering and conditional diffusion generation are combined, privacy protection and data diversity are taken into consideration, and the security and generalization ability of forged face image detection are remarkably improved.
Owner:HANGZHOU DIANZI UNIV

Sensor data recovery method and system based on mask perception space-time modeling

The invention provides a sensor data recovery method and system based on mask perception space-time modeling, and belongs to the technical field of Internet of Things and intelligent sensing data processing. The method comprises the following steps: modeling a missing position into a trainable potential vector representation through an adaptive missing representation module, and avoiding noise introduced by traditional zero filling or mean filling; a missing perception space-time decoding module is adopted, and a dynamic weight distribution mechanism of mask constraint is combined in the decoding process, so that information is effectively prevented from being excessively smooth; designing a space-time dual-channel feature aggregation unit, and capturing spatial dependence and time sequence dependence between the sensors at the same time; and finally realizing accurate completion of a large-scale space-time sensor matrix through an output recovery unit. According to the method, the accuracy and robustness of sensor data restoration can be remarkably improved, the reliability of subsequent monitoring, prediction and anomaly detection is enhanced, and the method has wide engineering application value.
Owner:SOUTHEAST UNIV

Object detection device incorporating quantum computing and game theoretic optimization and related methods

An object detection device may include a variational autoencoder (VAE) configured to encode image data to generate a latent vector, and decode the latent vector to generate new image data. The object detection device may also include a quantum computing circuit configured to perform quantum subset summing, and a processor. The processor may be configured to generate a game theory reward matrix for a plurality of different deep learning models, cooperate with the quantum computing circuit to perform quantum subset summing of the game theory reward matrix, select a deep learning model from the plurality thereof based upon the quantum subset summing of the game theory reward matrix, and process the new image data using the selected deep learning model for object detection.
Owner:EAGLE TECHNOLOGY LLC

Digital human expression generation method based on multi-modal feature fusion and emotion enhancement

The invention discloses a digital human expression generation method based on multi-modal feature fusion and emotion enhancement, and belongs to the technical field of generative artificial intelligence. Comprising the steps of extracting voice semantic features and voice emotion features based on driving voice, extracting text emotion features based on an emotion text, and extracting visual latent variables and image semantic features based on a reference image; fusing the voice emotion features and the text emotion features through a perception resampling mechanism to generate fused emotion control features; fusing emotion control features, voice semantic features, visual latent variables and image semantic features as conditions, inputting the conditions into a DiT video generation model, generating a denoised video potential vector, mapping the denoised video potential vector into a facial expression potential vector by a Transform adapter, and restoring the facial expression potential vector into a facial expression parameter sequence by a FaceVese decoder to drive a digital human. According to the invention, end-to-end generation from voice and text to high-fidelity and emotion-controllable facial expression parameters can be realized.
Owner:ZHEJIANG UNIV

Image Generation Method Based on Brownian Bridge Diffusion Model, Device and Medium

PendingUS20260087592A1Image enhancementImage analysisSatellite imageBrownian bridge
The present application relates to an image generation method and apparatus based on a Brownian bridge diffusion model, a device and a medium, wherein the method includes: receiving an image combination including a satellite image and a ground panoramic image; extracting shared features of the satellite image and the ground panoramic image; performing a polar coordinate transformation on the satellite image to obtain an initial latent vector of the satellite image and a latent vector of the ground panoramic image; gradually adding noise into the latent vector of the ground panoramic image to obtain a latent vector of the satellite image; gradually removing the noise in the latent vector of the satellite image to generate a target latent vector; and decoding the target latent vector to generate a target ground panoramic image. The efficiency and quality of conversion from the satellite image to the ground panoramic image are improved.
Owner:SHENZHEN UNIV

CSI feedback method based on Transform and entropy constraint vector quantization

The embodiment of the invention provides a CSI (Channel State Information) feedback method based on Transform and entropy constraint vector quantization. The method is applied to the technical field of wireless communication. The method comprises the following steps: receiving an input angle-time delay domain channel matrix, and performing feature extraction on the channel matrix through a convolutional layer to obtain a first feature tensor; the first feature tensor is input into a trained Transform encoder for analysis processing, and a latent vector is obtained; performing quantization processing on the latent vector through the trained vector quantization variational auto-encoder to obtain discretization representation of the latent vector; performing dequantization processing on the discretized representation of the latent vector to obtain a dequantized latent vector; performing feature extraction on the dequantized latent vector through a convolutional layer to generate an initial channel feature; the initial channel characteristics are input into a trained Transform decoder for analysis processing, and final channel characteristics are obtained; and the final channel feature is mapped to the target resolution to obtain a reconstructed channel matrix, so that the reconstruction precision of the channel matrix is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Digital human action stylized generation method based on text driving

The invention discloses a digital human action stylized generation method based on text driving, and belongs to the technical field of digital human action generation. The method comprises the following steps: extracting a 3D human body action parameter sequence from a target person video; performing automatic verification and physical optimization on the sequence to obtain a pure data set; extracting quantized style features and encoding the quantized style features into low-dimensional style latent vectors; encoding a natural language instruction into a text semantic vector, and fusing the text semantic vector with the style latent vector to generate a control signal; a diffusion model is driven by the control signal to generate an action parameter sequence conforming to instruction semantics and personalized styles; and finally mapping to a digital human model and rendering and outputting an animation video. According to the method, deep fusion of the text instruction and the personalized exercise style is realized, the problems of single style, unreasonable physics and the like of the generated action are solved through a systematic physical compliance guarantee system, and the sense of reality and expressive force of the digital human action are remarkably improved.
Owner:SHANGHAI IRIDIUM WEISI INTELLIGENT TECH CO LTD

Synthesizing sequences of 3D geometries for movement-based performance

A technique for generating a sequence of geometries includes converting, via an encoder neural network, one or more input geometries corresponding to one or more frames within an animation into one or more latent vectors. The technique also includes generating the sequence of geometries corresponding to a sequence of frames within the animation based on the one or more latent vectors. The technique further includes causing output related to the animation to be generated based on the sequence of geometries.
Owner:ETH ZURICH +1

Hidden space confrontation sample generation method and system based on multi-scale feature separation

The invention discloses a hidden space adversarial sample generation method and system based on multi-scale feature separation, and the method comprises the steps: employing a neural network quantization training method based on straight-through estimation, and training a hierarchical vector quantization variational auto-encoder; carrying out differentiable Haar wavelet transformation on the input image by adopting a wavelet packet transformation algorithm, decomposing the input image into a low-frequency component and a high-frequency component, and realizing multi-scale feature separation; inputting the high-frequency component into a hierarchical vector quantization variational auto-encoder, and extracting and quantizing global high-frequency features and local high-frequency detail features; in the potential space, a learnable disturbance variable is introduced, a potential vector after disturbance is constructed, and the potential vector is reconstructed into an adversarial sample through a decoder; and based on a preset disturbance target, carrying out iterative optimization on the disturbance vector until a confrontation sample which satisfies an attack success condition and is optimized in visual quality is generated. According to the method, a wavelet domain variational auto-encoder and a hidden space iterative attack algorithm are fused, and an adversarial sample with high fidelity and clear interpretation is generated.
Owner:XINJIANG UNIVERSITY

Generative reverse face recognition method based on text guidance

The invention provides a text guidance-based generative reverse face recognition method. The method comprises the following steps of: obtaining an initial latent vector and a fine tuning generator by using a pre-trained editing encoder and an original face image; taking the initial latent vector as an initial vector during gradient updating, and starting to circularly update until a final latent space code is obtained; and inputting the subsurface space code into the fine-tuned generator, and taking the generated image as a protected image corresponding to the original face image. The total loss function comprises an adversarial loss for explicitly promoting diversity and a loss function of a visual effect; the fine-tuned image generated by the generator is closer to an auxiliary image randomly selected from the auxiliary data set so as to realize protection of dynamic tracking of face recognition; meanwhile, it is ensured that the image generated by the fine-adjusted generator follows the specification of target text prompt and keeps the visual similarity with the original face image. The method has excellent performance in the aspect of preventing the face image of the user from being identified by the dynamic FR strategy.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Adapting simulated character interactions to different morphologies and interaction scenarios

Some implementations relate to methods, systems, and computer-readable media for adapting simulated character interactions to different morphologies and interaction scenarios. The system accesses a graph representing a control policy for a simulated character's movements in a virtual environment. This graph undergoes encoding and processing through a graph neural network to generate latent embeddings for the graph. A fixed-length latent vector is determined from the latent embeddings. This vector is input to a feedforward neural network, generating control signals for the character's actions. Through a reinforcement learning loop, the character's motions are continuously refined by iteratively adjusting the graph based on evaluating the actions of the simulated character via a reward function, adapting the control policy to different character morphologies and / or interaction scenarios.
Owner:ROBLOX CORP

Learning from imperfect data for anomaly detection

Apparatus and method of training Machine Learning (ML) models. In an embodiment, the apparatus performs initial training of an anomaly detection model based on training samples of a training dataset over multiple epochs, where the anomaly detection model comprises a variational autoencoder (VAE). For each training sample during an epoch, the initial training comprises inputting an original data sequence of the training sample into the VAE encoder to output a multivariant distribution in latent space, sampling the multivariant distribution to generate multiple latent vectors, inputting the latent vectors into the VAE decoder to output reconstructed data sequences, and computing an estimated sample weight for the training sample. The apparatus identifies, after multiple epochs, corrupted samples from the training dataset based on the estimated sample weights, removes the corrupted samples to generate a filtered training dataset, and performs final training of the anomaly detection model based on the filtered training dataset.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Method and system for generating image watermark based on potential variable optimization of diffusion model

PendingCN121458514AImage enhancementImage analysisWatermark robustnessWatermark method
The invention discloses an image watermark generation method and system based on a diffusion model optimization potential variable, and the method comprises the steps: training a secret encoder and a watermark extractor, and coding a message into a secret residual error; after the residual error is injected into the potential representation of the image, mapping the residual error into a potential vector with a watermark, which accords with the prior of a diffusion model, through DDIM inversion; performing iterative optimization on the potential vector by using a pre-trained diffusion model U-Net, introducing double constraints of image reconstruction loss and message fidelity loss in a denoising process, and generating an image with high visual quality and high watermark robustness; and finally, extracting the message through a watermark extractor and carrying out statistical significance verification to realize copyright authentication. According to the method, dynamic watermark embedding is achieved, the attack resistance is high, the image quality is well kept, an original model structure is not changed, and the method is suitable for copyright protection and user tracking of AIGC content.
Owner:ZHEJIANG UNIV OF TECH

Farmland layout method and system

The invention provides a farmland layout method and system. The farmland layout method comprises the steps of obtaining user intention interactive input and agricultural multi-source heterogeneous data about a to-be-planned area; performing data preprocessing on the agricultural multi-source heterogeneous data to form an agricultural feature tensor; performing multi-modal feature fusion on the user intention interaction input and the agricultural feature tensor to form an agricultural feature latent vector; generating a first farmland layout scheme by a pre-trained generative reasoning model under the condition of the agricultural feature latent vector; the first farmland layout scheme is optimized through reinforcement learning, and a second farmland layout scheme optimized through reinforcement learning is output; and rendering the second farmland layout scheme into an interactive three-dimensional farmland layout scene through a diffusion model. The generated interactive three-dimensional farmland layout scene restores real landforms and crop forms, reduces on-site secondary surveying and mapping and manual trial and error, and ensures that the layout can be directly implemented on the ground.
Owner:SHANGHAI ALL THINGS SHENGLONG TECHNOLOGY DEVELOPMENT CO LTD

A power image model knowledge migration method and system based on predicted divergence confrontation

The application relates to a power image model knowledge migration method and system based on predicted divergence confrontation, which comprises the following steps: performing feature description on an original power image through a visual model, performing hidden space sampling on the feature description of the original power image through a feature extractor, and generating an initial latent vector; inputting the initial latent vector into a generator to obtain a reconstructed power image, and generating an adversarial power image through gradient symbol method-based adversarial disturbance; inputting the adversarial power image into a target model and a substitute model respectively, and calculating a predicted divergence loss; training the substitute model to learn the knowledge of the target model through the predicted divergence loss; and when the performance of the substitute model on a verification set is stable and close to that of the target model, completing knowledge migration. Compared with the prior art, the application significantly reduces data labeling cost, improves the adaptability of knowledge migration in a resource-limited scene, and improves the learning efficiency of a substitute model for key features of a target model and the decision boundary exploration ability.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1

Deep learning framework for video remastering

Restoration methods and systems are disclosed for video remastering. Techniques disclosed include receiving a video sequence. For each frame of the video sequence, techniques disclosed include encoding, by a degradation encoder, a video content associated with the frame into a latent vector. The latent vector is a representation of the degradation present in the video content; the degradation present in the video content includes one or more degradation types. Based on the latent vector and the video content, techniques disclosed further include generating, by a backbone network, one or more feature maps, and, then, restoring the frame based on the one or more feature maps.
Owner:DISNEY ENTERPRISES INC

Power data anomaly detection method and model based on generative adversarial network

The power data anomaly detection method based on the generative adversarial network comprises the following steps: original data signals are divided into small sequence signals, and corresponding latent vectors are obtained through mapping; latent vectors are mapped one by one to obtain a group of pseudo-time sequence data; the true probability of each sub-sequence of the pseudo-time sequence data is calculated, and the true probability of each small sequence signal is calculated; the true probability of the sub-sequence of the pseudo-time sequence data is compared with the true probability of each small sequence signal, the discrimination loss of each group of sub-sequences is calculated, and the total discrimination loss is obtained by summation; the sub-sequences of each pseudo-time sequence data are compared with each small sequence signal to obtain the residual loss of each group of sub-sequences, and the total residual loss is obtained by summation; according to the residual loss and the discrimination loss of each group of sub-sequences, the residual score and the discrimination score are calculated, the residual score and the discrimination score are weighted, and an anomaly score is obtained; the anomaly score is compared with a preset threshold to obtain a discrimination result.
Owner:SOUTH CHINA NORMAL UNIV

A multi-party model training method, system and apparatus

The embodiment of the specification provides a multi-party participated model training method, system and device, the method comprises: a first party inputs a plurality of first features into a first feature extraction network, homomorphically encrypts an output result to obtain a plurality of first encrypted vectors and sends the plurality of first encrypted vectors to a second party; the second party inputs a plurality of second features into a second feature extraction network, homomorphically encrypts an output result to obtain a plurality of second encrypted vectors; the second party homomorphically calculates the plurality of first encrypted vectors and the plurality of second encrypted vectors to obtain a plurality of fusion encrypted vectors and sends the plurality of fusion encrypted vectors to the first party after being disordered, and records a first correspondence before and after disordering; the first party decrypts the plurality of fusion encrypted vectors and inputs the plurality of fusion encrypted vectors into a first part of a classification network to obtain a plurality of latent vectors; the first party uses the plurality of latent vectors and a classification label, and the second party uses network parameters of a second part of the classification network and the first correspondence to perform multi-party secure calculation and determine a first loss; the first party and the second party update the classification network according to the first loss.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

An experimental process intelligent error correction method and device

The present application relates to the technical field of machine vision, in particular to an experimental process intelligent error correction method and device. The method comprises: dynamically evaluating the importance of each channel of multi-channel time series data and performing dimension reordering; inputting the reordered data into an encoder to obtain first and second latent vectors of adjacent time steps, and constructing a spliced latent vector that fuses time series and causal information; inputting the spliced vector into a forward and backward decoder to calculate a reconstruction inconsistency score; when the score exceeds a threshold value, locking the first latent vector, optimizing the gradient of the second latent vector to minimize inconsistency, obtaining a correction vector and generating error correction data. The scheme of the present application improves the recognition accuracy of complex anomalies, enhances the ability to distinguish between normal changes and errors, and improves the reliability of error correction results.
Owner:JIANGNAN UNIV +1

Training method for 3D model completion network, 3D model completion method and device

ActiveCN117408910BAlgorithmSimulation
This disclosure provides a training method, a 3D model completion method, and an apparatus for a 3D model completion network, including: pre-constructing a 3D model completion network to be trained; the 3D model completion network to be trained includes a 3D variational autoencoder and a diffusion model; acquiring a 3D model to be trained for network training; inputting the 3D model to be trained into the 3D variational autoencoder, inputting the encoded latent vectors into the diffusion model, performing noise addition and denoising processing on the diffusion model, and then inputting the latent vectors into the decoder to obtain a predicted generated 3D model; based on the predicted generated 3D model, calculating the loss of the 3D variational autoencoder and the diffusion model, and training the 3D model completion network to be trained. The trained 3D model completion network is then used to complete the 3D model to be completed. This ensures the quality of the 3D model while improving production efficiency.
Owner:北京渲光科技有限公司