Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17974 results about "Encoder" patented technology

An encoder is a device, circuit, transducer, software program, algorithm or person that converts information from one format or code to another, for the purpose of standardization, speed or compression.

Integration of self-organizing maps with autoencoder-GAN frameworks for enhanced routing in capsule networks

A method is provided for enhanced data routing in neural networks using Self-Organizing Maps (SOM) integrated with Autoencoder-GAN. The method comprises training an autoencoder to encode input data into a latent space representation; applying a Self-Organizing Map (SOM) to organize the latent space representation into a topological map; refining the latent space representation using a Generative Adversarial Network (GAN), wherein the generator generates enhanced latent space representations and the discriminator evaluates their quality; using the refined latent space representations to update the SOM topology dynamically; generating routing coefficients based on the updated SOM topology to guide data routing in a capsule network; and dynamically adjusting routing within the capsule network using the generated routing coefficients to enhance performance based on the refined latent representations.
Owner:LEPTUDE INC

Mapping latent space of vector quantized variational autoencoders to functional basis vectors for enhanced data representation and manipulation

A method is provided for mapping the latent space of a Vector Quantized Variational AutoEncoder (VQ-VAE) to polynomial basis vectors. The method includes training a VQ-VAE model on a dataset to obtain a set of codebook vectors representing the latent space; defining a polynomial basis for the latent space, the polynomial basis containing terms up to a predetermined order; mapping each codebook vector to the polynomial basis by determining polynomial coefficients that represent each codebook vector in terms of the polynomial basis; and using the polynomial coefficients to reconstruct and manipulate latent space representations.
Owner:LEPTUDE INC

Systems and methods for enhancing autoencoder performance and interpretability through language-guided feature selection and encoding

A method for structuring the latent space of an autoencoder is provided. The method includes analyzing natural language descriptions related to input data; creating language-guided libraries that categorize and abstract data features based on the analyzed descriptions; mapping input data into the categorized and abstracted features within the latent space of the autoencoder; and training the autoencoder to minimize reconstruction loss while adhering to the structure imposed by the language-guided libraries.
Owner:LEPTUDE INC

Systems and Methods for Temporal Acceleration Encoding in Geodesic Latent Space for Event Forecasting

A system and method for temporal acceleration encoding in Lorentzian latent space enables real-time event forecasting within navigable spatiotemporal media. The system encodes media data into compact Lorentzian latent patches using variational autoencoders and organizes them within a multi-dimensional hyperspace spanning spatial, temporal, orientation, scale, and spectral coordinates. Temporal acceleration encoding computes velocity and acceleration vectors along geodesic trajectories, extracting event signatures through multi-scale aggregation over sliding windows. An acceleration-indexed memory stores dynamic descriptors with composite keys comprising hyperspace coordinates and motion characteristics. Event forecasting retrieves similar historical patterns and conditions a forecast head to produce event probabilities and time-to-event estimates with uncertainty calibration. The system streams forecast metadata to edge devices for real-time prediction and adaptive navigation, supporting applications in surveillance, autonomous systems, predictive media exploration, and anomaly detection where both temporal forecasting and multidimensional navigation capabilities are essential.
Owner:ATOMBEAM TECH INC

Multi-modal semantic and physical law driven remote sensing image generation method

The invention discloses a multi-modal semantic and physical law driven remote sensing image generation method, belongs to the technical field of computer vision and remote sensing image generation, and aims to solve the problems of insufficient cross-modal semantic alignment, low reliability of a generation result and insufficient physical mechanism fusion. The four-stage method comprises the following steps: firstly, rejecting low-quality samples from original data and unifying a spatial scale; then, extracting a multi-modal semantic vector by adopting a BLIP model and a CLIP model, and introducing a remote sensing physical rule to carry out vector optimization; then position coding and physical constraint conditions are embedded in the submerged space, and multi-source information joint modeling is achieved through a cross-modal encoder; and finally, by taking text description, physical priori knowledge and diffusion time steps as joint conditions, performing de-noising reasoning based on a Transform architecture, and completing back diffusion reconstruction by means of a trans-attention mechanism. According to the method, physical rationality and semantic consistency are improved, and a more reliable technical normal form is provided for remote sensing image generation in the fields of disaster monitoring, military simulation and the like.
Owner:CHINA UNIV OF MINING & TECH +2

Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Data center machine room AI energy-saving control method and system

The invention discloses a data center machine room AI energy-saving control method and system, a digital twin model of a machine room operation state is constructed through a holographic perception and heterogeneous data fusion technology, centimeter-level monitoring of an equipment state and environmental parameters is realized, and the system integrates a laser radar array, an acoustic sensor and a gas sensor network. The time-space alignment of multi-modal data is completed by combining edge computing nodes, holographic mapping including thermodynamic characteristics, vibration characteristics and gas leakage risks is formed, historical temperature control strategy characteristics are extracted by adopting a variational auto-encoder based on a dynamic strategy generation mechanism of generative artificial intelligence, and a load trend is predicted by combining a long-short-term memory network. Constructing a self-adaptive strategy pool; the multi-agent reinforcement learning framework enables temperature control, equipment scheduling and power grid response to form game optimization, the strategy robustness in a complex scene is improved, and the system innovatively fuses power grid real-time electricity price and carbon transaction data so as to establish a multi-target decision system.
Owner:SHENZHEN JITON INTELLIGENT TECH CO LTD

Dynamic Latent Space Adaptation Based on Spatiotemporal Kernal Context for Multiscale Rendering

A system for dynamic latent space adaptation using spatiotemporal kernel context for multiscale rendering with hierarchical and Lorentzian autoencoders. The Spatiotemporal Kernel Estimator (SKE) analyzes media through motion field, temporal recurrence, frequency band, and scene semantics analyzers to generate adaptive kernel parameters encoding content-specific importance distributions. The system dynamically adapts latent manifold geometry by modifying metric tensor properties according to kernel context, enabling content-aware compression that allocates representational capacity based on visual significance. A multiscale cache implements kernel-adaptive retention policies prioritizing important regions. An adaptive renderer provides intelligent level-of-detail selection based on zoom level and kernel-estimated importance, optimizing processing allocation. The self-optimizing architecture continuously refines kernel context and geometric adaptation based on user interaction and performance feedback, achieving superior compression ratios and perceptual quality. Applications include bandwidth-efficient video streaming, virtual reality, scientific visualization, and cognitive video analytics requiring intelligent context-aware visual processing.
Owner:ATOMBEAM TECH INC

Temporal dynamics simulation in matmul-free neural architectures

A neural network system is provided. The system includes an autoencoder configured to encode input data into a latent space representation; a generator neural network configured to receive a noise vector and the latent space representation and output a set of routing coefficients; a discriminator neural network configured to evaluate the effectiveness of the routing coefficients by measuring the performance of a capsule network utilizing said routing coefficients; and a capsule network comprising a first capsule layer and a second capsule layer, wherein the routing coefficients are used to dynamically route outputs from the first capsule layer to the second capsule layer.
Owner:LEPTUDE INC

Multi-modal dynamic fusion and incremental learning fault diagnosis method for deep vertical shaft equipment

The invention discloses a multi-modal dynamic fusion and incremental learning fault diagnosis method for deep vertical shaft equipment, which belongs to the technical field of industrial equipment fault diagnosis, and comprises the following four steps of: constructing a pre-training large model to perform feature extraction, and relying on a multi-layer Transformer encoder and a dual loss function, establishing a multi-modal dynamic fusion and incremental learning fault diagnosis model; mining cross-modal universal fault features from vibration, temperature and current multi-modal time sequence data; according to the method, multi-modal features are fused, multi-modal association is constructed, modal weights are dynamically adjusted through a modal gating unit and a time delay compensation attention mechanism to adapt to signal quality changes, and meanwhile time sequence deviation is corrected to achieve accurate association; incremental learning is realized by using a decoupling projection layer, and a lightweight projection module is designed for a newly added fault task to suppress disastrous forgetting; network training is optimized, pre-training loss, incremental learning loss and attention regularization loss are integrated through a multi-objective loss function, and model stability and diagnosis precision are improved. The method has the advantage that the model stability and the diagnosis precision are improved.
Owner:CHINA COAL NO 5 CONSTR +1

Road crack detection method and system based on improved RT-DETR, computer equipment and storage medium

The road crack detection method based on the improved RT-DETR comprises the following steps: shooting a road at a preset flight height by using an unmanned aerial vehicle to obtain an original road image containing a crack; a pre-trained crack detection model is utilized to carry out crack detection based on an original road image to obtain crack parameters, and the crack detection model is obtained through improvement and training based on an RT-DETR (Real-Time Detecting Transformer) model; the method for improving the RT-DETR model to obtain the crack detection model comprises the following steps: replacing a basic residual block at the tail end of a ResNet18 backbone network in the RT-DETR model with a dynamic snakelike convolution residual block (DSCRBlock); a cross-scale feature fusion module (CCFM) in a hybrid encoder in an RT-DETR model is replaced by a bidirectional diffusion focusing pyramid network (BDFPN), and the bidirectional diffusion focusing pyramid network comprises a primary focusing sub-network and a secondary focusing sub-network. The method can efficiently and accurately identify the road crack, can be applied to the unmanned aerial vehicle, and is easier to implement.
Owner:HENAN UNIVERSITY OF TECHNOLOGY

Windmill bridge coupling response analysis method

The invention relates to the field of bridge structure dynamic response analysis, and discloses a windmill bridge coupling response analysis method. According to the method, wind speed, wind direction and vehicle speed data are collected, and a data set is constructed by combining finite element and CFD coupling numerical simulation; a parallel encoder is adopted to fuse Transform feature extraction and LSTM time sequence processing to generate a hybrid prediction response; constructing a physical constraint and composite loss function based on a train-bridge motion equation, and optimizing neural network parameters through a subtraction average strategy; and finally, predicting dynamic response through forward propagation and verifying physical consistency to form a model optimization closed loop. According to the method, a deep learning method and physical equation constraints are fused, the analysis precision and calculation efficiency of windmill bridge coupling response are remarkably improved, and a more reliable dynamic evaluation means is provided for bridge wind resistance design.
Owner:CENT SOUTH UNIV +1

Welded pipe conveying abnormity prediction method and system based on large model reasoning

The invention discloses a welded pipe conveying abnormity prediction method and system based on large model reasoning, and aims to solve the problems that multi-source data is difficult to align, cross-station false correlation is caused, prediction lacks executable positioning and time sequence, and linkage control reliability is insufficient. Event alignment is carried out by taking a controller edge signal and an encoder zero position as time anchor points, a production line topology semantic graph containing time delay, capacity and interlocking attributes is constructed, and topology reachability and physical time delay constraints are applied in a self-attention long sequence model to carry out multi-step rolling prediction. And outputting a risk probability, refining the risk probability to spatial positioning of a roller way section or a shaft and the minimum executable intervention time, and generating a risk interval in combination with uncertainty estimation and calibration so as to drive an upstream beat self-adaptive speed reduction, shunting or stopping strategy. The technical effects of improving accuracy and interpretability, reducing false alarm and missing alarm, ensuring that linkage can be executed in advance and meeting edge time delay budget are achieved.
Owner:JIANGSU YINJIANG PRECISION TECH CO LTD

Multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion

The invention discloses a multi-modal remote sensing semantic segmentation method and system for learning frequency domain fusion. The method comprises the following steps: respectively extracting multi-scale features of two modal input images by adopting a double-branch encoder; sequentially executing frequency domain decoupling and fusion, mutual information constraint-based feature optimization and low-frequency guided cross-modal fusion processing on each scale feature to generate a fused semantic feature; and performing up-sampling and feature refining on the fused features through a decoder, and outputting a full-resolution segmentation prediction map. According to the multi-modal remote sensing image semantic segmentation method, modal sharing information and specific details are effectively separated through frequency domain decoupling, feature representation is optimized through mutual information constraint, adaptive feature fusion is achieved in combination with an attention mechanism, and the accuracy and robustness of multi-modal remote sensing image semantic segmentation are remarkably improved.
Owner:NORTHEAST FORESTRY UNIV

End-side multi-mode large model accelerated reasoning method and system

The invention provides an end-side multi-modal large model accelerated reasoning method and system, and the method comprises the steps: carrying out the two-stage screening and rearrangement of visual tokens based on the CLS attention and text-to-visual attention in a visual encoder and pre-filling stage, and constructing a sparse attention and sparse key value cache; in a decoding stage, an important neuron set is judged according to activation gating or historical statistics, only a corresponding feedforward network weight is pulled and calculated, missed weights are loaded on demand through asynchronous I / O, and hot neurons are maintained in a high-speed memory to utilize model sparsity, so that video memory / memory occupancy and calculation overhead are remarkably reduced on an end side; throughput and time delay performance are improved. According to the method, the internal memory and computing resources required by reasoning of the multi-modal large language model are reduced from two dimensions by utilizing the endogenous sparsity of the end-side large language model in input and the model, so that a higher reasoning speed is achieved by utilizing fewer resources on the premise of keeping the size of the model unchanged, and the performance of the whole system is improved.
Owner:SHANGHAI JIAOTONG UNIV

Digestive tract pathological diagnosis visual language large model construction method based on reinforcement learning and application thereof

The invention discloses an alimentary canal pathological diagnosis visual language large model construction method based on reinforcement learning and application thereof, and belongs to the technical field of medical image processing. The method comprises the following steps: firstly, extracting pathological information through layout analysis and adaptive threshold processing, and recombining the pathological information into a structured data set containing an inference chain; secondly, a visual encoder and a multi-branch classifier are used for extracting features and confidence coefficients, and dynamic structured cue words are generated; and finally, inputting the image and the cue word into a multi-modal large model, carrying out supervised fine-tuning hot start, and carrying out reinforcement learning training by adopting a group relative strategy optimization algorithm in cooperation with a composite reward function containing format, semantics and diagnosis dimensions. The problems that a general model is prone to generating illusion in a pathological scene and lacks reasoning logic are solved, and the accuracy and logicality of pathological report generation are remarkably improved.
Owner:SHENZHEN SHENGQIANG TECH

Information retrieval method and device based on multi-modal knowledge graph

The invention discloses an information retrieval method and device based on a multi-modal knowledge graph, and relates to the technical field of information retrieval, and the method comprises the following steps: obtaining multi-modal entity data; performing feature extraction on the multi-modal entity data to obtain a multi-modal feature vector; performing semantic unification on the multi-modal feature vectors to obtain multi-modal vectors with unified semantics; constructing a knowledge graph triple according to the multi-modal vectors with unified semantics; constructing a multi-modal knowledge graph according to the knowledge graph triple and the corresponding modal source information; inputting the natural language query of the user and the multi-modal knowledge graph into a preset graph enhanced generative retrieval large model, and searching a multi-modal entity related to the natural language query and a relation chain thereof in the multi-modal knowledge graph, and extracting multi-modal contents associated with the multi-modal entity and the relation chain, processing the multi-modal contents through respective encoders, injecting the processed multi-modal contents into an attention layer of the decoder, and outputting answers. According to the method, the high-precision and high-consistency intelligent question-answering capability oriented to complex tasks can be realized.
Owner:四川省文物交流和信息中心 +2

Complex scene-oriented end-to-end multi-modal content unified perception method and system

The invention belongs to the technical field of multi-modal data processing, and discloses a complex scene-oriented end-to-end multi-modal content unified perception method and system. The method comprises the following steps: performing intelligent sensing, identification, acquisition, screening and standardization processing on multi-modal data content, and outputting structured and standardized multi-modal content data; inputting a feature extraction model in parallel, and performing multi-modal content data feature unified modeling and preliminary fusion by adopting a multi-modal unified encoder which is internally integrated with a cross-modal attention layer and is based on a Transform architecture; cascade fusion high-order semantic representation is extracted step by step through a multi-stage and multi-level cascade cross-modal fusion structure; and performing deep semantic analysis on the extracted cascaded fusion high-order semantic representation by adopting a pre-trained semantic understanding model. According to the method, the fusion depth and perception precision of the multi-modal information in a complex scene are improved, and efficient and accurate understanding and interactive response of the multi-modal content are facilitated.
Owner:SHENZHEN WANGLIAN ANRUI NETWORK TECH CO LTD

Ocean wind field prediction method based on neural network

The invention provides an ocean wind field prediction method based on a neural network, and belongs to the technical field of ocean wind field prediction.The method comprises the steps that sparse ocean observation data are collected, a spatial covariance matrix is established, the spatial covariance matrix is converted into a graph structure, and then multi-hop neighborhood feature aggregation is conducted through a graph convolutional network; a tensor decomposition algorithm is combined for modeling high-order feature interaction to generate a gridding wind field, a bidirectional long-short-term memory network encoder is used for extracting space-time invariant features, a multi-layer perceptron predictor is used for directly mapping a future multi-step wind field, and a course learning strategy and a Shenchang differential equation boundary layer are matched for correction. The technical problem that sparse ocean observation data are difficult to accurately reconstruct into a high-resolution gridding wind field is solved.
Owner:自然资源部天津海洋中心(自然资源部天津海洋预报台)

Method and system for RAG type data item extraction and consistency checking oriented to engineering technology document

The invention relates to the technical field of intelligent processing of document information, discloses a method and a system for extracting RAG type data items and checking consistency for engineering technology documents, and solves the problem that an existing scheme is difficult to consider high-accuracy extraction, cross-document consistency, strong traceability and structured landing in engineering technology document data processing. The scheme comprises the following steps of: performing layout analysis and necessary OCR (Optical Character Recognition) processing on a received document, and dividing the document into multiple levels of text blocks; carrying out vectorization processing on the multi-level text blocks by adopting an embedding model subjected to two-stage fine adjustment, establishing a vector index and recalling candidate fragments to form a candidate fragment set; a cross encoder with efficient fine tuning of parameters is adopted to screen out high-correlation fragments; performing standardization processing by extracting a target variable to form a structured record; consistency checking is conducted on the target variables, and finally structured records and checking results are bound with minimum evidence fragments and corresponding position information to be exported through a format list.
Owner:CHINA HYDROELECTRIC ENGINEERING CONSULTING GROUP CHENGDU RESEARCH HYDROELECTRIC INVESTIGATION DESIGN AND INSTITUTE

Bipedal action model for humanoid robot

The present disclosure provides a system for generating motor control commands for a humanoid robot, comprising an alpha model with over 1 billion parameters that processes visual observations and language instructions at a first frequency to generate contextual embeddings, and a beta model operating at a higher second frequency. The beta model includes an embodiment-specific state encoder projecting robot state information into a shared embedding space, a diffusion transformer module generating denoised action sequences through iterative flow-matching that cross-attends to the alpha model's contextual embeddings, and an embodiment-specific action decoder converting denoised sequences into motor control commands. The beta model generates action chunks comprising future action sequences over a predetermined time horizon in a single inference step, with the complete system having less than 5 billion parameters.
Owner:FIGURE AI INC

Marine multi-mode environment perception and intelligent ship navigation decision-making method based on double-branch vision-semantic encoder

The invention discloses an ocean multi-mode environment perception and intelligent ship navigation decision-making method based on a double-branch vision-semantic encoder. The method comprises the following steps: S1, acquiring a multi-source data image containing a ship and a surrounding environment thereof from an existing public maritime data set or platform; s2, training a double-branch vision-semantic encoder by using the multi-source data image, and inputting a to-be-processed image extracted in real time into a multi-modal feature matrix in the trained double-branch vision-semantic encoder; s3, based on the multi-modal feature matrix, obtaining positioning information of the ship and surrounding environment elements, and constructing a dynamic security domain model; and S4, in combination with the dynamic security domain model and the multi-ship relative position relationship, carrying out quantitative evaluation on the navigation risk, and generating a self-adaptive navigation strategy based on an evaluation result. According to the invention, high-precision ship positioning and environment element identification under complex weather and illumination conditions are realized by using all-weather characteristics and multi-scale visual feature coding of SAR imaging.
Owner:HARBIN ENG UNIV

Abnormal data prediction and state evaluation method for battery

The invention discloses a battery abnormal data prediction and state evaluation method, and relates to the technical field of battery state prediction, and the method mainly comprises the steps: carrying out the preprocessing of an experiment data set, and obtaining multi-dimensional time series data; a combined feature encoder, a pre-response encoder and a memory analysis module are constructed to realize a battery abnormal data fault prediction model; training the model by using the multi-dimensional time sequence data to obtain a trained model, and predicting the to-be-predicted data to obtain a prediction result; and calculating a reconstruction error between a prediction result and original data, constructing an AUROC evaluation model, and evaluating the battery abnormal data fault prediction model. By implementing the battery abnormal data prediction and state evaluation method provided by the invention, the feature extraction efficiency, the abnormal recognition precision, the detection stability and the generalization ability can be improved.
Owner:WUHAN UNIV OF SCI & TECH

Intelligent metering method for remnant base paper rolls of corrugated paper board assembly line

The invention relates to the technical field of packaging, in particular to an intelligent metering method for remnant base paper rolls of a corrugated paper board assembly line. The intelligent metering method for the remnant base paper rolls of the corrugated paper board assembly line comprises steps as follows: a plurality of base paper feeding positions are arranged in the conveying direction of the assembly line at intervals, a base paper feeding frame is arranged in each base paper feeding position, and two base paper feeding rolls located on two sides of each base paper feeding frame are arranged on the base paper feeding frame; each base paper feeding roll is provided with lap recording sensors and incremental encoders, the lap recording sensors are used for monitoring the number of paper feeding laps of base paper clamped on the corresponding base paper feeding roll, the incremental encoders are used for monitoring the length of each lap after the base paper on the base paper feeding roll passes by one lap, and the lap recording sensors and the incremental encoders are connected with a controller. Intelligent metering can be realized, the operation cost of an enterprise is reduced, labor is saved, and the utilization rate of enterprise resources is increased.
Owner:DACHENG PACKAGING PROD SUZHOU

Generative adversarial network-based MRI-PET mode conversion method and system

The invention discloses an MRI-PET mode conversion method and system based on a generative adversarial network, and belongs to the technical field of artificial intelligence medical image generation. And the multi-scale structure representation injection module injects multi-scale anatomical prior information at different stages of the encoder, and overcomes the limitations of insufficient utilization of prior information and single injection scale. And the adaptive semantic residual fusion module adopts semantic attention guidance and double-branch attention weighting, adaptively fuses fine-grained local features and global context information, harmonizes the difference between the fine-grained local features and the global context information in an abstract level and a semantic category, and solves the problems of feature conflict and semantic fuzziness in a bottleneck region. The direction sensing space-frequency discriminator realizes multi-dimensional and fine-grained adversarial supervision through a space, frequency and local image block multi-branch collaborative discrimination mechanism, and improves the structural fidelity and spectrum authenticity of a synthetic image. And the generated image is superior to the existing method in indexes such as structural similarity and peak signal-to-noise ratio, and has higher clinical practical value.
Owner:NORTHEASTERN UNIV AT QINHUANGDAO

Image background region segmentation and matching method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a segmentation and matching method, device, equipment and medium for an image background region. Multi-scale features are obtained through cavity convolution branch, fusion features are generated through weighted fusion, the fusion features are input into a decoder network to obtain a background segmentation mask, a background area image is obtained by combining the preprocessed image, and a visual converter model is input to extract background feature vectors; and performing similarity matching with a target feature vector in a preset feature database to retrieve a target image. According to the method, the segmentation accuracy and robustness are improved through combination of multi-scale feature fusion and the visual converter, meanwhile, the retrieval discrimination and reliability are enhanced through feature matching, the problems that an existing method is insufficient in segmentation precision and limited in retrieval capacity are solved, and precise segmentation and efficient retrieval of the background in a complex scene are achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent digital human training method and system based on multi-modal interaction

The invention discloses an intelligent digital human training method and system based on multi-modal interaction, and belongs to the technical field of semantic indexing.The method specifically comprises the steps that voice, vision and text data are analyzed and converted into high-dimensional feature vectors through a modal exclusive encoder, the high-dimensional feature vectors are projected to a unified semantic space through a cross-modal semantic mapping model, and the high-dimensional feature vectors are obtained; generating a semantic primitive containing a modal identifier, a core semantic tag and a feature weight; semantic primitives are used as nodes, directed edges and edge weight table association strength are established based on semantic similarity, typical scene node connection weights are strengthened, and a mesh map containing intra-modal hierarchy and inter-modal cross association is formed; constructing a double-layer index on the basis of the mesh map; semantic primitives are extracted from newly added data, the position of a new node in an association graph is determined through a graph matching algorithm, an association edge with an existing node is automatically established, and a lower-layer modal exclusive index is synchronously updated.
Owner:JIANGXI INST OF FASHION TECH

Open vocabulary industrial defect detection method based on multi-modal prior prompt

The invention provides an open vocabulary industrial defect detection method based on multi-modal prior prompt, which comprises the following steps: S1, collecting and sorting defect images of industrial products to be detected, and constructing a large industrial reference data set; s2, building a visual priori prompt pool, and injecting refined and fine-tuned priori into multi-scale industrial features; s3, designing a significance Gaussian distribution modeling mechanism, and capturing a complex spatial mode of position prior, so that fine-grained prior has better expression ability, and the generalization of prior is improved; s4, establishing a decoupling LoRA text encoder, and extracting an industrial semantic basis in the industrial text template through a hierarchical prompt template and a hierarchical decoupling LoRA mechanism; and S5, constructing a visual text fusion unit, and performing multi-modal fusion on text prompt embedding and visual features. According to the method, high-precision and strong-generalization defect detection is realized by fusing visual and text prior prompts, and the robustness and adaptability in a complex open environment are remarkably improved.
Owner:CENT SOUTH UNIV

Network security event tracing method, system and device based on AI and medium

The invention discloses an AI-based network security event tracing method, system and device and a medium, and the method specifically comprises the steps: constructing a network entity association graph based on a multi-modal data set, mining the implicit association between entities through a graph convolutional network, recognizing an APT attack chain, and obtaining graph feature data; based on the multi-modal data set, an LSTM-Transform hybrid model is adopted to analyze time sequence characteristics of network traffic, slow penetration and low-frequency detection behaviors are detected, and time sequence characteristic data are obtained; based on the graph feature data and the time sequence feature data, high-value features are screened through a genetic algorithm, and cross-modal combination features are generated by using a depth auto-encoder; based on cross-modal combination features, a network environment digital twin is constructed, an attack diffusion path is simulated, and a service influence range is quantified. According to the method, accurate tracing of the network security event is realized, and the detection and tracking capabilities of complex network attacks and the intelligent level of a response strategy are comprehensively improved.
Owner:ANHUI SANQI JIYU NETWORK TECH CO LTD

Medical report generation method, model training method, equipment and medium

The invention discloses a medical report generation method, a model training method, equipment and a medium, and the model training method comprises the steps: constructing a medical report generation model framework which comprises a global semantic collaborative multi-modal enhancement module, a visual encoder, a text encoder, a medical insight analyzer and an LLM decoder; wherein the global semantic collaborative multi-modal enhancement module respectively enhances a medical image and a medical report by utilizing a selected image enhancement strategy and a text enhancement strategy, and the medical insight analyzer comprises a fine-grained structure learning device and a global context guide learning device which are connected in sequence so as to enhance the cross-modal alignment capability; and performing intelligent collaborative optimization by taking a strategy set formed by an image enhancement strategy and a text enhancement strategy and architecture configuration parameters of the medical insight analyzer as optimization targets to obtain an optimal medical report generation model. The medical report generation performance can be effectively improved.
Owner:CENT SOUTH UNIV