Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

818 results about "Model architecture" patented technology

Explanatory model architecture for image scoring reasoning

A method includes obtaining an image, the image associated with a mask corresponding to a portion of the image, generating a plurality of images based on the image and the mask, each image of the plurality of images depicting a different color in the portion of the image corresponding to the mask, executing a machine learning model to generate an image performance score for each of the plurality of images, ranking the plurality of images according to the image performance scores for the plurality of images, and generating a record comprising one or more images of the plurality of images based on the rankings of the plurality of images.
Owner:VIZIT LABS INC

Distribution network auxiliary decision-making method and system considering source load fluctuation relevance, and medium

The invention relates to the technical field of power systems and automation thereof, in particular to a distribution network auxiliary decision-making method and system considering source load fluctuation relevance and a medium. The method comprises the following steps: firstly, collecting related information of a distribution network area, quantifying a synchronization and hysteresis association rule of multi-source heterogeneous data fluctuation, and constructing a composite feature vector and a standardized risk perception data set; defining a state space and an action space of a reinforcement learning algorithm based on the composite feature vector, and realizing auxiliary decision-making optimization of the distribution network; constructing a scene feature library, calculating the fluctuation relevance similarity between a new scene and a historical scene, and multiplexing a deep reinforcement learning model architecture and carrying out transfer learning; building a power grid digital twinborn simulation platform, designing evaluation indexes, generating candidate schemes, deducing the candidate schemes, selecting recommendation strategies and storing the recommendation strategies in a strategy knowledge base.
Owner:SUQIAN POWER SUPPLY COMPANY OF JIANGSU PROVINCE POWER +2

Slope early warning method and system based on deep learning

The invention discloses a slope early warning method and system based on deep learning, and particularly relates to the technical field of slope early warning, and the method comprises the steps: S1, multi-source data collection, S2, dynamic graph construction, S3, meta-learning model initialization, S4, space-time fusion prediction, S5, dynamic risk assessment, and S6, graded early warning triggering. Through multi-modal data fusion, an innovative model architecture and an intelligent early-warning mechanism, the slope early-warning capability can be remarkably improved, multi-source data are fused, a cross-modal attention mechanism is utilized, the slope state is comprehensively and accurately reflected, the early-warning accuracy is improved, a dynamic graph structure is constructed to be combined with a meta-learning engine, different slopes are adapted, continuous optimization can be achieved, and the early-warning capability of the slope is improved. Meanwhile, a scientific grading early warning system is established, a historical case library and related equipment are linked, resources are efficiently allocated, life and property safety is guaranteed, and disaster losses are reduced.
Owner:CHINA SHANXI SIJIAN GRP

Finance report analysis method, device and equipment based on multi-source heterogeneous data processing

The invention relates to the technical field of artificial intelligence, and discloses a financial report analysis method, and the method comprises the steps: carrying out the data preprocessing of historical financial association data and historical business operation data, and obtaining to-be-analyzed historical data; performing multi-order feature engineering processing on the to-be-analyzed historical data to obtain historical feature vector data; constructing an initial financial analysis large model, and training the initial financial analysis large model by using the historical feature vector data to obtain a trained financial analysis large model; accessing the trained large financial analysis model into a target system, and triggering the large financial analysis model to operate; and driving the large financial analysis model to perform multi-dimensional semantic analysis and quantitative reasoning on the to-be-analyzed financial report data to obtain visual financial report analysis result data. The method can be applied to internal financial statements of enterprises with businesses of science and technology finance, medical health, old-age care and the like, and the financial statement analysis efficiency and comprehensiveness can be improved through the multi-modal data fusion and knowledge enhancement large model architecture technology.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Wind driven generator fault diagnosis method and system based on Mamba-ResNet

The invention relates to the technical field of fault diagnosis, in particular to a wind driven generator fault diagnosis method and system based on Mamba-ResNet. The method comprises the following steps: carrying out feature extraction and feature fusion by utilizing preprocessed data, namely constructing adaptive window short-time Fourier transform (AW-STFT) to carry out dynamic time-frequency resolution analysis, carrying out parallel feature extraction and constructing a multi-dimensional heterogeneous feature vector, and carrying out a cross-modal adaptive gating fusion mechanism based on a bidirectional cross gating unit; the method comprises the following steps: constructing a Mamba-ResNet hybrid deep network model architecture; performing model training on the constructed network model architecture; and performing fault diagnosis on the wind driven generator by using the trained model architecture. A tedious manual feature design process in a traditional method is avoided, and the automation level and adaptability of a diagnosis system are remarkably improved.
Owner:YANTAI UNIV

Ship trajectory prediction method based on graph attention mechanism and electronic chart

The invention provides a ship trajectory prediction method based on a graph attention mechanism and an electronic chart, and relates to the technical field of intelligent ships. By fusing dynamic AIS trajectory data and static navigation channel geographic information, high-precision modeling and prediction of the future motion trajectory of the ship are realized. According to the method, a traditional time sequence-based trajectory modeling mode is expanded into multi-modal joint modeling, channel structure vector representation is introduced for the first time, structural information of a channel environment is extracted on the basis of a graph neural network, and deep fusion is performed on the structural information and historical trajectory data of a ship; and constructing a trajectory expression mode containing time, space and environment triplex semantics at the same time. In the aspect of model architecture, a graph attention mechanism is adopted to enhance the node representation capability in a channel sub-graph, and meanwhile, a time sequence modeling module is introduced to model a trajectory evolution rule, so that the prediction precision and generalization capability of the model in a complex water area environment are effectively improved, and high-quality modeling, interpretable prediction and intelligent support of the ship trajectory are realized.
Owner:DALIAN MARITIME UNIVERSITY

Server running state monitoring method and system and medium

The invention relates to the technical field of computers, in particular to a server running state monitoring method and system and a medium. The method comprises the steps of obtaining time sequence operation and maintenance data of a system service; the method comprises the following steps: mapping an unstructured log text into a low-dimensional dense text vector by adopting an embedded learning method, splicing the text vector and a standardized numerical vector of a structured index to obtain a unified high-dimensional feature vector, and forming a cross-time vector database; constructing an AI model used for time sequence prediction, anomaly detection and cascade reasoning of classification decision based on a multi-model fusion architecture; training the AI model based on the vector database; and inputting the real-time feature vector into the trained AI model, and outputting to obtain a decision result of the system service operation state. Uniform expression of cross-modal features is effectively realized, cascade model architecture design is cooperated, the dynamic adjustment capability of the model and the accuracy of composite fault detection are improved, and rapid decision-making of fault types is realized.
Owner:HANGZHOU ROBAM APPLIANCES CO LTD

Intelligent scheduling method and system for Huaan Atlas heterogeneous computing resources based on dynamic load awareness

The invention belongs to the technical field of computing resource scheduling, and particularly relates to an intelligent scheduling method and system for Huaan Atlas heterogeneous computing resources based on dynamic load awareness. The method comprises the steps that the real-time state of multi-dimensional hardware data is collected, and a basic data source is provided for subsequent steps; dynamically adapting tasks and hardware characteristics through a matching degree matrix, modeling aiming at various basic data, and constructing a state vector required by reinforcement learning; predicting a fault risk score through a lightweight prediction model deployed at each computing node; and a deep Q network is adopted as a model architecture, a state vector and a fault risk score are input, reinforcement learning training is performed through a reward function in a multi-target vector form, a final scheduling model is obtained, and a task allocation decision is output. The problems that in the prior art, the hardware state cannot be sensed in real time, hardware characteristic matching is ignored, consequently, the computing resource utilization rate is insufficient, and fault recovery is passive are solved.
Owner:SHANDONG ZHIYANG ELECTRIC

Medical image segmentation method based on AFMHiFormer

The invention provides a medical image segmentation method based on an AFMHiFormer. The method comprises the steps that firstly, a multiple data enhancement module is provided, and the data distribution diversity is improved while the enhancement stability is guaranteed; secondly, a segmentation model AFHiMFormer is constructed, and the model architecture adopts a double-branch encoder and a multi-scale decoder; thirdly, a feature enhancement module is provided to construct a dynamic complementation mechanism of semantic enhancement and boundary modeling; fourthly, a multi-scale feature fusion module is introduced, multi-scale context information is captured through parallel hole convolution with different expansion rates, and self-adaptive fusion of global and local features is achieved; and fifth, a cross-scale fusion module is designed in the multi-scale decoder, so that the deep layer branch and the shallow layer branch are efficiently fused in a multi-level feature space. According to the method, the advantages of CNN and Transform are combined, dynamic fusion of local and global features is realized by providing a new module, and a remarkable performance advantage is shown in a medical image segmentation task.
Owner:CHANGCHUN UNIV OF TECH

Dynamic federal mutual learning method and system for balancing personalization and generalization

The invention relates to the technical field of federated learning, in particular to a dynamic federated mutual learning method and system for balancing individuation and generalization, and the method specifically comprises the following steps: each client carries out the preprocessing of data to be processed of a model, and carries out the strong enhancement and weak enhancement processing; inputting the data subjected to strong enhancement processing into a shared model, inputting the data subjected to weak enhancement processing into a private model, and performing iterative training on the two models; related parameters of the shared model after each round of iterative training and a difference item between two model parameters are uploaded to a federation server; the federated server adopts a multi-dimensional adaptive aggregation strategy to obtain an updated global model, and returns the updated global model to each client to replace the shared model in the next round of training; and finally generating a generalization result and a personalized result. According to the method, the private-shared model architecture is constructed, and dynamic federated mutual learning is carried out in combination with the federated server, so that balance and collaborative improvement of individuation and generalization performance can be realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +2

Systems and methods for underwater imagery enhancement

A computer-implemented method for training a generative adversarial network (GAN) for enhancing underwater images. An adversarial loss is computed for updating a discriminator model and a combined loss is calculated for updating a generator model. The combined loss is calculated based on loss components including the adversarial loss and at least one further loss component. Additionally disclosed herein is a generator network for processing underwater images that includes a novel encoder-decoder model architecture. Unlocking insights from Geo-Data, the present invention further relates to improvements in sustainability and environmental developments: together we create a safe and liveable world.
Owner:FNV IP BV

Hyperspectral image classification method of cross-hop node interaction graph attention network

The invention discloses a hyperspectral image classification method of a cross-hop node interaction graph attention network. The method comprises the following steps: step 1, constructing a DNIGAT-CFF overall model architecture; 2, differentiated features are extracted based on the coupled convolution blocks; step 3, cross-hop node interaction graph attention network enhanced spectrum-spatial feature learning; step 4, carrying out multi-scale cross guidance feature fusion CGFF; step 5, important fusion features are highlighted by a weighted attention mechanism; according to the method, the attention network of cross-hop node interaction is constructed, interaction between nodes with different hop counts is effectively utilized, and the extraction capability of spectrum and spatial features is enhanced; a multi-scale cross guide feature fusion module is adopted, complementarity and correlation between different scale features are fully considered, and effective fusion of the multi-scale features is achieved; and in combination with a weighted attention mechanism, important features in the fused multi-scale features are highlighted, so that the precision of hyperspectral image classification is improved.
Owner:QIQIHAR UNIVERSITY

Reinforced learning training method and system for relieving hallusion of multi-modal large model

The invention discloses a reinforcement learning training method and system for relieving illusion of a multi-modal large model, and belongs to the field of reinforcement learning training of a multi-modal large language model. Firstly, a planning and visual description generation step is introduced in an early stage to guide a model to perform structured reasoning, then a grouping relative strategy optimization algorithm is used, reward values are calculated for multiple candidate responses generated by the model after cold start, and particularly, a visual perception reward mechanism is set. The reward mechanism evaluates the consistency of the generated text description and the visual information by using an external large language model. Then, based on a vision description attention score advantage distribution method, learning of the model on key vision signals is dynamically enhanced, and the perception ability of the model on the vision signals is improved; and finally, the perception and reasoning performance of the model is further improved by adopting multiple rounds of rejection sampling and supervised fine tuning. The scheme does not depend on a model architecture, the extra overhead is small, the illusion problem caused by early image-text inconsistency is effectively solved, and the accuracy and the reliability are improved.
Owner:ZHEJIANG UNIV +1

Deep learning image data intelligent supervision system and method

The invention discloses an intelligent image data supervision system and method for deep learning, and relates to the technical field of image processing, and the method comprises the following steps: carrying out the combined deployment of a high-temperature-resistant conventional camera and an event camera, calibrating and obtaining parameters, and setting collection parameters to carry out the multi-modal collection of a foundation pit image; performing time domain weighted averaging and asynchronous event fusion on the acquired data, calculating optical flow by using an algorithm, and tracking ROI in combination with Kalman filtering; building a TSR-WGAN model architecture, constructing a sample set training model, inputting a preprocessing sequence, reasoning, and outputting a corrected image; configuring parameters such as a subset size of a DIC algorithm, and analyzing the corrected image sequence to generate displacement field and strain field data; and indexes such as absolute displacement are calculated, threshold values of all the indexes are set, and foundation pit abnormity early warning is triggered when data exceed the threshold values. The situation that a traditional image data monitoring scheme excessively depends on environmental parameters in a high-temperature scene can be effectively improved.
Owner:HEBEI PROVINCIAL COMM PLANNING & DESIGN INST +2

Modified large language model architecture with span-level attention mechanism for conversion of natural language text to structured knowledge graph

Various embodiments of the present disclosure provide machine learning architectures and data processing techniques for improving computer-based text comprehension. The techniques may include identifying a plurality of data entity tokens from a target section of a multi-section natural language document and generating, using an embedding layer of a semantic chunking model, a text span embedding for a text span of the target section. The techniques may include leveraging the semantic chunking model to generate an attended span representation for the text span based on the text span embedding and the plurality of data entity tokens. The techniques may include identifying an entity topic that corresponds to the text span based on the attended span representation and, responsive to an identification of the entity topic, generating a subgraph data object for a knowledge graph using the text span.
Owner:OPTUM INC

6G intelligent load balancing and fault self-healing method based on AI and network slice

The invention relates to the technical field of network slice resource allocation, in particular to a 6G intelligent load balancing and fault self-healing method based on AI and network slices, which comprises the step of building an AI decision-making layer model architecture, a self-adaptive optimization mechanism, a resource dynamic scheduling mechanism and a strategy execution guarantee architecture. According to the method, the network load trend is predicted in real time through the AI technology, the resource allocation between the slices is dynamically adjusted, and the problem of low resource utilization rate caused by static allocation and periodic adjustment is solved. Meanwhile, an automatic fault detection and recovery mechanism is designed, when a node fault is detected, standby resource takeover can be quickly triggered, session continuity can be kept, the service interruption time is remarkably shortened, and the requirements of a 6G network for high reliability and low time delay are met.
Owner:NANJING AIPULU SATELLITE COMMUNICATION TECHNOLOGY CO LTD

Weak supervision online video moment positioning method and system based on memory perception

The invention relates to a weak supervision online video moment positioning method and system based on memory perception, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the multi-modal feature fusion of a given video and a text query thereof, and obtaining the unified frame level representation at each stage; inputting the fused features into an offline module and an online module in an integral and frame-by-frame manner by using an offline guide online model architecture; in the off-line module, generating a Gaussian mask to reconstruct query of a covered part of words, and obtaining a proposal of an action starting moment; in the on-line module, the long-term historical memory in the window is used for enhancing the score, the attention weight of the score in the window is dynamically generated, and the score of the current frame is calculated in a weighted mode; taking the proposal obtained by the offline module as a pseudo tag, and providing supervision information for the score sequence of the online module; and high-performance weak supervision on-line moment positioning can be completed only by independently deducing the on-line module. The expansion capability and the application value of the model are remarkably improved.
Owner:SHANDONG UNIV

Machine-learned model architecture for predicting future object state

Predicting a future state, such as a future position and / or orientation (i.e., pose), of an object may comprise classifying, by a first machine-learned model, a lane the object may occupy and classifying, by a second machine-learned model, a target pose the object may occupy. A third machine-learned model may determine an offset from the target pose that may be used to determine a predicted (future) pose of the object by applying the offset to the target pose.
Owner:ZOOX INC

Shield tunnel land subsidence real-time prediction method considering spatio-temporal information

The invention discloses a shield tunnel ground subsidence real-time prediction method considering spatio-temporal information, and the method comprises the steps: constructing an input feature system, including classifying model input into geometric information, multi-ring geological condition information, multi-time step shield parameter information and historical subsidence information; a multi-source information fusion model architecture is designed, geometric information, geological condition information, shield operation parameter information and historical settlement information serve as input of the model, and feature extraction is conducted through multiple encoders. Then feature fusion is carried out, nonlinear mapping is carried out by using a residual network (ResNet), and finally a real-time settlement prediction value is output; according to the scheme, the space-time characteristics of ground subsidence induced by tunneling are deeply excavated, and real-time prediction of subsidence of any position in a disturbance range is realized through deep fusion of multi-source characteristics.
Owner:SOUTHEAST UNIV +1

Robot task data processing method based on double-model architecture and robot

The invention provides a robot task data processing method based on a double-model architecture and a robot, and belongs to the technical field of electric digital data processing. Generating a task sequence corresponding to the first user instruction through the decision-making layer text generation large model, and generating a first action list of each task in the task sequence through the execution layer multi-mode large model. And the ROS main control unit calls an execution unit to execute the first action list so as to complete the current task, obtain task feedback information and update the context information set. The ROS main control unit enters a preset state, when a preset condition is not met, the execution layer multi-mode large model generates a first action list of each task according to the execution layer multi-mode large model, when the preset condition is met, the ROS main control unit enters a waiting state, and after a temporary instruction is obtained, the execution layer multi-mode large model generates a second action list based on the updated context information set and executes the second action list. The defect that modal interference is prone to occurring in the process of processing complex tasks during single AI large model control can be overcome.
Owner:SHENZHEN YAHBOOM TECH CO LTD

Conversational agnostic matchmaking model architecture

A system for matchmaking using a conversational agnostic matchmaking model is described. The system can receive a first query indicating a request for document objects and including criteria for selection of the document objects. The system can identify named entities from portions of the first query. The system can generate a second query to obtain the document objects, in response to the named entities being indicative of a context for the first query. The system can obtain the document objects according to the second query. The system can generate a reply to the first query including a description object and the documents, the description object based on the first query. The system can cause a user interface to present the reply to the first query and the description object.
Owner:ADP INC

Gust front wind shear identification method based on artificial intelligence

The invention provides a gust front wind shear identification method based on artificial intelligence, and the method comprises the steps: carrying out the noise filtering, missing value supplementary measurement, data smoothing, wind shear value calculation and sample screening extraction of collected radial speed data, and obtaining a gust front wind shear sample; performing coordinate system conversion, sample set division and data annotation on the basis of gust and front wind shear samples to obtain an expanded data set; designing and training a Mask R-CNN model architecture to obtain a gust and front wind shear identification model; and inputting the expanded data set into the gust and front wind shear identification model to carry out gust and front wind shear detection. According to the method, dependence on reflectivity factor data can be reduced, an identification model is constructed based on gust front radial speed data, gust front wind shear can be accurately identified, pixel-level segmentation and positioning of a wind shear area can be realized, and identification efficiency is improved.
Owner:CHENGDU UNIV OF INFORMATION TECH

Methods and system for industrial defect identification

An inspection system includes a model architecture for industrial defect identification. The model architecture includes a text encoder model that receives a text object having free-form text and generates a text embedding. A visual encoder model receives a region of interest of an image and generates a region embedding. A cross-modality fusion layer acts between the text encoder model and the visual encoder model to fuse outputs of nodes within the models to be used as inputs to nodes in a subsequent layer. A cross-modality decoder model aligns the text embedding and the region embedding to generate a bounding box for the region if it is similar to the text object. A positional encoder generates a positional embedding based on the bounding box. A mask decoder model generates a segmentation mask based on the positional embedding within an output to highlight the region defined by the text object.
Owner:RTX CORP

Course analysis management system based on deep learning

The invention relates to the technical field of course analysis management, in particular to a course analysis management system based on deep learning. The method has the advantages that multi-modal data deep analysis realizes full-dimensional analysis of unstructured data by integrating a 3D-CNN model, a Transform architecture and a BERT model and synchronously extracting an attention hot area, a voice emotional state and a text knowledge point association network of a classroom video; nonlinear behavior modeling adopts an LSTM network and time convolutional network fusion model, a knowledge internalization path and forgetting curve prediction are dynamically generated, parameters are optimized in combination with incremental learning, and a transition rule across knowledge points is captured; a teaching scene-evaluation threshold mapping table is constructed based on a reinforcement learning algorithm through dynamic decision and resource collaboration, collaborative optimization under multi-campus data privacy protection is achieved in combination with a federated learning framework, GPU computing nodes are dynamically allocated through a heterogeneous resource scheduling engine, and the analysis efficiency is improved.
Owner:ZHUHAI QIYAO IND CO LTD

Front-end resource dynamic preloading method, system, equipment and medium

The invention provides a front-end resource dynamic preloading method, system and device and a medium, and belongs to the technical field of front-end engineering. The method comprises the following steps: constructing and training a user behavior prediction model based on a Transform time sequence model architecture, and deploying the user behavior prediction model to a browser side after compression processing; the method comprises the following steps: acquiring behavior data of a user by monitoring a user interaction event, constructing a time sequence operation sequence, and collecting network information and equipment performance information; inputting the time sequence operation sequence into the model, outputting probability distribution of the next operation, and determining related resources; based on the probability distribution, the resource type and the network information, calculating the priority of the resource, and generating a resource priority queue; according to the network information and the equipment performance information, determining a pre-loaded resource type, and adjusting a resource priority queue to generate a pre-loading strategy; a cache rule is defined, a cache configuration file is generated, and a CDN cache is used for preloading resources; and executing the preloading strategy, and carrying out rollback processing after loading fails.
Owner:SHANDONG INSPUR CLOUD GOVERNMENT INFORMATION TECHNOLOGY CO LTD

Tunnel group road section traffic capacity evaluation method based on intelligent network connection

The invention specifically relates to a tunnel group road section traffic capacity assessment method based on intelligent network connection, which relates to the technical field of traffic engineering, and comprises the steps of constructing an AI-driven dynamic assessment model based on fusion data and a digital twinborn model in combination with tunnel engineering constraints, and calculating the traffic capacity, congestion risk and bottleneck position of each road section of a tunnel group in real time. In the invention, a multi-source data acquisition system is combined with an edge and cloud two-stage fusion architecture to construct an LSTM-XGBoost double-model architecture; the LSTM model can calculate and predict the traffic capacity of four time nodes in the future 10 minutes in real time, the XGBoost model accurately locates the bottleneck position and quantifies the contribution degree of four types of causes, meanwhile, the precision is ensured through multi-dimensional verification, the tunnel group whole domain can be covered, instantaneous traffic changes can be captured, and a panoramic decision basis of real-time data, prediction trend and bottleneck causes is provided for traffic management.
Owner:FUJIAN CHUANZHENG COMM COLLEGE

Wind power cluster short-term power prediction method and device based on space-time diagram neural network

The invention relates to a wind power cluster short-term power prediction method and device of a space-time diagram neural network fused with physical information and computer equipment, and the method comprises the steps: obtaining related information data of each wind power plant in a wind power cluster, and carrying out the preprocessing; forming a physical prior data set through an engineering analysis model fusing the wake flow analysis model and the blocking effect model; taking each wind power plant as a node of the graph, constructing graph structure data for predicting the power of the wind power plant, and forming a dynamic adjacent matrix; constructing a space-time diagram neural network WB-STGNN model architecture comprising a diagram convolutional neural network module, a gating time convolutional network and a multi-layer perceptron; the method comprises the following steps: pre-training by using a physical prior data set, and then performing formal training based on historical power data and a dynamic adjacency matrix to obtain a space-time diagram neural network WB-STGNN model; inputting the wind speed of the prediction day, and predicting the active power of the whole wind power cluster in 24 hours of the prediction day. By adopting the method, the precision and efficiency of wind power cluster power prediction can be effectively improved.
Owner:HOHAI UNIV +1

Low-cost visual field large model for visual multi-modal information processing

The invention provides a low-cost visual field large model for visual multi-modal information processing, and the model comprises an image encoder module which is used for converting an input image into low-dimensional feature representation; the feature extraction module is used for extracting multi-scale visual features based on a hierarchical multi-task learning strategy and reducing redundant calculation through a cross-modal parameter sharing mechanism; the task specifying module is used for designing a lightweight sub-model for image classification, target detection and image generation tasks, and integrating pruning and quantification technologies to optimize calculation efficiency; the reasoning optimization module is used for reducing model reasoning complexity and energy consumption by adopting a low-rank decomposition LoRA and mixed precision calculation technology; and the multi-modal fusion module is used for integrating vision, text and sensor data through a cross-modal attention mechanism to generate cross-modal joint feature representation so as to improve task robustness in a complex scene. According to the method, through the design of model architecture, parameter quantity and calculation optimization, the requirement for hardware resources is remarkably reduced.
Owner:TONGJI UNIV

Large language model architecture for high quality and high throughput

The embodiments may include a computer readable medium including instructions that when executed by one or more processing devices cause the one or more processing devices to perform a method. The method may include inputting a prompt to a trained model. The trained model may be a Large Language Model comprising a plurality of neural network layers including: at least one first layer comprising an attention module associated with a Transformer deep learning architecture; and at least one second layer comprising a state space module associated with a Mamba deep learning architecture. The method may further include receiving an output of the trained model.
Owner:AL21 LABS

Systems and methods for link resolution for internal entities and documentation using pre-seeded language models

Systems and methods for an artificial intelligence model architecture that involves a first artificial intelligence model trained to map a plurality of entities to ranked documentation from a documentation source, and a second artificial intelligence model that comprises a language model trained to generate an additional query to run on the plurality of documents from the documentation source. By training the second model to generate additional queries as entities and / or links are discovered, the system may quickly and efficiently determine links and / or potential resolutions as well as received feedback thereon.
Owner:CAPITAL ONE SERVICES LLC