Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11 results about "Language module" patented technology

The language module, also known as the "language faculty", is a hypothetical structure in the human brain or cognitive system which contains innate capacities for language, originally posited by Noam Chomsky. There is ongoing debate about this in the fields of cognitive science and neuroscience.

Remote sensing visual language agent construction method and device based on hybrid strategy

The invention discloses a remote sensing visual language agent construction method and device based on a hybrid strategy. The method comprises the following steps: constructing a multi-modal data set in the field of remote sensing; taking the multi-modal data set as a supervision signal, and performing supervision fine tuning on the visual language model to obtain an initial visual language model; constructing a reward mechanism, sampling the initial visual language model for multiple times, and determining an advantage value of each sampling according to the reward mechanism; converting the advantage value of each sampling into a relative advantage value, and optimizing parameters of the initial visual language model based on the relative advantage value to obtain a remote sensing visual language module; and receiving a remote sensing task input by a user through a remote sensing visual language module, generating a natural language answer, and synchronously outputting a thinking chain. The problems that an existing visual language model is high in sample requirement and poor in generalization ability and performance are solved. Therefore, various downstream tasks in the remote sensing field are solved, and the generalization ability and the processing efficiency of the model in complex remote sensing tasks are remarkably improved.
Owner:XI AN JIAOTONG UNIV

Induced dialogue interaction system based on psychological intervention artificial intelligence model

The invention is applied to the technical field of dialogue interaction systems, and particularly discloses an induced dialogue interaction system based on a psychological intervention artificial intelligence model, which comprises a voice recognition module, a deep language module, a voice output module, a closed feedback module and an abnormal condition processing module, the voice recognition module is connected with the deep language module through wireless signals, and the deep language module is connected with the voice output module through wireless signals. According to the induced dialogue interaction system based on the psychological intervention artificial intelligence model, psychological theories such as a social emotion selection theory, a communication adaptation theory and the like are introduced in the system, and an old people age layering optimization deep language model is combined, so that the recognition and fitness reply ability of the model to common emotion expression of old people is enhanced; meanwhile, by means of intervention boundary setting, psychological intervention scenes of the old people are accurately adapted, and intervention safety and effectiveness are guaranteed.
Owner:SUZHOU UNIV

Explainable language model training method and system for knowledge-intensive tasks

PendingCN122451494ALanguage moduleRecognition algorithm
The application relates to an explainable language model training method and system for a knowledge-intensive task, and the method comprises the following steps: S1, acquiring original text data; the original text data is a question proposed according to a knowledge-intensive task; S2, inputting the original text data into a preset large language model for training; S3, adopting a hierarchical functional circuit recognition algorithm to recognize a functional circuit in the trained preset large language model; S4, calculating the correlation degree of the functional circuit and the knowledge-intensive task, wherein the correlation degree is a weighted combination based on a selectivity index and an influence index; and S5, outputting an explainable language module, wherein the explainable language module comprises the functional circuit and the corresponding correlation degree. The functional circuit inside the large language model is induced and recognized, mechanism-level transparency is provided, and the method is far superior to existing surface explanation methods.
Owner:HUNAN CREATOR INFORMATION TECH CO LTD

Flood remote sensing image language understanding method and system based on priori guidance large model

PendingCN122289953Aquality improvementsolid data foundationPattern recognitionLanguage understanding
This invention relates to the field of flood remote sensing image understanding, specifically disclosing a method and system for language understanding of flood remote sensing images based on a priori-guided large-scale model. The method includes: constructing a dataset containing multi-temporal flood remote sensing images and multi-granularity linguistic descriptions; constructing and training a flood prior model to extract prior information on the spatial distribution of floods and convert it into textual prior information; constructing a large-scale language model containing a visual encoder, a projection layer, and a language module, and training it by inputting the images, textual descriptions, and textual prior information together; inputting the image to be understood and the corresponding textual prior information into the trained large-scale language model, and outputting the language description result. This invention, by introducing a flood prior model to provide explicit spatial distribution guidance, effectively alleviates the semantic understanding bias and spatial positioning error of general visual language models in flood scenarios, significantly improving the accuracy of linguistic descriptions of flood images.
Owner:INFORMATION COMM COMPANY STATE GRID SHANDONG ELECTRIC POWER +1

Task-adaptive visual language large model collaborative pruning method

The invention relates to the field of language model pruning, and particularly discloses a task-adaptive visual language large model collaborative pruning method, which comprises the following steps of: S1, performing task perception, acquiring a training data set containing image text pairs, shielding a corresponding data set of a single mode, and calculating loss value fixed mode dependency so as to adaptively bias and finely adjust visual and language module shares; s2, calculating importance scores, dividing model weights, hierarchically calculating gradient norm approximate values, and multiplying the gradient norm approximate values with modal shares to obtain scores of all layers; s3, sparseness is distributed, summarizing scores are normalized, the sparseness of each layer is calculated in combination with a total parameter retention proportion, and a threshold value is set; and S4, structured sparseness is carried out, a pruning strategy and token screening are combined, a fusion score of modal correlation degrees is calculated, effective tokens are reserved to calculate channel scores, and key channels are reserved according to groups. According to the technical scheme, the problems that traditional pruning is poor in adaptation, task modal dependence is not considered, and key parameters are easily cut by mistake are solved, compression and reasoning efficiency is improved, cross-modal information loss is avoided, and performance loss is reduced.
Owner:SHENZHEN UNIV

Method and device for constructing remote sensing visual language agent based on hybrid strategy

ActiveCN121686220BLanguage moduleEngineering
This application discloses a method and apparatus for constructing a remote sensing visual language agent based on a hybrid strategy. The method includes: constructing a multimodal dataset in the remote sensing domain; using the multimodal dataset as a supervisory signal to fine-tune a visual language model under supervision, obtaining an initial visual language model; constructing a reward mechanism to sample the initial visual language model multiple times, and determining the advantage value for each sample based on the reward mechanism; converting the advantage value for each sample into a relative advantage value, and optimizing the parameters of the initial visual language model based on the relative advantage value, obtaining a remote sensing visual language module; receiving remote sensing tasks input by the user through the remote sensing visual language module, generating natural language answers, and simultaneously outputting the thought chain. This solves the problems of existing visual language models having high sample requirements and poor generalization ability and performance. Furthermore, it addresses various downstream tasks in the remote sensing domain, significantly improving the model's generalization ability and processing efficiency in complex remote sensing tasks.
Owner:XI AN JIAOTONG UNIV

Virtual assistant domain functionality

Aspects include methods, systems, and computer-program products providing virtual assistant domain functionality. A natural language query including one or more words is received. A collection of natural language modules is accessed. The collection natural language modules are configured to process sets of natural language queries. A natural language module, from the collection of natural language modules, is identified to interpret the natural language query. An interpretation of the natural language query is computed using the identified natural language module. A response to the natural language query is returned using the computed interpretation.
Owner:SOUNDHOUND AI IP LLC

Task-adaptive visual language large model collaborative pruning method

The present application relates to the field of language model pruning, and specifically discloses a task-adaptive visual language large model collaborative pruning method, comprising: S1, task perception, obtaining a training data set containing image text pairs, shielding the corresponding data set of a single mode, calculating the loss value and the mode dependency, and adaptively biasing the visual and language module share; S2, calculating the importance score, dividing the model weight, calculating the gradient norm approximation in layers, multiplying the mode share to obtain the score of each layer; S3, allocating sparsity, normalizing the score, calculating the sparsity of each layer combined with the total parameter retention ratio and setting a threshold; S4, structured sparsity, combining pruning strategy and token screening, calculating the fusion score of mode correlation, calculating the channel score by retaining effective tokens, and retaining key channels according to groups. The technical scheme provided by the present application solves the problems of poor adaptation of traditional pruning, not considering task mode dependency, and easy mis-pruning of key parameters, improves compression and inference efficiency, avoids cross-modal information loss, and reduces performance loss.
Owner:SHENZHEN UNIV

Visual language model asymmetric fine tuning method and system and storage medium

The invention discloses an asymmetric fine tuning method and system for a visual language model, and a storage medium, and the method employs an asymmetric decoupling fine tuning strategy for the structural heterogeneity of a visual and language module of a large visual language model: unfreezing a visual encoder, carrying out the fine tuning of a layer normalized affine parameter in a Transform block, and calibrating the feature distribution offset caused by high-resolution input; for the language model backbone, a learnable Switch GLU-Adapter module is inserted through residual parallel connection, and the fine-grained reasoning capability is enhanced; only a small number of key parameters of the model are updated, and other pre-training parameters are frozen; in a high-resolution visual question and answer task, the method not only adapts to differentiated requirements caused by heterogeneity of modules, but also remarkably reduces training calculation and storage overhead, avoids disastrous forgetting, and meanwhile ensures the adaptation performance of efficient fine tuning of parameters.
Owner:NANJING UNIV

Artificial intelligence device for common sense reasoning for visual question answering and control method thereof

A method for controlling an artificial intelligence (AI) device can include receiving, via a processor in the AI device, an input image and a query related to the input image, generating, via the processor, an answer prompt template based on the query, the answer prompt template including a sentence containing a mask token located at a position corresponding to an answer within the sentence, and combining the query and the answer prompt template to generate a string of text including the mask token. Also, the method can further include inputting the string of text to a pre-trained mask language module (MLM) and generating a plurality of scores respectfully corresponding to a plurality of answers, each of the plurality of answers being a candidate for replacing the mask token, determining a selected answer among the plurality of answers based on the plurality of scores, and outputting the selected answer.
Owner:LG ELECTRONICS INC

A speech recognition system and method based on a double-weighted directed graph

A speech recognition system and method based on a double-weight directed graph, comprising the following steps: 1) a language flow module divides continuous speech into multiple audio frames, and sends the obtained multiple audio frames to an acoustic module; 2) the acoustic module converts the audio frames into an acoustic model; 3) a language module generates a directed and weighted decoding graph according to a text knowledge base, and sends the weighted directed decoding graph to a double-weight wide search module; 4) the double-weight wide search module processes the acoustic model and the weighted directed decoding graph to preliminarily generate a recognition result; 5) an identification output module receives the recognition result and path total loss value sent by the double-weight wide search module, and further processes the two to obtain a language recognition result, and outputs the language recognition result. The application can consider the recognition of logical language output and non-logical speech, retain the use of a language model to remove the influence of pronunciation proximity, and improve the recognition rate of non-logical language such as a telephone number.
Owner:WUHAN XUJIAN TECHNOLOGY CO LTD