Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

37 results about "Language modelling" patented technology

Method and system for pre-training sign language understanding framework based on semantic enhancement of skeleton

The invention provides a method and system for pre-training a sign language understanding framework based on semantic enhancement of skeletons, and relates to the technical field of sign language recognizing.The method comprises the steps that sign language video data and skeleton sequences and text data matched with the sign language video data are obtained; extracting a skeleton key point sequence from the sign language video data, modeling to form skeleton features, and performing word segmentation processing on a text; in a pre-training stage, skeleton features and word segmentation texts are input into a fusion network to generate bidirectional enhanced features, and global and local similarities are obtained through double-layer semantic alignment; calculating the comparison loss based on the similarity, and coordinating the weight through the balance parameter to obtain the hierarchical loss; executing the matching task and the language modeling task to obtain corresponding loss, weighting and combining the three types of loss into pre-training total loss, and adjusting parameters to complete training; and finally, part of parameters are optimized in combination with specific task types in a fine tuning stage, and enhanced understanding of sign language semantics is realized. The sign language understanding accuracy is improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

MBSE-oriented multi-agent collaborative automatic modeling method and system

The invention relates to the technical field of model-driven system engineering modeling, in particular to an MBSE-oriented multi-agent collaborative automatic modeling method and system. The method comprises the following steps: acquiring a natural language modeling demand of a user, retrieving in a shared knowledge environment to obtain a knowledge sub-graph related to the demand, constructing a modeling context, and calling a generation agent to generate a current SysML v2 model; calling a verification agent to perform grammar verification and semantic verification based on the domain knowledge graph on the model, and outputting a verification result; when a verification result represents that grammar does not pass or semantic inconsistency exists or an engineering constraint rule is violated, calling a repair agent to execute deterministic repair and iterative verification according to the domain knowledge graph until a preset condition is met, and outputting a final SysML v2 model; notch marking or pocket bottom repairing can be carried out under the condition that the notch is covered. According to the method, the correctness, interpretability and convergence efficiency of complex system modeling are improved.
Owner:CHANGCHUN UNIV OF SCI & TECH

Data processing method and apparatus

PCT designated stageWO2026046246A1Semantic analysisNeural learning methodsData setLanguage modelling
The present application provides a data processing method and apparatus. The method comprises: acquiring a natural language processing (NLP) training data set; inputting the NLP training data set into a first large language model to train the first large language model, the trained first large language model comprising a first language modeling parameter; acquiring the first language modeling parameter; obtaining first text data on the basis of the NLP training data set; and obtaining a first prompt word set on the basis of the first text data, the first language modeling parameter, and a second large language model, wherein the first prompt word set and the first text data are used for training a third large language model, and the first large language model and the second large language model are based on the same model architecture. The method provided by the present application can obtain the first prompt word set of which the data volume is much smaller than that of the NLP training data set. In this way, using the first prompt word set to train the third large language model can improve the speed and accuracy of large language model training.
Owner:HUAWEI TECH CO LTD

Yi language speech recognition method based on self-supervision and attention feature fusion

The invention relates to the technical field of natural language processing, and discloses a Yi language speech recognition method based on self-supervision and attention feature fusion. The method comprises a feature encoder module, a comparative learning module, a mask language modeling module, a joint optimization and feature fusion module and a decoder module, the feature encoder module adopts a convolutional neural network structure and converts continuous waveform signals into feature representation suitable for subsequent modeling, and the comparative learning module performs feature fusion on the continuous waveform signals through a Gumbel-Softmax technology. The method comprises the following steps that: a mask language modeling module and a feature fusion module are integrated, deviation caused by manual definition or clustering is avoided, the mask language modeling module obviously enhances semantic understanding and tone modeling capabilities of a model in a low-resource scene, a self-attention feature fusion mechanism is introduced into the feature fusion module, continuous features, discrete unit representation and semantic context representation from an acoustic level are spliced, and a self-attention feature fusion mechanism is introduced into the self-attention feature fusion mechanism. The decoder module adopts a decoder structure based on connection time sequence classification, and the corresponding relation between the voice and the text can be achieved without strict alignment labeling.
Owner:KUNMING UNIVERSITY

Zero-shot domain transfer with a text-to-text model

Example solutions for zero-shot domain transfer with a text-to-text model train a text-to-text model for a target domain using unlabeled in-domain text training data, and concurrently train the model using labeled general-domain task training data. The in-domain training comprises masked language modeling (MLM) training, and the task training comprises both natural language generation (NLG) training and natural language understanding (NLU) training. The NLG training comprises natural language inference (NLI) training and the NLU training comprises summarization training. The trained model acquires domain-specific task competency, sufficient to perform a language task within the target domain. Suitable target domains include radiology, biomedical, and other medical, legal, and scientific domains. This approach leverages large volumes of general-domain task training data and plentiful unlabeled in-domain text, even as labeled in-domain training data may be unavailable or prohibitively expensive for certain specialized domains.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Multi-modal large language model training method and device, equipment and medium

The invention discloses a multi-modal large language model training method and device, equipment and a medium, which are applied to financial / medical scenes, and are characterized in that a target multi-modal data training set is obtained by preprocessing an obtained multi-modal data training set containing sample images and corresponding sample texts; according to the target multi-modal data training set, a visual encoder and a large language model are connected in series to obtain a first training structure, loss calculation is carried out on the first training structure, and language modeling loss is obtained; and performing loss calculation on a second training structure obtained by connecting the visual representation alignment middle layer, the visual basic model, the visual-language projection module and the alignment loss module in series according to the target multi-modal data training set to obtain alignment loss, and training a multi-modal large language model based on language modeling loss and the alignment loss to obtain the multi-modal large language model. And the target multi-modal large language model is obtained until the model training ending condition is reached, so that the alignment effect of the multi-modal information is improved, and the understanding ability of the model for the image visual information is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Systems and methods for utilizing topic models to weight mixture-of-experts for improvement of language modeling

ActiveUS12718016B2EngineeringLanguage modelling
Systems and methods are disclosed for predicting a next text. A method may include receiving one or more documents, such as a document associated with a healthcare provider. The document is then processed to generate one or more tokens which are representative of the document. The document is then processed with a machine-learning model, such as a topic model, and a topic vector is output for the document. Based at least in a part on this topic vector, the document is then processed by one or more expert machine-learning models, which each output a probability vector. The various probability vectors are then further processed to calculate a total probability vector for the document. Based at least in part on the total probability vector for the document, a text output is selected.
Owner:UNITEDHEALTH GROUP INC

Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning

A computer-implemented method may use a neural network-based language model for generating a 3D structure of a molecule conditioned by atomic connectivity or a binding of the molecule with a target protein pocket. The language model may be trained with the large-scale training data set to obtain a pre-trained language model, wherein the pre-trained language model may be trained to generate the 3D structures of molecules and / or protein pockets. The pre-trained language model may be trained with the finetuning training data to obtain a finetuned language model capable of at least one of: generating 3D structures of molecules capable of binding to the target protein pockets or generating 3D structures of molecules bound in the target protein pockets or generating 3D structures of the target protein pockets. The trained language model capable of generating the 3D structure of a generated molecule that binds with the protein pocket.
Owner:INSILICO MEDICINE AI LIMITED

Apparatuses, methods, and computer program products for processing service message data objects via large language modeling and classification machine learning to provide service message classifications

Methods, apparatuses, or computer program products that process service message data objects via large language modeling and classification machine learning to provide service message classifications. In some examples, a large language model is applied to a plurality of service message data objects associated with an application framework to generate a feature set for the plurality of service message data objects, a classification machine learning model is applied to the feature set to generate a plurality of classification data objects associated with the plurality of service message data objects that classify a respective service message data object as belonging to a predefined class of a plurality of predefined classes, and a rendering of a dashboard visualization is initiated via an electronic interface based at least in part on the plurality of classification data objects.
Owner:ATLASSIAN PTY LTD

MRNA codon sequence language modeling method based on bidirectional Mama architecture

The invention relates to an mRNA (messenger Ribonucleic Acid) codon sequence language modeling method based on a bidirectional Mama framework, which comprises the following steps of: acquiring an mRNA coding sequence, performing sequence tokenization on the mRNA coding sequence, and querying by adopting a word list to obtain an embedding characteristic; the embedded features are sequentially input into a convolution feature extraction module and a state space module of a bidirectional Mamba model to generate collaborative representation, and meanwhile, each time step serves as an affine pair and is used for carrying out inter-block synthesis on block-level pairs to obtain a synthesis result; and inputting the collaborative representation and synthesis result into a bidirectional branch and linear fusion module to obtain a linear transformation fusion result, projecting the linear transformation fusion result to a word list dimension, obtaining a cross entropy and a model parameter used for training a bidirectional Mama model, and generating a target dimension vector of each sequence by using the trained bidirectional Mama model. According to the invention, the accuracy and expandability of mRNA sequence analysis and design can be improved.
Owner:HANGZHOU INSTITUTE OF MEDICAL SCIENCES CHINESE ACADEMY OF SCIENCES

Training machine-learning model to generate an embedding for an online session

An online system trains a machine-learning model to generate an embedding in real time for a current session of a user with the online system. The machine-learning model is trained by applying a masked language modeling algorithm to training data including a training sequence of actions and a masked action to predict a user’s action that follows the training sequence of actions. The online system captures current session data describing a sequence of actions of the user performed during the current session. The online system applies the trained machine-learning model to predict a next user’s action and generate a session embedding that encodes information about the sequence of actions and the next action. Using the session embedding, the online system ranks a list of objects. The online system generates a user interface signal causing a user’s device to display a user interface with the ranked list of objects.
Owner:MAPLEBEAR INC

Power operation personal risk static assessment method, system and equipment based on comparative learning and meta-learning prompts, and medium

The invention discloses an electric power operation personal risk static assessment method, system, equipment and medium based on comparative learning and meta-learning prompts.The electric power operation personal risk static assessment method and system based on comparative learning and meta-learning prompts.The electric power operation personal risk static assessment method and system based on comparative learning and meta-learning prompts.The electric power operation personal risk static assessment method and system based on comparative learning and meta-learning prompts.The electric power operation personal risk static assessment method comprises the steps that deep semantic alignment of text and numerical bimodal features is achieved through comparative Introducing a meta-learning mechanism to enable the deep learning model to obtain cross-task generalization ability in a small sample scene; the classification task is reconstructed into a language modeling task through prompt learning, and end-to-end joint optimization is carried out by means of learnable prompt vectors and features. According to the method, an improved comparative learning framework is adopted, and a loss function of cross-modal attention and innovation is utilized, so that deep semantic alignment of text and numerical features in a unified semantic space is realized, and the problem of a shallow layer of traditional multi-modal fusion is solved.
Owner:YUNNAN POWER GRID CO LTD KUNMING POWER SUPPLY BUREAU

Function annotation abundance sequence-based base model training method and device

The invention relates to a base model training method and device based on a functional annotation abundance sequence. The method comprises the following steps: S1, carrying out function annotation on an open reading frame of a genome or metagenome sample; s2, counting the occurrence frequency of each annotation and constructing a sequence according to an abundance descending order; s3, after the sequence is subjected to token processing, inputting the sequence into a model based on Transform, and adopting joint training of language modeling, comparative learning and classification loss to obtain species-level and token-level fixed dimension embedding; s4, on the basis of the embedding, completing downstream tasks such as phylogenetic tree construction, species identification and phenotype prediction, BGC / MGC recognition and key gene positioning in three levels, namely a genome level, a gene cluster level and a gene / protein level. In the embodiment of the invention, good uniformity, expandability and interpretability are displayed, and the dependence on a reference database is reduced. The corresponding device comprises a data processing module, a model training module and an application module, and can be realized by program instructions in a computer readable storage medium.
Owner:ZHEJIANG LAB

A restaurant reasoning large language model training method based on task vector fusion

This invention pertains to the field of intelligent technology in the catering industry, specifically a training method for a large-scale reasoning language model based on task vector fusion. It involves collecting comprehensive business data from the catering industry, achieving cross-modal entity alignment based on fractional-order grey relational analysis, constructing a domain knowledge graph including a causal relationship layer, and enhancing the corpus. A domain-adaptive base model is obtained through pre-training with causal language modeling loss as the primary loss and combined with dual-auxiliary loss. Task vectors are then fine-tuned and purified through adaptive low-rank decomposition. Adaptive task weights are calculated based on Sobol global sensitivity analysis, and task vectors are nonlinearly fused using Caputo fractional derivatives to obtain a multi-task fusion model. Finally, multi-source causal evidence is fused based on Riemannian manifold geodesic causal distance and D-S evidence theory to enhance the model's causal reasoning training. This method improves the model's multi-task balance performance and causal reasoning accuracy.
Owner:BEIJING QUESHI TECH DEV CO LTD

Text error correction model training method and device, electronic equipment and storage medium

The invention provides a text error correction model training method and device, electronic equipment and a storage medium, and belongs to the technical field of natural language process.The method comprises the steps that sample data are obtained, the sample data comprise input data and output data, the input data comprise a text error correction instruction and an original text, and the output data comprise a standard text; the standard text is a text obtained by performing error correction on an original text through a pre-trained text error correction model, and the text error correction instruction is used for indicating a task needing to be executed by the model; based on the original text, performing weighted average on vocabulary units in the vocabulary to obtain a linguistic intention vector of the original text; based on a preset language modeling task, performing statement reasoning on the language intention vector to obtain a statement reasoning result; and adjusting model parameters of the text error correction model based on the statement reasoning result and the standard text. Through the scheme, unsupervised reinforcement learning is applied to the text error correction task.
Owner:HISENSE XINGHAI TECHNOLOGY (HANGZHOU) CO LTD

Recommendation method and system based on language modeling and elastic reasoning collaborative architecture

The invention provides a personalized recommendation method and system based on a language modeling and elastic reasoning collaborative architecture, and belongs to the technical field of recommendation. The system comprises a data module, a language representation construction module, an elastic sub-model generation module, a dynamic routing training module and a self-adaptive recommendation module. In the implementation process, the language representation construction module formats user behaviors and article information into a natural language sequence through a predefined personalized prompt template, unified coding is conducted through a large language model, and user and article language representations with rich semantics are generated. The elastic sub-model generation module carries out multi-scale structure segmentation on a feed-forward layer and multi-head attention in a unified Transform architecture, and a series of nested sub-models with gradually increased calculation complexity and consistent semantics are constructed. In the model training process, the dynamic routing training module randomly activates sub-models of different scales for forward and reverse propagation in each training step, introduces a task-aware routing mechanism, and dynamically selects the optimal sub-model according to the characteristics of the current recommendation task in the reasoning stage. And the self-adaptive recommendation module loads the trained elastic language recommendation model, and calls the sub-model selected by the dynamic route to generate a recommendation result according to the context information and the task type of the target user. Through wide experimental verification, the prediction performance superior to that of a full-size model is realized on various recommendation tasks, and the problem of performance fluctuation of a large language model in a recommendation scene due to mismatching of a model scale and task requirements is effectively solved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Multi-agent collaborative automation modeling method and system oriented to mbse

The application relates to the technical field of model-driven system engineering modeling, in particular to a multi-agent collaborative automatic modeling method and system for MBSE. The method obtains user natural language modeling requirements, retrieves and obtains a knowledge subgraph related to the requirements in a shared knowledge environment, constructs a modeling context, calls a generation agent to generate a current SysML v2 model, calls a verification agent to perform syntax verification and semantic verification based on a domain knowledge graph, and outputs a verification result; when the verification result represents that syntax is not passed or there is semantic inconsistency or violation of engineering constraint rules, a repair agent is called to perform deterministic repair according to the domain knowledge graph and iterative verification until a preset condition is met to output a final SysML v2 model; in the case of coverage gaps, gap marking or bottom repair can be performed. The application improves the correctness, explainability and convergence efficiency of complex system modeling.
Owner:CHANGCHUN UNIV OF SCI & TECH

Low complexity prefix processing in language modeling

Various embodiments include methods, and computing devices that perform the methods, of improving execution of a generative artificial intelligence model. Embodiment methods may include receiving an input prompt that is tokenized into a sequence of input tokens and converted into input embedding vectors. These input embedding vectors may be processed through a collection of i self-attention-based transformer layers extending from a first layer to a transition layer index (i), and a transitional output from the transition layer index (i) may be stored. The transitional output may be applied to a collection of (N−i) cross-attention-based transformer layers extending from the transition layer index (i+1) to the number of layers (N), and output tokens may be generated based on the final cross-attention based hidden state output from the final layer in the number of layers (N).
Owner:QUALCOMM INC

Methods of polypeptide design using combined masked language modeling and yeast surface display and sequences

The polypeptide design method and sequence of the present application combine mask language modeling and yeast surface display, comprising the following steps: cleaning a protein sequence database, selecting protein sequences meeting the requirements as a training set of a language model, and performing mask language modeling on the protein sequences contained in the training set; designing a set of downstream tasks on the basis of a pre-trained model, and fine-tuning the downstream tasks; randomly masking residues of a selected reference sequence, and predicting the masked residues; performing virtual screening on polypeptide candidates generated by the model; and determining the protein expression level and affinity of the screened polypeptides through yeast display technology. The present application can generate a large number of polypeptide candidates that may have specific properties or functions by transferring natural language processing technology to the polypeptide generation field, and combines artificial intelligence generation, virtual screening and wet experimental characterization to build a relatively complete "dry-wet combination" design process.
Owner:GUANGXI ZHONGMA PENCHENG PHARMACEUTICAL IND GROUP CO LTD

Aerospace system fault identification and modeling method based on model

The invention discloses a model-based aerospace system fault identification and modeling method, which comprises the following steps of: analyzing system requirements by using a Sysml language modeling tool, and determining input and output parameters according to system performance indexes to establish a system-level parameter model; determining a system structure composition model according to the performance index requirements and the parameter model; identifying an activity to be carried out by a system which meets the function performance requirement, and establishing a system function activity model; establishing a subsystem-level parameter model according to a system function activity model, and performing downward decomposition according to the ratio; establishing a fault element meta-model through a Sysml extension mechanism; and fault information is identified, fault analysis is carried out, a fault and fault propagation model is established, and forward design modeling containing the fault information is realized. According to the method, the fault information is recognized and the fault analysis model is established while the forward design of the system is carried out, the problem of collaborative modeling of the forward design and the reverse fault information is solved, and the system is helped to recognize risk optimization design in the design stage as a whole.
Owner:CHINA AEROSPACE STANDARDIZATION INST

Low complexity prefix processing in language modeling

Various embodiments include methods, and computing devices that perform the methods, of improving execution of a generative artificial intelligence model. Embodiment methods may include receiving an input prompt that is tokenized into a sequence of input tokens and converted into input embedding vectors. These input embedding vectors may be processed through a collection of self-attention-based transformer layers extending from a first layer to a transition layer index ( ), and a transitional output from the transition layer index ( ) may be stored. The transitional output may be applied to a collection of cross-attention-based transformer layers extending from the transition layer index ( +1) to the number of layers (N), and output tokens may be generated based on the final cross-attention based hidden state output from the final layer in the number of layers (N).
Owner:QUALCOMM INC

AI-based accurate advertisement putting method and system

The invention relates to the technical field of advertisement putting, in particular to an accurate advertisement putting method and system based on AI (artificial intelligence). A context semantic vector is generated by using a gating fusion method of two-channel language modeling and a visual semantic vector, and a semantic consistency correction item is introduced to solve the problem of dislocation of language description and picture content; a stability scoring mechanism and a rhythm suppression item based on a time window are introduced, so that the semantic matching precision is ensured, and meanwhile, a mistaken projection phenomenon caused by instantaneous fluctuation or context jump is suppressed; a context interference tolerance scoring method is put forward, advertisement display intensity is determined by combining anchor verbal skill density, picture jump degree and user question density, dynamic selection is carried out in three styles of bullet screen prompt, card display and anchor prompt through a style selection network, and structure matching regular terms of advertisement fields and context semantics are introduced. And the advertisement content is highly consistent with the display context.
Owner:GUANGZHOU HAIZHIYA MEDIA TECH CO LTD

GPU kernel characterization method

The invention discloses a GPU (Graphics Processing Unit) kernel characterization method. The method comprises the following steps: step 1, instruction standardization preprocessing: preprocessing an original GPU assembly instruction; 2, performing fine adjustment in a language model by utilizing the GPU assembly code corpus subjected to standardization preprocessing in the step 1, and retraining a special Tokenizer; step 3, construction of instruction semantic information; the method specifically comprises the steps of target instruction wrapping and context modeling; 4, performing fine tuning on the language model by using the shielding language modeling task, and generating a hidden state for the input target instruction by using the fine-tuned language model; and 5, averaging the embedded vectors of all the target instructions in each warp to obtain a warp embedded vector, and carrying out global fusion on all the warp embedded vectors to obtain a complete semantic vector representation of the GPU kernel. The GPU kernel characterization method supports multi-level modeling from an instruction level to a kernel level, and has good generalization ability and efficient reasoning performance.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Controllable text generation method and device

ActiveCN116881709BData setAlgorithm
The application provides a controllable text generation method and device, comprising: obtaining a to-be-processed task; inputting the to-be-processed task into a pre-constructed controllable text generation model to obtain a task reply; the controllable text generation model is obtained by training a pre-stored language model using a training sample data set; the training sample data set comprises a data set composed of comparison samples; the comparison samples at least include negative example reply samples, positive example reply samples and to-be-processed task samples; the positive example reply samples are reply samples without undesirable behaviors, and the negative example reply samples are reply samples with undesirable behaviors. The application does not need to modify the model architecture or introduce additional auxiliary models, simultaneously performs language modeling and comparison learning training using the training sample data set to obtain the controllable text generation model, the trained model is ready for use, and the application has higher universality, lower time complexity and more stable reply fluency.
Owner:TSINGHUA UNIVERSITY

Self-speculation decoding method and device based on dual advance quit, medium and equipment

The invention relates to the technical field of natural language processing, can be applied to business scenes such as financial analysis and medical report generation, and provides a self-speculation decoding method and device based on dual advanced exit, a medium and equipment. Connecting a lightweight adapter and reusing a language modeling head of the lightweight adapter to construct a self-draft model; for a to-be-decoded sequence, when the main model is calculated to the lth layer, executing hierarchy advance quit, obtaining a hidden state and inputting the hidden state into the self-drafting model; a candidate sequence is generated through iteration of the self-drafting model, and if the highest probability of semantic unit generation in a certain step is lower than a threshold value, semantic unit-level pre-exit iteration termination is executed; splicing the middle hidden state corresponding to each iteration step into a parallel tensor, and inputting the parallel tensor into a remaining layer of the main model for one-time parallel verification to obtain a verification result of each position; and determining a final decoding sequence according to the verification result and the candidate sequence. According to the embodiment, the deployment and calculation overhead is reduced, and the end-to-end reasoning efficiency is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

A visual language navigation method based on topological map pre-training

The application discloses a kind of visual language navigation methods based on topological map pretraining, belong to computer vision field, including the following steps: pretraining stage, baseline method utilizes mask language modeling task, mask region classification task and single-step action prediction task to pretrain the model of baseline network.Model of baseline method is pre-trained by three pre-training tasks respectively using MLM prediction head, MRC prediction head, SAP prediction head realization.Three new pre-training tasks are used to pretrain baseline network in pretraining stage;Topological map navigation order prediction task, topological map node orientation prediction task and topological map-instruction matching prediction task are realized by NOP prediction head, NSP prediction head and MIP prediction head respectively.The method shows the optimal performance in success rate, and the pre-training task proposed in the application can further improve the navigation performance of visual language navigation model.
Owner:BEIJING UNIV OF TECH

Method and system for semantic augmentation of a pre-trained sign language understanding framework based on bone

The application provides a method and system for enhancing a pre-training sign language understanding framework based on bone semantics, and relates to the technical field of sign language recognition, wherein the method comprises: obtaining sign language video data, paired bone sequences and text data thereof; extracting bone key point sequences from the sign language video data and modeling to form bone features, and performing word segmentation on the text; in a pre-training stage, inputting the bone features and the segmented text into a fusion network to generate bidirectional enhanced features, and obtaining global and local similarities through double-level semantic alignment; calculating a contrast loss based on the similarities, and obtaining a hierarchical loss by balancing parameters to coordinate weights; performing a matching task and a language modeling task to obtain corresponding losses, weighting and combining the three types of losses into a pre-training total loss, and adjusting parameters to complete training; and finally, in a fine-tuning stage, optimizing part of the parameters in combination with specific task types to realize enhanced understanding of sign language semantics. The application improves the accuracy of sign language understanding.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

A Chinese language understanding method and device that integrates the phonetic and semantic features of Chinese characters

PendingCN122311211ALanguage understandingLanguage modelling
This invention discloses a Chinese language understanding method and apparatus that integrates the phonetic and semantic features of Chinese characters. The method includes collecting and preprocessing Chinese text corpora, obtaining text sequences, converting the text sequences into radical embeddings, pinyin embeddings, and character shape embeddings, and integrating these to obtain integrated embeddings. A pre-trained understanding model is used to understand a Chinese language understanding task and predict the corresponding answer. The understanding model employs a three-level masked language modeling unit to obtain semantic information based on the text sequence and integrated embeddings, a phonetic loan character identification unit to identify phonetic loan characters, and a phonetic-semantic cross-prediction unit to predict the character shape and pinyin based on pinyin embeddings and character shape embeddings, respectively. This invention comprehensively captures the phonetic and semantic features of Chinese characters and integrates radical-level semantic information into the understanding model, improving the generalization ability of the understanding model and enabling it to understand language in complex scenarios.
Owner:WUHAN VOCATIONAL COLLEGE OF SOFTWARE & ENG (WUHAN OPEN UNIV)

Code characterization method based on operation code

The invention discloses a code characterization method based on an operation code. The code characterization method comprises the following steps: analyzing a Python source code into an operation code sequence; performing single-mode mask language modeling on the operation code sequence to obtain operation code semantic features; splicing the operation code sequence and the corresponding code annotations and then carrying out bimodal mask language modeling to obtain cross-modal semantic association features; executing multi-modal comparative learning based on the cross-modal semantic association features to obtain an optimized code representation vector; the code characterization vector is output for use by downstream tasks. According to the method, the bottom layer execution logic and semantic features of the Python code can be effectively mined. Meanwhile, by applying a mapping method from an operation code to a sequence and a multi-modal data collaboration mechanism, the dependence of the model on a complex analysis process is reduced, and the robustness of the model on low-quality or sparse code data is improved.
Owner:HUNAN UNIV OF SCI & TECH