Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41 results about "Word representation" patented technology

A representation term is a word, or a combination of words, that semantically represent the data type of a data element. A representation term is commonly referred to as a class word by those familiar with data dictionaries.

A method, device, and medium for processing NOTAM text based on semantic enhancement

This invention relates to the field of text processing technology, and in particular to a method, device, and medium for processing navigational notice text based on semantic enhancement. The method includes: first, acquiring a content carrier to be processed; then, acquiring the semantic vector and glyph feature vector of the content carrier; concatenating the two types of vectors to form an enhanced text representation; extracting temporal features from the enhanced text representation to obtain temporal features containing forward and backward logical relationships within the text; acquiring the weights of words and sentences in the temporal features and performing weighting to obtain weighted word representations and weighted sentence representations; performing correction processing on the weighted representations to generate corrected text; and finally, validating the corrected text and outputting the target text. This invention can improve the accuracy and efficiency of content carrier processing.
Owner:CIVIL AVIATION UNIV OF CHINA

Equipment health state prediction question-answering system and method for pre-training large language model

The invention discloses an equipment health state prediction question-answering system and method for pre-training a large language model. The system comprises an acquisition module, a marking module, an editing module, a prediction module and an interaction module. The acquisition module is used for acquiring time sequence data of industrial production equipment and preprocessing the time sequence data to obtain processed data; the marking module is used for dividing and marking the processed data to obtain a marked sequence; the editing module is used for reprogramming the mark sequence to obtain word representation of the time sequence data; the prediction module is used for collecting the health state of the industrial production equipment for prediction based on the word representation to obtain a prediction result; and the interaction module is used for interacting with the user to complete question and answer communication of the prediction result. According to the method, a large amount of structured and unstructured data can be processed, so that the key indexes and modes of the health state of the equipment are identified more accurately, the health state of the equipment is predicted, instant early warning information is provided, and an enterprise is helped to take measures before a fault occurs.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Navigation notice text processing method and device based on semantic enhancement and medium

The invention relates to the technical field of text processing, in particular to a navigation announcement text processing method and device based on semantic enhancement and a medium, and the method comprises the steps: firstly obtaining a to-be-processed content carrier; then obtaining a semantic vector and a font feature vector of a to-be-processed content carrier; splicing the two types of vectors to form enhanced text representation; performing time sequence feature extraction on the enhanced text representation to obtain a time sequence feature containing a forward and backward logic relationship of the text; obtaining weights of words and sentences in the time sequence features and completing weighting to obtain weighted word representation and weighted sentence representation; performing correction processing on the weighted representation to generate a corrected text; and finally, checking the corrected text, and outputting a target text. According to the invention, the processing accuracy and efficiency of the content carrier can be improved.
Owner:CIVIL AVIATION UNIV OF CHINA

Text representation method, word representation method, corresponding apparatus, medium and device

ActiveCN115392234Baccurate vector representationSolve the problem of inconsistent lengthsNatural language data processingSpecial data processing applicationsAlgorithmTheoretical computer science
The present disclosure relates to a text representation method, a word representation method, a corresponding device, a medium and equipment, the text representation method comprising: obtaining a text and performing word segmentation on the text to obtain a word sequence; generating a text vector of the text, and performing at least one round of iteration on the text vector, each round of iteration comprising multiple sub-iterations; in the i-th sub-iteration of each round of iteration, based on a sliding window with a length of k, after sliding the sliding window of the last sub-iteration on the word sequence backward, performing weighted combination on the real word vector of the k words in the current sliding window and the text vector optimized in the last sub-iteration, taking the vector after the weighted combination as the predicted word vector of the next word outside the current sliding window; optimizing the text vector according to the predicted word vector and the real word vector of the next word; and after completing the at least one round of iteration, obtaining a final text vector for representing the text. The present disclosure can accurately obtain the vector representation of the text.
Owner:NEUSOFT CORP

Multimodal large model image segmentation method and device based on hierarchical word element representation

The embodiment of the application provides a kind of multi-modal large model image segmentation method and device based on hierarchical word representation, mask image is encoded into word sequence by designing mask marker, and progressive generation from shape prototype to local detail is realized by causal attention mechanism.Three-stage training strategy is adopted, first, the marker is trained through the mask reconstruction task, then the mask word is integrated into the large multi-modal model and jointly trained, finally, it is fine-tuned using high-resolution data.Multi-level supervision is carried out based on hierarchical mask loss function, and accurate segmentation of natural language description target is realized.The method effectively solves the deficiencies of traditional technology in complex scene understanding, model training and mask generation, and significantly improves the performance of multi-modal image segmentation.
Owner:UNIVERSAL UBIQUITOUS TECH CO LTD

A multi-modal aspect-level sentiment analysis method, device, equipment and storage medium

ActiveCN117235261BSemantic analysisBiological modelsData packWord representation
The application discloses a kind of multi-modal aspect-level sentiment analysis method, device, equipment and storage medium.The application includes: obtaining multi-modal input data;The multi-modal input data includes input sentence and input image;The input image is input into pre-training conversion model, and the image caption of the input image is output;The context text representation of the input sentence and the context image caption description representation of the image caption are generated;Based on attention mechanism, the context text representation and the context image caption description representation are used to generate semantic information;The syntax mask matrix is constructed using the semantic information;Aspect word representation is obtained by performing graph convolution operation on the syntax mask matrix;The aspect word representation includes text representation and image representation;The text representation and image representation are interactively predicted to obtain the sentiment classification of the multi-modal input data.
Owner:SOUTH CHINA NORMAL UNIV

Emotion tetrad extraction method based on dialogue aspect

The invention discloses a dialogue aspect-based emotional tetrad extraction method, which comprises the following steps of: 1) extracting context-aware word representation for each utterance in a dialogue D; 2) carrying out overall modeling on the semantics of the whole dialogue, and refining the incidence relation between the utterances; 3) performing cross-utterance grammar analysis on a word level; 4) fusing multi-level associated information; and 5) based on the word level representation obtained in the step 4), processing all possible word pairs in the dialogue to obtain a complete emotion tetrad. According to the multi-level association optimization network framework, the association expression of the utterance level is remarkably enhanced through overall semantic modeling, meanwhile, by means of cross-utterance grammar analysis, the association modeling capacity of the word level is effectively improved, and the accuracy of overall tetrad extraction is improved.
Owner:HUAZHONG UNIV OF SCI & TECH

Method, device and electronic device for determining word representation vectors

Embodiments of the present application provide a method and device for determining word representation vectors, electronic equipment and computer readable storage medium, and belong to the field of natural language processing. The method comprises: obtaining a set of word units of a text; obtaining a context vector of the text based on the set of word units; and predicting a next word of the text based on the context vector. The method for determining word representation vectors can effectively obtain a corresponding set of word units even for pictographic characters or languages evolved from pictographic characters that are prone to out-of-vocabulary words, thereby improving the accuracy of determining word representation vectors.
Owner:BEIJING SAMSUNG TELECOM R&D CENT +1

Information retrieval method and device, electronic equipment and computer readable storage medium

Embodiments of the present application provide an information retrieval method and device, electronic equipment and computer readable storage medium, and relate to the field of artificial intelligence. The method comprises: dividing a question to obtain at least one question word, and dividing a target document containing an answer to obtain at least one document word; determining at least one document word representation corresponding to at least one document word based on at least one question word and at least one document word; dividing the target document into at least one short sentence, and determining at least one short sentence representation corresponding to at least one short sentence based on at least one document word representation; and determining at least one target short sentence from at least one short sentence as an answer corresponding to the question based on at least one short sentence representation. The present application solves the problems of incomplete semantics and missing extraction in the prior art when extracting long answers, thereby saving the time cost of user information retrieval and improving the user's retrieval experience.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Training method of cue word quality evaluation model and quality evaluation method

The invention discloses a training method of a cue word quality evaluation model and a quality evaluation method.The training method comprises the steps that a cue word training text is obtained, text sequence word segmentation is conducted on the cue word training text, and a semantic accuracy sequence, an expression integrity sequence and a logic continuity sequence of the cue word training text are obtained; according to the semantic accuracy sequence, the expression integrity sequence and the logic continuity sequence, performing multi-dimensional parallel coding fusion processing to obtain a target cue word representation vector; performing double-branch prediction fusion on the target cue word representation vector to obtain fused prediction data; and according to the fusion prediction data, performing parameter updating on the initialized cue word quality evaluation model to obtain a trained cue word quality evaluation model. According to the training method, a cue word quality evaluation model can be provided, and the model can be beneficial to improving the interpretability and accuracy of cue word quality evaluation. The invention relates to the technical field of natural language processing.
Owner:WUHAN UNIV OF TECH

An aspect-level sentiment analysis method based on multi-level knowledge enhancement

The application discloses a kind of aspect-level sentiment analysis method based on multilevel knowledge enhancement, belong to sentiment analysis field. Including the following steps: S1: text word vector is obtained using GloVe word embedding tool, the syntax dependency tree of sentence is constructed using Stanza, and the dependency graph is constructed accordingly;S2: the context representation of sentence is extracted by inputting word vector into BiLSTM, and the dependency graph is updated using sentiment dictionary and sensitive relationship set, to realize the sentiment and syntax enhancement of sentence;S3: the enhanced dependency graph is input into GCN to model node features, to obtain specific aspect representation;S4: aspect word is enhanced using concept atlas to obtain aspect word representation, which is fused with specific aspect representation to obtain aspect representation;S5: aspect representation and context representation are coordinated and optimized using interactive attention, to obtain the final representation of sentence, so as to determine aspect sentiment tendency.The application can effectively improve the accuracy of specific aspect sentiment classification, and help merchants accurately locate problems in products or services.
Owner:ANHUI UNIV OF SCI & TECH

Large model knowledge graph question answering method and device based on potential unit fine-tuning

The present application relates to a kind of large model atlas question answering method and device based on potential unit fine-tuning, the method includes the following steps: based on the knowledge graph of symbolization and relationship path and the variable of using high-dimensional prompt word representation potential unit, potential relationship learning is carried out between large language model and knowledge graph, and relationship reasoning path is generated, wherein the input layer of large language model is additionally trainable potential unit;And potential unit is randomly initialized by normal distribution, and is fine-tuned with large language model in hidden learning process;Relevance evaluation is carried out to relationship reasoning path, and the most relevant relationship reasoning path is selected;Input natural language question, and the answer in knowledge graph is obtained by the most relevant relationship reasoning path.The present application filters out irrelevant and misleading context fragments using these fine-tuned potential units, while improving response efficiency and performance.
Owner:ZHEJIANG UNIV

Limb representation method and device and electronic equipment

The invention discloses a lexical element representation method and device and electronic equipment. The method comprises the following steps: S1, performing word segmentation on an input text by adopting a first word segmentation device to obtain a first lexical element sequence arranged according to the sequence of the input text; s2, performing word segmentation on the input text by adopting a second word segmentation device to obtain a second lexical element sequence arranged according to the sequence of the input text; extracting lexical units, different from the first lexical unit sequence, in the second lexical unit sequence to form a combined phrase set; s3, converting the lexical elements corresponding to the first lexical element sequence into a first embedded sequence, converting the lexical elements corresponding to the phrases in the combined phrase set into embedded vectors, and fusing the embedded vectors into a second embedded sequence; and replacing the corresponding elements in the first embedding sequence with the elements in the second embedding sequence to obtain a compressed embedding sequence. On the premise that an original model embedding matrix can be reused, high-frequency cross-word boundary phrases are dynamically fused to shorten the sequence length, and therefore the calculation efficiency and the expression capacity of a subsequent model are improved.
Owner:HANGZHOU RAPID INTELLIGENT TECHNOLOGY CO LTD

Named entity recognition system and method based on parallel LSTM (Long Short Term Memory) and gating mechanism

The invention discloses a named entity recognition system and method based on parallel LSTM (Long Short Term Memory) and a gating mechanism, and relates to the technical field of computer data processing, and the system comprises a character modeling module which is used for generating character-level word representation from an input character sequence; the statement modeling module is used for splicing the character-level word representation and a word-level pre-training representation searched from an external pre-training word vector library according to a word index to obtain a word representation matrix, and obtaining a statement-level word representation based on the word representation matrix; and the sequence decoding module is used for performing sequence-level labeling decoding on the statement-level word representation and outputting an entity label sequence. The method can reduce redundant calculation and stabilize reasoning.
Owner:JIANGSU OPEN UNIVERSITY (THE CITY VOCATIONAL COLLEGE OF JIANGSU)

A patent technology prediction system and method fusing a time sequence knowledge graph and contrast learning

PendingCN122332624AGraph spectraEngineering
The application discloses a patent technology prediction system and method fusing a time sequence knowledge graph and contrast learning, relates to the patent technology prediction field, and is proposed in view of the problem of inaccurate technology prediction in the prior art. Patent elements are extracted and standardized; quadruples are constructed and heterogeneous relationships are defined; time slicing is performed and a graph snapshot sequence is generated; structural flow and heterogeneous graph structures are encoded respectively; time flow is encoded; patent semantics are represented and a unified space is aligned; patent representation and real associated technology keyword representation are explicitly semantically aligned based on a contrast learning enhancement mechanism; coarse retrieval results are obtained; fine retrieval results are obtained and rearranged; coarse ranking scores and fine correction scores are fused and output; multi-level ranking targets are jointly designed; a rearranger and a Gold Injection strategy are trained; and an overall loss function and end-to-end training are performed. The application has the advantages of being capable of depicting the dynamic evolution process of patent technology association changing over time, improving the overall quality of prediction results, and the like.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Question correction method and device, electronic equipment and storage medium

This invention provides a question error correction method, apparatus, electronic device, and storage medium. It employs a question error correction model, which includes an encoding layer and a decoding layer. The encoding layer comprises a first-type convolutional neural network layer and the encoding end of a Transformer layer. The decoding layer comprises a second-type convolutional neural network layer and the decoding end of a Transformer layer. The question text samples and their corresponding corrected question text samples used during training can be selected based on the required accuracy of the question error correction model, thus addressing the problem of excessively large parameter space in N-gram models when N is too large. Furthermore, since the question error correction model utilizes word vectors, which are a distributed feature representation, they take into account the semantics of words, effectively mitigating the high sparsity of traditional statistical word representations.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

A text classification method based on label semantic learning and attention adjustment mechanism

The application discloses a text classification method based on label semantic learning and attention adjustment mechanism, and mainly comprises the following steps: preprocessing text data, extracting text semantic features, text label graph embedding, using a multi-head adjustment attention mechanism to measure the semantic relationship between words and labels, then multi semantic integration and network training, thereby realizing multi-label text classification, training the model, and then using the trained model to predict the category of a text. The application proposes a multi-head adjustment attention hybrid BERT model for a multi-label text classification framework, which can effectively extract useful features from text content, establish semantic connection between labels and words, obtain label-specific word representation, and thus improve the performance of multi-label text classification.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Systems and methods for seeded neural topic modeling

ActiveUS12639512B2Natural language translationWord representationData mining
A method may include: receiving a seed topic word distribution; receiving a corpus of documents; generating bag of words representations for the corpus of documents; converting the corpus of documents to vector representations; training a topic modeling system using the seed topic word distribution and concatenated bag of words representations and the vector representations resulting in a topic word distribution and a document word distribution; generating a plurality of new generated topics based on the topic word distribution; precomputing a topic word distribution penalty and a topic word distribution reward for the plurality of topics; penalizing the topic modeling system in response to a divergence and rewarding the topic modeling system in response to a similarity; determining a total loss from a neural network loss, the topic word distribution penalty, and the topic word distribution reward; and training the topic modeling system based on the total loss.
Owner:JPMORGAN CHASE BANK NA

IMAGE RECONFIGURATION IN VIDEO CODING USING RATE-DISTIRSURES OPTIMIZATION.

Given a sequence of images in a first keyword representation, methods, processes, and systems are presented for image reconfiguration using distortion-rate optimization, where reconfiguration allows the images to be encoded in a second keyword representation that enables more efficient compression than using the first keyword representation. Syntax methods for signaling reconfiguration parameters are also presented.
Owner:DOLBY LABORATORIES LICENSING CORP

Financial product question and answer method, device and equipment, medium and program product

The invention provides a financial product question and answer method which can be applied to the technical field of artificial intelligence. The financial product question and answer method comprises the following steps: inputting an obtained question text into a predicate learning model based on a neural network to obtain a question predicate representation; inputting the question text into an entity learning model based on a neural network to obtain question entity representation; determining a target relationship, a target head entity and a target tail entity corresponding to the predicate representation and the question entity representation in a relationship representation, a head entity representation and a tail entity representation in a knowledge graph learning model; and converting the target relationship, the target head entity and the target tail entity into texts, and outputting the texts as answer texts. The invention further provides a financial product question and answer device and equipment, a storage medium and a program product.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Event argument extraction method and device

The application provides an event argument extraction method and device. The method comprises the following steps: encoding training data and event types respectively to obtain a trigger word context semantic representation and an event type representation, and then interacting the two representations to obtain a trigger word representation containing event type information and predicting an event type; generating an argument extraction question corresponding to the event type, and then splicing and encoding the argument extraction question and a text to be extracted to obtain a label context semantic representation, a context semantic representation of each word in the sentence to be extracted, and a context semantic representation of an argument role; splicing the context semantic representation of the label, the context semantic representation of each word in the sentence to be extracted, and the context semantic representation of the argument role to be extracted respectively, and then inputting the spliced result into a discrimination network to obtain a discrimination probability and a labeling probability respectively; and determining an extraction result corresponding to the argument role by combining the discrimination probability and the labeling probability. The method improves the event extraction performance.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A medical electronic medical record named entity recognition method based on word vector fusion

The application discloses a medical electronic medical record named entity recognition method based on word vector fusion, and analyzes the text content of a Chinese electronic medical record, studies the influence of different forms of word vectors on named entity recognition, and proposes a text named entity recognition method based on word vector fusion, which combines a static word vector representation model and a dynamic word vector pre-training model with context semantic information, is used for solving the static word representation of medical field words and different entities with multiple semantic information in the electronic medical record text, realizes targeted extraction of medical entities, and improves the performance of the electronic medical record text named entity recognition model.
Owner:CHINA THREE GORGES UNIV

Methods and systems for predicting an upcoming data point associated with an entity

PendingUS20260087253A1Semantic analysisPayment architectureWord representationLinguistic rule
Methods and systems for predicting an upcoming data point are disclosed. Method performed by a server system includes accessing a plurality of encoded features associated with each data point of a plurality of data points corresponding to an entity. Method includes generating a data point word representation for the entity based on the plurality of encoded features associated with each data point. The data point word representation includes one or more words. Method includes generating a data point sentence representation based on the one or more words associated with each data point. Method includes generating, by a Large Language Model (LLM), an upcoming data point representation in the data point sentence representation based on the data point sentence representation. Method includes decoding the upcoming data point representation to obtain one or more features associated with an upcoming data point based on the set of predefined language rules.
Owner:MASTERCARD INT INC

Dedicated field Chinese word segmentation method and system based on transfer learning

ActiveCN121279308ASemantic analysisBiological modelsWord identificationSemantic lexicon
The invention relates to a transfer learning-based special field Chinese word segmentation method and system, and aims to solve the problems that culture load words in international education field texts are difficult to recognize, the adaptability of a general model field is poor, and semantic consistency optimization is insufficient. Enhancing cultural load word representation in combination with an attention mechanism; eliminating domain differences by using adversarial training and feature distribution alignment technologies, and generating domain invariant features; fusing sequence labeling and boundary detection tasks to optimize word segmentation boundaries; and finally, dynamically adjusting the segmentation path based on the semantic similarity score. According to the method, the word segmentation accuracy of the composite terms in the education field is remarkably improved, the problems of wrong segmentation of culture load words, reduction of field migration performance and lack of semantic coherence are solved, and high-robustness word segmentation support is provided for an international education text processing system.
Owner:YANCHENG INST OF IND TECH

High-resolution remote sensing sample labeling method based on topic model

ActiveCN116363460BPattern recognitionWord representation
The application provides a high-resolution remote sensing sample labeling method based on a topic model, comprising the following steps: obtaining a training sample set and a sample set to be labeled; extracting traditional features and deep features of the training sample set; performing feature quantization on the traditional features and the deep features of the training sample set to obtain a bag-of-words representation of the training sample set; constructing a visual topic model, inputting the training sample set represented by the bag-of-words into the visual topic model to obtain a topic distribution of the training sample set; constructing and training a weak classifier model by using the topic distribution of the training sample set and labeling information, wherein the weak classifier model comprises at least two weak classifiers; and labeling the sample set to be labeled by using the weak classifier model. The method uses a small amount of labeled training sample set to assist in iteratively training multiple weak classifiers, the accuracy of the weak classifiers is improved, the generalization ability of the finally obtained weak classifier model is strong, and the weak classifier model can accurately classify remote sensing samples and remote sensing images.
Owner:BEIJING DATA INTELLIGENCE INFORMATION TECH CO LTD

Method, apparatus, device and storage medium for determining antigen specificity

ActiveCN117012281BBiostatisticsSequence analysisWord representationImmune receptor
The embodiment of the application provides a kind of antigen specificity determination method, device, equipment and storage medium, at least be applied to artificial intelligence field, medical field and adaptive immune receptor, wherein, method includes: double-chain biological information of cell receptor is carried out word encoding processing, obtains amino acid word sequence;Wherein, at least one amino acid word in the amino acid word sequence is indicated;Using pre-trained amino acid sequence prediction model, based on the amino acid word sequence, the cell receptor is extracted, and the amino acid sequence representation of the cell receptor is obtained;Wherein, the amino acid sequence prediction model is obtained by training the data obtained by shielding processing part of sample amino acid word representation in sample data;Based on the amino acid sequence representation, the antigen specificity of the cell receptor is determined;Through the application, the cell receptor can be accurately extracted, so that the antigen specificity of the cell receptor is accurately determined.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method, apparatus, and computer program product for large-scale language model

The invention relates to a method, an apparatus and a computer program product for a large-scale language model. The use of the cloud-assisted vehicle-mounted large-scale language model is performed by inputting a pre-edit cue including personal information relating to an occupant of the vehicle to a first LLM (large-scale language model) mounted on the vehicle, transmitting a post-edit cue to a second LLM mounted on the cloud, and displaying the post-edit cue to the second LLM. The post-editing cue represents a pre-editing cue without personal information, receiving an original score of a second LLM, inputting the post-editing cue to a third LLM installed on a vehicle or a cloud, the third LLM including the same parameter set as the first LLM, determining a difference between the original score of the first LLM and the original score of the third LLM, and outputting the difference between the original score of the first LLM and the original score of the third LLM. The sum of the difference and the original score of the second LLM is determined to obtain an adjusted original score, and an output is obtained from the adjusted original score.
Owner:TOYOTA JIDOSHA KK

Semantic feature extraction model training method and device, equipment and storage medium

The application discloses a training method and device of a semantic feature extraction model, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: obtaining training corpus of the semantic feature extraction model, wherein the training corpus comprises word text corpus of a target language and pronunciation annotation information thereof; obtaining word representation vector sequences of the word text corpus and pronunciation representation vector sequences of the pronunciation annotation information; extracting fused semantic features from the word representation vector sequences and the pronunciation representation vector sequences through the semantic feature extraction model; determining prediction results corresponding to a pre-training task of the semantic feature extraction model based on the fused semantic features; determining pre-training loss of the semantic feature extraction model based on the prediction results and real results, and adjusting parameters of the semantic feature extraction model according to the pre-training loss to obtain a pre-training completed semantic feature extraction model. The application can improve the semantic representation capability of the semantic feature extraction model.
Owner:TENCENT TECH (BEIJING) CO LTD

Malware detection method based on semantic analysis and bidirectional encoding representation

ActiveCN116432184BAlgorithmTerm memory
The present application aims at the problem of ambiguous word representation and lack of context semantics in traditional model detection of malicious code, and proposes a malware detection method based on semantic analysis and bidirectional encoding representation, which combines BERT with convolution recurrent network based on external attention mechanism, uses malware API function call sequence as the feature of model learning, and performs static analysis to detect existing malware; the API call function sequence has correlation in context and semantics, BERT is used for word representation task, and semantic information is received from the sequence; the convolutional neural network and the long short-term memory network are respectively used for completing secondary feature extraction and API function chain relationship mining; and the attention mechanism is added after the long short-term memory network, so that the key information in the text can be better focused, the influence of noise is reduced, and the accuracy in the text classification task is improved; the present application is not affected by the change and deformation of malicious code itself, and the accuracy reaches 98.81%.
Owner:SHENYANG LIGONG UNIV