Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

32 results about "Lexical set" patented technology

A lexical set is a group of words that all fall under a single category based on some shared phonological feature.

Natural language text data intelligent classification method and system based on deep learning

The invention provides a natural language text data intelligent classification method and system based on deep learning, and relates to the technical field of natural language processing, and the method comprises the steps: 1, employing a context awareness mechanism to analyze the real semantics of a target vocabulary according to an antagonistic variant existing in a text, and obtaining a target vocabulary; in combination with a word meaning library and a pre-training process of a dynamic learning rate adjustment strategy, generating a candidate replacement vocabulary set with consistent semantics; and step 2, based on the candidate replacement vocabulary set, performing multi-dimensional semantic similarity calculation and emotional tendency discrimination, determining applicable vocabularies conforming to an original culture background through a context adaptation strategy, and generating a standardized text sequence. According to the method, through multi-dimensional semantic analysis, cultural context fusion, cross-granularity feature construction and dynamic parameter correction, the accuracy and adaptability of natural language text classification are realized.
Owner:厦门知链科技有限公司

English teaching method based on somatosensory interaction

The invention discloses an English teaching method based on somatosensory interaction, and relates to the field of interactive teaching, and the method comprises the steps: collecting limb behaviors and voice data of a user, and generating a teaching task instruction based on a teaching theme selected by the user; matching actions and voices of a user with a target grammar structure and a vocabulary set by constructing a mapping library of limb behaviors and English communication intentions; analyzing the limb movement data through a space-time diagram convolutional network, recognizing movement features, and generating a voice intention identifier through voice recognition and semantic analysis; comparing the consistency of the limb behavior and the voice intention, and if a conflict exists, generating a correction instruction; and dynamically adjusting the task difficulty according to the historical performance of the user, and increasing visual and voice interference until all teaching objectives are completed. The method has the advantages that by collecting the limb movement and voice data of the user in real time and combining the dynamic task and the intelligent error correction system, the language learning experience with high interactivity is achieved, and the learning effect and the user participation degree are effectively improved.
Owner:陈秀兰

Interaction method and device, electronic equipment, storage medium and computer program product

The invention discloses an interaction method and device, electronic equipment, a storage medium and a computer program product, and belongs to the technical field of human-computer interaction. The method comprises the following steps: receiving query input of a user; in response to the query input, determining a first query vocabulary, and performing data enhancement on the first query vocabulary to obtain a target query vocabulary set; obtaining a specific score corresponding to the target query vocabulary set; based on the specific score corresponding to the target query vocabulary set, performing sparse retrieval and dense retrieval on each target query vocabulary in the target query vocabulary set to obtain a target retrieval result corresponding to the target query vocabulary set; and inputting the target retrieval result and the first query vocabulary into a pre-trained language model to obtain a target response result output by the language model, and outputting the target response result to the user. According to the method, the hallusion probability in the query process can be reduced, and the coverage range and the accuracy of the retrieval result are improved.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

A new engineering major Chinese knowledge concept extraction method based on a dictionary

ActiveCN116127954BSolve the shortcomings of being unable to fully utilize word informationreduce lossesBiological modelsNatural language data processingConcept recognitionEngineering
The application discloses a new engineering major Chinese knowledge concept extraction method based on a dictionary, and comprises the following steps: 1) obtaining a new engineering related subdivided major, and converting all course teaching materials and syllabuses into text data; 2) using relevant text data to obtain corresponding words through word segmentation processing, and training a word2vec word vector model and an original word vector on the basis; 3) obtaining a large number of relevant course major vocabulary sets through a crawler technology, selecting corresponding major keywords as seeds, inputting the seeds into the trained word2vec model, obtaining words with a similarity above a threshold, and jointly forming a new engineering knowledge concept dictionary with the segmented words; and 4) constructing an NECE model, recognizing knowledge concepts of original course materials, and storing a concept set. The application can use the word2vec model to construct corresponding course major vocabulary sets and word vectors, and use the NECE model to realize extraction of major course concepts, thereby laying a data foundation for construction of an education system knowledge graph.
Owner:YANGZHOU UNIV

Remote sensing open vocabulary object detection method based on multi-modal large language model

This application provides a remote sensing open vocabulary target detection method based on a multimodal large language model, belonging to the field of remote sensing image detection technology. The method includes: acquiring remote sensing images of a user-selected target area; performing remote sensing target recognition to obtain a target vocabulary set and a target vocabulary matching degree set; configuring the environmental remote sensing image division range to divide and acquire environmental remote sensing images within the environment of the target area, obtaining an environmental vocabulary set and an environmental vocabulary matching degree set; configuring language model recognition resources, randomly combining the environmental vocabulary set and the target vocabulary set, inputting them into the configured remote sensing large language model, and outputting a target open vocabulary set and a frequency set; calculating the open vocabulary confidence score, selecting the optimal open vocabulary as the target detection result for the target area. This solves the technical problem that existing remote sensing target recognition methods can only detect predefined categories and cannot address the recognition needs of non-predefined ground features.
Owner:SHAANXI TIRAIN TECH CO LTD

Domain-specific retrieval language model

The invention relates to a domain-specific retrieval language model. Various examples, systems, and methods are disclosed that relate to domain-specific document retrieval that combines custom vocabulary integration and embedded model update. A computing system may extract a plurality of segments from a collection of documents and generate a query corresponding to at least one segment. The computing system may identify terms that meet uniqueness criteria, and input these terms into a markup machine to create a vocabulary dataset. The vocabulary datasets, document segments, and queries can be used to update the embedding model to support retrieval and semantic alignment within private documents.
Owner:NVIDIA CORP

Sign language recognition system based on cross-modal graph alignment

The invention provides a sign language recognition system based on cross-modal diagram alignment, and relates to the technical field of intelligent sign language recognition. Respectively inputting the medical emergency sign language video vocabulary set into the medical sign language vocabulary feature extraction module and the medical sign language visual feature extraction module for training; and training the medical sign language visual feature extraction module by using the medical sign language vocabulary modal diagram corresponding to each medical first aid sign language video and the medical sign language video modal diagram corresponding to each medical first aid sign language video. According to the method, the medical sign language vocabulary features and the visual features are combined, cross-modal graph alignment is achieved, key information expressed by the sign language can be rapidly and accurately recognized, the generalization ability of the model is improved, and the requirements for the key information extraction speed and accuracy in first aid are met.
Owner:天津市急救中心

Multi-modal large model driven robot cluster collaborative awareness and decision-making method

The invention discloses a multi-modal large model driven robot cluster collaborative awareness and decision-making method, and the method comprises the steps: enabling a robot in a robot cluster to achieve the collection of multi-modal original data through a carried multi-type sensor, carrying out the denoising, alignment and standardization preprocessing of the original data of each modal, and obtaining the multi-modal preprocessing data of a uniform format, the invention relates to the technical field of artificial intelligence. According to the robot cluster collaborative awareness and decision-making method driven by the multi-modal large model, through keyword matching, context analysis and logic conjunction analysis, the problems that in the prior art, natural language instructions are fuzzy in understanding, and key information is prone to being missed are solved, associated word collection is constructed based on historical instructions, and the algorithm is simple and convenient to operate. Four task types of environment exploration, target search, collaborative operation and emergency response can be dynamically matched according to natural language keywords, and task allocation is ensured to be adaptive to robot functions, so that efficient interaction of robot cluster collaborative awareness is realized.
Owner:广州新华学院

A text abstract generation method and device, an electronic device, and a storage medium

The application discloses a text abstract generation method and device, an electronic device and a storage medium. The method comprises the following steps: obtaining a plurality of target texts to be processed; performing word segmentation on each target text to be processed based on a latest word segmentation dictionary to obtain a vocabulary set corresponding to each target text to be processed; processing the vocabulary set corresponding to each target text to be processed by using a latest word vector model to obtain a word vector corresponding to each target text to be processed; calculating the similarity between each two target texts to be processed based on the word vector corresponding to each target text to be processed; aggregating the target texts to be processed based on the similarity between the target texts to be processed to obtain a plurality of text categories; extracting corresponding keywords from the word segmentation set corresponding to each text category; and determining the text abstract corresponding to each text category from each target text belonging to the text category based on the corresponding keywords of the text category.
Owner:CHINA CONSTRUCTION BANK

Hotword extraction method and device for speech recognition, storage medium and product

One or more embodiments of the specification provide a hot word extraction method, device, storage medium and product for speech recognition. The method comprises: obtaining a first data set, the first data set containing documents related to a target field; performing speech synthesis based on the original text of the documents contained in the first data set to generate simulated speech data; and performing recognition processing on the simulated speech data to obtain recognized text corresponding to the documents; comparing the original text of the documents with the corresponding recognized text, and counting the error words contained in the original text and the corresponding error frequencies to obtain an error word set; selecting a first candidate hot word set from the error word set; wherein the error frequency of the words selected into the first candidate hot word set is higher than the error frequency of the words not selected; and extracting a hot word related to the target field from the first candidate hot word set, the hot word reflecting the high-frequency words prone to errors of the speech recognition system in the target field.
Owner:HANGZHOU ANT KUAI TECHNOLOGY CO LTD

Text2SQL semantic caching method based on context and mode matching

The invention provides a Text2SQL semantic caching method based on context and pattern matching, and relates to the field of natural language processing and database interaction.The method comprises the steps that user query is input into a semantic compression module, semantic significant keywords are extracted to obtain a candidate set, a vocabulary set is extracted and constructed through a large language model, and a database is obtained; obtaining a compressed query through a syntactic organizer; inputting the historical dialogue into a context encoder, and obtaining global context representation through encoding and two-stage attention; performing coarse-grained filtering on the compressed query to form a newest candidate set; performing fine-grained context matching on the global context representation and the newest candidate set to obtain context similarity; and comparing the context similarity with a preset threshold value, judging whether the cache is hit or not, if the cache is hit, interacting the hit SQL statement with a database, and returning a query result. According to the method and the device, the problems of low SQL multiplexing accuracy and high response delay are solved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Text feature-based file sensitive content detection method and device, equipment and medium

The invention discloses a text feature-based file sensitive content detection method and device, equipment and a medium, and relates to the technical field of information security. A to-be-detected file is preprocessed to obtain text blocks; the preprocessing comprises content extraction, character cleaning and text partitioning; segmenting the text block to obtain a vocabulary set, calculating a hash value of each vocabulary in the vocabulary set, mapping the hash values to different index positions of an initial feature vector to obtain a feature vector, and performing normalization and feature quantization processing on the feature vector to obtain a target feature vector; splicing the target feature vectors to obtain a text feature set; calculating text similarity between the text feature set and the file feature sample set; and carrying out file sensitive content detection on the to-be-detected file by utilizing the text similarity. By means of the technical scheme, the problems of safety, performance and precision in cross-network and cross-device file sensitive content detection can be solved, and the accuracy of file sensitive content detection is improved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Method and system for enhancing online speech recognition based on online OCR (Optical Character Recognition)

The invention provides a method and system for enhancing online speech recognition based on online OCR, and relates to the technical field of speech recognition, and the method comprises the steps: capturing a shared digital image in a collaborative session online in real time, and extracting a text in the shared digital image; performing word segmentation processing on the extracted text to obtain a first vocabulary set; training and generating a session exclusive language model corresponding to the collaborative session based on the first vocabulary set; fusing the session exclusive language model and a pre-created universal language model to obtain a corresponding fused language model; performing conversion processing on the fused language model to obtain a corresponding language model decoding graph; in the decoding stage, a language model decoding graph and a pre-created acoustic vocabulary decoding graph are loaded, the two loaded decoding graphs are dynamically decoded in real time, a text recognition result of the voice in the collaborative session is obtained, and the acoustic vocabulary decoding graph is created in advance based on a universal acoustic model and a universal language model. The objective of the invention is to improve speech recognition precision and efficiency.
Owner:CHINA TELECOM CLOUD TECH CO LTD

Song list generation method and device, computer equipment, storage medium and program product

The invention relates to a song list generation method and device, computer equipment, a storage medium and a program product. The method comprises the steps of obtaining a candidate song set of a target language; obtaining a vocabulary set of a target language; obtaining the number of uncovered words corresponding to each candidate song in the candidate song set; the uncovered words are vocabularies which are included in the vocabulary set and are not included in songs of the target song menu; and based on the number of the uncovered words corresponding to each candidate song in the candidate song set, selecting one or more candidate songs in the candidate song set and adding the selected candidate songs to the target song menu.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Software demand case normativity inspection method, tool and computer program product

The invention provides a normative checking method and tool for software requirement cases and a computer program product. The checking method comprises the following steps: extracting attributes from a software demand case, and converting the attributes into feature vectors by adopting one-hot coding according to a demand case related vocabulary set and a preset weight of each attribute; vector dimensions are equal to the size of the vocabulary set, vector positions are in one-to-one correspondence with words in the vocabulary set, and word frequencies are weighted and counted according to attribute weights and then are filled into the corresponding vector positions; on the basis of the feature vectors, classifying the software demand cases by utilizing a classification model; and matching the corresponding writing specification according to the category of the software demand case, and carrying out natural language structured analysis on the description of the software demand case by adopting an improved shift-protocol dependency analysis method to judge whether the description conforms to the writing specification or not. According to the method, the problem of low manual inspection efficiency is solved, the inspection efficiency is greatly improved through multi-dimensional technical optimization, and standardization and intelligentization of software demand review are promoted.
Owner:NO 15 INST OF CHINA ELECTRONICS TECH GRP

A text review topic sentiment analysis method, device and equipment

The application provides a text review subject sentiment analysis method, device and equipment, which comprises the following steps: preprocessing obtained review data about a museum to obtain a vocabulary set; performing importance sorting on the vocabulary in the vocabulary set based on a TF-IDF algorithm; screening first and second emotional seed words based on the sorting result; selecting target words with a required similarity to the emotional seed words from the vocabulary set based on a SO-PPMI algorithm; forming an emotional dictionary based on the target words and a general emotional dictionary; determining the weight of each word based on the similarity between each word and each emotional seed word recorded in the emotional dictionary; constructing a WLDA model based on the weight and an LDA model, and determining the subject words corresponding to each subject based on the model; and determining the emotional tendency of each subject based on the subject, the subject words and the emotional dictionary. The method can quickly and accurately perform sentiment analysis on museum reviews.
Owner:NAT UNIV OF DEFENSE TECH

Demand similarity determination method and apparatus, electronic device, and computer storage medium

The application discloses a demand similarity determination method and device, electronic equipment and a computer storage medium. The industry characteristic data of a first demand and the industry characteristic data of a second demand are compared to determine a vocabulary set. The target matching result of the first demand and the second demand is determined according to the vocabulary set, wherein the target matching result is used to indicate the similarity of the first demand and the second demand. Therefore, the similarity between the new demand and the plurality of preset demands in the database can be identified, and the repeated development, function conflict and the like are avoided, and the development efficiency is improved.
Owner:XIANGYANG BRANCH CHINA MOBILE GRP HUBEI CO LTD +1

Target language vocabulary substitution load control and pacing management system

PendingCN122655752AEngineeringLexical set
The application discloses a target language vocabulary replacement load control and learning rhythm management system, comprising: a familiar word state acquisition module, which is used for acquiring a set of mastered target language vocabularies of a user; a learning content analysis module, which is used for identifying replaceable semantic units or replaceable vocabulary positions in to-be-output learning content; a replacement capacity determination module, which is used for determining a target vocabulary replacement capacity in current learning content; a familiar word priority mapping module, which is used for matching the replaceable semantic units or replaceable vocabulary positions with the set of target language vocabularies, and preferentially determining familiar word replacement positions within the target vocabulary replacement capacity range; a new word injection window allocation module, which is used for determining a new word injection window in current learning content based on remaining replacement positions; a learning rhythm adjustment module, which is used for dynamically adjusting subsequent learning content; and a learning content output module, which is used for outputting learning content according to the familiar word replacement positions and the new word injection window.
Owner:CHUANGZHI YUNWEI (BEIJING) TECHNOLOGY CO LTD

Chinese text difficulty classification method and system and storage medium

The invention discloses a Chinese text difficulty classification method and system and a storage medium. The method comprises the following steps: performing text word segmentation processing on an obtained Chinese text to be processed; obtaining the Embedding expression of each Chinese vocabulary in the Chinese vocabulary set, and calculating linguistic indexes; the Embedding vectors and linguistic indexes are combined, and a feature matrix is constructed and obtained; processing the feature matrix by using a convolutional neural network, and extracting local features; processing the feature matrix by using a pre-trained Transform model, and extracting global features; and fusing the local features and the global features to form final feature representation of the text, and inputting the final feature representation into a classifier for text difficulty classification to obtain a Chinese text difficulty classification result. The embodiment of the invention can be suitable for various scenes needing to evaluate the text readability, is high in classification efficiency and high in classification accuracy, and can be widely applied to the technical field of computers.
Owner:SOUTH CHINA NORMAL UNIV

Test corpus generation method, related method, device, equipment and storage medium

The invention discloses a test corpus generation method, a related method, a related device, equipment and a storage medium, and the test corpus generation method comprises the following steps: determining whether to select a candidate vocabulary as an extended vocabulary of a seed vocabulary based on a semantic distance between the candidate vocabulary in a text corpus and the seed vocabulary with a to-be-tested function; based on the seed vocabularies and the extended vocabularies thereof, obtaining a vocabulary set of the function to be Obtaining a weight factor of each vocabulary based on the association closeness between each vocabulary in the vocabulary set and the to-be-tested function; based on the weight factor of each vocabulary, respectively weighting the semantic similarity between each vocabulary and the candidate test corpus of the function to be tested to obtain a semantic score of the candidate test corpus; and based on the semantic score of the candidate test corpus, determining whether to select the candidate test corpus as a target test corpus of the to-be-tested function. According to the scheme, the generation efficiency and the quality stability of the test corpus can be improved.
Owner:IFLYTEK CO LTD

Retrieval method and apparatus, electronic device, computer readable storage medium

The present disclosure provides a retrieval method and device, an electronic device and a computer readable storage medium. The method comprises: obtaining a first feature vocabulary set of a target object to be retrieved, wherein the first feature vocabulary in the first feature vocabulary set is used to describe an object attribute possessed by the target object; generating a plurality of first images of a first modality according to the first feature vocabulary set; and retrieving a target record set from a preset database according to the first feature vocabulary set and the plurality of first images. The target record set comprises at least one target object record, and the target object record is an object record in the preset database that satisfies a first preset condition. The object record that satisfies the first preset condition comprises a second image and / or a second feature vocabulary set contained in the object record, and at least one of the second image and / or the second feature vocabulary set matches at least one of the first feature vocabulary set and the plurality of first images. According to the embodiments of the present disclosure, the accuracy of the target record set obtained by retrieval can be improved.
Owner:MASHANG CONSUMER FINANCE CO LTD

Intelligent classification method and system for natural language text data based on deep learning

The application provides a natural language text data intelligent classification method and system based on deep learning, and relates to the technical field of natural language processing.The method comprises the following steps: in step 1, the real semantics of a target word is analyzed by using a context perception mechanism in view of an adversarial variant existing in the text, a pre-training process is combined with a word meaning library and a dynamic learning rate adjustment strategy, and a set of candidate replacement words with consistent semantics is generated; in step 2, multi-dimensional semantic similarity calculation and sentiment tendency discrimination are carried out based on the set of candidate replacement words, a suitable word conforming to the original cultural background is determined through a context adaptation strategy, and a standardized text sequence is generated. Through multi-dimensional semantic analysis, cultural context fusion, cross-granularity feature construction and dynamic parameter correction, the accuracy and adaptability of natural language text classification are realized.
Owner:厦门知链科技有限公司

Method and system for quickly removing duplicate of document based on multi-algorithm fusion

PendingCN121859883AQuick exclusionOvercome semantic blind spotsText processingEnergy efficient computingComputation complexityBloom filter
The invention discloses a rapid document duplicate removal method and system based on multi-algorithm fusion, and the method comprises the steps: carrying out the preprocessing of a document, and generating a duplicate removal vocabulary set; performing rapid member judgment on the de-duplicated vocabulary set through a Bloom filter; generating a locality sensitive hash fingerprint and a probability signature in parallel for the document passing through the Bloom filter; constructing a zoning locality sensitive hash index, and retrieving a candidate similar document set from a document library; calculating the similarity of the title dimension, the content dimension and the metadata dimension for each candidate similar document, and performing weighted fusion on the similarity of each dimension according to a preset weight to obtain a comprehensive similarity score corresponding to the document; and comparing the comprehensive similarity score with a set similarity threshold value to judge whether the document is repeated or not according to a comparison result. Through a collaborative mechanism of coarse screening, indexing and similarity fusion decision, the calculation complexity is remarkably reduced, and meanwhile, the accuracy of large-scale document deduplication is improved.
Owner:CAPINFO CO LTD

Multi-code large model security reinforcement method based on single training

The invention relates to a multi-code large model security reinforcement method based on single training, which comprises the following steps: selecting a code large model and a target model which are respectively configured with a corresponding lexical set and a vocabulary, and then selecting a shared lexical set C; training on the code large model by adopting a LoRA fine tuning method to obtain a trained safe small model; the Mtrg and the Msrc are subjected to characterization alignment, and M's rc is obtained; the vocabulary of the Mtrg and the vocabulary of the M's src are aligned, and a lexical element mapping table of mapping from the Mtrg to the M's src is obtained; defining a code task and a loop termination condition, and calculating fusion probability distribution pensession; and obtaining a final output code conforming to X through loop calculation and selection according to the pensession. Through the method, the security of code generation of any target model can be enhanced through the single-trained security small model.
Owner:CHONGQING UNIV

A data storage method, system, storage medium and electronic device

Embodiments of the present application provide a data storage method, system, storage medium and electronic device. The method comprises: parsing a first to-be-stored file to obtain first data text information; based on a preset vocabulary set of each preset data type, filtering out the vocabulary under the preset data type from the first data text information to obtain a first vocabulary set; calculating the frequency of the vocabulary in the first vocabulary set appearing in the first data text information to obtain a first word frequency, determining the correlation degree of the first data text information and the preset data type based on the correlation degree coefficient of the first word frequency and each preset word frequency; comparing the correlation degrees of the first data text information and each preset data type, determining the preset data type corresponding to the maximum correlation degree as the data type of the first data text information, and storing the first to-be-stored file based on the data type of the first data text information. The present application can improve the perfection and integration of data.
Owner:SUPCON TECH CO LTD

A text2sql semantic caching method based on context and pattern matching

The application provides a Text2SQL semantic caching method based on context and pattern matching, relates to the field of natural language processing and database interaction, and comprises the following steps: inputting a user query into a semantic compression module, extracting semantic significant keywords to obtain a candidate set, extracting and constructing a vocabulary set through a large language model, and obtaining a compressed query through a syntax organizer; inputting historical dialogues into a context encoder, obtaining a global context representation through coding and two-stage attention; performing coarse-grained filtering on the compressed query to form a latest candidate set; performing fine-grained context matching on the global context representation and the latest candidate set to obtain a context similarity; comparing the context similarity with a preset threshold to determine whether the cache hits, and if the cache hits, interacting with a database with the hit SQL statement and returning a query result. The application solves the problems of low SQL reuse accuracy and high response delay.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Text correction based topic modeling enhancement method and apparatus

The application discloses a text correction-based topic modeling enhancement method and device, which comprises the following steps: extracting time distribution features of words, measuring the similarity between words according to the time distribution features, calculating the time expression ability index of the words, and obtaining a time special vocabulary set; extracting spatial distribution features of the words, measuring the spatial distribution difference of the words according to the spatial distribution features, generating a word distance matrix, performing vocabulary clustering based on the word distance matrix, and obtaining a spatial special vocabulary set; fusing point of interest data, extracting semantic distribution features of the words according to the point of interest data, calculating the semantic expression ability index of the words, and obtaining a semantic special vocabulary set; and modifying the text content based on the time special vocabulary set, the spatial special vocabulary set and the semantic special vocabulary set, to obtain an enhanced topic model; the application can extract spatiotemporal semantic special vocabularies and calculate expression ability indexes, and further correct the text to enhance topic modeling.
Owner:ZHUOYU INTELLIGENT TECH CO LTD

Attention increment-based long context generation model inference acceleration method and system

A method and system for accelerating inference in a long-context generative model are disclosed. In the pre-filling inference stage of the generative model, an attention increment metric is calculated for the input lexical at a preset stop-calculation start layer, and an active lexical set is determined according to a preset retention ratio. The active lexical set is reused in the subsequent calculation of the network block, that is, only the lexical subsequences corresponding to the active lexical set are used to perform network block calculation, further calculation of inactive lexical is stopped, and its hidden representation is kept as the last update result, thereby reducing the overall computational load of the pre-filling stage.
Owner:SHANGHAI JIAOTONG UNIV

A vocabulary processing method, apparatus, device, and storage medium

ActiveCN115730594BNatural language data processingLexical setWord target
The present disclosure relates to a vocabulary processing method, device, equipment and storage medium. The vocabulary processing method comprises: determining a target vocabulary, wherein the target vocabulary comprises at least one vocabulary; determining, based on corpus data and the target vocabulary, vocabulary statistics data of a first vocabulary used as a first word set and vocabulary statistics data of the first vocabulary; determining, according to the vocabulary statistics data and the vocabulary statistics data, expected information of the first vocabulary used as the first word set; and determining, according to the expected information, a first characteristic of the first vocabulary, wherein the first characteristic is at least a word set with the largest expectation of the first vocabulary. Through the method of the present disclosure, the highest frequency word set of the vocabulary in the vocabulary set is provided, so that the user can master the actual usage of the vocabulary at a specific learning stage, thereby improving the learning efficiency of the user.
Owner:CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1