Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

47 results about "Lexical set" patented technology

A lexical set is a group of words that all fall under a single category based on some shared phonological feature.

Natural language text data intelligent classification method and system based on deep learning

The invention provides a natural language text data intelligent classification method and system based on deep learning, and relates to the technical field of natural language processing, and the method comprises the steps: 1, employing a context awareness mechanism to analyze the real semantics of a target vocabulary according to an antagonistic variant existing in a text, and obtaining a target vocabulary; in combination with a word meaning library and a pre-training process of a dynamic learning rate adjustment strategy, generating a candidate replacement vocabulary set with consistent semantics; and step 2, based on the candidate replacement vocabulary set, performing multi-dimensional semantic similarity calculation and emotional tendency discrimination, determining applicable vocabularies conforming to an original culture background through a context adaptation strategy, and generating a standardized text sequence. According to the method, through multi-dimensional semantic analysis, cultural context fusion, cross-granularity feature construction and dynamic parameter correction, the accuracy and adaptability of natural language text classification are realized.
Owner:厦门知链科技有限公司

Hot word extraction method and device for speech recognition, storage medium and product

One or more embodiments of the invention provide a hot word extraction method and device for speech recognition, a storage medium and a product. The method comprises the steps of obtaining a first data set, wherein the first data set comprises documents related to a target domain; performing voice synthesis on the basis of an original text of a document contained in the first data set to generate simulated voice data; and carrying out identification processing on the simulation voice data to obtain an identification text corresponding to the document, comparing the original text of the document with the corresponding recognition text, and counting the vocabularies which are contained in the original text and are wrongly recognized and the corresponding error frequency to obtain an error vocabulary set; screening out a first candidate hot word set from the error vocabulary set; wherein the error frequency of the vocabularies selected into the first candidate hot word set is higher than that of the vocabularies which are not selected into the first candidate hot word set; hot words related to the target domain are extracted from the first candidate hot word set, and the hot words reflect error-prone high-frequency vocabularies of the speech recognition system in the target domain.
Owner:HANGZHOU ANT KUAI TECHNOLOGY CO LTD

English teaching method based on somatosensory interaction

The invention discloses an English teaching method based on somatosensory interaction, and relates to the field of interactive teaching, and the method comprises the steps: collecting limb behaviors and voice data of a user, and generating a teaching task instruction based on a teaching theme selected by the user; matching actions and voices of a user with a target grammar structure and a vocabulary set by constructing a mapping library of limb behaviors and English communication intentions; analyzing the limb movement data through a space-time diagram convolutional network, recognizing movement features, and generating a voice intention identifier through voice recognition and semantic analysis; comparing the consistency of the limb behavior and the voice intention, and if a conflict exists, generating a correction instruction; and dynamically adjusting the task difficulty according to the historical performance of the user, and increasing visual and voice interference until all teaching objectives are completed. The method has the advantages that by collecting the limb movement and voice data of the user in real time and combining the dynamic task and the intelligent error correction system, the language learning experience with high interactivity is achieved, and the learning effect and the user participation degree are effectively improved.
Owner:陈秀兰

International trade contract intelligent auditing system based on natural language processing

The invention relates to the technical field of contract intelligent auditing, in particular to an international trade contract intelligent auditing system based on natural language processing, which comprises a data acquisition preprocessing module, a key vocabulary extraction module, a key vocabulary analysis screening module and a text similarity association module, the data acquisition preprocessing module carries out word segmentation and segmentation on a contract text and a regulation standard text, then the key word extraction module extracts a key word set in each text segment through a word frequency statistical method, and in order to prevent the situation that the key word set extracted by the key word extraction module is inaccurate, the key word set is extracted by the key word extraction module. The key vocabulary analyzing and screening module analyzes the paraphrases of the key vocabularies in the corresponding words, removes the key vocabularies which are not nouns in the key vocabulary set, provides a paraphrasing basis for the vocabulary same paraphrasing mapping unit, unifies same paraphrasing word vectors, and transmits the same paraphrasing word vectors to the key vocabulary same paraphrasing mapping unit; and the contract and the regulation text are deeply associated in terms of vocabulary semantics.
Owner:NANJING LUNAIZUN TECHNOLOGY CO LTD

Interaction method and device, electronic equipment, storage medium and computer program product

The invention discloses an interaction method and device, electronic equipment, a storage medium and a computer program product, and belongs to the technical field of human-computer interaction. The method comprises the following steps: receiving query input of a user; in response to the query input, determining a first query vocabulary, and performing data enhancement on the first query vocabulary to obtain a target query vocabulary set; obtaining a specific score corresponding to the target query vocabulary set; based on the specific score corresponding to the target query vocabulary set, performing sparse retrieval and dense retrieval on each target query vocabulary in the target query vocabulary set to obtain a target retrieval result corresponding to the target query vocabulary set; and inputting the target retrieval result and the first query vocabulary into a pre-trained language model to obtain a target response result output by the language model, and outputting the target response result to the user. According to the method, the hallusion probability in the query process can be reduced, and the coverage range and the accuracy of the retrieval result are improved.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

A new engineering major Chinese knowledge concept extraction method based on a dictionary

ActiveCN116127954BSolve the shortcomings of being unable to fully utilize word informationreduce lossesBiological modelsNatural language data processingConcept recognitionEngineering
The application discloses a new engineering major Chinese knowledge concept extraction method based on a dictionary, and comprises the following steps: 1) obtaining a new engineering related subdivided major, and converting all course teaching materials and syllabuses into text data; 2) using relevant text data to obtain corresponding words through word segmentation processing, and training a word2vec word vector model and an original word vector on the basis; 3) obtaining a large number of relevant course major vocabulary sets through a crawler technology, selecting corresponding major keywords as seeds, inputting the seeds into the trained word2vec model, obtaining words with a similarity above a threshold, and jointly forming a new engineering knowledge concept dictionary with the segmented words; and 4) constructing an NECE model, recognizing knowledge concepts of original course materials, and storing a concept set. The application can use the word2vec model to construct corresponding course major vocabulary sets and word vectors, and use the NECE model to realize extraction of major course concepts, thereby laying a data foundation for construction of an education system knowledge graph.
Owner:YANGZHOU UNIV

Remote sensing open vocabulary object detection method based on multi-modal large language model

This application provides a remote sensing open vocabulary target detection method based on a multimodal large language model, belonging to the field of remote sensing image detection technology. The method includes: acquiring remote sensing images of a user-selected target area; performing remote sensing target recognition to obtain a target vocabulary set and a target vocabulary matching degree set; configuring the environmental remote sensing image division range to divide and acquire environmental remote sensing images within the environment of the target area, obtaining an environmental vocabulary set and an environmental vocabulary matching degree set; configuring language model recognition resources, randomly combining the environmental vocabulary set and the target vocabulary set, inputting them into the configured remote sensing large language model, and outputting a target open vocabulary set and a frequency set; calculating the open vocabulary confidence score, selecting the optimal open vocabulary as the target detection result for the target area. This solves the technical problem that existing remote sensing target recognition methods can only detect predefined categories and cannot address the recognition needs of non-predefined ground features.
Owner:SHAANXI TIRAIN TECH CO LTD

Domain-specific retrieval language model

The invention relates to a domain-specific retrieval language model. Various examples, systems, and methods are disclosed that relate to domain-specific document retrieval that combines custom vocabulary integration and embedded model update. A computing system may extract a plurality of segments from a collection of documents and generate a query corresponding to at least one segment. The computing system may identify terms that meet uniqueness criteria, and input these terms into a markup machine to create a vocabulary dataset. The vocabulary datasets, document segments, and queries can be used to update the embedding model to support retrieval and semantic alignment within private documents.
Owner:NVIDIA CORP

Sign language recognition system based on cross-modal graph alignment

The invention provides a sign language recognition system based on cross-modal diagram alignment, and relates to the technical field of intelligent sign language recognition. Respectively inputting the medical emergency sign language video vocabulary set into the medical sign language vocabulary feature extraction module and the medical sign language visual feature extraction module for training; and training the medical sign language visual feature extraction module by using the medical sign language vocabulary modal diagram corresponding to each medical first aid sign language video and the medical sign language video modal diagram corresponding to each medical first aid sign language video. According to the method, the medical sign language vocabulary features and the visual features are combined, cross-modal graph alignment is achieved, key information expressed by the sign language can be rapidly and accurately recognized, the generalization ability of the model is improved, and the requirements for the key information extraction speed and accuracy in first aid are met.
Owner:天津市急救中心

Multi-modal large model driven robot cluster collaborative awareness and decision-making method

The invention discloses a multi-modal large model driven robot cluster collaborative awareness and decision-making method, and the method comprises the steps: enabling a robot in a robot cluster to achieve the collection of multi-modal original data through a carried multi-type sensor, carrying out the denoising, alignment and standardization preprocessing of the original data of each modal, and obtaining the multi-modal preprocessing data of a uniform format, the invention relates to the technical field of artificial intelligence. According to the robot cluster collaborative awareness and decision-making method driven by the multi-modal large model, through keyword matching, context analysis and logic conjunction analysis, the problems that in the prior art, natural language instructions are fuzzy in understanding, and key information is prone to being missed are solved, associated word collection is constructed based on historical instructions, and the algorithm is simple and convenient to operate. Four task types of environment exploration, target search, collaborative operation and emergency response can be dynamically matched according to natural language keywords, and task allocation is ensured to be adaptive to robot functions, so that efficient interaction of robot cluster collaborative awareness is realized.
Owner:广州新华学院

A text abstract generation method and device, an electronic device, and a storage medium

The application discloses a text abstract generation method and device, an electronic device and a storage medium. The method comprises the following steps: obtaining a plurality of target texts to be processed; performing word segmentation on each target text to be processed based on a latest word segmentation dictionary to obtain a vocabulary set corresponding to each target text to be processed; processing the vocabulary set corresponding to each target text to be processed by using a latest word vector model to obtain a word vector corresponding to each target text to be processed; calculating the similarity between each two target texts to be processed based on the word vector corresponding to each target text to be processed; aggregating the target texts to be processed based on the similarity between the target texts to be processed to obtain a plurality of text categories; extracting corresponding keywords from the word segmentation set corresponding to each text category; and determining the text abstract corresponding to each text category from each target text belonging to the text category based on the corresponding keywords of the text category.
Owner:CHINA CONSTRUCTION BANK

Robot Manipulation Learning Method and Device Through Prediction of Interaction

The present application relates to the field of computer vision technology, and discloses a robot manipulation learning method and device through predicted interaction. In terms of the encoder, the key frames of the initial state and the termination state in the interaction process are used as the input of the visual coding module to obtain a visual embedding sequence, and the language description in the interaction process is segmented according to the positions in the predefined vocabulary set to obtain a language embedding sequence. In the decoder, a predictor is used to predict unseen transition frames, representing the interaction between the initial state and the termination state, enabling the model to understand "how to interact"; a detector is used to infer the positions of the interacting objects, enabling the model to know "where to interact". Among them, information exchange between the transition frames and the positions of the interacting objects is achieved through bidirectional attention, promoting mutual support in the training process, which can improve the performance of downstream robot applications. And information irrelevant to the interaction is filtered out, emphasizing the key states, thereby realizing a concentrated learning process.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Data query method and device based on natural language, equipment and medium

The invention relates to the technical field related to data query, in particular to a data query method and device based on a natural language, equipment and a medium. The method comprises the steps of obtaining text data corresponding to a natural language; performing word segmentation processing on the text data to obtain a vocabulary set; extracting time vocabularies in the vocabulary set; performing vectorization processing on the vocabulary set based on a knowledge graph pre-established for an actual working scene to obtain a vocabulary vector set; determining a preset number of vocabularies with the highest weight in the text data based on the vocabulary vector set; and distilling the result with the highest weight step by step. On the basis of a preset number of vocabularies with the highest weight and the time vocabularies, generating prompt words annotated with query conditions; inputting the cue word and the text data into a preset large language model to obtain a query statement; and querying in a preset database based on the query statement, and obtaining a query result.
Owner:SUZHOU BLUE WING INTELLIGENT DIGITAL TECHNOLOGY CO LTD

Hotword extraction method and device for speech recognition, storage medium and product

One or more embodiments of the specification provide a hot word extraction method, device, storage medium and product for speech recognition. The method comprises: obtaining a first data set, the first data set containing documents related to a target field; performing speech synthesis based on the original text of the documents contained in the first data set to generate simulated speech data; and performing recognition processing on the simulated speech data to obtain recognized text corresponding to the documents; comparing the original text of the documents with the corresponding recognized text, and counting the error words contained in the original text and the corresponding error frequencies to obtain an error word set; selecting a first candidate hot word set from the error word set; wherein the error frequency of the words selected into the first candidate hot word set is higher than the error frequency of the words not selected; and extracting a hot word related to the target field from the first candidate hot word set, the hot word reflecting the high-frequency words prone to errors of the speech recognition system in the target field.
Owner:HANGZHOU ANT KUAI TECHNOLOGY CO LTD

Word weight ranking method, device and equipment and storage medium

The application relates to the technical field of artificial intelligence, and discloses a word weight sorting method, device and equipment and a storage medium, the sorting method comprising the following steps: obtaining a first target sentence, performing word segmentation on the first target sentence to obtain a first vocabulary set contained in the first target sentence; removing each first vocabulary in the first vocabulary set from the first target sentence to obtain a second target sentence; forming a sentence pair by combining any second target sentence and the first target sentence, inputting the sentence pair into a pre-trained language model to generate a vector representation corresponding to the sentence pair; determining a target similarity between the second target sentence and the first target sentence in the sentence pair according to the vector representation corresponding to the sentence pair; sorting the target similarities of the second target sentences and the first target sentence; and determining a weight order of each first vocabulary according to the sorting of the target similarities. The application solves the problem that the word weight cannot be accurately sorted in the prior art.
Owner:PING AN TECH (SHENZHEN) CO LTD

Text2SQL semantic caching method based on context and mode matching

The invention provides a Text2SQL semantic caching method based on context and pattern matching, and relates to the field of natural language processing and database interaction.The method comprises the steps that user query is input into a semantic compression module, semantic significant keywords are extracted to obtain a candidate set, a vocabulary set is extracted and constructed through a large language model, and a database is obtained; obtaining a compressed query through a syntactic organizer; inputting the historical dialogue into a context encoder, and obtaining global context representation through encoding and two-stage attention; performing coarse-grained filtering on the compressed query to form a newest candidate set; performing fine-grained context matching on the global context representation and the newest candidate set to obtain context similarity; and comparing the context similarity with a preset threshold value, judging whether the cache is hit or not, if the cache is hit, interacting the hit SQL statement with a database, and returning a query result. According to the method and the device, the problems of low SQL multiplexing accuracy and high response delay are solved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Text feature-based file sensitive content detection method and device, equipment and medium

The invention discloses a text feature-based file sensitive content detection method and device, equipment and a medium, and relates to the technical field of information security. A to-be-detected file is preprocessed to obtain text blocks; the preprocessing comprises content extraction, character cleaning and text partitioning; segmenting the text block to obtain a vocabulary set, calculating a hash value of each vocabulary in the vocabulary set, mapping the hash values to different index positions of an initial feature vector to obtain a feature vector, and performing normalization and feature quantization processing on the feature vector to obtain a target feature vector; splicing the target feature vectors to obtain a text feature set; calculating text similarity between the text feature set and the file feature sample set; and carrying out file sensitive content detection on the to-be-detected file by utilizing the text similarity. By means of the technical scheme, the problems of safety, performance and precision in cross-network and cross-device file sensitive content detection can be solved, and the accuracy of file sensitive content detection is improved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Method and system for enhancing online speech recognition based on online OCR (Optical Character Recognition)

The invention provides a method and system for enhancing online speech recognition based on online OCR, and relates to the technical field of speech recognition, and the method comprises the steps: capturing a shared digital image in a collaborative session online in real time, and extracting a text in the shared digital image; performing word segmentation processing on the extracted text to obtain a first vocabulary set; training and generating a session exclusive language model corresponding to the collaborative session based on the first vocabulary set; fusing the session exclusive language model and a pre-created universal language model to obtain a corresponding fused language model; performing conversion processing on the fused language model to obtain a corresponding language model decoding graph; in the decoding stage, a language model decoding graph and a pre-created acoustic vocabulary decoding graph are loaded, the two loaded decoding graphs are dynamically decoded in real time, a text recognition result of the voice in the collaborative session is obtained, and the acoustic vocabulary decoding graph is created in advance based on a universal acoustic model and a universal language model. The objective of the invention is to improve speech recognition precision and efficiency.
Owner:CHINA TELECOM CLOUD TECH CO LTD

A data processing method, apparatus, storage medium, and computer device

An embodiment of the present application discloses a data processing method, which includes obtaining a vocabulary set corresponding to a source sample and a label sample; obtaining a target word and its synonyms, and calculating a similarity score between the target word and its synonyms; converting the words and synonyms in the vocabulary set into a word vector set, and performing vector mixing on the target word and the corresponding synonyms to obtain a mixed word vector; replacing the word vector of the corresponding target word with the mixed word vector and inputting it into a preset model for training; generating a mixed label, obtaining the difference between the word probability distribution of the mixed label and the word prediction probability distribution of the mixed label output by the preset model, and iteratively training the model parameters of the preset model according to the difference to obtain a trained preset model. In this way, the efficiency of data processing is improved, and the diversity of the output of the trained model is increased.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Entity recognition method and device, electronic equipment, storage medium and program product

The invention provides an entity recognition method and device, electronic equipment, a storage medium and a program product, and belongs to the technical field of artificial intelligence. Inputting the to-be-recognized statement into an entity recognition model to obtain an entity recognition result output by the entity recognition model; the entity recognition model determines association feature representation of characters and vocabularies based on vector representation of each character in the to-be-recognized statement and a vocabulary set associated with each character, and fuses the association feature representation and dependency information representation of the to-be-recognized statement to obtain vector features; and performing entity recognition based on the vector features to obtain an entity recognition result. Vector features used for entity recognition simultaneously comprise associated feature representations of characters and vocabularies and dependency information representations of statements to be recognized, boundary features and feature information of the characters are enriched, the problem that entity boundaries are not clear can be solved, and the multi-scene entity recognition capability of an entity recognition model is effectively improved.
Owner:CHINA MOBILE GROUP ZHEJIANG +2

Song list generation method and device, computer equipment, storage medium and program product

The invention relates to a song list generation method and device, computer equipment, a storage medium and a program product. The method comprises the steps of obtaining a candidate song set of a target language; obtaining a vocabulary set of a target language; obtaining the number of uncovered words corresponding to each candidate song in the candidate song set; the uncovered words are vocabularies which are included in the vocabulary set and are not included in songs of the target song menu; and based on the number of the uncovered words corresponding to each candidate song in the candidate song set, selecting one or more candidate songs in the candidate song set and adding the selected candidate songs to the target song menu.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

Software demand case normativity inspection method, tool and computer program product

The invention provides a normative checking method and tool for software requirement cases and a computer program product. The checking method comprises the following steps: extracting attributes from a software demand case, and converting the attributes into feature vectors by adopting one-hot coding according to a demand case related vocabulary set and a preset weight of each attribute; vector dimensions are equal to the size of the vocabulary set, vector positions are in one-to-one correspondence with words in the vocabulary set, and word frequencies are weighted and counted according to attribute weights and then are filled into the corresponding vector positions; on the basis of the feature vectors, classifying the software demand cases by utilizing a classification model; and matching the corresponding writing specification according to the category of the software demand case, and carrying out natural language structured analysis on the description of the software demand case by adopting an improved shift-protocol dependency analysis method to judge whether the description conforms to the writing specification or not. According to the method, the problem of low manual inspection efficiency is solved, the inspection efficiency is greatly improved through multi-dimensional technical optimization, and standardization and intelligentization of software demand review are promoted.
Owner:NO 15 INST OF CHINA ELECTRONICS TECH GRP

A text review topic sentiment analysis method, device and equipment

The application provides a text review subject sentiment analysis method, device and equipment, which comprises the following steps: preprocessing obtained review data about a museum to obtain a vocabulary set; performing importance sorting on the vocabulary in the vocabulary set based on a TF-IDF algorithm; screening first and second emotional seed words based on the sorting result; selecting target words with a required similarity to the emotional seed words from the vocabulary set based on a SO-PPMI algorithm; forming an emotional dictionary based on the target words and a general emotional dictionary; determining the weight of each word based on the similarity between each word and each emotional seed word recorded in the emotional dictionary; constructing a WLDA model based on the weight and an LDA model, and determining the subject words corresponding to each subject based on the model; and determining the emotional tendency of each subject based on the subject, the subject words and the emotional dictionary. The method can quickly and accurately perform sentiment analysis on museum reviews.
Owner:NAT UNIV OF DEFENSE TECH

Demand similarity determination method and apparatus, electronic device, and computer storage medium

The application discloses a demand similarity determination method and device, electronic equipment and a computer storage medium. The industry characteristic data of a first demand and the industry characteristic data of a second demand are compared to determine a vocabulary set. The target matching result of the first demand and the second demand is determined according to the vocabulary set, wherein the target matching result is used to indicate the similarity of the first demand and the second demand. Therefore, the similarity between the new demand and the plurality of preset demands in the database can be identified, and the repeated development, function conflict and the like are avoided, and the development efficiency is improved.
Owner:XIANGYANG BRANCH CHINA MOBILE GRP HUBEI CO LTD +1

Self-learning method, device, equipment and medium for intent classification model

The present application discloses a self-learning method, apparatus, medium and equipment for an intent classification model. The method includes: inputting a target text and a text label into the intent classification model, wherein the target text includes one or more target terms and the text label includes a classification result of the target text; processing the target terms and text labels using the intent classification model to obtain a first a posteriori probability that the target terms belong to each target intent category, wherein the target intent categories include global intent categories and precise intent categories; if the first a posteriori probability is greater than a preset probability threshold, determining that the global intent category included in the target intent category is the target category, and adding the target term to the vocabulary set corresponding to the target category; determining the union and intersection of multiple vocabulary sets, and updating the global intent category vocabulary based on the difference between the union and the intersection. The method of the present application realizes self-learning of the text intent classification model, reduces manual operations, improves efficiency and reduces costs.
Owner:PING AN INT FINANCIAL LEASING CO LTD

A text watermark embedding and detection method based on model context learning

A text watermark embedding and detection method based on model context learning includes the following steps: S1. Segmenting a given text and performing part-of-speech analysis to form a vocabulary set; S2. Selecting suitable replacement words from the vocabulary set; S3. Selecting synonyms similar to the selected words from a synonym library to construct a new synonym library; S4. Creating a series of synonymous rewriting examples as corpus to form a text watermark synonym rewriting corpus; using the corpus to train a language model for contextual learning guided by thought chains; then using the language model to generate rewritten texts, selecting the rewritten text that is most semantically similar to the given text based on overall sentence semantic similarity and contextual word semantic similarity; S5. Encoding the watermark information into a binary sequence, selecting synonyms to replace the synonyms at corresponding positions in the rewritten text, and thus embedding the binary sequence into the rewritten text. This method maintains semantic consistency and concealment while improving watermark capacity and robustness.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Speech enabling system

PendingUS20250273214A1Speech recognitionTeaching apparatusDisplay deviceIncomprehensible sounds
A system and method for recognizing (reading) the tongue movements, vocalizations, and throat vibrations of a person and converting (translating) them into meaningful synthesized words, which could be pronounced by an electronic speaker and / or displayed on a display. Often a patient / person who has lost the ability to speak may still be able to move their tongues, or make unfathomable sounds, which cannot be recognized as intelligible words. The system and method can continuously record the movement of the patient's tongue, vocalizations, and throat sounds and extract small video segments corresponding to different words attempted by the patient. Each of these video segments can then be analyzed by AI software or other configured software to match the specific tongue movement with a pre-learned reference word, and once identified, the computer / system can speak or verbalize the word, and / or display it on a screen. The set of pre-learned reference words can be recorded and saved during the training session(s) for each individual patient. During the training session(s), the embedded AI software or other configured software can ask the patient to say a specific word multiple times and for each time, the system / software can record the associated tongue movements, and any verbalizations, and throat vibrations. Multiple recordings for a single word can be preferably performed so that the AI software or other configured software can capture all possible movement variations of the same word and can aggregate some common features as unique identification for that word, which can be saved in the system as a reference to that specific word.
Owner:WEISS JUSTIN BENJAMIN +1

Target language vocabulary substitution load control and pacing management system

PendingCN122655752AEngineeringLexical set
The application discloses a target language vocabulary replacement load control and learning rhythm management system, comprising: a familiar word state acquisition module, which is used for acquiring a set of mastered target language vocabularies of a user; a learning content analysis module, which is used for identifying replaceable semantic units or replaceable vocabulary positions in to-be-output learning content; a replacement capacity determination module, which is used for determining a target vocabulary replacement capacity in current learning content; a familiar word priority mapping module, which is used for matching the replaceable semantic units or replaceable vocabulary positions with the set of target language vocabularies, and preferentially determining familiar word replacement positions within the target vocabulary replacement capacity range; a new word injection window allocation module, which is used for determining a new word injection window in current learning content based on remaining replacement positions; a learning rhythm adjustment module, which is used for dynamically adjusting subsequent learning content; and a learning content output module, which is used for outputting learning content according to the familiar word replacement positions and the new word injection window.
Owner:CHUANGZHI YUNWEI (BEIJING) TECHNOLOGY CO LTD

Text-based document classification method and document classification device

The present disclosure relates to a text-based document classification method and a document classification device. A text-based document classification method according to an embodiment of the present disclosure is performed by a processor inside a computing device, and may include: extracting, from a document image that has been input, words included in the document image; generating, based on a degree of similarity between the words, a word set including a configured number of words; generating a word set image by individually turning the word set into an image; extracting an important keyword used for document classification among words included in the word set image; and classifying a type of the document image by using the important keyword.
Owner:SAMSUNG SDS CO LTD