Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

75 results about "Chinese word" patented technology

Method and device for intelligent semantic error correction and business term optimization of foreign trade letter electricity

The invention relates to the technical field of natural language processing, in particular to a foreign trade letter intelligent semantic error correction and business term optimization method and device, and the method comprises the steps: obtaining a target foreign trade letter, and constructing a target corpus; performing Chinese word segmentation and part-of-speech tagging on the target foreign trade letter; performing term optimization based on the word segmentation result and the knowledge graph; identifying the letter title by using a conditional random field model, and converting the letter title into structured data; a Bi-LSTM-CRF model is adopted to carry out risk point detection, including Bi-LSTM coding, feature engineering and Max-pooling technologies, a part-of-speech sequence is obtained through a softmax function and a Viterbi path, and sequence labeling is carried out to obtain a risk point detection result; and finally, performing Chinese error correction based on the word segmentation result after part-of-speech tagging and the knowledge graph. The recognition and correction accuracy of foreign trade terminologies is improved, and communication obstacles caused by nonstandard use of the terminologies are effectively reduced.
Owner:GUANGDONG VOCATIONAL COLLEGE OF SCI & TRADE

Intelligent questioning and answering method and system in material field based on vector database

The invention discloses a material field intelligent question and answer method and system based on a vector database, and the method comprises the steps: recognizing and extracting a question input by a user through a language large model to obtain a question keyword, determining Chinese vocabularies in the question keyword, determining whether the Chinese vocabularies have multi-language synonyms, and if yes, determining that the Chinese vocabularies have multi-language synonyms; replacing the Chinese vocabulary with the multi-language synonym to obtain a multi-language question; retrieving in a vector database based on the multi-language question to obtain a first literature list; performing cross matching on the multi-language question and the first literature list through a cross-language Reranker model, and sorting the first literature list to obtain a second literature list; and forming a question reply based on the literature of the second literature list through the language large model, aiming at the special requirements in the field of material science, the method optimizes the question and literature processing flow, can quickly and accurately retrieve the literature required by the user, and generates the question reply conforming to the habit of the user.
Owner:TIANMUSHAN LABORATORY

Storage format for chinese language and related processing method and apparatus

A storage format of Chinese language (“Readable Hanyu Expression” or “RHE”) and the related processing methods and systems. Unlike current Chinese processing methods which directly code Chinese characters into fonts for display, RHE takes an indirect approach by storing Chinese language in the RHE storage format that can be mapped to several display forms including simplified and traditional Chinese characters, Hanyu Pinyin, etc. In the RHE storage format, each Chinese word is stored as an RHE storage element having the format (Syllable+Tone)n+Mark, where n is the number of syllables (Chinese characters) in the word, Syllable represents the pronunciation (without the tone) of the character, Tone represents the tone of the pronunciation, and Mark is a value that differentiates different words having the same pronunciations and tones. Various mapping tables are used to map RHE storage elements to standard Chinese character codes (such as Unicode) and Pinyin expressions.
Owner:YANG MINGWEI

Rule-based enhanced chinese word segmentation and semantic unit parsing method and system

PendingCN122311199AEngineeringChinese word
This application relates to the field of data processing technology and discloses a method and system for Chinese word segmentation and semantic unit parsing based on rule enhancement. The method includes: constructing a multi-level domain rule base from forced matching rules, pattern template rules, and context constraint rules to generate a set of rule triples; loading the set of rule triples into an AC automaton to linearly scan the input text to obtain a set of candidate segments and trigger word position indices; constructing a candidate directed acyclic graph, and obtaining a structured word segmentation sequence and a pruned candidate pool through dynamic programming after dynamic modulation and fusion scoring; verifying the integrity of downstream slots, and locating backup candidate edges in the pruned candidate pool using gap signals when failure occurs, and obtaining a corrected word segmentation sequence through renegotiation and scoring. This application improves the completeness of professional terminology recognition in professional vertical scenarios and the self-correction capability of cross-domain word segmentation results.
Owner:TIANJIN FEIPENG SHENGYUAN TECHNOLOGY DEVELOPMENT CO LTD

A brain-computer interface system for recognizing the intention of Chinese oral language based on a sound-meaning integration double model

ActiveCN121560160BSpoken languageStereotaxis
The application provides a Chinese spoken language intention recognition brain-computer interface system based on a sound-meaning integration double model, belongs to the technical field of biomedical engineering, and relates to language brain-computer interface technology. Taking sound-meaning integration as the core, the stereotactic intracranial electroencephalogram (sEEG) technology is adopted to collect neural signals of the brain articulatory motor coding area and the semantic concept organization coding area. The system comprises a voice initiation decoder, a speech decoder, a semantic decoder and a Chinese word speech-semantic fusion synthesizer, the target decoder is constructed by extracting high gamma band features of key brain areas of the frontal lobe (left inferior frontal gyrus, premotor cortex, etc.), the temporal lobe (anterior temporal lobe, dorsolateral temporal lobe, etc.). At the same time, a visual and auditory induction training paradigm is matched, three tasks of listening to sound to group words, looking at words to group words and word association are set, and the subjects are supported to generate words independently. The system effectively solves the homonym and near homonym word ambiguity problem in Chinese spoken language recognition, and provides a precise interactive tool for ALS and other speech disorder patients.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Storage format for Chinese language and related processing method and apparatus

A storage format of Chinese language (“Readable Hanyu Expression” or “RHE”) and the related processing methods and systems. Unlike current Chinese processing methods which directly code Chinese characters into fonts for display, RHE takes an indirect approach by storing Chinese language in the RHE storage format that can be mapped to several display forms including simplified and traditional Chinese characters, Hanyu Pinyin, etc. In the RHE storage format, each Chinese word is stored as an RHE storage element having the format (Syllable+Tone)n+Mark, where n is the number of syllables (Chinese characters) in the word, Syllable represents the pronunciation (without the tone) of the character, Tone represents the tone of the pronunciation, and Mark is a value that differentiates different words having the same pronunciations and tones. Various mapping tables are used to map RHE storage elements to standard Chinese character codes (such as Unicode) and Pinyin expressions.
Owner:YANG MINGWEI

Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models

The invention provides a Chinese spoken language intention recognition brain-computer interface system based on pronunciation-meaning integration double models, belongs to the technical field of biomedical engineering, and relates to a language brain-computer interface technology. Sound-sense integration is taken as a core, and a stereotactic intracranial electroencephalogram (sEEG) technology is adopted to acquire neural signals of a brain phonetic motion coding region and a semantic concept organization coding region. The system comprises a sound production starting decoder, a voice decoder, a semantic decoder and a Chinese word voice-semantic fusion synthesizer, and a target decoder is constructed by extracting high gamma wave band characteristics of key brain regions of frontal lobe (left subfrontal gyrus, anterior cortex of motion and the like) and temporal lobe (anterior temporal lobe, dorsal lateral temporal lobe and the like). Meanwhile, an audio-visual induction training normal form is matched, three tasks of listening word combination, character reading word combination and vocabulary association are set, and subjects are supported to autonomously generate vocabularies. The system effectively solves the problem of ambiguity of homophonous and near-phonetic words in spoken Chinese recognition, and provides an accurate interaction tool for speech disorder patients such as ALS and the like.
Owner:BEIJING TIANTAN HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV

Dialect song generation method and dialect song display method and device

The invention discloses a dialect song generation method and device and a dialect song display method and device, and relates to the technical field of computers. The method comprises the steps of obtaining a first Chinese song and Chinese lyrics of the first Chinese song; dialect lyrics of the first Chinese song are generated through a language model according to the Chinese lyrics of the first Chinese song, and the dialect lyrics of the first Chinese song are used for describing semantics of the Chinese lyrics of the first Chinese song in a dialect language expression mode; a first dialect song corresponding to the first Chinese song is generated according to the first Chinese song, the dialect lyrics of the first Chinese song and the dialect pronunciation mapping relation through a song generation model, and the first dialect song is used for singing the first Chinese song through dialects. The dialect pronunciation mapping relation is established by adopting a pronunciation mode of at least one language to mark dialect pronunciation of the Chinese vocabulary. According to the method, a more accurate dialect pronunciation mapping relation is constructed, and the accuracy of dialect pronunciation in the first dialect song is improved.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

A big data topic analysis method based on an embedding model

This invention relates to a big data topic analysis method based on an embedding model. First, the Sentence-BERT model is used to perform sentence embedding representation on preprocessed Chinese text data. Then, the UMAP projection dimensionality reduction algorithm is used to reduce the dimensionality of the embedded vectors. Next, the HDBSCAN clustering algorithm is used to cluster the dimensionality-reduced vectors. Based on the assignment of each Chinese text in the target Chinese dataset to a corresponding topic class, the Chinese words with the highest c-TF-IDF scores are selected to represent each topic class. Finally, the DSG model is used to perform word embedding representation on the topic words, calculating the similarity between different topic words and between different topic classes, thereby detecting the volatility of newly emerging topic classes. The entire scheme design has higher topic consistency and topic diversity, and can detect new hot topics in a timely and accurate manner, providing early warnings.
Owner:HOHAI UNIV

A method and system for data pattern matching for interfaces

The application provides a data mode matching method and system for an interface, and belongs to the technical field of mode matching, and comprises the following steps: obtaining preprocessed data modes by preprocessing initial data modes of any two interfaces to be matched; obtaining language attributes of parameter names and parameter descriptions in the preprocessed data modes; if it is determined that the language attributes contain English words, calculating name similarity of the parameter names to obtain a name similarity result; if it is determined that the language attributes contain Chinese words, calculating semantic similarity of the parameter descriptions to obtain a semantic similarity result; and fusing the name similarity result and the semantic similarity result to output a data mode matching result of the any two interfaces to be matched. The application realizes automatic mode matching between heterogeneous mode interfaces provided by different systems, solves the problem that data mode matching between interfaces needs to be manually performed when heterogeneous data is integrated, and thus improves interoperability between systems.
Owner:WUHAN YANGTZE COMM ZHILIAN TECH +1

A method for generating a pinyin bucket word library of a four-level index

PendingCN122452553ADigital dataData integrity
The present application relates to the technical field of electronic digital data processing, and discloses a four-level index pinyin bucket word library generation method, acquires a Chinese word library text, extracts the first Chinese character and the first pinyin of each line of words as a classification key, and generates a triple; the triple is classified into a corresponding pinyin bucket according to the pinyin, the words in the bucket are sorted in descending order of word frequency, an independent Chinese character block is generated for each unique Chinese character, and the words are solidified according to the word length by using a first character multiplexing mechanism; a four-level static index of an initial letter statistical area and a pinyin index area and a Chinese character word index area and a word storage area is sequentially constructed, each index area uses fixed-length entries and absolute offset addressing; a metadata area containing a file length double-semantic field is constructed, a high-frequency-first deterministic truncation is performed on the word library according to a preset threshold, each data area is spliced and a data integrity check value is appended, and an embedded binary word library is generated. The problems of high storage redundancy and uncontrollable memory are solved, and the purposes of deterministic analysis, resource adaptation and high security are achieved.
Owner:SICHUAN HAIGE HENGTONG PRIVATE NETWORK TECH CO LTD

A machine translation software defect detection method based on combined semantics

The application discloses a machine translation software defect detection method based on combined semantics, comprising the following steps: S1, obtaining different Chinese translation sentences of the same English source sentence from different translation software; S2, using a word alignment model to correspond English words in the source sentence with segmented Chinese words in the Chinese translation sentences; S3, using a sentence compression model and a syntactic structure analysis method to obtain a main part and each additional part of the English source sentence respectively and form a sub-sentence set; S4, aligning each part of the English source sentence obtained in the step S3 with the corresponding translation part to obtain aligned translation of the main part and the additional part of the source sentence; and S5, detecting errors, including semantic similarity calculation and synonym checking.
Owner:TIANJIN UNIV

A professional word segmentation method applied to nuclear power industry

ActiveCN116306611BSolve the problem of inaccurate word segmentationreduce investmentData processing applicationsSemantic analysisProcess engineeringChinese word
The application belongs to the field of natural language processing in the nuclear power industry, and particularly relates to a professional word segmentation method applied to nuclear power industry corpus, which comprises nuclear power professional vocabulary construction, nuclear power stop word vocabulary construction, nuclear power synonym vocabulary construction, nuclear power same pronoun vocabulary construction, nuclear power field new word recognition, nuclear power field entity automatic recognition, nuclear power field synonym automatic recognition, nuclear power corpus Chinese accurate word segmentation and the like operations. The application has the beneficial effect of completely solving the problem of inaccurate Chinese word segmentation of nuclear power industry corpus, laying a solid foundation for the application of subsequent big data and artificial intelligence in the field of nuclear power natural language processing, and reducing the investment of other practitioners in the field of nuclear power in natural language processing.
Owner:CHINA NUCLEAR POWER OPERATION TECH CORP

Data processing method and system capable of improving efficiency of chinese vocabulary learning

The present invention provides a data processing method and system capable of improving efficiency of Chinese vocabulary learning. The method comprises: a data storage unit storing a vocabulary database related to Chinese language learning; segmenting each word into two or more sub-words; binding each sub-word to a balloon graphic, and assigning a color to the balloon; displaying four or more balloons on a client display screen, wherein one sub-word is displayed on each balloon, and the position of the balloon in the screen changes randomly; detecting an instruction for dragging one balloon in the client display screen so that the one balloon at least partially overlaps another balloon; upon detection of said instruction, indicating that phrase matching is performed on the sub-words bound to the selected balloons; and traversing the data storage unit, if there is a corresponding phrase in the data storage unit, indicating that the matching is successful, removing the balloons from the client display screen, and displaying the phrase in a set area of the client display screen, otherwise, retaining the balloons on the client display screen until all balloons in the client display screen are removed.
Owner:HONG KONG CHINESE COMPUTER SOCIETY CO LTD

Foreign Chinese education guidance system based on natural language processing

The invention discloses a foreign Chinese education guidance system based on natural language processing, which relates to the technical field of education and comprises an input module, a natural language processing module, a knowledge database, a learning state analysis module, a teaching strategy generation module, a teaching content pushing module and an output module. The system generates personalized teaching strategies according to specific conditions of learners, provides personalized learning paths and teaching contents according to Chinese basis, learning objectives and learning styles of different learners, integrates rich resources such as Chinese vocabularies, grammar rules, cultural knowledge, example sentences and exercise questions, intelligently recommends according to the conditions of the learners, and improves the learning efficiency. According to the method, the utilization efficiency of teaching resources is improved, the learning behaviors and achievements of learners are tracked in real time, the learning progress and the Chinese language ability are comprehensively evaluated and presented in a visual mode, the interactivity and adaptability of teaching are enhanced, in addition, multi-modal teaching is achieved, and richer and more vivid learning experience is provided for the learners.
Owner:厦门工学院

A Chinese text sorting system based on strong encoding and Chinese word segmentation

The present invention discloses a Chinese text sorting system based on strong encoding and Chinese word segmentation. This system implements Chinese text sorting based on a strong encoding model and Chinese word segmentation data. First, a database containing a large amount of Chinese text and corresponding labels is acquired. The labeled Chinese text data is used as input, and the Chinese text is word segmented and then encoded into a machine-readable format. The encoded sentences are then input into a Chinese text sorting model for model training. The trained model can then be used to automatically sort newly acquired Chinese text. This system achieves automated, highly accurate Chinese text sorting, taking into account the contextual relationships between Chinese words. This overcomes the low efficiency of manual text sorting and the low accuracy of traditional methods. The system is widely applicable and contributes to the intelligent development of military intelligence sorting, news topic classification, and film review classification, among other fields.
Owner:ZHEJIANG UNIV

A multi-language word vector weighting alignment method based on orthogonal pach analysis

This invention relates to the field of natural language processing technology and discloses a multilingual word vector weighted alignment method based on orthogonal Protodyakonov analysis. The method includes the following steps: S1: Randomly initialize an orthogonal matrix R, and randomly rotate the Chinese word vectors, i.e., X < -XR; substitute the weight matrix A, calculate the weighted orthogonal transformation matrix W according to the provided weighted Protodyakonov analysis, and obtain the weighted aligned Chinese word vector XW2. The Chinese word vectors are downloaded through a given word vector link to obtain Chinese and English word vectors X and Y. A given alignment dictionary L is used to weight sentiment words, with the goal of weighted alignment from X to Y. The method provided by this invention can perform weighted alignment according to the needs of downstream tasks, further improving the performance of downstream tasks, realizing a shared word vector space, and providing a mathematical proof of the weighted orthogonal Protodyakonov analysis.
Owner:GUANGZHOU UNIVERSITY

Method and apparatus for constructing a glossary of archaic words and phrases, electronic device, storage medium

Embodiments of the present disclosure provide a method and apparatus for constructing a vocabulary of archaic Chinese words, an electronic device, and a storage medium, belonging to the technical field of data processing. The method for constructing a vocabulary of archaic Chinese words includes: obtaining a data set of a preset archaic Chinese word list to be constructed; splitting each archaic Chinese character in the data set into radicals, obtaining the target radicals of each archaic Chinese character; encoding each of the target radicals to obtain a preliminary character sequence of each archaic Chinese character; counting the frequencies of consecutive radical pairs according to the preliminary character sequence; wherein the consecutive radical pairs include at least two adjacent target radicals; merging the preliminary character sequences corresponding to the consecutive radical pairs with the highest frequencies to obtain a target coding sequence number; constructing a target vocabulary according to the target coding sequence number. Through the technical solution provided by the embodiments of the present disclosure, an adaptive construction of a vocabulary of archaic Chinese words can be realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Translated content determination graphical user interface for an electronic device

1. The name of the design product: graphical user interface for judging translation content of electronic device. 2. The use of the design product: an electronic device. 3. The design points of the design product: the content of graphical user interface. 4. The picture or photo that best shows the design points: design 2 front view. 5. Design 2 is designated as the basic design. 6. The use of graphical user interface: for judging translation content. The human-computer interaction and change state of graphical user interface: click one Chinese word such as "avalanche" in design 1 front view to enter design 1 interface change state diagram 1; click the English word corresponding to "avalanche" such as "avalanche" in design 1 interface change state diagram 1 to enter design 1 interface change state diagram 2; keep for a preset time such as 1 second to enter design 1 interface change state diagram 3; click one Chinese word such as "apple" in design 1 interface change state diagram 3 to enter design 1 interface change state diagram 4; click the English word corresponding to "apple" such as "apple" in design 1 interface change state diagram 4 to enter design 1 interface change state diagram 5; keep for a preset time such as 1 second to enter design 1 interface change state diagram 6; wait until only the last Chinese word and the last English word are left to enter design 1 interface change state diagram 7; click the Chinese word "beef" in design 1 interface change state diagram 7 to enter design 1 interface change state diagram 8; click the English word "beef" corresponding to "beef" in design 1 interface change state diagram 8 to enter design 1 interface change state diagram 9; keep for a preset time such as 1 second to enter design 1 interface change state diagram 10. Click one Chinese word in the main view of design 2, such as "avalanche", to enter design 2 interface change state diagram 1; click an English word in design 2 interface change state diagram 1 that does not correspond to "avalanche", such as "apple", to enter design 2 interface change state diagram 2; keep for a preset time, such as 1 second, to enter design 2 interface change state diagram 3; click one Chinese word in design 2 interface change state diagram 3, such as "apple", to enter design 2 interface change state diagram 4; click an English word in design 2 interface change state diagram 4 that does not correspond to "apple", such as "mother", to enter design 2 interface change state diagram 5; keep for a preset time, such as 1 second, to enter design 2 interface change state diagram 6; click one Chinese word in design 2 interface change state diagram 6, such as "pencil", to enter design 2 interface change state diagram 7; click an English word in design 2 interface change state diagram 7 that does not correspond to "pencil", such as "mother", to enter design 2 interface change state diagram 8; keep for a preset time, such as 1 second, to enter design 2 interface change state diagram 9; click one Chinese word in design 2 interface change state diagram 9, such as "pencil", to enter design 2 interface change state diagram 10; click an English word "beef" in design 2 interface change state diagram 10 that does not correspond to "pencil" to enter design 2 interface change state diagram 11; keep for a preset time, such as 1 second, to enter design 2 interface change state diagram 12.
Owner:BEIJING WATERDROP TECH GRP CO LTD

An Automatic Evaluation Method for Course Teaching Cases Based on Bloom's Taxonomy

The present invention discloses an automatic evaluation method for course teaching cases based on the Bloom classification method. First, teaching cases from the computer science courses to be evaluated are collected and organized into a document containing several teaching cases. Then, the organized case data is annotated and standardized into the format of a standard data set. Next, a classification model based on a pre-trained language model and a convolutional neural network is built, and the pre-processed case data set is used to train the model. Finally, the data set of the course teaching cases to be evaluated is fed into the trained classification model based on the pre-trained language model and the convolutional neural network for text classification, and an accurate classification result of the data set of the test cases to be classified is obtained. The method of the present invention constructs a small-sample data set of course teaching cases. The built evaluation model enhances the ability to recognize and segment Chinese words, has better semantic representation ability and the ability to capture context semantics, and improves the performance of automatic classification of course teaching cases.
Owner:GUILIN UNIV OF ELECTRONIC TECH

A method and device for recognizing terms in the automobile industry, and a storage medium

The application relates to the field of speech recognition and discloses a vehicle industry term speech recognition method and device and a storage medium. The method comprises the following steps: constructing a Chinese word-phone pronunciation dictionary and a vehicle professional term BPE-phone pronunciation dictionary; training an initialized basic speech recognition model by using Chinese speech data; and migrating the parameter weight from the basic speech recognition model, and training a target speech recognition model by using the Chinese speech data on the migrated model. According to the embodiment of the application, a speech recognition system with good effect can be efficiently trained under the condition that only a small amount of high-quality vehicle professional term speech data set and a large amount of unlabeled vehicle professional term speech data are needed. The finally obtained speech recognition system does not use a pre-training model with excessive parameter quantity, ensures the speed of speech recognition, and reduces the requirement on hardware in actual application.
Owner:SHENZHEN SILICON MOUNTAIN TECH CO LTD

Virtual anchor video generation method, system, electronic device and storage medium

The present invention relates to a method, system, electronic device, and storage medium for generating a virtual anchor video, comprising receiving a broadcast text and performing polyphonic word processing on the broadcast text; setting virtual anchor actions based on the broadcast text after polyphonic word processing; rendering the virtual anchor image in response to a change operation on the anchor image layer; integrating the polyphonic word processing text, the virtual anchor actions, and the virtual anchor image into a unified data structure for rendering to generate a virtual anchor video. The present invention effectively solves the problem of polyphonic word recognition in Chinese words through word segmentation operations and pinyin rendering, ensuring the accuracy of the broadcast text, and uses the Laplace smoothing-naive Bayes algorithm to recommend appropriate virtual anchor actions based on text semantics, so that users can recommend appropriate anchor actions based on text semantics during actual use.
Owner:GUANGDONG SOUTH SMART MEDIA TECH CO LTD

An automobile accessory name word segmentation method and system

The application belongs to the technical field of computers, and particularly relates to a kind of automobile accessory name word segmentation method and system.The method comprises: creating a list of accessory names, creating a containment relationship graph, and using the containment relationship graph.The list of accessory names is created by manual annotation, the created list of accessory names is sorted by accessory name length, and a processed accessory name list with an initial state of empty is created;for each accessory name in the sorted list of accessory names, the following operations are performed: finding a containment relationship, generating a directed acyclic graph, setting a weight, finding the optimal path in all paths according to a dynamic programming algorithm, and updating the containment relationship graph.The application uses a containment relationship graph to save the word segmentation results in a graph structure, not only completes Chinese word segmentation, but also saves the knowledge of the automobile accessory field contained in the word segmentation results, provides an important source of knowledge for creating a knowledge graph in the field of automobile accessories, and can improve the accuracy and efficiency of Chinese word segmentation.
Owner:FUJIAN ZHONGCHUANG AUTOLINK NETWORK TECH CO LTD

A Chinese word segmentation method, device and storage medium

The application discloses a Chinese word segmentation method, device and storage medium, and belongs to the technical field of natural language processing. The Chinese word segmentation method comprises the following steps: S1, a second language translation sentence of a to-be-detected sentence is acquired; S2, a Chinese Bert pre-training language model is used to code the to-be-detected sentence, so as to acquire vector representation of semantic information of the whole sentence and a sentence vector representation sequence; S3, a second language Bert pre-training language model is used to code the translation sentence, so as to acquire vector representation of semantic information of the whole sentence; S4, semantic features of the to-be-detected sentence and the translation sentence are fused, so as to acquire a predicted category of each word of the to-be-detected sentence; and S5, the to-be-detected sentence is segmented according to the predicted category, so as to acquire a word segmentation result. The method improves the accuracy of word segmentation, and has a good word segmentation effect on foreign words in particular.
Owner:ZHONGKE FANYU (WUHAN) TECH CO LTD

Text sensitive word library extraction method, device and equipment based on neural network model

The application relates to a text sensitive word library extraction method, device and equipment based on a neural network model. The method comprises the following steps: constructing a sensitive word library extraction model; the sensitive word library extraction model comprises a self-defined rule algorithm, a character-by-character segmentation algorithm, a Chinese word segmentation algorithm and an N-gram algorithm; the sensitive word library extraction model is pre-trained according to a training data set; the pre-trained sensitive word library extraction model is used for word segmentation extraction and word library quantitative analysis on the text to be extracted; the full-amount word library is subjected to field structure design; the pre-trained sensitive word library extraction model is iteratively trained according to the designed word library; after the sensitive word extraction is carried out by using the trained sensitive word library extraction model, data analysis is carried out according to a word segmentation frequency filtering rule, a word segmentation character quantity filtering rule and a word segmentation type filtering rule, and a sensitive word library is obtained. The method can improve the sensitive word extraction accuracy and efficiency.
Owner:HUNAN INKE INTERACTIVE ENTERTAINMENT NETWORK INFORMATION CO

Chinese neck vessel ultrasound prompt generation method and system

The invention provides a Chinese neck vessel ultrasound prompt generation method and system, and belongs to the technical field of text data processing.The method comprises the steps that firstly, Chinese word segmentation is conducted on a text so that the text can be converted into a word sequence, and starting and ending marks are added to the head and tail of the sequence; replacing a separation mark between the sequences and making a guide sequence; secondly, building a multi-attention mechanism LSTM network which comprises a Bi-LSTM encoder part, an LSTM decoder part, a part for checking a seen attention mechanism, a part for prompting attention mechanism and a part for prompting internal attention mechanism; and finally, training a network based on cross entropy loss between a multi-attention mechanism LSTM prompt generation result and a manual annotation gold standard. According to the method, the potential correlation between the ultrasonic prompts and the attention represented by the features in the prompts are introduced, so that the purpose of improving the ultrasonic prompt generation effect is achieved.
Owner:BEIJING JIAOTONG UNIV

Deep learning-based Chinese word segmentation method

The invention provides a Chinese word segmentation method based on deep learning, and relates to the technical field of natural languages, and the method comprises the steps: dividing a Chinese word segmentation public data set according to a proportion, and carrying out the pre-training, and obtaining the multi-dimensional features of the word granularity; and constructing a Chinese word segmentation model TC-CRF based on deep learning, and performing Chinese word segmentation processing on the multi-dimensional features by using the Chinese word segmentation model TC-CRF. According to the method, the problem that long dependence cannot be obtained by an existing convolution feature extraction method is solved, the problem that the global transformer obtains too large noise is solved, and Chinese word segmentation is effectively realized by using deep learning.
Owner:SOUTHWEST JIAOTONG UNIV

Expressway traffic accident cause network construction method based on text mining

The invention discloses a text mining-based highway traffic accident cause network construction method, which comprises the following steps of: collecting highway traffic accident investigation reports, performing data selection, conversion and data cleaning, and constructing an accident text corpus; performing targeted Chinese word segmentation, loading and merging word lists, customizing professional dictionaries and stop word lists, and completing corpus datamation; the weights and word frequencies of the text feature words are considered, dimension reduction screening is further carried out, and accident cause feature keywords are extracted. According to the method, the expressway traffic accident unstructured text data is fully utilized, a highly interactive vocabulary mode in the obvious accident cause data is constructed, and compared with application of structured data, the situation when an accident occurs can be better restored; the accident cause network established according to the point mutual information value can effectively solve the problem of high-frequency co-occurrence but weak semantic relationship between lexical items, clearly displays the significance logic and association between causes, and improves the quality and interpretability of the highway traffic accident cause network.
Owner:YELLOW RIVER CONSERVANCY TECHN INST

Dependency graph-based open geological spatial relationship extraction method and system

The invention belongs to the cross technical field of geographic information science and natural language processing, and particularly discloses an open geological space relation extraction method and system based on a dependency graph, and the method comprises the steps: obtaining a sentence text set according to geological text data, screening sentences at least comprising two space entities from the sentence text set to form a subject-object pair, and obtaining a screened sentence text set; constructing an entity relationship triple, obtaining segmented words, part-of-speech and syntactic analysis, and forming a geological space relationship data set; training an extraction model by using the training and evaluation sample set, wherein the model comprises a context semantic coding module, a dependency graph attention network module and a two-stage triple extraction module; and testing by using the trained extraction model, performing a two-stage prediction process on an input test data set, firstly measuring a relationship and then predicting an entity, and generating a structured triple. The method can overcome Chinese word segmentation errors, effectively captures long-distance dependence, and is suitable for open relation extraction in the geological field.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Intraspinal anesthetic needle puncture virtuality and reality combination teaching system

PendingCN121306371AEducational modelsMedical referencesAnesthesia needleEpidural needles
The invention discloses an intraspinal anesthetic needle puncture virtuality and reality combined teaching system, which relates to the technical field of medical teaching evaluation, and comprises the following steps of: 1, compiling an index definition library, a Chinese word matrix and a report template into a rule net version through unit and dimensional reasoning and equation and interval checking, and carrying out property test; 2, aligning and normalizing the multi-source sequence and the event according to a conversion chain, and packaging into a fact package containing a source, a formula and a unit; 3, generating a rendering plan based on the fact package, preposing word use and display precision, and pre-scoring; 4, constrained generation and reverse checking are carried out, minimum correction is carried out under the condition that a fact numerical value is not changed, and a consistency text and a final consistency score are output; 5, signing and issuing a gating and embedding a pedigree and a signature; in case of data, threshold or template change, calculating a change influence domain, re-rendering the influenced sentences, and outputting differences and auditing; consistency, traceability and verifiability of facts and characters are realized.
Owner:THE SECOND AFFILIATED HOSPITAL TO NANCHANG UNIV