Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Morpheme" patented technology

A morpheme is the smallest meaningful unit in a language. A morpheme is not identical to a word. The main difference between them is that a morpheme sometimes does not stand alone, but a word, by definition, always stands alone. The linguistics field of study dedicated to morphemes is called morphology. When a morpheme stands by itself, it is considered as a root because it has a meaning of its own (such as the morpheme cat). When it depends on another morpheme to express an idea, it is an affix because it has a grammatical function (such as the –s in cats to indicate that it is plural). Every word comprises one or more morphemes.

A rule corpus-based text specification marking method and system

The application relates to the technical field of text label marking, and provides a text specification marking method and system based on a rule corpus, which comprises the following steps: analyzing a policy and regulation document, identifying and marking condition morphemes and conclusion morphemes in the policy and regulation document, constructing a logical relationship between the two by using a large language model, and forming a rule corpus composed of structured morpheme pairs; performing semantic embedding on the corpus to generate a semantic vector library; performing multi-label coding on a verification data set based on the rule corpus, and constructing a multi-label training data set; training a deep learning classification model by taking semantic vectors as features and multi-labels as targets, so that a text specification marking model is obtained; and automatically marking target text by using the model. The application significantly improves the accuracy, interpretability and business adaptability of text marking, improves the update quality of a system label data set, and reduces the system maintenance cost.
Owner:SSE INFORMATION NETWORK LTD

Text specification marking method and system based on rule corpus

The invention relates to the technical field of text label marking, and provides a text standard marking method and system based on a rule corpus, and the method comprises the steps: analyzing a policy and regulation file, recognizing and marking a condition morpheme and a conclusion morpheme in the policy and regulation file, constructing a logic relation between the condition morpheme and the conclusion morpheme by using a large language model, and marking the rule corpus; forming a rule corpus consisting of structured morpheme pairs; performing semantic embedding on the corpus to generate a semantic vector library; performing multi-label coding on the verification data set based on the rule corpus, and constructing a multi-label training data set; training a deep learning classification model by taking the semantic vector as a feature and taking multiple tags as a target to obtain a text specification marking model; and automatically marking the target text by using the model. According to the method, the accuracy, the interpretability and the service adaptability of text marking are remarkably improved, the updating quality of the system label data set is improved, and the system maintenance cost is reduced.
Owner:SSE INFORMATION NETWORK LTD

Multi-task processing oriented Soc chip server computing power distribution method

The invention relates to the technical field of computers, and discloses a multi-task processing-oriented Soc chip server computing power distribution method, which comprises the following steps of: acquiring execution primitive vectors of at least two tasks concurrently executed on an Soc chip; converting the execution primitive vector into an execution symbol sequence corresponding to each task; based on the execution symbol sequence, identifying an execution morpheme representing a task execution stage, and constructing a task grammar state machine, used for predicting a next morpheme, of each task; analyzing cross-task morpheme association between the execution morphemes of the concurrent tasks; and on the basis of the current morpheme of a certain task, the predicted next morpheme is associated with the cross-task morpheme, and a collaborative scheduling instruction is generated. By converting the execution primitive vector into the execution symbol sequence and identifying the execution morphemes based on the sequence to construct the task grammar state machine, the prediction of the next execution stage of the task is realized.
Owner:深圳麓麟科技有限公司

Method for segmenting sign language into morphemes, method for predicting morpheme positions, and method for augmenting data

ActiveUS12608980B2Natural language data processingBiometric pattern recognitionFrame (artificial intelligence)Morpheme
Provided are a method for segmenting sign language into morphemes, a method for predicting morpheme positions, and a method for augmenting data. A system for analyzing sign language according to an embodiment of the present invention comprises: a recognition unit which recognizes key points of a speaker from a sign language video; and a prediction unit which inputs the recognized key points into an artificial intelligence model, segments the sign language into morphemes, and predicts position information of the segmented morphemes. Accordingly, by recognizing the morphemes of the sign language video frame by frame on the basis of a skeletal model and thereby segmenting the sign language into morphemes and predicting morpheme positions, it is possible to lay the foundations for accurate sign language translation.
Owner:KOREA ELECTRONICS TECH INST

Boundary perception and pre-training enhanced low-resource Tibetan language recognition method and system

The invention discloses a boundary perception and pre-training enhanced low-resource Tibetan language recognition method and system, and the method comprises the following steps: 1, carrying out the standardization processing of data and texts, and enabling the language knowledge and pre-training model enhanced low-resource Tibetan language voice to be converted into text. Unified modeling of Tibetan language boundaries, normal characters and morphemes is realized under a low-resource condition; remarkable and stable alignment is realized through a tsheg-perception adapter and three-head joint learning; language knowledge is fused into search in a learnable mode through the micro WFST soft constraint, and illegal font and language / case marking errors are greatly reduced. Training data are effectively expanded through teacher semi-supervised and coverage rate active learning, and cross-domain and accent robustness is improved; the end side maintenance cost is remarkably reduced through'adapter + atlas' hot updating; iCR and a boundary F1 are introduced as closed-loop indexes, so that the quality of an optimization target is consistent with that of a Tibetan Chinese character method, and the whole method has the advantages in engineering landing, interpretability and maintainability.
Owner:INSTITUTE OF ETHNOLOGY & ANTHROPOLOGY CHINESE ACADEMY OF SOCIAL SCIENCES

Information processing apparatus method for responding to user inquiry via information source or operator

An information processing apparatus outputs answer information corresponding to inquiry information that is input. The information processing apparatus includes a memory and circuitry. The memory is configured to store a plurality of databases each having at least a first field and a second field. The circuitry is configured to: perform morphological analysis on the inquiry information, to divide the inquiry information into morphemes; perform a first matching process based on the morphemes and the first field of each of the plurality of databases, to determine whether to adopt the database as an extraction source from which the answer information is to be extracted; and perform a second matching process based on the morphemes and the first field of the database, which is determined to be adopted as the extraction source, to output, as the answer information, data in the second field corresponding to data in the first field.
Owner:RICOH CO LTD

A general english tweet preprocessing method and computer equipment

The present application relates to a kind of general english tweet preprocessing method and computer equipment, belong to data processing technical field;Solved the problem that a large number of subjective words and non-standard morphemes exist in english tweet, affect tweet preprocessing result and named entity recognition performance;The method of the present application includes: based on multiple field english text, subjective word table is obtained by construction;The non-standard morpheme in the english tweet to be processed is carried out semantic reduction and information extraction, and the tweet text after preprocessing is obtained;Based on the tweet text after preprocessing, double stack structure is constructed to extract clause;Based on subjective word table, using syntax dependency analysis model and tree parent-child level structure, the tweet text after preprocessing is carried out named entity extraction, and the named entity recognition result of english tweet is obtained;The preprocessing result of english tweet is obtained by outputting the tweet text after preprocessing, the clause contained in english tweet and the named entity recognition result of english tweet.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

Automated digital text optimization and modification

A system, method, and computer program product for implementing automated digital text optimization is provided. The method includes receiving during execution of an Internet search process, digital textual content comprising embedded words, morphemes, and phrases. A specified format for generating a summary for the digital textual content is identified and probability distribution attributes are generated with respect to target vocabulary of the digital textual content. Content of the digital textual content is replaced with replacement content and an automated digital summary of the digital textual content is generated thereby enabling optimized operational functionality of a recurrent neural network (RNN) encoder-decoder hardware device. A decoded version of the automated digital summary is presented via a specialized graphical user interface.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

A zero-sample-based song timbre rapid conversion method and device

This invention discloses a method and apparatus for rapid conversion of singing voice timbre based on zero-shot sampling. The method includes constructing a singing dataset containing dry vocals and lyrics; extracting the audio codebook index sequence of the dry vocals using a singing feature decoupler incorporating a Hubert model and residual quantization codebook; introducing a text encoder to extract morpheme features and morpheme index sequences from the lyrics; optimizing the singing feature decoupler through cross-prediction to improve the accuracy of speech content feature extraction; introducing pitch and timbre features to represent prosody; and enhancing the quality of synthesized vocals generated by the generator based on speech content features, pitch features, and timbre features through adversarial training. This achieves rapid conversion of song vocal timbre to the user's timbre while ensuring conversion quality.
Owner:BEIJING DUIJIUDANGGE TECH CO LTD

A morphologically enhanced tensorized word embedding compression system

The application discloses a morphological enhancement based tensorized word embedding compression system, which comprises a morpheme segmentation module, a morpheme index and embedding module and a word embedding generation module; the morpheme segmentation module divides each word in a word table of a text task into morphemes; the morpheme index and embedding module firstly generates a morpheme table by counting the division result of the morpheme segmentation module, then defines a morpheme index matrix and a plurality of trainable morpheme embedding matrices, each row of the morpheme index matrix represents the position of the morpheme of a corresponding word in the word table in the morpheme table, and each row of the morpheme embedding matrix represents an embedding vector of a corresponding morpheme in the morpheme table; the word embedding generation module indexes the morpheme vector from the morpheme embedding matrix and performs a tensor product on each word in the word table, adds the results of a plurality of tensor products to generate a word embedding vector; and the application overcomes the problems of a large amount of parameters and storage space occupation in general word embedding technology, and the problem of task effect loss when high compression word embedding is performed.
Owner:TIANJIN UNIV +1

Information processing apparatus method for responding to user inquiry via information source or operator

An information processing apparatus outputs answer information corresponding to inquiry information that is input. The information processing apparatus includes a memory and circuitry. The memory is configured to store a plurality of databases each having at least a first field and a second field. The circuitry is configured to: perform morphological analysis on the inquiry information, to divide the inquiry information into morphemes; perform a first matching process based on the morphemes and the first field of each of the plurality of databases, to determine whether to adopt the database as an extraction source from which the answer information is to be extracted; and perform a second matching process based on the morphemes and the first field of the database, which is determined to be adopted as the extraction source, to output, as the answer information, data in the second field corresponding to data in the first field.
Owner:RICOH CO LTD

Japanese evaluation system, japanese evaluation apparatus, japanese evaluation method, and program

To provide a Japanese evaluation system, a Japanese evaluation device, a Japanese evaluation method, and a program capable of efficiently creating a clear Japanese sentence suitable for foreign language translation.SOLUTION: A Japanese evaluation system 100 includes a morphological analysis unit 212 that divides an acquired Japanese target sentence into morphemes by morphological analysis, a measurement unit 213 that measures the number of times of detection of a predetermined particle in the morphemes of the target sentence, and a setting unit 214 that, when the predetermined morpheme is detected from the morphemes of the target sentence, divides the number of times of detection before and after the predetermined morpheme and sets the divided number of times of detection as the number of times of detection related to the target sentence.SELECTED DRAWING: Figure 4
Owner:KONICA MINOLTA INC

Sign language translation system and method based on time domain and representation alignment

The invention discloses a sign language translation system and method based on time domain and representation alignment. The method comprises the following steps: extracting sign language video visual features through a visual encoder; generating an alignment path of the visual feature sequence and the sign language morpheme sequence based on a discriminant dynamic time warping algorithm; training a sign language fragment segmentation model by using the aligned path; the sign language fragment features are aggregated into time domain alignment visual representation by using an adaptive time sequence fusion module; performing vision-semantic representation alignment training according to semantic embedding of morphemes; in the sign language translation step, visual features of a video to be translated are extracted through a visual encoder, and morpheme fragments are generated by using a sign language fragment segmentation model; a morpheme-level feature is obtained through the self-adaptive time sequence fusion module and then input into a translation decoder, and a corresponding natural language text is output; according to the method, the problems of speed difference and visual-semantic modal difference of sign languages can be solved at the same time, and translation errors caused by similar actions between sign language primitives are reduced, so that the translation accuracy is improved.
Owner:TIANJIN UNIV

Mongolian multi-modal sentiment analysis method based on cross-modal information enhancement and fusion

The invention discloses a Mongolian multi-modal sentiment analysis method based on cross-modal information enhancement and fusion. In order to solve the problems that emotion expression is dispersed and implicit clues in multi-modal emotion information are difficult to capture due to Mongolian adhesive language characteristics, a three-level innovation scheme is provided: (1) a cross-modal explicit enhancement layer: converting the implicit emotion clues such as rhythm, expression and the like in audios and videos into emotion description texts by using a multi-modal large language model; the modal information is enriched through a path of implicit, explicit and semantic fusion; (2) a semantic alignment layer based on a Gram matrix: extracting a second-order statistical structure of a text modal as a semantic reference, mapping heterogeneous audio and video features to a unified semantic space, and solving cross-modal alignment difficulty caused by Mongolian morphological change; and (3) a gating displacement fusion layer: a self-adaptive displacement mechanism is adopted, accurate injection of non-text information is realized through fine tuning of semantic positions of text features, and the problem of modal submerging in a Mongolian long morpheme sequence is avoided.
Owner:INNER MONGOLIA UNIV OF TECH

Lip language recognition method based on linguistic prior and progressive cross-modal contrast disambiguation network

The invention discloses a linguistic prior and progressive cross-modal contrast disambiguation network-based lip language recognition method, which is suitable for the field of computer vision speech recognition and comprises the following steps of: 1, on the basis of a linguistic vision-phoneme mapping rule, constructing a fuzzy text and realizing replacement by replacing similar phonemes in a pronunciation text; and 2, constructing a progressive comparison disambiguation network which is composed of a vision-phoneme disambiguation sub-module and a text semantic disambiguation sub-module, mapping in coding through vision-phoneme feature alignment and text semantic alignment in a mode, and realizing end-to-end ambiguity analysis in decoding. According to the method, through fusion of linguistics rules and a deep learning technology, the whole process optimization of fuzzy sample construction-progressive disambiguation-targeted evaluation is realized, so that the decoupling capability of the model for visual morpheme-phoneme asymmetric ambiguity can be improved, and the lip language recognition precision and accuracy of the fuzzy sample are improved.
Owner:HEFEI UNIV OF TECH

Coding and application method and system constructed based on rule file and corpus

ActiveCN121960503ASolving interpretability challengesabsolute certaintySemantic analysisBusiness practiceTheoretical computer science
The invention relates to the technical field of text label application management, and provides a coding and application method and system constructed based on a rule file and a corpus, and the method comprises the steps: carrying out the file coding of the rule file; performing chapter and section term coding on chapter and section terms in the rule file; constructing a corpus unit; file codes, chapter and clause codes and corpus units are associated and combined to form a traceable code combination, the traceable code combination is applied to standardized marking of business file key information, standard reference traceability is formed through coding, and a business corpus can be formed by combining business practice and generalization mapping to improve the application effect. According to the method, reverse tracing of original rule files, clause positions and morpheme semantics is realized, and forward retrieval can be carried out by utilizing texts which are matched with conditions through codes. According to the method, semantic disambiguation and accurate traceability of the rule text are realized through a multi-level coding system, meanwhile, the generalization understanding capability of an AI technology is compatible, and the accuracy of text marking is improved.
Owner:SSE INFORMATION NETWORK LTD

Mongolian multi-modal sentiment analysis method based on cross-modal information enhancement and fusion

A Mongolian multi-modal sentiment analysis method based on cross-modal information enhancement and fusion. To solve the problems of scattered sentiment expression caused by the agglutinative characteristics of Mongolian and the difficulty of capturing implicit clues in multi-modal sentiment information, a three-level innovative solution is proposed: (1) Cross-modal explicit enhancement layer: Use multi-modal large language models to convert implicit sentiment clues such as prosody and expression in audio and video into text descriptions of sentiment, enriching the information of each modality through the "implicit→explicit→semantic fusion" path; (2) Semantic alignment layer based on Gram matrix: Extract the second-order statistical structure of the text modality as the semantic reference, and map heterogeneous audio and video features to a unified semantic space to solve the cross-modal alignment difficulty caused by the morphological changes of Mongolian; (3) Gating displacement fusion layer: Use an adaptive displacement mechanism to fine-tune the semantic position of text features to achieve precise injection of non-text information, avoiding the modality submersion problem in long morpheme sequences of Mongolian.
Owner:INNER MONGOLIA UNIV OF TECH

A big data-based information processing method and system

This invention belongs to the field of big data information technology and discloses an information processing method and system based on big data: The method involves segmenting the sentences of the text to be detected to obtain a second set of sentences; filtering the second set of sentences using a sensitive word library to obtain a first set of candidate sensitive sentences and a third set of sentences; calculating the sentence similarity between sentences in the first set of candidate sensitive sentences and sentences in the sensitive word library, with sentences having a maximum similarity greater than or equal to a first threshold being considered sensitive sentences; recombining the morphemes of sentences in the third set of sentences; filtering the recombined sentences using a sensitive word library to obtain candidate sensitive sentences; and calculating the sentence similarity between candidate sensitive sentences and sentences in the sensitive word library, with the maximum similarity greater than or equal to a first threshold being considered sensitive sentences. 1 The statement is identified as a sensitive statement at that time; the maximum similarity is less than TH. 1 But greater than or equal to TH 2 In such cases, the information is then reviewed manually. This invention improves the detection rate and accuracy of sensitive information.
Owner:XINYANG AGRI & FORESTRY UNIV

Electronic apparatus for comparing similarity between texts on basis of sub-sentence extraction and method for comparing similarity between texts by using same

An electronic apparatus for comparing a similarity between texts on the basis of sub-sentence extraction, according to the present disclosure, may comprise: a memory storing at least one instruction; and at least one processor that executes the at least one instruction. The at least one processor may: perform analysis of dependence between a plurality of morphemes included in at least two pieces of text data; generate a dependent relationship graph for the plurality of morphemes, on the basis of the dependence analysis; extract at least one sub-sentence by using the dependent relationship graph; and calculate a similarity between the at least two pieces of text data by cross-comparing the at least one sub-sentence.
Owner:MUHAYU INC

A sign language generation method based on semantic guided diffusion model

The application discloses a sign language generation method based on a semantic guidance diffusion model, which comprises the following steps: inputting a sign language morpheme sequence and a noise posture sequence into a text encoder and a visual encoder respectively, extracting text features and posture features, inputting the posture features and the text features into a local enhancement module, each local enhancement layer of the local enhancement module being stacked and comprising adaptive sign language graph convolution and time convolution, the posture features being enhanced under the guidance of the text features, finally obtaining local enhanced posture features, converting the local enhanced posture features through a feature converter to obtain converted features, then passing the converted features through a global enhancement module to obtain global enhanced posture features, and inputting the global enhanced posture features into a Fiergi to obtain a predicted sign language posture sequence. The application takes into account the delicacy of local joint generation and the natural coherence of the overall sequence, and a semantic consistency guidance mechanism can effectively ensure that the generated posture sequence is highly consistent with the corresponding sign language morpheme sequence at the semantic level.
Owner:ZHEJIANG UNIV OF TECH

Information processing method, information processing system, and program

The information processing method includes: a step (S10) for acquiring information relating to a person's communication as text information (IT); a step (S20) in which morpheme analysis is performed on the text information (IT) to decompose the text information (IT) into a plurality of sentences (w), and sentiment analysis is performed on the plurality of sentences (w); and a step (S30) for forming a matrix table (M) by arranging items relating to a plurality of sentiment expressions indicating the sentiment of the person in a first direction and arranging items relating to attribute partitions indicating the attributes of the person or the organization in a second direction, and visualizing the analysis results of the sentiment analysis in correspondence with each row and each column in the matrix table (M).
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Information processing device, information processing method, and program

PCT designated stageWO2025263301A1Notation recordingInformation processingMorpheme
The present disclosure relates to an information processing device, an information processing method, and a program that make it possible to appropriately analyze music. In the present invention, on the basis of lyrics data of music data, the music data is divided into segments on the basis of line feed information, and after morpheme analysis, the segments are converted into a matrix according to the number of mora, the degree of similarity corresponding to the distance between the segments is calculated by using the matrix, and parts such as a chorus, A-melody, and B-melody constituting the musical piece are estimated by using the segments as units on the basis of the similarity between the segments. The present invention can be applied to a music analysis device.
Owner:SONY GROUP CORP

system

We provide the system. [Solution] Means for collecting publicly available document data, A method for extracting text data from collected document data, A method for preprocessing extracted text data to remove noise and normalize the text, A means for analyzing morphemes from preprocessed text data and extracting important topics, A means for generating a summary based on extracted topics, A means of saving the generated summary and providing it to the user, A system that includes this.
Owner:SOFTBANK GROUP CORP

Large language model evaluation method and device, equipment, storage medium and program product

The invention discloses a large language model evaluation method and device, equipment, a storage medium and a program product, and relates to the technical field of computers, the method comprises the steps of obtaining a target prediction query statement and a target reference query statement thereof, the target prediction query statement being generated by a first large language model for a target service; obtaining a first evaluation result of the target prediction query statement based on a morpheme comparison result between the target prediction query statement and the target reference query statement; under the condition that the first evaluation result represents that the target prediction query statement does not meet the preset morpheme requirement, performing semantic evaluation on the target prediction query statement based on the target reference query statement to obtain a second evaluation result of the target prediction query statement; and obtaining a performance evaluation result of the first large language model based on the first evaluation result and the second evaluation result. According to the method and the device, the problem of relatively low accuracy of SQL capability evaluation of the large language model can be solved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD