Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

58 results about "Common word" patented technology

Systems and methods for updating and executing language models using input from distributed computing environments

Described herein are interactive systems and methods for training language models using input from distributed computing environments. The system can receive prompts for a language model. Each prompt can include at least one common term corresponding to an intent relating to wagers. The system can determine that the language model has not been updated using training examples that include the at least one common term corresponding to the intent. The system can generate, using the prompts and additional information corresponding to the at least one common term, a set of training examples. Each set of training examples can include a respective prompt having the at least one common term. The system can update the language model using the set of training examples.
Owner:DK CROWN HOLDINGS INC

Using synthetic data to improve word error rate of differentially private ASR models

A method (600) includes pre-training an audio encoder (210) on a public training utterance set and from a corpus of text utterances (501), sampling a predetermined number of most frequent words (501). The method also includes randomly generating a predetermined number of transcripts (520), and for each corresponding transcripts, processing, using a TTS system (550), the corresponding transcript to generate a corresponding synthetic speech utterance (532). The corresponding transcript and the corresponding synthetic speech utterance form a corresponding synthetic training sample (502). During a first fine-tuning stage, the method also includes fine-tuning an ASR model (200) on the synthetic training samples. During a second fine-tuning stage, the method also includes fine-tuning, using a differentially private parameter-efficient-fine-tuning (DP-PEFT) technique, the ASR model on a plurality of private training samples (402), wherein the DP-PEFT technique updates only a subset of newly added or existing parameters of the pre-trained audio encoder.
Owner:GOOGLE LLC

IMPROVING SPEECH RECOGNITION TRANSCRIPTIONS

ActiveDE102021122068B4Speech recognitionAudiometry testCommon word
Computer-implemented method (500) for training a model to improve speech recognition, wherein the computer-implemented method comprises: Receiving (502) an utterance by one or more processors, wherein the receiving is performed by a virtual assistant in a specific node of the virtual assistant, wherein frequently occurring terms have been identified for the specific node over a period of time; Transcription (504) of the utterance into text by the one or more processors; Generating (506) a transcription confidence score based on transcription and audiometry by one or more processors; in response to the transcription confidence score being below a threshold, comparing (510) phonemes in the utterance with phonemes in at least one term from a list of frequently occurring terms by one or more processors; Generating a sound similarity score for phonemes in which at least one term from a list of frequently occurring terms is selected by one or more processors; and Replacing (512) the transcription with at least one term from the list of frequently occurring terms if the sound similarity score is above a threshold, by one or more processors.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Document classification apparatus, method, and storage medium

According to one embodiment, a document classification apparatus includes a processing circuit. The processing circuit is configured to: acquire text content for each of logical elements for semi-structured document data including text data stored for each of the logical elements; select logical elements from the logical elements and generating logical element sets each including the logical elements; analyze text contents for the respective logical element sets and constructing respective word embedded spaces; select a first word embedded space and a second word embedded space including a common word shared with the first word embedded space from the word embedded spaces, and update the first word embedded space based on similarity to the common word in the second word embedded space; and output a classification result of the document data using the first word embedded space and embedding information of a feature quantity of a classification target.
Owner:KK TOSHIBA

Methods, devices, electronic devices, readable storage media, and computer program products for protecting large language model datasets based on polluting lexical units.

This application discloses a method, apparatus, electronic device, readable storage medium, and computer program product for protecting a large language model dataset based on polluted lexical units, belonging to the field of artificial intelligence technology. The method includes: generating a polluted data statistical template by statistically analyzing each data item in the original dataset to obtain the high-frequency common words and grammatical structure characteristics of the original dataset; generating a polluted lexical sequence including high-frequency common words and polluted lexical units, wherein the polluted lexical units include: user-created words that do not exist in reality, generated by arranging common characters, and their corresponding parts of speech; generating polluted sentences using the polluted lexical sequence based on the grammatical structure characteristics of the original dataset; filling the corresponding data item of each data item in the original dataset with the polluted sentences according to the existence probability recorded in the polluted data statistical template, generating a polluted dataset; polluting the original dataset using the polluted sentences and the polluted dataset, and storing the pollution locations to generate a protected dataset.
Owner:CHINA MOBILE (XIONGAN) ICT CO LTD +4

Vertical fin gate transistor and memory device

PendingCN121665554ABit lineMemory cell
Embodiments of the present disclosure provide a memory device and a vertical fin gate transistor. A memory device includes a memory cell, wherein the memory cell includes a bit line, a word line over the bit line, and a semiconductor substrate on the bit line. The word line includes a plurality of fin-type words, a fin connector connecting the fin-type word line, and a common word line on the fin connector. The memory cell also includes a body line physically contacting one of the sidewalls of the semiconductor substrate and an insulating layer covering the body line, wherein the remaining sidewalls of the semiconductor substrate are partially covered by the fin word lines. The body line is grounded to direct accumulated charge out of the semiconductor substrate, thereby reducing floating body effects in the memory cell.
Owner:NAN YA TECH

Large language model data set protection method and device based on pollution lexical elements, electronic equipment, readable storage medium and computer program product

The invention discloses a large language model data set protection method and device based on polluted lemnes, electronic equipment, a readable storage medium and a computer program product, and belongs to the technical field of artificial intelligence. The method comprises the steps of performing statistics on data items of each piece of data in an original data set to generate a pollution data statistics template, and obtaining high-frequency common words and grammatical structure characteristics of the original data set; the polluted lexical element sequence comprises the high-frequency common words and polluted lexical elements, and the polluted lexical elements comprise self-created words which are generated by arranging the common characters and do not exist in reality and corresponding part of speech; generating a pollution statement by using the pollution lexical element sequence according to grammatical structure characteristics of the original data set; according to the existence probability recorded in the pollution data statistics template, filling the pollution statement into a data item corresponding to each piece of data in the original data set to generate a pollution data set; and polluting the original data set by using the pollution statement and the pollution data set, storing the pollution position, and generating a protection data set.
Owner:CHINA MOBILE (XIONGAN) ICT CO LTD +4

An entity matching method, system, device and medium based on word frequency

The application relates to an entity matching method, system, device and medium based on word frequency, wherein the method comprises the following steps: performing word segmentation on aliases in a plurality of entity data to obtain a first word segmentation list and an alias word set, counting word frequency data of words in the alias word set, removing city words existing in the first word segmentation list to obtain a second word segmentation list, judging whether the words in the second word segmentation list are common words according to the word frequency data, processing the words in the second word segmentation list according to the judgment result, obtaining alias keywords corresponding to the aliases, and querying whether the alias keywords exist in a text corpus; if the alias keywords exist, an entity corresponding to the alias keywords is recognized in the text corpus. Through the application, the problems of missed recognition caused by alias simplification and alias synonym misjudgment are solved. The alias keywords are obtained based on the word frequency of the words in the aliases, entity matching is more accurate, and the accuracy of entity matching is improved.
Owner:火石创造科技有限公司

Interview conversation audio processing method and device, electronic equipment and storage medium

The application provides an interview conversation audio processing method and device, electronic equipment and a storage medium, and relates to the technical field of AI interview. In the method, a pre-trained language model is used to extract first words and second words from a resume of an interview candidate and a standard answer of an interview question, wherein the first words and the second words include common words and non-common words related to a post to which the interview candidate applies; an intervention hot word table is generated according to the first words and the second words; and text conversion processing is performed on the interview conversation audio based on the intervention hot word table to obtain converted text corresponding to the interview conversation audio, so that when the interview candidate uses some professional terms or industry terms, the recognition error is reduced, and the accuracy of the obtained converted text is improved.
Owner:BEIJING NIUKE TECH CO LTD

Text similarity calculation method based on TF-IDF and word vectors

The invention relates to a TF-IDF and word vector-based text similarity calculation method. The method comprises the following steps of: 1, acquiring a first text and a second text to be compared, and performing word segmentation processing to form a first word set and a second word set; 2, calculating the TF-IDF value of each word in the word set by using a preset TF-IDF model, and solving a word vector by using a preset word vector model; 3, finding out a public word set and left and right adjacent word sets related to each public word through a semantic matching algorithm; and 4, inputting the TF-IDF values of the common words and the adjacent words of the common words into a first similarity model to calculate a first similarity, inputting the TF-IDF values of the word set, the word vectors, the common words and the like into a second similarity model to obtain a second similarity, and finally inputting the TF-IDF values of the word set, the word vectors, the common words and the like into a fusion model to obtain a text similarity. According to the method, multi-dimensional information is integrated, the text similarity degree is accurately measured, and the text processing efficiency and precision are effectively improved.
Owner:GUANGZHOU CITY UNIV OF TECH

Example sentence-driven machine translation result selecting apparatus and method thereof

To improve a filtering-type example sentence-driven analysis unit forming one input sentence analysis tree with using, as one input, an OR tree obtained by expressing a plurality of analysis trees obtained through analysis of an input sentence with use of grammar rules in the form of one tree structure and also with using, as the other input, a temporarily exclusive tree obtained by the example sentence or an exclusive three pre-equipped by the apparatus.SOLUTION: In a common word sequence acquisition between an input sentence and an example sentence, a method according to the present invention includes the steps of: acquiring a plurality of common word sequences; designating the minimum number of common words; determining a common word start position in an example sentence tree; and replacing the common word with a common part of speech.SELECTED DRAWING: Figure 1
Owner:榊 博史 +1

A Large-Scale Watermarking Method Based on Red-Green Tables

This invention discloses a large-scale watermarking method based on an improved red-green table, belonging to the field of large-scale security. Specifically, it first pre-divides commonly used words into a red list according to their frequency of occurrence, and includes commonly used standardized Chinese characters and uncommon characters not included in the list in a static exemption list. Then, when the user inputs a prompt, the large model generates an initial dynamic red list for token 0, and simultaneously generates candidate tokens for token 0. The final token 0 is then selected based on the dynamic red list. This process continues, generating the next dynamic red list for token 1 based on the final token 0, and selecting the final token 1. This continues until all user-input prompts have generated final tokens, forming a complete text with a watermark. When the user inputs text Q to be verified, the corresponding dynamic red list is reconstructed based on the initial key, and the frequency of tokens belonging to the dynamic red list in text Q is counted. If the actual frequency is <15% and z-score ≥2.58, the complete text Q is determined to contain a watermark. This invention improves the simplicity and security of large-scale watermarking.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Information retrieval system

To provide an information retrieval system configured to enable highly accurate information retrieval.SOLUTION: An information retrieval system is configured to: acquire a post text posted on a network, and analyze, for each information to be retrieved, the post text posted for the information, to extract words included in the post text; identify the frequency of the extracted words appearing in the post text and a co-occurrence relation between the extracted words, and set the words having the appearance frequency in the post text equal to or higher than a threshold, as frequent words; and set related words, which are words related to the frequent words, based on the co-occurrence relation. When retrieving information, the system retrieves information with the frequent words or related words set thereto, which matches a search word input by a user, and retrieves information while applying heavier weight to the frequent words than the related words.SELECTED DRAWING: Figure 13
Owner:AISIN CORP

PUF-based obfuscation scheme for in-memory architecture

It is proposed an in-memory computing (IMC) circuit (116), comprising:-an array comprising a matrix of memory cells (112) having a plurality of rows and columns wherein the memory cells (112) within the same column are connected by a common bit line and the memory cells within the same row are connected by a common word line, wherein the memory unit (112) is configured for storing weights of a trained neural network architecture, wherein the order of the weights is pre-disorganized; -at least one decoder (126) configured for outputting, using the key, a number of shift operations to be performed by each of the plurality of shift registers (128); -the plurality of shift registers (128) configured to shift an output of the array according to an output of the decoder (126).
Owner:ROBERT BOSCH GMBH

Stacked SRAM with shared wordline connection

Stacked static random-access memory (SRAM) circuits have doubled word length for a given SRAM cell area. An integrated circuit (IC) die includes stacked SRAM cells in vertically adjacent device layers with access transistors connected to a common wordline. The IC die with stacked SRAM cells having a common word line may be attached to a substrate and coupled to a power supply and, advantageously, to an active-cooling structure. SRAM cells may be formed in vertically adjacent layers of a substrate and electrically connected at their access transistor gate electrodes.
Owner:INTEL CORP

Three-dimensional memory device and preparation method thereof, and electronic equipment

The invention provides a three-dimensional memory device, a preparation method thereof and electronic equipment. The three-dimensional memory device comprises a substrate; a plurality of memory cell layers and a plurality of isolation layers; each storage unit layer at least comprises a plurality of storage unit groups, each storage unit group at least comprises a common word line and a plurality of storage units, each storage unit comprises a transistor and a capacitor, and the transistor is a surrounding gate transistor. The three-dimensional stacking technology is introduced, the memory units are stacked in the vertical direction to improve the overall density of the memory, meanwhile, the indium gallium zinc oxide with the wide-band-gap characteristic is used as a channel material to reduce transistor leakage current, so that the data retention time is prolonged, in addition, the structures of the memory units are optimized, and the memory performance is improved. The problem that capacitance signals of a bottom layer device are difficult to read is solved by utilizing a vertical bit line design, meanwhile, the grid control capability is enhanced by combining a surrounding grid transistor structure, the performance of a transistor device is improved, and performance optimization and stability improvement of a memory device are realized.
Owner:INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD

Semiconductor memory array, semiconductor memory and data processing method

The invention discloses a semiconductor storage array, a semiconductor memory and a data processing method, and the semiconductor storage array comprises M rows and N columns of storage units, a power supply voltage line, M common word lines, N first selection gate lines and N bit lines, the m row and n column storage unit comprises a control sub-circuit and a storage sub-circuit; and the control sub-circuit is electrically connected with the mth common word line, the nth first selection gate line, the nth bit line and the connection node respectively, and is configured to store a signal of the nth bit line in the connection node during writing under the control of the signal lines of the mth common word line and the nth first selection gate line, or store the signal of the nth bit line in the connection node during reading. Reading a signal of the connection node to the nth bit line; and a storage sub-circuit electrically connected to the connection node and the power supply voltage line, respectively, and configured to store a voltage difference between signals of the connection node and the power supply voltage line.
Owner:BEIJING SUPERSTRING ACAD OF MEMORY TECH

Semiconductor device and data storage system including the same

A semiconductor device includes a first structure having first and second memory regions, an extension region therebetween, and word lines; and a second structure having a circuit region overlapping the extension region. The word lines include first and second common word lines at different levels, and first and second intermediate individual word lines at a same level and spaced apart. Each of the first and second common word lines are in the first and second memory regions and the extension region. The first intermediate individual word line is in the first memory region and extends into the extension region at a level between the first and second common word lines. The second intermediate individual word line is in the second memory region and extends into the extension region. The circuit region includes pass transistors connected to the word lines. A pass transistor overlaps the word lines in the extension region.
Owner:SAMSUNG ELECTRONICS CO LTD

Automatic extraction of semantically similar question topics

A method, computer system, and computer program product are provided for automatically extracting semantically-similar question topics. A set of documents, wherein each document includes a plurality of words. One or more clusters of documents are identified in the set of documents based on a presence of common words in the documents of the one or more clusters. The one or more clusters are adjusted based on semantic similarity by adding or removing one or more documents from the one or more clusters. A topic is extracted from each adjusted cluster of documents.
Owner:CISCO TECHNOLOGY INC

Memory and access method therefor, and electronic device

PCT designated stageWO2026000630A9Digital storageBit lineMemory cell
A memory and an access method therefor, and an electronic device. The memory comprises at least one memory array, a plurality of common word lines, and a plurality of common bit lines. The memory array comprises a plurality of vertically stacked memory cell arrays and a plurality of vertical local bit lines. Each memory cell array comprises a plurality of local word lines extending in a second direction. The plurality of local word lines in a same memory cell array are divided into a plurality of word line groups, and different word line groups are connected to different common word lines. Each word line group comprises two local word lines connected to memory cells that are spaced one column apart from the word line group, each row of local bit lines corresponds to four common bit lines, and the common bit lines connected to the local bit lines connected to the memory cells connected to the local word lines in a same word line group are not adjacent.
Owner:BEIJING SUPERSTRING ACAD OF MEMORY TECH +1

PUF-based obfuscation scheme for in-memory architectures

An in-memory computation (IMC) circuit. The IMC circuit includes: an array comprising a matrix of memory cells having a plurality of rows and columns, wherein memory cells within the same column are connected through a common bit line and memory cells within the same row are connected through a common word line, wherein the memory cells are configured for storing weights of a trained neural network architecture, wherein the order of the weights is pre-scrambled; at least one decoder configured for outputting a number of shifting operations to be performed by each shifting register of a plurality of shifting registers using a secret key; the plurality of shifting registers configured for shifting outputs of the array according to the output of the decoder.
Owner:ROBERT BOSCH GMBH

Memory and access control method therefor, and electronic device

PCT designated stageWO2026065684A1Digital storageMemory cellCommon word
A memory and an access control method therefor, and an electronic device. The memory comprises: a plurality of memory arrays distributed on a substrate in the direction parallel to the substrate, a plurality of first word line gating sub-circuits (21), and a plurality of first word line gating control lines (HB_MAT_S); each memory array comprises a plurality of layers of memory cell arrays and a plurality of common word lines (CWLs), each memory cell array comprises a plurality of word lines extending in parallel, and the CWLs are connected to at least one word line; each CWL is connected to a corresponding word line driving terminal (HB_SWD) by means of the corresponding first word line gating sub-circuit (21); and a same word line driving terminal (HB_SWD) is separately connected to at least two CWLs of different memory arrays, and each first word line gating sub-circuit (21) is configured to electrically connect the corresponding word line driving terminal (HB_SWD) to the corresponding CWL or disconnect the corresponding word line driving terminal (HB_SWD) from the corresponding CWL under the control of the corresponding first word line gating control line (HB_MAT_S), wherein different CWLs of a same memory array are connected to different word line driving terminals (HB_SWD).
Owner:BEIJING SUPERSTRING ACAD OF MEMORY TECH

A Chinese spelling error correction method

This invention discloses a Chinese spelling error correction method. The method increases training difficulty by replacing common words and characters in the original dataset with easily confused words and characters, thereby increasing the accuracy of model predictions. Simultaneously, the error correction model provided by this invention is optimized by integrating glyph and pinyin information based on the Macbert model, resulting in a more accurate error detection score and thus a more accurate identification of the corrected character. When the error correction model cannot obtain a suitable corrected character for a specific position in a sentence, this invention introduces a binary classification model to predict each candidate character at that specific position to obtain the corrected character, thereby enabling more accurate and efficient error correction of Chinese spelling.
Owner:ZHEJIANG UNIV

Three-dimensional memory devices

The present invention provides a three-dimensional memory device, such as a three-dimensional AND flash memory device. The three-dimensional memory device includes a plurality of word lines, a plurality of first switches, a plurality of second switches, and N conductive layers, where N is a positive integer greater than 1. The word lines are divided into a plurality of word line groups. The first switches receive a common word line voltage. The second switches receive a reference ground voltage. The first word line group connects to the first conductive layer via the second conductive layer. The i-th word line group, in turn, connects to the first conductive layer via the i+1-th conductive layer to the second conductive layer, where N>i>1.
Owner:MACRONIX INTERNATIONAL CO LTD

Semiconductor memory device and refresh control method thereof

The invention provides a semiconductor memory device and a refresh control method thereof to suppress unnecessary power consumption. The semiconductor memory device includes: a memory cell array having a plurality of segments, each segment having a plurality of normal word lines and a plurality of redundant word lines; and a row redundancy circuit configured to activate a non-defective one of the normal word lines and a non-defective one of the redundant word lines during a refresh operation; wherein activation and replacement of defective ones in the common word lines and defective ones in the redundant word lines are not performed in each section.
Owner:WINBOND ELECTRONICS CORP

Cross-point architecture for pcram

A cell array of a memory device includes: a first deck of memory cells arranged in a first row and a second row and a plurality of columns, wherein the memory cells in the second row in the first deck is displaced in the first horizontal direction with respect to the memory cells in the first row in the first deck; a first common word line metal track; and a plurality of first bit line metal tracks, wherein the plurality of first bit line metal tracks comprises a first group of first bit line metal tracks and a second group of first bit line metal tracks interposed in the first horizontal direction, and each of the first group is disposed on only one of the memory cells in the first row, and each of the second group is disposed on only one of the memory cells in the second row.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Vertical fin-gate transistor and memory device

PendingUS20260068141A1Bit lineMemory cell
Embodiments of the present disclosure provide a memory device and a vertical fin-gate transistor. The memory device includes a memory cell including a bit line, a word line above the bit line, and a semiconductor substrate on the bit line. The word line includes a plurality of fin type word lines and a common word line connecting the fin type word lines, where the fin type word lines partially cover a first sidewall of the semiconductor substrate. The memory cell also includes a body line physically contacting a second sidewall of the semiconductor substrate opposite to the first sidewall of the semiconductor substrate, and an insulating layer embedding the body line. The body line is grounded to direct the accumulated charges out of the semiconductor substrate, thereby reducing the floating body effect in the memory cell.
Owner:NAN YA TECH

Intelligent place name and address translation method based on multilingual syllable segmentation

The present application relates to the field of intelligent translation technology, and specifically discloses a method for intelligent place name and address translation based on multilingual syllable segmentation, which introduces a deep learning algorithm to make adaptive translation strategy decisions for each vocabulary unit in the source language place name and address text, so as to automatically distinguish between proper nouns that should be transliterated and common vocabulary that should be translated using a dictionary, and performs fine-grained syllable-level segmentation and target language syllable mapping on proper nouns marked as transliterated to achieve accurate transliteration of proper nouns; at the same time, for common vocabulary marked as dictionary translation, a dictionary matching algorithm is used for accurate translation, and then the transliteration results of proper nouns and the translation results of common vocabulary are standardized and reorganized to obtain the place name and address translation results. The present application can effectively improve the translation effect of proper place names when mixed with common words, and enhance the translation accuracy and flexibility of place names and addresses in a multilingual environment.
Owner:SHAANXI TIRAIN TECH CO LTD

Sensitive data identification method and device, electronic equipment, medium and program product

The application discloses a sensitive data identification method and device, electronic equipment, medium and program product, and relates to the technical field of data security. The sensitive data identification method comprises the following steps: performing sentence segmentation on to-be-identified text data to obtain a plurality of short sentences; counting specific words in each short sentence which are not in a preset common word library, and recording the word frequency in a corresponding time residence matrix; performing word segmentation on each short sentence according to a preset word segmentation rule, determining a plurality of word segmentation paths corresponding to each short sentence; calculating the joint probability corresponding to each word segmentation path based on the word frequency of each time residence matrix, taking the word segmentation path with the highest joint probability as a target word segmentation path, and performing sensitive word retrieval on each word segmentation to obtain a sensitive data identification result. The technical scheme of the application solves the problem that the traditional sensitive word matching algorithm is prone to errors when processing Chinese text, thereby affecting the accuracy of sensitive data identification.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

Semiconductor device and manufacturing method thereof, electronic equipment and working method

The invention discloses a semiconductor device and a manufacturing method thereof, electronic equipment and a working method, and the semiconductor device comprises a plurality of storage units, and each storage unit comprises a first storage subunit and a second storage subunit which are distributed at different layers; a plurality of word lines, one column of storage units arranged along a second direction is respectively connected with two word lines, one word line is connected with one column of first storage subunits, and the other word line is connected with one column of second storage subunits; the plurality of common word lines extend along a second direction and are connected with the first storage subunits and the second storage subunits at the same time; wherein the word lines and the common word lines are alternately distributed in the direction perpendicular to the substrate, and the common word lines are located between the two word lines. According to the semiconductor device, common gate control of the storage subunits can be achieved through the common word line, the space in the vertical direction is saved, the stacking number of the storage units is increased, and the manufacturing cost is saved.
Owner:BEIJING SUPERSTRING ACAD OF MEMORY TECH