Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Delimiter" patented technology

A delimiter is a sequence of one or more characters for specifying the boundary between separate, independent regions in plain text or other data streams. An example of a delimiter is the comma character, which acts as a field delimiter in a sequence of comma-separated values. Another example of a delimiter is the time gap used to separate letters and words in the transmission of Morse code.

Data processing method and related device

PCT designated stageWO2026040635A1TransmissionAlgorithmThresholding
The present application relates to the field of communications, and provides a data processing method and a related device, which are capable of reducing data transmission bandwidth. The method comprises: a first apparatus acquiring first information, wherein the first information indicates to send a first sequence, the first sequence comprises an inter-frame gap, a first subsequence, and a start of frame delimiter, a length of the first sequence is less than a first threshold, a first byte of the first subsequence is used for indicating a start of a data sequence, and the data sequence comprises the first subsequence and the start of frame delimiter; and sending the sequence.
Owner:HUAWEI TECH CO LTD

Determining semantic and grammatical correctness of user-expanded sentence using integrated programmatic and specialized guided and constrained artificial intelligence

A system and method guide an Artificial Intelligence engine to determine the semantic and grammatical correctness of a user-expanded sentence in real-time. The sentence validation process involves receiving input from the user, the input includes sentence fragment that the user wishes to expand and user-expanded sentence that the user constructs on the fragment provided. The inputs are broken down into tokens. The word-level tokenization algorithm is used, which identifies tokens by splitting the text into spaces, punctuation marks, and other delimiters. Further, a token comparison algorithm is used to assess the relationship between the sentence fragment and the user-expanded sentence to analyze order and placement. Once the token comparison is complete, a prompt is generated using prompt generator to evaluate grammatical and semantic evaluation of the user-expanded sentence. Real-time feedback is provided to the user based on grammatical and semantic evaluation.
Owner:2HR LEARNING INC

Decoding device, decoding method, and program

A decoding device comprises a decoding processing unit that decodes a plurality of codewords included in an input codeword sequence while switching a codebook or an analysis tree that is to be used for decoding among a plurality of codebooks or analysis trees. The decoding processing unit determines a codebook or an analysis tree that is to be used for decoding a second codeword next to a first codeword included in the input codeword sequence in accordance with the first codeword and the codebook or the analysis tree used for decoding the first codeword. However, in cases where the first codeword is a delimiter codeword that is a specific codeword, the decoding processing unit determines the codebook or the analysis tree that is to be used for decoding the second codeword to be a specific codebook or analysis tree corresponding to the delimiter codeword regardless of the codebook or the analysis tree used for decoding the first codeword.
Owner:NT T INC

A syslog log automatic parsing method

This invention discloses an automatic Syslog log parsing method, comprising: collecting Syslog log messages; removing header information and retaining the message body content, parsing it into a string set, and determining a delimiter; using the delimiter to segment the Syslog log message body content into a string array; confirming whether a key exists in the string array; classifying the strings, if a key exists, classifying the logs according to the key and data length; if no key exists, checking the data type and classifying the logs according to the array length and data type; automatically generating parsing templates for different log categories, using the parsing templates to extract the value corresponding to each field, and performing subsequent field mapping or transformation normalization processing. Using this invention, regular expression parsing of Syslog logs can be generated quickly, saving significant manpower.
Owner:HUANENG LANCANG RIVER HYDROPOWER CO LTD +1

Sse-based streaming data processing display method and system, and related device

The application provides a kind of based on SSE's stream data processing display method, system and related equipment, method includes based on the preset parameter construction and initiates HTTP request, establishes the SSE connection with server;In the data stream of the SSE connection, continuously receive data segment and append to data buffer;The data buffer is scanned and identified pre-defined message separator, to extract one or more independent SSE message from the data buffer and store in message queue;From the message queue, the SSE message is obtained, and the SSE message is split into character sequence and stored in character queue;According to the preset timer, periodically take out single or multiple characters from the character queue, render the character to user interface.The application realizes the smooth typewriter display effect by constructing double-layer queue rendering technology, and improves the response performance of SSE stream data in complex business scenarios.
Owner:SHENZHEN MAIFENG TECH CO LTD

English long sentence analysis method and system based on hierarchical perception

PendingCN122334237AQuestion analysisSentence analysis
This invention discloses a hierarchical perception-based method and system for analyzing complex English sentences, relating to the interdisciplinary fields of natural language processing and intelligent education. By defining structural delimiters, sentences are divided into syntactic levels according to the number of effective delimiters, first eliminating invalid parallel structural delimiters; then, a syntactic encoding module completes text segmentation, hierarchical tagging, word embedding, and feature fusion; a hierarchical perception module calculates feature weights and generates target syntactic features; finally, a syntactic decoding module outputs the sentence's main body and modifier decomposition results. The hierarchical framework constructed by this invention fully covers the core modifiers of English, is suitable for the scenario of complex sentences in the College Entrance Examination (Gaokao), and can achieve accurate and lightweight decomposition of all types of complex sentences. It effectively solves the problems of irregular decomposition, poor generalization, and poor adaptability of traditional methods, and the analysis process is clear and teachable.
Owner:SHENZHEN TAITAIGE TECHNOLOGY CO LTD

End-to-end speech separation algorithm based on speech language model

PendingCN121583282ASpeech analysisAudio restorationVoice source
The invention discloses an end-to-end speech separation algorithm based on a speech language model, and the algorithm comprises the steps: discretizing a continuous audio into a 32-order discrete codebook sequence through a residual vector quantization coder-decoder, and introducing a transcription start symbol lt through an SOT strategy; sOSgt, SOSgt; a special separator is lt; sCgt; and a termination symbol lt; eOSgt, EOSgt; splicing a multi-person voice sequence; extracting audio depth features by using a pre-trained WavLM model, and guiding an autoregression decoder to output a separated zero-order codebook sequence in combination with a cross attention mechanism; predicting a high-order codebook sequence step by step through a non-autoregression model, configuring an independent embedding layer to fuse low-order information, and introducing a task embedding mechanism to optimize modeling; based on a special separator lt; sCgt; and slicing the multi-order discrete codebook sequence, and outputting an independent voice source through an Encodec decoder. According to the method, the intelligibility of voice separation and the audio restoration quality can be effectively improved, the decoding speed is high, the subjective hearing experiment result and the downstream task performance are excellent, the scene that the number of speakers is unknown is supported, and the industrialization application prospect is wide.
Owner:SHANGHAI JIAOTONG UNIV

Triple extraction method based on diffusion enhancement relation

The invention discloses a diffusion enhancement relation-based triple extraction method, which comprises the following steps of: 1) inputting a Chinese text into a BERT model, and converting the Chinese text into a corresponding index to obtain word embedding information; 2) inputting word embedding information into a triple diffusion model, and performing noise addition, feature fusion and de-noising processing to obtain a predicted triple; 3) inputting word embedding information into the character feature extraction network model, learning semantic information through a multi-layer perceptron and a double-affine attention mechanism in combination with relative position coding and a multi-head attention module, relieving a label imbalance problem by adopting a word embedding disturbance mechanism, and outputting a relation triple matrix; and 4) respectively decoding the triads obtained in the two steps, and taking union sets to obtain a final relation triad. According to the method, the boundary diffusion information and the character-level semantic information are fused, so that the problems that Chinese semantics are complex, natural separators are lacked, small-field data semantics extraction is insufficient, extraction of a traditional decoding method is incomplete, labels are unbalanced and the like are effectively solved, the accuracy and integrity of Chinese relation triple extraction and model stability are remarkably improved, and the method is suitable for large-scale popularization and application. The method is suitable for various natural language processing application scenes such as information questions and answers and search engine optimization.
Owner:ZHEJIANG UNIV OF TECH

Log compression method and device based on structured marking and hybrid coding

The invention discloses a log compression method and device based on structured marking and hybrid coding, and solves the technical problem that the compression rate is obviously low due to an existing log compression method. The method comprises the following steps: acquiring an original log message, and preprocessing the original log message by adopting a predefined separator set and two types of regular expressions to generate a dynamic tag sequence and a static tag sequence; classifying the dynamic label sequence, and outputting a plurality of structured labels and a plurality of unstructured labels; based on the plurality of structured markers, generating a refined skeleton and a simplified sub-marker matrix; generating a dictionary file and a binary coded data file according to the static mark sequence, the plurality of unstructured marks, the refined skeleton and the simplified sub-mark matrix; and compressing the dictionary file and the binary coded data file based on a compression algorithm to generate a complete compressed log file.
Owner:SUN YAT SEN UNIV

URL-based API asset merging method and system

The application provides a URL-based API asset merging method and system, which comprises the following steps: obtaining host access record log data; based on the access record log data, extracting a URL, using a specific separator to split the URI, and forming a word set array; according to the word set data, sequentially taking two words as a relationship, taking the words as nodes, constructing a word relationship graph, and calculating the out-degree and in-degree of the words; according to the out-degree and in-degree obtained in the foregoing step, selecting words with out-degree or in-degree less than a specified threshold value by specifying the threshold value; based on the words obtained in the foregoing step, screening out URIs containing the corresponding words, and calculating the URIs with high similarity in the full-amount URIs by using minhash; based on the URIs selected in the foregoing step, calculating the support degree of the selected corresponding words, and replacing the words with a wildcard string if the support degree is less than a specified threshold value, so as to realize API normalization. The application solves the technical problems of a large API asset list and high dependence on manual operation experience.
Owner:SHANGHAI GUAN AN INFORMATION TECH

Decoder-only extractive schema linking for text-to-sql

Extractive schema linking includes generating a tokenized schema from an SQL schema, a tokenized natural language question, and tokenized candidates. The tokenized candidates are generated by tokenizing candidates from the SQL schema. Each tokenized candidate is formed by a string of tokens having a first token representing an initial delimiter and a last token representing an end delimiter. Vectorial representations of the tokenized schema, the tokenized natural language question, and tokenized candidates are generated, and transformed vectorial representations generated by processing the vectorial representations through a decoder-only model. Concatenated vectors are generated from the transformed vectorial representations of the tokenized candidates, the concatenated vectors generated by concatenating a first transformed vectorial representation corresponding to the first token with a last transformed vectorial representation corresponding to the last token of each tokenized candidate. Quantitative relevancies of the plurality of candidates from the SQL schema are generated based on the concatenated vectors.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Self-programming system

Methods in a system for programming in natural language, comprising the steps - Machine capture and analysis (S100) of digital input information, which includes a plurality of input elements in natural language, - Determine (S200), by the system, based on analyzing the input information, an elementary data type of an input element of the input information, and - Generating (S300) source code for the elementary data type, the step of which includes capturing the input information: - Dividing the input information into sentences, where a sentence comprises a sequence of consecutive input elements in the input information and ends with a separator; and where the step of determining (S200) the elementary data type comprises the following substeps: - Recognizing a first set comprising a first sequence of input elements and a second set comprising a second sequence of input elements, wherein the first sequence and the second sequence differ only at one position, wherein the first sequence includes a first difference input element at this position and the second sequence includes a second difference input element at this position, - Determining the elementary data type for the difference input elements; assigning the specified elementary data type to each of the difference input elements.
Owner:AMRI MASOUD

Use of access unit delimiters and adaptive parameter sets

To provide a decoder and a method for deriving a necessary parameter from an access unit.SOLUTION: A video decoder 20 comprises an in-loop filter 90 which filters a reconfiguration version of a decoded picture, and a parameterizing unit. The parameterizing unit reads in-loop filter control information for parameterizing the in-loop filter from a parameter set located in an access unit of a decoded picture following a video encoding unit along the sequence of a data stream 14 and / or a part of a video encoding unit following data included in a video encoding unit carrying block-based prediction parameter data and predicted residue data along the sequence of the data stream, and also parameterizes the in-loop filter so as to filter the reconfiguration version of the encoded picture by a method corresponding to the in-loop filter control information.SELECTED DRAWING: Figure 2
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Small sample event extraction method based on T5 and knowledge graph data enhancement

The invention relates to the technical field of natural language processing, and discloses a T5 and knowledge graph data enhancement-based small sample event extraction method, which comprises the following steps of: segmenting an original text according to a preset segmentation bound character, and screening a longest fragment as a covering object; performing prediction generation and backfilling on the covering object by utilizing the T5 model after fine adjustment to obtain a preliminary enhanced text; performing similar entity replacement on nouns or pronouns except the key information in the preliminarily enhanced text based on a Chinese concept knowledge graph to obtain a new sample; and combining the new sample with the original text to form an enhanced data set, and driving the event extraction model to train. According to the method, text distribution consistency is kept by adopting generative fine tuning, semantic diversity is increased by combining knowledge graph replacement, and meanwhile, tagging information is directly inherited by protecting trigger words and arguments, so that the accuracy and generalization of event extraction under a small sample condition are effectively improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

String decoding method and device, electronic equipment and storage medium

Embodiments of the present application disclose a string decoding method and device, electronic equipment and a storage medium. The method comprises: in the case of receiving a to-be-decoded string parameter, traversing each byte in the to-be-decoded string parameter; when any byte is traversed, reading the code value of the byte, judging whether the byte is a high-position Chinese byte based on the code value of the byte, if it is determined that the byte is a high-position Chinese byte, offsetting the traversal pointer by a first preset offset value; if the byte is not a high-position Chinese byte, judging whether the byte is a byte storing a delimiter, if yes, determining the position of the byte in the to-be-decoded string parameter, and offsetting the traversal pointer by a second preset offset value; in the case of traversing a byte storing a preset end character, the to-be-decoded string parameter traversal ends, and the to-be-decoded string parameter is cut into to-be-decoded character units according to the determined position of the preset delimiter; and decoding each to-be-decoded character unit based on a preset decoding rule.
Owner:SI-TECH INFORMATION TECH CO LTD

Method and system for document chunking

A method and system for chunking a document are provided. The method according to some embodiments may include chunking a document including a plurality of sentences into a plurality of section chunks based on sentences including a section delimiter, determining whether a size of each of the plurality of section chunks exceeds a preset first threshold and chunking a section chunk, a size of the section chunk among the plurality of section chunks exceeds the preset first threshold, into a plurality of sub-chunks based on whether a similarity between sentences included in the section chunk is equal to or greater than a second threshold. The second threshold may be determined based on a similarity distribution of a query for the document, calculated by comparing a query generated from the document using a generative model with the document.
Owner:SAMSUNG SDS CO LTD

Delimiter insertion device and speech recognition system

A delimiter insertion device includes an inter-word time acquisition unit that acquires an inter-word time, which is the length of time until a next word is spoken for each word included in the uttered speech; and a delimiter insertion unit that inserts a delimiter into a target text, which is a text obtained by speech recognition of the uttered speech, based on a delimiter insertion model and the inter-word time. The delimiter insertion model outputs delimiter prediction information indicating a delimiter in response to the input of a delimiter-removed sentence. The delimiter insertion unit inserts a delimiter into the target text based on delimiter prediction information obtained by inputting the target text to the delimiter insertion model.
Owner:NTT DOCOMO INC