Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

16 results about "Unicode" patented technology

Unicode is a computing industry standard for the consistent encoding, representation, and handling of text expressed in most of the world's writing systems. The standard is maintained by the Unicode Consortium, and as of May 2019 the most recent version, Unicode 12.1, contains a repertoire of 137,994 characters covering 150 modern and historic scripts, as well as multiple symbol sets and emoji. The character repertoire of the Unicode Standard is synchronized with ISO/IEC 10646, and both are code-for-code identical.

Intelligent ai routing advisory platform with synthetic injection testing, bias detection digital twin, zero-copy pipeline, cryptographic compliance verification, and autonomous multi-tier coordination for heterogeneous ai provider ecosystems

A computer-implemented system for routing artificial intelligence (AI) queries. The system utilizes a zero-copy data pipeline, which processes prompts in memory-mapped buffers to eliminate at least one memory copy operation, thereby reducing latency relative to conventional serialization pipelines. The system continuously verifies AI provider compliance by injecting synthetic prompts containing invisible, Ed25519-signed Unicode watermarks. Algorithmic bias is detected by generating counterfactual “digital twin” prompts and applying Fisher exact statistical testing.Routing decisions for multi-tier autonomous systems are governed by safety-level requirements (ASIL-D, ASIL-B, QM) and may be constrained by external routing directives received via a meta-identifier. A hash-chained manifest, cryptographically signed using Ed25519 and consumed by downstream gateways, is generated for each routing decision, with its Merkle root asynchronously anchored to a blockchain to create a tamper-evident audit trail for regulatory compliance.
Owner:WEBER AXEL

Character marking and identifying method and device, equipment and storage medium

The invention discloses a character marking and identifying method and device, equipment and a storage medium, and the method comprises the steps: obtaining the text content of a current text input by a user, mapping the mantissa feature of each Unicode point value in the text content to a private area, generating a corresponding private area code point, and carrying out the recognition of the Unicode point value in the private area; the private area comprises a basic multi-language plane private area, a private area A and a private area B; inserting the private area code points into the text content as zero-width characters to obtain updated target text content; the original mantissa features of the code points of the private area are compared with the actual mantissa features of the target text content, and the current text is judged to be manually input or other source content including AI generation according to the comparison result, so that the reliability and accuracy of text identification can be remarkably improved; the method does not need to depend on a large language model for complex analysis, greatly reduces the technical threshold and detection cost, and has good cross-platform compatibility and practical application value.
Owner:蒋励勤

Method and system for generating a unified script code (USC) representation for multilingual text processing

PCT designated stageWO2026074589A1Natural language translationSyllableDevanagari
The present invention provides a computer-implemented method and system for generating a Unified Script Code (USC) representation of multilingual text to enable consistent, script-neutral, and phonemically accurate encoding across diverse languages and writing systems The system converts Unicode-encoded text from one or more Indic scripts into a Devanagari-based intermediate form using predefined or bitwise mapping. It then normalizes the text by inserting inherent vowels, converting dependent vowel signs (matras) into independent vowels, and removing halant characters. The resulting USC representation explicitly encodes consonant-vowel sequences, reduces script specific variation, and preserves phonetic integrity. This approach improves tokenization efficiency, enhances performance in natural language processing and machine learning tasks, and supports reversible conversion to original scripts.
Owner:SINGH NAU NIHAL

An embedded device font library dynamic replacement system and method based on Unicode and GB2312 encoding mapping

The application discloses a kind of embedded device font library dynamic replacement system and method based on Unicode and GB2312 encoding mapping, belong to embedded system character coding and font library management technical field.Its method is by constructing lightweight Unicode-GB2312 bidirectional index mapping table, realizes Unicode character on-demand loading and mixed rendering under the premise of reserving original font library.Hash table is used to accelerate query, virtual area code extension and incremental updating mechanism, so that storage occupancy is reduced by more than 95%, font library updating efficiency is improved by 100 times, and progressive adaptation of full Unicode character set is supported.The application can be widely applied to intelligent terminal, industrial equipment and other scenes, and solves the contradiction between multi-language compatibility and resource limitation.
Owner:NANJING GUODIAN NANZI POWER GRID AUTOMATION CO LTD

Encryption method, decryption method, leakage identification method and device

ActiveCN114741709BPlay a protective effectDoes not change document encodingInternet privacyEngineering
This application provides an encryption method, decryption method, leakage identification method, text transmission method, and apparatus, including: obtaining a string to be encrypted; determining a protection character in the string to be encrypted based on a preset association between the original code position of the protection character in the Unicode Conformity Table (Unicode Conformity Table) and the multiplexed code position in the Unicode Conformity Table; and replacing the original code position of the protection character with a target multiplexed code position corresponding to the original code position according to the association, thereby obtaining an encrypted string. The encryption technology described in this application is based on the text itself. The encrypted text will be displayed in normal plaintext form in a trusted environment, while it will be presented as garbled text in an untrusted environment. Even if the encrypted text is leaked, it can still protect the text content. The decrypted protection character can be displayed and edited in normal plaintext form in a trusted environment, achieving the purpose of normal use of the text content.
Owner:ALIBABA (CHINA) CO LTD

Systems and methods for detecting typographical errors in domain name entries

Systems and methods for detecting a typographical error in a domain name, including: receiving a domain name comprising Unicode characters; encoding each character, where the encoding includes: computing an integer index of the Unicode characters; converting the integer index into a binary representation; and multiplying the binary representation by a dense matrix to obtain a floating-point vector; and comparing the floating-point vector to a reference floating-point vector of a known domain name using a model to determine if the domain name contains the typographical error.
Owner:DNSFILTER INC

Multi-type HTTP (Hyper Text Transport Protocol) coding feature extraction method and system and medium

The invention discloses a multi-type HTTP (Hyper Text Transport Protocol) coding feature extraction method and system and a medium. The method comprises the following steps: constructing a multi-type HTTP coding data set; sample data of URL codes, Base64 codes, Unicode codes, Hex codes, HTML codes and combined codes and nested codes of the URL codes, the Base64 codes, the Unicode codes, the Hex codes and the HTML codes are collected respectively, and corresponding type labels are marked; the method comprises the following steps: reading a multi-type HTTP coded data set, sequentially reading samples, and carrying out feature extraction; constructing a multi-type coding feature matrix; constructing a coding sample feature vector, constructing and standardizing a feature matrix of the multi-type HTTP coding data set, constructing a coding type vector matrix according to type values in the multi-type HTTP coding data set, and performing transverse splicing operation on the coding type vector matrix and the standardized feature matrix to construct a multi-type coding feature matrix; according to the method, a multi-type coding data set which can be used for coding learning training, verification and testing is constructed, feature extraction is performed through unique features of multi-type coding, and the accuracy and efficiency of feature extraction are improved.
Owner:CHANGSHU INSTITUTE OF TECHNOLOGY

Pinyin supplementation labeling system and method for distinguishing six tones of'machine / product, odd / seven and sparse / west 'in Chinese pronunciation

The invention relates to the field of language and character information processing, in particular to a supplementation labeling system for uniquely distinguishing three groups of homomorphic abnormal sound phonemes of'machine / product, odd / seven and sparse / west 'through machine-readable symbols in a Chinese pinyin system, an input method, electronic equipment and a computer readable storage medium, and aims to provide a pinyin supplementation labeling system. On the premise of keeping the original spelling of the Pinyin scheme unchanged, the pronunciation of the product, the pronunciation of the seven and the pronunciation of the west are uniquely distinguished by using the recorded symbols (. J.q.x) of the Unicode, namely the prefix points and the original letters, so that the blind test of 100 thousand news corpora can be achieved, and after the method is adopted, the TTS misreading rate is reduced from 2.8% to 0.07%; the ambiguity of the ASR post-processing homomorphic and abnormal sound field is reduced by 62%, and the recognition accuracy of the whole sentence is improved by 3.4%; in a class test of Chinese as a foreign language, the correct rate of distinguishing initial consonants of learners is improved by 28%, and class hours are shortened by 30%; the symbol occupies 2 bytes, and the expansion rate is stored as lt; 0.5%, almost negligible; the method is completely compatible with Unicode, GB 18030 and ISO 10646, and a new character does not need to be made.
Owner:霍立远

Systems and methods for dynamically generating and mapping unicode for cross-distributed network data transmissions

ActiveUS12549196B2Code conversionSchema mappingNetwork data
Systems, computer program products, and methods are described herein for generating and mapping unicode for cross-distributed network data transmissions. The present disclosure is configured to identify a cross-distributed network data transmission comprising source ledger code; generate a core schema for a destination ledger, wherein the destination ledger is based on data of the cross-distributed network data transmission; generate, by a ledger transformation module, a data dictionary, wherein the ledger transformation module receives ledger code from a large language model (LLM); generate code mapping instructions for the cross-distributed network data transmission based on the core schema for the destination ledger; convert the source ledger code to a unicode; receive, by a receiver schema mapping module, the unicode and the code mapping instructions; and generate, by the receiver schema mapping module, an intermediary ledger comprising the unicode and based on the code mapping instructions.
Owner:BANK OF AMERICA CORP

A method, device, computer storage medium and terminal for realizing character recognition

A method, device, computer storage medium and terminal for realizing character recognition, the embodiment of the present application acquires a character image containing a preset number of characters based on a PDF document, and performs multi-word recognition on the character image; in the case that the result of multi-word recognition contains a preset number of characters, the final recognition result of the characters contained in the PDF document is determined according to the result of multi-word recognition. The embodiment of the present application can read and render the characters in the PDF file, can determine which character in the PDF the character image corresponds to, in the case that the result of multi-word recognition contains a preset number of characters, the quick confirmation of missed detection or multiple detection is realized, and the accuracy of character recognition is improved according to the result of multi-word recognition; the recognized result is converted into a uniform code (Unicode) and attached to the corresponding character of the PDF, and the accurate recognition of the characters contained in the PDF document is realized.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +1

Regular expression filter for Unicode transform format strings

The invention is entitled "Regular Expression Filter for Unicode Transformation Format Strings". The invention provides an apparatus comprising a memory system having a memory die, the memory system comprising circuitry configured to receive a plurality of bytes at a first input terminal, the plurality of bytes comprising an input stream of a plurality of first variable length encoded symbols, and the circuitry configured to determine a sequence of symbol byte lengths for each of the plurality of first variable length encoded symbols.
Owner:SANDISK TECHNOLOGIES LLC

Unicode confusion command attack detection system and method based on dual-granularity feature coupling model

The invention discloses a Unicode confusion command attack detection system and method based on a dual-granularity feature coupling model, and belongs to the technical field of network security. In order to solve the problem of accurate recognition of confused malicious commands, a data preprocessing module, a standardization module, a character-level confusion feature extraction module, a word-level semantic feature extraction module, a dual-granularity feature coupling module and a malicious command judgment module are connected in sequence; the data preprocessing module is used for cleaning command line data and executing command line text standardization operation; the standardization module is used for replacing a specific mode with a universal placeholder and performing word segmentation operation; the character-level confusion feature extraction module is used for extracting character-level features of command keywords; the word-level semantic feature extraction module is used for extracting word-level features; the dual-granularity feature coupling module is used for forming a mixed feature vector; and the malicious command judgment module is used for inputting the mixed feature vector into a pre-training random forest classifier to output a binary classification result.
Owner:HARBIN UNIV OF SCI & TECH

Detection of Machine Learning Model Attacks Obfuscated in Unicode

A prompt for a generative artificial intelligence (GenAI) model which contains unicode is received. The prompt is then tokenized to result in a plurality of tokens. Token forming part of a repeating sequence are identified and then removed to result in a modified set of tokens. The modified set of tokens are subsequently detokenized to result in a modified prompt. It is then determined, whether ingestion of the modified prompt by the GenAI model will result in the GenAI model behaving in an undesired manner. The modified prompt is passed to the GenAI model when it is determined that ingestion of the modified prompt will not result in the GenAI model behaving in an undesired manner. Otherwise, at least one remediation action is initiated when it is determined that ingestion of the modified prompt by the GenAI model will result in the GenAI model behaving in an undesired manner.
Owner:HIDDENLAYER INC

A multi-language display method, device and equipment for an automobile instrument and a storage medium

This application discloses a method, apparatus, device, and storage medium for multilingual display of automotive instrument panels. The technical solution provided in this application obtains multiple original Unicode encoding strings corresponding to the text to be displayed sent by the vehicle's host computer, as well as translation markers for the original Unicode encoding strings. Based on the translation markers, the original Unicode encoding strings are translated to obtain multiple target Unicode encoding strings. Multiple glyph data are determined from a preset character library based on the multiple target Unicode encoding strings, and the glyph data is rendered and displayed. This eliminates the need for complete sentence translation of the text to be displayed, effectively improving display accuracy in multilingual usage scenarios while maintaining automotive chip performance, thus enhancing the display effect of the automotive instrument panel.
Owner:GUANGZHOU ZHOULIGONG SCM DEV

Mixed language typesetting method, device and equipment for embedded equipment and medium

PendingCN121303055ANatural language data processingProgramming languageCharacter analysis
The invention discloses a mixed language typesetting method and device for embedded equipment, equipment and a medium, and the method comprises the steps: obtaining a to-be-typeset text stream of a mixed language, and dividing the to-be-typeset text stream into a plurality of text fields through a preset Unicode bidirectional algorithm to obtain a text field list; according to the input sequence of the to-be-typeset text stream, character-by-character analysis is conducted on the character fields in the character field list in sequence, so that a font bitmap corresponding to each character is obtained from a preset cache region, according to pre-generated typesetting parameters, a target embedded device is controlled to conduct typesetting rendering on the font bitmap of each character line by line, and typesetting rendering is conducted on the font bitmap of each character line by line. Unified typesetting of the text of the mixed language on the target embedded device is achieved; wherein the typesetting parameters are generated according to a preset alignment rule table; the alignment rule table comprises a baseline offset table and an equal-width proportion table. According to the method and the device, the real-time performance and the visual consistency of the mixed language typesetting of the embedded equipment can be improved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Systems and methods for detecting typographical errors in domain name entries

Systems and methods for detecting a typographical error in a domain name, including: receiving a domain name comprising Unicode characters; encoding each character, where the encoding includes: computing an integer index of the Unicode characters; converting the integer index into a binary representation; and multiplying the binary representation by a dense matrix to obtain a floating-point vector; and comparing the floating-point vector to a reference floating¬ point vector of a known domain name using a model to determine if the domain name contains the typographical error.
Owner:DNSFILTER INC