Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Unicode" patented technology

Unicode is a computing industry standard for the consistent encoding, representation, and handling of text expressed in most of the world's writing systems. The standard is maintained by the Unicode Consortium, and as of May 2019 the most recent version, Unicode 12.1, contains a repertoire of 137,994 characters covering 150 modern and historic scripts, as well as multiple symbol sets and emoji. The character repertoire of the Unicode Standard is synchronized with ISO/IEC 10646, and both are code-for-code identical.

Intelligent ai routing advisory platform with synthetic injection testing, bias detection digital twin, zero-copy pipeline, cryptographic compliance verification, and autonomous multi-tier coordination for heterogeneous ai provider ecosystems

A computer-implemented system for routing artificial intelligence (AI) queries. The system utilizes a zero-copy data pipeline, which processes prompts in memory-mapped buffers to eliminate at least one memory copy operation, thereby reducing latency relative to conventional serialization pipelines. The system continuously verifies AI provider compliance by injecting synthetic prompts containing invisible, Ed25519-signed Unicode watermarks. Algorithmic bias is detected by generating counterfactual “digital twin” prompts and applying Fisher exact statistical testing.Routing decisions for multi-tier autonomous systems are governed by safety-level requirements (ASIL-D, ASIL-B, QM) and may be constrained by external routing directives received via a meta-identifier. A hash-chained manifest, cryptographically signed using Ed25519 and consumed by downstream gateways, is generated for each routing decision, with its Merkle root asynchronously anchored to a blockchain to create a tamper-evident audit trail for regulatory compliance.
Owner:WEBER AXEL

Large-capacity file high-speed encryption method based on optimized SM4 / AES

The invention discloses a high-capacity file encryption method and system based on block processing. According to the method, a high-capacity file is read and encrypted in blocks according to the fixed size, the situation that the high-capacity file is wholly loaded to a memory is avoided, memory occupation is reduced, and efficient encryption is achieved; two symmetric encryption algorithms of SM4 and AES are supported, and the corresponding algorithm is automatically identified and called during decryption by embedding an algorithm identifier in an encrypted file; and generating a check code for the encrypted file by using SM3 for integrity check before decryption. A file path coding mode irrelevant to a platform is adopted, a Unicode path is supported, and mainstream operating systems such as Windows and Linux can be compatible; and meanwhile, a buffer area and a multi-thread mechanism are introduced, so that file I / O operation and encryption calculation are executed in parallel, and the encryption efficiency is improved by utilizing the advantages of a multi-core processor. According to the method, the problems of high memory occupation, inconvenience in algorithm switching, poor cross-platform compatibility and low encryption efficiency of high-capacity file encryption are solved, and the method has a wide application prospect.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Systems and methods for dynamically generating and mapping unicode for cross-distributed network data transmissions

ActiveUS20250330197A1Code conversionSchema mappingLinguistic model
Systems, computer program products, and methods are described herein for generating and mapping unicode for cross-distributed network data transmissions. The present disclosure is configured to identify a cross-distributed network data transmission comprising source ledger code; generate a core schema for a destination ledger, wherein the destination ledger is based on data of the cross-distributed network data transmission; generate, by a ledger transformation module, a data dictionary, wherein the ledger transformation module receives ledger code from a large language model (LLM); generate code mapping instructions for the cross-distributed network data transmission based on the core schema for the destination ledger; convert the source ledger code to a unicode; receive, by a receiver schema mapping module, the unicode and the code mapping instructions; and generate, by the receiver schema mapping module, an intermediary ledger comprising the unicode and based on the code mapping instructions.
Owner:BANK OF AMERICA CORP

Character marking and identifying method and device, equipment and storage medium

The invention discloses a character marking and identifying method and device, equipment and a storage medium, and the method comprises the steps: obtaining the text content of a current text input by a user, mapping the mantissa feature of each Unicode point value in the text content to a private area, generating a corresponding private area code point, and carrying out the recognition of the Unicode point value in the private area; the private area comprises a basic multi-language plane private area, a private area A and a private area B; inserting the private area code points into the text content as zero-width characters to obtain updated target text content; the original mantissa features of the code points of the private area are compared with the actual mantissa features of the target text content, and the current text is judged to be manually input or other source content including AI generation according to the comparison result, so that the reliability and accuracy of text identification can be remarkably improved; the method does not need to depend on a large language model for complex analysis, greatly reduces the technical threshold and detection cost, and has good cross-platform compatibility and practical application value.
Owner:蒋励勤

Embedded device word stock dynamic replacement system and method based on Unicode and GB2312 coding mapping

The invention discloses an embedded device word stock dynamic replacement system and method based on Unicode and GB2312 coding mapping, and belongs to the technical field of embedded system character coding and word stock management. According to the method, a lightweight Unicode-GB2312 bidirectional index mapping table is constructed, and on-demand loading and mixed rendering of Unicode characters are achieved on the premise that an original word stock is reserved. And by adopting a hash table accelerated query, virtual zone bit code expansion and incremental updating mechanism, the storage occupation is reduced by more than 95%, the word stock updating efficiency is improved by 100 times, and the progressive adaptation of a full Unicode character set is supported. The method can be widely applied to intelligent terminals, industrial equipment and other scenes, and the contradiction between multi-language compatibility and resource limitation is solved.
Owner:NANJING GUODIAN NANZI POWER GRID AUTOMATION CO LTD

Mongolian Baiti-based Mongolian conversion method and device for Mongolian conforming to GBT25914-2023 standard, readable storage medium and equipment

PendingCN120930595ANatural language data processingControl characterSoftware engineering
The invention provides a Mongolian Baiti-based method and device for converting Mongolian into Mongolian conforming to a GBT25914-2023 standard, a readable storage medium and equipment. The method comprises the following steps: S1, inputting Mongolian in a word form; s2, preprocessing the Mongolian words, including deleting other control characters except the first control character in the plurality of continuous control characters; s3, obtaining font IDs obtained by Mongolian words through a text rendering engine; s4, the obtained font ID is converted into a corresponding nominal character Unicode code string which conforms to the standard of GB / T25914-2023; s5, filtering the redundant control characters of the Mongolian words corresponding to the Unicode coding string of the nominal character converted in the step S4, and filtering the redundant control characters of the Mongolian words corresponding to the Unicode coding string of the nominal character converted in the step S4; and S6, finally outputting the Mongolian which accords with the standard of GB / T25914-2023. The invention provides a method for converting Mongolian established by Mongolian Baiti into Mongolian conforming to the GB / T25914-2023 standard, so as to conform to the policy requirement and meet the requirement that the Mongolian conforming to the GB / T25914-2023 standard needs to be used in reality, and provides a method for converting Mongolian established by Mongolian Baiti into Mongolian conforming to the GB / T25914-2023 standard.
Owner:INNER MONGOLIA UNIVERSITY

Method and system for generating a unified script code (USC) representation for multilingual text processing

PCT designated stageWO2026074589A1Natural language translationSyllableDevanagari
The present invention provides a computer-implemented method and system for generating a Unified Script Code (USC) representation of multilingual text to enable consistent, script-neutral, and phonemically accurate encoding across diverse languages and writing systems The system converts Unicode-encoded text from one or more Indic scripts into a Devanagari-based intermediate form using predefined or bitwise mapping. It then normalizes the text by inserting inherent vowels, converting dependent vowel signs (matras) into independent vowels, and removing halant characters. The resulting USC representation explicitly encodes consonant-vowel sequences, reduces script specific variation, and preserves phonetic integrity. This approach improves tokenization efficiency, enhances performance in natural language processing and machine learning tasks, and supports reversible conversion to original scripts.
Owner:SINGH NAU NIHAL

An embedded device font library dynamic replacement system and method based on Unicode and GB2312 encoding mapping

The application discloses a kind of embedded device font library dynamic replacement system and method based on Unicode and GB2312 encoding mapping, belong to embedded system character coding and font library management technical field.Its method is by constructing lightweight Unicode-GB2312 bidirectional index mapping table, realizes Unicode character on-demand loading and mixed rendering under the premise of reserving original font library.Hash table is used to accelerate query, virtual area code extension and incremental updating mechanism, so that storage occupancy is reduced by more than 95%, font library updating efficiency is improved by 100 times, and progressive adaptation of full Unicode character set is supported.The application can be widely applied to intelligent terminal, industrial equipment and other scenes, and solves the contradiction between multi-language compatibility and resource limitation.
Owner:NANJING GUODIAN NANZI POWER GRID AUTOMATION CO LTD

Encryption method, decryption method, leakage identification method and device

ActiveCN114741709BPlay a protective effectDoes not change document encodingInternet privacyEngineering
This application provides an encryption method, decryption method, leakage identification method, text transmission method, and apparatus, including: obtaining a string to be encrypted; determining a protection character in the string to be encrypted based on a preset association between the original code position of the protection character in the Unicode Conformity Table (Unicode Conformity Table) and the multiplexed code position in the Unicode Conformity Table; and replacing the original code position of the protection character with a target multiplexed code position corresponding to the original code position according to the association, thereby obtaining an encrypted string. The encryption technology described in this application is based on the text itself. The encrypted text will be displayed in normal plaintext form in a trusted environment, while it will be presented as garbled text in an untrusted environment. Even if the encrypted text is leaked, it can still protect the text content. The decrypted protection character can be displayed and edited in normal plaintext form in a trusted environment, achieving the purpose of normal use of the text content.
Owner:ALIBABA (CHINA) CO LTD

Systems and methods for detecting typographical errors in domain name entries

Systems and methods for detecting a typographical error in a domain name, including: receiving a domain name comprising Unicode characters; encoding each character, where the encoding includes: computing an integer index of the Unicode characters; converting the integer index into a binary representation; and multiplying the binary representation by a dense matrix to obtain a floating-point vector; and comparing the floating-point vector to a reference floating-point vector of a known domain name using a model to determine if the domain name contains the typographical error.
Owner:DNSFILTER INC

Flattening HTTP decoding method and system based on multi-classification learning, and medium

The invention discloses a flattening HTTP (Hyper Text Transport Protocol) decoding method and system based on multi-classification learning and a medium. A multi-classification HTTP coding detection model is generated by constructing an HTTP coding data set, extracting a multi-classification feature matrix of the coding data set and training the coding data set through multi-classification learning; extracting an HTTP original load from the HTTP traffic, extracting a multi-classification feature vector of the HTTP original load, loading a multi-classification HTTP coding detection model, and carrying out multi-classification HTTP coding detection; judging an HTTP coding detection result; if the code of the original load in the HTTP flow is detected, flattening HTTP decoding is carried out; recording and outputting an HTTP decoding result; according to the method, the original load is extracted from the HTTP flow, multi-classification machine learning training detection model is carried out aiming at multiple types of codes such as URL, Base64, Unicode, Hex and HTML, flattening decoding operation is carried out according to a model detection result, and finally the clear HTTP load is cleaned out.
Owner:CHANGSHU INSTITUTE OF TECHNOLOGY

Multi-type HTTP (Hyper Text Transport Protocol) coding feature extraction method and system and medium

The invention discloses a multi-type HTTP (Hyper Text Transport Protocol) coding feature extraction method and system and a medium. The method comprises the following steps: constructing a multi-type HTTP coding data set; sample data of URL codes, Base64 codes, Unicode codes, Hex codes, HTML codes and combined codes and nested codes of the URL codes, the Base64 codes, the Unicode codes, the Hex codes and the HTML codes are collected respectively, and corresponding type labels are marked; the method comprises the following steps: reading a multi-type HTTP coded data set, sequentially reading samples, and carrying out feature extraction; constructing a multi-type coding feature matrix; constructing a coding sample feature vector, constructing and standardizing a feature matrix of the multi-type HTTP coding data set, constructing a coding type vector matrix according to type values in the multi-type HTTP coding data set, and performing transverse splicing operation on the coding type vector matrix and the standardized feature matrix to construct a multi-type coding feature matrix; according to the method, a multi-type coding data set which can be used for coding learning training, verification and testing is constructed, feature extraction is performed through unique features of multi-type coding, and the accuracy and efficiency of feature extraction are improved.
Owner:CHANGSHU INSTITUTE OF TECHNOLOGY

A sensitive word detection method and device, computer equipment and a storage medium

The embodiment of the application belongs to the technical field of natural language processing, and relates to a sensitive word detection method and device, computer equipment and a storage medium. The method comprises the following steps: receiving a sensitive word detection request sent by a user terminal; calling a Chinese character Unicode encoding white list, and performing an interference character clearing operation on input text data according to the Chinese character Unicode encoding white list to obtain pure text data; performing a character merging operation on the pure text data to obtain to-be-matched text data; calling a sensitive word Trie tree, and performing a sensitive word similarity matching operation in a similarity matching chain between the characters of the to-be-matched text data and the root node character of the sensitive word Trie tree as a starting point to obtain matching result data; performing a comprehensive scoring operation on the matching result data to obtain a comprehensive scoring result; and outputting the comprehensive scoring result to the user terminal. The sensitive words in the user input text are automatically and efficiently identified and filtered, so that the safety and compliance of the system are protected.
Owner:ASPIRE INFORMATION TECH BEIJING

A method for converting information characters of a bus instrument matrix

The application discloses a passenger car instrument information character matrix conversion method, comprising the following steps: firstly, using a pixel matrix group to deconstruct a defined instrument section to be described, and allocating a character to each matrix block according to a basic pixel block; then, performing unicode coding on the character to be expressed; then, performing space transformation processing on the character matrix of the space; finally, combining the calculated character block matrix and transmitting the combined character block matrix to an instrument medium to be displayed. The method has low requirements on the computing power of a chip, simplifies the instrument graphic display performance, and solves the instrument display problem of a low-cost chip.
Owner:XIAMEN KING LONG UNITED AUTOMOTIVE IND CO LTD

Pinyin supplementation labeling system and method for distinguishing six tones of'machine / product, odd / seven and sparse / west 'in Chinese pronunciation

The invention relates to the field of language and character information processing, in particular to a supplementation labeling system for uniquely distinguishing three groups of homomorphic abnormal sound phonemes of'machine / product, odd / seven and sparse / west 'through machine-readable symbols in a Chinese pinyin system, an input method, electronic equipment and a computer readable storage medium, and aims to provide a pinyin supplementation labeling system. On the premise of keeping the original spelling of the Pinyin scheme unchanged, the pronunciation of the product, the pronunciation of the seven and the pronunciation of the west are uniquely distinguished by using the recorded symbols (. J.q.x) of the Unicode, namely the prefix points and the original letters, so that the blind test of 100 thousand news corpora can be achieved, and after the method is adopted, the TTS misreading rate is reduced from 2.8% to 0.07%; the ambiguity of the ASR post-processing homomorphic and abnormal sound field is reduced by 62%, and the recognition accuracy of the whole sentence is improved by 3.4%; in a class test of Chinese as a foreign language, the correct rate of distinguishing initial consonants of learners is improved by 28%, and class hours are shortened by 30%; the symbol occupies 2 bytes, and the expansion rate is stored as lt; 0.5%, almost negligible; the method is completely compatible with Unicode, GB 18030 and ISO 10646, and a new character does not need to be made.
Owner:霍立远

Systems and methods for dynamically generating and mapping unicode for cross-distributed network data transmissions

ActiveUS12549196B2Code conversionSchema mappingNetwork data
Systems, computer program products, and methods are described herein for generating and mapping unicode for cross-distributed network data transmissions. The present disclosure is configured to identify a cross-distributed network data transmission comprising source ledger code; generate a core schema for a destination ledger, wherein the destination ledger is based on data of the cross-distributed network data transmission; generate, by a ledger transformation module, a data dictionary, wherein the ledger transformation module receives ledger code from a large language model (LLM); generate code mapping instructions for the cross-distributed network data transmission based on the core schema for the destination ledger; convert the source ledger code to a unicode; receive, by a receiver schema mapping module, the unicode and the code mapping instructions; and generate, by the receiver schema mapping module, an intermediary ledger comprising the unicode and based on the code mapping instructions.
Owner:BANK OF AMERICA CORP

A method, device, computer storage medium and terminal for realizing character recognition

A method, device, computer storage medium and terminal for realizing character recognition, the embodiment of the present application acquires a character image containing a preset number of characters based on a PDF document, and performs multi-word recognition on the character image; in the case that the result of multi-word recognition contains a preset number of characters, the final recognition result of the characters contained in the PDF document is determined according to the result of multi-word recognition. The embodiment of the present application can read and render the characters in the PDF file, can determine which character in the PDF the character image corresponds to, in the case that the result of multi-word recognition contains a preset number of characters, the quick confirmation of missed detection or multiple detection is realized, and the accuracy of character recognition is improved according to the result of multi-word recognition; the recognized result is converted into a uniform code (Unicode) and attached to the corresponding character of the PDF, and the accurate recognition of the characters contained in the PDF document is realized.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +1

Regular expression filter for Unicode transform format strings

The invention is entitled "Regular Expression Filter for Unicode Transformation Format Strings". The invention provides an apparatus comprising a memory system having a memory die, the memory system comprising circuitry configured to receive a plurality of bytes at a first input terminal, the plurality of bytes comprising an input stream of a plurality of first variable length encoded symbols, and the circuitry configured to determine a sequence of symbol byte lengths for each of the plurality of first variable length encoded symbols.
Owner:SANDISK TECHNOLOGIES LLC

Unicode confusion command attack detection system and method based on dual-granularity feature coupling model

The invention discloses a Unicode confusion command attack detection system and method based on a dual-granularity feature coupling model, and belongs to the technical field of network security. In order to solve the problem of accurate recognition of confused malicious commands, a data preprocessing module, a standardization module, a character-level confusion feature extraction module, a word-level semantic feature extraction module, a dual-granularity feature coupling module and a malicious command judgment module are connected in sequence; the data preprocessing module is used for cleaning command line data and executing command line text standardization operation; the standardization module is used for replacing a specific mode with a universal placeholder and performing word segmentation operation; the character-level confusion feature extraction module is used for extracting character-level features of command keywords; the word-level semantic feature extraction module is used for extracting word-level features; the dual-granularity feature coupling module is used for forming a mixed feature vector; and the malicious command judgment module is used for inputting the mixed feature vector into a pre-training random forest classifier to output a binary classification result.
Owner:HARBIN UNIV OF SCI & TECH

Precise PDF font matching method and system based on multi-dimensional data fusion

The invention relates to a PDF font accurate matching method and system based on multi-dimensional data fusion, and the method comprises the steps: a data pre-collection layer collects a font matching multi-dimensional data source, carries out the caching, and updates the data in a user inactive time period; the feature calculation layer is used for carrying out character-by-character scanning, measurement similarity calculation and dynamic Unicode segmented font set optimization on original data; and the hierarchical decision engine layer performs font matching according to scenarized requirements, selects font resources for editing operation of a user on the premise of ensuring visual consistency of documents edited by the user, and performs embedding operation on the font resources. According to the PDF font accurate matching method and system based on multi-dimensional data fusion, the data pre-collection layer, the feature calculation layer and the hierarchical decision engine layer are included, efficient and accurate font matching is achieved through multi-dimensional data fusion and dynamic priority scheduling, the situation that fonts used in a PDF document are displayed as messy codes can be effectively avoided, and the accuracy of matching of the fonts is improved. And the copyright compliance risk can also be avoided.
Owner:SHENZHEN JINNIU TECH CO LTD

Detection of Machine Learning Model Attacks Obfuscated in Unicode

A prompt for a generative artificial intelligence (GenAI) model which contains unicode is received. The prompt is then tokenized to result in a plurality of tokens. Token forming part of a repeating sequence are identified and then removed to result in a modified set of tokens. The modified set of tokens are subsequently detokenized to result in a modified prompt. It is then determined, whether ingestion of the modified prompt by the GenAI model will result in the GenAI model behaving in an undesired manner. The modified prompt is passed to the GenAI model when it is determined that ingestion of the modified prompt will not result in the GenAI model behaving in an undesired manner. Otherwise, at least one remediation action is initiated when it is determined that ingestion of the modified prompt by the GenAI model will result in the GenAI model behaving in an undesired manner.
Owner:HIDDENLAYER INC

A multi-language display method, device and equipment for an automobile instrument and a storage medium

This application discloses a method, apparatus, device, and storage medium for multilingual display of automotive instrument panels. The technical solution provided in this application obtains multiple original Unicode encoding strings corresponding to the text to be displayed sent by the vehicle's host computer, as well as translation markers for the original Unicode encoding strings. Based on the translation markers, the original Unicode encoding strings are translated to obtain multiple target Unicode encoding strings. Multiple glyph data are determined from a preset character library based on the multiple target Unicode encoding strings, and the glyph data is rendered and displayed. This eliminates the need for complete sentence translation of the text to be displayed, effectively improving display accuracy in multilingual usage scenarios while maintaining automotive chip performance, thus enhancing the display effect of the automotive instrument panel.
Owner:GUANGZHOU ZHOULIGONG SCM DEV

Intelligent ancient Tibetan character dividing method based on character combination structure and Attention BiLSTM

The invention discloses an ancient Tibetan intelligent character dividing method based on a character combination structure and Attention BiLSTM, and relates to the technical field of natural language processing, and the method comprises the steps: collecting electronic literatures, ancient book digital achievements and non-divided Tibetan texts in a standard corpus; performing cleaning, format regularization and VCC sequence extraction on the text, labeling a configuration role and completing character standardization processing; through Unicode coding standardization and VCC structure verification, illegal combinations are corrected, and homomorphic and heteromorphic codes are unified; calculating a structure specification coefficient, a semantic aggregation coefficient and a structure legality coefficient, dynamically evaluating data cleaning, model semantic consistency and language path legality, and automatically adjusting a strategy for unqualified conditions; according to the method, an enhanced model fusing residual connection and a multi-head attention mechanism is combined with a CRF layer, based on a BMES label system and a Viterbi algorithm, high-precision word segmentation prediction and error correction are achieved, and key technical support is provided for digital processing of the ancient Tibetan.
Owner:ZHEJIANG UNIV +1

Mixed language typesetting method, device and equipment for embedded equipment and medium

PendingCN121303055ANatural language data processingProgramming languageCharacter analysis
The invention discloses a mixed language typesetting method and device for embedded equipment, equipment and a medium, and the method comprises the steps: obtaining a to-be-typeset text stream of a mixed language, and dividing the to-be-typeset text stream into a plurality of text fields through a preset Unicode bidirectional algorithm to obtain a text field list; according to the input sequence of the to-be-typeset text stream, character-by-character analysis is conducted on the character fields in the character field list in sequence, so that a font bitmap corresponding to each character is obtained from a preset cache region, according to pre-generated typesetting parameters, a target embedded device is controlled to conduct typesetting rendering on the font bitmap of each character line by line, and typesetting rendering is conducted on the font bitmap of each character line by line. Unified typesetting of the text of the mixed language on the target embedded device is achieved; wherein the typesetting parameters are generated according to a preset alignment rule table; the alignment rule table comprises a baseline offset table and an equal-width proportion table. According to the method and the device, the real-time performance and the visual consistency of the mixed language typesetting of the embedded equipment can be improved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Systems and methods for detecting typographical errors in domain name entries

Systems and methods for detecting a typographical error in a domain name, including: receiving a domain name comprising Unicode characters; encoding each character, where the encoding includes: computing an integer index of the Unicode characters; converting the integer index into a binary representation; and multiplying the binary representation by a dense matrix to obtain a floating-point vector; and comparing the floating-point vector to a reference floating¬ point vector of a known domain name using a model to determine if the domain name contains the typographical error.
Owner:DNSFILTER INC

Vector character recognition method and system based on bag-of-words model feature point retrieval

The application particularly relates to a vector character recognition method based on a bag-of-words model feature point retrieval, which comprises the following steps: S100, reading character contour information of any vector graph by reading vector graph data; S200, analyzing the character contour information into control point coordinates; S300, drawing the control point coordinates into a control point grayscale graph; S400, extracting an ORB feature vector according to the control point grayscale graph; S500, taking the ORB feature vector as input, and searching for a character ID with the highest similarity from a visual dictionary through a bag-of-words tree index; and S600, obtaining a font and a unicode code corresponding to the vector character through a character ID mapping relationship. Through the above scheme, the vector graph file can be directly subjected to character recognition without format conversion, and meanwhile has the following multiple advantages: first, the character recognition range is large, the accuracy is high, and the character set can be extended to be larger; second, the character recognition speed is fast, and the single character recognition speed is about 1.5 ms; and third, the font can be judged while the character recognition is performed.
Owner:HEFEI HIGH DIMENSIONAL DATA TECH CO LTD