Variable-length character coding method capable of infinitely expanding coding range
By adopting logo bit design and Unicode character set support in the character encoding method, the problems of low encoding efficiency, large storage space, slow transmission speed and poor compatibility in the prior art are solved, efficient and space-saving character encoding is achieved, and compatibility with existing standards is maintained.
Patent Information
- Application Number
- CN202510176778.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-30
AI Technical Summary
Existing character encoding methods have shortcomings in encoding efficiency, storage space occupation, transmission speed and compatibility, and are difficult to meet the growing character needs and requirements of big data applications.
Using unique logo bit design and Unicode character set support, identify character boundaries through logo bits and convert Unicode code points directly into binary form, optimize the encoding range and compatibility, and reduce storage space requirements by simplifying the encoding process.
It significantly improves character encoding efficiency, reduces storage space usage, speeds up network transmission speed, and maintains compatibility with existing encoding standards such as ASCII and UTF-8, ensuring infinite expansion and flexibility of future character encoding ranges.
Smart Images

Figure CN120074533A_ABST
Abstract
Claims
1. A character encoding method, characterized in that: * Use the first bit of each byte as a flag to indicate whether it is the starting byte of a character; * Encode Unicode code points / unspecified character set code points to bits 2 to 8 of each byte respectively.
2. A character encoding method, characterized in that: * Receive a character in a Unicode character set / unspecified character set as input; * Determine the number of bytes required for encoding based on the Unicode code point / unspecified character set code point of the character; * According to this encoding method, Unicode code points / unspecified character set code points are encoded into byte sequences; * Output the encoded byte sequence.
3. A character decoding method, characterized in that: * Receive a byte sequence encoded using this method as input; * Parse the byte sequence according to the encoding rules and restore the original characters; * Output the decoded characters.