Natural language character and numerical value general solution system and method based on 4HPU and application of natural language character and numerical value general solution system and method
Through the natural language character and numerical solution system based on 4HPU, the code of numbers, letters, Chinese characters and symbols is encoded from right to left, which solves the problem of lack of uniformity and nesting of encoding in the prior art, and realizes efficient data recovery and natural language processing.
Patent Information
- Application Number
- CN202510638608.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, natural language characters and numerical encoding lack uniformity, structural and semantic nesting capabilities, resulting in the inability to effectively handle complex nesting and reasoning in artificial intelligence models, especially in natural language, numerical logic and mixed symbol scenarios, which are inefficient in coding and severe information loss.
The natural language character and numerical solution system based on 4HPU is adopted. Through the input character module, category detection module, 4HPU encoder, nested path generator and output encoding string module, numbers, letters, Chinese characters and symbols are encoded from right to left, the main encoding path is constructed, and reverse recognition and logical expansion are achieved through the decoding path and the extended path.
It realizes the unified, universal and inference of natural language characters and numerical encoding, and provides clear encoding methods for structural expression and path inference, which are suitable for intelligent error correction, data recovery, natural language processing and low-power device application scenarios.
Smart Images

Figure CN120386847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data processing system and method of artificial intelligence and its applications, in particular to a system and method for character encoding, symbol structure modeling and intelligent computing and its applications. Background Art
[0002] Artificial intelligence computing systems in the prior art usually use encoding methods such as the American Standard Code for Information Interchange (ASCII) and Unicode to perform numerical mapping on characters and symbols. Such methods are essentially linear mapping encodings, lacking directionality, structure, and semantic nesting capabilities. In the processing of artificial intelligence (AI) models and structure recognition, the above linear encoding cannot carry complex nesting and reasoning (using a trained model to perform operations on new data). Especially in scenarios such as natural language, numerical logic, and mixed symbols, such as mathematical symbols, units, and punctuation, the existing encoding methods have the following defects: the encoding methods are not unified. For example, natural numbers use decimal, characters use ASCII, Chinese characters use Unicode, and mathematical symbols and logical symbols often use independent extended sets. The encoding between various language systems is severely fragmented and lacks a unified abstract model; there is a lack of a coordinated format for expressing context structures, that is, there is a lack of expressive structures with pathability, nesting, and relevance, resulting in the inability to compress and model complex structures at the sentence level and formula combination level; the encoding system itself lacks an internal direction flow and nesting structure because the existing encoding format is only a static symbol mapping without pre- and post-order semantics and cannot carry the cognitive expression method that "structure is a path"; it is impossible to compress and express polysemous words, composite expressions, or graphic structure information with a high-dimensional structure, lacking a path for nested topology and dynamic backtracking, making the encoding inefficient and causing serious information loss during intelligent reasoning and information backtracking. Therefore, in terms of AI model structure understanding, logical expression, and multi-modal symbol collaborative processing, the current encoding system has gradually exposed its bottlenecks and incompatibilities, and there is an urgent need for a new encoding architecture with structural path logic, compression ability, and cross-context general solution ability. Summary of the Invention
[0003] The object of the present invention is to provide a natural language character and numerical general solution system and method based on 4HPU and its applications, and the technical problem to be solved is to achieve unified, general, and inferable natural language character and numerical encoding.
[0004] The present invention adopts the following technical solutions: A natural language character and numerical general solution system based on 4HPU is used for encoding and processing raw data of numbers, letters, Chinese characters, or symbols at the bottom layer of artificial intelligence, and is provided with a main encoding path for realizing encoding generation and structure nesting.
[0005] The main coding path of the system of the present invention consists of an input character module, a category detection module, a 4HPU encoder, a nested path generator, and an output coding string module;
[0006] The input character module receives the original data of numbers, letters, Chinese characters, and / or symbols input by the artificial intelligence character interface and sends it to the category detection module;
[0007] After receiving the original data, the category detection module conducts identification, determines which category among numbers, letters, Chinese characters, or symbols it belongs to, gives a category flag, and sends it to the 4HPU encoder; the category flag consists of 4 digits, and each digit of the 4 digits is 0 or 1;
[0008] The 4HPU encoder encodes numbers, letters, Chinese characters, or symbols in a nested manner from right to left to obtain a coding string composed of at least one group of basic coding units, and then outputs the coding string and the category flag to the nested path generator; the basic coding unit is 4 digits, each digit of the 4 digits is 0 or 1, and only one digit of the 4 digits is 1;
[0009] The numerical coding adopts the structure of the arrangement of basic coding units;
[0010] The letter coding is represented by symmetric positive and negative numerical values according to the upper and lower cases of English letters;
[0011] The Chinese character coding takes strokes as units, maps each stroke to 4HPU basic coding units, and arranges and encodes them according to the stroke order to form a coding string structure;
[0012] The symbol coding adopts a coding string structure of 4 - bit, 8 - bit, or 12 - bit numbers from the idle sparse coding;
[0013] The nested path generator arranges the category flag and the basic coding units in the order from right to left, and combines them to form a complete coding string with a hierarchical structure relationship for the number, letter, Chinese character, or symbol: category flag + data coding string;
[0014] The output coding string module outputs the complete coding string to the artificial intelligence data processing part outside the main coding path.
[0015] The strokes of the system of the present invention are dot, horizontal, vertical, left - falling stroke, right - falling stroke, rising stroke, turning stroke, and hook stroke. The coding of dot, horizontal, vertical, and left - falling stroke and the coding of right - falling stroke, rising stroke, turning stroke, and hook stroke form an inverse relationship of basic coding units.
[0016] For the symbol coding of the system of the present invention, the idle 4 - bit sparse coding is assigned to arithmetic symbols, the idle 8 - bit sparse coding is assigned to comparison symbols, and the idle 12 - bit sparse coding is assigned to logical symbols and set symbols.
[0017] The natural language character and numerical general solution system based on 4HPU of the present invention system is provided with a decoding path and an expansion path. The decoding path is used to realize the reverse recognition of the encoding, and the expansion path is used to realize the expansion of the logic process.
[0018] The decoding path of the present invention system is composed of a nested path generator, a reverse path parser, and an identification result output module;
[0019] The reverse path parser obtains a complete encoded string from the nested path generator, unfolds according to the reverse process of encoding numbers, letters, Chinese characters or symbols by the 4HPU encoder, restores the basic encoding unit to the original data, and restores the category flag to numbers, letters, Chinese characters or symbols;
[0020] The identification result output module outputs the restored original data, the restored numbers, letters, Chinese characters or symbols outside the decoding path.
[0021] The expansion path of the present invention system is composed of a nested path generator, an idle path allocator, a symbol expansion output module, and a reverse parser;
[0022] The idle path allocator obtains the basic encoding units not used in the complete encoded string from the nested path generator, and expands and generates a new 4HPU code bit allocation rule according to the idle sparse code set;
[0023] The symbol expansion output module outputs the new 4HPU code bit allocation rule outward.
[0024] A natural language character and numerical general solution method based on 4HPU includes the following steps:
[0025] I. Construct a basic encoding unit
[0026] Use 4-bit digital codes to construct a basic encoding unit. Each bit of the 4-bit digital code can be 0 or 1, and only one bit of the 4-bit digital code is 1;
[0027] The basic encoding unit has four encoding states: S1, S2, S3, S4. The encoding corresponding to S1 is 0001, the encoding corresponding to S2 is 0010, the encoding corresponding to S3 is 0100, and the encoding corresponding to S4 is 1000;
[0028] II. Determine the categories of numbers, letters, Chinese characters or symbols
[0029] For the original data from a limited scale, respectively determine whether the original data belongs to the categories of numbers, letters, Chinese characters or symbols, and give the category flags of numbers, letters, Chinese characters or symbols;
[0030] The category flag consists of 4 digits, each digit in the 4 digits being 0 or 1. The encoding for the numeric category is 0001, the encoding for the alphabetic category is 0010, the encoding for the Chinese character category is 0100, and the encoding for the symbol category is 1000;
[0031] The numerical value is a natural number, such as 1, 2, 3. The letters are uppercase and lowercase English letters, such as A - Z, a - z. The Chinese characters are the recording symbols of Chinese, and the symbols are mathematical operation symbols, comparison symbols, set logic operation symbols, and punctuation control symbols;
[0032] III. Encoding
[0033] The natural numbers, letters, Chinese characters, and / or symbols are encoded in a nested manner from right to left to obtain a complete encoded string: category flag + data encoded string;
[0034] The numerical encoding uses a structural nested encoding, adopting the structure of arranging basic encoding units, that is, arranging another group of basic encoding units after a group of basic encoding units. The number of groups of basic encoding units is determined by the size of the number;
[0035] The alphabetic encoding is represented by symmetric positive and negative numerical values according to the uppercase and lowercase of English letters. It consists of 4 digits, and each digit in the 4 digits can be 0 or 1, and only one digit in the 4 digits is 1;
[0036] The Chinese character encoding takes strokes as units, maps each stroke to a 4HPU basic encoding unit, and arranges the encoding according to the stroke order to form the structure of the encoded string;
[0037] The symbol encoding adopts an available 4 - bit, 8 - bit, or 12 - bit encoded string structure from the idle sparse encoding.
[0038] The complete encoded string obtained by the method of the present invention is expanded in the reverse process of encoding, restoring the basic encoding units to the original data, and restoring the category flag to numbers, letters, Chinese characters, or symbols.
[0039] An application of a general solution method for natural language characters and numerical values based on 4HPU, where the general solution method for natural language characters and numerical values based on 4HPU is applied to intelligent error correction and completion, direct data recovery and reverse recognition at the local device end, lightweight natural language processing in a low - resource environment, intelligent indexing and similarity matching system based on the encoding path, or reverse path verification in an efficient storage and backup system.
[0040] Compared with the existing technology, the present invention adopts a nested method from right to left to perform structured encoding on natural numbers, letters, Chinese characters, and symbols to obtain a complete encoded string, and establishes a unified, universal, and inferable system and method for encoding numbers, letters, Chinese characters, and symbols. The structural expression and path reasoning are clear, and it can be dynamically expanded and compressed. It provides a more excellent basic encoding method for data recovery and reverse recognition, natural language processing, intelligent communication, and low-power device application scenarios, and can be widely used in the technical fields of natural language structured representation, multimodal intelligent processing, low-latency communication, and edge computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic diagram of the system structure of the present invention.
[0042] Figure 2 It is a schematic diagram of the 4HPU basic coding unit of the present invention.
[0043] Figure 3 It is a natural number coding schematic diagram of the present invention.
[0044] Figure 4 It is a schematic diagram of the nested structural path of letters Aa to Zz of the present invention.
[0045] Figure 5 It is a schematic diagram of the stroke coding structure of the present invention.
[0046] Figure 6 It is a schematic diagram of the positional coding structure of the Chinese characters "人", "日" and "大" of the present invention.
[0047] Figure 7 is a reverse route identification diagram.
[0048] Figure 8 This is a schematic diagram of reverse path identification derivation.
[0049] Figure 9 This is a schematic diagram comparing the 4HPU structure and ASCII encoding.
[0050] Figure 10 It is a schematic diagram of 4HPU encoding of the present invention.
[0051] Figure 11 It is a schematic diagram of the complex logic symbol structure path of the present invention. DETAILED DESCRIPTION
[0052] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.
[0053] The 4HPU-based natural language character and numerical value general solution system (system) of the present invention is used for data processing at the bottom layer of artificial intelligence, specifically for encoding processing of original data of numbers, letters, Chinese characters or symbols.
[0054] As Figure 1 shown, the system of the present invention is provided with an input character module, a category detection module, a 4HPU encoder, a nested path generator, an output encoding module, a reverse path parser, an identification result output module, an idle path allocator, and a symbol extension output module. Each of the modules constitutes a main encoding path, a decoding path, and an extension path. Among them,
[0055] The main encoding path is composed of an input character module, a category detection module, a 4HPU encoder, a nested path generator, and an output encoding string module. The main encoding path is used to implement encoding generation and structural nesting.
[0056] The decoding path is composed of a nested path generator, a reverse path parser, and an identification result output module. The decoding path is used to implement the reverse identification of the encoding and restore the original data.
[0057] The extension path is composed of a nested path generator, an idle path allocator, a symbol extension output module, and a reverse parser. The extension path is used to implement logical process extension.
[0058] The main encoding path, the decoding path, and the extension path constitute a structured general solution system for encoding generation, structural nesting, reverse identification, and logical extension, and fully closed-loop realizes encoding generation, structural nesting, reverse identification, and logical extension.
[0059] Next, the operations of each module in the main encoding path, the decoding path, and the extension path will be described separately.
[0060] 1. Each module in the main encoding path,
[0061] The input character module receives the original data input from the artificial intelligence character interface and sends it to the category detection module. The original data is numbers, letters, Chinese characters, or symbols input by the user.
[0062] After receiving the original data, the category detection module performs identification, determines which category it belongs to among numbers, letters, Chinese characters, or symbols, gives a category flag, and sends it to the 4HPU encoder. The category flag is composed of 4 digits, and each digit of the 4 digits is 0 or 1. The encoding for the number category is 0001, the encoding for the letter category is 0010, the encoding for the Chinese character category is 0100, and the encoding for the symbol category is 1000.
[0063] The 4HPU encoder encodes the numbers, letters, Chinese characters, or symbols to obtain an encoding string composed of at least one group of basic encoding units, and then outputs the encoding string and the category flag to the nested path generator.
[0064] The basic coding unit is a 4-bit digital code. Each of the 4 bits can be 0 or 1, and only one of the 4 bits is 1. The basic coding unit can be used to express status, directionality, or spatial nestability, and is the basic structure for the general solution of the system of the present invention.
[0065] The 4HPU coding adopts a nested structure coding method from right to left:
[0066] The numerical coding adopts a structural nested coding method, using the structure arranged by the basic coding units, that is, arranging another group of basic coding units after a group of basic coding units. The number of groups of basic coding units is determined by the size of the number. For example, Figure 3 as shown, the coding of the number "7" is 00010010, which is composed of the arrangement of two groups of 4HPU basic units. The first group represents the number 5, and the second group represents the number 2, 5 + 2 = 7.
[0067] The letter coding is represented by symmetric positive and negative numerical values according to the upper and lower cases of English letters. It is composed of 4-bit digital codes. Each of the 4 bits can be 0 or 1, and only one of the 4 bits is 1.
[0068] The Chinese character coding takes strokes as units, maps each stroke to the 4HPU basic coding unit, and arranges and codes them according to the stroke order to form a coded character string structure.
[0069] The symbol coding adopts an available 4-bit, 8-bit or 12-bit coded character string structure from the idle sparse coding.
[0070] The nested path generator arranges (nests) the category flag and the basic coding unit according to the complexity of the character structure, such as the number of strokes and the nesting level, in the order from right to left, and combines them to form a complete coded character string with a hierarchical structure relationship for the number, letter, Chinese character or symbol: category flag + data coding string.
[0071] The output coded character string module outputs the complete coded character string to the artificial intelligence data processing part outside the main coding path.
[0072] 2. Each module in the decoding path,
[0073] The reverse path parser obtains the complete coded character string from the nested path generator, and unfolds according to the reverse process of encoding numbers, letters, Chinese characters or symbols by the 4HPU encoder according to the 4HPU coding method set by the present invention, restores the basic coding unit to the original data, and restores the category flag to the number, letter, Chinese character or symbol.
[0074] The recognition result output module outputs the restored original data, the restored number, letter, Chinese character or symbol, outside the decoding path for display or further processing.
[0075] 3. Expand each module in the expansion path,
[0076] The idle path allocator is used to allocate idle coding bits for symbol expansion. The idle path allocator obtains the basic coding units (structural codes) that are not adopted (matched successfully) in the complete coding string from the nested path generator, and combines symbols according to the idle sparse code set to expand and generate a new 4HPU coding bit allocation rule.
[0077] The symbol expansion output module outputs the new 4HPU code bit allocation rule outward.
[0078] The symbol expansion output module and the reverse path parser form a look-up table dependency relationship. The look-up table dependency relationship means that during the reverse recognition process, the reverse path parser completes the restoration and recognition operation of special symbols by referring to the symbol coding mapping data provided by the symbol expansion output module. In Figure 1 it, the symbol expansion output module is connected to the reverse path parser through a virtual path, indicating that it is the look-up source module for the coding mapping table.
[0079] The natural language character and numerical general solution method (method) based on 4HPU of the present invention includes the following steps:
[0080] I. Construct a basic coding unit
[0081] Use 4-bit digital codes to construct a basic coding unit. Each bit in the 4-bit digital code can be 0 or 1, and only one bit in the 4-bit digital code is 1. The basic coding unit can be used to express status, directionality, or nestability.
[0082] As Figure 2 shown, map the basic coding unit expressed by the four-bit digital code to a two-dimensional plane to form a planar structure diagram of the basic coding unit expressed by numerical values.
[0083] In the planar structure diagram expressed by numerical values, the basic coding unit has four coding states: state S1 to state S4. The code corresponding to S1 is 0001, the code corresponding to S2 is 0010, the code corresponding to S3 is 0100, and the code corresponding to S4 is 1000.
[0084] When the four-bit digital codes of S1 to S4 express directionality, the directionality coding rule is: 0001 to the right, 0010 upward, 0100 downward, 1000 to the left, and the direction symbols are: → to the right, ↑ upward, ↓ downward, ← to the left, as shown in Table 1.
[0085] Table 1 Four coding states and direction expressions
[0086]
[0087] The planar structure diagram expressed by states S1 to S4 can be used to express directionality.
[0088] The main coding path is the process of encoding structure expression. The structure itself has a spatial expansion order, such as the stroke order of Chinese characters, syntactic tree structure or graph structure, so the structural flow can be marked with directionality.
[0089] The nesting of 4HPU codes is arranged from right to left, and S1 to S4 are used for different structural flow processes respectively. The directionality helps in path visualization, structure determination and reverse restoration.
[0090] Directionality enables the coding path graph to have a consistent logical flow, making it easier for the modules of the main coding path to perform path tracking, conflict detection, and backtracking solutions.
[0091] Directionality is also the basis for semantic recognition. For example, if the path points upwards, it means entering the upper nested level; if it points downwards, it means returning to the current node level; and if it points leftwards, it may point to the end of the structure or a closed loop.
[0092] The base layer is divided. The base layer is the smallest unit of the basic coding units arranged to form a multi-layer structure. In the coding structure from right to left, each basic coding unit forms a base layer, and each layer is a level.
[0093] The encoding of each natural number, letter, Chinese character, and symbol is composed of at least one basic coding unit arranged (nested) from right to left. Each valid state in the basic coding unit, such as 0001, 0010, 0100, or 1000, is a layer, forming a basic layer. The rightmost end is the innermost layer (deepest layer), and the leftmost end is the outermost layer (starting layer). If there are only two layers, it can be understood as the outer layer nesting the inner layer. If there are three layers, the layer on the left nests the layer on the right, and so on. For example:
[0094] The number 1 is encoded as 0001, which is a one-layer structure.
[0095] The number 7 is encoded as 0001 0010, which is a two-layer nested structure.
[0096] The number 24 is encoded as 0100 0000, which is a two-layer nested structure and not a place value accumulation.
[0097] The letter A is encoded as 0001 0001, and S1 is nested within S1, also forming a two-layer structure.
[0098] The stroke order of the Chinese character "大" is "horizontal, left-falling, right-falling", so its encoding path is 1011 1110 0001, which is a three-layer structure.
[0099] The base layer has the following functions:
[0100] The path construction unit, the basic layer is the node of the path map and also the constituent fragment of the encoded semantic structure, which can be used to construct nested expressions, structural synthesis, and logical tree diagrams.
[0101] The basic unit of recognition and backtracking. In reverse path recognition, each basic layer can be independently recognized and parsed layer by layer, enabling reverse recovery, path tracing, and structural comparison.
[0102] The scale of compression and expansion. In operations such as path compression, encoding simplification, and symbol merging, the number of basic layers is the basic basis for measuring the expression complexity and compression ratio.
[0103] The mapping bridge for cross-context general solutions. The basic layer enables structures from different sources, such as numbers, languages, and symbols, to be mapped as path units in a unified coordinate structure, with a high degree of abstract expressiveness.
[0104] Therefore, through the division of the basic layer and nested paths, a structural expression model different from traditional binary dense coding is established, realizing the general solution of path-driven natural language and numerical values.
[0105] The basic coding unit is the smallest unit for constructing all structures. The following explains how to determine the state or directionality of the basic coding unit. Four state units are selected as the first group of core basic codes, as shown in Table 2.
[0106] Table 2 The first group of core basic codes
[0107] Status 4HPU Encoding Structural Meaning S1 0001 Indicates that the path "extends or advances to the right" S2 0010 Indicates that the structure "enters a new nesting upward" S3 0100 Indicates that the path "forks or switches levels downward" S4 1000 Indicates that the path "recycles or closes to the left"
[0108] If it represents directionality, the encoding arranged from right to left is mapped to the position change of each state in the encoding structure, and the direction from the innermost layer to the outermost layer is the hierarchical expansion direction.
[0109] II. Judging the categories of numerical values, letters, Chinese characters, or symbols
[0110] For the original data from a limited scale, respectively judge whether the original data belongs to the categories of numerical values, letters, Chinese characters, or symbols, and respectively give the category flags of numerical values, letters, Chinese characters, or symbols according to the corresponding encoding rules of numerical values, letters, Chinese characters, or symbols.
[0111] The category flag consists of 4-digit digits. Each digit in the 4-digit digits is 0 or 1. The encoding for the numerical category is 0001, the encoding for the letter category is 0010, the encoding for the Chinese character category is 0100, and the encoding for the symbol category is 1000.
[0112] The numerical values are natural numbers, such as 1, 2, 3; the letters are English uppercase and lowercase letters, such as A-Z, a-z; the Chinese characters are the recording symbols of Chinese; the symbols are mathematical operation symbols, comparison symbols, set logic operation symbols, and punctuation control symbols.
[0113] 3. Coding
[0114] After determining the state, directionality, and division of the base coding unit, encoding is performed.
[0115] The coding is expressed in 4HPU, which is the abbreviation of 4-level Hierarchical Plane Unit.
[0116] 4 (Four) means that the basic coding unit uses 4-bit state as the basic granularity;
[0117] H (Hierarchical) refers to the "nested structure" in the coding process, which can be compressed layer by layer from the outside to the inside, or expanded layer by layer from the inside to the outside;
[0118] P (Plane) means that the coding system naturally has the ability to expand the two-dimensional plane structure and can be mapped in the XY coordinate system;
[0119] U (Unit) means that all codes are built on the combination of "basic units" and are directional and combinatorial.
[0120] Each 4HPU encoding unit consists of 4 digits, which is called a layer. For a code consisting of two groups of 4 digits, the right group of 4HPU encoding units is the first layer, and the left group of 4HPU encoding units is the second layer. For a code consisting of three groups of 4 digits, the right group of 4HPU encoding units is the first layer, the middle group of 4HPU encoding units is the second layer, and the left group of 4HPU encoding units is the third layer. This method is repeated for each layer.
[0121] The 4HPU encoding method is a structured encoding of natural numbers, letters, Chinese characters and / or symbols in a nested manner from right to left:
[0122] Numerical coding adopts structural nested coding, which uses a structure in which basic coding units are arranged, that is, another group of basic coding units is arranged after one group of basic coding units. The number of groups of basic coding units is determined by the size of the number.
[0123] The letter code is represented by symmetrical positive and negative values according to the uppercase and lowercase English letters. It consists of 4 digits. Each digit in the 4 digits can be 0 or 1, and only one digit in the 4 digits is 1.
[0124] Chinese character encoding is based on strokes, each stroke is mapped to a 4HPU basic encoding unit, and the encoding is arranged in stroke order to form a coding string structure.
[0125] Symbol encoding uses available 4-bit, 8-bit, or 12-bit codeword structures from the available sparse codes.
[0126] The encoding is completed to obtain a complete encoded string: category flag + data encoded string.
[0127] The following specifically describes the encoding of the basic encoding units for numerical values, letters, Chinese characters, and symbols respectively.
[0128] 1. Natural number encoding
[0129] For the encoding of natural numbers, the encoding object is the natural number of any non - negative integer, and it is expressed using a multi - layer nested structure based on 4HPU. As Figure 3 shown, the natural number encoding includes the following steps:
[0130] Convert a natural number value into a hierarchical expression of a set of nested paths, that is, expressed by arranging at least one basic encoding unit. Each basic encoding unit is a layer. The layer on the right is the first layer, also known as the inner layer, and the left - hand side is sequentially called the second layer, the third layer,.... If there are only two layers, the second layer can be called the outer layer. The first layer (the right - most) is the path end or kernel. Each time a layer is added to the left, it represents the wrapping and advancement of the structure. Each layer is represented by 4 - bit 0 or 1 digits (called basic status codes). When represented by the digit 1 in the basic encoding unit, it is called activated, and the activation starts from the first digit on the right and proceeds bit - by - bit to the left.
[0131] Only one position in each basic encoding unit can be activated, that is, each basic encoding unit has four digital expression states. When the activation bit reaches the left - most last bit and the natural number increases further, a new layer is added to the left. The first digit on the right of the second - level layer is marked as 1, and all digits in the first - level layer are 0. When the natural number increases further, the first - level layer is activated bit - by - bit from the first digit on the right. When the activation bit reaches the left - most last bit and the natural number still increases, the second digit on the right of the second - level layer is 1, and all digits in the first - level layer are 0. The natural number encoding continues in this way to form a complete encoded string for each natural number. Specific examples are as follows:
[0132] One - layer 4 - bit encoding
[0133] The natural number 1 corresponds to the encoding 0001 (activation bit is the 1st bit, the first digit is marked as 1),
[0134] The natural number 2 corresponds to the encoding 0010 (activation bit is the 2nd bit),
[0135] The natural number 3 corresponds to the encoding 0100 (activation bit is the 3rd bit),
[0136] The natural number 4 corresponds to the encoding 1000 (activation bit is the 4th bit).
[0137] Two - layer 8 - bit encoding
[0138] The natural number 5 corresponds to the encoding 0001 0000 (activation bit is the 5th bit),
[0139] The natural number 6 corresponds to the code 0001 0001 (the 5th and 1st activation bits),
[0140] The natural number 7 corresponds to the code 0001 0010 (the 5th and 2nd activation bits),
[0141] The natural number 8 corresponds to the code 0001 0100 (the 5th and 3rd activation bits),
[0142] The natural number 9 corresponds to the code 0001 1000 (the 5th and 4th activation bits),
[0143] The natural number 10 corresponds to the code 0010 0000 (the 6th activation bit),
[0144] The natural number 11 corresponds to the code 0010 0001 (the 6th and 1st activation bits),
[0145] The natural number 12 corresponds to the code 0010 0010 (the 6th and 2nd activation bits),
[0146] The natural number 13 corresponds to the code 0010 0100 (the 6th and 3rd activation bits),
[0147] The natural number 14 corresponds to the code 0010 1000 (the 6th and 4th activation bits),
[0148] The natural number 15 corresponds to the code 0100 0000 (the 7th activation bit),
[0149] The natural number 16 corresponds to the code 0100 0001 (the 7th and 1st activation bits),
[0150] The natural number 17 corresponds to the code 0100 0010 (the 7th and 2nd activation bits),
[0151] The natural number 18 corresponds to the code 0100 0100 (the 7th and 3rd activation bits),
[0152] The natural number 19 corresponds to the code 0100 1000 (the 7th and 4th activation bits),
[0153] The natural number 20 corresponds to the code 1000 0000,
[0154] The natural number 21 corresponds to the code 1000 0001,
[0155] The natural number 22 corresponds to the code 1000 0010,
[0156] The natural number 23 corresponds to the code 1000 0100,
[0157] The natural number 24 corresponds to the code 1000 1000.
[0158] Three - layer 12 - bit encoding
[0159] The natural number 25 corresponds to the encoding 000100000000,
[0160] The natural number 50 corresponds to the encoding 001000000000,
[0161] The natural number 75 corresponds to the encoding 010000000000,
[0162] The natural number 100 corresponds to the encoding 100000000000.
[0163] Summarizing the above, the encoding rules of natural numbers are summarized in Table 3.
[0164] Table 3 Natural number encoding
[0165] Natural Number Range Required Number of Layers Example Encoding Path 0~4 1 layer Only one 4-bit activation bit is required 5~24 2 layers Such as 7: 0001 0010 25~124 3 layers Such as 63: 0100 0010 0001 … n layers Multiply each layer by 5 to expand to any integer
[0166] Taking the natural number 7 as an example to illustrate the encoding: The overall encoding string of the natural number 7 is arranged from right to left in sequence: 00010010, the first - layer encoding of 7 is 0010, and the second - layer encoding is 0001.
[0167] The natural number encoding adopts a nested cumulative method from right to left. This directional setting is the core logical basis for constructing structural paths, expressing nested levels, and realizing reversible path recognition.
[0168] The natural number encoding can not only represent the numerical value itself, but also represent directionality. The first layer is located at the right - most end of the encoding string and is the core or end - node of the whole structure. Then come the second layer, the third layer... arranged in sequence from right to left, indicating that the outer structure wraps the inner layer layer by layer. All structural paths are reflected as growing from right to left in the physical encoding, but are reflected as expanding from inside to outside in the structure. This path direction of the arrangement is called "structural directionality". Each layer not only contains structural information, but also conveys path direction and depth. Eventually, the overall encoding can be completely reversely recognized and structurally restored.
[0169] For encoding reverse parsing, encoded by 4HPU, all natural number encodings have structural nesting and traceable path properties. Therefore, each encoding string can be completely reversely parsed to restore its original structural semantics and corresponding numerical value.
[0170] Taking the encoding structure of the overall encoding string 00010010 of the natural number 7, the natural number can be reversely parsed layer by layer from it: Read the encoding string from right to left, that is, read each 4 - bit encoding unit from the innermost layer to the outermost layer,
[0171] The state of the first layer: The right - most 4 bits are 0001, recognized as state S1, direction → right;
[0172] Second - layer status: on the left is 0010, recognized as status S2, with the direction upward ↑;
[0173] According to the path splicing order, from the inside to the outside: S1←S2, restore the coding fragment sequence (path diagram) represented by this value, and finally determine the status - chain coding fragment sequence corresponding to the natural number and the natural number to be encoded.
[0174] Press in the matching items, and the corresponding natural number is obtained as 7.
[0175] 2. Alphabet coding
[0176] As Figure 4 shown, the English alphabet coding adopts a coding structure of positive and negative numbers. Capital letters A - Z are mapped to + 1 to + 26, and lowercase letters a - z are mapped to - 1 to - 26. As shown in Table 4.
[0177] Table 4 Mapping comparison between letters and numbers
[0178]
[0179]
[0180] In order to keep each letter corresponding to 8 bytes, the upper - and lower - case of YZ are encoded with idle codes, which are set to occupy a unified 8 - bit representation, that is, expressed by 2 four - bit basic coding units, forming an 8 - bit (2×4) structure path coding string. The idle code is the idle structure path (idle empty space) that does not participate in the basic coding. The idle path refers to the sparse status bits in the four - bit status combination space that are not occupied by natural numbers, regular letters, or control characters. The method of the present invention, through the mapping rule, assigns these idle coding bits as extended paths to the letters Y and Z, thus completing the unified coverage of the paths of 52 English letters, ensuring that all English letters are represented in a unified 8 - bit structure without compression or additional marking.
[0181] In high - frequency scenarios such as AI model input and chip control signal transmission, adopting a unified length for the coding structure can significantly improve processing efficiency, memory mapping speed, and matrix alignment performance. The structure path with a unified length is convenient for batch input, path comparison, and dynamic analysis, avoiding structural misalignment or matching failure caused by different character lengths. Therefore, the present invention standardizes the coding lengths of upper - and lower - case English letters to each character occupying a fixed two - layer nested structure, that is, 8 bits.
[0182] From - 26 to + 26, each number corresponds to a value encoded by 4HPU, arranged from right to left. Capital letters A - Z correspond to the 4HPU coding from + 1 to + 26. Lowercase letters a - z correspond to the 4HPU coding from - 1 to - 26. As shown in Table 5 and Table 6, the 4HPU coding is arranged from right to left.
[0183] Negative numbers from -1 to -26 represent lowercase letters and are represented by reverse paths, which are not only the reverse identifiers of the structural paths but also can be regarded as the structural mapping expressions of decimal negative numbers. The reverse path and the forward path form a negation relationship in terms of logical relationship. Negation means taking the opposite value of 0 and 1. If the forward is 0, the reverse is 1. Similarly, if the forward is 1, the reverse is 0. For example, for the capital letter A, the decimal number is 1 and the 4HPU code is 0001. For the lowercase letter a, the decimal number is -1 and the 4HPU code is 1110. The structural mapping formed by the positive and negative paths is shown in Table 7.
[0184] Table 5 4HPU Codes for Capital Letters
[0185]
[0186]
[0187] Table 6 4HPU Codes for Lowercase Letters
[0188] Letter Decimal Number 4HPU Encoding (from right to left) a -1 1110 b -2 1101 c -3 1011 d -4 0111 e -5 1110 1111 f -6 1110 1110 … … …
[0189] Table 7 Structural Mapping Formed by Positive and Negative Paths
[0190] Positive Numbering Letter Encoding (→ direction) Reverse Numbering Letter Encoding (← direction) +1 A 0001 0001 -1 a 1110 1110 +2 B 0001 0010 -2 b 1110 1101 +3 C 0001 0100 -3 c 1110 1011
[0191] Through the structural path mapping of uppercase and lowercase letters, an English letter coding structure with a unified length of 8 bits for each letter, a positive and negative symmetric structure, the ability of reversible recognition and direction reasoning, and numerical semantic docking is constructed, providing a stable coding basis for AI input optimization, chip structural path execution, and language model structure compression in practical applications. The coding levels of 26 English letters are shown in Table 8.
[0192] Table 8 Coding Levels of English Letters
[0193]
[0194] The method of the present invention adopts the directional nested expression of letters and numbers to jointly form the nested expression of structural paths to convey the structural intention of composite semantic units.
[0195] The directional nesting is not only the embodiment of the physical coding order but also implies the semantic structure function: the expression of the primary and secondary structure relationship, where the outer path coding usually represents the main semantic unit and the inner layer represents the attached structure or secondary attributes; the semantic context nesting, where letters and numbers can wrap each other to express composite semantics such as "the nth of the mth category" and "which one in which item"; the reverse path differentiates the semantic direction. If the path is from a number to a letter, it represents the attribution of "quantity and category"; if from a letter to a number, it represents the expansion of "category and quantity".
[0196] 3. Chinese Character Coding
[0197] Use Chinese character structure nested expression, such as Figure 5 As shown, a simplified code is generated by the first stroke, the last stroke and the identification code. Each Chinese character is regarded as a structure composed of multiple basic stroke units, and its stroke order information is converted into a 4HPU path encoding string to reflect the combination logic of Chinese characters in terms of space, structure and direction, and to show the stroke order hierarchy and nested encoding path of Chinese characters. The eight basic strokes of Chinese characters are: dot (丶), horizontal (一), vertical (丨), left-falling stroke (丿), right-falling stroke (捺). carry The eight strokes of fold (乛), hook (亅) are coded as shown in Table 9.
[0198] Table 9 Chinese character stroke 4HPU encoding
[0199]
[0200]
[0201] The TI~T4 states have an inverse relationship with the S1~S4 states. For example, S1 is forward expansion, and the 4HPU code is 0001; T1 is reverse expansion, and the 4HPU code is 1110.
[0202] In the method of the present invention, carry The fold (乛) and hook (亅) are states S1 to S4, and the 4HPU codes are 0001, 0010, 0100, 1000. The dot (丶), horizontal (一), vertical (丨), and left-falling stroke (丿) are states T4 to T1, and the 4HPU codes are 0111, 1011, 1101, 1110. carry The fold (乛), hook (亅) 4HPU codes constitute the inverse relationship of the basic coding unit.
[0203] Chinese character encoding includes full-code representation and simplified code representation.
[0204] The full code representation of Chinese characters is to nest and superimpose all strokes in the actual stroke order, that is, to arrange the 4HPU codes of the strokes from right to left in stroke order to form a code string nested from right to left.
[0205] like Figure 6 As shown in the figure, the encoding structure of the Chinese characters "人", "日", and "大" are explained.
[0206] For example, the Chinese character "大" (big) is encoded in the stroke order of horizontal, left-falling, and right-falling strokes. The corresponding stroke codes are: T3, T1, and S1. The 4HPU path code is arranged from right to left as: 0001 (right-falling stroke), 1110 (left-falling stroke), 1011 (horizontal stroke), and the encoding string is: 0001 1110 1011. In terms of the structural hierarchy: the first layer (right end): the right-falling stroke is the starting point of the structure (S1); the second layer: the left-falling stroke is the middle nested stroke (T1); the third layer (left end): the horizontal stroke is the outermost main structure (T3);
[0207] The overall structure shows the logical path of the character "大", which expands from the bottom to both sides.
[0208] Chinese character simplified codes are used to represent some Chinese characters with more strokes or overlapping stroke orders. They are represented by Chinese character simplified codes and are also used for encoding compression and fast recognition.
[0209] The Chinese character code consists of three parts: first stroke + last stroke + identification code. Identification code: used to distinguish Chinese characters with the same first stroke + last stroke, but different characters, such as "大" and "丈". See Table 9. For example:
[0210] The simple code of the character "大" is: the first stroke is 1011 (horizontal), the last stroke is 0001 (downward stroke), and the identification code is 0001; the simple code encoding string is: 1011 0001 0001.
[0211] The stroke order of the character "丈" is also: horizontal, left-falling, right-falling, but the identification code is different, identification code 0010; the simple code encoding string is: 10110001 0010.
[0212] Identification codes ensure that different glyphs remain uniquely identifiable even when compressed. The identification code is determined by sorting a 3,500-character standard character library. A comparison of full and simplified codes is shown in Table 10.
[0213] Table 10 Overview of Chinese character encoding
[0214]
[0215] 4. Symbol encoding
[0216] Symbols are mathematical operation symbols, comparison symbols, set logic operation symbols and punctuation control symbols, see Table 11. Symbol encoding represents the sparse coding mapping relationship of symbols in 4HPU.
[0217] Table 11 Symbol classification
[0218]
[0219] In the method of the present invention, symbol encoding follows an independent encoding principle different from natural numbers, letters, and Chinese characters to ensure its distinctiveness, extensibility, and logical clarity in expression. All symbol encodings are incorporated into the 4HPU sparse structure, and the high-order idle code segments are used for unified configuration and classification planning to ensure the non-ambiguity and structural consistency of symbol recognition.
[0220] Symbol encoding does not conflict with numbers, letters, and Chinese characters. Sparse code segments that are originally idle in the 4-bit encoding system and not used by natural numbers, letters, and Chinese characters are selected for encoding. For example, 1100, 1010, 0101, 0011 belong to the 4HPU sparse codes not occupied by the encodings of numbers, letters, and Chinese characters.
[0221] For concise expression, all basic symbols, including arithmetic symbols, are expressed using 1-layer 4-bit digital encoding. This avoids the complication of path nesting, making it more efficient to recognize in the structural path and facilitating quick recognition by the chip decoder and the AI inference engine.
[0222] Each symbol is logically classified according to its semantic category.
[0223] Idle 4-bit sparse encodings are assigned to arithmetic symbols. The arithmetic symbol encoding uses idle 4-bit codes, which are sparse and highly mutually exclusive, without overlapping with numerical values, letters, and Chinese characters, and can be directly combined in the 4HPU encoding. The arithmetic symbol encoding is shown in Table 12.
[0224] Table 12 Assignment of idle 4-bit sparse encodings to arithmetic symbol encoding
[0225] Symbol 4HPU Encoding Number of Layers Category + 1100 1 layer (4 bits) Arithmetic - 1010 1 layer (4 bits) Arithmetic × 0101 1 layer (4 bits) Arithmetic ÷ 0011 1 layer (4 bits) Arithmetic
[0226] Idle 8-bit sparse encodings are assigned to comparison symbols. The combined 8-bit unused encodings are assigned to comparison symbols (bidirectional). The comparison symbol encoding uses idle 8-bit codes, as shown in Table 13. The combined 8-bit has both directionality and symmetric and opposing characteristics, and very naturally maps to logical symbols such as = and ≠ that require symmetric or opposing expressions.
[0227] Table 13 Assignment of idle 8-bit sparse encodings to comparison symbols
[0228] Idle 8-bit Encoding Specified Use Remarks 11000011 "=" (equal sign) Activated simultaneously on both the left and right, symmetric and stable 10100101 "≠" (not equal sign) Cross activation, symbolizing confrontation and inequality
[0229] Idle 12-bit sparse encodings are assigned to logical symbols and set symbols. The more complex unused 12-bit encodings are used for logical symbols and set symbols (with multiple meanings). The idle 12-bit codes are assigned to logical symbols and set symbols, as shown in Table 14.
[0230] Table 14 Assignment of idle 12-bit sparse encodings to logical symbols and set symbols
[0231]
[0232]
[0233] Positive and negative symmetry. If it is necessary to express reverse logic, such as belonging and not belonging, it can be achieved through positive and negative mapping, that is, a negation relationship is formed in terms of logical relationship. Ensure that all extended encodings will not damage the normal space of the original natural numbers, letters, and Chinese characters.
[0234] As Figure 11 shown, positive and negative symmetry can also be understood as follows: for the structural path of a certain symbol, if its logical semantics has a distinction between "positive" and "negative", such as "belonging ∈" and "not belonging ", then through the reversal of the activation direction of each structural unit in the path, such as S1 and T1, S2 and T2, the reverse logic expression of the encoding is realized without changing its overall structural layer number and position distribution. See Table 15.
[0235] Table 15 Positive and negative symmetry
[0236]
[0237] In this way, the positive and negative meanings of all logical symbols can be symmetrically mapped in the encoding structure, without occupying the space of existing natural numbers, letters, and Chinese characters, and can be extended to conditional judgments or semantic branches under the hierarchical nesting of structural paths, with path reasoning, serving AI structured recognition and language logic calculation scenarios.
[0238] In the 4HPU encoding of the present invention, as Figure 10 shown, all numbers, letters, and strokes are based on the basic encoding unit of 4-bit numerical values, that is, each layer is represented by 4-bit numerical values. After allocating structures for the encoding of numbers, letters, and strokes, there are still a large number of unused 4-bit combined code segments in an idle state, forming an idle code space available for extension. The allocation of 4-bit, 8-bit, and 12-bit codes using sparse activation encoding is shown in Table 16.
[0239] Table 16 Allocation of 4-bit, 8-bit, and 12-bit codes activated by sparse encoding
[0240]
[0241]
[0242] The process of restoring the original character structure from the encoded string through reverse layering is applicable not only to natural numbers but also to Chinese characters with complex structures, symbol expression structures with multiple layers of nesting, or logical combination formulas, such as: ∈, ∑. Start point recognition: Read the first group of 4-bit 4HPU codes from the rightmost starting layer to confirm the path start point; Layer-by-layer reverse traceback: Read the codes layer by layer from right to left, with every 4 bits as a group; Structural action restoration: Each group of codes corresponds to a stroke action, character structure, or semantic logic unit; Path chain restoration: Combine each action in the path order to finally restore the complete character or symbol expression.
[0243] The method of the present invention can perform symmetric expression, negative number representation, combined symbols, and logical connection on numerical values, letters, Chinese characters, and symbols. The basic coding path is defined in a paired manner using a forward path (+) and a reverse path (-), forming a symmetric structure under 4HPU coding.
[0244] Forward path +: Incrementally nested from right to left, representing the natural expansion direction from the start point to the target;
[0245] Reverse path -: Retain the forward path structure, but symmetrically flip the activated bit directions in each layer. For example: S1 to T4, S2 to T3, corresponding to the exchange of positive and negative positions.
[0246] For example, in letter coding, A to Z correspond to +1 to +26, and lowercase letters a to z correspond to -1 to -26, forming a completely symmetric structure. As Figure 7 shown, the coding is superimposed layer by layer from right to left. Each layer can be regarded as a structural path, and the positive and negative values are only expressed as mirror images of the path direction or layer direction.
[0247] Symmetric expression means that for any coding unit (natural number, letter, symbol), there exists a coding item with a path structure symmetric to it, and their direction structures are mirror images of each other and are logically equivalent. For example, +7 and -7 correspond to the reverse of the coding direction of the path structure but the same number of layers.
[0248] Negative number representation means that the path activation mode in the coding adopts the T series structure (3 activation or 4 activation), representing a logically "negative" expression, which can be used to express the negative state of lowercase letters and symbols, such as: -a, -∈.
[0249] Combined symbols refer to composite structures generated by combining multiple coding paths, such as "≥", "≠", "→" in mathematics, and each part of the structure is superimposed and expressed in the form of a 4HPU path combination.
[0250] Logical connection means that there is a correlation between path structures, and some activation sequences themselves imply Boolean relationships, such as "∧", "∨", "⊕", etc., which can be defined with specific nested combinations in the 4HPU structure.
[0251] Each 4-digit 4HPU encoding unit represents a clear structural action, including: Chinese character strokes (such as left-falling strokes, right-falling strokes, and right-falling strokes), numerical jumps (such as hierarchical advancement), letter alignments (such as Z and z mapping), and symbol nodes (such as path starting points or logical judgment forks).
[0252] Since the 4HPU encoding has a cumulative nested structure from right to left, the reverse expression (reverse recognition) or inverse is: starting from the rightmost starting layer, read layer by layer in units of 4 bits, and each time a group is read, a group of structural actions is restored, and the path logic is rebuilt in sequence, and finally the original characters, letters, numbers or symbols are restored.
[0253] For example, the reverse expression of Chinese characters is deduced, and the 4HPU code string is: 0001 1110 1011. There are three layers from right to left. The third layer code 1011 represents the horizontal stroke (一), the second layer code 1110 represents the left stroke (丿), and the first layer code 0001 represents the right stroke (捺). The stroke order is: horizontal, left-falling stroke, right-falling stroke, which leads to the Chinese character: "大".
[0254] The reverse derivation method is as follows: Starting from the rightmost end (the lowest bit) of the 4HPU codeword string, read the first 4-bit state, parse it, and translate the 4-bit binary number into the corresponding 4HPU basic state S1S4 or T1T4. Based on the basic state, deduce the meaning of the number, letter, stroke, or symbol at that layer, and restore the number, letter, stroke, or symbol represented by the 4 bits. Then proceed to the previous layer and repeat the above steps until all layers are parsed.
[0255] In Chinese character stroke recognition based on 4HPU coding, there is a special case where different Chinese characters have exactly the same stroke order path, such as horizontal, left-falling, and right-falling strokes. Simply relying on the full stroke code cannot distinguish different Chinese characters. To resolve this conflict problem of characters with the same stroke order but different meanings, the method of the present invention uses an identification code. The identification code is: after completing the normal full stroke code encoding, a 4-digit identification code is added as an additional layer, located at the leftmost end (highest bit) of the stroke order code string, to distinguish Chinese characters with the same stroke order but different shapes and meanings.
[0256] For example: Figure 7 and Figure 8 As shown in the figure, the Chinese characters "大" (big) and "丈" (foot) both have the stroke order: horizontal, left-falling, right-falling. Their 4HPU encoding string is: 0001 1110 1011. Without the identification code, the stroke order alone cannot distinguish "大" from "丈". Therefore, an identification code is used to distinguish them. This identification code is added after the basic stroke order code, as shown in Table 17.
[0257] Table 17 Significance of adding identification code
[0258]
[0259]
[0260] Among them, 0001 is used to recognize the character "大", and 0010 is used to recognize the character "丈".
[0261] The identification code is a supplementary information bit added to the end of the basic stroke order code, using a 4-bit or 8-bit 4HPU coding structure to resolve recognition ambiguity caused by stroke order conflicts. Identification code allocation rules:
[0262] Natural order rule: assign numbers in the order of first appearance starting from 0001 and increasing in sequence;
[0263] Simple numbering rule: When there are multiple Chinese characters with the same stroke order but different semantics, the priority is defined according to semantics or commonness, and the priority number is assigned;
[0264] Global interval reservation: Reserve a set of expandable identification code number intervals for characters with the same stroke order that may be added in the future. For example, 0001 0001 is a common character, and 0001 0001 1100 is an extended character. The code length can be from 4 to 12 bits to ensure that there is no conflict when the character library grows in the future.
[0265] In the reverse deduction, starting from the rightmost end of the encoded string, the basic stroke order encoding string is performed layer by layer, the stroke order path is confirmed, and the stroke order is analyzed. After the stroke order encoding is completed, it is determined whether the stroke order has ambiguity. If so, the tail identification code is read, and the corresponding Chinese character is uniquely determined based on the stroke order path + identification code combination.
[0266] When using 4HPU simple code for Chinese character recognition, the 4HPU simple code usually consists of the following parts:
[0267] The first level: the simple code flag, indicating the use of the simple code mode,
[0268] The second layer: the first stroke code, 4 bits,
[0269] The third layer: the last stroke code, 4 bits,
[0270] The fourth layer: identification code, 4 or 8 digits, indicating the number of strokes or identification features.
[0271] Reverse deduction steps: Starting from the rightmost end of the encoded string, read the 4-bit encoding units in sequence, confirm the first stroke and the last stroke information, extract the identification information, read the identification code, and extract the number of strokes or feature marks.
[0272] For example, taking the character "大" as an example, the encoding string is: 0100 1011 0001 0100, and the parsing process is as follows:
[0273] 0100 (Simplified code flag, indicating the use of simplified code),
[0274] 1011 (first horizontal stroke),
[0275] 0001 (last stroke),
[0276] 0100 (auxiliary identification code, indicating the number of strokes is 3).
[0277] Reverse deduction, read the number of strokes 3, search all 3-stroke Chinese characters in the standard character library, and select the Chinese characters with a horizontal stroke as the first stroke and a downward stroke as the last stroke. Preliminary screening shows that it may be "大" or "丈". Further, according to the recognition code, match the "大" character code and determine it as "大".
[0278] In rare cases, if the identification code is still not sufficient to make a complete distinction, contextual text retrieval is used, such as searching for the combination of "tall" and "vast", or the frequency priority principle of the language vocabulary, or prompting manual assisted selection.
[0279] The method of the present invention can be applied in the following fields:
[0280] 1. Intelligent error correction and completion, such as missing character recovery in optical character recognition (OCR). The 4HPU encoding structure constructed using the method of the present invention can perform reverse path reasoning on incomplete encoded character strings. Based on the first and last strokes, structural nesting, and identification codes, the most likely original Chinese character structure can be retrieved. After combining it with the language context, logical repairs or suggestions can be made, thereby improving the error correction capability and accuracy of OCR.
[0281] 2. Lightweight Natural Language Processing (NLP) in low-resource environments. Traditional large-scale models cannot run on terminals with limited computing resources, such as microcontrollers, embedded chips, and low-power AI devices. The 4HPU encoding structure of the present invention features a structural hierarchy and reversible logic, replacing the traditional token and word vector representation. It expresses language structure with "nested paths" and uses path distance and structural similarity to perform basic NLP reasoning, matching, and classification, without relying on high-dimensional embeddings or cloud computing power.
[0282] 3. Direct data recovery and reverse identification on the local device. In traditional systems, data recovery often relies on network transmission and cloud-based model inference, which poses latency and privacy risks. The method of the present invention provides an "encode → decode" mapping. After receiving the encoded string on the device, the method of the present invention can be used to restore the original characters, phrases, and values through a local built-in reverse path parsing module, completely eliminating reliance on the cloud and enabling semantic understanding and data reconstruction in edge computing or offline conditions.
[0283] 4. Intelligent Indexing and Similarity Matching System Based on Encoding Path. In tasks such as file retrieval, knowledge graph, and semantic similarity matching, the 4HPU structure encoding can play an important role. All information, such as characters, phrases, and logical expressions, is structurally encoded into nested paths. Using the "hierarchical pattern + structural direction + state activation" in the paths, a high-dimensional structure index space is constructed. Path analysis is performed on the newly input information, and rapid structural comparison and fuzzy matching are carried out with the encoded paths. It is applied to scenarios such as semantic search, similar formula matching, and Chinese character structure comparison.
[0284] 5. Reverse Path Verification in High-Efficiency Storage and Backup Systems. Data storage systems often need to verify data integrity and consistency. Traditional methods rely on hashing or redundant encoding. The nested path structure obtained by the method of the present invention is reversible, and each segment of the encoding can be restored. During the storage process, only the 4HPU encoded string needs to be saved, and the original information structure can be reversely deduced and verified for matching. For compressed data, anti-tampering verification at the structural level can be performed through the logical consistency of the structural paths, significantly reducing redundant fields. It is applicable to the fields of distributed storage, log tracking, and data compliance verification.
[0285] As Figure 9 shown, compared with the ASCII unidirectional linear numbering of the prior art, the encoding obtained by the method of the present invention has clear structural expression and path reasoning, can be dynamically extended and compressed, and provides a more excellent basic encoding method for application scenarios such as natural language processing, intelligent communication, and low-power devices. It can be widely applied in the technical fields of natural language structured representation, multi-modal intelligent processing, low-latency communication, and edge computing.
[0286] Specifically:
[0287] 1. Expandability,
[0288] Due to the use of linear numbering, the character set capacity of ASCII encoding is limited. It natively only supports 128 characters, and even after expanding to 256 characters, it is still insufficient to support the global multilingual system and must rely on external systems such as Unicode. 4HPU encoding: Based on a sparse 4-bit (or higher-bit) nested structure, the encoding space can be dynamically expanded as needed without destroying the existing encoding. Theoretically, it can be infinitely expanded and naturally adapts to multi-language, multi-symbol, and large-scale data environments.
[0289] 2. Structure and flatness. ASCII encoding has a unidirectional linear numbering, and there is no structural connection between characters, making it unable to express context, directionality, and combinatorial logic. 4HPU encoding: Each 4HPU unit has a clear directionality, such as →, ↑, ↓, ←. The encoding process naturally forms a structural path, and the whole shows a flat expansion, supporting multi-dimensional applications of natural language, graphic data, and logical reasoning.
[0290] 3. Directionality and path reasoning. ASCII encoding has no direction and no path, and cannot perform automatic reasoning or reverse derivation. 4HPU encoding: nested paths are generated during the encoding process, supporting reverse path recognition and derivation, and having natural advantages in intelligent systems such as natural language processing, visual recognition, and structural reasoning.
[0291] 4. Compressibility and storage efficiency. ASCII encoding has a fixed 8-bit character. Even when expressing simple characters, it occupies the full bit width, resulting in redundancy. 4HPU encoding flexibly encodes according to the number of nested layers. Simple characters only require one layer of 4 bits, and complex structures are extended through levels. As a whole, it has a high compression ability and has natural adaptability in low-power devices and edge computing scenarios.
[0292] 5. Application prospects and future trends. With the rapid development of globalization, intelligence, Internet of Things (IoT), and natural language processing (NLP) technologies, the limitations of traditional linear encoding systems such as ASCII in terms of flexibility and intelligence have become increasingly obvious. 4HPU encoding, with its infinite expandability, planar structure, path reasoning ability, and high storage efficiency, has become an important development direction for the underlying systems of future natural language encoding, structured communication, and intelligent computing; in emerging fields such as AI large models, smart cities, intelligent hardware, and quantum communication, the 4HPU encoding system can provide strong support for multi-modal data fusion, low-power consumption processing, and large-scale structure modeling.
[0293] The system of the present invention is a data expression that can replace traditional tokenization and binaryization. It is suitable for both AI model embedding and structural reasoning, and for hardware embedding and edge deployment, with strong practicality and scalability, laying a foundation for building future structural language systems, 4HPU computing devices, and AI operating systems.
Claims
1. A natural language character and numerical general solution system based on 4HPU, characterized in that: The natural language character and numerical general solution system based on 4HPU is used for encoding and processing the original data of numbers, letters, Chinese characters or symbols at the bottom layer of artificial intelligence, and is provided with a main encoding path for realizing encoding generation and structural nesting.
2. The natural language character and numerical general solution system based on 4HPU according to claim 1, characterized in that: The main encoding path consists of an input character module, a category detection module, a 4HPU encoder, a nested path generator and an output encoding string module; The input character module receives the original data of numbers, letters, Chinese characters and / or symbols input by the artificial intelligence character interface and sends it to the category detection module; After receiving the original data, the category detection module conducts identification, determines which category among numbers, letters, Chinese characters or symbols it belongs to, gives a category flag, and sends it to the 4HPU encoder; the category flag consists of 4 digits, and each digit of the 4 digits is 0 or 1; The 4HPU encoder encodes numbers, letters, Chinese characters or symbols in a nested manner from right to left to obtain an encoding string composed of at least one group of basic encoding units, and then outputs the encoding string and the category flag to the nested path generator; the basic encoding unit is 4 digits, and each digit of the 4 digits is 0 or 1, and only one digit of the 4 digits is 1; The numerical encoding adopts the structure of the arrangement of basic encoding units; The letter encoding is represented by symmetric positive and negative numerical values according to the upper and lower cases of English letters; The Chinese character encoding takes strokes as units, maps each stroke to a 4HPU basic encoding unit, and arranges and encodes them according to the stroke order to form an encoding string structure; The symbol encoding adopts an encoding string structure of 4 bits, 8 bits or 12 bits from the idle sparse encoding; The nested path generator arranges the category flag and the basic encoding unit in the order from right to left, and combines them to form a complete encoding string with a structural hierarchical relationship for the number, letter, Chinese character or symbol: category flag + data encoding string; The output encoding string module outputs the complete encoding string to the artificial intelligence data processing part outside the main encoding path.
3. The natural language character and numerical general solution system based on 4HPU according to claim 2, characterized in that: The strokes are dot, horizontal, vertical, left-falling stroke, right-falling stroke, rising stroke, turning stroke, and hook stroke. The encoding of dot, horizontal, vertical, and left-falling stroke and the encoding of right-falling stroke, rising stroke, turning stroke, and hook stroke form an inverse relationship of the basic encoding unit.
4. The natural language character and numerical general solution system based on 4HPU according to claim 3, wherein: For the symbol encoding, the idle 4-bit sparse encoding is assigned to arithmetic symbols, the idle 8-bit sparse encoding is assigned to comparison symbols, and the idle 12-bit sparse encoding is assigned to logical symbols and set symbols.
5. The natural language character and numerical general solution system based on 4HPU according to claim 4, wherein: The natural language character and numerical general solution system based on 4HPU is provided with a decoding path and an expansion path. The decoding path is used to realize the reverse recognition of the encoding, and the expansion path is used to realize the expansion of the logical process.
6. The natural language character and numerical general solution system based on 4HPU according to claim 5, wherein: The decoding path consists of a nested path generator, a reverse path parser and an identification result output module; The reverse path parser obtains the complete encoding string from the nested path generator, unfolds according to the reverse process of encoding numbers, letters, Chinese characters or symbols by the 4HPU encoder, restores the basic encoding unit to the original data, and restores the category flag to numbers, letters, Chinese characters or symbols; The recognition result output module outputs the restored original data, which are restored into numbers, letters, Chinese characters or symbols, outside the decoding path.
7. The natural language character and numerical general solution system based on 4HPU according to claim 5, wherein: The extended path consists of a nested path generator, an idle path allocator, a symbol extension output module and a reverse parser; The idle path allocator obtains the basic coding units not adopted by the complete coding string from the nested path generator, and expands and generates a new 4HPU coding bit allocation rule according to the idle sparse code set; The symbol extension output module outputs the new 4HPU code bit allocation rule outward.
8. A general solution method for natural language characters and numerical values based on 4HPU, comprising the following steps: I. Construct basic coding units Use 4-bit digital codes to construct basic coding units. Each bit in the 4-bit digital code can be 0 or 1, and only one bit in the 4-bit digital code is 1; The basic coding unit has four coding states: S1, S2, S3, S4. The coding corresponding to S1 is 0001, the coding corresponding to S2 is 0010, the coding corresponding to S3 is 0100, and the coding corresponding to S4 is 1000; II. Judge the categories of numerical values, letters, Chinese characters or symbols For the original data from a limited scale, respectively judge whether the original data belongs to the categories of numerical values, letters, Chinese characters or symbols, and give the category flags of numerical values, letters, Chinese characters or symbols; The category flag consists of 4-bit digital codes. Each bit in the 4-bit digital code is 0 or 1. The coding for the numerical category is 0001, the coding for the letter category is 0010, the coding for the Chinese character category is 0100, and the coding for the symbol category is 1000; The numerical values are natural numbers, such as 1, 2, 3, the letters are English uppercase and lowercase letters, such as A-Z, a-z, the Chinese characters are recording symbols of Chinese, and the symbols are mathematical operation symbols, comparison symbols, set logic operation symbols and punctuation control symbols; III. Coding Adopt a nested method from right to left to code natural numbers, letters, Chinese characters and / or symbols to obtain a complete coding string: category flag + data coding string; The numerical coding adopts a structural nested coding, adopting the structure of arranging basic coding units, that is, arranging another group of basic coding units after a group of basic coding units. The number of groups of basic coding units is determined by the size of the number; The letter coding is represented by symmetric positive and negative numerical values according to the uppercase and lowercase of English letters. It consists of 4-bit digital codes. Each bit in the 4-bit digital code can be 0 or 1, and only one bit in the 4-bit digital code is 1; The Chinese character coding takes strokes as units, maps each stroke to a 4HPU basic coding unit, and arranges and codes them according to the stroke order to form a coding string structure; The symbol coding adopts an available 4-bit, 8-bit or 12-bit coding string structure from the idle sparse coding.
9. The natural language character and numerical general solution method based on 4HPU according to claim 8, wherein: For the complete coding string, expand it according to the reverse process of coding, restore the basic coding units to the original data, and restore the category flag to numbers, letters, Chinese characters or symbols.
10. An application of the natural language character and numerical general solution method based on 4HPU described in claim 8, characterized in that: The natural language character and numerical general solution method based on 4HPU is applied to intelligent error correction and completion, direct data recovery and reverse recognition on the local device side, lightweight natural language processing in a low-resource environment, intelligent indexing and similarity matching system based on the coding path, or reverse path verification in an efficient storage and backup system.
Citation Information
Cited By
Intelligent data processing system based on large model
CN121071023A
A large model-based data intelligent processing system
CN121071023B
Literal translation method and system based on 4HPU structure coding
CN121579076A