Literal translation method and system based on 4HPU structure coding
By using the 4HPU structure encoding method, Chinese semantics, symbol sequences, graphic data, and data blocks are directly mapped into sparse activation path units, which solves the problems of redundant calculations and low efficiency in the execution of Chinese semantics at the computer level in the existing technology, and realizes efficient and stable direct mapping and execution.
Patent Information
- Application Number
- CN202511773910.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies suffer from redundant computation and inefficiency in the mapping process from Chinese semantics to the underlying execution of computers, making it impossible to directly execute natural language instructions or structured semantic information.
By adopting a 4HPU-based structure encoding method, Chinese semantics, symbol sequences, graphic data or data blocks are directly converted into sparse activation path units by setting basic encoding units, Chinese characters, symbols, graphics and data block encoding, generating an executable instruction chain and achieving a cross-domain consistent instruction chain.
It achieves efficient, low-power, and traceable direct mapping from Chinese semantics to computer operating systems, electro-optical display control, and communication and storage, avoiding intermediate interpretation layers and improving execution efficiency and stability.
Smart Images

Figure CN121579076A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an encoding method and system for natural language input to operation execution in computer systems and artificial intelligence, and in particular to a direct translation method and system based on 4HPU structure encoding. Background Technology
[0002] In scenarios involving natural language interaction, operating system instruction execution, electro-optical display control, and integrated communication and storage encoding processing, a direct mapping from Chinese semantics to the underlying computer execution is required. Existing computer architecture and artificial intelligence methods for instruction parsing and task execution (literal translation or computation) primarily rely on tokenization and probability distribution mechanisms, requiring intermediate steps such as word segmentation, embedding, and probability prediction to complete the literal translation task. In complex semantic scenarios like Chinese, this approach often leads to redundant computation, inefficiency, and semantic bias. Furthermore, at the underlying execution levels of computer operating systems, display devices, and communication and storage, tokenization methods require an additional interpretation layer between the language and execution layers, making it impossible to execute commands simply by inputting natural language instructions or structured semantic information. Summary of the Invention
[0003] The purpose of this invention is to provide a direct translation method and system based on 4HPU structure encoding, and the technical problem to be solved is to improve the efficiency from input to execution.
[0004] This invention adopts the following technical solution: a direct translation method based on 4HPU structure encoding, comprising the following steps: I. Setting up the 4HPU structure code A set of basic coding units is constructed using four digits. Each digit in the four digits is either 0 or 1. After the four digits are combined, only one digit is 1 and the rest are 0. II. Setting Chinese characters, symbols, graphics, and data block encoding The Chinese encoding is divided into three levels according to the semantics of Chinese characters: characters, words, and sentences. Specifically, characters are mapped using basic encoding units based on strokes or components; words are combined according to the basic encoding units of the individual characters that make up the word; and sentences connect the basic encoding units of words according to grammatical relationships. The symbol encoding is based on the symbol's punctuation class, mathematical class, logical class, path class, and control class. After generating the corresponding basic encoding units and forming path units, a category label and a symbol identification code are assigned to the end node position of the path unit. The graphic encoding decomposes the elements constituting the graphic into a three-level structure encoding in spatial order: points, lines, and surfaces; or strokes, components, and the whole. The data block encoding is as follows: the data block is divided into independent blocks according to the logical boundaries or logical length of the computer system and artificial intelligence field. Before encoding each independent block, a category label, data block number and integrity verification information are added to the beginning of the block. The data content of the independent block is mapped byte by byte into multiple sets of basic encoding units. III. Receiving Input Information Receive Chinese semantics, symbol sequences, graphic data, or data blocks represented in text encoding, symbol sequences, vector paths, or binary block formats from the human-computer interaction interface, application system interface, or external device sensing end of the computer system and artificial intelligence; IV. Encoding Mapping Transform Chinese semantics, symbol sequences, graphical data, or data blocks into basic coding units. The path unit obtained after the Chinese semantic code is composed of a four-digit combination of status number, direction label and semantic label in sequence, where only one digit is 1 and the rest are 0; The boundary ±0 identifier of the symbol is obtained, which can be combined with the Chinese semantics and data block to form a path unit in which only one bit is 1 and the rest are 0; The obtained graphical data consists of a four-digit combination of status number, direction label and semantic label, in which only one digit is 1 and the rest are 0; The path sequence obtained is a four-digit combination of a data block containing a path number, a category identifier, and a check digit, where only one digit is 1 and the rest are 0. V. Path Analysis Obtain the state number layer from the path unit and path sequence to confirm the activation direction and node position; Obtain the directional label layer from the path unit and path sequence, and calculate the path flow direction and connection relationship; Semantic tag layers are obtained from path units and path sequences, and semantic nodes are mapped to execution instructions or state descriptions. The action category labels are mapped to execution instructions, the object and parameter category labels are mapped to input parameters, and the modifier category labels are mapped to state descriptions, ultimately generating an executable instruction set. VI. Instruction Mapping The instruction set is directly generated into an executable instruction chain, which is a Chinese semantic operating system call instruction, electro-optical display instruction, or communication and storage instruction based on the basic coding unit; VII. Implementation and Feedback The results of instruction mapping are output to the kernel interface of computer system and artificial intelligence, electro-optical display control module or communication and storage execution module.
[0005] In the method of the present invention, all four bits of the code are 0, which represents a fully static inactive state and is marked as +0. All four bits of the code are 1, which represents a fully active saturated state and is marked as −0.
[0006] The basic coding unit of the method of this invention has four states: State 1, coded as 0001 State 2, coded as 0010 State 3, coded as 0100 State 4, coded as 1000.
[0007] The category label of the method of the present invention identifies the semantic type of the path unit, and the symbol identification code is a unique number assigned to each symbol.
[0008] The graphic elements of the method of this invention are encoded according to their geometric direction order: Right → Code is 0001 The top right corner ↗ is coded as 0010. The code for "Up" (↑) is 0100. The top left corner is coded as 1000. The left arrow (←) is encoded as 1110. The bottom left corner (↙) is coded as 1101. The down arrow (↓) is encoded as 1011. The bottom right corner is coded as 0111.
[0009] The method of this invention adds ±0 markers at the end of the encoding position for closed area graphics or graphic boundaries.
[0010] The data block sequence number in the method of the present invention is a number assigned sequentially when the data block is divided into blocks, and the integrity verification information is a file class, network class, and / or cache class identifier.
[0011] The data block of the method of the present invention is an information unit existing in the form of binary data or structured data, carrying file content, network transmission packets, sensor data, database records or cached fragment content.
[0012] The basic coding unit of the method of this invention constitutes the basis of the 4HPU structure coding.
[0013] A system based on 4HPU structure encoding for implementing the method of the present invention comprises an input interface module, a 4HPU encoding module, a path parsing module, an instruction mapping module, and an execution and backtracking module.
[0014] Compared with existing technologies, this invention encodes and parses Chinese semantics, symbol sequences, graphic data, and data blocks using the same sparse activation rules and path mapping mechanism. This enables cross-domain reuse of the same structure encoding system across the operating system instruction layer, electro-optical display control layer, and communication and storage transmission layer, achieving a consistent instruction link across domains and ensuring uniqueness, traceability, and low-power operation. Attached Figure Description
[0015] Figure 1 This is a basic form of the 4HPU structure encoding and a schematic diagram of sparse activation of the present invention.
[0016] Figure 2 This is a schematic diagram of the S- and Z-shaped path expansion of the present invention.
[0017] Figure 3 This is a three-layer structure diagram of a non-token path parser.
[0018] Figure 4 This is a flowchart of the execution of the direct translation calculation of the present invention at the operating system layer.
[0019] Figure 5 This is a schematic diagram of the link in the electro-optic display layer for the direct translation calculation of the present invention.
[0020] Figure 6 This is a schematic diagram of the direct translation computation of the present invention in the communication and storage layers. Detailed Implementation
[0021] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The direct translation method based on 4HPU structure encoding of the present invention includes the following steps: I. Setting up the 4HPU structure code like Figure 1 As shown, a set of basic coding units (4HPU state units, 4HPU encoding) is constructed using four digits. These basic coding units form the basis of the 4HPU structure encoding. Each digit in the four digits can be 0 or 1. After the four digits are combined, only one digit is 1, and the rest are 0. Thus, the basic coding unit forms four states: State 1 S1, coded as 0001, is illustrated by four dots arranged at the four corners of a square, with a black dot at the bottom left.
[0022] State 2 S2, coded as 0010, is illustrated by four dots arranged at the four corners of a square, with a black dot in the lower right corner.
[0023] State 3 S3, coded as 0100, is illustrated by four dots arranged at the four corners of a square, with the black dot in the upper left.
[0024] State 4 S4, coded as 1000, is illustrated by four dots arranged at the four corners of a square, with a black dot in the upper right corner.
[0025] The basic coding unit has only one bit set to 1, also known as sparse activation.
[0026] In this invention, a structured sequence composed of several 4HPU state units arranged in order is called a path. The path is used to represent the trajectory or logical order of the input information's state in a two-dimensional (or three-dimensional) structural space. Each path constitutes a set of changing states, which can be written as: path P=(Si→Sj→Sk→…→Sn), where the symbol "→" indicates the directional connection between states, and i, j, ..., n represent the states represented by the 4HPU state units (nodes) in the trajectory or logical order.
[0027] The smallest constituent element of a path is called a path unit. Each path unit corresponds to an independent sparse activation state (S1~S4) and its direction attribute.
[0028] When multiple path units are activated consecutively, a complete path is formed, which is used to express semantic changes or execution flow.
[0029] Relationship between states and paths: Each state S1~S4 represents a fixed directional base point in space: S1 corresponds to the bottom left direction, S2 corresponds to the bottom right direction, S3 corresponds to the top left direction, and S4 corresponds to the top right direction.
[0030] like Figure 2 As shown, when the state changes continuously along S1 → S2 → S3 → S4, the path in the structural space manifests as a "Z-shaped or S-shaped" connected trajectory. This connected trajectory is the structured propagation method of semantics, symbols, or instructions using 4HPU state units. Therefore, the state is the static expression of the path nodes, and the path is the dynamic combination of states. The state provides spatial location, and the path provides temporal or logical order; together, they constitute the structural basis for semantic translation.
[0031] In a four-bit code, only one bit is 1, while the other three are 0, indicating that the code is in a normal activation state, such as 0001, 0010, 0100, 1000. This represents a positive activation state of the path, called positive sparse activation, and the code is called positive sparse activation encoding. In positive sparse activation encoding, a "1" represents that the node is activated and participates in semantic signal transmission within the current computation cycle, while a "0" represents that the bit is in a resting state and does not participate in the energy or information flow of the current path. The unique "1" in the four-bit code forms a directional activation vector, indicating the direction of information flow propagation and the focus of computation. Therefore, positive activation not only represents the geometric direction of the path but also corresponds to the energy transfer process of information flowing from the input source to the output in computer systems and artificial intelligence systems. Semantically, positive activation represents the advancement of semantics from the starting node to the target node. Logically, it represents the unidirectional execution flow of the computation task from receiving → parsing → executing. Through the positive sparse activation mechanism, computer architectures and artificial intelligence systems can activate paths with minimal energy, achieving efficient, loosely coupled, and controllable structural execution.
[0032] If the positive sparse activation code is inverted bitwise, changing the original 1 to 0 and the other three bits to 1, such as 1110, 1101, 1011, 0111, this represents the reverse (compensation) activation state of the path, called negative sparse activation. Negative sparse activation indicates that after executing the positive information flow, the path enters a reverse feedback and stabilization compensation phase. Changing from "0" to "1" means the corresponding node is reactivated to receive or balance the residual information or state feedback from the previous cycle; changing from "1" to "0" indicates that the node has completed its information output and entered a resting state. In the overall computer system and artificial intelligence system, negative sparse activation forms a reverse path symmetrical to positive sparse activation, used for result feedback, path verification, and state reset in the computer system and artificial intelligence system, ensuring that the semantic flow and execution flow can form a closed-loop feedback and dynamic balance after completing one translation. This symmetrical setting gives the 4HPU structure coding system directional complementarity and path reversibility, ensuring the stability and traceability of the computer system and artificial intelligence system while maintaining execution efficiency.
[0033] If all four bits of the code are 0: 0000, it indicates a completely static inactive state, marked as +0. If all four bits of the code are 1: 1111, it indicates a fully active saturated state, marked as −0. Both the completely static inactive state and the fully active saturated state are marked as ±0. ±0 is a special marker in sparse activation, used for path start / end, boundary, or check identifiers.
[0034] The encoding formed by arranging at least two sets of basic coding units is called 4HPU structure encoding. The paths and path units composed of basic coding units are collectively referred to as the 4HPU structure encoding system.
[0035] II. Setting Chinese characters, symbols, graphics, and data block encoding 1. Chinese encoding In natural language information, Chinese characters are the smallest semantic units, which can be used to form phrases and sentences to express input content such as actions and objects.
[0036] The Chinese encoding rules are based on the semantics of Chinese characters, divided into three hierarchical structures: character, word, and sentence. Among them, The characters are mapped using basic coding units based on strokes or components, and are encoded according to the Chinese character encoding method with application number 202510638608.5.
[0037] Words are combined according to the basic encoding units of the characters that make up the word, such as action, object, parameter, and modifier, and arranged to form an encoding. Each character first generates its corresponding basic encoding using the 202510638608.5 Chinese character encoding method, and then logically arranges them according to semantic functional relationships (such as action class, object class, parameter class, and modifier class) to form a word-level encoding sequence.
[0038] The word encoding is arranged in a four-layer structure: subject, object, complement, and modifier. Each layer corresponds to a combination of one or more character-level encoding units. For example: "Open" is an action word, encoded as [S1→S2], which represents the action path from initiation (S1) to execution (S2).
[0039] "File" is an object class term, encoded as [S3→S4], representing the mapping path of a static target.
[0040] "Fast" is a modifier, encoded as [S1→S4→S2], representing the accelerated path from the initial state to the final state.
[0041] "For user" is a parameter-type phrase, encoded as [S2→S3→S1], which indicates the action pointing from the execution node to the target node.
[0042] When a word consists of multiple characters, its character-level encoding is concatenated and semantically weighted to generate a word-level sparse activation path sequence. For example, the path concatenation and semantic weighting of "open file" is: [S1→S2]+[S3→S4], and the generated word-level sparse activation path sequence is: [S1→S2→S3→S4]. This combined structure reflects the semantic hierarchy and logical order within the word, ensuring that each word has a unique representation in the 4HPU encoding space, thus allowing it to be directly mapped to action, object, or modifier-type instructions during translation (execution).
[0043] Sentences are formed by connecting words according to grammatical relationships, creating a directed path chain. A directed path chain is a sequence of paths formed by connecting multiple word-level codes in the order of syntactic structure: subject, predicate, object, and modifier. This path sequence is directional, with each connection point linking the terminating state of the preceding word with the starting state of the following word, thus expressing the semantic order and logical dependencies within the sentence.
[0044] The encoding rule for a directed path chain C is: C = (P1→P2→P3→…→Pn), where P1, P2, P3… are word-level paths. The symbol “→” remains consistent with the previous text, representing the directional connection between path nodes or word-level paths. At the sentence level, grammatical relations such as “subject-verb,” “verb-object,” and “modifier-headword” are further refinements of directional connections at the semantic level; that is, grammatical dependency structures are abstracted from the directional connections of paths. Therefore, sentence-level path chains can be directly converted into semantic execution chains during execution. For example, the sentence: The phrase "System quickly opens folder" uses "System" as the subject (path [S3→S4]), "Quickly" as the modifier (path [S1→S4→S2]), "Open" as the action (path [S1→S2]), and "Folder" as the object (path [S3→S4→S2]). Therefore, the directed path chain C for "System quickly opens folder" is: [S3→S4]→[S1→S4→S2]→[S1→S2]→[S3→S4→S2]. The rearranged directed path chain C is: [S3→S4→S1→S4→S2→S1→S2→S3→S4→S2]. This invention can directly generate the corresponding Chinese execution instruction chain based on the path chain: speed, open, folder. Each action node can correspond to the path direction, achieving a direct Chinese translation from semantics to execution, without needing English symbols or intermediate interpretation layers. In this way, the syntactic path structure and execution logic maintain a one-to-one correspondence, realizing a direct Chinese translation computation process where "statement is path, path is execution." From semantics to execution, the actions, objects, and parameters in the semantics are mapped to corresponding Chinese instructions according to the path sequence, and the system interfaces of the computer system and artificial intelligence are directly invoked to execute them in a predetermined order, so that the semantic meaning is directly executed. At the end of each level, a category label and path number are appended.
[0045] Category labels are used to identify the semantic type of a path unit. In Chinese encoding, category labels include "character class," "word class," and "sentence class," used to distinguish hierarchical semantics. In symbolic encoding, category labels include "operator class," "logical symbol class," and "punctuation mark class," used to identify operations and logical relationships. While the category labels in Chinese encoding and symbolic encoding are not entirely identical, they both map to a unified structure field.<type>, to maintain the consistency of the hierarchy identification.
[0046] Path number is used to identify the unique position and execution order of the path. Each path number is composed of a hierarchy prefix and a serial number, for example: word-level path number: Z-01, Z-02; word-level path number: C-01, C-0; sentence-level path number: J-01, J-02. Among them, the prefixes Z, C, J represent word, word, and sentence respectively, and the numbers represent the execution order of the hierarchy path. Path number is bound to trace_id, used to maintain consistent indexing in execution and backtracking.
[0047] Each hierarchy is converted into a sparse activation path unit, which refers to the smallest semantic unit in 4HPU encoding, composed of an activated encoding bit and its corresponding category label and path number. In the smallest semantic unit, the activation state is used to represent the position and role of the current encoding unit in the four basic directions, the category label is used to identify the belonging category of the smallest semantic unit in the semantic layer or functional layer, such as word, word, sentence or symbol, and the path number is used to identify the unique position of the unit in the path chain and support backtracking and execution order management. Each sparse activation path unit can be identified, transmitted or executed independently. A plurality of smallest semantic units are combined in semantic order to form a path chain from input to output. The path chain directly reflects the execution process of the semantics and is the basic structure of the direct translation of language information in the computing layer. In this way, the structure, function and role of the path unit can be clearly defined without relying on symbolic formulas, thereby ensuring the natural language integrity and engineering reproducibility of the specification.
[0048] The process of converting hierarchy to sparse activation path unit is as follows: 1. Receive the sequence of encoding hierarchy (word, word, sentence), 2. Add category label and path number at the end of the hierarchy, 3. Write the sequence as 4HPU sparse activation path [S1→S2→S3→S4], 4. Determine the activation direction according to the context dependency (such as subject-predicate, modification, object), 5. Output sparse activation path unit as computer system and artificial intelligence input.
[0049] The final output is: semantic path of word; semantic combination and functional state of word; logical chain and execution instruction of sentence. Together they form the Chinese semantic direct translation (execution) instruction stream of the input computer system and artificial intelligence system. Chinese semantic execution instruction stream can be directly called, displayed, drawn or communicated by the operating system of computer system and artificial intelligence, realizing seamless mapping from language to execution.
[0050] 2. Symbolic encoding Symbol is a symbol of auxiliary written language, non-textual identification of operation relationship, logical structure, file identification, path separator and control instruction, including punctuation, mathematical symbol, logical operator, file symbol and specific control character, used to assist the expression of Chinese semantics and the accurate definition of operation instruction. The specific control character is the path indicator and the control instruction symbol.
[0051] Symbol coding generates its corresponding basic coding unit according to the function category of the symbol, punctuation, mathematics, logic, path, and control, and forms a path unit. Then, a category label and a symbol identification code are assigned to the end node of the path unit. The symbol identification code is a unique number assigned to each symbol, used to maintain consistency during coding, identification and execution. Each symbol is mapped to a fixed sparse activation template, which is the standard activation form of the corresponding symbol in 4HPU coding, used to unify the activation mode of the same type of symbol. The standard activation form in 4HPU coding is: only one bit is 1 and the others are 0 in four-bit coding, in order 0001, 0010, 0100, 1000, corresponding to the positive sparse activation of four basic directions; their reverse complements are 1110, 1101, 1011, 0111, corresponding to the reverse sparse activation of four directions; in addition, boundary states +0 (0000) and -0 (1111) are used to represent static or closed state.
[0052] The steps of symbol mapping to a fixed sparse activation template are as follows: 1. Identify the category of the symbol, 2. Call the standard sparse activation template corresponding to the category, 3. Activate the corresponding bit state at the end node of the path unit, which refers to the bit set to "1" in the 4HPU sparse activation template used by the symbol, used to indicate the activation position of the symbol in the path structure.
[0053] 4. Add symbol identification code and category label at the end node of the path unit, complete the mapping.
[0054] Through the above steps, each symbol is fixedly mapped to a unique sparse activation form in the coding stage, realizing the corresponding consistency of symbol and path structure, and ensuring that the symbol can be directly identified and executed between the two layers of language and calculation.
[0055] If the symbol appears at the boundary position of the semantic chain, ±0 mark is added at the beginning or end node of the coding path to distinguish the level. The semantic chain is a path sequence connected according to the syntax and logic order, used to express the complete semantic relationship. The boundary of the semantic chain refers to the starting point and the end point of the semantic unit, i.e. the beginning and end nodes of the path.
[0056] When the symbol is mixed with Chinese or numbers, the sparse activation path unit is arranged in the order of symbol coding first, then Chinese or number coding.
[0057] 3. Graph coding The graph is a visual information unit composed of interface element (image element) points, lines, surfaces, strokes and / or vector paths, used to express the content of shape, structure, direction and spatial relationship, which can represent both static images and dynamic drawing tracks. The graph is a geometric extension of Chinese characters and symbols.
[0058] The graph coding is to decompose the elements constituting the graph into points, lines, surfaces, or strokes, components and whole in spatial order, and each graph element is mapped to 4HPU sparse activation coding according to its geometric direction order: Rightward → coded as 0001, Right up ↗ coded as 0010, upward ↑ coded as 0100, Left up ↖ coded as 1000, Leftward ← coded as 1110, Left down ↙ coded as 1101, Downward ↓ coded as 1011, Right down ↘ coded as 0111.
[0059] When strokes are continuous or connected across regions, as shown in Figure 2 , S and Z type path extension is adopted to maintain the continuity of path direction. Among them, Continuous strokes are the connection form of adjacent strokes connected head to tail in the same region with consistent direction.
[0060] Cross-region connection is the connection of strokes or paths from one coding region to adjacent regions. The region is the four-directional quadrant space divided in 4HPU coding, and the cross-region connection refers to the connection of strokes with the starting point and the ending point belonging to different quadrants.
[0061] S-type path extension is a continuous path that bends in the shape of "S" between adjacent quadrants. S-type path extension is the extension rule of the path when multiple strokes are continuous.
[0062] Z-type path extension is a path that returns in the shape of "Z" between adjacent quadrants. Z-type path extension is the extension rule of the path when multiple regions are connected.
[0063] For closed region graphs or graph boundaries, ±0 markers are added at the end of the coding to represent the start and end state.
[0064] All graphic elements are converted into sparse path units after 4HPU encoding, and category labels, position parameters and hierarchical identifiers are attached to the end node positions of the encoding to form structured path data that can be coordinated with Chinese and symbolic semantics. Coordination refers to the process of identifying and operating together in the same path order during semantic analysis and execution of graphic, Chinese and symbolic encoding.
[0065] The final format of graphic elements, Chinese and symbolic encoding is represented as a unified path sequence structure: category label, path number, sparse activation state.
[0066] Sparse path unit is the smallest path structure unit with only one active bit "1" after 4HPU encoding. Its boundary state includes ±0: positive zero (0000) represents static starting point or unexcited state, and negative zero (1111) represents full excitation or path closure state. Normal activation state and ±0 together constitute the complete state set of path execution.
[0067] Position parameter is the coordinate information of the unit in the two-dimensional or three-dimensional encoding matrix, used for positioning and combination.
[0068] Hierarchical identifier is its boundary state including ±0: positive zero (0000) represents static starting point or unexcited state, and negative zero (1111) represents full excitation or path closure state. Normal activation state and ±0 together constitute the complete state set of path execution.
[0069] Structured path is a path chain formed by the sequential connection of multiple sparse path units, used to realize the unified representation and analysis of multi-modal information. Multi-modal information refers to the combined expression of different types of input in the same encoding system, including Chinese, symbols, graphics and data blocks.
[0070] 4. Data block encoding Data block is a unit of information in the form of binary data or structured data, used to carry file content, network transmission packet, sensor data, database record or cache fragment content. Binary data or structured data has fixed or variable length, and belongs to the smallest data entity that can be recognized, stored and transmitted by computer system and artificial intelligence. Structured data is a collection of data organized in the form of fields or key-value pairs, each field has a name, type and value, and is a standardized data format for parsing and indexing. Key-value is a pair of records composed of "field name (key)" and its corresponding "field content (value)".
[0071] Formally organized data collection is a group of data items arranged in order or hierarchical relationship.
[0072] Standardized data format is a data layout method that can be stably parsed and accessed by computer system and artificial intelligence system according to the specified field structure.
[0073] Data block encoding is: the data block to be input is divided into independent blocks according to the logical boundary or logical length in the field of computer system and artificial intelligence, and a category label, a data block serial number and integrity check information are attached to the head position of each independent block before encoding.
[0074] The logical boundary is a demarcation point that can be independently identified in content, protocol or task division.
[0075] The logical length is a fixed or variable byte length set according to the data type or processing capacity.
[0076] The data block serial number is a number assigned in sequence when the data block is divided.
[0077] The integrity check information is a file class, network class and / or cache class identifier. The file class is taken from the file header or the check code field, the network class is taken from the communication protocol header, and the cache class is taken from the check area of the memory block or cache page.
[0078] The data content of the independent block is mapped byte by byte into a plurality of groups of basic encoding units (4HPU sparse activation units) according to a predetermined segment length, and the predetermined segment length is set according to the data size and cache capacity in the initialization or task loading stage.
[0079] The sparse activation unit is the smallest encoding unit with only one bit being "1" after 4HPU encoding. When the head and tail of the block are data boundaries, ±0 markers are added to the head and tail of the independent block to distinguish the boundaries. If the data volume of the independent block is large, and the head and tail of the block are data boundaries, then multiple four-bit combination expansion is used to realize continuous path encoding. The multiple four-bit combination expansion arranges and combines multiple groups of 4HPU encoding units to form an expanded code to represent continuous data. The continuous path encoding is a sparse activation path sequence that maintains the same direction and can be sequentially executed after multiple expansion. A structured sparse path sequence with path number, category identifier and check bit is generated. It can be directly involved in direct translation execution and backtracking verification in communication, storage or calculation process.
[0080] The category identifier is a field indicating the category to which the data block belongs, and is generated according to the data source. The check bit is a parity or hash check result calculated according to the data content. The structured sparse path sequence is a set of continuous sparse activation paths with path number, category identifier and check bit, which is used for direct translation execution and backtracking verification in communication, storage or calculation process.
[0081] The method of the present application, direct translation refers to the process of directly parsing and executing the input information according to the 4HPU structure path without intermediate conversion layer. The 4HPU structure path sequentially parses the state number, direction label and semantic label according to the order of sparse activation encoding, and directly generates the corresponding instruction according to the parsing result and executes it in sequence.
[0082] III. Receiving input information Receiving Chinese semantics, symbol sequences, graphical data or data blocks represented in text encoding, symbol sequences, vector paths or binary block format from computer system, artificial intelligence human-computer interaction interface, application system interface or external device sensor input.
[0083] IV. Encoding mapping Chinese semantics, symbol sequences, graphical data or data blocks are converted into basic encoding units (sparse activation encoding), specifically: 1. When the input information is Chinese semantics, the input content is divided into three levels of word, word and sentence according to the semantic hierarchy: Word level: Based on the strokes or components of Chinese characters, each stroke direction is mapped to eight groups of positive and negative sparse activation units, respectively, 0001 for horizontal, 0010 for vertical, 0100 for stroke, and 1000 for n. The corresponding inverse strokes are bitwise complement.
[0084] Word level: According to the semantics of Chinese words, such as action, object, parameter, modification, at least two Chinese characters that constitute a word are combined in order. Sparse units are the smallest semantic encoding units that only activate one state bit after 4HPU encoding, used to represent the basic structure of a single Chinese character or symbol.
[0085] Sentence level: At least two word level paths are connected into a complete semantic chain according to the grammatical relationship, and ±0 markers are added at the beginning and end of the sentence to identify the start and end of the semantics and the syntactic boundary.
[0086] After 4HPU encoding, sparse activation path units composed of state number, direction label and semantic label are formed, realizing the unique mapping of Chinese semantics to structural paths.
[0087] State number is the serial number that identifies the active position of the current encoding unit in the 4HPU structure, used to distinguish different state points.
[0088] Direction label is an identification field that represents the running direction of the path in the four basic directions left, right, up and down.
[0089] Semantic label is an identification information that identifies the function of the encoding unit in the semantic layer, such as subject, predicate, object and modification.
[0090] Structural path is a directed path formed by sequentially connecting multiple sparse activation path units, used to represent the execution route of semantics.
[0091] 2. When the input information is symbols, the symbols are classified into punctuation, mathematical, logical operator, file symbol and special control character according to their functions. Each symbol is assigned a class label and a symbol identification code at the start of the path unit before encoding, which is mapped to a fixed sparse activation template: The fixed sparse activation template for punctuation is 0001, The fixed sparse activation template for mathematical symbol is 0010, The fixed sparse activation template for logical operator is 0100, The fixed sparse activation template for file symbol is 1000, The fixed sparse activation template for path indicator is 1110, The fixed sparse activation template for control instruction class symbol is 0111.
[0092] When the symbols appear in sequence, the paths of the symbols are combined in order of appearance, with the boundaries marked by ±0, to form sparse activation path units that can be combined with Chinese semantic and data blocks.
[0093] 3. When the input information is graphics, the input graphics are decomposed into points, lines, surfaces or strokes according to their spatial structure, and their geometric direction, start and end position and closure relationship are extracted.
[0094] The direction information (→,↗,↑,↖,←,↙,↓,↘) of each point or stroke is mapped to eight sets of positive and negative 4HPU sparse activation units, where the positive strokes use single set activation code: 0001, 0010, 0100, 1000, and the negative strokes use its bitwise complement: 1110, 1101, 1011, 0111.
[0095] When strokes or line segments are connected continuously across regions, S and Z type path expansion is used to maintain directional continuity. For closed regions or graphic boundaries, ±0 markers are added at the start and end positions to distinguish the start and end states.
[0096] If the graphic data is a dot matrix or vector composite input, each element is converted into a sparse activation path unit in sequence, with position coordinates and level labels added. The specific steps are: 1. Scan the graphic data in row or vector order, 2. Convert each point, line or surface element into a sparse activation path unit in sequence, 3. Add position coordinates and level labels to the converted units, 4. Combine the complete path sequence in the order of scanning.
[0097] After encoding, the final result is a sparse activation path sequence composed of state numbers, direction labels, and semantic labels, which is used for subsequent path parsing and instruction mapping. The sparse activation path sequence is a continuous path chain formed by connecting multiple sparse activation path units in temporal or spatial order. The difference between it and a single sparse activation path unit is that the former represents the overall structure or graph sequence, while the latter only represents a single activation node.
[0098] 4. When the input information is a data block, the data block format is either a binary stream or a key-value pair format. Based on logical boundaries or logical length, the data block is divided into independent blocks, and the additional category label, data block sequence number, and integrity verification information of the data block are extracted from the data block header. After extraction, these are written into the encoding table header and participate in the subsequent path mapping and verification process.
[0099] Before each data block enters the encoding process, a category label is appended to the beginning of the block. The category label is used to identify the type of data block, such as file, network, or cache. It corresponds to the category source of the aforementioned integrity verification information, but the functions are different. The former is used for type identification, while the latter is used for data verification. The data block sequence number is used for sequential identification.
[0100] The data content of independent blocks is sequentially mapped to 4HPU sparse activation units, either by bytes or fixed segment length. Each 8 bits of data is mapped to at least two sets of four-bit sparse activation codes. When the data stream is transmitted in the forward direction, the forward bits are 0001, 0010, 0100, and 1000; during readback, verification, or reverse parsing, the reverse bits are 1110, 1101, 1011, and 0111. This maintains the symmetry between the bit-level correspondence of the coding layer and the structural layer. The bit-level correspondence is a one-to-one mapping between data bits and 4HPU code bits. The structural symmetry is a mirror image of the two sets of sparse activation codes in the path direction and state distribution.
[0101] When data blocks are input continuously or when large amounts of data are segmented and sliced, ±0 markers are added at the beginning and end of the data blocks to indicate the boundaries and the sequential reorganization relationship of the data blocks.
[0102] When the data size of a data block exceeds the upper limit of the segment length of an independent block, such as 256 bytes, or when encoding cannot be completed within one cache cycle, multiple four-bit combination extensions are used to achieve continuous path encoding.
[0103] After encoding, a structured sparse activation path sequence with path number, category identifier, and check bit is generated. The structured sparse activation path sequence is a complete path chain built upon the sparse activation path sequence, with the addition of path number, category identifier, and check information, used for identification and execution. The difference between it and the sparse activation path sequence is that the former includes traceability and check fields, possessing execution and verification functions; the latter is merely a basic activation state sequence. The structured sparse activation path sequence can directly participate in the translated (execution) instruction mapping and trace_id backtracking verification during communication or storage execution.
[0104] V. Path Analysis Path parsing is the process of sequentially reading and semantically reconstructing the sparse activation units in a structural path, used to extract executable semantic information from the encoded path. The relationship between path parsing and structural path is that the former is the parsing and execution process of the latter, while the latter is its static encoded form.
[0105] 1. Path and Path Unit A path is a structured sequence formed by connecting at least two node units after 4HPU sparse activation encoding in a temporal or logical order. It is used to express the trajectory of Chinese characters, symbols, graphics, or data block input information in the structure space.
[0106] The node unit is the smallest coded point with a unique state number after 4HPU sparse activation coding.
[0107] A structured sequence is a directed coding chain formed by sequentially connecting at least two node units.
[0108] The structural space is the multidimensional coordinate domain (the area covered by the coordinate system) of the path mapped by 4HPU encoding.
[0109] The trajectory is the direction and order in which the input information propagates along the path sequence.
[0110] A path unit is the smallest constituent node in a path, containing three basic attributes: state number, direction label, and semantic label. Each path unit can independently represent a local semantic state or be combined with adjacent units to form a complete semantic flow or instruction chain.
[0111] Semantic state refers to the semantic meaning or functional position of a single path unit within its context. Semantic flow is a continuous semantic expression process formed by connecting multiple path units in a logical order. Instruction chain is an executable instruction sequence generated after mapping the semantic flow. The difference between a path unit and sparse path units or sparse activation path units is as follows: a path unit is a general definition that emphasizes structural composition; a sparse path unit is a simplified form in which only one state bit is activated; a sparse activation path unit is a path unit with a clear activation state and execution attributes, used for direct interpretation and execution.
[0112] 2. Path Representation Rules The path is represented node-to-node, denoted as: P = (S1→S2→S3...→Sn), where each node Sᵢ corresponds to a sparse activation code. Here, S is the state number defined earlier, and i represents the state's number in the path sequence. Nodes represent the specific location and state of the corresponding semantics or operations within the structure space.
[0113] If the path is a cycle or a symmetric structure, then the reverse path Nᵢ is formed, where Nᵢ is the bitwise inversion of Sᵢ, completing the closed loop. When the path includes start and end boundaries, add +0 and −0 to the first and last nodes respectively to form a complete path.
[0114] At least two paths can be grouped together using their index numbers (path_id) to enable multi-semantic parallel or multi-task parsing.
[0115] 3. Path Unit Generation Steps Based on the aforementioned encoding and mapping steps for Chinese characters, symbols, graphics, and data blocks, corresponding sparse activation units are generated respectively.
[0116] Arrange the corresponding sparse activation units sequentially according to the input order and add path numbers to form a logically ordered node chain.
[0117] For continuous, multi-layered, or intersecting paths, S-shaped and Z-shaped path extensions are used to maintain directional continuity.
[0118] The final output of each path node is a path unit, which is structured into three layers: state number, direction label, and semantic label. A path node is a basic constituent point in the path, encoded by 4HPU, and has an independent state number, used to carry single semantic or operational information. The difference between a path unit and a sparsely activated path unit is that a path unit is a combination of path nodes and their associated labels, representing semantic positional relationships; a sparsely activated path unit is the smallest operational unit in the path unit where only one state bit is activated, and it can directly participate in execution or mapping.
[0119] 4. Three-layer structure like Figure 3 As shown, the state numbering layer is used to record the position and state of path units in the 4HPU sparse activation coding. The coding format is four bits, where positive sparse activation coding uses a single "1" to indicate direction, and negative sparse activation coding is the bitwise inverse of positive sparse activation coding, such as 0001 being represented as 1110. The state numbering ensures the unique addressability and activation symmetry of path nodes.
[0120] Directional label layer: such as Figure 2 As shown, it is used to represent the spatial or logical direction of a path, including the eight directions of the graphic elements: →, ↗, ↑, ↖, ←, ↙, ↓, ↘, and can be extended to S and Z type multidimensional paths to describe the extensibility and hierarchical jump relationship of the path.
[0121] Semantic tag layer: Used to record semantic information corresponding to path nodes, such as: open, display, transmit, store, extracted from Chinese semantics, symbol meaning, or data block context. The semantic tag layer provides semantic guidance for path execution.
[0122] 5. Path Unit Conversion Rules and Steps A parsing module is set up at the computer system and artificial intelligence input end. The path units and path sequences obtained through encoding and mapping are input into the parsing module in sequence: Step 1: Parse (obtain) the state number layer to confirm the activation direction and node position. The specific process is as follows: Read the state number of each node in the path sequence in sequence, determine the activation direction based on the "1" bit in the 4HPU encoding, such as bottom left, bottom right, top left, top right, and at the same time determine the coordinate index of node i in the structure space based on the position of node number i in the sequence.
[0123] Step 2: Parse (obtain) the direction label layer and calculate the path flow and connection relationships. The specific process is as follows: Read the direction labels of each node, and determine the connection direction between adjacent nodes based on the direction identifiers, such as "→", "↑", "↓", and "←". Then, calculate the overall flow direction of the path according to the node number order to generate a directed connection matrix between nodes, which is used to confirm the subsequent path connectivity and order relationship.
[0124] Step 3: Parse (obtain) the semantic tag layer and map the semantic nodes to execution instructions or state descriptions. The specific process is as follows: The system reads the semantic tags of each node in sequence, and matches the preset instruction template according to the tag type, such as action, object, parameter, and modifier.
[0125] The action category labels are mapped to execution instructions, the object and parameter category labels are mapped to input parameters, and the modifier category labels are mapped to state descriptions, ultimately generating a set of interpreted instructions that can be executed in subsequent systems.
[0126] Ultimately, a direct translation relationship structure of "semantics, state, and instruction" is formed.
[0127] 6. Definition of semantic, state, and instruction translation relationships and inter-layer correspondence. "Semantic, state, and instruction direct translation relationship" means that the input natural language or structured data no longer undergoes traditional word segmentation and probability prediction, but is directly mapped into an executable instruction chain after path parsing. "Natural language or structured data" includes the four types of input defined in this invention: Chinese, symbols, graphics, and data blocks. Among them, Chinese belongs to natural language, symbols and graphics can be input according to structured descriptions, and data blocks are a type of structured data.
[0128] The steps for directly mapping a path to an executable instruction chain are as follows: First, the state number layer and direction label layer are parsed to obtain a directed structural path. Second, the semantic label layer is read, and semantic nodes are converted into instruction elements according to "action, object, parameter, and modifier". Third, the instruction elements are arranged into an instruction chain according to the path order, and parameters such as object, range, and threshold are bound. Fourth, a trace_id and checksum are added to the instruction chain to form an executable interpreted instruction chain for subsequent execution and backtracking verification. An instruction element is the smallest executable instruction unit generated from a single semantic node, containing the action type and its associated object or parameters, and is used to form the basic components of a complete instruction chain.
[0129] The semantic layer corresponds to the semantic tag layer, providing the source and type of semantic content; the state layer corresponds to the state number layer, defining the execution state under the semantic node; and the instruction layer corresponds to the direction tag layer and maps to the interface of the computer architecture and artificial intelligence operating system or embedded control system, determining the transmission of instruction flow in the structural path, the execution direction, and the invocation of the kernel execution module or peripheral driver module. These three layers form a closed-loop mapping, making the input semantics "i.e., path, state, and instruction," achieving a direct translation process from semantics to execution without the need for an intermediate interpretation layer.
[0130] To ensure the traceability of path units in execution and result feedback, this invention assigns a unique path number, trace_id, to each path sequence generated by 4HPU encoding.
[0131] Path numbers are generated randomly or sequentially during the path unit generation phase, using a 64-bit or 128-bit structure identifier format. Trace_id=(path_type+timestamp+sequence).
[0132] Here, `path_type` represents the path type (Chinese characters, symbols, graphics, or data blocks), `timestamp` represents the generation timestamp, and `sequence` is an auto-incrementing sequence number. Each `trace_id` corresponds one-to-one with the path sequence, structure state sequence, and execution result, and is used for execution logging, error backtracking, and result verification.
[0133] VI. Instruction Mapping 1. Generation of structural state sequences A structural state sequence is a structured data sequence formed by arranging at least two sparse activation units after 4HPU encoding and path parsing according to semantic logic or execution order. It is used to represent the dynamic evolution process of input information in a structural intelligent system.
[0134] The structural intelligence system is a non-tokenized information processing system built on 4HPU encoding. It has path parsing, semantic translation and instruction mapping functions to realize the direct mapping process from input semantics to execution.
[0135] Structured data sequences are ordered data sets organized according to semantics, state, and path order after being encoded by 4HPU, and are used for direct translation, path backtracking, and execution mapping.
[0136] Each structural intelligence system contains three core fields: state_id: records the activation state of a node in the path using a four-bit sparse activation code; dir_tag: identifies the directional attribute of the node, such as →, ↗, ↑, ↖, ←, ↙, ↓, ↘; and sem_tag: identifies the semantics or instructions represented by the node.
[0137] The specific steps to obtain the structured data sequence are: Seq=[(state_id1,dir_tag1,sem_tag1),(state_id2,dir_tag2,sem_tag2),...,(state_idn,dir_tagn,sem_tagn)].
[0138] Extract three pieces of information from the path units output from the path parsing stage: status number, direction label, and semantic label; sort them according to the temporal order or logical dependency order of the input semantics; perform S-shaped and Z-shaped path expansion on consecutive nodes to maintain directional continuity; Add +0 and -0 markers to the first and last nodes of the structured data sequence, respectively, to indicate the start and end states; Integrate all node information into a unified JSON key-value pair structured state sequence, for example: [{state_id:1, dir_tag:"→", sem_tag:"open"}, {state_id:2, dir_tag:"↑", sem_tag:"file"}]. The structured state sequence serves as the "structured path instruction stream," providing the basic input for subsequent literal translation mapping. The structured path instruction stream is a set of executable instructions generated according to the node order in the structured state sequence, used to drive the semantic-to-execution literal translation operation.
[0139] The structure state sequence (instruction set) is directly translated into Chinese semantic operating system call instructions, electro-optical display instructions, or communication and storage instructions based on 4HPU encoding.
[0140] Literal translation refers to the process of directly generating an executable instruction chain from three layers of information—state number, direction label, and semantic label—in a structural state sequence without going through traditional compilation, word segmentation, or probabilistic interpretation layers.
[0141] The translation process uses semantic tags as semantic sources, status numbers as execution locations, and direction tags as flow sequences to achieve a structural mapping from input semantics to execution, that is, "semantics is instructions, and path is execution".
[0142] Through direct translation, it is possible to directly execute operations corresponding to the input semantics in the operating system kernel, electro-optical display control module, and communication and storage channels in computer systems and artificial intelligence, forming a structural intelligent operation mode that integrates semantics, state, and instructions.
[0143] The steps to translate a structure state sequence into an operation call instruction are as follows: 1. Input reception: Receive the structural state sequence S = [S1, S2, …, S] output from the path resolution phase. n ]; 2. Semantic parsing: Read the semantic tag sem_tag of each node and identify its corresponding semantics, such as open, read, write, close; 3. Determine the execution module: Based on the group identifier of the state number state_id, determine the corresponding execution module in the operating system, such as: file system, process management, I / O management, or kernel interface; 4. Determine the execution order: Generate an execution chain based on the order information of the direction tag dir_tag to ensure that the instruction flow is consistent with the semantic logic; 5. Generate instruction format: Combine semantic, state, and direction information to generate a system-recognizable calling format: sys.exec[(sem_tag1→sem_tag2→…→sem_tagn)] For example: Input "open file", output sys.exec[open→read→close].
[0144] like Figure 4 As shown, the above steps enable the direct mapping of Chinese semantics to the executable call chain at the operating system layer under 4HPU encoding.
[0145] The steps to translate the structural state sequence into electro-optical display instructions are as follows: 1. Input reception: Receives the structure-encoded graphical path sequence or text semantic path sequence; 2. Semantic parsing: Read the semantic tag sem_tag to identify the display operation type, such as: displaying text, drawing graphics, refreshing the screen; 3. Determine the display channel: Match the display device type based on the state number state_id, such as LED array, LCD controller, e-ink screen; 4. Determine the drawing path: Generate drawing trajectories (points, lines, and surfaces) according to the direction labels dir_tag in sequence. If an S / Z type path exists, maintain directional continuity. 5. Generate instruction format: Integrate semantics, direction, and status into a display control instruction set: display.exec[(sem_tag1→sem_tag2→…→sem_tagn)] For example: Input "Display text: Spring Breeze", output display.exec[open→draw→refresh].
[0146] like Figure 5 As shown, the above steps realize the direct path mapping of semantic information to the display layer, and complete the electro-optical direct translation display of text and graphics.
[0147] The steps to translate the structural state sequence into communication and storage instructions are as follows: 1. Input reception: Receive encoded data block path sequences or communication semantic sequences; 2. Semantic parsing: Read the semantic tag sem_tag to identify the operation type, such as: send, receive, write, read, check; 3. Determine the target channel: Map the communication port, storage address, or data block number based on the state number (state_id); 4. Determine the transmission direction: Generate a data stream path based on the direction tag dir_tag: send, write, check, and maintain consistency in the flow direction; 5. Generate instruction format: Combine semantic, status, and direction information to generate communication or stored execution format: comm.exec[(sem_tag1→sem_tag2→…→sem_tagn)] For example: Input "Send data packet A and store the result", output comm.exec[send→write→verify].
[0148] like Figure 6 As shown, the above steps realize the direct mapping of the structural state sequence to the communication and storage channels, and support the full-path direct interpretation execution of the data transmission and verification process.
[0149] As can be seen from the above examples, the direct translation of the present invention refers to the process of directly converting the input Chinese characters, symbols, graphics or data block information into system executable instructions through 4HPU structure encoding and path parsing, without going through word segmentation, probability prediction or intermediate interpretation layers.
[0150] VII. Implementation and Feedback The result of the instruction mapping is output to the kernel interface of the computer system and artificial intelligence, the electro-optical display control module, or the communication and storage execution module, and the translation engine calls the corresponding module for execution.
[0151] After execution, a path number (trace_id) is generated and attached, which is used to identify the structure state sequence of this execution.
[0152] The trace_id is created during the path unit generation phase and remains unchanged throughout the translation and execution process.
[0153] The execution result, along with the trace_id, is returned to the upper-level control module or user interface to ensure the traceability and verifiability of the result.
[0154] If an error or interruption occurs during execution, the corresponding structural path and status node can be quickly located based on the trace_id, and the relevant sequence can be retrieved again for recalculation or error correction, thereby realizing a closed-loop feedback mechanism.
[0155] The present invention provides a direct translation method based on 4HPU structure encoding. A direct translation system is set up within the computer architecture and artificial intelligence framework. This system comprises an input interface module, a 4HPU encoding module, a path parsing module, an instruction mapping module, and an execution and feedback block. The backtracking mechanism includes three parts: trace_id path number recording, execution result comparison, and exception recovery. In the embodiments of the present invention, the direct translation system is implemented using Python combined with C++ extensions.
[0156] VIII. Backtracking and Security Verification 1. Backtracking Backtracking refers to the process of reversing and re-verifying the execution results of the structural state sequence based on the path number `trace_id` during the execution of translated instructions by the system. Reversing the path involves tracing back to the corresponding path node using `trace_id` to determine the location of execution anomalies or result deviations. Re-verification involves the system re-executing or verifying the calculation results at that node to ensure that the output is consistent with the expected semantics.
[0157] When the system's execution and backtracking module detects an abnormal state, execution interruption, or inconsistent output, it calls the path record corresponding to the `trace_id` and reloads the three layers of information—state number, direction label, and semantic label—from the structured state sequence, recalculating or correcting errors according to the original execution order. Abnormal states include data corruption, instruction loss, and broken path chains. Execution interruptions may be caused by device interruptions, communication timeouts, or kernel errors. Inconsistent output is triggered when backtracking is activated when at least two consecutive execution results do not match the expected output.
[0158] The backtracking process does not rely on external logs but leverages the reversible encoding characteristics of the 4HPU structure path itself to achieve reverse derivation between semantics, states, and instructions, ensuring the reproducibility and logical consistency of the execution process. The reversible encoding characteristic means that the 4HPU structure encoding has a one-to-one correspondence between forward and reverse states. That is, each forward sparse activation code such as 0001, 0010, 0100, 1000 has a unique corresponding inverse complement such as 1110, 1101, 1011, 0111. The system can use this mapping to recover the original path and execution sequence without relying on external logs.
[0159] 2. Security Verification Security verification refers to the verification of path integrity, data consistency, and execution permissions during the instruction execution and result output stages.
[0160] Each path is embedded with a verification field during generation. The verification field includes the path number trace_id, the state hash state_hash, and the check bit.
[0161] When outputting results, the system calculates the structure digest value `result_hash` of the execution result and compares it with the original `state_hash`. If they match, the execution is confirmed as safe and written to the acknowledgment log. A match means that the structure digest value `result_hash` calculated by the system is completely identical to the original digest value `state_hash` generated at the start of the task in terms of hash algorithm such as SHA-256. In other words, there is no difference in the numerical comparison between the two, which is considered to mean that the execution process has not been tampered with and the output result is correct.
[0162] If the verification fails, a backtracking process is immediately triggered to reload the original path for secondary verification or manual review. Verification failure occurs when the system-calculated execution result digest value result_hash is inconsistent with the original digest value state_hash, or when the result data is missing, corrupted, or out of bounds, making it impossible to verify its correctness through a one-time automatic verification. In such cases, verification is deemed a failure and a backtracking process is triggered.
[0163] The security verification mechanism ensures that each direct translation execution has a unique and verifiable fingerprint, making the system tamper-resistant and traceable.
[0164] 3. The synergistic relationship between backtracking and security verification Backtracking provides reverse recovery capabilities at the logical level, while security verification provides authenticity guarantees at the result level. The combination of the two constitutes the execution closed loop of this invention: trace_id location, verification detection, anomaly triggering, path replay, and result confirmation, realizing a verifiable, reproducible, and error-correctable intelligent security system.
[0165] The execution result includes a path ID (trace_id) for backtracking and security verification. The final result is a structured execution receipt set, which includes three parts: execution output, path summary, and security signature. 1. Execution output: Represents the actual result after the system completes. This can be a text result, such as "Operation completed" or "File saved"; a graphical display, such as "Interface updated" or "Image drawn"; a communication signal, such as "Data packet sent confirmation"; or storage feedback, such as "Write completed indicator".
[0166] 2. Path Digest: A summary of information generated from the structure state sequence, recording the path number, number of nodes, execution time and resource usage used during execution, for result auditing and performance tracking.
[0167] 3. Security Signature: A structured hash signature automatically generated by the system based on the backtracking and verification mechanism. The format is: sig=hash(trace_id+result_hash+timestamp), which is used to verify the uniqueness and integrity of the execution result.
[0168] The final output execution receipt dataset is stored with trace_id as the index and can be called, queried or verified in system logs, display modules or communication channels, realizing a structural intelligent closed loop of "identifiable input, verifiable execution and traceable results".
[0169] To facilitate understanding and implementation of the present invention, the symbols, letters, and foreign abbreviations appearing in the specification and accompanying drawings are uniformly defined below: 1. Path node Pᵢ: Represents the i-th path node unit in the structural state sequence, containing three pieces of information: state number (state_id), direction label (dir_tag), and semantic label (sem_tag). It is the smallest execution unit of path resolution.
[0170] 2. State ID: The node state value represented by 4-bit sparse activation coding, used to identify the activation position of the current node in the 4HPU structure coding.
[0171] 3. Direction tag dir_tag: Represents the direction attribute of a path node. The value range includes eight direction symbols: →, ↗, ↑, ↖, ←, ↙, ↓, and ↘, which are used to express the geometric or logical direction of the path.
[0172] 4. Semantic tag sem_tag: Represents the semantic function corresponding to the node, such as: open, display, transmit, store, used for semantic instruction generation.
[0173] 5. Path ID (trace_id): A unique ID generated for each path sequence, composed of path type (path_type), timestamp (timestamp), and auto-incrementing sequence number (sequence), used for execution result backtracking and security verification.
[0174] 6. System execution instruction format sys.exec[…]: This represents the sequence of operating system call instructions generated by the 4HPU structure translation system. The part in square brackets is the semantic execution chain, such as sys.exec[open→read→close].
[0175] 7. Display execution instruction format display.exec[…]: This represents the electro-optical display instruction sequence generated in the 4HPU structure direct translation system, used to control the display output process.
[0176] 8. Communication execution instruction format comm.exec[…]: Represents the sequence of communication and storage execution instructions generated in the 4HPU structure translation system, used for data transmission and result verification.
[0177] 9. Result Hash / State Hash: A digest value used for security verification, calculated using a hash function, to compare consistency before and after execution.
[0178] 10. Positive and Negative Zero (±0): In the 4HPU sparse activation system, these represent the boundary or static state of a path. Fully static +0 is encoded as 0000, and fully activated −0 is encoded as 1111, used for path start and end identification. The 4HPU sparse activation system is a four-bit binary encoding system, where each encoding unit has only one bit as "1" and the rest as "0", used to represent the unique activation state of a path node and its directional distribution in the structure space.
[0179] 11. 4HPU (Four Hierarchical Processing Unit): The four-bit sparse activation coding method adopted by the present invention is the core architecture for implementing structural path coding and non-token literal translation.
[0180] 12. Execution Receipt Set: A structured data set of execution result outputs, including three parts: execution output, path summary, and security signature, which is used for system confirmation and external verification.
[0181] Example 1, Chinese semantic literal translation operating system Using the system of the present invention, as Figure 3 shown, input "Open file test.txt". After being processed by the 4HPU structure encoder, the input semantics are decomposed into action units, object units, and parameter units, and category flags are attached to ensure uniqueness.
[0182] The parsing result is literally translated as the operating system call: fd = open("test.txt", O_RDONLY, NULL).
[0183] The execution result has a trace_id, which is convenient for backtracking and verification.
[0184] Example 2, electro-optical literal translation display Using the system of the present invention, as Figure 4 shown, input "Display text: Spring breeze". The input Chinese characters are mapped stroke by stroke to a path sequence: The stroke order of the character "春" is: [Horizontal (1 = 0001) → Vertical (2 = 0010) → Left-falling stroke (3 = 0100) → Right-falling stroke (4 = 1000) → Horizontal (-1 = 1110) → Vertical (-2 = 1101) → Horizontal (-4 = 0111)]; The stroke order of the character "风" is: [Left-falling stroke (3 = 0100) → Horizontal (1 = 0001) → Vertical fold (-3 = 1011) → Right-falling stroke (4 = 1000)]. Finally, it is parsed as the display driver instruction draw(stroke_i), and a trace_id is generated to support backtracking. The stroke order of the character "春" is: [Horizontal (1 = 0001) — Vertical (2 = 0010) — Left-falling stroke (3 = 0100) — Right-falling stroke (4 = 1000) — Horizontal (-1 = 1110) — Vertical (-2 = 1101) — Horizontal (-4 = 0111)]; The stroke order of the character "风" is: [Left-falling stroke (3 = 0100) — Horizontal (1 = 0001) — Vertical fold (-3 = 1011) — Right-falling stroke (4 = 1000)].
[0185] The final parsing result is the display driver instruction draw(stroke_i), which generates a trace_id to support backtracking.
[0186] Example 3, Communication and Storage Translation The system employing the present invention, such as Figure 5 As shown, the input is "transfer file data.txt". The action unit, object unit, and parameters (data.txt encoded character by character) are directly translated into a communication and storage instruction chain: (1) Communication: net.send("data.txt"); (2) Storage: storage.write("data.txt", block_X).
[0187] During execution, a trace_id is generated. If an error occurs, the process can be traced back and re-executed.
[0188] The present invention has the following advantages: 1. In steps four and five, the present invention directly converts Chinese semantics, symbols, graphic data or data blocks into sparse activation path units through 4HPU encoding, and realizes direct mapping from semantics to instructions in the path parsing step with a three-layer structure of state number layer, direction label layer and semantic label layer, fundamentally avoiding the dependence of traditional token segmentation and probability prediction.
[0189] The above steps use the structural path as the smallest computational and execution unit. Instead of using semantic slicing or probabilistic models for prediction, they directly generate structural state sequences through fixed sparse activation encoding. This achieves high adaptability and accurate direct translation for polysemous Chinese contexts and complex syntax. Therefore, the method of this invention can maintain semantic integrity, consistency, and traceability under non-tokenized conditions, significantly improving the execution efficiency and accuracy of Chinese semantic processing.
[0190] 2. In step six, this invention directly maps the 4HPU structure state sequence to instantly transform semantic paths into operating system, electro-optical display, and communication and storage instructions, skipping the word segmentation, rendering, and intermediate interpretation layers in traditional computing systems. Because sparse activation paths are used as the smallest execution unit, semantic parsing and instruction generation are completed within the same structural space, no longer relying on multi-layered language interpreters or explicit compilation processes, thus avoiding redundant computations in the semantic transfer and graphics rendering stages.
[0191] In typical application scenarios, such as command parsing in Chinese operating systems, dynamic interface display, and data communication, it can significantly reduce execution latency by about 30-60% and reduce computing power consumption by 40%, achieving lightweight direct translation execution.
[0192] 3. In step four, this invention employs sparse activation paths instead of traditional dense matrix calculations. The combined encoding of semantics, direction, and state is completed within the smallest activation unit of the 4HPU encoding, triggering local calculations only at necessary nodes, rather than performing item-by-item operations on the global matrix. This sparse activation method, during path parsing and instruction mapping, only requires access to and energy allocation to effective nodes, significantly reducing memory access and computational power consumption, reducing the overall computational complexity from O(n²) to approximately O(k), where k is the actual number of activated nodes. It is particularly suitable for deployment in embedded systems, edge computing units, and low-power terminal devices, enabling efficient Chinese semantic parsing and structured execution with limited resources, achieving low power consumption and high responsiveness.
[0193] 4. In step seven, this invention assigns a unique path number, `trace_id`, to each path during the execution of the structural state sequence. This number is used throughout the entire execution, verification, and output process. After the system completes instruction mapping and direct interpretation, the result data is bound to the `trace_id` and synchronously written to the receipt set, ensuring that each execution has a unique traceable identifier. When an abnormal state or inconsistent results are detected, the system can quickly locate the corresponding structural path based on the `trace_id`, call the original state number, direction label, and semantic label information to reload the execution flow, and achieve reverse backtracking and anomaly recovery. Simultaneously, the system calculates and compares `state_hash` and `result_hash` through a built-in security verification mechanism. If the verification passes, the result is confirmed to be complete; if it fails, backtracking and error correction are automatically triggered. This ensures the uniqueness, verifiability, and tamper resistance of the execution results, forming a closed-loop execution system of "locating `trace_id`, verification detection, path replay, and result confirmation."
[0194] 5. The unified architecture based on the 4HPU encoding system of this invention encodes and parses Chinese semantics, symbol sequences, graphic data, and data blocks using the same sparse activation rules and path mapping mechanism. This enables cross-domain reuse of the same structural encoding system across the operating system instruction layer, electro-optical display control layer, and communication and storage transmission layer. In the operating system domain, this encoding system is used for direct translation of semantics into system calls; in the display domain, it is used for fast mapping of graphic paths into electro-optical drawing instructions; and in the communication and storage domain, it is used for path-based transmission and backtracking verification of data blocks into transmission instructions. Therefore, this invention maintains unified path definitions, status numbers, and semantic mapping logic across different hardware and software modules, forming a structural intelligent foundation layer that is universal across various terminals, operating systems, and network environments. It possesses consistency and scalability across modules, scenarios, and domains, thereby overcoming the problems of severe token dependence, low execution efficiency, and lack of backtracking mechanisms in existing technologies.< / type>
Claims
1. A direct translation method based on 4HPU structure encoding, comprising the following steps: I. Setting up the 4HPU structure code A set of basic coding units is constructed using four digits. Each digit in the four digits is either 0 or 1. After the four digits are combined, only one digit is 1 and the rest are 0. II. Setting Chinese characters, symbols, graphics, and data block encoding The Chinese encoding is divided into three levels according to the semantics of Chinese characters: characters, words, and sentences. Specifically, characters are mapped using basic encoding units based on strokes or components; words are combined according to the basic encoding units of the individual characters that make up the word; and sentences connect the basic encoding units of words according to grammatical relationships. The symbol encoding is based on the symbol's punctuation class, mathematical class, logical class, path class, and control class. After generating the corresponding basic encoding units and forming path units, a category label and a symbol identification code are assigned to the end node position of the path unit. The graphic encoding decomposes the elements constituting the graphic into a three-level structure encoding in spatial order: points, lines, and surfaces; or strokes, components, and the whole. The data block encoding is as follows: the data block is divided into independent blocks according to the logical boundaries or logical length of the computer system and artificial intelligence field. Before encoding each independent block, a category label, data block number and integrity verification information are added to the beginning of the block. The data content of the independent block is mapped byte by byte into multiple sets of basic encoding units. III. Receiving Input Information Receive Chinese semantics, symbol sequences, graphic data, or data blocks represented in text encoding, symbol sequences, vector paths, or binary block formats from the human-computer interaction interface, application system interface, or external device sensing end of the computer system and artificial intelligence; IV. Encoding Mapping Transform Chinese semantics, symbol sequences, graphical data, or data blocks into basic coding units. The path unit obtained after the Chinese semantic code is composed of a four-digit combination of status number, direction label and semantic label in sequence, where only one digit is 1 and the rest are 0; The boundary ±0 identifier of the symbol is obtained, which can be combined with the Chinese semantics and data block to form a path unit in which only one bit is 1 and the rest are 0; The obtained graphical data consists of a four-digit combination of status number, direction label and semantic label, in which only one digit is 1 and the rest are 0; The path sequence obtained is a four-digit combination of a data block containing a path number, a category identifier, and a check digit, where only one digit is 1 and the rest are 0. V. Path Analysis Obtain the state number layer from the path unit and path sequence to confirm the activation direction and node position; Obtain the directional label layer from the path unit and path sequence, and calculate the path flow direction and connection relationship; Semantic tag layers are obtained from path units and path sequences, and semantic nodes are mapped to execution instructions or state descriptions. The action category labels are mapped to execution instructions, the object and parameter category labels are mapped to input parameters, and the modifier category labels are mapped to state descriptions, ultimately generating an executable instruction set. VI. Instruction Mapping The instruction set is directly generated into an executable instruction chain, which is a Chinese semantic operating system call instruction, electro-optical display instruction, or communication and storage instruction based on the basic coding unit; VII. Implementation and Feedback The results of instruction mapping are output to the kernel interface of computer system and artificial intelligence, electro-optical display control module or communication and storage execution module.
2. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: When all four bits of the code are 0, it indicates a completely static, inactive state, marked as +0. When all four bits of the code are 1, it indicates a fully active, saturated state, marked as −0.
3. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: The basic coding unit has four states: State 1, coded as 0001 State 2, coded as 0010 State 3, coded as 0100 State 4, coded as 1000.
4. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: The category label identifies the semantic type of the path unit, and the symbol identifier is a unique number assigned to each symbol.
5. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: The graphic elements are encoded according to their geometric orientation sequence: Right → Code is 0001 The top right corner ↗ is coded as 0010. The code for "Up" (↑) is 0100. The top left corner is coded as 1000. The left arrow (←) is encoded as 1110. The bottom left corner (↙) is coded as 1101. The down arrow (↓) is encoded as 1011. The bottom right corner is coded as 0111.
6. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: For closed regions or graphic boundaries, add ±0 markers at the end of the code.
7. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: The data block sequence number is a number assigned sequentially when the data block is divided into blocks, and the integrity verification information is a file class, network class, and / or cache class identifier.
8. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: The data block is an information unit existing in the form of binary data or structured data, carrying file content, network transmission packets, sensor data, database records, or cached fragments.
9. The direct translation method based on 4HPU structure encoding according to claim 1, characterized in that: The basic coding units form the basis of the 4HPU structure coding.
10. A system for implementing the method of claim 1 based on 4HPU structure encoding, characterized in that: It consists of an input interface module, a 4HPU encoding module, a path parsing module, an instruction mapping module, and an execution and backtracking module.
Citation Information
Patent Citations
Natural language character and numerical value general solution system and method based on 4HPU and application of natural language character and numerical value general solution system and method
CN120386847A