Data transmission method, apparatus, device, medium, and product

CN122845690APending Publication Date: 2026-09-29BEIJING YIHUI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610915009.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而此类协议的传输报文中包含大量字段名、分隔符、标签等冗余文本信息,不仅大幅占用传输带宽,造成带宽资源浪费,同时冗余信息会增加接收方的解析运算量,导致数据解析延迟高、吞吐性能差,无法满足高并发、低延时的传输需求

Benefits of technology

[0013]本发明实施例的技术方案,通过获取待传输的业务数据,业务数据包含业务数据标识、数值信息及语义描述信息;确定本地缓存中是否存在与业务数据标识对应的结构描述符;若否,则将语义描述信息编码为结构描述符,对数值信息进行编码生成数值字段,构建包含头部字段、结构描述符及数值字段的第一数据包,将头部字段中结构描述符对应的指示标志位置为有效;结构描述符用于指示数值字段的数据解析规则;若是,则构建包含头部字段及数值字段的第二数据包,将头部字段中结构描述符对应的指示标志位置为无效;将第一数据包或第二数据包发送至接收方,解决了现有技术中通用自解析协议数据传输冗余高、解析速率低,以及二进制协议依赖外部配置、可维护性差的问题,实现了在数据包中携带数据解析规则、无需外部信息即可独立解码的通用自解析能力,同时通过结构描述符的缓存复用避免重复传输,降低冗余数据与带宽占用,通过紧凑二进制编码与顺序化解析提升编解码效率,实现了数据传输的通用性、自解析能力与高效性,提升了数据交互的整体性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845690A_ABST
    Figure CN122845690A_ABST
Patent Text Reader

Abstract

A data transmission method, device, equipment, medium and product are disclosed. The method comprises: obtaining service data to be transmitted, the service data comprising a service data identifier, numerical information and semantic description information; determining whether a structure descriptor corresponding to the service data identifier exists in a local cache; if not, encoding the semantic description information into the structure descriptor, encoding the numerical information to generate a numerical field, constructing a first data packet comprising a header field, the structure descriptor and the numerical field, and setting an indication flag corresponding to the structure descriptor in the header field to valid; if yes, constructing a second data packet comprising a header field and a numerical field, and setting an indication flag corresponding to the structure descriptor in the header field to invalid; and sending the first data packet or the second data packet to a receiver, so as to enable the data packet to have a general self-analysis capability of being independently decoded without external information, while reducing transmission redundancy and improving data transmission and analysis efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data transmission method, apparatus, device, medium, and product. Background Technology

[0002] In a distributed system architecture, service nodes need to interact with each other frequently and in real time. The self-parsing capability and transmission parsing efficiency of the data transmission protocol directly affect the overall performance of the system.

[0003] Current mainstream data transmission protocols are mainly divided into two categories: general text protocols and traditional binary encoding protocols. General text protocols such as JSON and XML transmit data containing complete field names, structural semantics, and other information, allowing the receiver to decode and parse the data without relying on additional configurations, template files, or other third-party information outside the protocol. However, the transmission messages of these protocols contain a large amount of redundant text information such as field names, delimiters, and tags. This not only significantly consumes transmission bandwidth, wasting bandwidth resources, but also increases the parsing computation load on the receiver, resulting in high data parsing latency and poor throughput, failing to meet the requirements of high concurrency and low latency transmission.

[0004] Traditional binary encoding protocols use a compact binary data format, requiring the receiver to rely entirely on pre-agreed external templates or protocol rules to decode the data, lacking self-parsing capabilities. When the business data structure changes, the configurations at both ends need to be updated synchronously, resulting in high maintenance costs and poor scalability. Summary of the Invention

[0005] This invention provides a data transmission method, apparatus, device, medium, and product to enable data packets to have a universal self-parsing capability that can be independently decoded without external information, while reducing transmission redundancy and improving data transmission and parsing efficiency.

[0006] According to one aspect of the present invention, a data transmission method is provided, applied to a sender, the method comprising: Acquire the service data to be transmitted, the service data including service data identifier, numerical information and semantic description information; Determine whether a structure descriptor corresponding to the business data identifier exists in the local cache; If not, the semantic description information is encoded into a structure descriptor, the numerical information is encoded to generate a numerical field, a first data packet containing a header field, the structure descriptor, and the numerical field is constructed, and the position of the indicator flag corresponding to the structure descriptor in the header field is set to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field; If so, then construct a second data packet containing the header field and the numeric field, and set the indicator flag corresponding to the structure descriptor in the header field to invalid; Send the first data packet or the second data packet to the recipient.

[0007] According to another aspect of the present invention, a data transmission apparatus is provided, configured on a sender, the apparatus comprising: The acquisition module is used to acquire the business data to be transmitted, the business data including business data identifier, numerical information and semantic description information; The detection module is used to determine whether a structure descriptor corresponding to the business data identifier exists in the local cache; The first data packet construction module is used to encode the semantic description information into a structure descriptor if no, encode the numerical information to generate a numerical field, construct a first data packet containing a header field, the structure descriptor, and the numerical field, and set the position of the indicator flag corresponding to the structure descriptor in the header field to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field. The second data packet construction module is configured to construct a second data packet containing a header field and the numerical field if the condition is met, and to set the indicator flag corresponding to the structure descriptor in the header field to invalid. The sending module is used to send the first data packet or the second data packet to the receiver.

[0008] According to one aspect of the present invention, a data transmission method is provided, applied to a receiver, the method comprising: Receive data packets sent by the sender, parse the header fields of the data packets, and obtain the descriptor indicator flags; If the descriptor indicator flag is valid, then the structure descriptor is read from the data packet, and the structure descriptor is associated with and stored with the service data identifier of the data packet; If the descriptor indicator flag is invalid, then the stored structure descriptor is located based on the business data identifier; The data packets are decoded based on the structure descriptor to restore the business data.

[0009] According to another aspect of the present invention, a data transmission apparatus is provided, configured at a receiver, the apparatus comprising: The parsing module is used to receive data packets sent by the sender, parse the header fields of the data packets, and obtain the descriptor indicator flags; The storage module is used to read the structure descriptor from the data packet if the descriptor indicator flag is valid, and to associate and store the structure descriptor with the service data identifier of the data packet. The lookup module is used to look up the stored structure descriptor based on the business data identifier if the descriptor indicator flag is invalid. The decoding module is used to decode the numerical fields in the data packet based on the structure descriptor and restore the business data.

[0010] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to said at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data transmission method described in any embodiment of the present invention.

[0011] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data transmission method described in any embodiment of the present invention.

[0012] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the data transmission method as described in any embodiment of the present invention.

[0013] The technical solution of this invention, through acquiring the business data to be transmitted, which includes a business data identifier, numerical information, and semantic description information; determining whether a structure descriptor corresponding to the business data identifier exists in the local cache; if not, encoding the semantic description information into a structure descriptor, encoding the numerical information to generate a numerical field, constructing a first data packet containing a header field, a structure descriptor, and a numerical field, and setting the position of the indicator flag corresponding to the structure descriptor in the header field to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field; if yes, constructing a second data packet containing a header field and a numerical field, and setting the position of the indicator flag corresponding to the structure descriptor in the header field to invalid; and sending the first or second data packet to the receiver, solves the problems of high data transmission redundancy, low parsing rate, and poor maintainability of binary protocols in the prior art, which rely on external configuration. It realizes the general self-parsing capability of carrying data parsing rules in the data packet and being able to decode independently without external information. At the same time, it avoids repeated transmission by caching and reusing the structure descriptor, reducing redundant data and bandwidth occupation, and improves encoding and decoding efficiency through compact binary encoding and sequential parsing. It achieves the universality, self-parsing capability, and high efficiency of data transmission, and improves the overall performance of data interaction.

[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart of a data transmission method provided according to an embodiment of the present invention; Figure 2 This is a flowchart of a data transmission method provided according to an embodiment of the present invention; Figure 3 This is a schematic diagram for characterizing header fields according to an embodiment of the present invention; Figure 4 This is a schematic diagram for representing a numerical field according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the encoding format for representing semantic option fields provided in an embodiment of the present invention; Figure 6This is a schematic diagram of a semantic option field provided according to an embodiment of the present invention; Figure 7 This is a schematic diagram illustrating a sender-based data transmission method according to an embodiment of the present invention; Figure 8 This is a schematic diagram illustrating a receiver-based data transmission method according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of a data transmission device provided according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a data transmission device provided according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of an electronic device that implements the data transmission method of this invention. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0019] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to maintain user personal information security and network security. It should also be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein are all conducted with the user's knowledge and consent, and comply with relevant privacy protection regulations.

[0020] Before introducing this technical solution, its application scenarios can be described first. The technical solution provided by this invention can be applied to any scenario that requires efficient transmission of structured data and involves repetitive transmission needs.

[0021] For example, when a user logs into the system, the client reports user information to the server. Data transmission between the sender and receiver can be achieved based on the technical solution provided in this embodiment. When the same user logs in multiple times or different users log in, the data structure is the same but the specific values ​​are different, and some fields (such as device fingerprint and geographical location) are not required every time. This solution can avoid the repeated transmission of all fields.

[0022] When an e-commerce order is submitted, the client sends order data to the server. Data transmission between the sender and receiver can be achieved using the technical solution provided in this embodiment. Different orders have the same data structure, but fields such as product information and shipping address are filled in as needed. Null or default values ​​do not need to be repeatedly encoded and transmitted, thereby reducing request load.

[0023] When mobile applications report event data, the client sends user behavior data to the server. Data transmission between the sender and receiver can be achieved using the technical solution provided in this embodiment. The data structure for similar events is fixed, but the extended attribute fields change dynamically. Reusing structure descriptors can reduce redundant data transmission and save mobile data consumption.

[0024] In financial transaction systems, transaction terminals send message data to processing terminals. Data transmission between the sender and receiver can be achieved based on the technical solution provided in this embodiment. While similar transactions share the same data structure, some fields (such as additional information and remarks) only appear in specific scenarios. By adaptively carrying structure descriptors, simplified transmission is achieved, improving transmission efficiency.

[0025] Figure 1This is a flowchart of a data transmission method according to an embodiment of the present invention. This embodiment is applicable to any situation requiring efficient transmission of structured data. The method can be applied to the sender and can be executed by a data transmission device, which can be implemented in hardware and / or software and can be configured in a computing device. Figure 1 As shown, the method includes: S110. Obtain the service data to be transmitted. The service data includes service data identifier, numerical information and semantic description information.

[0026] Among them, business data refers to the original data content to be transmitted, which includes specific business information.

[0027] For example, when performing a user login authentication task, the business data can be login request data composed of fields such as user account, password hash, device identifier, geographical location, and login timestamp; when performing an e-commerce order creation task, the business data can be order submission data composed of fields such as buyer identifier, product list, shipping address, payment method, and discount information; when performing an IoT sensor reporting task, the business data can be environmental monitoring data composed of fields such as device number, temperature reading, humidity reading, battery level, and collection timestamp.

[0028] A business data identifier can be an identifier used to uniquely identify a type of business data. Data with the same data structure shares the same business data identifier, facilitating the association and reuse of data structures by the sender and receiver. Numerical information refers to the specific values ​​of each field in the business data, that is, the data part that carries the actual business semantic content. Semantic description information refers to the information used to describe the definition of the business data structure, including but not limited to: field names, data types, hierarchical nesting relationships, and container structures, which are used to explain the organization and parsing rules of numerical information.

[0029] In one implementation, the sender can obtain the business data to be transmitted from the application layer's business interface call. This business data is dynamically generated by the application runtime and includes three parts: business data identifier, numerical information, and semantic description information. The business data identifier is pre-assigned by the application according to the business type and embedded in the business data. The numerical information is the specific field value input by the user or generated by the system. The semantic description information is automatically derived from the application's data model definition and includes the name, type, and hierarchical relationship of each field.

[0030] For example, business data may include login request data composed of fields such as user account, password hash, device identifier, geographical location and login timestamp, or order submission data composed of fields such as buyer identifier, product list, shipping address, payment method and coupon information.

[0031] In another implementation, the sender can act as a forwarding node, receiving business data to be transmitted from an external data source or message queue. This business data is pre-stored in a structured format in a database or message middleware, and the sender reads the business data through a data access interface. The business data identifier is generated by mapping the table name or message topic of the data source, the numerical information is the field values ​​in the database record or message body, and the semantic description information is derived from the table structure definition of the data source or the message schema convention.

[0032] For example, in scenarios such as IoT gateways aggregating sensor data and message buses forwarding business events, the business data may include environmental monitoring data composed of fields such as device number, temperature reading, humidity reading, battery level and collection timestamp, or financial message data composed of fields such as transaction account, transaction amount, transaction type, additional information and remarks.

[0033] In another implementation, the sender can statically define the business data structure to be transmitted through a configuration file or protocol agreement. The business data identifier is specified by the identifier of the configuration item, the numerical information is filled in at runtime according to the actual business scenario, and the semantic description information is provided by the predefined structure template in the configuration file.

[0034] For example, in scenarios where the business data structure is fixed and frequently repetitive, such as periodic communication tasks like heartbeat packet reporting and status synchronization, the business data may include heartbeat packet data composed of fields such as node identifier, running status, load index, heartbeat sequence number and timestamp, or status synchronization data composed of fields such as session identifier, synchronization version, incremental data and verification information.

[0035] S120. Determine whether a structure descriptor corresponding to the business data identifier exists in the local cache.

[0036] The structure descriptor refers to compact binary data encoded based on semantic description information, containing template fields and semantic option fields, used to describe the mapping relationship between the type structure of numerical information and field names. The local cache refers to a storage area in the sender's storage medium used to store the structure descriptor; it can be located in memory or a high-speed storage device, and is used to avoid repeatedly generating existing structure descriptors.

[0037] In one implementation, the sender can maintain a local cache based on a hash table, using the business data identifier as the hash key and the storage address or content digest of the structure descriptor as the hash value. When business data is received, the sender can calculate the hash value of the business data identifier and perform a lookup operation in the hash table. If an entry corresponding to the hash value exists in the hash table and the corresponding structure descriptor has not exceeded its preset validity period, it is determined that the cache has been hit, i.e., the structure descriptor exists. If the entry does not exist in the hash table or the corresponding entry has expired, it is determined that the cache has not been hit, i.e., the structure descriptor does not exist. This triggers the encoding process of semantic description information to generate a new structure descriptor, and the mapping relationship between the structure descriptor and the business data identifier is written into the hash table.

[0038] In another implementation, a fixed-capacity cache space can be pre-defined to store the mapping relationship between business data identifiers and structure descriptors, with all mapping relationships forming an entry list. Each time business data is retrieved, the entry list in the cache is traversed, comparing the business data identifier with the identifiers of each entry in the entry list. If a matching entry exists, it is determined that a structure descriptor exists; if no matching entry is found after traversal, it is determined that a structure descriptor does not exist, a new structure descriptor is generated and inserted at the head of the entry list, and simultaneously, it is checked whether the cache capacity has exceeded the limit. If it has, the least used entry at the tail of the list is removed.

[0039] In another implementation, a hierarchical storage architecture can be used in the local cache. The structure descriptors corresponding to frequently used business data identifiers are placed in the memory cache layer, while less frequently used or older structure descriptors are persisted to the disk storage layer. After retrieving business data, the business data identifier is first retrieved from the memory cache layer. If the retrieval is successful, it is considered a cache hit. If the memory cache layer does not find the identifier, the disk storage layer is further queried. If the disk storage layer finds the identifier successfully, it is considered that a structure descriptor exists. If the disk storage layer still does not find the identifier, it is considered that a structure descriptor does not exist, triggering the semantic description information encoding process to generate a new structure descriptor, which is then written to both the memory cache layer and the disk storage layer.

[0040] S130. If not, the semantic description information is encoded into a structure descriptor, the numerical information is encoded to generate a numerical field, a first data packet containing a header field, a structure descriptor, and a numerical field is constructed, and the position of the indicator flag corresponding to the structure descriptor in the header field is set to valid.

[0041] The structure descriptor indicates the data parsing rules for numeric fields, guiding their decoding and reconstruction. These parsing rules, carried by the structure descriptor, define the type, hierarchy, field order, and name mapping information, serving as the basis for the receiver to reconstruct structured business data from numeric fields. Template fields are sequences of symbols composed of type descriptors and structure descriptors in a hierarchical nested order, defining the type structure and hierarchy of numeric information. Semantic option fields are serialized data containing key-value pairs of symbol numbers and field names, establishing a mapping between field names and symbol positions in the template field. Numeric fields are binary data encoded from numeric information according to the symbol sequence defined in the template field; they do not contain any self-descriptive information such as field names or type identifiers, and their parsing relies entirely on the structure descriptor.

[0042] The header field refers to the control information located at the beginning of the data packet. It contains flags indicating the presence status of components such as structure descriptors and numeric fields, as well as basic metadata such as service data identifiers. The indicator flags are specific segments in the header field used to explicitly declare the presence status of the structure descriptor in the current data packet. The receiver uses these flags to determine whether it needs to read the structure descriptor. The indicator flags have two values: valid and invalid. Valid: Indicates that the current data packet carries a complete structure descriptor; the receiver needs to read and parse the structure descriptor to obtain the data parsing rules. Invalid: Indicates that the current data packet does not carry a structure descriptor; the receiver should search for a stored structure descriptor in its local cache based on the service data identifier.

[0043] The first data packet refers to the complete data packet constructed when the structure descriptor does not exist in the local cache. It consists of three parts: header fields, structure descriptor, and numeric fields, and is used for data synchronization during the first transmission or after the cache expires.

[0044] In this embodiment, semantic description information can be split into multiple substructures according to a hierarchical structure. Each substructure is independently encoded as a substructure descriptor fragment, and each fragment contains independent template subfields and semantic sub-option fields. The substructure descriptor fragments are assembled into a complete structure descriptor according to the hierarchical nesting relationship. During the assembly process, a fragment reference index table is established to record the position and hierarchical relationship of each fragment in the complete structure descriptor. Numerical fields are generated by encoding the assembled complete template field symbol sequence. When constructing the header fields, the position of the indicator flag corresponding to the structure descriptor is set to valid, and the fragment count and reference index table are appended. During parsing, the receiver reads the complete structure descriptor based on the valid flag, identifies the fragment reference index table, loads the substructure descriptor fragments as needed, assembles and restores them to the complete structure descriptor, and then parses the numerical fields based on the restored structure descriptor.

[0045] It can also maintain a template field fragment library for commonly used substructures, pre-encoding frequently occurring hierarchical substructures as template field fragments. When parsing the hierarchical structure of semantic description information, it identifies hierarchical substructures that match the fragment library, references the corresponding template field fragments, and performs encoding on unmatched independent substructures to generate custom template field fragments. Following the hierarchical nesting order, it assembles the referenced template field fragments and custom template field fragments into a complete template field, and generates corresponding semantic option fields, together forming the structure descriptor. When handling header fields, it sets the indicator flags corresponding to the structure descriptor to valid and appends a fragment reference identifier list. During receiver parsing, it reads the structure descriptor based on the valid flags, identifies the fragment reference identifier list, loads the corresponding template field fragment from the local fragment library, and concatenates it with the received custom template field fragments to reconstruct the complete template field.

[0046] It can also parse the hierarchical structure of semantic description information, identify the repetitive arrangement patterns of type descriptors and structure delimiters in template fields, construct a compressed dictionary of symbol sequences, replace frequently occurring symbol subsequences with dictionary indices, and generate compressed template fields. Simultaneously, it extracts all field names to construct a name dictionary, and replaces field names in semantic option fields with dictionary indices. The compressed dictionary and name dictionary are appended to the structure descriptor as dictionary information. Numerical information is encoded according to the order of the compressed template field symbol sequences to generate numerical fields. When constructing header fields, the indicator flags corresponding to the structure descriptor are set to valid, and a compression flag is appended. During parsing, the receiver reads the structure descriptor based on the valid flags, identifies the compression flag, loads the appended dictionary information, restores the dictionary indices to the original symbol sequences and field names, and then parses the numerical fields based on the restored structure descriptor.

[0047] A version identifier can be attached to the newly generated structure descriptor, which includes a generation timestamp and a structure hash value. The structure descriptor and version identifier are written together to the local cache, and a version history of this business data identifier is established. When receiving structure descriptors with the same business data identifier but different versions, the receiver determines the version difference based on the version identifier, performs incremental updates on the differing parts, and maintains the existing cache for the same parts, thus realizing dynamic updates of the structure descriptor.

[0048] S140. If so, construct a second data packet containing header fields and numeric fields, and set the indicator flag corresponding to the structure descriptor in the header field to invalid.

[0049] The second data packet refers to a simplified data packet constructed when the structure descriptor already exists in the local cache, which includes header fields and numeric fields.

[0050] In this embodiment, if it is determined that a structure descriptor corresponding to the business data identifier exists in the local cache, the numerical information can be encoded to generate a numerical field, and the indicator flag corresponding to the structure descriptor in the header field can be set to invalid. Then, a second data packet containing the header field and the numerical field is constructed.

[0051] S150, Send the first data packet or the second data packet to the receiver.

[0052] It should be noted that the sender can be used to represent the network communication entity that constructs and transmits data packets, and is responsible for data encoding, data packet construction, and network transmission. The receiver can be used to represent the network communication entity that receives and parses data packets, and is responsible for data reception, structure descriptor caching, and decoding and restoring numerical fields.

[0053] In this embodiment, after the sender completes the construction of the first data packet or the second data packet, it can send the data packet to the receiver through the established transport layer connection.

[0054] When the first data packet is sent, the indicator flag in the header field is in a valid state. After the receiver parses the header field, it can read the structure descriptor from it, associate the structure descriptor with the business data identifier and store it in the local cache, and then parse the numerical field according to the structure descriptor to restore the original business data.

[0055] When sending the second data packet, the indicator flag in the header field is invalid. After parsing the header field, the receiver looks up the local cache based on the service data identifier, retrieves the stored structure descriptor, and then parses the numerical fields based on the structure descriptor to reconstruct the service data. This method uses the indicator flag for explicit differentiation, enabling the receiver to accurately identify the data packet type and adopt the corresponding parsing strategy, thus ensuring the reliability of data transmission and the correctness of parsing.

[0056] Furthermore, before sending the first or second data packet, the sender can dynamically select the transmission protocol based on network environment characteristics. For example, in a low-latency network environment, a datagram-oriented transmission protocol can be used to reduce connection establishment overhead; in a network environment with high reliability requirements, a connection-oriented transmission protocol can be used, with the addition of sequence numbers and acknowledgment mechanisms to ensure that data packets arrive in an orderly manner and are transmitted reliably. The receiver adopts an appropriate receiving strategy according to the transmission protocol type. After parsing the header fields, it performs structure descriptor reading or cache lookup according to the status of the indicator flags, and finally completes the decoding of the numerical fields and the restoration of the service data. This method, through dynamic adaptation of the transmission protocol, meets the requirements for transmission efficiency and reliability in different network environments.

[0057] The technical solution of this invention, through acquiring the business data to be transmitted, which includes a business data identifier, numerical information, and semantic description information; determining whether a structure descriptor corresponding to the business data identifier exists in the local cache; if not, encoding the semantic description information into a structure descriptor, encoding the numerical information to generate a numerical field, constructing a first data packet containing a header field, a structure descriptor, and a numerical field, and setting the position of the indicator flag corresponding to the structure descriptor in the header field to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field; if yes, constructing a second data packet containing a header field and a numerical field, and setting the position of the indicator flag corresponding to the structure descriptor in the header field to invalid; and sending the first or second data packet to the receiver, solves the problems of high data transmission redundancy, low parsing rate, and poor maintainability of binary protocols in the prior art, which rely on external configuration. It realizes the general self-parsing capability of carrying data parsing rules in the data packet and being able to decode independently without external information. At the same time, it avoids repeated transmission by caching and reusing the structure descriptor, reducing redundant data and bandwidth occupation, and improves encoding and decoding efficiency through compact binary encoding and sequential parsing. It achieves the universality, self-parsing capability, and high efficiency of data transmission, and improves the overall performance of data interaction.

[0058] Based on the above technical solution, optionally, the structure descriptor includes a template field and a semantic option field; the template field is a sequence of symbols composed of type descriptors and structure delimiters, and the semantic option field contains key-value pairs of symbol ordinal numbers and field names, with the symbol ordinal number used to indicate the reference index of the field name in the template field.

[0059] In this embodiment, a template field refers to a sequence of symbols composed of type descriptors and structure delimiters arranged in a hierarchical nesting order, used to define the type structure and container hierarchy of numerical information. A type descriptor is a symbol used to identify the data attribute type, corresponding to the basic data type in the semantic description information, such as integer, floating-point, boolean, and string types, and is used to indicate the decoding method of the data at the corresponding position in the numerical field. A structure delimiter is a symbol used to identify the boundary of the container structure, corresponding to container types such as object start, object end, array start, and array end in the semantic description information, and is used to indicate the hierarchical nesting relationship of the data at the corresponding position in the numerical field. A symbol sequence is an ordered string of symbols formed by arranging type descriptors and structure delimiters in a specific traversal order. The template field uses this sequence to fully express the type definition and hierarchical layout of the data structure.

[0060] Semantic option fields are serialized data containing key-value pairs of symbol indices and field names. They establish a mapping between field names and the positions of symbols in the template field, and are the part of the structure descriptor that carries the semantic meaning of the names. Symbol indices are incremental numbers assigned to each field name, corresponding to the field's position in the template field's symbol sequence. They serve as reference indexes for field names, decoupling names from structures. Key-value pairs are data units composed of symbol indices and field names, with the symbol indices as keys and the field names as values. Multiple key-value pairs arranged in an ordered manner constitute the main content of the semantic option field. The reference index refers to the positioning function of the symbol indices during data parsing. The receiver uses the symbol indices to determine the type and hierarchy information of the corresponding field from the template field's symbol sequence, achieving an index lookup from name to parsing rules.

[0061] To reduce data transmission overhead and improve data transmission efficiency, semantic description information is encoded into template fields, including: parsing the hierarchical structure of semantic description information and identifying the data type of each data node in the hierarchical structure; data types include data attribute types and container structure types; based on a preset type mapping relationship, data attribute types are converted into corresponding type descriptors, and container structure types are converted into corresponding structure delimiters; according to the nesting order of the hierarchical structure, the type descriptors and structure delimiters are serialized and combined to generate template fields composed of symbol sequences that indicate data parsing rules.

[0062] In this context, the hierarchical structure refers to the parent-child nesting relationship between data nodes in the semantic description information. It can be presented as a tree topology, with the root node corresponding to the overall data structure and child nodes corresponding to nested fields or container elements. Data nodes are the basic units in the hierarchical structure, corresponding to a single field or container in the business data. The data characteristics they carry can be categorized into data attribute types or container structure types. The type mapping relationship is a predefined correspondence rule between data attribute types and type descriptors, used to ensure that the sender and receiver use consistent symbolic representations for the same data type.

[0063] Data attribute types are the basic data types carried by data nodes, including integer, floating-point, boolean, and string types, and are used to define the encoding method of field values. Type descriptors are encoding symbols that identify the data attribute type; each basic data type corresponds to a unique type descriptor, and the receiver can determine the decoding method of the data at the corresponding position in the numeric field based on the type descriptor.

[0064] Container structure types are composite structure types that organize multiple data nodes, including object types and array types: object types represent collections of key-value pairs, and array types represent ordered sequences of elements. Structure delimiters are coded symbols that identify the boundaries of the container structure, including object start delimiters, object end delimiters, array start delimiters, and array end delimiters, which define the scope of the container by appearing in pairs.

[0065] Nesting order refers to the traversal order of data nodes in a hierarchical structure, reflecting the arrangement of type descriptors and structure delimiters in a symbol sequence. Serialization composition is the process of arranging and concatenating type descriptors and structure delimiters into a continuous string of symbols according to a specified traversal order. Data parsing rules are information such as type definitions, hierarchical relationships, and field order carried by template fields through symbol sequences, used by the receiver to restore numeric fields into structured business data.

[0066] In one implementation, after obtaining the semantic description information, the sender can identify the root node of the hierarchical structure and determine whether it is a data attribute type or a container structure type. If the root node is a basic data node, its data attribute type is converted into the corresponding type descriptor according to the type mapping relationship, serving as the unique element of the symbol sequence. If the root node is a container structure node, its container structure type is converted into an object start descriptor or an array start descriptor. All child nodes are recursively traversed, and the type identification and conversion operations are repeatedly performed on each child node, converting the data attribute type of the child node into a type descriptor and the nested container structure within the child node into paired structure delimiters, until all leaf nodes are processed. The root node container is then supplemented with the corresponding object end descriptor or array end descriptor. All type descriptors and structure delimiters are arranged sequentially in a depth-first traversal nested order, concatenated to form a complete symbol sequence, which serves as a template field. This template field implicitly expresses the hierarchical nesting relationship of the data through the order of the symbols. The receiver reads the symbols in the same order to reconstruct the data's type structure and hierarchical layout.

[0067] In another implementation, when the sender parses the hierarchical structure of the semantic description information, a breadth-first traversal strategy can be used to process container nodes. Starting from the root node, the container structure type of the root node is converted to an object start symbol or an array start symbol. All direct child nodes of the current level are traversed, and the data attribute types of each child node are converted to type descriptors and arranged sequentially. Placeholders are inserted for child nodes containing container structures. After completing the traversal of the current level, the next level is entered, expanding all container nodes layer by layer until the leaf nodes. Corresponding structure delimiters are added to each container node to form closed boundaries. The type descriptors and structure delimiters are combined into a symbol sequence according to the nested order of the hierarchical traversal, serving as a template field. This template field clearly reflects the width characteristics of the data through hierarchical segmentation, allowing the receiver to parse in segments according to the hierarchy, reducing the processing complexity of deeply nested structures.

[0068] In another implementation, when the sender parses the hierarchical structure, for container nodes of variable-length array type, a length descriptor indicating the number of elements can be inserted after the array start descriptor, followed by the type descriptors of each element arranged sequentially, and closed with the array end descriptor. For optional fields, an existence flag can be inserted before the corresponding type descriptor to indicate whether the field actually exists in the numeric field. The structure delimiter, length descriptor, and existence flag are combined in a nested order to form a symbol sequence, forming a template field. This template field, by expanding the symbol set, can express the parsing rules for variable-length structures and optional fields, and the receiver dynamically adjusts the parsing strategy based on the length descriptor and existence flag.

[0069] The above implementation transforms the hierarchical structure of semantic description information into a symbol sequence composed of type descriptors and structure delimiters, achieving a compact binary expression of data structure parsing rules. Specifically, the type mapping relationship ensures unified encoding of basic data types and eliminates redundant information in type text; the paired setting of structure delimiters clearly defines container boundaries, accurately presenting hierarchical nesting relationships in the symbol sequence, compressing complex hierarchical structure definitions into a compact symbol sequence, while fully preserving type and hierarchical information. This avoids the transmission and storage of redundant content such as field names and type text, reduces the transmission and storage overhead of structure description, and improves the processing efficiency of data serialization and deserialization.

[0070] To reduce the redundant overhead of structured data during transmission and parsing, and to improve the efficiency of semantic indexing and field location, semantic description information is encoded into semantic option fields. Specifically, this includes: extracting the names of each field contained in the semantic description information and counting the total number of field names; traversing each field name and obtaining the symbol sequence number corresponding to each field name in the template field; and constructing serialized data containing the total number and several key-value pairs as semantic option fields.

[0071] Each key-value pair consists of a symbolic index and a field name. The symbolic index indicates the reference index of the corresponding field name during data parsing. The receiver can quickly match the field name using the symbolic index, achieving accurate location and retrieval of business data fields without repeatedly transmitting the complete field name, thus further reducing data transmission volume while ensuring the accuracy and consistency of field mapping relationships during data parsing. The field name refers to the semantic identifier of each field in the business data, which is stripped during the encoding of numeric fields and retained in the semantic option field. In other words, the field name can be the identifier name of each data node in the semantic description information, used to distinguish the semantic meaning of different fields at the business level, and serves as the value part of the key-value pair in the semantic option field. The total count is a statistical count of all field names in the semantic description information, serving as the header metadata of the semantic option field, allowing the receiver to pre-allocate the required storage space before parsing. The symbolic index is an incrementally assigned number according to the order of appearance of the field name in the template field symbol sequence, serving as the key part of the key-value pair in the semantic option field, used to establish the association between the field name and the position of the template field.

[0072] Serialized data refers to continuous binary data formed by arranging key-value pairs in a specific format and adding header metadata, serving as the storage and transmission medium for semantic option fields. Semantic option fields consist of serialized data comprising a total number of header metadata and several key-value pairs, used to establish a mapping relationship between field names and symbol numbers.

[0073] In one implementation, when generating the semantic option field, the sender can traverse the hierarchical structure of the semantic description information, extract the field names corresponding to all leaf nodes, and use the total number of extracted field names as the total quantity. Following the order of the symbol sequence in the template field, each field name is assigned a symbol index starting from 0 and incrementing, ensuring a one-to-one correspondence between the symbol index and the field's position in the symbol sequence. Then, serialized data containing the total quantity and several key-value pairs is constructed, where each key-value pair uses the symbol index as the key and the field name as the value. The key-value pairs are arranged in ascending order of symbol index to form an ordered sequence, serving as the semantic option field. This semantic option field establishes a direct mapping from field names to template field positions through symbol indexes. When the receiver parses the data, it can look up the corresponding symbol index based on the field name, locate the corresponding parsing rule from the symbol sequence of the template field, and achieve an indirect index from name to type.

[0074] In another implementation, hash indexes can be used to optimize the lookup efficiency of key-value pairs when generating semantic option fields. Specifically, after extracting field names and counting the total number, a hash value is calculated for each field name, and a symbolic index is assigned to the field name based on the hash value, establishing a mapping between the symbolic index and the hash value. When constructing semantic option fields, key-value pairs are sorted by hash value, and metadata such as the number of hash buckets and collision list pointers are appended to the header. During parsing, the receiver can quickly locate the range of symbolic indices using the hash value of the field name, reducing the overhead of traversal lookup. This semantic option field accelerates the mapping process from field names to symbolic indices by introducing a hash index.

[0075] In another implementation, for data structures where field names share a common prefix or follow hierarchical naming conventions, after extracting the field names and counting their total number, the common prefix portion of the field names can be identified, a prefix dictionary can be constructed, and a prefix identifier can be assigned to each unique prefix. The field names in key-value pairs are replaced with a combination of "prefix identifier + suffix," while the symbolic sequence number remains unchanged. Simultaneously, a prefix dictionary is appended to the header of the semantic option field, recording the mapping relationship between the prefix identifier and the complete prefix. During parsing, the receiver first reconstructs the complete field name using the prefix dictionary, and then locates the parsing rule based on the symbolic sequence number. This semantic option field further compresses its data volume by eliminating redundant prefix storage in field names.

[0076] The above implementation, by declaring the total number of field names in advance, allows the receiver to pre-allocate the storage resources required for parsing, avoiding the additional overhead of dynamic expansion. The sequential allocation of symbol indices ensures a one-to-one correspondence between key-value pairs and the template field symbol sequence, guaranteeing the accuracy of parsing rule location. Furthermore, the ordered arrangement of key-value pairs or hash index optimization reduces the time complexity of field name lookup, improving parsing efficiency in large-scale field scenarios. Overall, the semantic option field, as the name semantic carrier of the structure descriptor, complements the type structure description of the template field, ensuring the traceability of field name semantics and achieving a compact and efficient expression of structural description information.

[0077] To remove redundant semantic information such as field names, further compress data transmission volume, and improve serialization and parsing efficiency, numerical information is encoded to generate numerical fields. This includes: sequentially obtaining the corresponding data to be encoded in the numerical information according to the arrangement order of the symbol sequence in the template field; converting the data to be encoded into target encoded data based on the data type of the data to be encoded; and concatenating the target encoded data in order to generate a numerical field that does not contain field names.

[0078] In this sequence, each symbol corresponds to a data node in the numerical information, used to identify the data attribute type or container boundary of that node. The order of arrangement represents the sequence of symbols in the symbol sequence, determined by the hierarchical traversal method of the semantic description information. The encoding of numerical fields must strictly follow this order to ensure a one-to-one correspondence with the template fields. The data to be encoded refers to the specific field value in the numerical information corresponding to a symbol position in the template field, and can be read sequentially according to the symbol sequence. The data type is the basic data type of the data to be encoded, including integer, floating-point, boolean, and string types, specified by the type descriptor at the corresponding position in the template field. The target encoded data is the result of converting the data to be encoded into a specific binary format according to its corresponding data type. The numerical field is binary data composed of the target encoded data concatenated sequentially according to the symbol sequence.

[0079] In one implementation, when generating the numeric field, the sender can read the symbol sequence of the template field and parse it sequentially from left to right. When a type descriptor is encountered, the data to be encoded corresponding to that symbol position is extracted from the numeric information and converted into target encoded data according to the encoding rules specified by the type descriptor in the template field. For example, integer data is converted to fixed-width two's complement, floating-point data is converted to IEEE standard binary format, Boolean data is converted to single-bit representation, and string data is converted to a byte sequence. When a structure delimiter is encountered, the data to be encoded is not read; it is skipped as a hierarchical boundary marker. All target encoded data are concatenated in the order of the symbol sequence to generate continuous binary data without any field names, which is the numeric field. This numeric field eliminates all self-descriptive information through pure binary concatenation, and the receiver can extract binary fragments of corresponding length from the numeric field based on the symbol sequence of the template field to reconstruct the original data.

[0080] In another implementation, for scenarios where the template field contains optional fields, the actual existence status of each field can be marked in the numerical information. The system traverses the data in the order of the symbol sequence. When an existence marker indicates that a field exists, the data to be encoded is read and converted into the target encoded data; when the field does not exist, no data is read or storage space is allocated, and the current position is skipped. The concatenated target encoded data sequence is tightly arranged with no redundant intervals. By dynamically omitting optional fields that do not appear, the data volume is further compressed.

[0081] In another implementation, when generating a numeric field, if the sender encounters a large number of repetitive values, it can identify the repetitive value patterns after obtaining the data to be encoded in symbol sequence order. A numeric dictionary is then constructed for the frequently occurring values, and dictionary indices are assigned. During encoding, the dictionary indices replace the original repetitive values ​​as the target encoded data. The original values ​​are fully encoded only upon their first appearance and are placed as dictionary content at the beginning of the numeric field. Subsequent repetitive values ​​are represented by their indices. All target encoded data or dictionary indices are concatenated in symbol sequence order, with the numeric dictionary appended to the header as decoding metadata. This method reduces the transmission volume in high-frequency repetitive data scenarios by eliminating redundant storage of repetitive values.

[0082] The above implementation method, based on the order constraints of the template field symbol sequence and data type encoding rules, achieves compact binary encoding of numerical information, ensuring a precise correspondence between the numerical field and the template field symbol positions. The encoding process decouples the type structure description from the actual data content, minimizing the amount of data transmitted while ensuring accurate and complete data parsing, effectively reducing network transmission overhead and improving end-to-end data interaction efficiency.

[0083] Optionally, based on the data type of the data to be encoded, it is converted into target encoded data, including: for numeric data to be encoded, it is converted into target encoded data with a preset bit width; for string data to be encoded, the byte length of its string content is obtained, the byte length is converted into an indefinite-length integer, and the binary data of the string content is concatenated to generate target encoded data.

[0084] Numeric types are data attribute types of the data to be encoded, including integer and floating-point types. Their values ​​belong to a set of numbers, and after encoding, they occupy a fixed-length binary space. String types are also data attribute types of the data to be encoded, with values ​​being character sequences. After encoding, they occupy a variable-length binary space, requiring additional length information to define data boundaries. Preset bit width refers to the fixed number of bits predefined when binary encoding numeric data. It can be determined based on the value range and precision requirements of the numeric type, ensuring consistent encoding length for data of the same type. Byte length refers to the total number of bytes occupied after encoding string data into binary, used to identify the actual length of the string content. Variable-length integers refer to integer values ​​represented using variable-length encoding. Smaller values ​​occupy fewer bytes, and larger values ​​occupy more bytes, often used to compactly represent length information such as byte length. String content refers to the actual character sequence carried by string data. Binary data refers to the bit stream formed after data encoding, and is the final storage and transmission carrier for numeric fields and target encoded data.

[0085] In this embodiment, when processing numeric data to be encoded, the sender can convert the data to be encoded into target encoded data according to the preset bit width specified by the corresponding type descriptor in the template field. For example, integers can be converted into two's complement binary codes with a fixed bit width, negative numbers can be encoded by extending the sign bit, and positive numbers can be filled with high bits according to the bit width, so that numeric values ​​of the same type always maintain a consistent encoding length.

[0086] When processing string-type data to be encoded, the character sequence can be converted into a byte sequence according to a unified encoding standard, and the total number of bytes obtained can be used as the byte length. The byte length is then encoded into a variable-length integer, using a variable-length encoding method with consecutive high bits. Smaller lengths are represented by a single byte, and larger lengths are represented by multiple bytes in a progressive manner. The variable-length integer and the byte sequence of the string content are concatenated in order to generate the target encoded data.

[0087] The above implementation achieves efficient binary conversion between different data types by using fixed-width encoding for numeric types and variable-length prefix encoding for string types. Fixed-width encoding ensures the deterministic length of numeric encodings, allowing the receiver to accurately extract fixed-length binary segments based on symbol positions, reducing parsing computational overhead. Variable-length encoding provides a compact representation of length information, enabling the receiver to determine string boundaries without the need for delimiters, avoiding escaping and scanning overhead. This approach ensures accuracy in parsing different data types while minimizing encoding size, effectively reducing network bandwidth consumption and improving the overall efficiency of data serialization and deserialization.

[0088] Figure 2 This is a flowchart of a data transmission method according to an embodiment of the present invention. This embodiment is applicable to any situation requiring efficient transmission of structured data. The method can be applied to the receiving party and can be executed by a data transmission device, which can be implemented in hardware and / or software and can be configured in a computing device. Figure 2 As shown, the method includes: S210. Receive the data packet sent by the sender, parse the header field of the data packet, and obtain the descriptor indicator flag.

[0089] In this context, a data packet refers to a binary transmission unit encoded by the sender, including two types: first data packet and second data packet. It is also the basic object that the receiver needs to parse and process. The header field is a control information area located at the beginning of the data packet, containing basic metadata such as the descriptor indicator flag and service data identifier. The descriptor indicator flag is a special segment in the header field used to explicitly identify the presence status of structural descriptors in the current data packet. The receiver determines the subsequent parsing strategy based on this flag.

[0090] In this embodiment, the receiver can receive data packets transmitted by the sender through the network port, read a fixed-length header field from the beginning of the data packet, parse the binary bit information therein, and extract the current value of the descriptor indicator flag. When receiving data packets, the receiver can also use an asynchronous event-driven mechanism to handle network data arrival events; when the data packet arrives, a parsing callback is triggered to extract the fixed-length header field from the beginning of the data packet buffer and complete the parsing of the descriptor indicator flag.

[0091] S220. If the descriptor indicator flag is valid, read the structure descriptor from the data packet and associate the structure descriptor with the service data identifier of the data packet for storage.

[0092] The business data identifier is a unique identifier for a type of business data structure, serving as the index key for the structure descriptor in the local cache. Association storage is the operation of establishing a mapping relationship between the structure descriptor and the corresponding business data identifier and writing it into the local cache, allowing subsequent data packets carrying the same business data identifier to directly reuse that structure descriptor.

[0093] If the descriptor indicator flag is valid, the current data packet is determined to be the first data packet. The structure descriptor can continue to be read from the header fields, associated with the business data identifier in the header fields and stored in the local cache, and the subsequent numerical fields can be parsed based on the structure descriptor.

[0094] In the specific implementation, when the descriptor indicator flag is valid, the binary content of the structure descriptor can be read from the position after the header fields of the data packet. The read boundary is determined according to the preset length of the structure descriptor or the internal end marker, and the template field and semantic option field are completely extracted. The business data identifier is obtained from the header fields, and the structure descriptor and the business data identifier are written into the local cache hash table storage area in the form of key-value pairs. A direct mapping is established with the business data identifier as the key and the storage address of the structure descriptor as the value. This method achieves efficient storage and fast lookup of the structure descriptor through hash indexing.

[0095] Furthermore, the receiver can read the symbol sequence of the template fields, parse the order of type descriptors and structure delimiters, and construct a type resolution rule table; simultaneously, it can read the total number and key-value pair sequence in the semantic option fields, parse the mapping relationship between symbol numbers and field names, and construct a name index table. The symbol sequence of the template fields, the type resolution rule table, the name index table, and the business data identifier are uniformly encapsulated into an internal structure description object, written to a local cache object pool, and indexed using the business data identifier. When parsing data with the same structure subsequently, this internal structure description object can be directly loaded, eliminating the need to repeatedly parse the original binary structure descriptor. This method, through pre-parsing conversion, effectively reduces the computational overhead in the subsequent decoding process.

[0096] S230. If the descriptor indicator flag is invalid, then look up the stored structure descriptor based on the business data identifier.

[0097] Among them, the stored structure descriptor refers to the structure descriptor that the receiver previously received and cached. It has been associated with the corresponding business data identifier and stored in the local cache, and can be directly reused by subsequent data packets carrying the same business data identifier.

[0098] If the descriptor indicator flag is invalid, the current data packet is determined to be the second data packet. In this case, the business data identifier can be extracted from the header fields, the structure descriptor matching the identifier can be searched in the local cache, and the subsequent numerical fields can be parsed based on the found structure descriptor.

[0099] This implementation prioritizes parsing header fields, enabling the receiver to quickly determine the data packet type and select the corresponding parsing path, thus ensuring the timeliness and accuracy of data processing.

[0100] In the specific implementation, the business data identifier can be extracted from the header field and its hash value can be calculated. A hash lookup is then performed in the local cache hash table storage area to locate the cache entry that matches the hash value. If the hash lookup is successful, the corresponding structure descriptor is obtained, and subsequent numerical fields are parsed based on it to restore the business data. If the hash lookup fails, it is determined that the corresponding structure descriptor does not exist in the local cache, and a cache miss response is returned to the sender, triggering the sender to resend the first data packet containing the structure descriptor.

[0101] When the receiver determines that the descriptor indicator flag is invalid, it can also use a hierarchical search strategy to retrieve the structure descriptor.

[0102] Specifically, the search can be performed in the memory cache layer based on the business data identifier. If a match is found, the structure descriptor is obtained and the numeric fields are parsed. If the memory cache layer does not find the record, the search continues in the disk cache layer for the persistently stored structure descriptor record. If a match is found, it is loaded into the memory cache layer and the numeric fields are parsed. If the disk cache layer still does not find the record, the structure descriptor is determined to be missing, the cache invalidation process is executed, and the sender is triggered to resend the first data packet carrying the structure descriptor.

[0103] Furthermore, while initiating a local cache lookup based on the business data identifier, the receiver can also predict the types of data packets that may arrive later and preload relevant structure descriptors into the cache. If the current business data identifier lookup is successful, the corresponding structure descriptor is used to parse the numerical field; if the lookup fails but the preloaded content is successful, the preloaded structure descriptor is reused to avoid waiting for retransmission from the sender; if both the lookup and preloading fail, a structure descriptor request process is triggered. This method, through a predictive preloading mechanism, effectively reduces the waiting latency caused by cache misses.

[0104] S240. Decode the numerical fields in the data packet based on the structure descriptor to restore the business data.

[0105] Decoding refers to the process by which the receiver, according to the parsing rules in the structure descriptor, reverse-engineers the binary data within the numeric fields into structured business data containing field names and their corresponding values. The business data is the final data form obtained by the receiver after decoding, containing structured information such as field names, data types, and specific values, and semantically consistent with the original business data to be transmitted by the sender. The template field provides binary data type determination and length boundary rules for the decoding process, while the semantic option field is used to restore the business semantics corresponding to the field names during decoding.

[0106] In one implementation, after obtaining the structure descriptor, the receiver can parse the symbol sequence of the template field and read each symbol sequentially from left to right. When a type descriptor is read, based on the data type it indicates, the receiver extracts the corresponding bit-width binary data from the current position of the numeric field and converts it into the original value according to the decoding rules of that type. Simultaneously, it searches for the symbol index corresponding to the current symbol position in the semantic option field, obtains the associated field name, and combines the field name with the original value into a key-value pair. When a structure delimiter is read, the receiver determines the start or end boundary of the container based on the delimiter type and adjusts the hierarchical parsing state. During this process, the receiver does not read the binary data in the numeric field. After a complete traversal of the symbol sequence, all key-value pairs are organized hierarchically to restore the structured business data.

[0107] In another implementation, the receiver can pre-compile the structure descriptor before decoding, converting the symbol sequence of the template fields into a sequence of directly executable parsing instructions. Each parsing instruction includes the data type, bit width or length acquisition method, field name index, and hierarchical state transition operation.

[0108] Decoding is performed sequentially according to the instructions. Binary data of the corresponding length is extracted from the numeric field. Based on the type identifier in the instruction, the corresponding decoding function is called to convert it into the original value. Field names are quickly retrieved from the semantic option field using the field name index and then combined into key-value pairs. When an object start instruction is encountered, a new level is created and pushed onto the stack; when an object end instruction is encountered, the current level is exited and the stack is popped. This method, through pre-compilation, moves the symbol resolution overhead forward, reducing the computational cost of repeated decoding.

[0109] The receiver can also use an on-demand decoding strategy for numeric fields.

[0110] During decoding, the symbol sequence of the template field and the key-value pair mapping of the semantic option field can be fully parsed, and an index table from field name to symbol position can be established, but the binary data in the numeric field is not extracted at this time. When the upper-layer application needs to obtain the value of a certain field, the corresponding symbol number is located from the index table according to the field name, its offset position and data width in the numeric field are calculated, and then the binary data is extracted from the corresponding position and decoded into the original value and returned. For fields that are not requested, no decoding operation is performed. This method effectively reduces unnecessary computational overhead through delayed decoding.

[0111] Optionally, the numerical fields in the data packet are decoded based on the structure descriptor, including: parsing the template field in the structure descriptor to obtain the symbol sequence; determining the field name corresponding to each symbol number in the symbol sequence based on the semantic option field in the structure descriptor; and extracting binary data from the numerical fields sequentially according to the order of the symbol sequence and converting it into original data, thereby achieving efficient self-parsing decoding without the need for external information.

[0112] Specifically, the receiver can parse the template field in the structure descriptor, read the type descriptor and structure delimiter sequentially according to a fixed length of binary bits, arrange them into a symbol sequence according to the reading order, and record the position index of each symbol in the sequence. It parses the semantic option field, reads the total number of fields, extracts each key-value pair sequentially, establishes a mapping table between symbol indices and field names, and stores the symbol indices as lookup keys and field names as lookup values ​​in a hash index structure. During decoding, it traverses each symbol according to the order of the symbol sequence. When a symbol is a type descriptor, it extracts the corresponding bit-width binary data from the current offset position of the numeric field based on the data type indicated by the type descriptor, converts the binary data into raw data, and simultaneously uses the position index of the symbol as the symbol indices to query the mapping table to obtain the field name, combining them into field data with a name and value. When a symbol is a structure delimiter, it adjusts the current nesting state and does not extract binary data. Following the complete traversal order of the symbol sequence, it performs binary data extraction, conversion, and name association for each symbol, restoring the complete business data.

[0113] Alternatively, after parsing the template field to obtain the symbol sequence, the receiver can pre-merge the key-value pairs of the symbol sequence and the semantic option field, arranging the key-value pairs in ascending order of symbol number to align the order of the key-value pairs with the position index of each symbol in the symbol sequence, generating a sequential mapping array of symbol sequences and field names. During decoding, each symbol is processed sequentially according to the order of the symbol sequence. When the symbol is a type descriptor, the corresponding bit-width binary data is extracted from the numeric field and converted into the original data. Simultaneously, the field name is obtained from the current position in the sequential mapping array, without additional hash calculations or index lookups. When the symbol is a structure delimiter, the position is skipped without extracting data. This implementation eliminates the computational overhead of lookup operations through pre-merging and alignment, improving decoding speed.

[0114] The above implementation ensures that the extraction order of binary data in the numerical field is consistent with the encoding order of the sender by sequentially traversing the symbol sequence, thus guaranteeing the determinism of decoding.

[0115] The technical solution of this embodiment involves receiving data packets sent by the sender, parsing the header fields of the data packets to obtain the descriptor indicator flag; if the descriptor indicator flag is valid, the structure descriptor is read from the data packet and associated with the service data identifier of the data packet for storage; if the descriptor indicator flag is invalid, the stored structure descriptor is searched based on the service data identifier; and the numerical fields in the data packet are decoded based on the structure descriptor to restore the service data. This allows the receiver to complete data decoding according to the parsing rules carried by the data packet itself, without relying on external protocol templates or preset configurations, ensuring the universality and self-parsing capability of the parsing process; at the same time, by caching and reusing the structure descriptor, repeated parsing and redundant data processing are avoided, reducing decoding computation overhead and improving data parsing efficiency; by distinguishing the data packet type and matching the corresponding structure information through the flag, the accuracy and reliability of data decoding are ensured, and the data restoration error rate is reduced.

[0116] As an optional embodiment of the above embodiments, specific application scenario examples are provided to enable those skilled in the art to further understand the technical solutions of the embodiments of the present invention. Specifically, please refer to the following detailed content.

[0117] This technical solution implements data transmission between the sender and receiver through a self-parsing binary data exchange protocol. The protocol describes the complete data structure using data templates (i.e., template fields), defines data semantic information through semantic option fields, and further reduces the redundancy of encoded data through a dynamic data exchange mechanism.

[0118] Based on this, the sender can encode the business data to be transmitted using a binary data exchange protocol to obtain data packets.

[0119] The data packet may contain one or more of the following fields: Header field, Id field, template field, Options field, and Payload field.

[0120] The header field is used to indicate whether the subsequent fields exist.

[0121] The identifier field is used to define the business data identifier for business data.

[0122] Template fields are used to define the complete data structure of business data.

[0123] The semantic option field is used to define the semantic field names of each component in the business data. It consists of the total number of field names (indefinite-length integers) and a set of key-value pairs (index + field name).

[0124] Numeric fields are used to define the coded values ​​of business data, and can consist of the number of data bytes (indefinite-length integers) and a numeric byte string.

[0125] The header field can occupy 1 byte and includes multiple indicator flags. For example, see [link to header field documentation]. Figure 3 R is located in bits 4-7 and is a reserved bit; I is used to indicate whether the identifier field exists; T is used to indicate whether the template field exists; O is used to indicate whether the semantic option field exists; P is used to indicate whether the numeric field exists; B is used to specify endianness, 0 indicates big endianness; 1 indicates little endianness (endianness determines the encoding order of multi-byte numeric values, and this article will not describe the handling of endianness in detail).

[0126] Variable-length integers are used to record quantities of varying lengths, such as string length, number of options, array length, etc. Variable-length integers are dynamically encoded based on their numerical value, effectively saving encoding space. They occupy between 1 and 4 bytes, and can represent up to 1GB, meeting the needs of various scenarios.

[0127] The highest two bits of the first byte of a variable-length integer are used to identify the total number of bytes occupied by the variable-length integer. Its encoding method and corresponding numerical range can be found in Table 1.

[0128] Table 1: In Table 1, the encoding structure consists of a 2-bit byte identifier and subsequent numerical bits, with different identifiers corresponding to different byte counts and numerical ranges: When encoded as “00+[6]”, this variable-length integer occupies 1 byte. Among them, the two highest bits “00” indicate that the total number of bytes is 1, and the remaining 6 bits are used to store the value, corresponding to a value range of 0x0 to 0x40; When encoded as “01+[6][8]”, this variable-length integer occupies 2 bytes. The two highest bits “01” indicate that the total number of bytes is 2. The remaining part consists of the lower 6 bits of the first byte and the 8 bits of the second complete byte, forming a 14-bit value, corresponding to a value range of 0x41 to 0x4000. When encoded as “10+[6][8][8]”, this variable-length integer occupies 3 bytes. Among them, the two highest bits “10” indicate that the total number of bytes is 3, and the remaining part consists of the lower 6 bits of the first byte and the 8 bits of the second and third bytes, forming a 22-bit value, corresponding to a value range of 0x4001 to 0x400000; When encoded as “11+[6][8][8][8]”, this variable-length integer occupies 4 bytes. The two highest bits “11” indicate that the total number of bytes is 4. The remaining part consists of the lower 6 bits of the first byte and 8 bits of each of the second to fourth bytes, forming a 30-bit value range of 0x400001 to 0x40000000.

[0129] The above encoding method dynamically indicates the number of bytes by using the highest two bits of the first byte, allowing smaller values ​​to be stored with fewer bytes and larger values ​​to be stored up to 4 bytes. This achieves dynamic optimization of the encoding size while ensuring the efficiency of the decoding process and the continuity of the numerical range. It can accurately parse variable-length integers without additional separators.

[0130] Template fields are data structures described by strings composed of letters and symbols, fully recording the data's structure, numeric type, and size information. Template fields can be generated based on data templates. For example, see Table 2 for a data template.

[0131] Table 2: Continuing from Table 2 above; Table 2 lists the template symbols for each data type after conversion. Here, `int8`, `uint8`, `int16`, `int32`, `uint32`, `int64`, `uint64`, `uint64`, `double`, and `string` represent data attribute types. The corresponding template symbols are the converted type descriptors. `Map` and `array` represent container structure types. The corresponding template symbols are the converted structure delimiters.

[0132] For example, an example will be used to illustrate the generation of template fields.

[0133] Assume the service data to be transmitted is represented as follows: { var1:<int32>, name1: <string>, obj1:{ Var2: <int32>, name2: <string>, }, arr1: }; Where, int32 is a 32-bit integer, and its converted type descriptor is i; string is a string, and its converted type descriptor is s; { represents the start symbol of an object, and its converted structure delimiter is {;} represents the end symbol of an object, and its converted structure delimiter is}; [ <int32>[] represents a 32-bit integer array. The converted type descriptor and structure delimiter combined are represented as [i]. The type descriptor and the structure delimiter are serialized and combined to generate a template field represented as: is{is}[i]. Since the data as a whole is treated as an object, the symbol for the global data object in the template can be omitted.

[0134] The following example demonstrates business data and its corresponding template fields: Assume the service data to be transmitted is represented as follows: { var1:100, name1:"soo", obj1:{ var2:200, name2:"too", }, arr1:[100,200] }; The corresponding template field is represented as: is{is}[i].

[0135] Continuing with the example above, the process of generating numeric fields will be illustrated below. Numeric fields are strictly coded according to the format of the template fields, as follows: For numeric data types (such as intN, uintjN, float, double), the target encoded data is encoded using fixed-byte encoding.

[0136] For string-type data to be encoded, the target encoded data consists of a variable-length integer and a string sequence, where the variable-length integer records the string length.

[0137] For data to be encoded of array type (such as array), the target encoded data consists of a variable-length integer and a sequence of array elements, where the variable-length integer records the array length.

[0138] For data to be encoded that is an object type (such as a map), the composition range of the object members is defined to generate the target encoded data; the overall data is treated as an object by default, and the outermost object symbol is omitted in the template field.

[0139] Combination Figure 4 The numerical field diagram shown illustrates the encoding methods for different data types as follows: The integer 100 corresponding to the template field symbol i is stored consecutively in fixed bytes; the string soo corresponding to the template field symbol s is composed of an integer of variable length 3 and the character sequence s, o, o in sequence; the object corresponding to the template field symbol {is} contains an integer 200 and the string too, where too is composed of an integer of variable length 3 and the character sequence t, o, o; the array corresponding to the template field symbol [i] is composed of an integer of variable length 2 and integer elements 100 and 200 in sequence.

[0140] Specifically, each symbol can be parsed sequentially according to the order of the symbol sequence in the template field. When a type descriptor is parsed, the corresponding data to be encoded is obtained from the numerical information, and based on the data type of the data to be encoded, it is converted into target encoded data. When a structure delimiter is parsed, no target encoded data is generated; it is only used as a container level boundary marker. All target encoded data are concatenated in order to generate a numerical field without field names, thereby reducing transmission bandwidth consumption and improving parsing efficiency. The template field guides the encoding order, data type, and boundaries of the numerical field. Each data unit is arranged closely according to the order of the template symbol sequence, without redundant separator information. The receiver can accurately parse all data in the numerical field based on the same template field.

[0141] Continuing with the example above, the process of generating semantic option fields will be illustrated below. Semantic option fields are used to store field names and can also be expanded to record other information. Semantic description information from business data is recorded in the semantic option field as an array of key-value pairs (name index: field name), the encoding format of which can be found in [reference needed]. Figure 5 .

[0142] like Figure 5 As shown, the overall structure of the semantic option field is as follows: First, the number of field names is encoded using a variable-length integer, used to record the total number of subsequent field name key-value pairs, i.e., the total number of field names in the business data. When the receiver parses the data, it can first read this number value, pre-allocate storage space, and then read the subsequent key-value pair sequence (1 to n) in sequence.

[0143] The key-value pair sequence consists of n field names, each field name comprising two parts: a symbol index (i.e., a name index) and the field name. The symbol index can be encoded using a variable-length integer, corresponding to the symbol's index in the template field, i.e., the field's position index in the symbol sequence, used to establish the association with the template field. The field name is the field name itself, which can be encoded as a string, consisting of a variable-length integer representing the length of the string and a sequence of characters.

[0144] Field names are stored consecutively in index order, without redundant separators. During parsing, the receiver first reads the number of field names, then sequentially reads the symbolic index and name of each key-value pair, establishing an "index-field name" mapping table. Subsequent decoding of numeric fields allows for quick matching of field names using the symbolic index in the template fields, reconstructing the complete structured business data. This encoding method, through the dynamic encoding characteristics of variable-length integers, achieves compact storage of the number of field names, indexes, and field names, reducing the transmission volume of semantic option fields while ensuring high efficiency in the parsing process.

[0145] The symbol index can be the symbol number corresponding to the field name in the template field. For example, referring to Table 3, the valid field names include: Var1, name1, obj1, var2, name2, and arr1. The symbol numbers corresponding to each valid field name can be represented as: 0, 1, 2, 3, 4, and 6.

[0146] Table 3: A structural diagram of the semantic option field can be found here. Figure 6 This diagram illustrates a binary encoding instance of a semantic option field, visually presenting the organization of the field names in the form of a specific byte stream.

[0147] like Figure 6 As shown, the semantic options field begins with a variable-length integer that records the total number of subsequent field names. In this example, the value is 6, indicating that there are 6 key-value pairs (represented as key:value pairs).

[0148] Each key-value pair consists of two parts: a symbolic number (i.e., the key) and a field name (i.e., the value); The symbol number is stored as a variable-length integer and is used to identify the symbol number of the field in the template field, i.e., the field position index; Field names are stored as strings and encoded using a format of "variable-length integer length + character sequence". The variable-length integer records the byte length of the field name, followed by the actual character sequence of the field name.

[0149] by Figure 6 Taking the first three key-value pairs as an example: The first key-value pair (0: var1): The symbol index is 0, the corresponding field name is var1, where 4 is the length of the field name, and the following bytes are the character sequence var1; The second key-value pair (1: name1): The symbol number is 1, the corresponding field name is name1, where 5 is the length of the field name, and the following bytes are the character sequence name1; The third key-value pair (2:obj1): The symbol number is 2, the corresponding field name is obj1, where 4 is the length of the field name, and the following bytes are the character sequence obj1; Subsequent key-value pairs are arranged sequentially in the same way, such as symbol number 3 corresponding to var2, symbol number 4 corresponding to name2, and symbol number 6 corresponding to arr1. All key-value pairs are stored consecutively in symbol index order, with no redundant separator information, resulting in a compact overall structure. When parsing, the receiver can first read the total number in the semantic option field, and then sequentially obtain each index and name to establish a complete "symbol number-field name" mapping table. Subsequently, when decoding numeric fields, the field name can be quickly matched using the symbol number in the template field to reconstruct the structured business data containing the field semantics.

[0150] In this technical solution, by separating semantic description information and numerical information, redundant semantic information directly appended to the numerical information is reduced. Furthermore, a dynamic data exchange mechanism further reduces the amount of semantic information repeatedly transmitted during communication. Typically, during communication, data with the same service data identifier carrying different values ​​are transmitted multiple times. Apart from the different numerical parts, their semantic description information is identical. Therefore, during communication, the sender can encode the semantic description information into the first data packet when sending the same type of data for the first time. The receiver caches the structure descriptor locally, and when subsequently receiving data with the same service data identifier, it can query the corresponding structure descriptor from the local record to decode the numerical field.

[0151] Next, we can introduce the data exchange method between the sender (encoder) and the receiver (decoder).

[0152] See Figure 7 When the sender receives the service data to be transmitted, it can determine whether a structure descriptor corresponding to the service data identifier exists in its local cache based on the service data identifier in the service data. If so, it indicates that the data with the same service data identifier is not being sent for the first time, so there is no need to encode the semantic description information into the data packet, and a second data packet containing header fields and numeric fields is constructed. If not, it indicates that the data with the same service data identifier is being sent for the first time, and the structure descriptor can be saved in the local cache to construct a first data packet containing header fields, structure descriptor, and numeric fields. The first or second data packet is then sent to the receiver.

[0153] See Figure 8 When the receiver receives an encoded data packet, it can read the header field and the identifier field of the first byte of the packet. It checks the header field to see if the descriptor indicator flag corresponding to the structure descriptor is valid. If yes, it indicates that this is the first time data with the service data identifier has been received; the receiver reads the structure descriptor from the packet and associates and stores it with the service data identifier of the packet. If no, it indicates that data with the same service data identifier has already been received; the receiver looks up the stored structure descriptor based on the service data identifier. Based on the structure descriptor, the receiver decodes the numerical fields in the data packet to reconstruct the service data.

[0154] This protocol employs binary encoding, which offers higher parsing efficiency compared to structured text formats. The encoded binary data is self-descriptive, capable of carrying complete data structure and semantic information. Simultaneously, the template fields are readable, facilitating an understanding of the overall data structure. The protocol as a whole is self-parsing; the decoding process does not rely on additional semantic description information or external configuration information. Furthermore, by employing variable-length integer encoding, separating the generation of semantic description information and numerical information, and using a dynamic data exchange mechanism, data redundancy is reduced, data encoding size is decreased, and the protocol's adaptability in network bandwidth-constrained scenarios is improved.

[0155] Figure 9 This is a schematic diagram of a data transmission device according to an embodiment of the present invention. Figure 9 As shown, the device is configured on the sender and includes: an acquisition module 310, a detection module 320, a first data packet construction module 330, a second data packet construction module 340, and a sending module 350.

[0156] The system includes: an acquisition module 310 for acquiring service data to be transmitted, the service data including a service data identifier, numerical information, and semantic description information; a detection module 320 for determining whether a structure descriptor corresponding to the service data identifier exists in the local cache; a first data packet construction module 330 for, if not, encoding the semantic description information into a structure descriptor, encoding the numerical information to generate a numerical field, constructing a first data packet including a header field, the structure descriptor, and the numerical field, and setting the position of the indicator flag corresponding to the structure descriptor in the header field to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field; a second data packet construction module 340 for, if yes, constructing a second data packet including a header field and the numerical field, and setting the position of the indicator flag corresponding to the structure descriptor in the header field to invalid; and a sending module 350 for sending the first data packet or the second data packet to the receiver.

[0157] The technical solution of this embodiment obtains the business data to be transmitted, which includes a business data identifier, numerical information, and semantic description information; determines whether a structure descriptor corresponding to the business data identifier exists in the local cache; if not, the semantic description information is encoded into a structure descriptor, the numerical information is encoded to generate a numerical field, and a first data packet containing a header field, a structure descriptor, and a numerical field is constructed, with the indicator flag corresponding to the structure descriptor in the header field set to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field; if yes, a second data packet containing a header field and a numerical field is constructed, with the indicator flag corresponding to the structure descriptor in the header field set to invalid; the first or second data packet is sent to the receiver. This solves the problems of high data transmission redundancy and low parsing rate in existing general self-parsing protocols, as well as the dependence on external configuration and poor maintainability of binary protocols. It realizes the general self-parsing capability of carrying data parsing rules in the data packet and being able to decode independently without external information. At the same time, it avoids repeated transmission by caching and reusing the structure descriptor, reducing redundant data and bandwidth occupation, and improves encoding and decoding efficiency through compact binary encoding and sequential parsing. This achieves the universality, self-parsing capability, and high efficiency of data transmission, and improves the overall performance of data interaction.

[0158] Based on the above-described apparatus, optionally, the structure descriptor includes a template field and a semantic option field; the template field is a sequence of symbols composed of a type descriptor and a structure delimiter, and the semantic option field contains key-value pairs of symbol indices and field names, wherein the symbol indices are used to indicate the reference index of the field name in the template field.

[0159] Optionally, based on the above-mentioned device, the first data packet construction module 330 includes: a first unit, used to parse the hierarchical structure of the semantic description information and identify the data type of each data node in the hierarchical structure; the data type includes data attribute type and container structure type; a second unit, used to convert the data attribute type into a corresponding type descriptor and the container structure type into a corresponding structure delimiter based on a preset type mapping relationship; and a third unit, used to serialize and combine the type descriptor and the structure delimiter according to the nesting order of the hierarchical structure to generate the template field composed of a sequence of symbols used to indicate data parsing rules.

[0160] Based on the above-described device, optionally, the first data packet construction module 330 includes: a fourth unit, used to extract the names of each field contained in the semantic description information and count the total number of the field names; a fifth unit, used to traverse each of the field names and obtain the symbol index corresponding to each field name in the template field; and a sixth unit, used to construct serialized data containing the total number and several key-value pairs as the semantic option field; wherein each key-value pair consists of the symbol index and the field name, and the symbol index is used to indicate the reference index of the corresponding field name during data parsing.

[0161] Based on the above-mentioned device, optionally, the first data packet construction module 330 includes: a seventh unit, used to sequentially obtain the corresponding data to be encoded in the numerical information according to the arrangement order of the symbol sequence in the template field; an eighth unit, used to convert the data to be encoded into target encoded data based on the data type of the data to be encoded; and a ninth unit, used to concatenate the target encoded data in sequence to generate a numerical field that does not contain field names.

[0162] Based on the above-mentioned device, optionally, the eighth unit is used to convert numeric data to be encoded into target encoded data with a preset bit width; and to obtain the byte length of the string content of the data to be encoded, convert the byte length into an indefinite-length integer, and concatenate the binary data of the string content to generate target encoded data.

[0163] The data transmission device provided in the embodiments of the present invention can execute the data transmission method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0164] Figure 10 This is a schematic diagram of a data transmission device according to an embodiment of the present invention. Figure 10 As shown, the device is configured on the receiving end and includes: a parsing module 410, a storage module 420, a lookup module 430, and a decoding module 440.

[0165] The parsing module 410 is used to receive data packets sent by the sender, parse the header fields of the data packets, and obtain the descriptor indicator flag. The storage module 420 is used to read the structure descriptor from the data packet if the descriptor indicator flag is valid, and associate the structure descriptor with the service data identifier of the data packet for storage. The lookup module 430 is used to look up the stored structure descriptor based on the service data identifier if the descriptor indicator flag is invalid. The decoding module 440 is used to decode the numerical fields in the data packet based on the structure descriptor to restore the service data.

[0166] The technical solution of this embodiment involves receiving data packets sent by the sender, parsing the header fields of the data packets to obtain the descriptor indicator flag; if the descriptor indicator flag is valid, the structure descriptor is read from the data packet and associated with the service data identifier of the data packet for storage; if the descriptor indicator flag is invalid, the stored structure descriptor is searched based on the service data identifier; and the numerical fields in the data packet are decoded based on the structure descriptor to restore the service data. This allows the receiver to complete data decoding according to the parsing rules carried by the data packet itself, without relying on external protocol templates or preset configurations, ensuring the universality and self-parsing capability of the parsing process; at the same time, by caching and reusing the structure descriptor, repeated parsing and redundant data processing are avoided, reducing decoding computation overhead and improving data parsing efficiency; by distinguishing the data packet type and matching the corresponding structure information through the flag, the accuracy and reliability of data decoding are ensured, and the data restoration error rate is reduced.

[0167] Optionally, the decoding module 440 includes: a tenth unit, used to parse the template field in the structure descriptor to obtain a symbol sequence; an eleventh unit, used to determine the field name corresponding to each symbol number in the symbol sequence based on the semantic option field in the structure descriptor; and a twelfth unit, used to extract binary data from the numerical field and convert it into original data in the order of the symbol sequence.

[0168] The data transmission device provided in the embodiments of the present invention can execute the data transmission method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0169] Figure 11 This is a schematic diagram of the structure of an electronic device implementing the data transmission method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0170] like Figure 11 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or a computer program loaded from storage unit 18 into the random access memory 13. The random access memory 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0171] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0172] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data transfer methods.

[0173] In some embodiments, the data transfer method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the data transfer method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data transfer method by any other suitable means (e.g., by means of firmware).

[0174] Various implementations of the systems and techniques described above herein can be implemented in digital circuit systems, integrated circuits, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0175] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0176] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0177] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0178] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data transmission of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0179] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0180] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from read-only memory 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.

[0181] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the data transmission method provided in any embodiment of this invention.

[0182] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0183] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0184] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention. < / string> < / string>

Claims

1. A data transmission method, applied to a sender, characterized in that, The method includes: Acquire the service data to be transmitted, the service data including service data identifier, numerical information and semantic description information; Determine whether a structure descriptor corresponding to the business data identifier exists in the local cache; If not, the semantic description information is encoded into a structure descriptor, the numerical information is encoded to generate a numerical field, a first data packet containing a header field, the structure descriptor, and the numerical field is constructed, and the position of the indicator flag corresponding to the structure descriptor in the header field is set to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field; If so, then construct a second data packet containing the header field and the numeric field, and set the indicator flag corresponding to the structure descriptor in the header field to invalid; Send the first data packet or the second data packet to the recipient.

2. The method according to claim 1, characterized in that, The structure descriptor includes a template field and a semantic option field; the template field is a sequence of symbols consisting of a type descriptor and a structure delimiter, and the semantic option field contains key-value pairs of symbol indices and field names, wherein the symbol indices are used to indicate the reference index of the field name in the template field.

3. The method according to claim 2, characterized in that, Encoding the semantic description information into template fields includes: The hierarchical structure of the semantic description information is parsed to identify the data type of each data node in the hierarchical structure; the data type includes data attribute type and container structure type. Based on the preset type mapping relationship, the data attribute type is converted into the corresponding type descriptor, and the container structure type is converted into the corresponding structure delimiter; Based on the nesting order of the hierarchical structure, the type descriptor and the structure delimiter are serialized and combined to generate the template field, which consists of a sequence of symbols and is used to indicate the data parsing rules.

4. The method according to claim 2, characterized in that, Encoding the semantic description information into semantic option fields includes: Extract the names of each field contained in the semantic description information, and count the total number of the field names; Iterate through each of the field names and obtain the symbol index corresponding to each field name in the template field; Construct serialized data containing the total quantity and several key-value pairs as the semantic option field; Each key-value pair consists of a symbol number and a field name, and the symbol number is used to indicate the reference index of the corresponding field name during data parsing.

5. The method according to claim 2, characterized in that, The process of encoding the numerical information to generate a numerical field includes: According to the arrangement order of the symbol sequence in the template field, the corresponding data to be encoded in the numerical information is obtained sequentially; Based on the data type of the data to be encoded, it is converted into target encoded data; The target encoded data is concatenated in sequence to generate a numeric field that does not contain field names.

6. The method according to claim 5, characterized in that, The process of converting the data to be encoded into target encoded data based on its data type includes: For numeric data to be encoded, convert it into target encoded data with a preset bit width; For string-type data to be encoded, obtain the byte length of its string content, convert the byte length into a variable-length integer, and concatenate the binary data of the string content to generate the target encoded data.

7. A data transmission method applied to a receiver, characterized in that, The method includes: Receive data packets sent by the sender, parse the header fields of the data packets, and obtain the descriptor indicator flags; If the descriptor indicator flag is valid, then the structure descriptor is read from the data packet, and the structure descriptor is associated with and stored with the service data identifier of the data packet; If the descriptor indicator flag is invalid, then the stored structure descriptor is located based on the business data identifier; The data packets are decoded based on the structure descriptor to restore the business data.

8. The method according to claim 7, characterized in that, Decoding the numeric fields in the data packet based on the structure descriptor includes: Parse the template field in the structure descriptor to obtain the symbol sequence; Based on the semantic option field in the structure descriptor, determine the field name corresponding to each symbol number in the symbol sequence; Binary data is extracted from the numerical field and converted into raw data in the order of the symbol sequence.

9. A data transmission device, configured at the sender, characterized in that, The device includes: The acquisition module is used to acquire the business data to be transmitted, the business data including business data identifier, numerical information and semantic description information; The detection module is used to determine whether a structure descriptor corresponding to the business data identifier exists in the local cache; The first data packet construction module is used to encode the semantic description information into a structure descriptor if no, encode the numerical information to generate a numerical field, construct a first data packet containing a header field, the structure descriptor, and the numerical field, and set the position of the indicator flag corresponding to the structure descriptor in the header field to valid; the structure descriptor is used to indicate the data parsing rules of the numerical field. The second data packet construction module is configured to construct a second data packet containing a header field and the numerical field if the condition is met, and to set the indicator flag corresponding to the structure descriptor in the header field to invalid. The sending module is used to send the first data packet or the second data packet to the receiver.

10. A data transmission device, configured at a receiver, characterized in that, The device includes: The parsing module is used to receive data packets sent by the sender, parse the header fields of the data packets, and obtain the descriptor indicator flags; The storage module is used to read the structure descriptor from the data packet if the descriptor indicator flag is valid, and to associate and store the structure descriptor with the service data identifier of the data packet. The lookup module is used to look up the stored structure descriptor based on the business data identifier if the descriptor indicator flag is invalid. The decoding module is used to decode the numerical fields in the data packet based on the structure descriptor and restore the business data.