Data processing method and apparatus, and electronic device
Patent Information
- Application Number
- CN202510359215.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2026-09-29
AI Technical Summary
在线存储往往采用高效的序列化数据进行存储,相关技术中,是由后台人员为不同的数据开发对应的序列化转换程序,需要耗费大量的人力,且数据处理的效率较低
[0017]本申请实施例提供的一种数据处理方法、装置及电子设备,方法包括:获取用于对离线存储中的目标数据进行格式转换的格式转换语句;格式转换语句用于将目标数据从原数据格式,转换为在线服务所需要的目标数据格式;对格式转换语句进行解析,得到供执行引擎执行的中间格式转换语句,和适用于将目标数据在目标数据格式和序列化格式之间进行转换的目标序列化协议;由执行引擎加载并执行中间格式转换语句,以将从离线存储中获取的目标数据进行格式转换,得到格式为目标数据格式的第一中间数据;由执行引擎按照目标序列化协议,将第一中间数据进行序列化处理,得到目标序列化数据;将目标序列化数据进行在线存储;由在线服务在从在线存储中获取到目标序列化数据后,按照目标序列化协议将目标序列化数据解析为第一中间数据,并使用第一中间数据;本申请提供的方法,通过对格式转化语句进行解析,来得到中间格式转换语句和目标序列化协议,无需单独为目标数据编写对应的序列化协议代码;同时,在线服务使用同样的目标序列化协议对目标序列化数据进行解析,避免了在线服务和在线存储的协议不一致导致的数据解析错误。
Smart Images

Figure CN122838482A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a data processing method, apparatus, and electronic device. Background Technology
[0002] In the data processing workflow, data developers need to process and transform offline data and publish the results to online storage for use by online services. Online storage often uses efficient serialized data for storage. In related technologies, backend personnel develop corresponding serialization conversion programs for different types of data, which requires a lot of manpower and has relatively low data processing efficiency. Summary of the Invention
[0003] In view of this, embodiments of this application propose a data processing method, apparatus, and electronic device that can automatically generate target serialization protocols without manual writing.
[0004] The embodiments of this application are implemented using the following technical solutions:
[0005] In a first aspect, embodiments of this application provide a data processing method, comprising: acquiring a format conversion statement for converting target data in offline storage; the format conversion statement being used to convert the target data from its original data format to a target data format required by an online service; parsing the format conversion statement to obtain an intermediate format conversion statement for execution by an execution engine, and a target serialization protocol suitable for converting the target data between a target data format and a serialization format; the execution engine loading and executing the intermediate format conversion statement to convert the target data obtained from offline storage to obtain first intermediate data in the target data format; the execution engine serializing the first intermediate data according to the target serialization protocol to obtain target serialized data; storing the target serialized data online; and the online service, after obtaining the target serialized data from the online storage, parsing the target serialized data into the first intermediate data according to the target serialization protocol, and using the first intermediate data.
[0006] Secondly, embodiments of this application provide a data processing apparatus, comprising: an acquisition module, configured to acquire a format conversion statement for converting target data in offline storage; the format conversion statement is used to convert the target data from its original data format to a target data format required by an online service; a parsing module, configured to parse the format conversion statement to obtain an intermediate format conversion statement for execution by an execution engine, and a target serialization protocol suitable for converting the target data between a target data format and a serialization format; a first execution module, configured to load and execute the intermediate format conversion statement by the execution engine to convert the target data acquired from offline storage to obtain first intermediate data in the target data format; a second execution module, configured to serialize the first intermediate data by the execution engine according to the target serialization protocol to obtain target serialized data; a storage module, configured to store the target serialized data online; and a service module, configured to parse the target serialized data into the first intermediate data according to the target serialization protocol after the online service acquires the target serialized data from the online storage, and use the first intermediate data.
[0007] In some implementations, the format conversion statement includes a query statement, which includes a field definition statement for at least one original field in the target data. The field definition statement for one original field defines the field type of the corresponding original field, the attributes of the original field, and the new field mapped by the original field under the target data format. The target serialization protocol includes serialization mapping statements for the new fields mapped to each of the original fields. A serialization mapping statement for a new field defines the encoding of the new field under the serialization format. The parsing module includes: a conversion unit for performing an abstract syntax tree conversion on the format conversion statement to obtain a target abstract syntax tree; a code generation unit for generating intermediate format conversion statements for execution by the execution engine based on the target abstract syntax tree; and a query protocol generation unit for generating serialization mapping statements for the new fields mapped to each of the original fields based on the attributes defined for each original field in the query statement expressed in the target abstract syntax tree, the new fields mapped to each original field under the target data format, the field types of each new field, and the sequence encoding assigned to the new field under the serialization format.
[0008] In some implementations, the query protocol generation unit performs the following processing on each original field: based on the attribute defined for the original field, it determines the attribute keyword representing the attribute from the attribute keywords provided by the serialization format; it combines the attribute keyword, the new field mapped to the original field under the target data format, the field type of the new field, and the sequence code assigned to the new field under the serialization format to obtain the serialization mapping statement of the new field mapped to the original field.
[0009] In some implementations, the format conversion statement further includes a string splitting definition statement; the string splitting definition statement describes the field type defined for at least one splitting element and the new element field mapped by each of the splitting elements under the target data format; the at least one splitting element is selected from the splitting results obtained by splitting the original string field, and the original string field exists in the target data; the target serialization protocol further includes an element serialization mapping statement corresponding to the string splitting definition statement; the parsing module further includes a splitting protocol generation unit, used to generate an element serialization mapping statement corresponding to the string splitting definition statement based on the statement type of the string splitting definition statement expressed in the target abstract syntax tree, the field type defined for each splitting element in the string splitting definition statement, the new element field mapped by each of the splitting elements under the target data format, and the sequence encoding assigned to the new element field under the serialization format.
[0010] In some implementations, the string splitting definition statement includes a first string splitting definition statement of a first type, which is used to perform single-level splitting and select a splitting element from the splitting results. The splitting protocol generation unit is specifically used to combine the first attribute keyword, the field type defined for the splitting element, the new element field mapped by the splitting element under the target data format, and the sequence code assigned to the new element field under the serialization format, if the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the first type, to obtain a serialization mapping statement of the new element field mapped by the field type of the splitting element; wherein, the first attribute keyword is an attribute keyword indicating a non-mandatory attribute.
[0011] In some implementations, the string segmentation definition statement includes a second string segmentation definition statement of a second type. The second string segmentation definition statement is used to perform single-level segmentation and select multiple segmentation elements from the segmentation results. All of the multiple segmentation elements are mapped to a first reference field under the target data format. Specifically, if the statement type of the string segmentation definition statement expressed in the target abstract syntax tree is the second type, the second attribute keyword, the field type defined for the multiple segmentation elements, the first reference field mapped to the multiple segmentation elements under the target data format, and the sequence code assigned to the first reference field under the serialization format are combined to obtain a serialization mapping statement for the first reference field mapped to the field type of the segmentation element. The second attribute keyword is an attribute keyword indicating repeated selection of the attribute.
[0012] In some implementations, the string segmentation definition statement includes a third string segmentation definition statement belonging to the third type. The third string segmentation definition statement is used to perform single-level segmentation and select multiple segmentation elements from the segmentation results. The multiple segmentation elements are mapped to at least two second reference fields under the target data format. The segmentation protocol generation unit is specifically used to combine the first attribute keyword, the first composite field name defined for the at least two second reference fields, the first composite field defined for the at least two second reference fields, and the sequence code assigned to the first composite field under the serialization format if the statement type of the string segmentation definition statement expressed in the target abstract syntax tree is the third type, to obtain a serialization mapping statement for the first composite field. For each segmentation element, the first attribute keyword, the field type defined for the segmentation element, the second reference field mapped by the segmentation element under the target data format, and the sequence code assigned to the second reference field under the serialization format are combined to obtain a serialization mapping statement for the second reference field mapped by the field type of the segmentation element. Wherein, the first attribute keyword is an attribute keyword indicating a non-mandatory attribute.
[0013] In some implementations, the string splitting definition statement includes a fourth string splitting definition statement belonging to a fourth type. This fourth string splitting definition statement is used for two-level splitting and selects multiple splitting elements from the splitting results. These multiple splitting elements are mapped to at least two third reference fields under the target data format. Specifically, the splitting protocol generation unit is used to, if the statement type of the string splitting definition statement expressed in the target abstract syntax tree is a fourth type, combine the second attribute keyword, the second composite field name defined for the at least two third reference fields, the second composite field defined for the at least two third reference fields, and the sequence code assigned to the second composite field under the serialization format to obtain a serialization mapping statement for the second composite field. For each splitting element, the first attribute keyword, the field type defined for the splitting element, the third reference field mapped by the splitting element under the target data format, and the sequence code assigned to the third reference field under the serialization format are combined to obtain a serialization mapping statement for the third reference field mapped by the field type of the splitting element. Wherein, the first attribute keyword is an attribute keyword indicating a non-mandatory attribute; the second attribute keyword is an attribute keyword indicating a repeatedly selected attribute.
[0014] Thirdly, embodiments of this application provide an electronic device, including: a processor; and a memory storing computer instructions, which, when executed by the processor, implement the above-described method.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the above-described method.
[0016] Fifthly, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the above-described method.
[0017] This application provides a data processing method, apparatus, and electronic device. The method includes: acquiring a format conversion statement for converting target data in offline storage; the format conversion statement converting the target data from its original data format to a target data format required by an online service; parsing the format conversion statement to obtain an intermediate format conversion statement for execution by an execution engine, and a target serialization protocol suitable for converting the target data between the target data format and a serialization format; the execution engine loading and executing the intermediate format conversion statement to convert the target data acquired from offline storage to obtain first intermediate data in the target data format; and the execution engine then... The target serialization protocol serializes the first intermediate data to obtain the target serialized data; the target serialized data is then stored online; after retrieving the target serialized data from the online storage, the online service parses the target serialized data into the first intermediate data according to the target serialization protocol and uses the first intermediate data. The method provided in this application obtains the intermediate format conversion statement and the target serialization protocol by parsing the format conversion statement, eliminating the need to write corresponding serialization protocol code separately for the target data; at the same time, the online service uses the same target serialization protocol to parse the target serialized data, avoiding data parsing errors caused by inconsistencies between the protocols of the online service and the online storage.
[0018] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A schematic diagram of an application scenario involving an embodiment of this application is shown.
[0021] Figure 2 A schematic flowchart of a data processing method provided in an embodiment of this application is shown.
[0022] Figure 3 An embodiment of this application is shown. Figure 2 A flowchart of step S120.
[0023] Figure 4 An embodiment of this application is shown. Figure 3 A flowchart of step S230.
[0024] Figure 5 The following is an example of a query statement provided in an embodiment of this application.
[0025] Figure 6 An embodiment of this application is shown. Figure 5 The serialization mapping statement corresponding to the query statement shown.
[0026] Figure 7 An application architecture diagram related to an embodiment of this application is shown.
[0027] Figure 8 A schematic diagram of a data processing apparatus provided in one embodiment of this application is shown.
[0028] Figure 9 A schematic diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation
[0029] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0031] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0032] In this document, "multiple" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following associated objects are in an "or" relationship. In the following description, references to "some embodiments or some embodiment methods" describe a subset of all possible embodiments. However, it is understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0033] To facilitate understanding of this application, some terms will be explained below.
[0034] A statement, in computer programming, refers to a complete, independently executable unit. It is a fundamental component of a programming language and is used to express specific operations or instructions.
[0035] Serialization: The process of converting data into a binary string.
[0036] Deserialization: The process of converting a binary string read from online storage back into the original data.
[0037] In the data processing workflow, data developers need to process and transform offline data and publish the results to online storage for use by online services. Online storage often uses efficient serialized data formats. In related technologies, backend personnel develop corresponding serialization conversion programs for different types of data, which requires a lot of manpower and has relatively low data processing efficiency.
[0038] Please see Figure 1 , Figure 1 A schematic diagram of an application scenario involving an embodiment of this application is provided, including a terminal 10 and a server 20, wherein the terminal 10 and the server 20 are connected via a wired or wireless network.
[0039] Users can construct format conversion statements for converting target data through terminal 10, and then send the format conversion statements to server 20. After receiving the format conversion statements, server 20 processes the target data by executing the method provided in this application.
[0040] For example, the process of server 20 executing the method provided in this application includes: obtaining a format conversion statement for converting the target data in offline storage; the format conversion statement is used to convert the target data from its original data format to the target data format required by the online service; parsing the format conversion statement to obtain an intermediate format conversion statement for execution by the execution engine, and a target serialization protocol suitable for converting the target data between the target data format and the serialization format; the execution engine loads and executes the intermediate format conversion statement to convert the target data obtained from offline storage to obtain first intermediate data in the target data format; the execution engine serializes the first intermediate data according to the target serialization protocol to obtain target serialized data; the target serialized data is stored online; and then, after the online service obtains the target serialized data from the online storage, it parses the target serialized data into the first intermediate data according to the target serialization protocol and uses the first intermediate data.
[0041] Terminal 10 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, wearable device (such as a smartwatch), virtual reality device, in-vehicle terminal, smart TV, etc., but is not limited to these.
[0042] Server 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0043] The present application will now be described in detail with reference to the embodiments.
[0044] Please see Figure 2 , Figure 2 A flowchart illustrating the data processing method provided in this application embodiment is given. The data processing method includes steps S110-S160:
[0045] S110. Obtain a format conversion statement for converting the target data in offline storage; the format conversion statement is used to convert the target data from its original data format to the target data format required by the online service.
[0046] The format conversion statements are pre-written by programmers for the target data. The original data format refers to the data format of the target data in offline storage. The syntax rules followed by the format conversion statements can be the syntax rules of SQL (Structured Query Language); because the syntax rules of SQL statements are simple in design and have a low learning threshold, they are more conducive to data processing personnel to construct format conversion statements.
[0047] Understandably, different target data have different original data formats, so it is necessary to construct corresponding format conversion statements for different target data.
[0048] The target data stored in offline storage can be data collected from various data sources, or data obtained after preprocessing (such as data cleaning, data statistics, etc.) the data collected from various data sources. There are no specific restrictions here. For example, in business scenarios such as search, recommendation, and advertising, it is necessary to collect business-related data for offline processing. The processed data needs to be written to offline storage. For example, feature data extracted from video analysis can be used for video recommendation. The data in offline storage needs to be format converted and serialized before being imported into online storage and then used by online services.
[0049] S120. Parse the format conversion statements to obtain intermediate format conversion statements for the execution engine to execute, and a target serialization protocol suitable for converting target data between target data format and serialization format.
[0050] As mentioned above, certain grammatical rules must be followed during the construction of format conversion statements. Therefore, when parsing format conversion statements, it is also necessary to parse them according to the grammatical rules followed by the format conversion statements in order to obtain the correct parsing results.
[0051] The execution engine refers to the component used to load and execute machine instructions. It should be noted that format conversion statements are a language geared towards data processing personnel. That is to say, format conversion statements cannot be directly executed by the execution engine to achieve data processing.
[0052] Therefore, in some implementations, the format conversion statement can be converted to obtain the abstract syntax tree corresponding to the format conversion statement. Then, by parsing the abstract syntax tree, intermediate format conversion statements and the target serialization protocol can be obtained. The intermediate format conversion statement is a code statement that can be parsed and executed by the execution engine.
[0053] Furthermore, for intermediate format conversion statements, the abstract syntax tree (API) can be parsed to determine the data conversion logic defined within it. Then, based on the determined data conversion logic, corresponding executable code is generated to obtain the intermediate format conversion statements. For the target serialization protocol, the API can be parsed to determine the objects to be converted (e.g., fields) involved in the data conversion logic within the API. Then, for each object to be converted, a corresponding serialization protocol is generated according to the serialization protocol generation rules. Finally, based on the serialization protocols of each object to be converted, the target serialization protocol is obtained.
[0054] S130. The execution engine loads and executes intermediate format conversion statements to convert the target data obtained from offline storage into the target data format, resulting in first intermediate data in the target data format.
[0055] The first intermediate data is the data in the format required by the online service. However, the use of the first intermediate data by the online service and the first intermediate data obtained through data processing are not synchronized. Therefore, the first intermediate data obtained through processing often needs to be stored online.
[0056] S140. The execution engine serializes the first intermediate data according to the target serialization protocol to obtain the target serialized data.
[0057] It should be noted that the target data format required by the online service may not be the same for different target data. If the first intermediate data is directly stored in the online storage, the online storage will need to store and transmit the first intermediate data with different target data formats, which will reduce the transmission efficiency of the online storage. Therefore, by serializing the first intermediate data and storing the serialized data in the online storage, it is beneficial to improve the storage efficiency and transmission efficiency of the online storage, thereby improving the response efficiency of the online service.
[0058] S150. Store the target serialized data online.
[0059] S160. After obtaining the target serialized data from the online storage, the online service parses the target serialized data into first intermediate data according to the target serialization protocol and uses the first intermediate data.
[0060] In some implementations, the online service may retrieve the target serialized data from online storage in response to a service request; or the online service may periodically retrieve the target serialized data from online storage, and the specific method is not limited here.
[0061] In some implementations, after generating the target serialization protocol, the target serialization protocol can also be synchronized with the online service, so that the online service can parse the target serialized data into first intermediate data according to the target serialization protocol.
[0062] In other implementations, after generating the target serialization protocol, the target serialization protocol and the target serialization data can be associated and stored online. This allows the online service to obtain the target serialization protocol associated with the target serialization data when retrieving the target serialization data from the online storage, thus facilitating the online service to parse the target serialization data into first intermediate data according to the target serialization protocol.
[0063] In the above implementation, based on obtaining the format conversion statement used to convert the target data in offline storage, the format conversion statement is parsed to obtain intermediate format conversion statements for the execution engine to execute, and a target serialization protocol suitable for converting the target data between the target data format and the serialization format. Then, the execution engine loads and executes the intermediate format conversion statement to convert the target data obtained from offline storage to obtain first intermediate data in the target data format. The execution engine then serializes the first intermediate data according to the target serialization protocol to obtain the target serialized data. The target serialized data is then stored online. After obtaining the target serialized data from online storage, the online service parses the target serialized data into the first intermediate data according to the target serialization protocol and uses the first intermediate data. In the above data processing, the intermediate format conversion statement and the target serialization protocol are obtained by parsing the format conversion statement, eliminating the need to write corresponding serialization protocol code separately for the target data. Simultaneously, the online service uses the same target serialization protocol to parse the target serialized data, avoiding data parsing errors caused by inconsistencies between the online service and online storage protocols.
[0064] In some implementations, please refer to Figure 3 , Figure 3 Exemplary embodiments provided in this application are given. Figure 2 The flowchart of step S120 shows that the format conversion statement includes a query statement, which includes a field definition statement for at least one original field in the target data. The field definition statement for an original field defines the field type, attributes, and new fields mapped to the original field under the target data format. The target serialization protocol includes serialization mapping statements for the new fields mapped to each original field. The serialization mapping statement for a new field defines the encoding of the new field under the serialization format. Step S120 includes:
[0065] S210. Perform abstract syntax tree transformation on the format conversion statement to obtain the target abstract syntax tree.
[0066] Abstract syntax tree (API) transformation can be achieved through a parsing engine, which can be based on parsing code (such as Python code).
[0067] It is understandable that the format conversion statement is transformed into an abstract syntax tree to obtain the target abstract syntax tree. Since the target abstract syntax tree is a data form that can be recognized and edited by machines, development tools can be used to perform a series of processes on the target abstract syntax tree, such as code parsing, code generation, and information addition.
[0068] S220. Generate intermediate format conversion statements for the execution engine to execute based on the target abstract syntax tree.
[0069] Specifically, the target abstract syntax tree can be traversed to determine the data transformation logic defined in the target abstract syntax tree. Then, based on the determined data transformation logic, code is generated to obtain intermediate format transformation statements for the execution engine to execute. The code generation process includes collecting code fragments from each node in the target abstract syntax tree, and combining the code fragments from each node according to the data transformation logic to obtain complete intermediate format transformation statements.
[0070] S230. Based on the attributes defined for each original field in the query statement expressed in the target abstract syntax tree, the new fields mapped to each original field under the target data format, the field types of each new field, and the sequence codes assigned to the new fields under the serialization format, generate the serialization mapping statement for the new fields mapped to each original field.
[0071] A complete field typically includes a field name, field type, and field attributes. Field types include, for example, floating-point numbers (Float, Double), unsigned 64-bit integer data types (unit64), and unsigned 32-bit integer data types (unit32). Field attributes are defined for the data corresponding to the field and are used during data processing. Examples of data attributes include the repeated (rpt) attribute, used to define data that can be selected repeatedly; the optional (opt) attribute, used to define data that is not required; and the required (req) attribute, used to define data that must be selected.
[0072] The original field refers to the field involved in the target data in offline storage. The attributes defined for the original field are the attributes defined for the original field in the target data in offline storage according to the syntax rules followed by the original data format. The new field mapped to each original field under the target data format refers to the new field specified in the format conversion statement to map the original field to the target data format. In some embodiments, if some fields can be shared under the original data format and the target data format, it is a new field mapped to an original field under the target data format, or it can be the original field itself.
[0073] The field type of a new field refers to the field type of a new field defined according to the syntax rules of the target data format, based on the field type of the original field in the original data format. In other words, the field type of the new field is the field type provided by the syntax rules of the target data format, and this field type is actually the same as the field type of the original field in the original data format, only the form of expression may be different.
[0074] Because a sequence code is assigned to a new field under the serialization format, the corresponding new field is represented by the corresponding sequence code under the serialization format. It is understood that this sequence code follows the syntax rules of the serialization format.
[0075] It is understandable that by defining attributes for each original field, mapping each original field to a new field in the target data format, and defining the field type of each new field, the original data format can be converted to the target data format. For example, mapping an original field to a new field can convert the original field to a new field in the target format; defining the field type for a new field can convert the field type of the original field to the field type of the new field in the target format; and redefining attributes for the new field in the target format.
[0076] In the above implementation, by generating a serialization mapping statement that maps the original field to the new field, the mapping relationship between the sequence code and the field is determined, so that the serialization mapping statement can be used to serialize the first intermediate data to obtain the target serialized data.
[0077] In some implementations, please refer to Figure 4 , Figure 4 Exemplary embodiments provided in this application are given. Figure 3 The flowchart for step S230 shows the processing steps S310-S320 for each original field:
[0078] S310. Based on the attributes defined for the original field, determine the attribute key representing the attribute from the attribute keys provided by the serialization format.
[0079] The serialization format is supported by online storage. Furthermore, the serialization format syntax provides various attribute keywords and corresponding attributes for each keyword. Therefore, under a given serialization format, the corresponding attribute keywords can be determined based on the attributes defined for the original field. Commonly used serialization formats in this field include those under the Protobuf (Protocol Buffers) framework, the Thrift framework, and the Jackson framework.
[0080] It should be noted that when processing different original fields of the target data, all fields need to share the same serialization format.
[0081] Taking the serialization format under the Protobuf framework as an example, the attribute keywords provided by this serialization format include optional, repeated, and required. Based on this, if the attribute defined for the original field is not required, the corresponding attribute keyword is optional; if the attribute defined for the original field is required, the corresponding attribute keyword is required; and if the attribute defined for the original field is a repeated selection attribute, the corresponding attribute keyword is repeated.
[0082] S320. Combine the attribute keywords, the new field mapped to the original field in the target data format, the field type of the new field, and the sequence code assigned to the new field in the serialization format to obtain the serialization mapping statement of the new field mapped to the original field.
[0083] In some implementations, multiple original fields in a query statement can form an original composite field, meaning that there is a nested composite structure among the original fields in the query statement. When performing data conversion, the original composite field can be mapped to a new composite field under the target data format. Obviously, the field type of a composite field cannot be uniquely determined. In this case, the attribute keywords, the new composite field mapped to the original composite field under the target data format, the original composite field, and the sequence code assigned to the new composite field under the serialization format can be combined to obtain the serialization mapping statement of the new composite field mapped to the original composite field.
[0084] Furthermore, if there are nested complex structures among the original fields in the query statement, it means that different original fields may have different levels. In this case, to ensure that the nesting relationship between the original fields is preserved in the processed new fields, the original fields can be first divided into levels before processing. Then, the above steps are performed on each original field under each level to obtain the serialization mapping statement of the new fields mapped by the original fields. It should be noted that for different original fields belonging to the same level, the sequence codes in their corresponding serialization mapping statements are also different, thus ensuring the uniqueness of the sequence codes under that level; while for different original fields belonging to different levels, since there is already a difference in levels, the sequence codes in their corresponding serialization mapping statements can be the same.
[0085] To facilitate understanding, the following will combine... Figure 5 and Figure 6 To explain, Figure 5 An exemplary query statement provided in an embodiment of this application is given below. Figure 6 An example is given Figure 5 The serialization mapping statement corresponding to the query statement shown.
[0086] Please see Figure 5 ,Depend on Figure 5 The query statement shown illustrates that a complete query statement includes the original field, the attributes defined for the original field, the new field mapped to the original field in the target data format, and the field type of the new field; Figure 5 The query statement "opt(uin,"uint32")as uin_user_tag" is used as an example to illustrate this. The original field is "uin", the attribute defined for the original field is "opt", and the new field mapped by the original field under the target data format is "uin_user_tag". The field type of the new field is "uint32". Through this query statement, the "uin" field can be converted to the unit32 type and redefined as the "uin_user_tag" field.
[0087] In some implementations, the query statement may also specify only the original field, which in turn represents a direct selection of the original field; Figure 5 Taking the query statement "tag_group" as an example, it only indicates that the original field is "tag_group". It means that the new field mapped by the original field "tag_group" in the target data format is also "tag_group". The field type of the new field is consistent with the field type of the original field, and the attributes of the new field are consistent with the attributes of the original field.
[0088] In other implementations, the query statement may include a composite field, which is a field composed of multiple composite fields. Different composite fields can form a nested relationship, meaning that one composite field can include another composite field. This will be discussed in conjunction with... Figure 5 For more detailed information, please refer to the following: Figure 5 , Figure 5 The query statement shown contains two original composite fields: "_UserTagGroupsWithfeed_tag" and "_UserTagGroupsWithfeed". The original composite field "_UserTagGroupsWithfeed_tag" is nested from four original fields: "feedid", "timestamp", "1.0", and an empty field. The original composite field "_UserTagGroupsWithfeed" is nested from three original fields: "group_rank", "groupscore", and "tag_group" and one original composite field "_UserTagGroupsWithfeed_tag".
[0089] Based on this, the fields in the query statement can be divided into three levels, with the first level including uin,
[0090] The first level includes two fields: _UserTagGroupsWithfeed; the second level includes four fields: _UserTagGroupsWithfeed_tag, group_rank, groupscore, and tag_group; the third level includes four fields: "feedid, timestamp, 1.0, and an empty field".
[0091] For the four fields in the third level, all of which are non-composite fields, the attribute keywords, the new fields mapped to the original fields in the target data format, the field type of the new fields, and the sequence codes assigned to the new fields in the serialization format can be combined to obtain the serialization mapping statement for the new fields mapped to the original fields. The resulting serialization mapping statement is as follows: Figure 6 As shown in Figure a.
[0092] Let's take one query statement as an example. The query statement is "opt(feedid,'uint64')as feed_id". The original field it refers to is feedid, the attribute it defines is opt, the attribute keyword corresponding to opt is optional, and the new field mapped under the target data format is feed_id. The field type of the new field is unit64. If the sequence code assigned to it is 1, then we can get the serialization mapping statement "optional uint64feed_id=1" which maps feedid to feed_id.
[0093] For the four fields in the second level, the resulting serialization mapping statement is as follows: Figure 6 As shown in b.
[0094] Of the four fields in the second level, group_rank, groupscore, and tag_group are non-composite fields, while _UserTagGroupsWithfeed_tag is a composite field. The serialization mapping statements for the non-composite fields are as described above and will not be repeated here. For the composite field _UserTagGroupsWithfeed_tag, the corresponding query statement is "rpt(XXX,"_UserTagGroupsWithfeed_tag")as tag_list". It indicates the original composite field _UserTagGroupsWithfeed_tag, the defined attribute is rpt, and the attribute keyword corresponding to rpt is repeated. The new composite field mapped to the original composite field under the target data format is tag_lis, where XXX is the content included in the original composite field _UserTagGroupsWithfeed_tag, namely "feedid, timestamp, 1.0 and empty field". If the sequence code assigned to it is 4, then the serialization mapping statement "repeated_UserTagGroupsWithfeed_tag tag_list=4" can be obtained to map _UserTagGroupsWithfeed_tag to tag_lis.
[0095] Similarly, for the two fields at the first level, namely the non-composite field uin and the composite field _UserTagGroupsWithfeed, the corresponding serialization mapping statements are as follows: Figure 6 As shown in c, it will not be explained in detail here.
[0096] It should be noted that, Figure 5The query statements in the system are constructed according to the syntax rules of SQL statements, and it supports the select clause in the SQL syntax rules.
[0097] In the SELECT clause, the selected object can be the original field, such as... Figure 5 In this case, you can directly select the "tag_group" field. There's no need to convert the original field's format; the original data format will be preserved. Alternatively, you can declare the selected original field using an attribute definition function and then declare the new field using the `as` syntax, such as... Figure 5 The expression `opt(uin,"uint32") as uin_user_tag` means defining the field attribute of the original field `uin` as `opt`, converting `uin` to `uint32` type, and mapping it to the new field `uin_user_tag`. Of course, if no `as` is used to define a new field for the original field, the original field remains unchanged. Figure 5 The `opt(group_rank, 'double')` method converts the `group_rank` field to `double` while retaining its original type `group_rank`. Additionally, empty fields can be declared using attribute definition functions for protocol formatting, such as... Figure 5 In the function `opt('float')as sampletype`, the original field is empty, and a new field `sampletype` is defined with the field type `float`.
[0098] The attribute definition functions include `rpt`, used to define `repeated` attributes; `opt`, used to define `optional` attributes; and `req`, used to define `required` attributes. Furthermore, the attribute definition functions can include the following four usages:
[0099] Define an empty field for protocol format filling, such as Figure 5 In the context of `opt('float') as sampletype`;
[0100] Map the original field to the new field, such as Figure 5 in opt(uin,"uint32")as uin_user_tag;
[0101] Define a new field using a constant as the default value, such as Figure 5 in opt(1.0,'float')as score;
[0102] Define a new composite field, such as Figure 5 The definition of the tag_list field is rpt()as tag_list.
[0103] In addition, the select clause also supports the concat function (concatenation function) to concatenate multiple strings; it can accept multiple input parameters, which can be original fields or constants.
[0104] In some implementations, the format conversion statement further includes a string splitting definition statement; the string splitting definition statement describes the field type defined for at least one splitting element and the new element field mapped by each splitting element under the target data format; at least one splitting element is selected from the splitting results obtained by splitting the original string field, and the original string field exists in the target data; the target serialization protocol also includes an element serialization mapping statement corresponding to the string splitting definition statement; step S120 may further include step S400:
[0105] Based on the statement type of the string splitting definition statement expressed in the target abstract syntax tree, the field type defined for each splitting element in the string splitting definition statement, the new element field mapped by each splitting element under the target data format, and the sequence encoding assigned to the new element field under the serialization format, generate the element serialization mapping statement corresponding to the string splitting definition statement.
[0106] The string splitting definition statement typically includes the splitting function, the field to be split, the delimiter, the selected splitting element, the new element field mapped by the splitting element under the target data format, and the field type of the new element field.
[0107] For example, a complete string splitting definition statement can be "split(f1,"#",0,"uint32")as uin", which means splitting the f1 field using "#", taking the 0th splitting element as the uin field, and the field type of the uin field is uint32; where split is the splitting function, f1 is the field to be split, # is the splitting symbol, 0 represents selecting the 0th splitting element, and as uin means determining the selected 0th splitting element as the uin field.
[0108] Furthermore, the statement type of a string splitting definition statement can be determined based on the number of delimiters, the number of selected splitting elements, and the number of new element fields. Therefore, the statement type of a string splitting definition statement can include the following four types:
[0109] The first type refers to performing a single-level segmentation (i.e., the number of segments is 1) and selecting a segmentation element from the segmentation results (i.e., the number of segmentation elements and the number of new element fields are both 1).
[0110] The second type refers to performing a single-level segmentation (i.e., the number of segments is 1), and selecting multiple segmentation elements from the segmentation results (i.e., the number of segmentation elements is greater than 1), and all multiple segmentation elements are mapped to the first reference field under the target data format (i.e., the number of new element fields is 1).
[0111] The third type refers to performing a single-level segmentation (i.e., the number of segments is 1), selecting multiple segmentation elements from the segmentation results (i.e., the number of segmentation elements is greater than 1), and mapping multiple segmentation elements to at least two second reference fields under the target data format (i.e., the number of new element fields is greater than 1).
[0112] The fourth type refers to performing two-level segmentation (i.e., the number of segments is 2), selecting multiple segmentation elements from the segmentation results (i.e., the number of segmentation elements is greater than 1), and mapping multiple segmentation elements to at least two third reference fields under the target data format (i.e., the number of new element fields is greater than 1).
[0113] In some implementations, the string splitting definition statement includes a first string splitting definition statement of the first type, which is used to perform single-level splitting and select a splitting element from the splitting result; based on this, step S400 specifically includes:
[0114] If the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the first type, the first attribute keyword, the field type defined for the splitting element, the new element field mapped by the splitting element under the target data format, and the sequence code assigned to the new element field under the serialization format are combined to obtain the serialization mapping statement of the new element field mapped by the field type of the splitting element; wherein, the first attribute keyword is the attribute keyword that indicates a non-mandatory attribute.
[0115] The first string splitting definition statement can be summarized in the following format:
[0116] The split function is defined as follows: (field to be split, delimiter, split element, field type of new element field) new element field.
[0117] For ease of understanding, the following example defines the first string splitting statement:
[0118] split(f1,"#",0,"uint32")as uin;
[0119] Its representative uses the split function to split the field f1 to be split using "#", takes the 0th split element as a uin field, and converts it to uint32 type.
[0120] Based on the first string splitting definition statement above, the following serialization mapping statement can be obtained:
[0121] optional uint32 uin = 1; where optional is the first attribute keyword, indicating that the new element field uin is a field that is not required.
[0122] In some implementations, the string splitting definition statement includes a second string splitting definition statement belonging to the second type. This second string splitting definition statement is used for single-level splitting and selects multiple splitting elements from the splitting results. All multiple splitting elements are mapped to a first reference field under the target data format. Based on this, step S400 specifically includes:
[0123] If the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the second type, the second attribute keyword, the field type defined for multiple splitting elements, the first reference field mapped by multiple splitting elements under the target data format, and the sequence code assigned to the first reference field under the serialization format are combined to obtain the serialization mapping statement of the first reference field mapped by the field type of the splitting element; wherein, the second attribute keyword is the attribute keyword indicating repeated selection of the attribute.
[0124] The second string splitting definition statement can be summarized in the following format:
[0125] The split function is defined as follows: (field to be split, delimiter, field type of new element field) new element field.
[0126] For ease of understanding, the following example defines the second string splitting statement:
[0127] split(f1,"#","uint32")as uin;
[0128] Its representative uses the split function to split the field f1 to be split using "#", maps all the split elements to uin fields, and converts them to uint32 type.
[0129] Based on the second string splitting definition statement above, the following serialization mapping statement can be obtained:
[0130] repeated uint32 uin = 1; where optional is the second attribute keyword, indicating that the new element field uin is a field with the repeated selection attribute.
[0131] In some implementations, the string splitting definition statement includes a third string splitting definition statement belonging to the third type. The third string splitting definition statement is used to perform single-level splitting and select multiple splitting elements from the splitting results. The multiple splitting elements are mapped to at least two second reference fields under the target data format. Step S400 specifically includes:
[0132] Step 1: If the string splitting definition statement expressed in the target abstract syntax tree is of type 3, combine the first attribute keyword, the first composite field name defined for at least two second reference fields, the first composite field defined for at least two second reference fields, and the sequence code assigned to the first composite field in the serialization format to obtain the serialization mapping statement of the first composite field.
[0133] Step 2: For each segmentation element, combine the first attribute keyword, the field type defined for the segmentation element, the second reference field mapped to the segmentation element in the target data format, and the sequence code assigned to the second reference field in the serialization format to obtain the serialization mapping statement of the second reference field mapped to the field type of the segmentation element; wherein, the first attribute keyword is an attribute keyword that indicates a non-mandatory attribute.
[0134] The third string splitting definition statement can be summarized in the following format:
[0135] The splitting function (field to be split, delimiter, first composite field name) first composite field; where the first composite field name is used to refer to the data transformation clause of each splitting element. The data transformation clause of a splitting element includes the first attribute keyword, the field type defined for the splitting element, the second reference field mapped by the splitting element under the target data format, and the sequence code assigned to the second reference field under the serialization format.
[0136] For ease of understanding, the following example defines a third string splitting statement:
[0137]
[0138] Based on the above third string splitting definition statement, the serialization mapping statement for the first composite field can be obtained as follows;
[0139] optional Msg1 level1_split=1;
[0140] Here, optional is the first attribute keyword, representing a non-required attribute; Msg1 represents the name of the first composite field; and level1_split represents the first composite field.
[0141] And, obtain the serialization mapping statement for the second reference field mapped to the field type of the delimiter element:
[0142] message Msg1{
[0143] optional uint32 uin = 1;
[0144] optional string desc = 2;
[0145] }
[0146] In this context, message Msg1 is the serialization protocol name, optional is the first attribute keyword, representing a non-required attribute, uint32 uin=1 means that the field type of uin is uint32 and its sequence encoding is 1. Similarly, stringdesc=2 means that the field type of desc is string and its sequence encoding is 2.
[0147] In some implementations, the string splitting definition statement includes a fourth string splitting definition statement belonging to the fourth type. This fourth string splitting definition statement is used for two-level splitting and selects multiple splitting elements from the splitting results. These multiple splitting elements are mapped to at least two third reference fields under the target data format. Based on this, step S400 specifically includes:
[0148] Step 3: If the string splitting definition statement expressed in the target abstract syntax tree is of type 4, combine the second attribute keyword, the name of the second composite field defined for at least two third reference fields, the second composite field defined for at least two third reference fields, and the sequence code assigned to the second composite field in the serialization format to obtain the serialization mapping statement of the second composite field.
[0149] Step 4: For each segmentation element, combine the first attribute keyword, the field type defined for the segmentation element, the third reference field mapped to the segmentation element in the target data format, and the sequence code assigned to the third reference field in the serialization format to obtain the serialization mapping statement of the third reference field mapped to the field type of the segmentation element; wherein, the first attribute keyword is the attribute keyword indicating a non-mandatory attribute; the second attribute keyword is the attribute keyword indicating a repeated selection attribute.
[0150] The fourth string splitting definition statement, along with the third type, can be summarized in the following format:
[0151] The splitting function is defined as follows: (field to be split, first delimiter, second delimiter, second composite field name) second composite field. The second composite field name refers to the data transformation clause of each splitting element. A data transformation clause for a splitting element includes a first attribute keyword, a field type defined for the splitting element, a third reference field mapped to the splitting element in the target data format, and a sequence code assigned to the third reference field in the serialization format.
[0152] For ease of understanding, the following example provides a fourth string splitting definition statement:
[0153]
[0154] Based on the fourth string splitting definition statement above, the serialization mapping statement for the second composite field can be obtained as follows;
[0155] repeated Msg1 level1_split=1;
[0156] In this context, "repeated" is the second attribute keyword, representing repeated selection of the attribute; "Msg1" represents the name of the second composite field; and "level1_split" represents the second composite field.
[0157] And, the serialization mapping statement for obtaining the third reference field mapped to the field type of the delimiter element:
[0158] message Msg1{
[0159] optional uint32 uin = 1;
[0160] optional string desc = 2;
[0161] }
[0162] In this context, message Msg1 is the serialization protocol name, optional is the first attribute keyword, representing a non-required attribute, uint32 uin=1 means that the field type of uin is uint32 and its sequence encoding is 1. Similarly, stringdesc=2 means that the field type of desc is string and its sequence encoding is 2.
[0163] In some implementations, the format conversion statement also includes a grouping statement (such as the group by statement in SQL syntax rules), which may include multi-level grouping fields to group the first intermediate data according to the grouping fields.
[0164] For example, a grouping statement can be:
[0165] group by
[0166] uin
[0167] taggroup_list.tag_group
[0168] Here, uin is the first-level grouping field, representing first-level grouping based on the uin field; taggroup_list.tag_group is the second-level grouping field, representing second-level grouping based on the taggroup_list.tag_group field after the first-level grouping based on the uin field.
[0169] It should be noted that since grouping statements only involve grouping data and do not involve the transformation of field types, there is no need to generate corresponding serialization mapping statements for grouping statements.
[0170] Please see Figure 7 , Figure 7 An application architecture diagram of this application is given as an example. Figure 7 In the application architecture shown, the format conversion statements are constructed according to the syntax rules of SQL statements.
[0171] The application architecture includes offline storage, a syntax parsing engine, an execution engine, online storage, and online services. These will be described in detail below.
[0172] Offline storage involves processing the initial data using data processing statements (such as SQL statements) after acquisition to obtain the target data, and then writing the target data to offline storage. It's important to note that data processing modifies the initial data itself, such as data completion and filtering, without altering the data format. Therefore, the original data format contained in the target data is necessarily also present in the initial data.
[0173] The syntax parsing engine is used to parse format conversion statements to obtain intermediate format conversion statements and target serialization protocols for the execution engine to execute. Before this, it is necessary to construct format conversion statements. Since format conversion statements only involve the conversion of data formats, and the original data format contained in the target data also exists in the initial data, the format conversion statements can be constructed based on the initial data after obtaining the initial data, or they can be constructed based on the target data in offline storage.
[0174] In some implementations, the obtained target serialization protocol can also be synchronized to the online service so that the online service can use the target serialization protocol to deserialize the target serialized data to obtain the first intermediate data.
[0175] The execution engine is used to retrieve target data from offline storage, load and execute intermediate format conversion statements to convert the target data retrieved from offline storage into first intermediate data in the target data format; then, according to the target serialization protocol, the first intermediate data is serialized to obtain target serialized data, and the target serialized data is written to online storage.
[0176] Online storage is used to store target serialized data.
[0177] Online services are used to retrieve target serialized data from online storage, parse the target serialized data into first intermediate data according to the target serialization protocol, and then use the first intermediate data.
[0178] In the above application architecture, since format conversion statements (such as those constructed using SQL syntax rules) are a language geared towards data processing personnel, they have a low learning curve, are easy to use, and can directly generate target serialization protocols through the parsing engine. Compared to existing technologies where data processing personnel manually write target serialization protocols, data processing efficiency is greatly improved. Furthermore, in the above application architecture, the generated target serialization protocol is synchronized to the execution engine and online services, enabling the execution engine and online services to use the same target serialization protocol. This facilitates the management of target serialization protocols and avoids data errors caused by protocol inconsistencies. Compared to the existing technology's data conversion and online service multi-party negotiation and alignment model, the process is more secure.
[0179] In some implementations, please refer to Figure 8 , Figure 8 A schematic diagram of a data processing apparatus according to this embodiment is provided. The data processing apparatus 600 includes:
[0180] The acquisition module 610 is used to acquire format conversion statements for converting the target data in offline storage; the format conversion statements are used to convert the target data from the original data format to the target data format required by the online service.
[0181] Parsing module 620 is used to parse format conversion statements to obtain intermediate format conversion statements for execution engine to execute, and target serialization protocol suitable for converting target data between target data format and serialization format.
[0182] The first execution module 630 is used by the execution engine to load and execute intermediate format conversion statements to convert the target data obtained from offline storage into first intermediate data in the target data format.
[0183] The second execution module 640 is used by the execution engine to serialize the first intermediate data according to the target serialization protocol to obtain the target serialized data.
[0184] Storage module 650 is used to store target serialized data online.
[0185] Service module 660 is used by the online service to parse the target serialized data into first intermediate data according to the target serialization protocol after obtaining the target serialized data from the online storage, and then use the first intermediate data.
[0186] In some implementations, the format conversion statement includes a query statement, which includes a field definition statement for at least one original field in the target data. The field definition statement for an original field defines the field type, attributes, and new fields mapped to the original field in the target data format. The target serialization protocol includes serialization mapping statements for the new fields mapped to each original field. A serialization mapping statement for a new field defines the encoding of the new field in the serialization format. The parsing module 620 includes: a conversion unit for performing an abstract syntax tree conversion on the format conversion statement to obtain a target abstract syntax tree; a code generation unit for generating intermediate format conversion statements for execution by the execution engine based on the target abstract syntax tree; and a query protocol generation unit for generating serialization mapping statements for the new fields mapped to each original field based on the attributes defined for each original field in the query statement expressed in the target abstract syntax tree, the new fields mapped to each original field in the target data format, the field types of each new field, and the sequence encoding assigned to the new field in the serialization format.
[0187] In some implementations, the query protocol generation unit processes each original field as follows: based on the attributes defined for the original field, it determines the attribute keywords representing the attributes from the attribute keywords provided by the serialization format; it combines the attribute keywords, the new field mapped to the original field in the target data format, the field type of the new field, and the sequence code assigned to the new field in the serialization format to obtain the serialization mapping statement of the new field mapped to the original field.
[0188] In some implementations, the format conversion statement further includes a string splitting definition statement; the string splitting definition statement describes the field type defined for at least one splitting element and the new element field mapped by each splitting element under the target data format; at least one splitting element is selected from the splitting results obtained by splitting the original string field, and the original string field exists in the target data; the target serialization protocol also includes an element serialization mapping statement corresponding to the string splitting definition statement; the parsing module 620 further includes a splitting protocol generation unit, used to generate an element serialization mapping statement corresponding to the string splitting definition statement based on the statement type of the string splitting definition statement expressed in the target abstract syntax tree, the field type defined for each splitting element in the string splitting definition statement, the new element field mapped by each splitting element under the target data format, and the sequence encoding assigned to the new element field under the serialization format.
[0189] In some implementations, the string splitting definition statement includes a first string splitting definition statement of the first type, which is used to perform single-level splitting and select a splitting element from the splitting results. The splitting protocol generation unit is specifically used to combine the first attribute keyword, the field type defined for the splitting element, the new element field mapped by the splitting element under the target data format, and the sequence code assigned to the new element field under the serialization format, if the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the first type, to obtain the serialization mapping statement of the new element field mapped by the field type of the splitting element; wherein, the first attribute keyword is an attribute keyword indicating a non-mandatory attribute.
[0190] In some implementations, the string splitting definition statement includes a second string splitting definition statement of the second type. The second string splitting definition statement is used to perform single-level splitting and select multiple splitting elements from the splitting results. The multiple splitting elements are all mapped to a first reference field under the target data format. The splitting protocol generation unit is specifically used to combine the second attribute keyword, the field type defined for the multiple splitting elements, the first reference field mapped by the multiple splitting elements under the target data format, and the sequence code assigned to the first reference field under the serialization format if the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the second type, to obtain a serialization mapping statement of the first reference field mapped by the field type of the splitting element. The second attribute keyword is an attribute keyword that indicates repeated selection of attributes.
[0191] In some implementations, the string splitting definition statement includes a third string splitting definition statement belonging to the third type. The third string splitting definition statement is used to perform single-level splitting and select multiple splitting elements from the splitting results. The multiple splitting elements are mapped to at least two second reference fields under the target data format. Specifically, if the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the third type, the first attribute keyword, the first composite field name defined for at least two second reference fields, the first composite field defined for at least two second reference fields, and the sequence code assigned to the first composite field under the serialization format are combined to obtain the serialization mapping statement of the first composite field. For each splitting element, the first attribute keyword, the field type defined for the splitting element, the second reference field mapped by the splitting element under the target data format, and the sequence code assigned to the second reference field under the serialization format are combined to obtain the serialization mapping statement of the second reference field mapped by the field type of the splitting element. Wherein, the first attribute keyword is an attribute keyword indicating a non-mandatory attribute.
[0192] In some implementations, the string splitting definition statement includes a fourth string splitting definition statement belonging to the fourth type. The fourth string splitting definition statement is used to perform two-level splitting and select multiple splitting elements from the splitting results. The multiple splitting elements are mapped to at least two third reference fields under the target data format. The splitting protocol generation unit is specifically used to combine the second attribute keyword, the second composite field name defined for at least two third reference fields, the second composite field defined for at least two third reference fields, and the sequence code assigned to the second composite field under the serialization format if the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the fourth type, to obtain the serialization mapping statement of the second composite field. For each splitting element, the first attribute keyword, the field type defined for the splitting element, the third reference field mapped by the splitting element under the target data format, and the sequence code assigned to the third reference field under the serialization format are combined to obtain the serialization mapping statement of the third reference field mapped by the field type of the splitting element. Wherein, the first attribute keyword is an attribute keyword indicating a non-mandatory attribute; the second attribute keyword is an attribute keyword indicating a repeated selection attribute.
[0193] Figure 9 A schematic diagram of a computer system suitable for implementing the electronic device of the embodiments of this application is shown. The electronic device can be the terminal described above, used to implement the data processing method provided in this application. It should be noted that... Figure 9 The computer system 1300 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0194] like Figure 9 As shown, the computer system 1300 includes a Central Processing Unit (CPU) 1301, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in Read-Only Memory (ROM) 1302 or programs loaded from storage portion 1308 into Random Access Memory (RAM) 1303. The RAM 1303 also stores various programs and data required for system operation. The CPU 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An Input / Output (I / O) interface 1305 is also connected to the bus 1304.
[0195] The following components are connected to I / O interface 1305: an input section 1306 including a keyboard, mouse, microphone, etc.; an output section 1307 including a cathode ray tube (CRT), liquid crystal display (LCD), and speakers, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card and a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to I / O interface 1305 as needed. Removable media 1311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1310 as needed so that computer instructions read from them can be loaded into storage section 1308 as needed.
[0196] In particular, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising computer instructions. When these computer instructions are executed by the central processing unit (CPU) 1301, various functions defined in the system of this application are performed.
[0197] This application also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the methods described in any of the above method embodiments.
[0198] It should be noted that the computer-readable storage medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0199] In the embodiments of this application, the terms "module" or "unit" refer to computer instructions or a portion of computer instructions that have a predetermined function and work together with other related parts to achieve a predetermined goal. These instructions can be implemented, wholly or partially, using software, hardware (e.g., processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that functions as a whole.
[0200] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.
Claims
1. A data processing method, characterized in that, include: Obtain a format conversion statement for converting the target data in offline storage; the format conversion statement is used to convert the target data from its original data format to the target data format required by the online service; The format conversion statement is parsed to obtain intermediate format conversion statements for the execution engine to execute, and a target serialization protocol suitable for converting the target data between the target data format and the serialization format; The execution engine loads and executes the intermediate format conversion statement to convert the target data obtained from offline storage into a first intermediate data in the target data format. The execution engine serializes the first intermediate data according to the target serialization protocol to obtain the target serialized data; The target serialized data is stored online; After the online service obtains the target serialized data from the online storage, it parses the target serialized data into the first intermediate data according to the target serialization protocol and uses the first intermediate data.
2. The method according to claim 1, characterized in that, The format conversion statement includes a query statement, which includes a field definition statement for at least one original field in the target data. The field definition statement for the original field defines the field type of the corresponding original field, the attribute of the original field, and the new field mapped by the original field under the target data format. The target serialization protocol includes serialization mapping statements for the new fields mapped to each of the original fields; a serialization mapping statement for a new field defines the encoding of the new field under the serialization format; The process of parsing the format conversion statement to obtain intermediate format conversion statements for the execution engine to execute, and a target serialization protocol suitable for converting the target data between the target data format and the serialization format, includes: The format conversion statement is transformed into an abstract syntax tree to obtain the target abstract syntax tree; Based on the target abstract syntax tree, intermediate format conversion statements are generated for the execution engine to execute; Based on the attributes defined for each original field in the query statement expressed in the target abstract syntax tree, the new fields mapped to each original field under the target data format, the field type of each new field, and the sequence code assigned to each new field under the serialization format, a serialization mapping statement for the new fields mapped to each original field is generated.
3. The method according to claim 2, characterized in that, The step of generating a serialization mapping statement for the new fields mapped to the original fields based on the attributes defined for each original field in the query statement expressed in the target abstract syntax tree, the new fields mapped to each original field under the target data format, the field types of each new field, and the sequence codes assigned to the new fields under the serialization format, includes: The following processing is performed on each original field: Based on the attributes defined for the original field, the attribute keywords representing the attributes are determined from the attribute keywords provided by the serialization format; The attribute keyword, the new field mapped to the original field under the target data format, the field type of the new field, and the sequence code assigned to the new field under the serialization format are combined to obtain the serialization mapping statement of the new field mapped to the original field.
4. The method according to claim 2 or 3, characterized in that, The format conversion statement also includes a string splitting definition statement; the string splitting definition statement describes the field type defined for at least one splitting element and the new element field mapped by each splitting element under the target data format; The at least one segmentation element is selected from the segmentation results obtained by segmenting the original string field, and the original string field exists in the target data; The target serialization protocol also includes element serialization mapping statements corresponding to string splitting definition statements; The step of parsing the format conversion statement to obtain intermediate format conversion statements for the execution engine to execute, and the target serialization protocol suitable for converting the target data between the target data format and the serialization format, further includes: Based on the statement type of the string segmentation definition statement expressed in the target abstract syntax tree, the field type defined for each segmentation element in the string segmentation definition statement, the new element field mapped by each segmentation element under the target data format, and the sequence encoding assigned to the new element field under the serialization format, an element serialization mapping statement corresponding to the string segmentation definition statement is generated.
5. The method according to claim 4, characterized in that, The string splitting definition statement includes a first string splitting definition statement belonging to the first type. The first string splitting definition statement is used to perform single-level splitting and select a splitting element from the splitting result. The step of generating an element serialization mapping statement corresponding to the string segmentation definition statement based on the statement type of the string segmentation definition statement expressed in the target abstract syntax tree, the field type defined for each segmentation element in the string segmentation definition statement, the new element field mapped by each segmentation element under the target data format, and the sequence encoding assigned to the new element field under the serialization format includes: If the statement type of the string segmentation definition statement expressed in the target abstract syntax tree is the first type, the first attribute keyword, the field type defined for the segmentation element, the new element field mapped by the segmentation element under the target data format, and the sequence code assigned to the new element field under the serialization format are combined to obtain the serialization mapping statement of the new element field mapped by the field type of the segmentation element. The first attribute keyword is an attribute keyword that indicates a non-mandatory attribute.
6. The method according to claim 4, characterized in that, The string splitting definition statement includes a second string splitting definition statement belonging to the second type. The second string splitting definition statement is used to perform single-level splitting and select multiple splitting elements from the splitting results. The multiple splitting elements are all mapped to the first reference field under the target data format. The step of generating an element serialization mapping statement corresponding to the string segmentation definition statement based on the statement type of the string segmentation definition statement expressed in the target abstract syntax tree, the field type defined for each segmentation element in the string segmentation definition statement, the new element field mapped by each segmentation element under the target data format, and the sequence encoding assigned to the new element field under the serialization format includes: If the statement type of the string segmentation definition statement expressed in the target abstract syntax tree is the second type, the second attribute keyword, the field type defined for multiple segmentation elements, the first reference field mapped by multiple segmentation elements under the target data format, and the sequence code assigned to the first reference field under the serialization format are combined to obtain the serialization mapping statement of the first reference field mapped by the field type of the segmentation element. The second attribute keyword is an attribute keyword that indicates the repeated selection of the attribute.
7. The method according to claim 4, characterized in that, The string splitting definition statement includes a third string splitting definition statement belonging to the third type. The third string splitting definition statement is used to perform single-level splitting and select multiple splitting elements from the splitting results. The multiple splitting elements are mapped to at least two second reference fields under the target data format. The step of generating an element serialization mapping statement corresponding to the string segmentation definition statement based on the statement type of the string segmentation definition statement expressed in the target abstract syntax tree, the field type defined for each segmentation element in the string segmentation definition statement, the new element field mapped by each segmentation element under the target data format, and the sequence encoding assigned to the new element field under the serialization format includes: If the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the third type, the first attribute keyword, the first composite field name defined for the at least two second reference fields, the first composite field defined for the at least two second reference fields, and the sequence code assigned to the first composite field under the serialization format are combined to obtain the serialization mapping statement of the first composite field. For each segmentation element, the first attribute keyword, the field type defined for the segmentation element, the second reference field mapped by the segmentation element under the target data format, and the sequence code assigned to the second reference field under the serialization format are combined to obtain the serialization mapping statement of the second reference field mapped by the field type of the segmentation element; The first attribute keyword is an attribute keyword that indicates a non-mandatory attribute.
8. The method according to claim 4, characterized in that, The string splitting definition statement includes a fourth string splitting definition statement belonging to the fourth type. The fourth string splitting definition statement is used to perform two-level splitting and select multiple splitting elements from the splitting results. The multiple splitting elements are mapped to at least two third reference fields under the target data format. The step of generating an element serialization mapping statement corresponding to the string segmentation definition statement based on the statement type of the string segmentation definition statement expressed in the target abstract syntax tree, the field type defined for each segmentation element in the string segmentation definition statement, the new element field mapped by each segmentation element under the target data format, and the sequence encoding assigned to the new element field under the serialization format includes: If the statement type of the string splitting definition statement expressed in the target abstract syntax tree is the fourth type, the second attribute keyword, the second composite field name defined for the at least two third reference fields, the second composite field defined for the at least two third reference fields, and the sequence code assigned to the second composite field under the serialization format are combined to obtain the serialization mapping statement of the second composite field; For each segmentation element, the first attribute keyword, the field type defined for the segmentation element, the third reference field mapped by the segmentation element under the target data format, and the sequence code assigned to the third reference field under the serialization format are combined to obtain the serialization mapping statement of the third reference field mapped by the field type of the segmentation element. Wherein, the first attribute keyword is an attribute keyword that indicates a non-mandatory attribute; the second attribute keyword is an attribute keyword that indicates a repeatedly selected attribute.
9. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire format conversion statements for converting target data in offline storage; the format conversion statements are used to convert the target data from its original data format to the target data format required by the online service. The parsing module is used to parse the format conversion statement to obtain intermediate format conversion statements for the execution engine to execute, and a target serialization protocol suitable for converting the target data between the target data format and the serialization format; The first execution module is used to load and execute the intermediate format conversion statement by the execution engine to convert the target data obtained from offline storage into a first intermediate data in the format of the target data. The second execution module is used by the execution engine to serialize the first intermediate data according to the target serialization protocol to obtain the target serialized data; A storage module is used to store the target serialized data online; The service module is used to parse the target serialized data into the first intermediate data according to the target serialization protocol after the online service obtains the target serialized data from the online storage, and then use the first intermediate data.
10. An electronic device, characterized in that, include: processor; A memory storing computer instructions that, when executed by the processor, implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-8.
12. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-8.