Data encoding and storage methods, devices, electronic equipment and storage media

CN117472912BActive Publication Date: 2026-09-01AGRICULTURAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311669448.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2026-09-01
Estimated Expiration
2043-12-07

AI Technical Summary

Technical Problem

[0004]本发明提供了一种数据的编码入库方法、装置、电子设备及存储介质,以解决相关技术中的数据的编码入库方法的效率较低以及通用性较差的技术问题

Benefits of technology

[0020]本发明实施例的技术方案,通过获取待编码数据集,其中,所述待编码数据集包括多个数据来源的待编码数据和/或多种数据结构的待编码数据;确定预设编码域,基于所述预设编码域对所述待编码数据集中的每个所述待编码数据进行编码得到目标编码数据;通过分流模型确定每个所述目标编码数据对应的目标分流编码,并将每个所述目标编码数据分流至所述目标分流编码对应的目标分流;通过提取入库模型确定所述目标分流编码对应的数据表编码,并将每个所述目标分流存入所述数据表编码对应的目标数据表。本申请基于数据编码对多源异构的大数据进行编码分析入库,可快速整合多源异构数据,提高了数据的编码入库的效率,且本申请适用于多种来源的数据,提高了数据的编码入库应用的广泛性和通用性。应当理解,本部分所描述的内容并非旨在标识本发明的实施例的关键或重要特征,也不用于限制本发明的范围。本发明的其它特征将通过以下的说明书而变得容易理解。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117472912B_ABST
    Figure CN117472912B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, electronic device, and storage medium for encoding and storing data. The method includes: acquiring a dataset to be encoded, wherein the dataset includes data from multiple data sources and / or data with multiple data structures; determining a preset encoding domain; encoding each piece of data in the dataset to be encoded based on the preset encoding domain to obtain target encoded data; determining a target distribution code corresponding to each target encoded data using a distribution model, and distributing each target encoded data to the target distribution code corresponding to the target distribution code; determining a data table code corresponding to the target distribution code using an extraction model, and storing each target distribution in the target data table corresponding to the data table code. Based on the technical solution of this invention, the efficiency of encoding and storing multi-source heterogeneous data can be improved, and the breadth and versatility of big data encoding and storage applications can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data encoding technology, and in particular to a method, apparatus, electronic device, and storage medium for encoding and storing data. Background Technology

[0002] Currently, big data analytics has been gradually applied to various platforms. By mining and analyzing massive amounts of data, a large amount of valuable information can be obtained. However, as the amount of data continues to increase, the data structure becomes more and more complex, and the data sources become more and more numerous, the data exhibits fragmented characteristics such as being scattered, heterogeneous, and of low quality. Therefore, the aggregation, analysis, and application of big data are becoming increasingly complex.

[0003] In related technologies, the process typically involves first grouping and classifying heterogeneous data from multiple sources by professionals in the relevant field. Then, the classified data is manually encoded layer by layer, and finally, the encoded data is stored in a relevant database for later application. Clearly, existing data encoding and storage methods involve too much manual work, are inefficient, and suffer from poor universality due to differences in data grouping and classification methods across different fields. Summary of the Invention

[0004] This invention provides a data encoding and storage method, apparatus, electronic device, and storage medium to solve the technical problems of low efficiency and poor versatility in data encoding and storage methods in related technologies.

[0005] According to one aspect of the present invention, a method for encoding and storing data is provided, wherein the method includes:

[0006] Obtain the dataset to be encoded, wherein the dataset to be encoded includes data to be encoded from multiple data sources and / or data to be encoded with multiple data structures;

[0007] A preset encoding domain is determined, and each piece of data to be encoded in the dataset to be encoded is encoded based on the preset encoding domain to obtain target encoded data;

[0008] The target split code corresponding to each target encoded data is determined by the split model, and each target encoded data is split to the target split corresponding to the target split code;

[0009] The data table code corresponding to the target traffic split code is determined by extracting the data entry model, and each target traffic split is stored in the target data table corresponding to the data table code.

[0010] According to another aspect of the present invention, a data encoding and storage apparatus is provided, wherein the apparatus comprises:

[0011] The data acquisition module is used to acquire the dataset to be encoded, wherein the dataset to be encoded includes data to be encoded from multiple data sources and / or data to be encoded with multiple data structures;

[0012] A data encoding module is used to determine a preset encoding domain and encode each piece of data to be encoded in the dataset to be encoded based on the preset encoding domain to obtain target encoded data.

[0013] The data splitting module is used to determine the target splitting code corresponding to each target encoded data through a splitting model, and to split each target encoded data to the target splitting code corresponding to the target splitting code;

[0014] The data entry module is used to determine the data table code corresponding to the target traffic split code by extracting the entry model, and to store each target traffic split into the target data table corresponding to the data table code.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the data encoding and storage method described in any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the data encoding and storage method described in any embodiment of the present invention.

[0020] The technical solution of this invention involves: acquiring a dataset to be encoded, wherein the dataset includes data from multiple data sources and / or data with multiple data structures; determining a preset encoding domain; encoding each piece of data in the dataset to be encoded based on the preset encoding domain to obtain target encoded data; determining a target split encoding corresponding to each target encoded data through a splitting model; and splitting each target encoded data to a target split corresponding to the target split encoding; determining a data table encoding corresponding to the target split encoding through an extraction and storage model; and storing each target split into a target data table corresponding to the data table encoding. This application performs encoding analysis and storage of multi-source heterogeneous big data based on data encoding, which can quickly integrate multi-source heterogeneous data, improve the efficiency of data encoding and storage, and is applicable to data from multiple sources, improving the breadth and versatility of data encoding and storage applications. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a data encoding and storage method according to Embodiment 1 of the present invention;

[0023] Figure 2 This is a data map for determining the source code of data according to an embodiment of the present invention;

[0024] Figure 3 This is a data diagram for determining the encoding of operation steps according to an embodiment of the present invention;

[0025] Figure 4 This is a flowchart illustrating a data splitting model provided by an embodiment of the present invention.

[0026] Figure 5 This is a data diagram for determining candidate split codes according to an embodiment of the present invention;

[0027] Figure 6 This is a flowchart of a data extraction and storage process according to an embodiment of the present invention.

[0028] Figure 7This is a flowchart of a data encoding and storage method according to Embodiment 2 of the present invention;

[0029] Figure 8 This is a flowchart illustrating data access using a data access model provided in an embodiment of the present invention;

[0030] Figure 9 This is a funnel analysis data graph for establishing a data access model provided by an embodiment of the present invention;

[0031] Figure 10 This is an event analysis data graph for establishing a data access model provided by an embodiment of the present invention;

[0032] Figure 11 This is a schematic diagram of the structure of a data encoding and storage device according to Embodiment 3 of the present invention;

[0033] Figure 12 This is a schematic diagram of the structure of an electronic device that implements the data encoding and storage method of this invention. Detailed Implementation

[0034] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0036] Example 1

[0037] Figure 1This is a flowchart illustrating a data encoding and storage method according to Embodiment 1 of the present invention. This embodiment is applicable to data encoding situations. The method can be executed by a data encoding and storage device, which can be implemented in hardware and / or software and can be configured in a computer. Figure 1 As shown, the method includes:

[0038] S110. Obtain the dataset to be encoded, wherein the dataset to be encoded is...

[0039] The dataset to be encoded can be understood as a large dataset to be encoded. Optionally, the dataset to be encoded may include data from multiple data sources and / or data with multiple data structures.

[0040] The data to be encoded can be understood as data to be encoded. In this embodiment of the invention, the data to be encoded can be preset according to scenario requirements, and is not specifically limited here. Optionally, the data to be encoded can be user behavior data throughout its entire lifecycle. The data to be encoded can include general fields, attribute fields, and business fields, wherein the general fields include scenario type, data source, operation steps, and operation time, etc.

[0041] S120. Determine a preset encoding domain, and encode each piece of data to be encoded in the dataset to be encoded based on the preset encoding domain to obtain target encoded data.

[0042] The preset encoding field can be understood as an encoding field used to encode the data to be encoded. In this embodiment of the invention, the preset encoding field can be preset according to scenario requirements, and is not specifically limited here. Optionally, the preset encoding field may include at least one of a general encoding field, an attribute encoding field, and a business encoding field.

[0043] The target encoded data can be understood as the encoded data to be encoded. Optionally, the target encoded data includes at least one of general encoding, attribute encoding, and business encoding. The general encoding is a fixed-length encoding and includes at least one of scenario encoding, data source encoding, operation step encoding, and operation time encoding.

[0044] The general encoding can be understood as the encoding corresponding to the general fields of the data to be encoded.

[0045] The attribute encoding can be understood as the encoding corresponding to the attribute field of the data to be encoded.

[0046] The business code can be understood as the code corresponding to the business field of the data to be encoded.

[0047] The scenario encoding can be understood as an encoding that characterizes the scenario type of the data to be encoded. The data source encoding can be understood as an encoding that characterizes the data source of the data to be encoded. The operation step encoding can be understood as an encoding that characterizes the operation steps of the data to be encoded. The operation time encoding can be understood as an encoding that characterizes the operation time of the data to be encoded.

[0048] For example, specifically, encoding each piece of data to be encoded in the dataset to be encoded based on the preset encoding domain to obtain the target encoded data can be:

[0049] The preset encoding domain is determined, which includes three parts: general encoding domain, attribute encoding domain, and business encoding domain.

[0050] 1. The general fields corresponding to the data to be encoded use fixed-length encoding (e.g., 35-bit length) to identify the scenario type, data source, operation steps, and operation time of the data to be encoded.

[0051] 1) The scene type is set using a 4-digit code, which is the abbreviation of the scene type.

[0052] 2) Data sources use a 5-digit encoding setting. The first 4 digits are the abbreviation of the data source, and the last digit is the abbreviation based on the data upload method. For example: P for page tracking uploads; I for API log uploads; F for file uploads, etc. (see reference) Figure 2 ), Figure 2 This is a data map for determining the source code of data, provided according to an embodiment of the present invention.

[0053] 3) The operation steps use a 12-digit encoding setting, which consists of a 2-digit abbreviation of the data source, a 4-digit abbreviation of the operation node, and a 4-digit abbreviation of the operation type. The operation type can include click, browse, log, and transaction history, separated by hyphens (see reference). Figure 3 ), Figure 3 This is a data diagram for determining the encoding of operation steps according to an embodiment of the present invention.

[0054] 4) The operation time is set using a 14-bit encoding, concatenating the date and time. The date format is YYYYMMDD, and the time format is HHMMSS, such as: date 20230101, time 123030.

[0055] 2. The attribute field mainly identifies user attribute information and uses variable-length encoding, with a length of 11 to 50 bits.

[0056] 3. Business fields are mainly data analysis-related fields proposed based on business scenarios, using variable-length encoding, with a length of 12 to 60 bits.

[0057] S130. Determine the target split code corresponding to each target encoded data through the split model, and split each target encoded data to the target split corresponding to the target split code.

[0058] The splitting model can be understood as a model capable of matching the target identification code and the target splitting code corresponding to the target encoded data. The input data of the splitting model is the target identification code corresponding to the target encoded data, and the output data is the target splitting code corresponding to the target identification code.

[0059] The target offloading code can be understood as the code corresponding to the target offloading.

[0060] The target splitting can be understood as file splitting or topic splitting corresponding to the target encoded data.

[0061] Optionally, before determining the target split code corresponding to each target encoded data through the split model, the method further includes:

[0062] Based on the scene encoding and the data source encoding, determine the target identification code corresponding to each target encoded data;

[0063] Determine candidate streams and their corresponding candidate stream codes, wherein the candidate streams include multiple file streams and / or multiple topic streams;

[0064] A first correspondence is established between the target identification code and the candidate splitting code, and the splitting model is determined based on the first correspondence.

[0065] The target identification code can be understood as a unique identification code corresponding to each target coded data. Specifically, the target identification code is obtained by superimposing the scene code and the data source code. In this embodiment of the invention, both the scene code and the data source code are unique for each target coded data; therefore, the target identification code corresponding to each target coded data is also unique.

[0066] The candidate splitting can be understood as splitting candidates. Optionally, the candidate splitting can be file splitting or topic splitting.

[0067] The candidate stream splitting code can be understood as the code corresponding to the candidate stream splitting. In this embodiment of the invention, each candidate stream splitting can correspond to one candidate stream splitting code.

[0068] The first correspondence can be understood as the correspondence between the target identification code and the candidate splitting code. In embodiments of the present invention, the first correspondence can be a one-to-one relationship, a one-to-many relationship, or a many-to-many relationship, etc.

[0069] Figure 4 This is a flowchart illustrating a data splitting model provided by an embodiment of the present invention. For example... Figure 4 As shown, exemplarily and specifically, determining the target split code corresponding to each target encoded data through a split model, and splitting each target encoded data to the target split code corresponding to the target split code, can be:

[0070] 1. It is understandable that different scenario types and data sources have corresponding unique codes. By establishing the correspondence between "scenario code + data source code" and the preset candidate diversion codes, a diversion model is formed.

[0071] 2. When data is split and transmitted, the corresponding relationship configuration can be queried according to the scenario code and the data source code to obtain the target split code. Then, the corresponding target split code can be found according to the target split code to perform data split and transmission.

[0072] For example, when transmitting data in real time via Kafka, topic encoding (topic split encoding) can be used as a candidate split encoding. According to the split model, the data is split and transmitted to the corresponding topic (topic split). If it is a non-real-time file transmission, file name can be used as a candidate split encoding. According to the split model, the target encoded data is split and written to the corresponding file split.

[0073] Taking topic encoding as an example, using a variable-length setting, the candidate split encoding consists of three parts: the first part is the data layer, in a 3-digit fixed-length format; the second part is the database table name, which is variable-length; the third part consists of the number of Kafka replicas and partitions, which is variable-length, such as R2P6, where R represents replicas, 2 is the number of replicas, P represents partitions, and 6 is the number of partitions (see reference). Figure 5 ), Figure 5 This is a data diagram for determining candidate split codes according to an embodiment of the present invention.

[0074] S140. Determine the data table code corresponding to the target split code by extracting the database model, and store each target split into the target data table corresponding to the data table code.

[0075] The extraction and input model can be understood as a model capable of matching the target split code with the data table code. The input data of the extraction and input model is the target split code, and the output data is the data table code corresponding to the target split code.

[0076] The data table encoding can be understood as the encoding corresponding to each candidate data table.

[0077] The target data table can be understood as the data table that stores the target stream.

[0078] Optionally, the extraction and storage model includes a data extraction model and a data storage model. The data extraction model is determined based on the correspondence between candidate split codes and target code data. The data storage model is determined based on the correspondence between operation step codes and candidate table codes.

[0079] Optionally, before determining the data table code corresponding to the target traffic splitting code by extracting the database model, the method further includes:

[0080] Determine the candidate data table and the candidate table code corresponding to the candidate data table;

[0081] A second correspondence is established between the candidate table encoding and the candidate split encoding, and the extraction and storage model is determined based on the second correspondence.

[0082] The candidate data table can be understood as a list of candidate data tables. In this embodiment of the invention, the candidate data table may be one or more.

[0083] The candidate table encoding can be understood as the encoding corresponding to the candidate data table.

[0084] The second correspondence can be understood as the correspondence between the candidate table encoding and the candidate split encoding. In embodiments of the present invention, the second correspondence can be a one-to-one relationship, a one-to-many relationship, or a many-to-many relationship, etc.

[0085] Figure 6 This is a flowchart illustrating data extraction and storage according to an embodiment of the present invention. Figure 6 As shown, the data extraction and storage process can be as follows:

[0086] 1. Establish data extraction and data entry models.

[0087] Data extraction model: Establish the correspondence between candidate split codes and target encoded data.

[0088] Data entry model: Establish the correspondence between operation step codes and candidate table codes.

[0089] 2. Based on real-time or non-real-time settings, the data extraction model is used to obtain the uploaded data under the target split encoding, such as consuming the topic split of Kafka corresponding to the target split encoding or reading the file split of the corresponding file name.

[0090] 3. Based on the data entry model, find the data table code corresponding to the operation step code, and enter the data into the target data table.

[0091] The technical solution of this invention involves: acquiring a dataset to be encoded, wherein the dataset includes data from multiple data sources and / or data with multiple data structures; determining a preset encoding domain; encoding each piece of data in the dataset to be encoded based on the preset encoding domain to obtain target encoded data; determining a target split encoding corresponding to each target encoded data through a splitting model; and splitting each target encoded data to a target split corresponding to the target split encoding; determining a data table encoding corresponding to the target split encoding through an extraction and storage model; and storing each target split into a target data table corresponding to the data table encoding. This application performs encoding analysis and storage of multi-source heterogeneous big data based on data encoding, which can quickly integrate multi-source heterogeneous data, improve the efficiency of data encoding and storage, and is applicable to data from multiple sources, improving the breadth and versatility of data encoding and storage applications.

[0092] Example 2

[0093] Figure 7 This is a flowchart of a data encoding and storage method provided in Embodiment 2 of the present invention. This embodiment appends data to the target data table corresponding to the data table encoding described in the above embodiments. Figure 7 As shown, the method includes:

[0094] S210. Obtain the dataset to be encoded, wherein the dataset to be encoded includes data to be encoded from multiple data sources and / or data to be encoded with multiple data structures.

[0095] S220. Determine a preset encoding domain, and encode each piece of data to be encoded in the dataset to be encoded based on the preset encoding domain to obtain target encoded data.

[0096] S230. Determine the target split code corresponding to each target encoded data through the split model, and split each target encoded data to the target split corresponding to the target split code.

[0097] S240. Determine the data table code corresponding to the target split code by extracting the database model, and store each target split into the target data table corresponding to the data table code.

[0098] S250. Determine the target access code for the data to be accessed based on the scenario code, the data source code, and the operation step code.

[0099] The data to be accessed can be understood as data to be accessed. In this embodiment of the invention, the data to be accessed may be the data to be encoded that has been stored in the target data table.

[0100] The target access code can be understood as the input code during data access. In this embodiment of the invention, the target access code can be determined based on the scenario code, the data source code, and the operation step code.

[0101] S260. The target access code is queried using a data access model to obtain and access the data to be accessed.

[0102] The data access model can be understood as a model capable of performing data queries based on the target access code. The input data of the data access model is the target access code, and the output data is the data to be encoded corresponding to the target access code.

[0103] Optionally, the step of querying the input target access code through a data access model to obtain and access the data to be accessed includes:

[0104] The target access code is matched with the input using a data access model to obtain the data table code corresponding to the target access code;

[0105] The data to be accessed is obtained and accessed based on the data table encoding corresponding to the data table.

[0106] Figure 8 This is a flowchart illustrating data processing using a data access model provided by an embodiment of the present invention. For example... Figure 8 As shown, the data access process can be as follows:

[0107] 1. Establish a data access model. For example, funnel analysis and event analysis models will be used. The following example demonstrates data access in a single-scenario funnel analysis (see reference). Figure 9 ), Figure 9 This is a funnel analysis data diagram for establishing a data access model according to an embodiment of the present invention. Specifically, 1) define the funnel hierarchy; 2) determine the coding range of the associated steps for each hierarchy; 3) combine the funnels, separating each hierarchy with a | and separating the step codes within the same hierarchy with a *; 4) set the model ID and establish the correspondence between the model ID and the model configuration. Setting the funnel: Open product page - Buy Now (submit product page, submit shopping cart) - Submit order - Complete payment, the corresponding step codes are TB-OPEN-CL01, TB-BUYS-CL01, TB-BUYS-CL02, TB-SUBM-CL01, TB-PAYS-CL01.

[0108] For example, consider an event analysis tree diagram in a search scenario (see reference). Figure 10 ), Figure 10 This is an event analysis data graph for establishing a data access model according to an embodiment of the present invention. Specifically, 1) Define a tree-like layer hierarchy, taking three layers as an example; 2) Determine the step code corresponding to each tree node; 3) Assemble the model, assembling it according to the hierarchy. First, define the layer sequence number, such as the first layer sequence number 2, the second layer sequence number 2, etc.; then determine the step code of the parent node, the first layer parent node step code is empty; then assemble the child nodes under the parent node step code, with each child node separated by |, and the same child node involving multiple step codes separated by *; 4) Set the model ID, and establish the correspondence between the model ID, layer sequence number, parent node step code and child node code assembly. Set the model: the first layer is the total search access volume, the second layer is the access volume of each section of the search homepage, and the third layer is the access volume of the second layer sub-items.

[0109] 2. Accessing Data. 1) Obtain the data access model based on the model ID, and decompose the model configuration layer by layer to obtain the step code set for each layer; 2) Obtain the data table code corresponding to the target access code based on the data access model; 3) Query the data table corresponding to the data table code to obtain the data details of the corresponding operation step code.

[0110] The technical solution of this invention determines the target access code of the data to be accessed based on the scenario code, the data source code, and the operation step code; it then uses a data access model to query the input target access code to obtain and access the data to be accessed. This improves the convenience of accessing encoded data already stored in the database.

[0111] Example 3

[0112] Figure 11 This is a schematic diagram of a data encoding and storage device provided in Embodiment 3 of the present invention. Figure 11 As shown, the device includes: a data acquisition module 310, used to acquire a dataset to be encoded, wherein the dataset to be encoded includes data to be encoded from multiple data sources and / or data to be encoded with multiple data structures; a data encoding module 320, used to determine a preset encoding domain, and encode each piece of data to be encoded in the dataset to be encoded based on the preset encoding domain to obtain target encoded data; a data splitting module 330, used to determine the target splitting code corresponding to each piece of target encoded data through a splitting model, and split each piece of target encoded data to the target splitting code corresponding to the target splitting code; and a data storage module 340, used to determine the data table code corresponding to the target splitting code through an extraction storage model, and store each target splitting data into the target data table corresponding to the data table code.

[0113] The technical solution of this invention involves: acquiring a dataset to be encoded, wherein the dataset includes data from multiple data sources and / or data with multiple data structures; determining a preset encoding domain; encoding each piece of data in the dataset to be encoded based on the preset encoding domain to obtain target encoded data; determining a target split encoding corresponding to each target encoded data through a splitting model; and splitting each target encoded data to a target split corresponding to the target split encoding; determining a data table encoding corresponding to the target split encoding through an extraction and storage model; and storing each target split into a target data table corresponding to the data table encoding. This application performs encoding analysis and storage of multi-source heterogeneous big data based on data encoding, which can quickly integrate multi-source heterogeneous data, improve the efficiency of data encoding and storage, and is applicable to data from multiple sources, improving the breadth and versatility of data encoding and storage applications.

[0114] Optionally, the preset encoding domain includes at least one of a general encoding domain, an attribute encoding domain, and a business encoding domain.

[0115] Optionally, the target encoded data includes at least one of general encoding, attribute encoding, and business encoding. The general encoding is a fixed-length encoding and includes at least one of scenario encoding, data source encoding, operation step encoding, and operation time encoding.

[0116] Optionally, the data encoding and storage device further includes an identification encoding determination module, a candidate splitting determination module, and a splitting model determination module; wherein,

[0117] The identification code determination module is used to determine the target identification code corresponding to each target coded data based on the scene code and the data source code before determining the target split code corresponding to each target coded data through the split model;

[0118] The candidate splitting determination module is used to determine candidate splitting and the candidate splitting code corresponding to the candidate splitting, wherein the candidate splitting includes multiple file splitting and / or multiple topic splitting;

[0119] The traffic splitting model determination module is used to establish a first correspondence between the target identification code and the candidate traffic splitting code, and to determine the traffic splitting model based on the first correspondence.

[0120] Optionally, the data encoding and storage device further includes a candidate table determination module and a storage model determination module; wherein,

[0121] The candidate table determination module is used to determine candidate data tables and candidate table codes corresponding to the candidate tables before determining the data table code corresponding to the target diversion code by extracting the database model.

[0122] The database entry model determination module is used to establish a second correspondence between the candidate table code and the candidate split code, and to determine the extraction database entry model based on the second correspondence.

[0123] Optionally, the data encoding and storage device further includes an access encoding determination module and a data query module; wherein,

[0124] The access code determination module is used to determine the target access code of the data to be accessed based on the scenario code, the data source code, and the operation step code.

[0125] The data query module is used to query the input target access code through the data access model to obtain and access the data to be accessed.

[0126] Optionally, the data query module is used for:

[0127] The target access code is matched with the input using a data access model to obtain the data table code corresponding to the target access code;

[0128] The data to be accessed is obtained and accessed based on the data table encoding corresponding to the data table.

[0129] The data encoding and storage device provided in the embodiments of the present invention can execute the data encoding and storage method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0130] Example 4

[0131] Figure 12 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0132] like Figure 12As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0133] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0134] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data encoding and storage methods.

[0135] In some embodiments, the data encoding and storage method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data encoding and storage method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data encoding and storage method by any other suitable means (e.g., by means of firmware).

[0136] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0137] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0138] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0139] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0140] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0141] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0142] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0143] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for encoding and storing data, characterized in that, include: Obtain the dataset to be encoded, wherein the dataset to be encoded includes data to be encoded from multiple data sources and / or data to be encoded with multiple data structures; A preset encoding domain is determined, and each piece of data to be encoded in the dataset to be encoded is encoded based on the preset encoding domain to obtain target encoded data; The target identification code corresponding to each target coded data is determined based on the scene code and the data source code; for each target coded data, the scene code and the data source code are unique, and the target identification code corresponding to each target coded data is also unique; Determine candidate streams and their corresponding candidate stream codes, wherein the candidate streams include multiple file streams and / or multiple topic streams; Establish a first correspondence between the target identification code and the candidate diversion code, and determine the diversion model based on the first correspondence. The target split code corresponding to each target encoded data is determined by the split model, and each target encoded data is split to the target split corresponding to the target split code; The data table code corresponding to the target traffic split code is determined by extracting the data entry model, and each target traffic split is stored in the target data table corresponding to the data table code. The target access code for the data to be accessed is determined based on the scenario code, the data source code, and the operation step code. The target access code is queried using a data access model to obtain and access the data to be accessed.

2. The method according to claim 1, characterized in that, The preset encoding domain includes at least one of the following: a general encoding domain, an attribute encoding domain, and a business encoding domain.

3. The method according to claim 1, characterized in that, The target encoded data includes at least one of general encoding, attribute encoding, and business encoding. The general encoding is a fixed-length encoding and includes at least one of scenario encoding, data source encoding, operation step encoding, and operation time encoding.

4. The method according to claim 1, characterized in that, Before determining the data table code corresponding to the target traffic splitting code by extracting the database model, the method further includes: Determine the candidate data table and the candidate table code corresponding to the candidate data table; A second correspondence is established between the candidate table encoding and the candidate split encoding, and the extraction and storage model is determined based on the second correspondence.

5. The method according to claim 1, characterized in that, The step of querying the input target access code through a data access model to obtain and access the data to be accessed includes: The target access code is matched with the input using a data access model to obtain the data table code corresponding to the target access code; The data to be accessed is obtained and accessed based on the data table encoding corresponding to the data table.

6. A data encoding and storage device, characterized in that, include: The data acquisition module is used to acquire the dataset to be encoded, wherein the dataset to be encoded includes data to be encoded from multiple data sources and / or data to be encoded with multiple data structures; A data encoding module is used to determine a preset encoding domain and encode each piece of data to be encoded in the dataset to be encoded based on the preset encoding domain to obtain target encoded data. The identification code determination module is used to determine the target identification code corresponding to each target coded data based on the scene code and the data source code; for each target coded data, the scene code and the data source code are unique, and the target identification code corresponding to each target coded data is also unique; A candidate splitting module is used to determine candidate splitting and the candidate splitting code corresponding to the candidate splitting, wherein the candidate splitting includes multiple file splitting and / or multiple topic splitting; The traffic splitting model determination module is used to establish a first correspondence between the target identification code and the candidate traffic splitting code, and determine the traffic splitting model based on the first correspondence. The data splitting module is used to determine the target splitting code corresponding to each target encoded data through a splitting model, and to split each target encoded data to the target splitting code corresponding to the target splitting code; The data entry module is used to determine the data table code corresponding to the target traffic code by extracting the entry model, and to store each target traffic into the target data table corresponding to the data table code. The access code determination module is used to determine the target access code of the data to be accessed based on the scenario code, the data source code, and the operation step code. The data query module is used to query the input target access code through the data access model to obtain and access the data to be accessed.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data encoding and storage method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the data encoding and storage method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-terminal heterogeneous data distribution storage method and system based on data medium station

    CN116955743A

  • Data element tag-based data distribution method and device and readable medium

    CN116991842A