Label data processing method, computer equipment and storage medium
By obtaining the entity view table of the tag entity and generating the first tag table, the problem of low efficiency in tag data processing is solved, and efficient processing of large amounts of data is achieved.
Patent Information
- Application Number
- CN202510645818.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-30
AI Technical Summary
The existing label data processing efficiency is low and it is difficult to meet the efficient processing requirements under large data volumes.
By obtaining the entity view table of the tag entity, the operation on the original data is reduced, and the first tag table and the integrated entity view table are generated to improve the processing efficiency.
It reduces the amount of data required for subsequent calculations, improves the processing efficiency of label data, and adapts to the demand for efficient processing of large amounts of data.
Smart Images

Figure CN120723764A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of label processing technology, and in particular to a label data processing method, computer equipment, and storage medium. Background Art
[0002] With the rapid development of computer technology, in order to effectively manage data, data labeling can be used to process labeled data.
[0003] At present, due to the large amount of label data, the existing related technologies lead to low efficiency in processing label data. Therefore, how to improve the processing efficiency of labels is a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The main technical problem solved by this application is to provide a label data processing method, computer equipment and storage medium, which can improve the processing efficiency of label data.
[0005] The first aspect of the present application provides a method for processing label data, which includes: obtaining an entity view table of a label entity, the entity view table containing a mapping relationship to the original data of the label entity; obtaining label data of the label entity, processing the label data, and generating a first label table; integrating the first label table and the entity view table to obtain a second label table.
[0006] A second aspect of the present application provides a computer device, which includes a memory and a processor coupled to each other, wherein the memory stores program data, and the processor is used to execute the program data to implement any step of the above-mentioned tag data processing method.
[0007] A third aspect of the present application provides a computer-readable storage medium, which stores program data that can be executed by a processor, and the program data is used to implement any step of the above-mentioned label data processing method.
[0008] The above scheme obtains the entity view table of the label entity. Since the entity view table contains the mapping relationship of the original data of the label entity, subsequent processing can be performed on the entity view table, reducing the operation of the original data, improving the processing efficiency of the label data, and reducing the amount of data calculated in the subsequent processing process; in addition, the label data of the label entity is obtained, the label data is processed to generate a first label table, and different label data can be processed simultaneously to generate the first label table, and then the first label table and the entity view table are integrated to obtain the second label table, which can improve the processing efficiency of the label data.
[0009] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of this application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be derived from these drawings without inventive effort. Among them:
[0011] Figure 1 This is a flowchart of an embodiment of a method for processing tag data of the present application;
[0012] Figure 2 This application Figure 1 A flow chart of an embodiment of step S12;
[0013] Figure 3 This is an example schematic diagram of an embodiment of multi-level tag data of the present application;
[0014] Figure 4 This application Figure 1 A flow chart of an embodiment of step S13;
[0015] Figure 5 This is a structural diagram of an embodiment of a device for processing label data of the present application;
[0016] Figure 6 This is a schematic structural diagram of an embodiment of a computer device of the present application;
[0017] Figure 7 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0019] The terms "first" and "second" in this application are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or devices.
[0020] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0021] The term "and / or" in this article is simply a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0022] This application provides the following embodiments, and each embodiment is described in detail below.
[0023] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of a method for processing tag data of the present application. The method may include the following steps:
[0024] S11: Obtain an entity view table of the tag entity, where the entity view table includes a mapping relationship to the original data of the tag entity.
[0025] Among them, the label entity can be an entity object that is the source of the label data to be processed. For example, the label entity can include living things, entities, equipment, areas, etc. The label entity can be an independent business object. Each label entity can have an independent label system. Each label entity contains a number of original data. The original data can be stored in a database, etc. The database can be an online or offline database. For example, for the application scenario of offline processing of label data, the database can be an offline data warehouse, etc. This application does not impose any restrictions on label entities.
[0026] Label data can be a form of data organization. In some business scenarios, it can be calculated according to preset logical rules, and can be a symbol (description) that relatively accurately reflects the business characteristics of the entity. For example, it can be the characteristics of the entity object collected and analyzed from the attributes, behavioral data, etc. of the entity object. Label data is a feature, a generalization, and an abstraction of business logic. Generally speaking, the value of label data has the characteristics of being highly generalized, independent of each other, and exhaustively enumerable. For example, label data can include "category", "duration", etc. The label value is the data value associated with the label data, which specifically describes the characteristics represented by the label data. For example, the label value of "category" label data is "first category", "second category", etc. This application does not limit label data.
[0027] Optionally, the classification of label data may be differentiated based on the dimension of label objects, and label data corresponding to different label entities cannot reference each other.
[0028] An entity view table of the original data of the tag entity may be obtained, and the entity view table may include a mapping relationship to the original data of the tag entity, so that the entity view table can serve as a mapping of the original data.
[0029] In some embodiments, the original data of the tag entity is obtained. The user can select the original data, and in response to the selection of the original data, the selected original data is determined. Specifically, the original data of the tag entity can be stored in a data table, that is, the original table of the entity. The original table can contain several data fields of the tag entity, and the amount of data is relatively large. In the process of creating the tag entity, the original data (that is, the data fields) required to create the tag data can be selected. The required data fields can be selected without using all the data fields, which can reduce the amount of data for subsequent calculations.
[0030] After determining the selected original data, use the selected original data to create an entity view table for the label entity. Determining the selected original data means determining the data fields for subsequent creation of label data. Entity view nodes can be created for the selected data fields. The entity view nodes can be represented as view tables, i.e., entity view tables, which can be used as a mapping of the original table and do not occupy storage space. In subsequent processing, data operations on the original table will not be performed, reducing the amount of invalid data caused by the use of the original table and reducing the amount of subsequent calculations. In addition, the entity view table can isolate the user's operating behavior (or subsequent processing) from the data in the original table. When the user operates the entity view table, it will not affect the original table, thereby protecting the original data of the original table.
[0031] In some embodiments, in response to updating a tag entity, such as in response to updating selected raw data, this approach may invalidate the entity view table, and a new entity view table may be recreated based on the updated selected raw data.
[0032] In the above method, creating an entity view table for the tag entity can reduce the amount of data calculated in the subsequent processing process.
[0033] S12: Acquire label data of the label entity, process the label data, and generate a first label table.
[0034] After creating the tag entity, the tag data of the tag entity can be created. In this way, the tag data of the tag entity can be obtained, and then the tag data can be processed to generate a first tag table of each tag data. The first tag table can be represented as a tag temporary table.
[0035] In some embodiments, the tag data is divided into multiple levels of tag data, and dependencies exist between the tag data at different levels. The semantic expressiveness of the tag data at different levels varies, with higher-level tag data having higher semantic expressiveness. Dependency relationships can represent reference or referenced relationships, and this application does not impose any restrictions on this.
[0036] In some embodiments, see Figure 2 , step S12 of the above embodiment can be further expanded. Acquiring the tag data of the tag entity, processing the tag data, and generating a first tag table, this embodiment may include the following steps:
[0037] S121: Creating tag data for the tag entity.
[0038] The label data can be divided into multiple levels, and there may be dependencies between the label data of each level. Optionally, the label data of the subsequent level may reference at least one label data of the previous level, and / or, the label data of the subsequent level may be generated by at least one label data of the previous level. Among them, the previous level may represent at least one level that is located before the subsequent level. The label data of the first level is composed of the original data of the label entity, such as the selected original data (selected field data). The label entity is the basis of label processing, and a label entity is the data source of the label. For example, the label entity can be an object, around which a lot of label data can be created, such as labels for attribute classifications, attribute ranges, and other labels.
[0039] See also Figure 3 Label data is divided into layers. Taking three layers of label data as an example, all label data in the first layer (such as the first layer label 1) can reference the selected original data of the label entity. The label data in the first layer can include atomic labels or fact labels, and the processed data is derived from the original field data of the label entity. Atomic labels or fact labels represent quantifiable independent labels directly mapped from the original data. They are used to describe the basic attributes of the business object. The data itself has the characteristics of independent enumeration.
[0040] All label data in the second-layer label (such as second-layer label 1, second-layer label 2), the processed data depends on the first-layer label data or the original data selected by the label entity. The label data of the second layer can contain derived labels. Among them, the derived label can represent a rule-based label generated based on an atomic label or a field in a data table. Optionally, its label data can classify attributes or determine the range to which its attributes belong, etc. For example, the derived label "high label value object" corresponds to an object whose label value of a certain attribute is greater than or equal to a preset value.
[0041] All the labels of the third layer (such as the third layer label 1), their processing data comes from the label data of the first layer or the second layer. In addition to the labels of the first layer, the second and third layer labels can depend on one or more labels. The label data of the third layer can include combined labels or compound labels, etc., wherein the combined labels or compound labels can represent label data obtained by combining or compounding the label data of the previous layer. Optionally, the data range of the combined label comes from the "label entity", and the combined label can be composed of atomic labels and / or derived labels, etc. For example, a "high label value object" can be used as a combined label.
[0042] That is, in the process of creating label data at each level, creating the first-level label data only requires a label entity and does not rely on anything else. Creating the second-level label data requires relying on the first-level label data and / or label entity. Creating the third-level label data is based on the creation of the first-level label data or the second-level label data, and can rely on the first-level label data and / or the second-level label data. If the corresponding first-level label data or second-level label data is not created under the current label entity, the creation of the corresponding label data for the third level is not allowed.
[0043] The label data at each level of this application can be determined based on the specific application scenario or label entity, and there is no restriction on the label data at each level of this application.
[0044] The above approach sets up label layers and label dependencies for label data. This allows the current label data and its downstream label data to be updated during the label data update process, eliminating the need for a full update. This facilitates management and updates, and allows for faster and more flexible processing. Furthermore, by dividing label data into multiple levels, different label data can be combined or composited into new label data, enriching the label data's meaning. Label data at different levels can express different levels of meaning.
[0045] S122: Obtain label processing rules; wherein the label processing rules include preset label fields in the label data that need to be processed.
[0046] Users can customize label processing rules for label data according to their needs. Label processing rules can define preset label fields in label data that need to be processed, such as label name, label value and other fields. This application does not limit label processing rules.
[0047] S123: Process the label data using label processing rules to generate a first label table.
[0048] After the label processing rules are defined, the label data is processed using the label processing rules to obtain a first label table.
[0049] In some embodiments, a fourth processing statement is generated for each piece of label data using label processing rules. This fourth processing statement can be a statement such as SQL (Structured Query Language) or Python. For example, SQL can represent the execution SQL statement for the label data processing task, generating a comprehensive SQL statement for the processed data that can be used when executing subsequent tasks. Based on the generated fourth processing statement, the label data is processed according to a preset execution order to generate a first label table. The preset execution order represents the order of dependencies between the label data at each layer.
[0050] Optionally, after the label data is created, the label metadata may be stored to provide basic information for subsequent processing corresponding to the first label table.
[0051] Optionally, after obtaining the tag data of the tag entity, a temporary tag table can be created for each tag data. The temporary tag table is used to store the tag result data of the tag data during the processing process. The temporary tag table of each tag data can be independent of each other. After the processing of any tag data is completed, the tag result data of the tag data can be written into the temporary tag table, thereby obtaining a first tag table.
[0052] Optionally, the first label table includes at least a label temporary table, and each label data may correspond to a label temporary table.
[0053] Optionally, the first label table includes at least one of the following: a unique identifier, a label name, and a label partition.
[0054] Optionally, the tag metadata includes at least one of the following: tag name, tag temporary table name, tag temporary table partition, dependent tag, tag level, and tag entity.
[0055] For example, the tag metadata is shown in Table 1 below:
[0056]
[0057] The specific information of the tag metadata is as follows:
[0058] Tag name: This is the field name of the tag during the use of the tag data. The tag name is globally unique under the tag entity and can be used to identify globally unique tag data.
[0059] Temporary Label Table Name: This represents the temporary label table created for the label data after the label data is created. It is used to store the label result data during the label data processing process. After each label data is created, a temporary label table is generated. Each temporary label table can be independent and only stores the data corresponding to that label data. When the processing task for a specific label data is completed, the corresponding label result data for this label data is populated into the label intermediate table (also known as the temporary label table or the first label table).
[0060] Optionally, the temporary tag table can contain multiple columns of information data, such as an ID (Identity Document, a unique identifier), a tag name, and a tag partition. The ID corresponds one-to-one with the ID of the tag entity. Each row of data processed from the tag entity and each tag data item are in a corresponding relationship. The tag name is the tag name in the tag metadata. The tag partition represents the partition of the temporary tag table in the tag metadata.
[0061] Tag temporary table partitions: These partitions record the time data in the tag temporary table was written. Partitions are divided by time. For example, a day can be represented as a partition, a week can be represented as a partition, and so on. This application does not impose any restrictions on partitions.
[0062] Dependent tags: The tag data that the tag data depends on, that is, the tag data that the tag data references. For example, the second-layer tag data can depend on the first-layer tag data and / or tag entities, etc., and all the tag data that the tag data depends on can be recorded and stored.
[0063] Tag level: records the specific level to which the tag data belongs and categorizes the tags. For example, the tag level can include the first level, second level, third level, etc.
[0064] Label entity: Indicates the entity to which the label data belongs. One piece of label data belongs to only one entity. It can distinguish the domain and category of the label data.
[0065] In some embodiments, the fourth processing statement referred to above includes at least one of a delete statement (drop), a create statement (create), and an insert statement (insert). A delete statement is used to delete data, a create statement is used to create data, and an insert statement is used to insert data. Taking SQL statements as an example, a delete statement, a create statement, and an insert statement can be represented as a drop statement, a create statement, and an insert statement, respectively, meaning that executing the corresponding statements can accomplish different tasks.
[0066] The drop statement can perform data deletion tasks. For example, in this step, you can delete the temporary label table in the system, that is, you can clear the system's previously processed or existing temporary label table to prevent data processing disorder.
[0067] The create statement is used to execute data creation tasks. In this step, you can create a temporary tag table, that is, create a temporary tag table in the system based on the definition of tag data.
[0068] The insert statement is used to insert data. In this step, the tag data can be inserted into the temporary tag table. In addition, the SQL statement generated by the defined rules can be combined to generate the final insert SQL statement for inserting the tag data.
[0069] Optionally, when a tag task is running, the aforementioned drop, create, and insert statements are executed in a preset order for each tag data item. The preset order is: first execute the drop statement, then the create statement, and finally the insert statement. For each tag data item, a corresponding SQL statement can be generated. After generating each SQL statement, each SQL statement can be stored so that it is executed sequentially when the task for the corresponding tag data is executed.
[0070] The above steps can refresh the definition of the latest label temporary table. After completing the processing of each label data, the latest label temporary table can be created for each label data to mark the label data as being in an available state. The available state can mean that other label data can be created using the label data, and the label data is referenceable or dependent. In some application scenarios, when the label temporary table is created but no data is filled in, if the downstream label data needs to rely on this label data, it can rely on this label temporary table. If the processing task of the downstream label data is started at this time, although the downstream label data cannot be processed to the data, it does not affect the operation of the overall system and no errors or abnormalities will occur.
[0071] The above method, by obtaining the above first label table, can provide an intermediate basis for the maintenance of the label result data of the processed label data, thereby facilitating the maintenance of individual label data.
[0072] Continue reading Figure 1 After the above step S12, the following steps are also included:
[0073] S13: Integrate the first label table and the entity view table to obtain a second label table.
[0074] After the label data processing is completed, the label data processing can be summarized, that is, the data of the first label table generated by all the label data in the above entire process is summarized into the second label table, which can be used as the final label result table.
[0075] The first label table and the entity view table may be integrated to obtain a second label table.
[0076] The above scheme obtains the entity view table of the label entity. Since the entity view table contains the mapping relationship of the original data of the label entity, subsequent processing can be performed on the entity view table, reducing the operation of the original data, improving the processing efficiency of the label data, and reducing the amount of data calculated in the subsequent processing process; in addition, the label data of the label entity is obtained, the label data is processed to generate a first label table, and different label data can be processed simultaneously to generate the first label table, and then the first label table and the entity view table are integrated to obtain the second label table, which can improve the processing efficiency of the label data.
[0077] In some embodiments, see Figure 4 , step S13 of the above embodiment can be further expanded. The first label table and the entity view table are integrated to obtain a second label table. This embodiment may include the following steps:
[0078] S131: Based on the tag metadata of the tag data, write the data fields in the first tag table into the tag narrow table.
[0079] All tag data that needs to be aggregated under the current tag entity can be obtained, as well as the tag metadata stored for each tag data. The tag metadata provides information for subsequent processing statements. Then, the first tag table corresponding to each tag data is obtained, and based on the tag metadata of each tag data, the data fields in the first tag table are written into the tag narrow table.
[0080] In some embodiments, a first processing statement for processing a narrow tag table can be generated based on tag metadata in the first tag table. The first processing statement can then be used to write data fields from the first tag table into the narrow tag table. The narrow tag table can be used to store data fields from a temporary tag table. The first processing statement can be used to write data fields from the first tag table into the narrow tag table.
[0081] Optionally, the temporary tag table may include at least one of the following data fields: ID, tag name, and tag partition.
[0082] Optionally, the tag narrow table may include at least one of the following data fields: ID, tag name, tag value, tag partition, and other information.
[0083] In some embodiments, the first processing statement includes at least one of a delete statement (drop), a create statement (create), and an insert statement (insert). The delete statement is used to delete data, the create statement is used to create data, and the insert statement is used to insert data. This section can be referred to in the fourth processing statement for the label temporary table above and will not be further described here.
[0084] In some embodiments, the drop statement and create statement can refer to the fourth processing statement of the above-mentioned label temporary table, which is similar to the above and will not be repeated here. For the insert statement, an insert statement can be generated for each label data to insert the data in the label temporary table corresponding to the label data into the label narrow table. The label metadata of the label data stores the label temporary table name, and the data fields can be inserted into the label narrow table in sequence based on this label temporary table name.
[0085] Alternatively, you can first delete the historical label narrow table, create a new label narrow table, and then insert the data fields of the label temporary table corresponding to each label data into the label narrow table. For example, using the first processing statement, you can sequentially insert the data fields of the label temporary table corresponding to all label data of the label entity (such as unique identifier, label name, label partition, etc.) into the label narrow table according to the label name.
[0086] For example, if the two tag data are tag data age_level and tag data consum_level, two tag temporary tables are generated in the system. The corresponding tag temporary tables can be inserted into a tag narrow table.
[0087] For example, for the tag data age_level (tag 1) and the tag data consum_level (tag 2), for example, the structure and sample data of the tag temporary table corresponding to the tag data age_level are as follows:
[0088] ID age_level partition 0109a198d2888085fbc5147ecc426cb0 X 20240630 c4a7921bd84baefa0ab5ec4cafee90f3 Y 20240630
[0089] Where ID represents the unique identifier of the tag data. age_level represents the tag name of the tag data, which can indicate the tag value range or category of the attribute of the tag entity, for example, the range category of attribute 1. partition represents the tag partition, which can also indicate the time when the tag data was written.
[0090] For example, the structure and sample data of the temporary tag table corresponding to the tag data cons_level are shown in Table 3 as follows:
[0091] ID consum_level partition 0109a198d2888085fbc5147ecc426cb0 High tag value objects 20240630
[0092] Here, consum_level represents the label name of the label data, for example, the classification of the attribute or the label value range (such as the evaluation classification of attribute 2).
[0093] The data fields in the label temporary table of the above two label data can be summarized and written into the label narrow table. The structure and sample data of the label narrow table are shown in Table 4 as follows:
[0094] ID tag_name tag_value partition 0109a198d2888085fbc5147ecc426cb0 age_level X 20240630 c4a7921bd84baefa0ab5ec4cafee90f3 age_level Y 20240630 0109a198d2888085fbc5147ecc426cb0 consum_level High tag value objects 20240630
[0095] tag_name indicates the tag name, and tag_value indicates the tag value.
[0096] The above insert statement can insert all data fields of the label temporary table of the label entity into the label narrow table.
[0097] S132: Convert the narrow label table into a wide label table.
[0098] The data fields of the narrow label table may be format-converted to convert the narrow table data fields of the narrow label table into wide table data fields to obtain the wide label table.
[0099] In some implementations, a second processing statement for processing a wide label table can be generated based on the narrow label table. The narrow label table can then be converted to a wide label table using the second processing statement. The second processing statement includes at least one of a delete statement, a create statement, and an insert statement. The delete statement is used to delete data, the create statement is used to create data, and the insert statement is used to insert data. This section can be referenced with the fourth processing statement for the temporary label table described above and will not be further elaborated here.
[0100] Optionally, the narrow tag table can be converted into a wide tag table based on the tag information in the narrow tag table. For example, based on the unique identifier in the narrow tag table, the data fields (such as tag name, tag partition, etc.) belonging to the same unique identifier are summarized to obtain the wide tag table.
[0101] Optionally, the wide label table may contain multiple data fields, including a unique identifier ID, a label name, a label partition, etc., and this ID corresponds one-to-one to the ID of the label entity. The label name can be obtained based on the situation of the created label data. The drop statement and create statement in the second processing statement can refer to the fourth processing statement of the above-mentioned label temporary table, which is similar to the above and will not be repeated here. The insert statement is used to convert the data of the above-mentioned narrow label table into the data of the wide label table, and then input it into the specified wide label table.
[0102] For example, the narrow label table in Table 4 can be converted into a wide label table. The structure and example data of the wide label table are shown in Table 5 as follows:
[0103] id age_level consum_level partition 0109a198d2888085fbc5147ecc426cb0 X High tag value objects 20240630 c4a7921bd84baefa0ab5ec4cafee90f3 Y 20240630
[0104] S133: Integrate the entity view table and the label wide table to obtain a second label table.
[0105] The entity view table and the wide tag table can be combined to generate a second tag table. Optionally, the entity view table includes at least one of the following data fields: a unique identifier, at least one attribute, a data partition, etc. Optionally, the wide tag table includes at least one of the following data fields: a unique identifier, a tag name of at least one tag data item, a tag partition, etc.
[0106] In some embodiments, a summarized third processing statement can be generated based on the entity view table and the label wide table, and then the entity view table and the label wide table can be integrated using the third processing statement to obtain a second label table. The third processing statement includes at least one of a delete statement, a create statement, and an insert statement; wherein the delete statement is used to execute the data deletion task, the create statement is used to execute the data creation task, and the insert statement is used to execute the data insertion task. The drop statement and the create statement in the third processing statement can refer to the fourth processing statement of the above-mentioned label temporary table, which is similar to the above and will not be repeated here. The insert statement is used to integrate the data of the above-mentioned label wide table and the data of the entity view table to obtain a second label table.
[0107] Optionally, according to the unique identifiers in the entity view table and the label wide table, data fields belonging to the same unique identifier may be integrated to obtain a second label table.
[0108] For example, the structure and sample data of the entity view table of the tag entity are shown in Table 6 as follows:
[0109] ID age height partition 0109a198d2888085fbc5147ecc426cb0 20 170 20240630 c4a7921bd84baefa0ab5ec4cafee90f3 50 168 20240630 bba579a03e48610dbd54781ede52be43 30 175 20240630
[0110] Among them, age represents attribute 1 of the tag entity, height represents attribute 3 of the tag entity, ID represents the unique identifier of the tag entity, and partition represents the data partition corresponding to each attribute of the tag entity, that is, the time when the data was written.
[0111] By integrating the entity view table (such as Table 6) and the label wide table (such as Table 5) of the above label entity, a second label table can be obtained.
[0112] For example, the structure and sample data of the second tag table are as follows, see Table 7:
[0113] ID age height age_level consum_level partition 0109a198d2888085fbc5147ecc426cb0 20 170 X High tag value objects 20240630 c4a7921bd84baefa0ab5ec4cafee90f3 50 168 Y 20240630 bba579a03e48610dbd54781ede52be43 30 175 20240630
[0114] The second label table can be used as the final label result table. In the label result table, if there is no label value corresponding to a data field, the corresponding label value in the table is empty.
[0115] In some embodiments, after integrating the first label table and the entity view table to obtain the second label table, the second label table can be pushed to the acceleration engine. Among them, the acceleration engine can be an acceleration engine such as ElasticSearch, Doris, etc., and this application does not limit the acceleration engine. In the subsequent process, the acceleration engine can perform corresponding preset processing on the label data or the second label table. For example, the preset processing includes label preview, label query, label statistics, label circle group and other processing, which can obtain better processing performance.
[0116] The above solution can accelerate the processing of label data and improve processing efficiency by using the label narrow table, the label wide table and the entity view table to obtain the final second label table.
[0117] In some embodiments, after the above steps, the above-mentioned label data, label tables, label entities, etc. can be updated. Exemplarily, for example, updating the label data may include adding label data, modifying label data, deleting label data, etc. For adding, modifying, deleting, etc. of label data, it is necessary to maintain the status, processing statements, etc. of each node of the corresponding label data in the system. For adding or deleting, it is necessary to add or delete the corresponding processing of the label data under the corresponding label entity. For adding, you can refer to the description of the above embodiment for updating. For modifying, you can refer to the above embodiment for modifying the corresponding stages of each label data.
[0118] After updating the tag data of the tag entity, the tag data of the tag entity can be updated, and the steps of obtaining the tag data of the tag entity, processing the tag data to generate a first tag table, and integrating the first tag table with the entity view table to obtain a second tag table can be re-executed in response to the updated tag data, thereby obtaining an updated second tag table. In some application scenarios, if the selected raw data corresponding to the tag entity is updated, the step of obtaining the entity view table of the tag entity needs to be re-executed to obtain the updated entity view table.
[0119] Optionally, referring to the above embodiment, all the latest tag data of the tag entity and its tag metadata may be obtained to obtain an updated first tag table, that is, to obtain the latest temporary tag table.
[0120] Optionally, a corresponding first processing statement (such as an insert statement) may be generated according to the updated tag metadata to insert the data fields of the tag temporary table into the tag narrow table to obtain an updated tag narrow table.
[0121] Optionally, based on the updated narrow tag table (due to the change in the generated data field), an updated second processing statement may be generated to obtain an updated wide tag table.
[0122] Optionally, an updated third processing statement may be obtained based on the updated label width table and the latest entity view table, and the updated label width table and the latest entity view table may be integrated to obtain an updated second label table.
[0123] In some embodiments, the above-mentioned update includes at least one of the following: adding new label data, modifying label data, and deleting label data. Among them, the deleted label data is in a non-referenced state. That is, for deleting label data, before deleting the label data, it can be determined whether the label data to be deleted is in a referenced state, that is, whether the label data to be deleted is referenced by other label data. If it is still in a referenced state, the label data cannot be deleted. For modifying label data, modifying the label data requires synchronously modifying the referenced downstream label data. Downstream label data refers to other label data that references the label data. After modifying the label data, it is necessary to determine whether the current label data has other referenced label data (that is, downstream label data). If so, its downstream label data needs to be modified synchronously.
[0124] Optionally, the task of updating the label data is to run the SQL processing statements defined in the label data in sequence. The update of the label data can be divided into the following three categories: updating the label data according to the entire entity, updating a certain label data separately, and updating a certain label data and downstream label data.
[0125] That is, the above-mentioned updated label data includes at least one of the following: full update of label data, separate update of label data, separate update of label data and downstream label data. Optionally, full update of label data means updating according to all label data of the label entity. That is, if there is an update to the label data of the label entity, it is necessary to execute in the above manner, with the updated full label data as the updated label data. Optionally, separate update of label data means updating a certain updated label data separately. For example, if there is an update to a certain label data of the label entity, only the updated label data is updated separately. Optionally, if there is an update to the label data, the label data and downstream label data can be updated separately, and other label data may not be updated. Thus, the above steps are re-executed with this as the updated label data to obtain updated label tables, label metadata, etc.
[0126] In some implementations, the step of obtaining the first tag table may be updated in the following manner.
[0127] Optionally, in response to the updated tag data including the full update tag data, the full update tag data is processed according to a preset execution order to generate a first tag table. In this manner, each tag data item can be processed according to the preset execution order of the updated rights update tag data of the tag entity, thereby generating the first tag table. Furthermore, the corresponding steps described above can then be sequentially executed to obtain the narrow tag table, the wide tag table, the second tag table, and so on.
[0128] Optionally, in response to the updated label data including the individually updated label data, the individually updated label data is processed to generate a first label table corresponding to the individually updated label data. This method can process the updated label data individually according to a predetermined execution sequence to obtain the first label table corresponding to the updated label data. Furthermore, the aforementioned corresponding steps can then be sequentially executed to obtain the narrow label table, the wide label table, the second label table, and so on.
[0129] Optionally, in response to the updated label data including individually updated label data and downstream label data, the individually updated label data and downstream label data are processed in a preset execution order to generate a first label table. This method can obtain all downstream label data corresponding to a particular individually updated label data, and sequentially process the individually updated label data and downstream label data in the preset execution order to obtain the first label table corresponding to each updated label data. Furthermore, thereafter, the aforementioned corresponding steps can be sequentially executed to obtain a narrow label table, a wide label table, a second label table, and so on.
[0130] The above solution can update the above tables of tag data, improve the maintenance efficiency of each tag data, and can update a certain tag data separately, or update its downstream tag data, which can speed up the update efficiency.
[0131] In some implementations, after obtaining the updated second tag table, the updated second tag table can be pushed to the acceleration engine, which can significantly improve the speed of tag query and facilitate better performance in subsequent tag statistics, circle grouping, and other processing.
[0132] It is understandable that in the above method of the specific implementation method, the writing order of each step does not mean a strict execution order and constitutes any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0133] In some embodiments, the present application also provides a label data processing device for implementing the label data processing method of any of the above embodiments.
[0134] See also Figure 5 , Figure 5 2 is a schematic diagram of a structure of an embodiment of a label data processing device of the present invention. The label data processing device 20 includes an entity view module 21, a first label module 22 and a second label module 23. The modules are interconnected.
[0135] The entity view module 21 is used to obtain an entity view table for a tag entity. The entity view table contains a mapping relationship to the tag entity's original data. The first tag module 22 is used to obtain the tag entity's tag data, process the tag data, and generate a first tag table. The second tag module 23 is used to integrate the first tag table with the entity view table to generate a second tag table.
[0136] It should be noted that the label data processing device described in the above embodiment and the label data processing method described in the above embodiment are of the same concept, wherein the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here. In actual applications, the label data processing device described in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this application does not limit this.
[0137] It is understood that the tag data processing method of this application can be executed by a computer device, which can be any device with processing capabilities, such as a mobile device, a computer, a server, etc., and this application does not limit this. In some possible implementations, the tag data processing method can be implemented by a processor calling program data stored in a memory.
[0138] For the above embodiment, this application provides a computer device, see Figure 6 , Figure 6 1 is a schematic diagram of the structure of an embodiment of a computer device of the present application. The computer device 30 includes a memory 31 and a processor 32, wherein the memory 31 and the processor 32 are coupled to each other, the memory 31 stores program data, and the processor 32 is used to execute the program data to implement the steps of any embodiment of the tag data processing method described above.
[0139] In this embodiment, the processor 32 may also be referred to as a CPU (Central Processing Unit). The processor 32 may be an integrated circuit chip having signal processing capabilities. The processor 32 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor 32 may be any conventional processor.
[0140] The method of the above embodiment can be implemented in the form of a computer program, so this application proposes a computer readable storage medium, please refer to Figure 7 , Figure 7 The computer-readable storage medium 40 stores program data 41 that can be executed by a processor, and the program data 41 can be executed by the processor to implement the steps of any embodiment of the tag data processing method described above.
[0141] The computer-readable storage medium 40 in this embodiment can be a medium that can store program data 41, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or it can also be a server that stores the program data 41. The server can send the stored program data 41 to other devices for execution, or it can also execute the stored program data 41 by itself.
[0142] In some embodiments, the functions or modules included in the device provided in the above embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, this application will not go into details here.
[0143] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, this application will not go into details here.
[0144] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0145] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0146] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0147] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium, which is a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application.
[0148] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device. They can be concentrated on a single computing device or distributed across a network consisting of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a computer-readable storage medium and executed by the computing device, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0149] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for processing label data, characterized in that: include: Obtaining an entity view table of a tag entity, wherein the entity view table includes a mapping relationship to original data of the tag entity; Acquire label data of the label entity, process the label data, and generate a first label table; The first label table and the entity view table are integrated to obtain a second label table.
2. The method according to claim 1, characterized in that The integrating the first label table and the entity view table to obtain a second label table includes: Based on the tag metadata of the tag data, writing the data fields in the first tag table into a tag narrow table; Converting the narrow label table into a wide label table; The entity view table and the label wide table are integrated to obtain the second label table.
3. The method according to claim 2, characterized in that The step of writing the data fields in the first tag table into the tag narrow table based on the tag metadata of the tag data includes: generating a first processing statement based on the tag metadata of the tag data, and writing the data field in the first tag table into a tag narrow table using the first processing statement; The converting the narrow label table into the wide label table comprises: generating a second processing statement based on the narrow label table, and converting the narrow label table into a wide label table using the second processing statement; The step of integrating the entity view table and the label wide table to obtain the second label table includes: A third processing statement is generated based on the entity view table and the label width table, and the entity view table and the label width table are integrated using the third processing statement to obtain the second label table.
4. The method according to claim 3, characterized in that The first processing statement, the second processing statement, and the third processing statement each include at least one of a delete statement, a create statement, and an insert statement; wherein the delete statement is used to execute a data deletion task, the create statement is used to execute a data creation task, and the insert statement is used to execute a data insertion task; And / or, the label data is divided into multiple levels of label data, and there is a dependency relationship between label data at different levels; And / or, the tag metadata includes at least one of the following: tag name, tag temporary table name, tag temporary table partition, dependent tag, tag level, tag entity.
5. The method according to claim 1, wherein The acquiring the label data of the label entity, processing the label data, and generating a first label table includes: Creating tag data for the tag entity; Obtaining label processing rules; wherein the label processing rules include preset label fields in the label data that need to be processed; The label data is processed using the label processing rules to generate the first label table.
6. The method according to claim 5, characterized in that The step of processing the label data using the label processing rule to generate the first label table includes: Generating a fourth processing statement for each tag data using the tag processing rule; Based on the fourth processing statement, the label data is processed in a preset execution order to generate the first label table; And / or, the first label table at least includes a label temporary table, and each label data corresponds to a label temporary table; And / or, processing the label data using the label processing rule to generate the first label table further includes: After any tag data is processed, the tag result data of the tag data is written into the tag temporary table.
7. The method according to claim 1, characterized in that The entity view table of the tag entity is obtained, including: Obtaining original data of the tag entity; In response to the selection of the original data, determining the selected original data; Creating an entity view table for the tag entity using the selected original data; And / or, after integrating the first label table and the entity view table to obtain the second label table, the steps include: Push the second label table to the acceleration engine.
8. The method according to claim 1, characterized in that After integrating the first label table and the entity view table to obtain the second label table, the method includes: In response to updating the tag data of the tag entity, at least re-performing the steps of obtaining the tag data of the tag entity with the updated tag data, processing the tag data to generate a first tag table; and integrating the first tag table with the entity view table to obtain a second tag table, thereby obtaining an updated second tag table; The update includes at least one of the following: adding new tag data, modifying tag data, and deleting tag data; wherein, deleting tag data is in an unreferenced state, and modifying tag data requires synchronously modifying the referenced downstream tag data; and / or, the updated label data includes at least one of the following: full update of label data, individual update of label data, individual update of label data and downstream label data; The processing of the label data to generate a first label table includes: In response to the updated label data including the full-updated label data, processing the full-updated label data according to a preset execution order to generate a first label table; In response to the updated label data including the separately updated label data, processing the separately updated label data to generate a first label table; In response to the updated label data including the separately updated label data and the downstream label data, the separately updated label data and the downstream label data are processed according to a preset execution order to generate a first label table.
9. A computer device, characterized in that: The method comprises a memory and a processor coupled to each other, wherein the memory stores program data, and the processor is configured to execute the program data to implement the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that Program data capable of being executed by a processor is stored, and the program data is used to implement the steps of the method according to any one of claims 1 to 8.