Label data processing method and device, medium, equipment and program product

By constructing a label data processing method, obtaining and hierarchically storing the association relationship of the label source table, the problems of system bloat and link congestion in the label data management system in multiple business scenarios are solved, and fast, lightweight label data management and personalized output are achieved.

CN120611205APending Publication Date: 2025-09-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410260295.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The existing tag data management system leads to system bloat due to frequent addition and deletion of tags and multiple business scenarios, consumes a lot of manpower and material resources, and easily causes upstream and downstream interactions, leading to link blockage.

Method used

By obtaining the source table data information of the tag source table, the preset association relationship between business party information, entity category, tag information, data source information and metadata information is determined, and it is stored in layers in the data asset library to build weak association relationships, avoid being stored as table data, and achieve decoupling between business, entities, and tags.

Benefits of technology

It achieves rapid addition and reduction of labels and entity data at almost zero cost, avoids link congestion, reduces redundant data tables, implements lightweight management and personalized output, and improves data service efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611205A_ABST
    Figure CN120611205A_ABST
Patent Text Reader

Abstract

The invention discloses a label data processing method and device, a medium, equipment and a program product, and relates to the technical field of computers.The method comprises the steps that source table data information of a label source table of a service end is obtained, and the label source table comprises multiple pieces of entity identification information of a service party and label data corresponding to the entity identification information; the source table data information comprises data source information of a label source table, metadata information, entity categories to which multiple pieces of entity identification information belong and label information to which label data belong; determining a preset association relationship among the business party information, the entity category, the label information, the data source information and the metadata information corresponding to the label source table; and storing the preset association relationship, the business side information and the source table data information in a data asset library, wherein the business side information, the label information and the entity category in the data asset library are stored in a layered manner. According to the method and the device, customized table acquisition can be realized, redundant data tables are reduced, and link blockage is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to methods, devices, media, equipment, and program products for processing tag data. Background Art

[0002] With the development of internet technology, label data has become a crucial data asset for businesses to conduct market analysis, business promotion, and R&D. Existing label data management systems typically categorize data by label dimension and establish strong binding relationships between labels and user entities through tables. Each type of label is associated with a user entity and stored as an attribute dimension table based on the source table. When adding labels, fields must be expanded in the corresponding dimension table for the user entity. In the case of multiple business scenarios, a set of dimension tables must be maintained for each business scenario. This label data storage method is suitable for stable business scenarios where labels are infrequently added or removed. However, real-world applications often involve numerous user entities and label dimensions across multiple businesses, with frequent data additions and deletions. Using this dimension table approach requires generating base dimension tables for multiple entities across multiple businesses, as well as derived dimension tables customized for specific application scenarios. This bloats the entire system, necessitating high costs for table data maintenance and manpower. It also easily creates cross-talk between upstream and downstream processes, and when a single feature fails, the entire chain can be blocked. Summary of the Invention

[0003] This application provides a tag data processing method, apparatus, medium, device, and program product. The technical solution is as follows:

[0004] In one aspect, the present application provides a tag data processing method, the method comprising:

[0005] Obtaining source table data information of a tag source table on the business side, the tag source table including multiple entity identification information of the business side and tag data corresponding to the entity identification information, the source table data information including data source information, metadata information, entity categories to which the multiple entity identification information belongs, and tag information to which the tag data belongs;

[0006] Determining a preset association relationship among the business party information corresponding to the tag source table, the entity category, the tag information, the data source information, and the metadata information;

[0007] The preset association relationship, the business party information and the source table data information are stored in a data asset library, and the business party information, the tag information and the entity category are stored in a hierarchical manner in the data asset library.

[0008] On the other hand, the present application provides a tag data processing device, the device comprising:

[0009] A first acquisition module is configured to acquire source table data information of a tag source table on the business side, wherein the tag source table includes multiple entity identification information of the business side and tag data corresponding to the entity identification information. The source table data information includes data source information, metadata information, entity categories to which the multiple entity identification information belongs, and tag information to which the tag data belongs;

[0010] A first determining module is used to determine a preset association relationship among the business party information corresponding to the tag source table, the entity category, the tag information, the data source information, and the metadata information;

[0011] Storage module: used to store the preset association relationship, the business party information and the source table data information in the data asset library, where the business party information, the tag information and the entity category are stored in a hierarchical manner.

[0012] On the other hand, the present application provides a computer-readable storage medium, which stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the label data processing method as described above.

[0013] On the other hand, the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the label data processing method as described above.

[0014] On the other hand, the present application provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, implements the label data processing method as described above.

[0015] The tag data processing method, apparatus, medium, device, and program product provided in this application have the following technical effects:

[0016] The technical solution of the present application obtains the source table data information of the tag source table on the business side, and uses the data source information, metadata information, entity categories to which multiple entity identification information belongs, and tag information to which the tag data belongs, included in the source table data information, as the data basis for subsequent determination of the association relationship, and then obtains the preset association relationship between the business party information, entity category, tag information, data source information and metadata information corresponding to the tag source table, and then stores the preset association relationship, business party information and source table data information in the data asset library, and the business party information, tag information and entity category in the data asset library are stored in layers to construct a divergent weak association relationship between the business party information, entity category and tag information, and corresponding other information. There is no need to store it as table data on the ground, and only the logical relationship between each element is maintained in the data asset library to achieve mutual decoupling between business, entity, and tag, and can quickly add and reduce tags and entity data at almost zero cost to avoid link congestion, and can customize the output of personalized user tag data according to actual business needs, reduce redundant data tables, and achieve lightweight management and personalized output of user tag data.

[0017] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;

[0020] Figure 2 This is a flowchart of a tag data processing method provided in an embodiment of the present application;

[0021] Figure 3 This is a flowchart of another tag data processing method provided in an embodiment of the present application;

[0022] Figure 4 This is a flowchart of another tag data processing method provided in an embodiment of the present application;

[0023] Figure 5 This is a flowchart of another tag data processing method provided in an embodiment of the present application;

[0024] Figure 6This is a schematic diagram of the principle of a label entry process provided by an embodiment of the present application;

[0025] Figure 7 This is a schematic diagram of the structural framework of a tag data processing system provided in an embodiment of the present application;

[0026] Figure 8 This is a principle example diagram of a tag data processing flow provided in an embodiment of the present application;

[0027] Figure 9 This is a schematic diagram of the structural framework of a tag data processing device provided in an embodiment of the present application;

[0028] Figure 10 This is a schematic diagram of the hardware structure of a device for implementing a tag data processing method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. Examples of the embodiments are shown in the accompanying drawings, in which the same or similar numbers throughout represent the same or similar elements or elements with the same or similar functions.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0031] See also Figure 1 , Figure 1 This is a schematic diagram of an application environment provided by an embodiment of the present application, such as Figure 1As shown, the application environment may include at least a service end 01, a tag data middle station 02, and a request end 03. Each of the service end 01, the tag data middle station 02, and the request end 03 may include a terminal and a server. In actual applications, the request end 01, the service end 02, and the tag data middle station 03 may be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions thereon.

[0032] The server in the embodiments of the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0033] Specifically, cloud technology refers to a managed technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to enable data computing, storage, processing, and sharing. Cloud technology can be applied in a variety of fields, such as healthcare cloud, cloud IoT, cloud security, cloud education, cloud conferencing, artificial intelligence cloud services, cloud applications, cloud calling, and cloud social networking. Based on the cloud computing business model, cloud technology distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides resources is called a "cloud." To users, the resources in the cloud appear infinitely scalable and can be accessed at any time, used on demand, and expanded at any time, with a pay-per-use policy. Providers of cloud computing infrastructure establish a cloud computing resource pool (referred to as a cloud platform, commonly referred to as IaaS (Infrastructure as a Service)) and deploy various types of virtual resources within the resource pool for external clients to choose from. The cloud computing resource pool primarily includes computing devices (virtualized machines, including operating systems), storage devices, and network devices.

[0034] Specifically, the server mentioned above may include a physical device, which may specifically include a network communication submodule, a processor, a memory, etc., and may also include software running in the physical device, which may specifically include an application program, etc.

[0035] Specifically, the terminal may include physical devices such as smartphones, desktop computers, tablet computers, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, smart wearable devices, and vehicle-mounted terminal devices, and may also include software running in physical devices, such as applications.

[0036] In an embodiment of the present application, the business end 01 is used to provide a label source table, and the label data middle station 02 is used to receive the source table data information of the label source table and generate preset association relationships between businesses, entities, labels and other related data in the source table data information based on the source table data information, store businesses, entities, labels and other related data in the data information in layers, and store preset association relationships; the request end 03 is used to initiate a data service request to the label data middle station 02, so that the label data middle station 02 extracts entity identification information and label data from the relevant label source table based on the entity or label information carried by the data service request, combined with the preset association relationship, and then generates a target label data table to achieve personalized customized data output. It can be understood that the requester of the request end can be the business party of the business end or other third party.

[0037] Furthermore, it is understandable that Figure 1 What is shown is merely an application environment of a tag data processing method. The application environment may include more or fewer nodes, and this application does not impose any limitation thereto.

[0038] It is understandable that in the specific implementation of this application, data related to entities, businesses, and labels are involved. When the embodiments of this application are applied to specific products or technologies, it is necessary to obtain the permission or consent of the user and business parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0039] Most existing tag data management systems categorize tags according to the tag dimension into attribute tags, user interest tags, and user action tags. User entities are strongly bound to tags, and each tag type is then associated with the user entity and stored in an attribute dimension table. Adding tags requires expanding the corresponding fields in the user entity's dimension table. Multiple business scenarios require maintaining a set of dimension tables for each scenario. This strong binding of user entities and tags according to the tag dimension is suitable for stable business scenarios where tags are infrequently added or removed. However, a single enterprise tag data middleware can connect to over ten business departments, over thirty user entities, and thousands of tags, and this number continues to expand as more business scenarios are added. Therefore, in complex business contexts, adopting a dimension table approach can lead to significant system bloat. Generating base dimension tables for multiple entities across multiple businesses, along with derived dimension tables customized for specific application scenarios, and maintaining these multiple dimension tables simultaneously consumes significant resources and manpower. It can also easily cause cross-talk between upstream and downstream nodes, and if a single feature fails, the entire chain can be blocked.

[0040] In view of this, in order to solve at least one of the above problems, this application introduces a tag data processing method provided by this application, Figure 2 This is a flowchart of a tag data processing method provided by an embodiment of the present application. The present application provides method operation steps such as the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (for example, in a parallel processor or multi-threaded processing environment). Please refer to Figure 2 A tag data processing method provided in an embodiment of the present application may include the following steps S201-S205:

[0041] S201: Acquire source table data information of a tag source table on the business side.

[0042] Specifically, the service end can be a device located on a service provider that can access the tag data middle platform, used to provide service services to users and receive user-authorized registration information and operation data. Service providers can be, for example, video applications, news applications, browser applications, instant messaging applications, short video applications, sports applications, music applications, or input method applications. The tag source table is generated by the service provider based on the entity identification information of the user entity and the tag data of each user entity. It includes multiple entity identification information of the service provider and the tag data corresponding to the entity identification information. The entity identification information is the unique identifier of the user entity, which can be, for example, the user entity ID. The tag data includes the tag value of the user entity in at least one tag information dimension. The tag information dimension can be set based on actual business needs, such as demographic attributes, user activity, content preferences, commercialization preferences, user value, device attributes, geographic location, channel attributes, monthly active service duration, and service click-through rate. The tag data can exemplarily include monthly active service duration, user activity value, and content preference category. The tag source table records and stores the tag values ​​of each tag information possessed by each user entity belonging to the service provider. The tag source table may be a wide table including multiple columns (number fields). For example, the multiple columns may be entity category, entity identification information, tag information 1, tag information 2, ..., tag information n, etc.

[0043] Specifically, the source table data information includes the data source information of the tag source table, metadata information, the entity category to which multiple entity identification information belongs, and the tag information to which the tag data belongs. The entity category is used to indicate the category to which the entity identification information belongs, and can be based on the application classification of the business party, such as account category, etc. Exemplarily, the entity category can include social application user accounts, instant messaging application user accounts, video application user accounts, etc., and each entity category corresponds to an entity series, such as the QQ account entity category. The tag information is the unique identifier of the tag, which can be a tag name, tag ID, etc., such as population attribute ID, user activity ID, content preference ID, commercialization preference ID, user value ID, device attribute ID, geographic location ID, channel attribute ID, business monthly active duration ID, business click-through rate ID, etc. Each tag information comes from a category of tags. The data source information and metadata information are data that define and describe the information of the tag source table, including but not limited to data describing the data structure, constraints, index and other information of the table. In some embodiments, data source information can include metadata about the tag source table, such as header information, which can specifically include but is not limited to library table information, task information for producing the library table, filtering conditions, and access fields. Each access field can correspond to one or more tag information. Metadata information can also include metadata about the tag, which can specifically include but is not limited to the tag's primary key field, data format, tag usage, field processing logic, and an enumeration dictionary. The enumeration dictionary includes the enumeration values ​​for each tag information. In addition, source table data information can also include other basic information, such as the tag source table's task name and registrant, which can be set based on actual business needs.

[0044] In some cases, the tag data middle platform can receive source table data information entered by the business party. Accordingly, S201 includes: providing an access data registration page, the access data registration page includes an entry operation bar for source table data information; and obtaining source table data information based on the submitted data for the entry operation bar.

[0045] After the business party generates the label source table, it can quickly access the label in the label data middle platform according to actual business needs. The label data middle platform provides an access data registration page, including an entry operation column for each data in the source table data information, and stores each data submitted for the entry operation column to obtain the source table data information. In some embodiments, the access data registration page of the label data middle platform may include a basic information form, a data source information form, and a metadata information form, etc., to respectively enter basic information such as entity category and label information, data source information, and metadata information. In this way, providing a registration page allows each business party to enter labels directly through page operations, so that subsequent business parties can use the label data middle platform to manage and filter related data of their own labels, and can easily obtain label-related data authorized by other business parties, while improving the convenience of label data processing and expanding the range of available data to optimize business applications.

[0046] In some embodiments, obtaining source table data information based on submitted data for an input operation column may include: after obtaining the submitted data, performing data quality verification and review on the submitted data, and generating source table data information based on the submitted data for which verification and review results indicate passing.

[0047] refer to Figure 6 , in response to the data submission operation of the tag access registration on the access data registration page by the business party, the submitted data is checked and reviewed for data quality based on the preset data validation rules. The review can be manual or automatic, and the results of the data quality check can be used as a data reference for the review process. After the review is passed, the submitted data is entered into the tag data middle platform, and subsequent association relationship construction and hierarchical storage are performed, which can be displayed on the page of the data asset library; if it fails, a rejection message is generated, including the reason for rejection, etc., and sent to the business party so that the business party can modify the data based on the rejection information. After the data entry is completed, the tag data middle platform divides the storage space based on the business party. Each business party can manage its own tags and related data separately. For example, after filtering the entity category, all tag information under the entity category can be displayed.

[0048] In some embodiments, the tag data center can also periodically obtain source table data information for each tag source table stored by each business party to update the data in the data asset library, add, delete, or replace corresponding data. Accordingly, S201 includes: establishing a dependency relationship between the data asset library and the business end; and periodically obtaining source table data information from the business end based on the dependency relationship.

[0049] Dependencies are used to associate the data asset library with the tag source tables on the business side. The business side can authorize the tag data middle platform to have periodic access rights to the tag source tables it stores, so as to read the data items corresponding to the source table data information of the business side's tag source tables every other period, and then obtain the source table data information. In this way, the synchronization of the tags, entities and related data between the data middle platform and each business side is achieved through periodic access, which is beneficial to the data management and sharing of each business side, as well as the optimization of tag data services. In addition, multi-business, multi-entity, and multi-tag user entity management and information processing based on source table data information greatly reduces the access and management costs of user tag data. The tag access is lightweight, reducing the daily access cost of the original dimension table method to the hourly level.

[0050] S203: Determine the preset association relationship among the business party information, entity category, tag information, data source information and metadata information corresponding to the tag source table.

[0051] Specifically, the business party information is used to refer to the business party and can be a unique identifier of the business party or the business party interface, such as a business party ID, a business party interface ID, etc. The preset association relationship can be obtained by abstracting the logical structure and physical structure of the source table data information, and can include the mapping relationship between the business party information, entity category, and tag information, the mapping relationship between the entity category and each information item in the data source information and each information item in the metadata information, such as the retrieval field and library table information corresponding to entity category 1, and the mapping relationship between the tag information and each information item in the data source information and each information item in the metadata information, such as the retrieval field and field processing logic corresponding to the tag information.

[0052] S205: Storing the preset association relationship, business party information, and source table data information in the data asset library.

[0053] Specifically, the business party information, label information and entity categories in the data asset library are stored in layers. The data asset library only maintains the logical relationship between the business party information and the information in the source table data information, without generating table data for each entity, each business or each label dimension. The addition, deletion and modification of the business party information, label information and entity categories will not affect each other. The information items of other source table data information can also be stored separately, and they are only weakly associated with each other.

[0054] In this way, the source table data information includes the data source information, metadata information, entity categories to which multiple entity identification information belongs, and label information to which the label data belongs, of the label source table, as the data basis for the subsequent determination of the association relationship, and then the preset association relationship between the business party information, entity category, label information, data source information and metadata information corresponding to the label source table is obtained, and then the preset association relationship, business party information and source table data information are stored in the data asset library, and the business party information, label information and entity category in the data asset library are stored in layers to construct a divergent weak association relationship between the business party information, entity category and label information, and other corresponding information. There is no need to store it as table data, and only the logical relationship between each element is maintained in the data asset library to achieve decoupling between business, entity and label. It can quickly add and reduce labels and entity data at almost zero cost to avoid link congestion, and can customize the output of personalized user label data according to actual business needs, reduce redundant data tables, and achieve lightweight management and personalized output of user label data.

[0055] In some embodiments, reference Figure 7 The overall architecture of the system includes data layer, tag layer, entity layer, business layer and application layer. Among them, the tag data middle platform includes tag layer, entity layer and business layer, which are located in the data asset library, forming a BET (Business-Entity-Tag) hierarchical model to realize hierarchical storage; the data layer is used to provide underlying data support for tag production, and can provide raw data module and data aggregation module. The raw data module stores raw data of various data types, such as online buried data, business system data, offline manual data, third-party data and content consumption details data, etc. The data aggregation module stores aggregated data of various data types, such as user data, browsing data, content data, operation data, member data, geographic location data, commercial consumption data, etc.; the application layer is used to undertake the data asset library and the requester business, relying on the tag data middle platform capabilities to provide user tag data services to the requester. The application layer includes a variety of tag services, such as crowd insights, crowd marketing analysis, marketing delivery applications, general API services, crowd export APIs, portrait laboratories, offline information extraction services, shared data assets, etc.

[0056] It can be understood that the label data middle platform is connected to multiple business parties, and the data asset library includes multiple business party information, multiple entity categories and multiple label information. The multiple entity categories and multiple label information are respectively derived from the label source table of the corresponding business party. Accordingly, multiple business party information is stored in the business layer, multiple entity categories are stored in the entity layer, and multiple label information is stored in the label layer. There are pre-configured associations between multiple business party information, multiple entity categories and multiple label information, that is, they are constructed based on the preset associations determined by each label access. Through a hierarchical model, business, entity, and label related information and data are stored in a hierarchical manner. With the entity category of the user entity as the core, user label information and business party information are divergently associated, so that there is a weak correlation between data between layers without affecting each other. Business, entities, labels, and corresponding preset associations can be quickly added and reduced at zero cost. Data display and processing processes can also be freely expanded based on the associations between data between layers. It is suitable for different business scenarios with multiple businesses, multiple entities, and multiple labels, and improves data service efficiency.

[0057] In some embodiments, S205 may specifically include: hierarchically storing business party information, tag information, entity categories, data source information, and metadata information in a data asset library, and configuring the associations between business party information, tag information, and entity categories in the data asset library based on preset association relationships, as well as separately configuring the associations between tag information and data source information and metadata information, and between entity categories and data source information and metadata information. Specifically, the business party information, tag information, entity categories, database table information, data source information such as retrieval fields, and metadata information of the tag source table, as well as primary key fields, field processing logic, and enumeration dictionaries, are stored separately, and the associations or mappings between each piece of data information are configured so that when any one piece of information is specified, the other associated information items can be obtained, so that the corresponding data information can be displayed in the corresponding user entity tag directory and other pages, providing tag management capabilities and enabling rapid online and offline access of tags, entities, and related data.

[0058] The label data middle platform relies on the hierarchical storage data of the data asset library to realize the ability of rapid label service, and then combines different labels according to actual business needs to carry out various forms of application service, which can include general label data service and crowd label data service. General label data service refers to the ability of the label data middle platform to provide general API services related to user descriptions, and crowd label data service refers to the service capability derived from the crowd selection provided by the label data middle platform.

[0059] Accordingly, based on some or all of the above embodiments, in some embodiments, reference Figure 3 The method further includes S301-S307:

[0060] S301: Receive a data service request.

[0061] The data service request carries at least one specified tag information and / or at least one specified entity category of the requester. It can be a request for a general tag data service. The requester can select the tag information needed for the bill of lading based on actual business needs, that is, circle out one or more specified tag information in the data asset library, or can select one or more entity categories in the data asset library based on the entity category needed for the bill of lading, or can also circle out the required specified tag information and entity category at the same time to perform corresponding data screening.

[0062] S303: Determine target data source information and target metadata information corresponding to at least one specified tag information and / or at least one specified entity category in the data asset library based on a preset association relationship.

[0063] By associating the data information items based on the preset association relationships in the data asset library, the data source information items (library table information, fetch fields, etc.) and metadata information items (field processing logic, enumeration dictionary, etc.) associated or mapped with the specified tag information are determined, and the target data source information and target metadata information corresponding to each specified tag information are obtained. In addition, the data source information items and metadata information items associated or mapped with each specified entity category are determined, and the target data source information and target metadata information corresponding to each specified entity category are obtained. It can be understood that when the data service request carries both the specified entity category and the specified tag information, the target data source information and target metadata information finally obtained are the intersection of all target data source information corresponding to each specified tag information and all target data source information corresponding to each specified entity category, as well as the intersection of all target metadata information corresponding to each specified tag information and all target metadata information corresponding to each specified entity category; taking the number field as an example, the data service request carries specified tag information 1, specified tag information 2 and entity category 1, then the number field set of the target data source information finally obtained is the union of the first number field set corresponding to the specified tag information 1 and the second number field set corresponding to the specified tag information 2, and the intersection of the third number field set corresponding to entity category 1.

[0064] S305: Determine a matching tag source table based on the target data source information and the target metadata information, and obtain at least one specified tag information and / or entity identification information and tag data corresponding to at least one specified entity category from the matching tag source table.

[0065] An offline task can be created to determine the target data source information and target metadata information, that is, to determine one or several tag source tables required for the current data service request, as well as the data segments required in the above tag source tables, thereby extracting the entity identification information and tag data in the data segments.

[0066] S307: Generate a target label data table according to the entity identification information and label data corresponding to at least one specified label information and / or at least one specified entity category.

[0067] The target tag data table includes the acquired entity identification information and the tag data corresponding to the specified tag information under the entity identification information. Specifically, based on the created offline task, the selected and extracted data segments can be merged, compressed, and other operations can be performed to import the corresponding entity identification information and tag data from offline to online storage. Figure 8 , can be KV storage, such as importing redis storage, etc., to realize data landing. The generated target label data table can be, for example, a basic portrait table, an interest portrait table or a geographic portrait table, etc., and then can provide corresponding general label data services based on the feature data in the target label data table, such as user description query service, crowd determination service, phone replacement and retrieval service, etc. User description query service refers to querying all label values ​​of user entities, such as age, gender and preferences authorized by users, etc. Crowd determination service refers to judging whether user entities have interest preferences for certain types of content, such as preference for entertainment or preference for literary works, etc. It can also be applied to AB experiments for business launch. The same user can include entity identification information of two or more entity categories. Based on this association relationship, in scenarios such as user replacement of devices or mobile phones, the user's historical label data can be found for the newly registered entity identification information to facilitate information migration and solve the cold start problem of new device entities or account entities. In this way, combined with the hierarchical storage and relationship-configured data structure in the data asset library, not only can the rapid access to tag-related data be achieved, but the required data can also be quickly obtained to meet the data processing and service needs of personalized tag selection and entity selection of different requesters, and to customize the output of personalized user description data, reduce redundancy, save computing and storage costs, and improve management convenience.

[0068] In some embodiments, when there are multiple designated entity categories, the target tag data table may include a sub-table corresponding to each designated entity category to implement separate storage of tag data for multiple entity categories, which is beneficial for user screening of the same entity category.

[0069] It can be understood that the same entity identification information (such as user entity ID) can be logged in and used in the applications of multiple business parties, that is, the entity identification information of different business applications overlaps. If a label data table is maintained for each business party, when obtaining label data belonging to the same entity identification information, it is necessary to extract and merge data from the label data tables of multiple business parties; or in the application of the same business party, there may be entity identification information of multiple entity categories for login and operation. Accordingly, in the label source table of the business party, the same label information may correspond to entity identification information of different entity categories. Then, when obtaining the entity identification set under a certain entity category corresponding to the same label information, it is also necessary to search the label source tables of multiple business parties, or generate a dimension table for each entity category. However, by adopting the solution of the present application, by decoupling the business, entity category and label and establishing a weak association, it is possible to circle the label information required for customization through label selection, or circle the entity category required for customization through entity category, and then find the label source table corresponding to the entity or the label source table corresponding to the label based on the configured association relationship combined with the circled information, and then directly extract the corresponding data segment from the source table to generate the target label data table. There is no need to pre-generate and maintain multiple dimension tables. When the source table data changes, there is no need to wait for the dimension table to be generated. Only the preset association relationship in the data asset library needs to be updated to obtain the updated target label data table lightly and quickly. If the label information or entity category that needs to be obtained is adjusted later, it can be increased or decreased on the circled page. The whole process is very lightweight, convenient and fast to expand, and reduces redundant storage.

[0070] In some embodiments, based on the service capability of tag information selection or entity category selection, the tag data middle platform can automatically obtain and process the data of the tag source table based on business needs after storing the data in the data asset library. The authority of tag information selection or entity category selection can be authorized by the business party when the tag is accessed, such as selecting the option to support selection capability on the access data registration page. Figure 4 After S205, the method may further include S401-S405:

[0071] S401: Based on a preset association relationship and source table data information, obtain entity identification information, tag data, and at least one tag enumeration value corresponding to tag information from a tag source table.

[0072] The label data segment corresponding to the hit label information in the label source table can be automatically obtained, including all entity identification information of the hit label information, the label value of the entity identification information in the label information dimension and the enumeration dictionary of the label information. Label data is generated based on the label value of the entity identification information in the label information dimension. The enumeration dictionary includes at least one label enumeration value. For example, if the label information is an age label, the labels hit by a user entity ID include an age label and its label data (age label value) is 25-30 years old. The enumeration values ​​of each label in the enumeration dictionary of the age label are under 18 years old, 18-25 years old, 26-30 years old, 31-40 years old, 41-50 years old and over 51 years old.

[0073] S403: For each tag enumeration value, determine a preset mapping relationship between the tag enumeration value and the entity identification information based on the tag data.

[0074] S405: Generate label bitmap data based on the preset mapping relationship.

[0075] Under the same label information dimension, each label data (label value) belongs to a label enumeration value, such as the age of 25 belongs to the label enumeration value 18-25 years old. Accordingly, the label enumeration value corresponding to each entity identification information can be determined to abstract the above-mentioned preset mapping relationship, determine the entity identification information set corresponding to each label enumeration value, and then generate label bitmap data. The label bitmap data includes at least one label enumeration value and the entity identification information set corresponding to each label enumeration value. In this way, when the tag selection or entity selection capabilities are supported, after storing the preset association relationship, business party information and source table data information in the data asset library to realize tag entry, the offline import task can be automatically started to import all entity identification information of each tag information under the corresponding entity category into the preset storage engine, such as OLAP (On-Line Analysis Processing) storage engine, and perform ID mapping on the entity identification information corresponding to all tag enumeration values ​​under the tag information to obtain the preset mapping relationship, and then perform bitmap compression on the preset mapping relationship to obtain tag bitmap data. Each tag enumeration value is a collection of all hit entity identification information, and the tag bitmap data provides data intersection and difference capabilities to facilitate joint screening through multiple tags. The tag bitmap data here can be stored based on clickhouse storage media, refer to Figure 8, specifically CK storage. Accordingly, the requester can also manage the tags for its own crowd services in the tag data center, such as obtaining capabilities such as crowd selection services, operation category selection services, crowd determination services, and crowd intersection and difference services. Crowd selection services refer to obtaining entity identification information that matches one or more tag information, while operation category selection services refer to obtaining entity identification information that has one or more operation preferences.

[0076] In some embodiments, the tag information is used to represent atomic tags or non-atomic tags. Atomic tags refer to tags of original data dimensions, such as basic attributes of entities, such as gender, age, sex, etc., or they can be processed data, such as membership level, annual consumption amount, etc., which can be directly obtained by looking up the table without additional calculations. Non-atomic tags may include derived tags or combined tags, etc., which need to be obtained through atomic tags or other non-atomic tags based on preset calculation rules. Correspondingly, the metadata information also includes calculation caliber information of non-atomic tags, and the calculation caliber information is used to represent the various associated tag information and calculation functions required for the tag data calculation of non-atomic tags; in the case where the tag information represents a non-atomic tag, corresponding tag data calculations need to be performed in the process of reading the tag source table and obtaining the tag data, refer to Figure 5 , the method for obtaining the tag data corresponding to the tag information includes S501-S503:

[0077] S501: Obtaining calculation caliber information from metadata information corresponding to tag information based on a preset association relationship;

[0078] S503: Acquire label data of each associated label information from the label source table based on the calculation caliber information, and perform label data calculation on the label data of each associated label information based on the operation function to obtain label data of the label information.

[0079] For example, taking tag information as user value, its associated tag information includes the total consumption amount and forwarding rate. The operation function defines the operation rules for calculating the user value based on the total consumption amount and forwarding rate. In this way, when the tag is entered, the relationship between the business, entity, and tag is only weakly associated. The association between the entity and the tag only needs to record the association between the tag information and the entity category. In the case of the existence of calculation caliber information, it is only necessary to record the association between the entity category, tag information, and calculation caliber information. There is no need for actual binding. Based on actual business needs, the required tags or entities can be customized in the form of circle selection or automatic acquisition. The offline task can combine the tags and entities circled by the business to find the tag source table and calculation caliber corresponding to the tag or entity to extract, calculate and store entity identification information and tag data. The whole process is lightweight and convenient.

[0080] In summary, this application provides an entity-based data presentation and processing method for multi-business, multi-entity, and multi-label scenarios. By constructing a business-entity-label (BET) hierarchical model, one-stop management of the entire label life cycle is achieved, reducing business demand communication costs and development costs, and maximizing the value potential of user label data assets. In different business scenarios, with user entities as the core, user entities and labels can be divergently associated. Both data presentation and processing processes can be freely expanded based on the preset association relationship, greatly reducing the access cost and management cost of user label data, and realizing personalized output of user label data.

[0081] The present application also provides a label data processing device, such as Figure 9 As shown, Figure 9 The following is a schematic diagram showing the structure of a tag data processing device provided in an embodiment of the present application. The device may include the following modules:

[0082] The first acquisition module 10 is used to acquire source table data information of a tag source table on the business side. The tag source table includes multiple entity identification information of the business side and the tag data corresponding to the entity identification information. The source table data information includes data source information of the tag source table, metadata information, entity categories to which the multiple entity identification information belongs, and tag information to which the tag data belongs.

[0083] The first determination module 20 is used to determine the preset association relationship between the business party information, entity category, tag information, data source information and metadata information corresponding to the tag source table;

[0084] Storage module 30: used to store the preset association relationship, business party information and source table data information in the data asset library, where the business party information, tag information and entity categories are stored in a hierarchical manner.

[0085] In some embodiments, the storage module 30 can be specifically used to: hierarchically store business party information, label information, entity categories, data source information and metadata information in the data asset library, and configure the association between business party information, label information and entity categories in the data asset library based on preset association relationships, and separately configure the association between label information and data source information, metadata information, and the association between entity categories and data source information, metadata information.

[0086] In some embodiments, the data asset library includes multiple business party information, multiple entity categories and multiple tag information, the multiple business party information is stored in the business layer, the multiple entity categories are stored in the entity layer, and the multiple tag information is stored in the tag layer; there is a pre-configured association relationship between the multiple business party information, multiple entity categories and multiple tag information.

[0087] In some embodiments, the apparatus further comprises:

[0088] Request receiving module: used to receive a data service request, the data service request carries at least one specified tag information and / or at least one specified entity category of the requester;

[0089] A second determining module is used to determine target data source information and target metadata information corresponding to at least one specified tag information and / or at least one specified entity category in the data asset library based on a preset association relationship;

[0090] Target data acquisition module: used to determine a matching tag source table based on target data source information and target metadata information, and obtain at least one specified tag information and / or entity identification information and tag data corresponding to at least one specified entity category from the matching tag source table;

[0091] Data table generation module: used to generate a target label data table based on at least one specified label information and / or entity identification information and label data corresponding to at least one specified entity category.

[0092] In some embodiments, the apparatus further comprises:

[0093] A second acquisition module is configured to, after storing the preset association relationship, business party information, and source table data information in the data asset library, acquire entity identification information, tag data, and at least one tag enumeration value corresponding to the tag information from the tag source table based on the preset association relationship and source table data information;

[0094] A third determining module is configured to determine, for each tag enumeration value, a preset mapping relationship between the tag enumeration value and the entity identification information based on the tag data;

[0095] The bitmap generation module is used to generate label bitmap data based on a preset mapping relationship. The label bitmap data includes at least one label enumeration value and a set of entity identification information corresponding to each label enumeration value.

[0096] In some embodiments, the tag information is used to represent an atomic tag or a non-atomic tag, and the metadata information includes calculation caliber information of the non-atomic tag, and the calculation caliber information is used to represent each associated tag information and operation function required for the tag data calculation of the non-atomic tag; in the case of representing a non-atomic tag, the device further includes:

[0097] Calculation caliber acquisition module: used to obtain calculation caliber information from metadata information corresponding to tag information based on a preset association relationship;

[0098] The label data calculation module is used to obtain the label data of each associated label information from the label source table based on the calculation caliber information, and perform label data calculation on the label data of each associated label information based on the operation function to obtain the label data of the label information.

[0099] In some embodiments, the first determining module 20 may include:

[0100] Page display submodule: used to provide the access data registration page, which includes the input operation column of source table data information;

[0101] Data receiving submodule: used to obtain source table data information based on the submitted data for the input operation column.

[0102] In some embodiments, the first determining module 20 may include:

[0103] Dependency building submodule: used to build the dependency relationship between the data asset library and the business end. The dependency relationship is used to associate the data asset library with the tag source table of the business end.

[0104] Periodic acquisition submodule: used to periodically obtain source table data information from the business end based on dependency relationships.

[0105] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0106] An embodiment of the present application provides a computer device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a label data processing method provided in the above method embodiment.

[0107] Figure 10 The hardware structure diagram of a device for implementing a tag data processing method provided in an embodiment of the present application is shown. The device may participate in or include the apparatus or system provided in an embodiment of the present application. Figure 10As shown, the device 10 may include one or more (illustrated as 1002a, 1002b, ..., 1002n in the figure) processors 1002 (the processor 1002 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 10 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 10 More or fewer components than shown, or with Figure 10 Different configurations shown.

[0108] It should be noted that the one or more processors 1002 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the device 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0109] The memory 1004 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method in the embodiment of the present application. The processor 1002 executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, that is, to implement the above-mentioned label data processing method. The memory 1004 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1004 may further include a memory remotely located relative to the processor 1002, and these remote memories may be connected to the device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0110] Transmission device 1006 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of device 10. In one embodiment, transmission device 1006 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 1006 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0111] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of device 10 (or mobile device).

[0112] An embodiment of the present application also provides a computer-readable storage medium, which can be set in a server to store at least one instruction or at least one program related to a label data processing method in a method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement a label data processing method provided in the above method embodiment.

[0113] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0114] An embodiment of the present invention further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a tag data processing method provided in any of the above-mentioned optional embodiments.

[0115] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0116] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.

[0117] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0118] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A tag data processing method, characterized in that: The method comprises: Obtaining source table data information of a tag source table on the business side, the tag source table including multiple entity identification information of the business side and tag data corresponding to the entity identification information, the source table data information including data source information, metadata information, entity categories to which the multiple entity identification information belongs, and tag information to which the tag data belongs; Determining a preset association relationship among the business party information corresponding to the tag source table, the entity category, the tag information, the data source information, and the metadata information; The preset association relationship, the business party information and the source table data information are stored in a data asset library, and the business party information, the tag information and the entity category are stored in a hierarchical manner in the data asset library.

2. The method according to claim 1, characterized in that Storing the preset association relationship, the business party information, and the source table data information in the data asset library includes: The business party information, the label information, the entity category, the data source information and the metadata information are stored in the data asset library in a hierarchical manner, and the association between the business party information, the label information and the entity category is configured in the data asset library based on the preset association relationship, and the association between the label information and the data source information and the metadata information, and the association between the entity category and the data source information and the metadata information are configured respectively.

3. The method according to claim 1, characterized in that The data asset library includes a plurality of business party information, a plurality of entity categories and a plurality of tag information, wherein the plurality of business party information is stored in the business layer, the plurality of entity categories are stored in the entity layer, and the plurality of tag information is stored in the tag layer; There is a preconfigured association relationship between the multiple business party information, the multiple entity categories, and the multiple tag information.

4. The method according to claim 1, wherein The method further comprises: receiving a data service request, wherein the data service request carries at least one specified tag information and / or at least one specified entity category of the requesting party; Determine, based on the preset association relationship, target data source information and target metadata information corresponding to the at least one specified tag information and / or the at least one specified entity category in the data asset library; Determining a matching tag source table based on the target data source information and the target metadata information, and obtaining the at least one specified tag information and / or entity identification information and tag data corresponding to the at least one specified entity category from the matching tag source table; A target label data table is generated according to the entity identification information and label data corresponding to the at least one specified label information and / or at least one specified entity category.

5. The method according to claim 1, wherein After storing the preset association relationship, the business party information, and the source table data information in the data asset library, the method further includes: Based on the preset association relationship and the source table data information, acquiring entity identification information, label data and at least one label enumeration value corresponding to the label information from the label source table; For each of the tag enumeration values, determining a preset mapping relationship between the tag enumeration value and the entity identification information based on the tag data; Tag bitmap data is generated based on the preset mapping relationship, where the tag bitmap data includes the at least one tag enumeration value and a set of entity identification information corresponding to each tag enumeration value.

6. The method according to any one of claims 1 to 5, characterized in that The tag information is used to represent an atomic tag or a non-atomic tag, and the metadata information includes calculation caliber information of the non-atomic tag, and the calculation caliber information is used to represent each associated tag information and operation function required for tag data calculation of the non-atomic tag; In the case of representing the non-atomic tag, a method for obtaining the tag data corresponding to the tag information includes: Acquiring the calculation caliber information from the metadata information corresponding to the tag information based on the preset association relationship; The label data of each associated label information is acquired from the label source table based on the calculation caliber information, and label data calculation is performed on the label data of each associated label information based on the operation function to obtain the label data of the label information.

7. The method according to any one of claims 1 to 5, characterized in that The source table data information of the tag source table of the business end is obtained, including: Providing an access data registration page, the access data registration page including an input operation column for the source table data information; The source table data information is obtained based on the submitted data for the input operation column.

8. The method according to any one of claims 1 to 5, characterized in that The source table data information of the tag source table of the business end is obtained, including: Constructing a dependency relationship between the data asset library and the business end, wherein the dependency relationship is used to associate the data asset library with the tag source table of the business end; The source table data information is periodically obtained from the business end based on the dependency relationship.

9. A tag data processing device, characterized in that: The device comprises: A first acquisition module is configured to acquire source table data information of a tag source table on the business side, wherein the tag source table includes multiple entity identification information of the business side and tag data corresponding to the entity identification information. The source table data information includes data source information, metadata information, entity categories to which the multiple entity identification information belongs, and tag information to which the tag data belongs; A first determining module is used to determine a preset association relationship among the business party information corresponding to the tag source table, the entity category, the tag information, the data source information, and the metadata information; Storage module: used to store the preset association relationship, the business party information and the source table data information in the data asset library, where the business party information, the tag information and the entity category are stored in a hierarchical manner.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the label data processing method according to any one of claims 1 to 8.

11. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the label data processing method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the tag data processing method according to any one of claims 1 to 8 is implemented.