Data labeling method, device and storage medium
Through the data labeling method, target labels are generated, which solves the problem of large amount of data and lack of correlation in business analysis of electronic devices, improves analysis efficiency and simplifies the data processing process.
Patent Information
- Application Number
- CN202210311708.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-03-28
AI Technical Summary
In the prior art, electronic devices face the problems of large amount of data and lack of correlation when conducting business analysis, resulting in complex analysis process and low processing efficiency.
Through the data labeling method, target tags are generated, and the association relationship between data is established. The tag device is used to analyze, generate and push data, and the tag processing is realized.
It improves the efficiency of business analysis, simplifies the data processing process, realizes the correlation between data, and reduces the cost of data connection.
Smart Images

Figure CN114840519B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to a method, device and storage medium for data labeling. Background Art
[0002] In various fields such as work, production, and management, there is a large amount of underlying data, such as source data, business data, log data, reported data, and third-party log data. Electronic devices can perform corresponding business analysis based on this underlying data.
[0003] After acquiring various underlying data, electronic devices conduct business analysis through the business operations side. However, not only is this underlying data voluminous, but the data itself does not directly reflect the business functions it can implement or correspond to. Furthermore, due to data sources, as well as technical, management, and institutional factors, different underlying data lacks correlation and cannot be connected. Therefore, during the analysis process, the business operations side needs to analyze each piece of data involved in the business analysis to determine the correlations between different underlying data, as well as the correlations between the business operations side and the underlying data, and then conduct business analysis on the underlying data. This shows that when electronic devices conduct business analysis based on existing underlying data, the analysis process is complex and processing efficiency is low. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a data labeling method, device and storage medium for solving the problems in the prior art of complex business analysis process and low business processing efficiency when electronic devices use data for business analysis.
[0005] A first aspect of an embodiment of the present application provides a method for data labeling, which includes: obtaining target data; parsing the target data to determine a target tag for the target data, where the target tag is used to represent the association relationship between the target data and a data analysis business, and / or the association relationship between different target data; and adding a target tag to the target data.
[0006] In combination with the first aspect, in a first possible implementation of the first aspect, parsing the target data and determining the target tag of the target data includes: parsing the target data in combination with preset tag content to generate executable tag code for the target data, wherein the tag code includes the tag source code, tag operation code and tag meaning code of the target data; generating a tag value of the target object based on the tag code; generating a target tag for the target data based on the tag value, a super primary key value and the tag code, wherein the super primary key value is generated based on the target data.
[0007] In combination with the first possible implementation method of the first aspect, in the second possible implementation method of the first aspect, the method also includes: establishing a correspondence between the data source, label type and label meaning in the label content; parsing the target data in combination with the preset label content, generating the label source code of the target data according to the data source, generating the label operation code of the target data according to the label type, and generating the label meaning code of the target data according to the label meaning.
[0008] In combination with the first aspect, in a third possible implementation of the first aspect, the method further includes: dividing the data source corresponding to the target object into a subject domain table and a dimension domain table according to subject data and dimension data, and presetting the table names, field types, and field meanings of the subject domain table and the dimension domain table.
[0009] In combination with the first aspect, in a fourth possible implementation of the first aspect, the label type includes: at least one of a rule type label, a statistical type label, and a machine learning type label.
[0010] In combination with the first aspect, in a fifth possible implementation of the first aspect, the super primary key value is generated based on the target data, including: based on the target data of the target object, searching for all primary keys used to identify the target object; merging all the primary keys, setting edge conditions, and constructing a graph model of the target object; and generating the super primary key of the target object based on the data set in the graph model.
[0011] In combination with the first aspect, in the sixth possible implementation of the first aspect, after generating the target tag of the target object based on the tag value, super primary key value and tag code of the target object, the method also includes: storing the target tag of the target data according to a preset storage method; and / or pushing the target tag of the target data to the corresponding business device.
[0012] In combination with the first aspect, in the seventh possible implementation of the first aspect, the method also includes: recording the generation process of the target label by means of burying points, obtaining data information of the target label, the data information including the data source, label type and label meaning of the target label; using the data source in the data information as the input value and the label type and label meaning as the output values, generating label information of the target label; and storing the label information.
[0013] A second aspect of an embodiment of the present application provides a data labeling device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in any one of the first aspects are implemented.
[0014] A third aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the first aspects are implemented.
[0015] The beneficial effects of the embodiments of the present application compared with the prior art are: the technical solution of the present application dynamically labels the corresponding data associated with the target object by defining rules and label operations and rule parsing, and generates target labels, so that the business equipment can perform business analysis by reading the required target label content, and then execute the corresponding business. By labeling the data, real-time labeling, and automatic labeling, the problem of complex business analysis process and low business processing efficiency when electronic devices use data for business analysis in the prior art is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 This is a schematic diagram of an application scenario of a data labeling method provided in an embodiment of the present application;
[0018] Figure 2 This is a structural block diagram of a label device provided in an embodiment of the present application;
[0019] Figure 3 This is a flow chart of a data labeling method provided in an embodiment of the present application;
[0020] Figure 4 This is a flowchart of a data marking process provided by an embodiment of the present application;
[0021] Figure 5 This is a schematic diagram of a data labeling device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0023] In order to illustrate the technical solution described in this application, specific embodiments are provided below.
[0024] In various fields such as work, production, and management, there is a large amount of underlying data, such as source data, business data, log data, reporting data, and third-party log data. Among them, source data includes system organizational structure data and device metadata. Business data includes user data (such as user basic information), behavioral data (such as user browsing history, historical operation records, etc.), product data (such as product name, product category, product reviews, inventory, etc.), and data generated during remote operation and maintenance. Log data includes data generated when the computer operating system or application software is running. Log data is helpful for future system maintenance. Reporting data includes data generated by the terminal when uploading or importing information. Third-party log data includes data generated when the system is connected to third-party devices and is associated with intrusion detection systems (IDS) or intrusion prevention systems (IPS). Based on the above underlying data, electronic devices can perform corresponding business analysis. For example, based on data related to the company, industry, equipment type, etc. of the production equipment, combined with IP address positioning technology and threat intelligence data, the user information corresponding to the production equipment can be identified, and then the production equipment can be used as the analysis object, and a portrait of the production equipment can be established in combination with the user information.
[0025] However, this underlying data is not only voluminous, but also incapable of directly reflecting the business functions it can achieve or correspond to. For example, in the field of industrial control security, when an electronic device's network is attacked, the device will record data related to the attack. This data serves as underlying data for the business operations side. However, after the business operations side receives this data, they cannot understand the specific content reflected by the data or the business functions it can achieve.
[0026] In addition, due to the different data sources, technology, management, and institutional reasons of different underlying data, the same type of data is stored and maintained independently in different underlying data layers, becoming isolated from each other. Alternatively, different underlying data layers interpret and define data in unique ways, resulting in some of the same type of data being assigned different meanings. As a result, after acquiring this data, electronic devices are unable to understand the lack of relevance between each piece of data and are unable to connect them. For example, e-commerce shopping application A (hereinafter referred to as application A) and e-commerce shopping application B (hereinafter referred to as application B) can provide similar shopping functions, but they use different systems and business models to record business data. In other words, the way application A records business data is different from the way application B records business data. When an electronic device needs to integrate and analyze the business data of application A and application B, it needs to repeatedly build, maintain, and analyze the data of these two applications, resulting in a complex data processing process and high data connection costs.
[0027] To summarize, currently, electronic devices, after acquiring various underlying data, conduct business analysis through the business operations side. However, not only is this underlying data voluminous, but the data itself does not directly reflect the business functions it can implement or correspond to. Furthermore, due to data sources, as well as technical, management, and institutional factors, different underlying data lacks correlation and cannot be connected. Therefore, during the analysis process, the business operations side needs to analyze each piece of data involved in the business analysis to determine the correlations between different underlying data, as well as the correlations between the business operations side and the underlying data, and then conduct business analysis on the underlying data. This shows that when electronic devices conduct business analysis based on existing underlying data, the analysis process is complex and processing efficiency is low.
[0028] Based on this, an embodiment of the present application provides a method for data labeling, which can add target tags to data, thereby establishing an association relationship between data and generating target tags, so that business equipment can execute corresponding business by reading the required target tag content.
[0029] Figure 1 This is a schematic diagram of an application scenario of a data labeling method proposed in an embodiment of the present application, such as Figure 1 As shown, this scenario involves data source devices, label devices, and business devices.
[0030] The data source device is used to provide data sources for the tag devices. The data in the data source device usually includes the above-mentioned source data, business data, log data, reported data, and third-party log data and other underlying data.
[0031] The labeling device is used to label data and obtain data labels. In one example, Figure 2 As shown, the label device includes a data metadata management module, a label definition module, a label rule parsing module, a super primary key graph generation module, a real-time label calculation module and a label lineage module.
[0032] The data metadata management module has preset subject domain tables and dimension domain tables.
[0033] The subject domain table represents the relationship between data topics and data. It not only comprehensively defines table names, field types, and field meanings, but also includes a large amount of underlying data from various sources. The table name can also be called the subject name, and the field type indicates the information contained within a specific subject. The field meaning indicates the specific meaning of the information within that subject.
[0034] After acquiring this underlying data, the data source device first extracts, cleans, transforms, and loads (Extract-Transform-Load, ETL) it, dividing it into data of different themes. Each theme is divided according to the type of underlying data generated, such as attacks suffered by the system, problems encountered, or shopping orders and payment records generated on e-commerce platforms. For example, when the theme is a system attack, the corresponding data is all data generated when the system was attacked. The data metadata management module obtains data of different themes and stores it in the theme domain table according to the corresponding relationship between the themes.
[0035] Dimensional domain tables represent the relationships between data dimensions and data. They not only fully define the table name, field types, and field meanings, but also include a large amount of underlying data from various sources. Table names can also be called topic names, while field types indicate the information contained within a specific topic. Field meanings indicate the specific meaning of the information within that topic.
[0036] After acquiring this underlying data, the data source device first extracts, cleans, transforms, and loads (ETL) it, dividing it into data of different dimensions. Each dimension is divided according to relatively fixed and unchanging entity objects, such as equipment and organizational structures that are not related to deeper business operations. For example, a dimension representing basic equipment information would correspond to data containing all basic information about the device, including model, manufacturer, and age. The data metadata management module, upon acquiring data of different dimensions, stores it in a dimensional domain table based on the corresponding relationships between the dimensions.
[0037] The label definition module is preset with label content, which includes the correspondence between labels and application scenarios, label types, and the corresponding positions and data fields of the data used in the subject domain table and dimension domain table when each label type is applied, as well as the meaning of the fields.
[0038] For example, the application scenarios can be in different fields, including the steel industry, rare metals, smart parks, smart transportation, petroleum and petrochemicals, and smart pipeline networks. Each application scenario is preset with a corresponding tag type. The data source used by each tag type is the table name, field type, and field meaning of the subject domain table and dimension domain table in the data metadata management module.
[0039] The above-mentioned label types include: at least one of rule-type labels, statistical-type labels, and machine learning-type labels. Among them, rule-type labels are used to define the detailed operation rules, data source tables, data fields, and the meaning of the data in the fields during the labeling process. Statistical-type labels are used to define the data statistical cycle, statistical methods, comparison rules, and filtering conditions during the labeling process, as well as the data source tables, data fields, and the meaning of the data in the fields in data metadata management. Machine learning-type labels are used to define the data source tables and data fields used for training data, the data source tables, data fields, and the meaning of the data in the fields in test data metadata management, the specific algorithms required, and the label meanings corresponding to the algorithm labels.
[0040] The super primary key generation module is used to generate the super primary key of the target object.
[0041] The tag rule parsing module is used to parse the tag rules and generate executable tag codes for the target data.
[0042] The label operation module is used to generate the label value of the target object according to the label code.
[0043] The label lineage module is used to record label lineage, monitor the entire process from reading data to generating labels, and trace data quality.
[0044] The tag value push module is used to push the above tag value, tag code and super primary key value to the corresponding business device.
[0045] Business equipment, used to execute corresponding business according to target tags.
[0046] The following is an illustrative description of the data labeling method provided in the embodiment of the present application based on the labeling device provided in the above embodiment.
[0047] Figure 3This is a flow chart of a data labeling method provided by an embodiment of the present application, which specifically includes the following steps S1-S6.
[0048] S1. The tag device obtains target data.
[0049] In this embodiment, the target data is the data of the target object in a specific application scenario. In different application scenarios, personnel, equipment, or factory buildings can all serve as target objects for data labeling. For example, if the application scenario is to collect statistics on personnel income, the target object is the relevant personnel, and the data source of the target data can be the National Bureau of Statistics or the natural person tax system.
[0050] S2. The tag device parses the target data and generates an executable tag code for the target data according to the preset tag content. The tag code includes a tag source code, a tag operation code, and a tag meaning code of the target data.
[0051] The label definition module based on the label device is pre-set with the label application scenario, label type, and the data source table, data field and field meaning used under each label type. Therefore, the label device can determine the data source, label type and label meaning corresponding to the target data by parsing the target data.
[0052] The data source is to determine the specific location of the target data in the subject domain table and dimension domain table in the data metadata management module.
[0053] The tag type is the tag type corresponding to the target data.
[0054] The label means the meaning of each data in the subject domain table and dimension domain table.
[0055] For example, the target data is analyzed based on the application scenario of the target data and the preset tag content in the tag definition module.
[0056] When the tag type corresponding to the target data parsed by the tagging device is a rule-based tag, the tagging device further determines the target data's data source information and generates a tag source code. The tag source code indicates the target data's location in the subject and dimension domain tables. Furthermore, the tagging device uses the rule-based tag to determine the operation rule corresponding to the target data and generates a tag operation code. Furthermore, the tagging device further determines the tag meaning within the target data's data source information and generates a tag meaning code, which represents the tag's specific meaning.
[0057] Similarly, when the tag type corresponding to the target data parsed by the tagging device is a statistical type tag, the tagging device further determines the target data's data source information and generates a tag source code. The tag source code is used to indicate the target data's acquisition location in the subject domain table and dimension domain table. Furthermore, the tagging device uses the rule type tag to determine the statistical period, statistical period, statistical method, and comparison rule corresponding to the target data and generates a tag operation code. Furthermore, the tagging device further determines the tag meaning within the target data's data source information and generates a tag meaning code, which represents the tag's specific meaning.
[0058] Similarly, when the label type corresponding to the target data parsed by the labeling device is a machine learning type label, the labeling device further determines the target data source information and generates a label source code. The label source code is used to indicate the acquisition location of the target data in the subject domain table and the dimension domain table. Furthermore, the labeling device determines the specific algorithm used for the target data based on the rule type label and generates label operation code for algorithm training and verification. Furthermore, the labeling device further determines the label meaning in the target data source information and generates a label meaning code. This label code is used to refer to the specific meaning of the label.
[0059] S3. The tag device generates a tag value for the target object based on the tag code.
[0060] The tagging device generates a computation task based on the computation rules corresponding to the tag computation code, as well as the tag source code, tag computation code, and tag meaning code. By executing this computation task, the tagging device obtains the tag value of the target object. This tag value represents the tagging result for the target object.
[0061] S4. The tag device generates a super primary key value of the target object through the target data.
[0062] In one example, the super primary key value is generated based on the target data, specifically including the following:
[0063] First, the labeling device searches for all primary keys used to identify the target object in the subject domain table and dimension domain table based on the target data.
[0064] In this embodiment, the primary key is the collection of all data associated with the target object, and each primary key can be uniquely associated with the target object. For example, when a person is the object, their ID number, mobile phone number, medical insurance card number, etc. can all be used as the primary key of the person.
[0065] Specifically, when a person is the target object, their age is labeled. In different scenarios, such as travel scenarios, the data source includes travel information from ride-hailing records or bus cards, or medical scenarios, the data source includes medical insurance card information stored in the medical system. Therefore, for the same target person, in different scenarios, when the data source is different, a unique primary key is required to represent the target person.
[0066] Then, the tagging device merges all primary keys, sets edge conditions, and builds a graph model (i.e., a graph model) of the target object.
[0067] The labeling device identifies all primary keys related to the target object, that is, the business and data tables of each source / end, and merges them after identification to form a graphic model with points and edges as boundary conditions.
[0068] For example, when targeting a person, data such as their ID number or mobile phone number is stored in the personnel file system. Registering via mobile phone number on an e-commerce platform also generates a unique account. The medical system also contains a corresponding medical insurance card number and ID number. The tagging device then merges these three types of data. If the ID number and mobile phone number in the personnel file system are associated with the mobile phone number in the e-commerce platform, then the object corresponding to the mobile phone number belongs to the same person. At the same time, if the medical system also has the same ID number as the one in the personnel file system, then the object corresponding to the ID number belongs to the same person. Therefore, the tagging device merges all the data related to the target object to construct a graph model of the target object. All primary keys related to the person are the points of the graph model, and the boundaries formed by all primary key information are the edges of the graph model. The overlapping information is represented by a single primary key.
[0069] After the tag device collects all data information related to the target object and constructs a graphic model, the graphic model includes all data set information representing the target object.
[0070] Finally, the labeling device generates a super primary key for the target object based on the data set in the above graph-text model.
[0071] Based on the above data set information, a unique super primary key that can represent the target object is generated. Based on the above example, in specific applications, in the personnel file system, the super primary key represents the object's ID number or mobile phone number and other data information; on the e-commerce operation platform, it represents the object's unique account number; in the medical system, it represents the object's medical insurance card number or ID number.
[0072] In this embodiment, the super primary key can uniquely identify a labeled object and can be applied in any scenario. Ultimately, the generated super primary key value and the above-mentioned label value are pushed to the business device together.
[0073] S5. The tag device generates a target tag for the target data based on the tag value, the super primary key value, and the tag code.
[0074] The super primary key value is generated by the labeling device according to the subject domain table and dimension domain table corresponding to the target data, and the super primary key value is used to uniquely label the target object.
[0075] The tag code is generated by the tag device through parsing the target data. The tag code is used to indicate the data source, data field, and meaning of the data in the field of the target object.
[0076] S6. The tag device adds a target tag to the target data.
[0077] In the data compliance method provided in the embodiment of the present application, after the target object is labeled as a user or device from various dimensions, subsequent device recommendations, personnel divisions, and target customer divisions can be performed based on the label information. For example, for a certain type of personnel object, it can be determined what type of target customer it is. In the industrial control security industry, it is more about device profiling to determine whether the device is a high-risk device, a high-quality device, or a general device.
[0078] The following is an exemplary description of the process of determining the super primary key value involved in the above embodiment.
[0079] When labeling, the ID of the same object is different for different business needs, different business equipment, and different data sources. Therefore, it is necessary to summarize the data tables in different business data to form a graphic model, thereby generating a super primary key. The super primary key is used to uniquely label the target object, and the super primary key is associated with all data information of the target object in all business devices.
[0080] In one implementation of this embodiment, the method further includes:
[0081] During the target label generation process, the generation process is monitored using a tracking method. The entire process from reading data to generating labels is monitored, and the read data metadata information, read data information, generated label information, and generated rules are recorded. The recorded data metadata information and read data information are used as input values, the generated rules are used as edge conditions, and the generated label information is used as output values. They are stored in the graph database for data traceability and fault tracing.
[0082] In an embodiment of the present application, after generating a target tag of the target object according to the tag value, super primary key value, and tag code of the target object, the method further includes:
[0083] The target tag of the target object is stored according to a preset storage method, such as storing the generated tag value and super primary key value in the database in a non-relational manner with a key-value storage structure. The target tag of the target object can also be pushed to the corresponding business device, such as pushing personnel tags to personnel statistics business devices, and pushing equipment tags to equipment portrait business devices, or both at the same time. It can be selected according to specific needs, and no specific restrictions are made in this embodiment.
[0084] Figure 4 An embodiment of the present application provides a data labeling process flow chart, which corresponds to steps S1-S6 of the above-mentioned data labeling method and the associated data source devices and service devices.
[0085] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0086] Figure 5 FIG. 1 is a schematic diagram of a data tagging device provided in an embodiment of the present application. Figure 5 As shown, the data labeling device 4 of this embodiment includes: a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40, such as a data labeling program. When the processor 40 executes the computer program 42, it implements the steps of the aforementioned data labeling method embodiments. Alternatively, when the processor 40 executes the computer program 42, it implements the functions of the modules / units in the aforementioned device embodiments.
[0087] For example, the computer program 42 may be divided into one or more modules / units, which are stored in the memory 41 and executed by the processor 40 to implement the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 42 in the data tagging device 4.
[0088] The data tagging device 4 can be a computing device such as a tablet computer, a tablet computer, a desktop computer, a notebook computer, a PDA, a cloud server, etc. The data tagging device can include, but is not limited to, a processor 40 and a memory 41. It will be understood by those skilled in the art that Figure 4It is only an example of the data labeling device 4 and does not constitute a limitation of the data labeling device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the data labeling device may also include input and output devices, network access devices, buses, etc.
[0089] The processor 40 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0090] The memory 41 can be an internal storage unit of the data tagging device 4, such as a hard disk or memory of the data tagging device 4. The memory 41 can also be an external storage device of the data tagging device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the data tagging device 4. Furthermore, the memory 41 can also include both the internal storage unit of the data tagging device 4 and an external storage device. The memory 41 is used to store the computer program and other programs and data required by the data tagging device. The memory 41 can also be used to temporarily store data that has been output or is about to be output.
[0091] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0092] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0093] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0094] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0095] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0096] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0097] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, which can also be completed by hardware related to computer program instructions. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0098] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A data labeling method, applied to a labeling device, characterized in that: The method comprises: Get target data; Parsing the target data to determine a target tag for the target data, where the target tag is used to indicate an association between the target data and a data analysis service, and / or an association between different target data; adding the target tag to the target data; The step of parsing the target data and determining a target tag for the target data includes: The target data is parsed in combination with preset tag content to generate executable tag code for the target data, wherein the tag code includes a tag source code, a tag operation code, and a tag meaning code of the target data. The tag source code is used to indicate the acquisition location of the target data in the subject domain table and the dimension domain table. The tag operation code is used to indicate the operation rule corresponding to the target data. The subject domain table is used to indicate the corresponding relationship between the subject of the data and the data, and the dimension domain table is used to indicate the corresponding relationship between the dimension of the data and the data. Generate a label value of the target object according to the label code; generating a target tag for target data according to the tag value, the super primary key value, and the tag code, wherein the super primary key value is generated according to the target data; The super primary key value is generated according to the target data, including: According to the target data of the target object, searching for all primary keys for identifying the target object; Merge all the primary keys, set edge conditions, and build a graph model of the target object; Generate a super primary key of the target object based on the data set in the graph model.
2. The method according to claim 1, characterized in that The method further comprises: Establish the corresponding relationship between data source, tag type and tag meaning in tag content; The target data is parsed in combination with the preset tag content, the tag source code of the target data is generated according to the data source, the tag operation code of the target data is generated according to the tag type, and the tag meaning code of the target data is generated according to the tag meaning.
3. The method according to claim 1, characterized in that The method further comprises: The data source corresponding to the target object is divided into a subject domain table and a dimension domain table according to subject data and dimension data, and the table names, field types and field meanings of the subject domain table and the dimension domain table are preset.
4. The method according to claim 2, characterized in that The tag types include: At least one of a rule-type label, a statistics-type label, and a machine learning-type label.
5. The method according to claim 1, wherein After generating a target tag for target data according to the tag value, the super primary key value, and the tag code, the method further includes: Storing the target tag of the target data according to a preset storage method; and / or, Push the target tag of the target data to the corresponding business device.
6. The method according to claim 2, characterized in that The method further comprises: Record the generation process of the target tag by embedding data points to obtain data information of the target tag, including the data source, tag type, and tag meaning of the target tag; Generate tag information of the target tag using the data source in the data information as input value and the tag type and tag meaning as output value; The tag information is stored.
7. A data labeling device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Service data processing method and device, electronic equipment and storage medium
CN111724063A
Method and device for constructing metadata tag library
CN113360496A