Annotating method, annotating device, computer device and storage medium
By obtaining the annotation configuration table and data dictionary to automate the data annotation process, the problem of low efficiency of manual annotation is solved, and efficient and accurate data annotation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2026-03-24
AI Technical Summary
The existing technology of manually annotating data increases the workload of developers and is prone to errors, resulting in low data annotation efficiency.
By obtaining a pre-configured annotation configuration table, the system filters the original data from the original database based on the annotation category and validity information, extracts the initial annotation content using a data dictionary, performs valid filtering, obtains the target annotation content, and automatically annotates the data.
It automates data annotation, reduces manual workload, and improves the efficiency and accuracy of data annotation.
Smart Images

Figure CN115034187B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an annotation method, annotation device, computer equipment, and storage medium. Background Technology
[0002] Currently, when processing data, it is necessary to annotate all tables and fields of big data. However, manually annotating data not only increases the workload of developers, but is also prone to errors, thus affecting the efficiency of data annotation. Summary of the Invention
[0003] The main objective of this application is to provide a data annotation method, data annotation apparatus, computer device, and storage medium, which aims to improve the efficiency of data annotation and reduce the workload of developers.
[0004] To achieve the above objectives, a first aspect of this application provides an annotation method, the method comprising:
[0005] Obtain a pre-configured annotation configuration table; wherein the annotation configuration table includes annotation categories and effectiveness information, and the effectiveness information is used to characterize the occurrence of the annotation category;
[0006] Based on the annotation category and the activation information, raw data is filtered from a preset raw database; wherein, the raw data is the data to be annotated, and the raw data includes the raw data name;
[0007] Initial annotation content is extracted from a preset data dictionary based on the original data name;
[0008] The initial annotation content is then subjected to a validity screening process to obtain the validity screening results;
[0009] If the legal filtering result is legal, then the target annotation content is filtered out from the initial annotation content;
[0010] The original data is annotated according to the target annotation content to obtain the target data.
[0011] In some embodiments, the step of filtering raw data from a preset raw database based on the annotation category and the validity information includes:
[0012] Based on the effective information, the target category is filtered from the annotation category;
[0013] The original data is filtered from the original database according to the target category.
[0014] In some embodiments, the raw data includes: a raw data table and raw fields of the raw data table; the step of filtering the raw data from the raw database according to the target category includes:
[0015] The original data table is obtained from the original database according to the target category;
[0016] The original data table is traversed to obtain the original fields.
[0017] In some embodiments, the original data includes: an original data table and the original fields of the original data table; the step of annotating the original data according to the target annotation content to obtain the target data includes:
[0018] Obtain the original annotation content of the original data;
[0019] The original annotation content is validated to obtain the validation result;
[0020] If the verification result indicates that the original annotation content is empty or contains garbled characters, then the original field is annotated according to the target annotation content to obtain the target data.
[0021] In some embodiments, after extracting initial annotation content from a preset data dictionary based on the original data name, the method further includes:
[0022] Updating the data dictionary specifically includes:
[0023] Obtain the original annotation content whose verification result is normal to obtain the legal annotation content;
[0024] If the data dictionary does not contain the legal annotation content, the original data name and the legal annotation content are mapped to obtain a mapping relationship;
[0025] The mapping relationship is stored in the data dictionary to update the data dictionary.
[0026] In some embodiments, after performing a legality screening on the initial annotation content and obtaining the legality screening results, the method further includes:
[0027] If the valid filtering result is invalid, alternative annotation content is generated according to the preset annotation rules. The alternative annotation content is used to annotate the original data to obtain the target data.
[0028] In some embodiments, the step of generating alternative annotation content according to preset annotation rules if the valid screening result is invalid includes:
[0029] If the valid filtering result is invalid, the alternative annotation content is generated based on the random number and the field name of the original field.
[0030] To achieve the above objectives, a second aspect of this application provides an annotation apparatus, the apparatus comprising:
[0031] An acquisition module is used to acquire a pre-configured annotation configuration table; wherein, the annotation configuration table includes annotation categories and effectiveness information, and the effectiveness information is used to characterize the occurrence of the annotation category;
[0032] The data filtering module is used to filter raw data from a preset raw database according to the annotation category and the effectiveness information; wherein, the raw data is data to be annotated, and the raw data includes the raw data name;
[0033] The extraction module is used to extract initial annotation content from a preset data dictionary based on the original data name;
[0034] The valid filtering module is used to perform valid filtering on the initial annotation content to obtain valid filtering results;
[0035] An annotation filtering module is used to filter out target annotation content from the initial annotation content if the valid filtering result is valid;
[0036] The annotation module is used to annotate the original data according to the target annotation content to obtain the target data.
[0037] To achieve the above objectives, a third aspect of this application provides a computer device, the computer device including a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for implementing communication between the processor and the memory, wherein the program, when executed by the processor, implements the method described in the first aspect.
[0038] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium for computer-readable storage, wherein the storage medium stores one or more programs that can be executed by one or more processors to implement the method described in the first aspect.
[0039] The annotation method, apparatus, computer device, and storage medium proposed in this application filter raw data from a pre-defined raw database based on annotation category and validity information, obtain the original data name of the raw data, extract initial annotation content from a data dictionary based on the original data name, and perform a validity screening on the initial annotation content to obtain a validity screening result. If the validity screening result is valid, the initial annotation content is determined as the target annotation content, and the raw data is annotated according to the target annotation content to obtain the target data. Therefore, this application realizes automated data annotation, which not only reduces the workload of manual labor but also improves the efficiency of data annotation. Attached Figure Description
[0040] Figure 1 This is a flowchart of the annotation method provided in the embodiments of this application;
[0041] Figure 2 yes Figure 1 The flowchart of step S102 in the document;
[0042] Figure 3 yes Figure 2 The flowchart of step S202 in the document;
[0043] Figure 4 yes Figure 1 The flowchart of step S106 in the process;
[0044] Figure 5 This is a flowchart of an annotation method provided in another embodiment of this application;
[0045] Figure 6 This is a block diagram of the annotation device provided in the embodiments of this application;
[0046] Figure 7 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0048] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0050] First, let's analyze some of the terms used in this application:
[0051] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0052] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0053] Comments: In Python, comments are used to describe the functionality of code or functions. Simply put, they are a record of the code, making it easier to understand when modifying or debugging it later, and also helping others understand the meaning of the code. There are typically three types of comments: single-line comments, multi-line comments, and Chinese character encoding declaration comments.
[0054] A data dictionary is a user-accessible catalog that records metadata about a database and application. An active data dictionary is one whose contents are automatically updated by the DBMS when the database or application structure is modified. A passive data dictionary requires manual updates when modifications are made. A data dictionary is a collection of descriptions of data objects or items in a data model, which is beneficial for programmers and others who need to refer to it. The first step in analyzing a system of user-exchanged objects is to identify each object and its relationships with other objects. This process is called data modeling, and it produces an object relationship diagram. After each data object and item is given a descriptive name, its relationships are described (or become part of a structure of potential descriptive relationships), then the data type is described (e.g., text, image, or binary value), all possible predefined values are listed, and simple textual descriptions are provided. This collection, organized into a book for reference, is called a data dictionary.
[0055] Configuration tables: Configuration tables are generally divided into built-in configuration tables and user-defined configuration tables. Built-in configuration tables include app.config, web.config, Settings.settings, etc. User-defined configuration tables typically store configuration information in XML files or the registry. This configuration information generally includes program settings, records runtime information, and information about controls (such as position and style).
[0056] Traversal: Traversal refers to visiting each node in a tree (or graph) sequentially along a search path. The operations performed on each node depend on the specific application problem; these operations might include checking or updating a node's value. Different traversal methods result in different order of node visits. Traversal is one of the most important operations on binary trees and forms the basis for other operations on binary trees. The concept of traversal also applies to multi-element collections, such as arrays.
[0057] Key-value distributed storage systems offer fast query speeds, large data storage capacities, and high concurrency support, making them ideal for primary key queries. However, they cannot handle complex conditional queries. When supplemented with a Real-Time Search Engine for complex conditional and full-text searches, they can replace relational databases like MySQL, which have lower concurrency performance, achieving high concurrency and high performance while saving tens of times the number of servers. Key-value distributed storage systems, such as MemcacheDB and TokyoTyrant, easily complete high-speed queries under tens of thousands of concurrent connections. In contrast, MySQL typically crashes with only a few hundred concurrent connections.
[0058] With the advent of the big data era, data needs to be annotated for subsequent use and management. However, in related technologies, manually annotating data is not only more labor-intensive due to the massive amount of data, but also prone to errors and reduces the efficiency of data annotation.
[0059] Based on this, embodiments of this application provide an annotation method, annotation apparatus, computer equipment, and storage medium, which not only reduce manual workload but also improve the efficiency of data annotation by automating data annotation.
[0060] This application provides an annotation method, annotation apparatus, computer device, and storage medium, which are specifically described through the following embodiments. First, the annotation method in the embodiments of this application is described.
[0061] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0062] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0063] The annotation method provided in this application relates to the field of artificial intelligence technology. The annotation method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the annotation method, but is not limited to the above forms.
[0064] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0065] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data for the proper functioning of the embodiments of this application obtained.
[0066] Figure 1 This is an optional flowchart of the annotation method provided in the embodiments of this application. Figure 1 The annotation method may include, but is not limited to, steps S101 to S106.
[0067] Step S101: Obtain a pre-configured annotation configuration table; wherein, the annotation configuration table includes annotation categories and effectiveness information, and the effectiveness information is used to characterize the occurrence of annotation categories;
[0068] Step S102: Select raw data from the preset raw database according to the annotation category and validity information; wherein, the raw data is the data to be annotated, and the raw data includes the raw data name;
[0069] Step S103: Extract initial annotation content from the preset data dictionary based on the original data name;
[0070] Step S104: Perform a valid filtering on the initial annotation content to obtain the valid filtering results;
[0071] Step S105: If the valid filtering result is valid, then filter out the target annotation content from the initial annotation content;
[0072] Step S106: Annotate the original data according to the target annotation content to obtain the target data.
[0073] Steps S101 to S106 of this embodiment involve obtaining the annotation category and effectiveness information from the annotation configuration table, and filtering the original data from the original database based on the annotation category and effectiveness information to determine the original data that needs to be annotated. After determining the original data to be annotated, the data dictionary is retrieved based on the original data, and initial annotation content is extracted from the data dictionary based on the original data name. The initial annotation content is then subjected to a valid filtering process to obtain a valid filtering result. The valid initial annotation content obtained from the valid filtering result is used as the target annotation content, and the original data is annotated based on the target annotation content to obtain the target annotation. Therefore, by automatically extracting the original data and automatically obtaining the target annotation content that matches the original data, and then annotating the original data with the target annotation content, automated data annotation is achieved, which not only reduces the workload of manual work but also improves the efficiency of data annotation.
[0074] In step S101 of some embodiments, the user pre-determines which data needs to be annotated in the pre-configured annotation configuration table, and then sets the corresponding content on the annotation configuration table according to the required annotation content. For example, if the annotation configuration table is as shown in Table 1:
[0075]
[0076] Table 1
[0077] Table 1 shows that the comment categories specify the categories that need to be commented. In the effectiveness information, "Y" indicates effective, while "N" indicates ineffective. From the effectiveness information in Table 1, we know that the comment category corresponding to number 1 is effective, while the others (numbers 2, 3, and 4) are ineffective. Therefore, the comment category that needs to be commented is the comment category corresponding to number 1, which is "Full Database Table Comments and Field Comments." Based on the comment category, it is clear that comments need to be added to all tables and table fields in the entire database.
[0078] It should be noted that the annotation configuration table is App_cfg. Therefore, when reading App_cfg, the annotation categories and effectiveness information are read one by one in ascending order of the sequence number to determine the effectiveness information corresponding to each annotation category. For example, as shown in Table 1, the effectiveness information "Y" corresponds to sequence number 1, and the annotation category for sequence number 1 is "full database table annotation and field annotation". In this case, it is necessary to obtain the original data tables of the entire database and annotate the original data tables and original fields.
[0079] In step S102 of some embodiments, the annotation category and validity information can be used to determine which annotation category needs to be annotated, so that the data to be annotated can be extracted from the original database according to the annotation category and validity information to obtain the original data. The preset original database pre-stores the original data matching the annotation category. Therefore, retrieving the original data from the original database by the annotation category and validity information to determine the data to be annotated simplifies the extraction of the original data.
[0080] In step S103 of some embodiments, after determining the original data, i.e., determining the data to be annotated, the data dictionary is retrieved based on the original data. The data dictionary pre-stores matching information between the original data name and the annotation content. Therefore, after determining the original data, the system automatically reads the corresponding annotation content from the annotation database, matches the annotation content with the original data name to obtain matching information, and constructs the data dictionary based on the matching information. Thus, the original data name of the original data is obtained, and the corresponding annotation content is extracted from the data dictionary based on the original data name to obtain the initial annotation content. The initial annotation content can be empty, garbled text, non-text annotation content, or compliant annotation content. Compliant annotation content is non-empty, non-garbled text, and Chinese annotation content. Therefore, the obtained initial annotation content cannot be directly used as the target annotation content for the original data; further filtering of the initial annotation data is required to obtain the target annotation content.
[0081] It should be noted that the data dictionary uses a key-value data format to store matching information between the original data name and the comment content. Furthermore, the key-value data format allows for quick retrieval of the corresponding initial comment content based on the original data name, thereby improving the efficiency of finding the initial comment content.
[0082] In step S104 of some embodiments, since the initial annotation content extracted directly from the data dictionary cannot be directly used as the target annotation content, it is necessary to perform a legal screening of the initial annotation content to obtain a legal screening result. The legal screening mainly verifies the initial annotation content for non-empty, non-garbled characters, and Chinese annotations to determine whether the initial annotation content is empty, garbled, or contains Chinese annotations, thus obtaining a legal screening result. Based on the legal screening result, it is determined whether the initial annotation content can be used as the target annotation content. The legal screening result includes: legal and illegal. In this embodiment, a legal screening result means that the initial annotation content is non-empty, non-garbled, and contains Chinese annotations; if the legal screening result is illegal, it means that the initial annotation content is empty, garbled, or does not contain Chinese annotations.
[0083] In step S105 of some embodiments, the legal filtering results include: legal and illegal. Since illegal initial annotation content cannot be used as the target annotation content of the original data, it is necessary to filter out legal initial annotation content based on the legal filtering results. Therefore, obtaining legal initial annotation content as the target annotation content, the target annotation content is used as the annotation content of the original data to provide subsequent annotation operations, making the annotation of the original data more accurate.
[0084] In step S106 of some embodiments, if the target annotation content is determined, the original data is annotated according to the target annotation content to obtain the target data, thereby realizing automated data annotation, reducing manual workload and improving the efficiency of data annotation.
[0085] Please see Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S202:
[0086] Step S201: Filter the target category from the annotation category based on the effective information;
[0087] Step S202: Filter the original data from the original database according to the target category.
[0088] In step S201 of some embodiments, the annotation categories and activation information on the annotation configuration table correspond, and the activation information is used to determine which annotation categories are target categories. Specifically, the annotation categories whose activation information is active are the target categories, that is, the categories that need to be annotated are determined.
[0089] For example, referring to Table 1, if the effective information is "N", then there is no need to obtain the corresponding annotation category. If the effective information is "Y" and the annotation category is "full database table annotations and field annotations", then the target category is determined to be "full database table annotations and field annotations".
[0090] In step S202 of some embodiments, the original database stores original data that matches the annotation category. Therefore, original data matching the target category can be filtered from the original database based on the target category. Thus, filtering original data from the original database by target category simplifies the acquisition of original data, thereby determining the data to be annotated.
[0091] For example, if the target annotation is "full database table annotation and field annotation", then the original data of the entire database is retrieved from the original database, that is, the original data tables of the entire database are retrieved. If the original data tables of the entire database include: TableName1, TableName2, TableName3, and TableName4, then four original data tables are retrieved from the original database. Therefore, by determining the target category, the corresponding original data can be extracted from the original database according to the target category, making the retrieval of original data simple.
[0092] Please see Figure 3 In some embodiments, the original data includes: an original data table and the original fields of the original data table, and step S202 may include, but is not limited to, steps S301 and S302.
[0093] Step S301: Obtain the original data table from the original database according to the target category;
[0094] Step S302: Traverse the original data table to obtain the original fields.
[0095] In step S301 of some embodiments, the raw data includes raw data tables and raw fields of the raw data tables. When the content of the target annotation is a table annotation and a field annotation, and the raw database stores raw data tables that match the annotation category, the raw data tables are retrieved from the raw database according to the target annotation to determine which raw data table needs to be annotated.
[0096] In step S302 of some embodiments, since the original fields are fields in the original data table, after obtaining the original data table, the original fields in the original data table are extracted by traversing the original data table, which makes the operation of obtaining the original fields simple.
[0097] It should be noted that if the target category is "to annotate tables and table fields in the specified list", then the original data tables corresponding to the specified tables in the list are retrieved from the original database according to the target category. Each original data table is then traversed to obtain the original fields. The original data name includes both the table name and the field name. Therefore, the initial annotation content for annotating the original data table is obtained from the data dictionary based on the table name, and the initial annotation content for annotating the original fields is obtained from the data dictionary based on the field names.
[0098] For example, if the list specifies tables as TableName1, TableName2, TableName3, and TableName4, then the original data tables are obtained as TableName1, TableName2, TableName3, and TableName4, respectively. Each of these tables is then iterated through to obtain its original fields. For instance, if iterating through TableName1 yields Field1, Field2, and Field3, then the original fields obtained are Field1, Field2, and Field3. Therefore, by obtaining the original data table through the target category and then iterating through each field in the original data table to obtain the original fields, the original data table and its original fields are obtained completely.
[0099] Please see Figure 4 In some embodiments, the original data includes: an original data table and the original fields of the original data table; step S106 may include, but is not limited to, steps S401, S402, and S403:
[0100] Step S401: Obtain the original annotation content of the original data;
[0101] Step S402: Perform a validity check on the original annotation content and obtain the check result;
[0102] Step S403: If the verification result is that the original annotation content is empty or the original annotation content contains garbled characters, then annotate the original field according to the target annotation content to obtain the target data.
[0103] In step S401 of some embodiments, when annotating the original data, the original annotation content of the original data is first obtained, so as to determine whether the original data needs to be annotated again based on the original annotation content, so that the original field annotations in the original data table are more accurate and comprehensive.
[0104] For example, please refer to Table 2:
[0105] Serial Number field name Field comments 1 Field1 insurance policy 2 Field2 3 Field3 %%##$
[0106] Table 2
[0107] As shown in Table 2, after traversing each original data table, the original data table is determined to be tableName1. TableName1 is then traversed to obtain the original fields. The original comment content corresponding to the field name "Field1" is "insurance policy", the original comment content of the field name "Field2" is empty, and the original comment content of the field name "Field3" is "%%##$".
[0108] In step S402 of some embodiments, after obtaining the original annotation content of the original data, the original annotation content is validated to obtain a validation result. The validation of the original annotation content mainly checks whether the original annotation content is empty or contains garbled characters, to obtain a validation result. The validation result includes: the original annotation content is empty, the original annotation content contains garbled characters, or it is normal. Therefore, the validation result determines whether the original annotation content is a normal annotation, and based on the validation result, it is determined whether the annotation content of the original data needs to be updated. If the original data is being annotated for the first time, the validation result is that the original annotation content is empty, that is, there is no original annotation content, and therefore the original data needs to be annotated.
[0109] For example, after determining the original annotation content, if the original annotation content is "insurance policy", the corresponding verification result is normal; if the original annotation content is empty, the verification result is that the original annotation content is empty; if the original annotation content is "%%##$", the verification result is that the original annotation content contains garbled characters. Therefore, by verifying the original annotation content to obtain the verification result, it is possible to determine whether to update the original annotation content of the original data based on the verification result.
[0110] In step S403 of some embodiments, if the verification result is that the original annotation content is empty or the original annotation content contains garbled characters, it means that the original annotation content of the original data is invalid. In this case, it is necessary to obtain the annotation content again, that is, to obtain the target annotation content, and to annotate the original data according to the target annotation content to obtain the target data, so as to update the annotation content on the original data to obtain the target data.
[0111] For example, if the original comment content for the field named "Field2" is empty, and the target comment content obtained from the data dictionary is "Insurance Type", then a database command is used to annotate the original data. The database command is "commenton tableName1.field2 is "Insurance Type"", thus adding the target comment content through the database command to obtain the target data.
[0112] In some embodiments, after step 104, the annotation method further includes:
[0113] If the valid screening result is invalid, alternative annotation content is generated according to the preset annotation rules. The alternative annotation content is used to annotate the original data to obtain the target data.
[0114] It should be noted that after obtaining the initial annotation content from the data dictionary, the initial annotation content undergoes a validity screening process to obtain valid results. If the valid screening result is invalid, it indicates that the initial annotation content obtained from the data dictionary is empty, contains garbled characters, or is not Chinese. In this case, alternative annotation content is generated according to the user's pre-set annotation rules. This alternative annotation content is then used to annotate the original data to obtain the target data. Therefore, generating alternative annotation content according to preset annotation rules and then using this alternative annotation content to annotate the original data to obtain the target data makes the data annotation more comprehensive.
[0115] For example, if the original comment content for the field named "Field3" contains garbled characters, and the initial comment content corresponding to "Field3" is found to be empty in the data dictionary, that is, no initial comment content is found, then alternative comment content is generated according to the preset comment rules, and the alternative comment content is annotated on the original data according to the database command to obtain the target data.
[0116] In some embodiments, if the valid filtering result is invalid, alternative annotation content is generated according to preset annotation rules, which may include, but is not limited to, the following steps:
[0117] If a valid filter result is invalid, generate alternative comment content based on the random number and the field name of the original field.
[0118] It should be noted that when a valid filter result is invalid, alternative annotation content is generated according to the preset annotation rules. The preset annotation rules are a random number and the field name of the original field. This makes it easy to generate alternative annotation content.
[0119] For example, if the preset annotation rule is "random two-digit number + field name", and if the initial annotation content for "Field3" in the data dictionary is empty, then the alternative annotation content "02_Field3" is generated based on the random two-digit number and the original field name. The database command "comment on tableName1.Field3 is "02_Field3"" is then used to annotate "Field3" with "02_Field3". Therefore, by generating alternative annotation content according to the preset annotation rule and annotating the original data based on these alternative annotations to obtain the target data, the data annotation becomes more comprehensive.
[0120] In some embodiments, after step 103, the annotation method further includes:
[0121] Update the data dictionary.
[0122] It should be noted that, in order to improve the comprehensiveness of the data dictionary and to obtain more comprehensive annotation content from it, the data dictionary needs to be updated regularly. This means that the data dictionary is updated based on the original annotation content of the original data to continuously improve it.
[0123] Specifically, when the data dictionary needs to be updated, the entire database's original data tables are retrieved, and each table is traversed to obtain the original fields. Then, the original comments for these fields are iterated over, and their validity is validated to obtain the validation result. If the validation result is valid, it is checked whether the original comment content exists in the data dictionary corresponding to the field name. If it does, no update is needed. If the original comment content does not exist in the data dictionary, it is added to the data dictionary, and the original comment content is matched with the field names in the data dictionary to update the data dictionary. Therefore, by updating the data dictionary based on the original comment content, the data dictionary is improved, making the search for initial comment content from the data dictionary more comprehensive and resulting in more complete data annotations.
[0124] Please see Figure 5 In some embodiments, updating the data dictionary may include, but is not limited to, steps S501, S502, and S503:
[0125] Step S501: Obtain the original comment content whose verification result is normal, and obtain the legal comment content;
[0126] Step S502: If the data dictionary does not contain any valid comment content, the original data name and the valid comment content are mapped to obtain a mapping relationship;
[0127] Step S503: Store the mapping relationship in the data dictionary to update the data dictionary.
[0128] In step S501 of some embodiments, when the data dictionary needs to be updated, the original data tables of the entire database are obtained, and each original data table is traversed to obtain the original fields. The original comment content of the original fields is then validated to obtain a validation result. The obtained original comment content is considered valid if the validation result is normal, meaning it meets the preset comment requirements. If the preset comment requirements are non-empty and non-garbled Chinese characters, then the obtained original comment content is non-empty, non-garbled, and consists of Chinese characters, thus obtaining valid comment content.
[0129] For example, if the original comment content corresponding to field name "Field1" is "insurance policy", the original comment content of field name "Field2" is empty, and the original comment content of field name "Field3" is "%%##$", then the validation result for the original comment content "insurance policy" is normal, the validation result for field name "Field2" is that the original comment content is empty, and the validation result for the original comment content "%%##$" is that the original comment content contains garbled characters. Therefore, "insurance policy" is obtained as a valid comment content.
[0130] In step S502 of some embodiments, after obtaining the valid annotation content, the data dictionary is scanned according to the valid annotation content to search for annotation content that matches the valid annotation content. If no valid annotation content exists in the data dictionary, the original data name and the valid annotation content are mapped to obtain a mapping relationship. Specifically, searching the data dictionary for valid annotation content mainly involves searching for the initial annotation content corresponding to the original data name to determine if there is any initial annotation content that matches the valid annotation content. If not, it indicates that no valid annotation content exists in the data dictionary. If there is an initial annotation content that matches the valid annotation content, then there is no need to update the initial annotation content corresponding to that data name in the data dictionary.
[0131] Specifically, since each field name in the data dictionary corresponds to one initial comment content, after obtaining the valid comment content, it is determined whether the valid comment content is consistent with the initial comment content in the data dictionary. If they are consistent, no update is needed; otherwise, a mapping relationship between the original field name and the valid comment content is constructed to add the initial comment content corresponding to the field name in the data dictionary.
[0132] In step S503 of some embodiments, after obtaining the mapping relationship, the mapping relationship is stored in the data dictionary, and a corresponding valid comment is added to the original field name of the data dictionary.
[0133] For example, if the valid annotation is "insurance policy" and the corresponding original field name is "Field1", and the initial annotation for field name "Field1" in the data dictionary is "insurance type", then there is no valid annotation in the data dictionary. In this case, "insurance policy" and "Field1" are mapped to obtain a mapping relationship, and this mapping relationship is stored in the data dictionary. Therefore, the initial annotation for field name "Field1" in the data dictionary will then be "insurance type" and "insurance policy". Thus, by updating the data dictionary, it is continuously improved, making the data annotation more complete.
[0134] Combining steps S101 to S106 above, a preset annotation configuration table is obtained. Based on the annotation category and effectiveness information in the annotation configuration table, the effectiveness information is determined, and the corresponding annotation category is identified as the target category. The original data table is then retrieved from the preset original database based on the target category. The original data table is then traversed to obtain the original fields. The table name and field names of the original data tables are obtained. Initial annotation content is retrieved from the data dictionary based on the table name, and non-empty, non-garbled, and Chinese characters of the initial annotation content are selected as the target annotation content. This target annotation content is then annotated onto the original data table. Simultaneously, the initial annotation content is retrieved from the data dictionary based on the field names. It is determined whether the initial annotation content is non-empty, non-garbled, and Chinese characters. If the initial annotation content is empty or garbled, alternative annotation content is generated based on a random number and the field name of the original field. When annotating the original field based on the alternative annotation content, the original annotation content of the original field is retrieved, and it is determined whether the original annotation content contains garbled characters or is empty. If the original annotation content contains garbled characters, the alternative annotation content is annotated onto the original field to obtain the target data. Therefore, by automatically annotating the original data tables and fields, we can reduce the workload of manual annotation and improve the efficiency of data annotation.
[0135] Please see Figure 6 This application also provides an annotation apparatus that can implement the above-described annotation method. The apparatus includes:
[0136] The acquisition module 601 is used to acquire a pre-configured annotation configuration table; wherein, the annotation configuration table includes annotation categories and effectiveness information, and the effectiveness information is used to characterize the occurrence of annotation categories;
[0137] The data filtering module 602 is used to filter raw data from a preset raw database according to the annotation category and validity information; wherein, the raw data is the data to be annotated, and the raw data includes the raw data name;
[0138] Extraction module 603 is used to extract initial annotation content from a preset data dictionary based on the original data name;
[0139] The valid filtering module 604 is used to perform valid filtering on the initial annotation content to obtain the valid filtering results;
[0140] The annotation filtering module 605 is used to filter out the target annotation content from the initial annotation content if the valid filtering result is valid;
[0141] The annotation module 606 is used to annotate the original data according to the target annotation content to obtain the target data.
[0142] The annotation device illustrated in this application obtains an annotation configuration table, filters original data from the original database according to the annotation category and validity information in the annotation configuration table, extracts initial annotation content from the data dictionary according to the original data name of the original data, and performs a validity filter on the initial annotation content to obtain a validity filter result. The validity filter result is used as the validity initial annotation content as the target annotation content, and the original data is annotated according to the target annotation content to obtain the target data. Therefore, by automatically obtaining the original data and annotating the original data according to the target annotation content to obtain the target data, automated data annotation is achieved, which can both reduce manual workload and improve the efficiency of data annotation.
[0143] The specific implementation of this annotation device is basically the same as the specific embodiment of the annotation method described above, and will not be repeated here.
[0144] This application also provides a computer device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the above-described annotation method. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0145] Please see Figure 7 , Figure 7 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes:
[0146] The processor 701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0147] The memory 702 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 702 and is called and executed by the processor 701 using the commented methods of the embodiments of this application.
[0148] The input / output interface 703 is used to implement information input and output;
[0149] The communication interface 704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0150] Bus 705 transmits information between various components of the device (e.g., processor 701, memory 702, input / output interface 703, and communication interface 704);
[0151] The processor 701, memory 702, input / output interface 703, and communication interface 704 are connected to each other within the device via bus 705.
[0152] This application also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described annotation method.
[0153] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0154] This application provides an annotation method, annotation device, computer equipment, and storage medium. It obtains original data by filtering data to be annotated from a preset original database according to annotation category and validity information. Then, it obtains the original data name of the original data, extracts initial annotation content from a preset data dictionary based on the original data name, and performs a valid filtering on the initial annotation content to obtain a valid filtering result. The valid initial annotation content obtained from the valid filtering result is used as the target annotation content. Finally, the original data is annotated according to the target annotation content to obtain the target data. This achieves automated data annotation, reduces manual workload, lowers the probability of errors during manual annotation, and improves the efficiency of data annotation.
[0155] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0156] It will be understood by those skilled in the art that Figure 1-5 The technical solutions shown do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0157] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0159] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0160] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0161] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0162] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0163] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0164] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0165] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An annotation method, characterized in that, The method includes: Obtain a pre-configured annotation configuration table; wherein the annotation configuration table includes annotation categories and effectiveness information, and the effectiveness information is used to characterize the occurrence of the annotation category; Based on the effective information, the target category is filtered from the annotation category; Original data is filtered from a preset original database according to the target category; wherein, the original data is data to be annotated, and the original data includes the original data name, the original data table, and the original fields of the original data table; Initial annotation content is extracted from a preset data dictionary based on the original data name; wherein, the data dictionary uses a key-value data format to store matching information between the original data name and the annotation content; The initial annotation content is then subjected to a validity screening process to obtain the validity screening results; If the legal filtering result is legal, then the target annotation content is filtered out from the initial annotation content; The original data is annotated according to the target annotation content to obtain the target data; If the valid filtering result is invalid, alternative annotation content is generated based on the random number and the field name of the original field. The alternative annotation content is used to annotate the original data to obtain the target data.
2. The method according to claim 1, characterized in that, The step of filtering the original data from the original database according to the target category includes: The original data table is obtained from the original database according to the target category; The original data table is traversed to obtain the original fields.
3. The method according to any one of claims 1 to 2, characterized in that, The step of annotating the original data according to the target annotation content to obtain the target data includes: Obtain the original annotation content of the original data; The original annotation content is validated to obtain the validation result; If the verification result indicates that the original annotation content is empty or contains garbled characters, then the original field is annotated according to the target annotation content to obtain the target data.
4. The method according to claim 3, characterized in that, After extracting initial annotation content from a preset data dictionary based on the original data name, the method further includes: Updating the data dictionary specifically includes: Obtain the original annotation content whose verification result is normal to obtain the legal annotation content; If the data dictionary does not contain the legal annotation content, the original data name and the legal annotation content are mapped to obtain a mapping relationship; The mapping relationship is stored in the data dictionary to update the data dictionary.
5. An annotation device, characterized in that, The device includes: An acquisition module is used to acquire a pre-configured annotation configuration table; wherein, the annotation configuration table includes annotation categories and effectiveness information, and the effectiveness information is used to characterize the occurrence of the annotation category; The data filtering module is used to filter target categories from the annotation categories based on the effective information, and to filter original data from a preset original database based on the target categories; wherein, the original data is the data to be annotated, and the original data includes the original data name, the original data table, and the original fields of the original data table; The extraction module is used to extract initial annotation content from a preset data dictionary based on the original data name; wherein the data dictionary uses a key-value data format to store matching information between the original data name and the annotation content; The valid filtering module is used to perform valid filtering on the initial annotation content to obtain valid filtering results; An annotation filtering module is used to filter out target annotation content from the initial annotation content if the valid filtering result is valid; The annotation module is used to annotate the original data according to the target annotation content to obtain the target data; The device is further configured to: if the valid filtering result is invalid, generate alternative annotation content based on a random number and the field name of the original field, wherein the alternative annotation content is used to annotate the original data to obtain the target data.
6. A computer device, characterized in that, The computer device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory, wherein the program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 4.
7. A storage medium, said storage medium being a computer-readable storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method, device and equipment for automatically generating Chinese annotation, and storage medium
CN108509199A
Database annotation method and device and terminal equipment
CN110110067A
Code annotation generation method and device
CN112416429A