Method and device for labeling data points
Through the machine learning-based data point labeling mechanism, the adaptive labeling mechanism is used to reduce manual labeling labor, solving the problem of inefficient data point labeling in complex business application scenarios, and achieving efficient and accurate data point labeling.
Patent Information
- Application Number
- CN201980099124.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-08-22
AI Technical Summary
In complex business application scenarios, the business instances represented by data points are becoming increasingly complex, and a large number of manual annotations are required by business experts to perform data analysis, resulting in labor-intensive and inefficient.
A machine learning-based data point annotation mechanism is provided. By dividing the target data set into subsets, receiving user input to specify marks for some data points, and using artificial intelligence to judge mark similarity, adaptively setting marks for data points in the entire data set.
It effectively reduces the user's manual labeling labor, improves the efficiency of data point labeling, and maintains high classification accuracy.
Smart Images

Figure CN114270336B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to information processing, and more particularly, to methods and apparatus for labeling data points. Background Art
[0002] Data visualization refers to the visual representation of data, which aims to convey the information contained in the data clearly and efficiently by graphical means. A data set consisting of multiple data points is the basis of data visualization. In an exemplary typical application scenario, a data visualization tool can obtain data collected by an Internet of Things (IoT) sensor at a certain frequency, such as temperature data, pressure data, humidity data, etc., visualize a large number of data points constructed from such sensor data, draw them into charts (e.g., data distribution charts) and present them on a visual user interface.
[0003] As the number of application scenarios increases, the business instances represented by data points become increasingly complex. For example, a data point may represent a combination of data collected by more than one IoT sensor at a time. Usually, only business domain experts who are familiar with specific businesses can understand and distinguish the different situations that different data points can reflect, such as whether the corresponding business instance is normal. Therefore, it is necessary for business domain experts, or at least with the assistance of business domain experts, to annotate the various data points obtained for data visualization and add tags to illustrate the corresponding situations, so as to facilitate relevant data analysis including data point classification. Summary of the invention
[0004] The present invention summary section is provided to introduce some selected concepts in a simplified form, which will be further described in the following detailed description section. The present invention summary section is not intended to identify any key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.
[0005] According to one aspect of the present disclosure, a method for labeling data points is provided, comprising: performing a labeling operation on a target data set, the target data set comprising a plurality of data points, each of the plurality of data points representing a business instance, the labeling operation comprising: dividing the target data set into a plurality of first subsets; for each of the plurality of first subsets: receiving user input, the user input being used to specify a tag for at least one data point in the first subset, wherein the tag specified for the data point is used to describe the situation of the business instance represented by the data point; determining whether the similarity between the tag and a tag previously specified for at least one data point in the target data set meets a preset condition; in response to determining that the similarity does not meet the preset condition, re-performing the labeling operation with the first subset as the target data set; in response to determining that the similarity meets the preset condition, setting a tag associated with the tag previously specified for at least one data point in the target data set for each data point in the first subset.
[0006] According to another aspect of the present disclosure, a device for labeling data points is provided, comprising: a module for performing a labeling operation on a target data set, the target data set comprising a plurality of data points, each of the plurality of data points representing a business instance, the module for performing a labeling operation on the target data set comprising: a module for dividing the target data set into a plurality of first subsets; for each of the plurality of first subsets: a module for receiving user input, the user input being used to specify a tag for at least one data point in the first subset, wherein the tag specified for the data point is used to illustrate the situation of the business instance represented by the data point; a module for determining whether the similarity between the tag and a tag previously specified for at least one data point in the target data set meets a preset condition; a module for, in response to determining that the similarity does not meet the preset condition, re-performing the labeling operation with the first subset as the target data set; and a module for, in response to determining that the similarity meets the preset condition, setting a tag associated with the tag previously specified for at least one data point in the target data set for each data point in the first subset.
[0007] According to another aspect of the present disclosure, a computing device is provided, comprising: a memory for storing instructions; and at least one processor coupled to the memory, wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform a labeling operation on a target data set, the target data set comprising a plurality of data points, each of the plurality of data points representing a business instance, the labeling operation comprising: dividing the target data set into a plurality of first subsets; for each of the plurality of first subsets: receiving user input, the user input being used to specify a tag for at least one data point in the first subset, wherein the tag specified for the data point is used to illustrate the situation of the business instance represented by the data point; determining whether the similarity between the tag and a tag previously specified for at least one data point in the target data set meets a preset condition; in response to determining that the similarity does not meet the preset condition, re-performing the labeling operation with the first subset as the target data set; in response to determining that the similarity meets the preset condition, setting a tag associated with the tag previously specified for at least one data point in the target data set for each data point in the first subset.
[0008] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which instructions are stored. When the instructions are executed by at least one processor, the at least one processor performs a labeling operation on a target data set, wherein the target data set includes a plurality of data points, each of the plurality of data points represents a business instance, and the labeling operation includes: dividing the target data set into a plurality of first subsets; for each of the plurality of first subsets: receiving user input, the user input is used to specify a tag for at least one data point in the first subset, wherein the tag specified for the data point is used to illustrate the situation of the business instance represented by the data point; determining whether the similarity between the tag and the tag previously specified for at least one data point in the target data set meets a preset condition; in response to determining that the similarity does not meet the preset condition, re-executing the labeling operation with the first subset as the target data set; in response to determining that the similarity meets the preset condition, setting a tag associated with the tag previously specified for at least one data point in the target data set for each data point in the first subset.
[0009] Various aspects of the present disclosure provide a data point labeling mechanism based on machine learning. Advantageously, the provided adaptive labeling mechanism only requires users (business domain experts, or at least operators with the assistance of business domain experts) to specify labels for a small number of data points in the target data set, and relies on artificial intelligence to apply the manually specified labels to all data points in the target data set, which can effectively reduce the user's manual labor while still maintaining a high classification accuracy.
[0010] In addition, in an example of any of the aforementioned aspects, optionally, it may also include: dividing the initial data set into a plurality of second subsets; for each of the plurality of second subsets, receiving user input, wherein the user input is used to specify a label for at least one data point in the second subset; and selecting one of the plurality of second subsets as the target data set.
[0011] Advantageously, the operations on the initial data set in the above example are also implemented using some operations similar to the aforementioned adaptive labeling mechanism. In implementation, the set operation process and corresponding modules can be partially reused, and the user's manual labor can be further reduced, thereby improving the user experience.
[0012] In addition, in an example of any of the foregoing aspects, in response to determining that the similarity satisfies the preset condition, setting a tag associated with a tag previously specified for at least one data point in the target data set for each data point in the first subset, optionally, it may include: setting the tag previously specified for at least one data point in the target data set as the tag of each data point in the first subset.
[0013] Advantageously, in the above example, setting the label previously specified for at least one data point in the target data set as the label of each data point in its current subset can simplify the operation while ensuring the labeling accuracy of the labeling operation.
[0014] In addition, in an example of any of the foregoing aspects, when determining whether the similarity between the tag and a tag previously assigned to at least one data point in the target data set satisfies a preset condition, it may optionally include: determining the semantic similarity between the tag and a tag previously assigned to at least one data point in the target data set; and determining whether the semantic similarity is greater than a preset threshold.
[0015] Advantageously, in the above example, the judgment is made based on semantic similarity, and more complex tags can be processed, so more complex business situations can be handled, and the requirements for user input are also lower. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Implementations of the present disclosure are illustrated by way of example and not limitation in the accompanying drawings in which like reference numerals designate the same or similar parts and in which:
[0017] Figure 1 An exemplary environment is shown in which some implementations of the present disclosure may be implemented;
[0018] Figure 2 An example of a data distribution graph including a plurality of data points is shown;
[0019] Figure 3 An example of a data distribution graph including a plurality of data points is shown;
[0020] Figure 4 An example of a data distribution graph including a plurality of data points is shown;
[0021] Figure 5 is a flow chart of an exemplary method according to one implementation of the present disclosure;
[0022] Figure 6 Some exemplary operations according to one implementation of the present disclosure are shown;
[0023] Figure 7 An example of a data distribution graph including a plurality of data points is shown;
[0024] Figure 8 is a block diagram of an exemplary apparatus according to one implementation of the present disclosure; and
[0025] Fig. 9 is a block diagram of an exemplary computing device according to one implementation of the present disclosure.
[0026] Reference numerals list
[0027] 110: device 120: at least one data source 130: network
[0028] 510: Specify a portion of the initial dataset as the target dataset
[0029] 520: Perform labeling operations on the target dataset
[0030] 521: Divide the target data set into a plurality of first subsets
[0031] 522: Receive user input for specifying a label for at least one data point in the current first subset
[0032] 523: Determine the similarity between the label and a label previously assigned to at least one data point in the target data set
[0033] 524: If the similarity does not meet the preset condition, re-execute the labeling operation using the current first subset as the target data set
[0034] 525: If the similarity meets the preset conditions, set a mark for each data point in the current first subset
[0035] 810-850: Module
[0036] 910: Processor
[0037] 920: Memory DETAILED DESCRIPTION
[0038] In the following description, a large number of specific details are set forth for the purpose of explanation. However, it is to be understood that the implementation of the present invention can be implemented without these specific details. In other examples, well-known circuits, structures and techniques are not shown in detail to avoid affecting the understanding of the description.
[0039] References throughout the specification to "an implementation," "implementation," "exemplary implementations," "some implementations," "various implementations," etc., indicate that the implementation of the invention being described may include certain features, structures, or characteristics, however, it does not mean that every implementation must include these certain features, structures, or characteristics. Furthermore, some implementations may have some, all, or none of the features described for other implementations.
[0040] In the following description and claims, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not intended as synonyms for each other. Instead, in particular implementations, "connected" is used to indicate that two or more components are in direct physical or electrical contact with each other, while "coupled" is used to indicate that two or more components cooperate or interact with each other, but they may or may not be in direct physical or electrical contact.
[0041] Data visualization tools can draw data distribution graphs for a large number of data points to present them to users. With the increasing number of application scenarios, the business instances represented by data points are becoming more and more complex. In order to conduct business-meaningful data analysis, including the correct classification of data points, business domain experts, or at least with the assistance of business domain experts, need to add labels to the acquired data points one by one to illustrate the corresponding situation. However, this requires a lot of manual labor.
[0042] The present disclosure aims to provide a data point labeling mechanism based on machine learning to solve the above problems. With the help of this mechanism, users (business domain experts, or at least operators with the assistance of business domain experts) are only required to specify labels for a small number of data points in the target data set, and rely on artificial intelligence to apply the manually specified labels to all data points in the target data set. This can effectively reduce the user's manual labor while still maintaining a high classification accuracy.
[0043] Refer to the following Figure 1 , which illustrates an exemplary operating environment 100 in which some implementations of the present disclosure may be implemented. Operating environment 100 may include device 110 and at least one data source 120. In some implementations, device 110 and data source 120 may be communicatively coupled to each other via network 130.
[0044] In some examples, a data visualization tool may be run on the device 110, which is used to visualize data obtained from at least one data source 120. In some examples, the machine learning-based data point annotation mechanism provided in the present disclosure may be implemented as part of the data visualization tool, for example, as a plug-in thereof. In other examples, the mechanism may be implemented as a separate component on the device 110.
[0045] Examples of device 110 may include, but are not limited to, a mobile device, a personal digital assistant (PDA), a wearable device, a smart phone, a cellular phone, a handheld device, a messaging device, a computer, a personal computer (PC), a desktop computer, a laptop computer, a notebook computer, a handheld computer, a tablet computer, a workstation, a minicomputer, a mainframe computer, a supercomputer, a network appliance, a web appliance, a processor-based system, a multi-processor system, a consumer electronic device, a programmable consumer electronic device, a television, a digital television, a set-top box, or any combination thereof.
[0046] At least one data source 120 is used to provide data for manipulation by a data visualization tool on the device 110. By way of example and not limitation, the data source 120 may include various types of sensors, such as image sensors, temperature sensors, pressure sensors, humidity sensors, current sensors, and the like. In some examples, the sensor 120 may be configured to collect data at a fixed frequency, while in other examples, the data sampling frequency of the sensor 120 is adjustable, for example, in response to an indication signal from an external source (e.g., the device 110). In addition, in some examples, the data source 120 may be a database, a storage device, or any other type of device for providing data. The present disclosure is not limited to a particular type of data source.
[0047] In addition, data collected by at least one data source 120 can be directly provided to the device 110 for data visualization operations, or can be first stored in the device 110 (e.g., in a memory contained therein) or in a database / server (not shown) communicatively coupled to the device 110 and / or the data source 120 via the network 130, and retrieved when needed.
[0048] The network 130 may include any type of wired or wireless communication network, or a combination of wired and wireless networks. In some examples, the network 130 may include a wide area network (WAN), a local area network (LAN), a wireless network, a public telephone network, an intranet, an Internet of Things (IoT), etc. In addition, although a single network 130 is shown here, the network 130 may be configured to include multiple networks.
[0049] In addition, despite the above combination Figure 1 The exemplary operating environment according to some implementations of the present disclosure is described. In other implementations, the communication between the device 110 and at least one data source 120 may also be directly coupled without going through a network. In some examples, the device 110 may be deployed at an industrial site, and the data collected by various industrial sensors at the industrial site may be regarded as the data source 120. The present disclosure is not limited to Figure 1 Specific architecture shown.
[0050] In addition, in some examples, the data visualization tools mentioned above and the data point annotation mechanism based on machine learning in the present disclosure can be deployed in a distributed computing environment and can also be implemented using cloud computing technology.
[0051] Figure 2 An example of a data distribution graph including a plurality of data points is shown. Figure 2 The data distribution diagram shown in is for the operation of a numerical control machine tool (CNC). Figure 2 In , each data point in the data set represents a business instance, which is specifically a vector composed of two characteristic elements. The two characteristic elements are used to represent the current values of the two spindles of the CNC at a moment. The current values can come from current sensors deployed in the corresponding spindle circuits. Figure 2 In the figure, the horizontal direction (abscissa) represents the current value of spindle 1, and the vertical direction (ordinate) represents the current value of spindle 2.
[0052] For ordinary data analysts who lack professional knowledge in the CNC field, Figure 2From the data distribution pattern of the data points shown in , the data points contained in the area circled by ellipse 210 are usually considered to correspond to the normal machine operation status of the CNC, while the data points contained in the areas circled by the other two ellipses 220 and 230 are considered to correspond to the abnormal conditions of the CNC. However, in the eyes of experts in the field familiar with the CNC, it may be obvious that ellipse 220 corresponds to a true abnormality, while ellipse 230 actually corresponds to a false abnormality that occurs in some cases. It can be seen that it is very important to add tags (knowledge) to these data points to illustrate the specific circumstances of the business instances represented by the corresponding data points, such as distinguishing between true abnormalities and false abnormalities and normal CNC machine operation conditions. The classification of data points based on such important tags and related data analysis have practical business value.
[0053] Although the above-mentioned business instance is illustrated by combining the current values of the two spindles of the CNC, those skilled in the art will understand that the business instance for the CNC may include a more complex structure. For example, a business instance may include a combination of more sensor data that can be used to describe the operating status of the CNC at a moment.
[0054] In addition, it is understandable that as the number of data points contained in a data set increases, it will become a very time-consuming and laborious task for business domain experts to label these data points one by one.
[0055] Figure 3 An example of a data distribution graph including a plurality of data points is shown. Figure 3 It can correspond to a view displayed on a graphical user interface (GUI) of a data visualization tool. Figure 3 In the example, the horizontal direction represents the abscissa and the vertical direction represents the ordinate. No specific limitation is made here on the actual physical meanings represented by the abscissa and ordinate, respectively. In some implementations of the present disclosure, Figure 3 Each data point in the data set shown in can represent a more complex business instance. In one example, each data point can represent a relationship diagram, such as a curve diagram, which is used to represent the relationship between several variables related to a specific business. As another example, each data point can also represent an image, such as an appearance image of a circuit board produced at a fixed position on an industrial assembly line, collected by an image sensor in order to find production defects at the fixed position.
[0056] For complex business instances, some appropriate mechanisms can be used to reduce their dimensionality. For example, in the previous example, the average value of all pixels in the relationship graph or image can be calculated to compress the relationship graph or image to a data point, so that such a data point can be displayed in a low-dimensional (such as two-dimensional) data distribution graph. In addition, as the complexity of the business instance represented by a data point increases, the content of the tags added to the data point may also be more and more complex, because more explanations are needed for the situation of the business instance, such as the specific operating status interpretation of the corresponding machine, the cause analysis of the failure, and matters needing attention, etc.
[0057] Continue to see Figure 3 , which shows the correct labeling of data points in the entire data set, where each data point is manually labeled by a business domain expert to illustrate the situation of the business instance represented by the data point. And, depending on the different situations of the corresponding business instances described by the labels added to the data points, different shapes are used to distinguish these data points. In other words, intuitively, the difference in shape can reflect the difference in labeling.
[0058] Figure 4 An example of a data distribution graph containing multiple data points is shown. In this example, Figure 4 and Figure 3 The data set is the same, but the difference is that the business domain experts did not Figure 4 Each data point in the is labeled to describe the business instance that the data point represents. Figure 4 The results of automatically classifying these data points based on a conventional unsupervised clustering algorithm are shown. Figure 4 In the graph, different shapes indicate that the corresponding data points belong to different clusters.
[0059] Comparison Figure 3 It can be clearly seen that Figure 4 The classification of data points based solely on unsupervised clustering algorithms is obviously different from Figure 3 The classification according to the correct labeling is shown. Figure 4 Some data points are incorrectly partitioned.
[0060] Next reference Figure 5 , Figure 5 FIG. 5 is a flowchart of an exemplary method 500 according to one implementation of the present disclosure. For example, the method 500 may be Figure 1The method 500 may be implemented in the device 110 shown in FIG. 1 or any similar or related entity. In one example, the method 500 may be part of an operation performed by a data visualization tool running on the device 110. The exemplary method 500 may be used to adaptively label data points, whereby the data points can be correctly classified accordingly, and only a small amount of manual labor by the user (business domain expert and / or related operator) is required, which greatly reduces the burden on the user and improves the user experience.
[0061] See also Figure 5 The method 500 starts at step 510, in which a portion of an initial data set is set as a target data set. The initial data set includes a plurality of data points, each of which represents a business instance.
[0062] In some examples, the business instance represented by each data point in the data set may be a vector consisting of more than one business-related feature element for a business. For example, each feature element is data collected by a corresponding sensor. Figure 2 For example, for a CNC, such a business instance can be a vector composed of current values from two current sensors deployed in two spindle circuits of the CNC. In addition, a business instance represented by each data point may also have a more complex structure. However, the present disclosure is not limited to a specific type of data point / business instance. The purpose of collecting data points / business instances may include but is not limited to monitoring the operating status of the business to discover possible problems, etc.
[0063] The method 500 then proceeds to step 520, in which a labeling operation is performed on the target data set to set a label for each data point in the target data set. As described above, the label is used to illustrate the situation of the business instance represented by the corresponding data point. According to an implementation of the present disclosure, the labeling operation of step 520 can be iteratively performed until all data points in the target data set are marked.
[0064] The operation in step 520 is described in detail below.
[0065] Specifically, performing a labeling operation on a target data set may include: in step 521, dividing the target data set into a plurality of first subsets. In some examples, cluster analysis may be used to divide the target data set into a plurality of first subsets. Cluster analysis is a machine learning technique that belongs to unsupervised learning, which attempts to divide the data points in a data set into several different subsets, each of which can be called a cluster. Among the many algorithms used in cluster analysis, the K-means clustering algorithm is one of the most commonly used and most important clustering algorithms. The algorithm is simple and efficient, can still work effectively for a large number of data points, and can meet actual production needs. In a preferred implementation, the K-means clustering algorithm can be used in step 521 to divide the data points in the target data set into a plurality of first subsets.
[0066] The relevant parameters of the K-means clustering algorithm can be set depending on the actual situation. For example, the relevant parameters may include the number of clusters / subsets, and the setting of the number of subsets may depend on the number of data points contained in the processed data set; additionally or alternatively, the setting of the number of subsets may also depend on the needs of the business represented by the data points in the data set; additionally or alternatively, the setting of the number of subsets may also depend on the experience of business domain experts and / or data analysts. Other considerations are also feasible.
[0067] In addition, step 521 may also adopt any other suitable clustering algorithm other than the K-means clustering algorithm, including but not limited to those clustering algorithms that use the number of clusters / subsets as parameters.
[0068] In a specific example, the graphical user interface of the data visualization tool can present a data distribution diagram of data points collected for a business (including the target data set), and for the division of a data set (target data set), such as the division using the above-mentioned K-means clustering algorithm, controls can be provided so that the user can view / input / modify the relevant parameters of the K-means clustering algorithm. In addition, the graphical user interface of the data visualization tool can also provide an animation effect display of the division process so that the user can visually monitor the entire process.
[0069] After step 521, the operation process proceeds to step 522, in which, for each of the plurality of first subsets divided by the target data set, a user input is received, wherein the user input is used to specify a label for at least one data point in the current first subset.
[0070] In some examples, the selection of at least one data point in the first subset can be automatic, that is, at least one data point is automatically selected from the first subset by, for example, a data visualization tool, and the selected data points are presented to the user together or one by one for the latter to assign labels to them respectively. In other examples, the user himself can manually select at least one data point from the presented first subset and then assign labels to them respectively. The present disclosure is not limited to the above specific examples. In a preferred implementation, the selection of at least one data point in the first subset is random. In addition, in some of the following examples, for the sake of ease of description, it is assumed that the user input received in step 522 is to assign a label to a data point in the first subset, and it can be understood in combination with the following description that this can minimize the user's manual labor.
[0071] The tag is used to illustrate the situation of the business instance represented by the corresponding data point. In some implementations of the present disclosure, such a tag can be a simple symbol, a word, a phrase, a more complex sentence, or even a paragraph, etc. In some examples, the tag can come from a preset tag set, and the user can select a suitable tag from the preset tag set when specifying a tag for a data point. In some other examples, the tag can also come from the user's free input, such as the user inputting a piece of text, etc. The present disclosure is not limited to the above-mentioned specific examples, and a combination of these and other examples is also possible, as long as it can help to clearly and accurately illustrate the situation of the corresponding business instance.
[0072] In addition, in some examples, in order to facilitate users to annotate selected data points, some relevant information of the data point can be presented to the user, such as presenting the user with the real data of the business instance represented by the data point, such as the values of each sensor, relationship diagrams, images, etc., as explained in the previous examples.
[0073] In a specific example, on the graphical user interface of the data visualization tool, for the current first subset, the selected data points to be annotated can be highlighted by highlighting and / or magnifying. For example, in the case where the user manually selects the data points, the user can use a pointing tool such as a mouse to hover over, click on, or otherwise select a data point, and the selected data point is highlighted and / or magnified. Additionally, after the data point is selected, the relevant information of the selected data point can be presented in the form of a pop-up menu or bubble, etc., to facilitate the user's reference for annotation. Accordingly, at least one control, such as a drop-down menu, a text box, etc., is also provided on the graphical user interface of the data visualization tool to facilitate the user to select / enter the content of the mark to be added.
[0074] Furthermore, to facilitate subsequent use, in some examples, the tag specified for the data point in the first subset may also be stored in association with an identifier of the first subset and / or an identifier of the data point.
[0075] Next, the operation process proceeds to step 523, in which it is determined whether the similarity between the label specified by the user input for at least one data point in the current first subset and the label previously specified for at least one data point in the target data set in step 522 satisfies a preset condition. The preset condition may be one condition or a combination of a series of conditions, and the series of conditions may also include other constraints on the current first subset and / or the target data set. It can be understood that, from the perspective of set relationship, the target data set is the parent set of the first subset.
[0076] In some examples, the judgment operation in step 523 may include determining whether the tag specified for at least one data point in the first subset is equal to the tag previously specified for at least one data point in the target data set. This may be particularly applicable to some occasions where the tags are relatively simple. For example, if the user specifies a tag "N" (assuming that N represents normal here) selected from a preset tag set for a data point in the current first subset, and the tag previously specified by the user for a data point in the target data set (that is, the parent set of the first subset) is also a tag "N" selected from the preset tag set, then the two can be directly determined to be equal, and it is considered that the similarity between them meets the preset condition.
[0077] In some other examples, the judgment operation may include determining whether the similarity between the tag specified for at least one data point in the first subset and the tag previously specified for at least one data point in the target data set is greater than a preset threshold. Here, any form of similarity measurement available between two tags is possible, and the present disclosure is not limited to a specific implementation. The selection of the threshold can be set by a standard for similar businesses, or can be manually specified by the user, for example, depending on the specific needs of the business, the experience of business domain experts and / or data analysts, etc. In a preferred implementation of the present disclosure, especially for a more complex situation such as a sentence or a paragraph, the operation of step 523 can be implemented by determining the semantic similarity between the two tags and then judging whether the semantic similarity is greater than a preset threshold (for example, 80%). By using a judgment method based on semantic similarity, the tag content that can be processed can be more complex, so it can cope with more complex business situations. On the other hand, the requirements for user input are also lower, and it is no longer required to select from a preset tag set, but the user can be allowed to enter freely. In addition, in some examples, for the calculation of semantic similarity, natural language processing technologies such as word embedding (Word Embedding) or sentence embedding (Sentence Embedding) can be used to convert two tags into vector representations respectively, and the semantic similarity between the two can be calculated accordingly.
[0078] Furthermore, in a specific example, after the user inputs a label specified for at least one data point in the first subset, the result of the similarity determination may be displayed on a graphical user interface of the data visualization tool.
[0079] Next, the operation process proceeds to step 524. In this step, in response to determining that the similarity does not satisfy the preset condition (that is, the judgment result in step 523 is "No"), the current first subset is used as the target data set. That is to say, at this time, the target data set as the parent set has been updated to the first subset that does not satisfy the preset condition in the judgment of step 523, and then the operation process jumps back to step 521 to re-execute the operation of this step and its subsequent steps.
[0080] It can be understood that the above operations included in step 520 are a series of iterative operations until it is found that in a certain iteration, the similarity judgment performed on the current first subset meets the preset condition. In addition, the preset condition may also include, for example: relative to the initially set target data set, the number of current iterations has exceeded a specified number (for example, the specified number is 10); in the current iteration, the number of data points in the first subset is no more than a specified number (for example, the specified number is 1), etc. Those skilled in the art can understand that the preset condition may include any combination of the above and other constraints.
[0081] On the other hand, in response to determining that the similarity satisfies the preset condition (i.e., the judgment result in step 523 is "yes"), in step 525, a mark associated with a mark previously specified for at least one data point in the target data set is set for each data point currently in the first subset.
[0082] In some examples, setting a mark in step 525 may include: setting a mark previously specified for at least one data point in the target data set as a mark for each data point in the current first subset. For example, if the mark previously specified for a data point in the target data set is X1, and the mark specified for a data point in the current first subset (which is one of the subsets of the target data set) is X2, where X2 is the same as X1 or the two are not the same but the similarity still meets the preset conditions, then X1 can be set as the mark for each data point in the first subset. In other words, if the judgment result in step 523 is "yes", it can be considered that the mark previously specified for a data point in the target data set (parent set) is appropriate. In this case, the above mark can be directly applied to each data point in the current first subset as its mark, without the need to use different marks separately for the latter.
[0083] In some other examples, the tag set in step 525 may also be determined based on a tag previously specified for at least one data point in the target data set (parent set), a tag specified for at least one data point in the current first subset, and / or a tag specified for at least one data point in a sibling subset of the current first subset (i.e., at least one other first subset in a plurality of first subsets divided by the target data set), etc. The present disclosure is not limited to the above or other specific examples.
[0084] It can be understood that after performing the operation in step 525 on the current first subset, the operations in 522-525 are also applied to all brother subsets of the first subset (i.e., other first subsets in the multiple first subsets divided by the target data set in step 521), so that appropriate labels are set for all data points in the target data set in each iteration, and ultimately appropriate labels can be set for all data points in the initially set target data set.
[0085] In addition, those skilled in the art can understand that if, in each iteration of the annotation operation for the target data set, only one data point is selected from each of the plurality of first subsets of the target data set and a label is assigned to it by the user, then the adaptive annotation mechanism according to the present disclosure can minimize the manual labor of the user. On the other hand, the user can also assign labels to more than one data point in the current first subset in step 522. In this case, the subsequent steps 523-525, if involving the label assigned to at least one data point in the first subset, and similarly involving the label previously assigned to at least one data point in the target data set, can refer to the result of a specific operation on the labels of the corresponding more than one data point, and the operation includes but is not limited to calculating the average value of more than one label.
[0086] In addition, in a specific example, the graphical user interface of the data visualization tool can also provide an animation effect display for classifying the data points in the data set according to the labels set for the data points, using various possible forms such as different colors / grayscales / shapes / patterns or any combination thereof to indicate that the corresponding data points belong to different clusters.
[0087] In a preferred implementation, setting a portion of the initial data set as a target data set in step 510 may include: dividing the initial data set into a plurality of second subsets; receiving user input for each of the plurality of second subsets, the user input being used to specify a label for at least one data point in the second subset; and selecting one of the plurality of second subsets as the target data set.
[0088] Among them, the operation process of dividing the initial data set into a plurality of second subsets is similar to the operation process of dividing the target data set into a plurality of first subsets in step 521 discussed above, and the details can be found in the above description. In some examples, cluster analysis can be used to divide the initial data set into a plurality of second subsets. In a preferred implementation, the K-means clustering algorithm can be used in this step to divide the initial data set into a plurality of second subsets. It should be noted that although the K-means clustering algorithm is also illustrated as an example of the use of the specific example of step 521, since the operation objects in this step and step 521 are different, the specific parameter settings of the K-means clustering algorithm in the two processes may not be the same.
[0089] In addition, for each of the plurality of second subsets, the operation process of receiving user input for specifying a label for at least one data point in the second subset is similar to the operation process of receiving user input for specifying a label for at least one data point in the first subset in step 522 discussed above, and the details can also be referred to the above description. In some examples, the selection of at least one data point in the second subset to be specified with a label can be performed automatically or manually by the user, and the present disclosure is not limited to this. In addition, in a preferred implementation, the selection of at least one data point in the second subset is random.
[0090] Furthermore, it is understandable that appropriate labels can be set for all data points in the initial data set by sequentially setting each of the multiple second subsets divided from the initial data set as a target data set and performing the labeling operation 520 described in detail above.
[0091] In the above example, the operations on the initial data set can also be implemented using some operations similar to the aforementioned annotation operations on the target data set. Therefore, in the implementation, the set operation process and corresponding modules can be partially reused, and the user's manual labor can be further reduced, thereby improving the user experience.
[0092] Although a specific operation example of step 510 is shown above, those skilled in the art may understand that it is also possible to set the target data set from the initial data set in other ways, and the present disclosure is not limited to the above specific example.
[0093] Next, combine Figure 6 The specific examples are used to further illustrate the processes of some of the operations described above.
[0094] exist Figure 6In the example, the initial data set is divided into three subsets using a clustering algorithm, namely, the first-level subsets A, B, and C shown at the top. In a similar manner, each subset is divided into a plurality of second-level subsets using a clustering algorithm. In addition, some second-level subsets may be further divided into a plurality of third-level subsets, and so on. For the convenience of illustration, the letter of each subset is the label assigned to a data point in the subset by user input in the manner described in this article, which is referred to as a set label here.
[0095] by Figure 6 Take the example on the left side of the figure as an example, the first-level subset A is divided into two second-level subsets. Based on the user's annotation, the set labels of the two second-level subsets are both A, which is the same as the set label of the upper-level set, and meets the preset conditions. In this case, there is no need to use a clustering algorithm to further divide the two second-level subsets. Instead, for example, all data points in the two second-level subsets (that is, all data points in the upper-level set) can be set with label A.
[0096] Let's look at it again Figure 6 In the middle example, the first-level subset B is divided into two second-level subsets B and D. The set label of the first second-level subset is B, which is the same as the set label of its upper-level set. Therefore, it is not further divided. Instead, the label B is set for all data points in the second-level subset. The set label of the second second-level subset is D, which is different from the set label B of its upper-level set. Therefore, it does not meet the preset conditions. In this case, it is necessary to use a clustering algorithm to further divide it. As shown in the figure, the second-level subset D is divided into two third-level subsets. Now, the set labels of the two third-level subsets are both D, which are the same as the set labels of their upper-level sets and meet the preset conditions. Therefore, there is no further division of the two third-level subsets. Instead, the label D is set for all data points in the two third-level subsets (that is, all data points in their upper-level sets).
[0097] Figure 6 The example on the right side of the figure shows the division related to the first-level subset C, with corresponding set labels C, E, F, and G. Its processing logic is consistent with the previous two examples, so it will not be described in detail here.
[0098] Figure 7 An example of a data distribution graph containing multiple data points is shown. In this example, Figure 7 and Figure 3 The data set corresponds to the same data set, but Figure 7The result after applying a method according to an implementation of the present disclosure to the data set is shown in FIG. Here, different shapes are used to distinguish the different clusters to which the corresponding data points belong, indicating that they are marked differently. It can be clearly seen that Figure 7 The results shown in Figure 3 Overall, only 3% of the data points in the entire data set need to be annotated by users, that is, by business domain experts, while the remaining data points are automatically annotated by machines, and the mechanism disclosed in the present invention can achieve 85% classification accuracy.
[0099] Reference below Figure 8 , Figure 8 is a block diagram of an exemplary apparatus 800 according to one implementation of the present disclosure. For example, the apparatus 800 may be Figure 1 The invention may be implemented in the device 110 shown in FIG. 1 or any similar or related entity.
[0100] The exemplary apparatus 800 is used to label data points. More specifically, the function implemented by the apparatus 800 may include performing a labeling operation on a target data set, wherein the target data set includes a plurality of data points, each of which represents a business instance. Figure 8 As shown, the exemplary apparatus 800 may include a module 810, which is used to divide the target data set into a plurality of first subsets. In addition, the apparatus 800 may also include a module 820, for each of the plurality of first subsets, the module 820 is used to receive user input, the user input is used to specify a tag for at least one data point in the first subset, wherein the tag specified for the data point is used to illustrate the situation of the business instance represented by the data point. In addition, the apparatus 800 may also include a module 830, which is used to determine whether the similarity between the tag and the tag previously specified for at least one data point in the target data set meets a preset condition. In addition, the apparatus 800 may also include a module 840, which is used to, in response to determining that the similarity does not meet the preset condition, re-execute the labeling operation with the first subset as the target data set. In addition, the apparatus 800 may also include a module 850, which is used to, in response to determining that the similarity meets the preset condition, set a tag associated with the tag previously specified for at least one data point in the target data set for each data point in the first subset.
[0101] It should be noted that although the apparatus 800 is shown as including modules 810-850, the apparatus 800 may include more or fewer modules to implement the various functions described herein. In some examples, the modules 810-850 may be included in one module for performing a labeling operation on a target data set, such as corresponding to Figure 5 In addition, in some examples, the apparatus 800 may further include additional modules for performing other operations described in the specification, such as Figure 5 Those skilled in the art will appreciate that the exemplary apparatus 800 may be implemented using software, hardware, firmware, or any combination thereof.
[0102] Now go to Fig. 9 , a block diagram of an exemplary computing device 900 according to an implementation of the present disclosure is shown here. As shown, the exemplary computing device 900 may include at least one processing unit 910. The processing unit 910 may include any type of general-purpose processing unit / core (for example, but not limited to: CPU, GPU), or a dedicated processing unit, core, circuit, controller, and the like. In addition, the exemplary computing device 900 may also include a memory 920. The memory 920 may include any type of medium that can be used to store data. In one implementation, the memory 920 is configured to store instructions that, when executed, cause at least one processing unit 910 to perform the operations described herein, such as the annotation operation 520 performed on the target data set, the entire exemplary method 500, and the like.
[0103] In addition, in some implementations, the computing device 900 may also be equipped with one or more peripheral components, which may include but are not limited to a display, a speaker, a mouse, a keyboard, and the like. In addition, in some implementations, the computing device 900 may also be equipped with a communication interface, which may support various types of wired / wireless communication protocols to communicate with an external communication network. Examples of communication networks may include but are not limited to: a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a public telephone network, the Internet, an intranet, the Internet of Things, an infrared network, a Bluetooth network, a near field communication (NFC) network, and the like.
[0104] In addition, in some implementations, the above-mentioned and other components can communicate with each other via one or more buses / interconnects, which can support any suitable bus / interconnect protocols, including Peripheral Component Interconnect (PCI), PCI Express, Universal Serial Bus (USB), Serial Attached SCSI (SAS), Serial ATA (SATA), Fibre Channel (FC), System Management Bus (SMBus), or other suitable protocols.
[0105] Various implementations of the present disclosure may be implemented using hardware units, software units, or a combination thereof. Examples of hardware units may include devices, components, processors, microprocessors, circuits, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), storage units, logic gates, registers, semiconductor devices, chips, microchips, chipsets, etc. Examples of software units may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether an implementation is implemented using hardware elements and / or software elements can vary depending on a variety of factors, such as desired computational rates, power levels, thermal tolerances, processing cycle budgets, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints as desired for a given implementation.
[0106] Some implementations of the present disclosure may include articles of manufacture. Articles of manufacture may include storage media for storing logic. Examples of storage media may include one or more types of computer-readable storage media capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and the like. Examples of logic may include various software units, such as software components, programs, applications, computer programs, applications, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. In some implementations, for example, articles of manufacture may store executable computer program instructions, which, when executed by a processing unit, cause the processing unit to perform the methods and / or operations described herein. Executable computer program instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. Executable computer program instructions can be implemented according to a predefined computer language, method or syntax for commanding a computer to perform a specific function. The instructions can be implemented using any appropriate high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language.
[0107] What has been described above includes examples of the disclosed architecture. It is certainly not possible to describe every conceivable combination of components and / or methods, but those skilled in the art will appreciate that many other combinations and arrangements are possible. Therefore, the novel architecture is intended to encompass all such substitutions, modifications and variations that fall within the spirit and scope of the appended claims.
Claims
1. A method for labeling data points, comprising: Performing a labeling operation on a target data set, the target data set comprising a plurality of data points, each of the plurality of data points representing a business instance, the labeling operation comprising: Dividing the target data set into a plurality of first subsets; For each first subset of the plurality of first subsets: receiving a user input, the user input being used to assign a label to at least one data point in the first subset, wherein the label assigned to the data point is used to describe a situation of a business instance represented by the data point; determining whether a similarity between the label and a label previously assigned to at least one data point in the target data set satisfies a preset condition; In response to determining that the similarity does not satisfy the preset condition, re-performing the labeling operation with the first subset as the target data set; In response to determining that the similarity satisfies the preset condition, a tag associated with a tag previously assigned to at least one data point in the target data set is set for each data point in the first subset.
2. The method according to claim 1, further comprising: Dividing the initial data set into a plurality of second subsets; For each of the plurality of second subsets, receiving a user input for specifying a label for at least one data point in the second subset; A second subset among the plurality of second subsets is selected as the target data set.
3. The method according to claim 1 or 2, wherein: In response to determining that the similarity satisfies the preset condition, setting a tag associated with a tag previously assigned to at least one data point in the target data set for each data point in the first subset includes: A label previously assigned to at least one data point in the target data set is set as a label for each data point in the first subset.
4. The method according to claim 1 or 2, wherein: Determining whether the similarity between the label and a label previously assigned to at least one data point in the target data set satisfies a preset condition includes: determining a semantic similarity between the label and a label previously assigned to at least one data point in the target data set; and It is determined whether the semantic similarity is greater than a preset threshold.
5. A device for marking data points, comprising: A module for performing a labeling operation on a target data set, wherein the target data set includes a plurality of data points, each of the plurality of data points represents a business instance, and the module for performing a labeling operation on the target data set includes: A module for dividing the target data set into a plurality of first subsets; For each first subset of the plurality of first subsets: a module for receiving user input, the user input being used to assign a label to at least one data point in the first subset, wherein the label assigned to the data point is used to describe the situation of the business instance represented by the data point; a module for determining whether the similarity between the tag and a tag previously assigned to at least one data point in the target data set satisfies a preset condition; A module for re-performing the labeling operation with the first subset as a target data set in response to determining that the similarity does not satisfy the preset condition; Means for setting, in response to determining that the similarity satisfies the preset condition, for each data point in the first subset a label associated with a label previously assigned to at least one data point in the target data set.
6. The apparatus according to claim 5, further comprising: A module for partitioning the initial data set into a plurality of second subsets; means for receiving, for each of the plurality of second subsets, a user input for specifying a label for at least one data point in the second subset; A module is configured to select a second subset of the plurality of second subsets as the target data set.
7. The device according to claim 5 or 6, wherein: The module for setting, in response to determining that the similarity satisfies the preset condition, a tag associated with a tag previously assigned to at least one data point in the target data set for each data point in the first subset comprises: Means for setting a label previously assigned to at least one data point in the target data set as a label for each data point in the first subset.
8. The device according to claim 5 or 6, wherein: The module for determining whether the similarity between the label and the label previously assigned to at least one data point in the target data set satisfies a preset condition comprises: means for determining a semantic similarity between the label and a label previously assigned to at least one data point in the target data set; and A module for determining whether the semantic similarity is greater than a preset threshold.
9. A computing device comprising: A memory for storing instructions; as well as At least one processor is coupled to the memory, wherein when the instructions are executed by the at least one processor, the at least one processor performs a labeling operation on a target data set, the target data set includes a plurality of data points, each of the plurality of data points represents a business instance, and the labeling operation includes: Dividing the target data set into a plurality of first subsets; For each first subset of the plurality of first subsets: receiving a user input, the user input being used to assign a label to at least one data point in the first subset, wherein the label assigned to the data point is used to describe a situation of a business instance represented by the data point; determining whether a similarity between the label and a label previously assigned to at least one data point in the target data set satisfies a preset condition; In response to determining that the similarity does not satisfy the preset condition, re-performing the labeling operation with the first subset as the target data set; In response to determining that the similarity satisfies the preset condition, a tag associated with a tag previously assigned to at least one data point in the target data set is set for each data point in the first subset.
10. The computing device of claim 9, wherein: When the instructions are executed by the at least one processor, the at least one processor further causes: Dividing the initial data set into a plurality of second subsets; For each of the plurality of second subsets, receiving a user input for specifying a label for at least one data point in the second subset; A second subset among the plurality of second subsets is selected as the target data set.
11. The computing device according to claim 9 or 10, wherein: In response to determining that the similarity satisfies the preset condition, when setting a tag associated with a tag previously assigned to at least one data point in the target data set for each data point in the first subset, the at least one processor is configured to: A label previously assigned to at least one data point in the target data set is set as a label for each data point in the first subset.
12. The computing device according to claim 9 or 10, wherein: When determining whether the similarity between the label and a label previously assigned to at least one data point in the target data set satisfies a preset condition, the at least one processor is configured to: determining a semantic similarity between the label and a label previously assigned to at least one data point in the target data set; as well as It is determined whether the semantic similarity is greater than a preset threshold.
13. A computer-readable storage medium having instructions stored thereon, wherein when the instructions are executed by at least one processor, the at least one processor performs a labeling operation on a target data set, wherein the target data set includes a plurality of data points, each of the plurality of data points represents a business instance, and the labeling operation includes: Dividing the target data set into a plurality of first subsets; For each first subset of the plurality of first subsets: receiving a user input, the user input being used to assign a label to at least one data point in the first subset, wherein the label assigned to the data point is used to describe a situation of a business instance represented by the data point; determining whether a similarity between the label and a label previously assigned to at least one data point in the target data set satisfies a preset condition; In response to determining that the similarity does not satisfy the preset condition, re-performing the labeling operation with the first subset as the target data set; In response to determining that the similarity satisfies the preset condition, a tag associated with a tag previously assigned to at least one data point in the target data set is set for each data point in the first subset.
14. The computer-readable storage medium of claim 13, wherein: When the instructions are executed by at least one processor, the at least one processor further causes the at least one processor to: Dividing the initial data set into a plurality of second subsets; For each of the plurality of second subsets, receiving a user input for specifying a label for at least one data point in the second subset; A second subset among the plurality of second subsets is selected as the target data set.
15. The computer-readable storage medium according to claim 13 or 14, wherein: In response to determining that the similarity satisfies the preset condition, when setting a tag associated with a tag previously assigned to at least one data point in the target data set for each data point in the first subset, the at least one processor is configured to: A label previously assigned to at least one data point in the target data set is set as a label for each data point in the first subset.
16. The computer-readable storage medium according to claim 13 or 14, wherein: When determining whether the similarity between the label and a label previously assigned to at least one data point in the target data set satisfies a preset condition, the at least one processor is configured to: determining a semantic similarity between the label and a label previously assigned to at least one data point in the target data set; as well as It is determined whether the semantic similarity is greater than a preset threshold.
Citation Information
Patent Citations
Method and system for mining of patterns in a data set
CN104636422A
Accurate tag relevance prediction for image search
CN107085585A