Self-supervised learning method, device, equipment and medium for tabular data
By performing multiple data augmentation processing in the data augmentation layer of the self-supervised learning model, diversified target augmentation data is generated, which solves the problem of poor data augmentation effect in self-supervised learning, and improves the representation ability and accuracy of table tasks.
Patent Information
- Application Number
- CN202210820800.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-13
AI Technical Summary
The data enhancement effect of table data in self-supervised learning is poor, resulting in poor diversity of enhancement data, which in turn affects the effectiveness of self-supervised learning methods.
By introducing a data enhancement layer, including a first data enhancement unit and an embedding layer unit, a variety of data enhancement processes are performed, such as noise enhancement, mask enhancement and multi-view enhancement, and a second data enhancement process during the mapping process of the embedding layer unit, a diversified target enhancement data is generated.
It improves the diversity of tabular data, makes up for the insufficient correlation between time, space and context, and thus improves the representation ability of table tasks and the accuracy of downstream tasks.
Smart Images

Figure CN115130670B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and particularly to self-supervised learning methods, devices, equipment, and media for tabular data. Background Art
[0002] Self-supervised learning methods (Self-supervised Learning) have been widely applied to the fields of images, speech, and text. Self-supervised learning methods aim to, for unlabeled data, mine the representation characteristics of the data itself as supervision information by designing auxiliary tasks to improve the feature extraction ability of the model. Self-supervised learning methods can make full use of the temporal, spatial, and context information of images, speech, and text. However, for tabular data, such as click-through rate prediction data, claim data, insurance application data, etc., there are problems with insufficient correlation of temporal information, spatial information, or context information. In self-supervised learning, the data augmentation effect for tabular data is not good, resulting in poor diversity of the augmented data generated from tabular data, and thus the effect of self-supervised learning methods applied to tabular data is not good. Summary of the Invention
[0003] An object of this application is to solve at least to some extent one of the technical problems existing in the related art.
[0004] To this end, an object of an embodiment of this application is to provide a self-supervised learning method, device, equipment, and medium for tabular data. By performing various data augmentation processes on tabular data, the augmented data becomes more diverse, making up for the problem of insufficient association of time, space, and context of tabular data, providing a powerful representation ability for improving tabular tasks, and being conducive to improving the accuracy of downstream tabular tasks.
[0005] To achieve the above object, a first aspect of an embodiment of this application provides a self-supervised learning method for tabular data, including:
[0006] Obtain original tabular data, and input the original tabular data into a self-supervised learning model. The self-supervised learning model includes a data augmentation layer, and the data augmentation layer includes a first data augmentation unit and an embedding layer unit;
[0007] Perform various first data augmentation processes on the tabular data through the first data augmentation unit to obtain multiple augmented tabular data, where the dimensions of the augmented tabular data are the same as those of the original tabular data;
[0008] Map the tabular data and multiple augmented tabular data through the embedding layer unit, and perform a second data augmentation process during the mapping of the embedding layer unit to obtain multiple target augmented data, where the dimensions of the target augmented data are different from those of the augmented tabular data;
[0009] Performing self-supervised learning training on the self-supervised learning model according to multiple pieces of the target augmented data.
[0010] In some embodiments, the self-supervised learning model includes a backbone network layer and a loss function layer. Performing self-supervised learning training on the self-supervised learning model according to multiple pieces of the target augmented data includes:
[0011] Inputting each piece of the target augmented data into the backbone network layer, mining supervision information from the target augmented data, and obtaining representation information through training according to the supervision information;
[0012] Inputting multiple pieces of the representation information into the loss function layer, calculating a target loss function value, where the target loss function value represents the correlation between multiple classification results;
[0013] Adjusting the parameters of the self-supervised learning model according to the target loss function value.
[0014] In some embodiments, the target loss function value is expressed as: where τ is a temperature hyperparameter used to control the difficulty of the target loss function; z i is the i-th piece of representation information, z j is the j-th piece of representation information, z k is the k-th piece of representation information; N is the total number of pieces of representation information; II [k≠i] indicates selecting the representation information where k≠i.
[0015] In some embodiments, the tabular data includes index values and data values corresponding to the index values; the first data augmentation process includes noise augmentation processing. Performing noise augmentation processing on the tabular data includes:
[0016] Selecting at least one of the data values as first target data;
[0017] Adding noise information to the first target data to obtain the augmented tabular data.
[0018] In some embodiments, the tabular data includes index values and data values corresponding to the index values; the first data augmentation process includes mask augmentation processing. Performing mask augmentation processing on the tabular data includes:
[0019] Selecting at least one of the data values as second target data;
[0020] Performing a masking process on the second target data to obtain masked data;
[0021] Performing data prediction on the masked data according to the unmasked data in the table data to obtain predicted data;
[0022] The mask data is replaced by the prediction data to obtain the enhanced table data.
[0023] In some embodiments, the table data includes at least two index values and data values corresponding to the index values; the first data enhancement processing includes multi-view enhancement processing; performing multi-view enhancement processing on the table data includes:
[0024] Randomly select a number of at least two index values as first index values and the others as second index values;
[0025] For the table data, the first index value and the data value corresponding to the first index value are retained, and the second index value and the data value corresponding to the second index value are deleted to obtain the enhanced table data.
[0026] In some embodiments, the embedding layer includes an input neuron, a plurality of hidden neurons, and an output neuron; the table data and the plurality of enhanced table data are mapped by the embedding layer unit, and a second data enhancement process is performed during the mapping process of the embedding layer unit to obtain a plurality of target enhanced data, including:
[0027] Selecting a plurality of hidden neurons from the plurality of hidden neurons as target hidden neurons;
[0028] Fully connecting the input neuron with the target hidden neuron, and fully connecting the target hidden neuron with the output neuron to obtain a target embedding layer;
[0029] The table data and the plurality of enhanced table data are mapped through the target embedding layer to obtain the target enhanced data.
[0030] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a self-supervised learning device for tabular data, comprising:
[0031] An input module, used to obtain tabular data, and input the tabular data into a self-supervised learning model, wherein the self-supervised learning model includes a data enhancement layer, and the data enhancement layer includes a first data enhancement unit and an embedding layer unit;
[0032] A first data enhancement module, configured to perform a plurality of first data enhancement processes on the table data through the first data enhancement unit to obtain a plurality of enhanced table data, wherein the dimension of the enhanced table data is the same as the dimension of the original table data;
[0033] A second data augmentation module, configured to map the tabular data and the multiple augmented tabular data through an embedding layer unit, and perform second data augmentation processing during the mapping of the embedding layer unit to obtain multiple target augmented data, where the dimension of the target augmented data is different from the dimension of the augmented tabular data;
[0034] A training module, configured to perform self-supervised learning training on the self-supervised learning model according to the multiple target augmented data.
[0035] To achieve the above object, a third aspect of the embodiments of the present application further provides an electronic device, where the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory. When the program is executed by the processor, the self-supervised learning method for tabular data described above is implemented.
[0036] To achieve the above object, a fourth aspect of the embodiments of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the self-supervised learning method for tabular data described above.
[0037] The self-supervised learning method, device, equipment and medium for tabular data disclosed in the embodiments of the present application obtain original tabular data, input the original tabular data into a self-supervised learning model, and the self-supervised learning model includes a data augmentation layer, and the data augmentation layer includes a first data augmentation unit and an embedding layer unit; perform multiple first data augmentation processes on the tabular data through the first data augmentation unit to obtain multiple augmented tabular data; map the tabular data and the multiple augmented tabular data through the embedding layer unit, and perform second data augmentation processing during the mapping of the embedding layer unit to obtain multiple target augmented data, where the dimension of the target augmented data is lower than the dimension of the augmented tabular data; perform self-supervised learning training on the self-supervised learning model according to the multiple target augmented data; through performing multiple data augmentation processes on the tabular data, the augmented data is made more diverse, making up for the lack of association in time, space and context of the tabular data, providing a powerful representation ability for improving tabular tasks, and being beneficial to improving the accuracy of downstream tabular tasks. Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the attached drawings related to the technical solutions in the embodiments of the present application or the prior art. It should be understood that the attached drawings below are only for conveniently and clearly presenting some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a structural diagram of the self-supervised learning model provided by the embodiments of the present application;
[0040] Figure 2 It is a schematic diagram of the second data augmentation processing provided by the embodiments of the present application;
[0041] Figure 3 It is a step diagram of the self-supervised learning method for tabular data provided by the embodiments of the present application;
[0042] Figure 4 It is a step diagram of the noise augmentation processing provided by the embodiments of the present application;
[0043] Figure 5 It is a step diagram of the mask augmentation processing provided by the embodiments of the present application;
[0044] Figure 6 It is a step diagram of the multi-view augmentation processing provided by the embodiments of the present application;
[0045] Figure 7 It is a step diagram of step S300 provided by the embodiments of the present application;
[0046] Figure 8 It is a step diagram of step S400 provided by the embodiments of the present application;
[0047] Figure 9 It is a structural diagram of the self-supervised learning device for tabular data provided by the embodiments of the present application;
[0048] Figure 10 It is a structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0049] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0050] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. The terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.
[0052] First, several terms involved in this application are analyzed:
[0053] Artificial Intelligence (AI): It is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0054] Natural Language Processing (NLP): It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. The natural language involved in this field is the language used by people in daily life, so it is also closely related to the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0055] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning (deep learning) usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0056] In recent years, with the rapid development of artificial intelligence technology, various types of machine learning models have achieved relatively good application effects in fields such as image classification, face recognition, and autonomous driving. Among them, self-supervised learning methods have been widely applied to the fields of images, speech, and text. The self-supervised learning method aims to, for unlabeled data, design auxiliary tasks to mine the inherent representation characteristics of the data as supervision information to improve the feature extraction ability of the model. The self-supervised learning method can make full use of the temporal, spatial, and context information of images, speech, and text. For example, for a picture, there is similar spatial information in the surrounding area of any pixel point in the picture; for a piece of speech, there is time series information in each part of the speech; for a piece of text, there is context information between each sentence or word in the text, and the above information can help understand the following information. Temporal, spatial, and context information can help self-supervised learning mine the inherent representation characteristics of the data.
[0057] However, for tabular data, such as click-through rate prediction data, claim data, insurance application data, etc., there are problems with insufficient relevance of temporal information, spatial information, or context information. In self-supervised learning, the data augmentation effect for tabular data is poor, resulting in poor diversity of the augmented data generated from tabular data, and thus the effect of the self-supervised learning method applied to tabular data is poor.
[0058] To solve the problems in the related art, an object of an embodiment of the present application is to provide a self-supervised learning method, device, equipment, and medium for tabular data. By obtaining the original tabular data and inputting the original tabular data into a self-supervised learning model, the self-supervised learning model includes a data augmentation layer, and the data augmentation layer includes a first data augmentation unit and an embedding layer unit; performing multiple first data augmentation processes on the tabular data through the first data augmentation unit to obtain multiple augmented tabular data; mapping the tabular data and the multiple augmented tabular data through the embedding layer unit, and performing a second data augmentation process during the mapping process of the embedding layer unit to obtain multiple target augmented data, where the dimension of the target augmented data is lower than the dimension of the augmented tabular data; performing self-supervised learning training on the self-supervised learning model according to the multiple target augmented data; through performing multiple data augmentation processes on the tabular data, making the augmented data more diverse, making up for the problem of insufficient association of time, space, and context of the tabular data, providing a powerful representation ability for improving tabular tasks, and being beneficial to improving the accuracy of downstream tabular tasks.
[0059] The implementation environment of a self-supervised learning method for tabular data provided by an embodiment of this application is as follows. The main software and hardware entities of this implementation environment mainly include an operating terminal or a server. Among them, this self-supervised learning method for tabular data can be configured to execute on the operating terminal alone, or configured to execute on the server alone, or executed based on the interaction between the operating terminal and the server. Specifically, it can be appropriately selected according to the actual application situation, and this embodiment does not make specific limitations on this. In addition, the operating terminal and the server can be nodes in the blockchain, and this embodiment does not make specific limitations on this.
[0060] Specifically, the operating terminal in this application may include, but is not limited to, any one or more of a smart watch, a smart phone, a computer, a personal digital assistant (PDA), a smart voice interaction device, a smart home appliance, or a vehicle-mounted terminal. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. A communication connection can be established between the operating terminal 101 and the server 102 through a wireless network or a wired network. The wireless network or the wired network uses standard communication technologies and / or protocols. The network can be set to the Internet or any other network, such as including, but not limited to, any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network, or a virtual private network.
[0061] In addition, the present application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0062] Embodiments of the present application provide a self-supervised learning method for tabular data, which performs self-supervised learning on the original tabular data through a self-supervised learning model. Refer to Figure 1 , Figure 1 which is a structural diagram of the self-supervised learning model. The self-supervised learning model includes a data augmentation layer 100, a backbone network layer 200, and a loss function layer 300. The data augmentation layer 100, the backbone network layer 200, and the loss function layer 300 are connected in sequence. The data augmentation layer 100 includes a first data augmentation unit 110 and a second data augmentation unit 120, and the second data augmentation unit 120 includes an embedding layer unit 130.
[0063] Refer to Figure 3 , Figure 3 which is a step diagram of the self-supervised learning method for tabular data. The self-supervised learning method for tabular data includes but is not limited to the following steps:
[0064] Step S100, obtain the original tabular data and input the original tabular data into the self-supervised learning model;
[0065] Step S200, perform multiple first data augmentation processes on the tabular data through the first data augmentation unit 110 to obtain multiple augmented tabular data;
[0066] Step S300, map the tabular data and the multiple augmented tabular data through the embedding layer unit 130, and perform a second data augmentation process during the mapping of the embedding layer unit 130 to obtain multiple target augmented data;
[0067] Step S400, perform self-supervised learning training on the self-supervised learning model according to the multiple target augmented data.
[0068] For step S100, the original tabular data can be obtained by receiving tabular data input by the user through an input device, or by retrieving it from a large database.
[0069] The original tabular data is data stored in tabular form, which can include click-through rate prediction data, claim settlement data, insurance application data, etc. The tabular form includes multiple rows and columns, with the rows and columns neatly arranged.
[0070] The original tabular data includes at least two index values and at least one data value corresponding to the index values. Usually, the first row of the table stores multiple index values, and then the subsequent rows store the data values corresponding to each index value respectively; or, the first column of the table stores multiple index values, and then the subsequent columns store the data values corresponding to each index value respectively.
[0071] For example, the form of the original tabular data is shown in Table 1.
[0072] Table 1 Example Table of Original Tabular Data
[0073] Initial diagnosis disease Warning line Whether it is a designated hospital Department code Reason for application Insurance type code N81.1 Y Hospitals within the scope 17.0 Accident, disease 521,1020
[0074] For this original tabular data, the first row of the table stores multiple index values, including the first diagnosed disease, warning line, whether it is a designated hospital, department code, application reason, and insurance type code.
[0075] The second row of the table stores multiple data values. The data value corresponding to the index value of the first diagnosed disease is "N81.1", the data value corresponding to the index value of the warning line is "Y", the data value corresponding to the index value of whether it is a designated hospital is "hospital within the range", the data value corresponding to the index value of the department code is "17.0", the data value corresponding to the index value of the application reason is "accident, disease", and the data value corresponding to the index value of the insurance type code is "521, 1020".
[0076] For step S200, after the tabular data is input into the self-supervised learning model, the tabular data first undergoes preprocessing through the data augmentation layer 100, that is, data augmentation processing. In the data augmentation layer 100, the tabular data first undergoes the first data augmentation processing through the first data augmentation unit 110. The first data augmentation unit 110 performs various first data augmentation processes on the tabular data respectively to obtain multiple augmented tabular data corresponding to the first data augmentation processing methods.
[0077] Among them, in this embodiment, the first data augmentation processing includes noise augmentation processing, mask augmentation processing, and multi-view augmentation processing. Of course, in other embodiments, more first data augmentation processing methods can also be adopted, or other different types of first data augmentation processing methods can be adopted according to actual needs.
[0078] It should be noted that the first data augmentation process directly performs data augmentation on the original tabular data at the original tabular data level; the dimension of the output augmented tabular data is the same as that of the input original tabular data.
[0079] Refer to Figure 4 , Figure 4 is the step diagram of the noise augmentation process. For the first data augmentation process being the noise augmentation process; then performing noise augmentation on the tabular data includes, but is not limited to, the following steps:
[0080] Step S211, select at least one data value as the first target data;
[0081] Step S212, add noise information to the first target data to obtain the augmented tabular data.
[0082] For example, perform noise augmentation on the original tabular data in Table 1. Select "N81.1" as the first target data, add a noise information to "N81.1" so that "N81.1" becomes "N81.2", and obtain the augmented tabular data. The augmented tabular data obtained by performing noise augmentation on the original tabular data is shown in Table 2.
[0083] Table 2 Example table of the augmented tabular data obtained by performing noise augmentation on the original tabular data
[0084] Initial diagnosis disease Warning line Whether it is a designated hospital Department code Reason for application Insurance type code N81.2 Y Hospitals within the scope 17.0 Accident, disease 521,1020
[0085] Of course, in other embodiments, other noise information can be added to "N81.1" selected as the first target data, so that "N81.1" becomes data such as "N81.4", "N82.4", or "M81.1", and obtain multiple augmented tabular data obtained by performing noise augmentation.
[0086] Of course, in other embodiments, noise augmentation can also be performed on one or more of other data values, that is, the data value of "Y", the data value of "hospitals within the range", the data value of "17.0", the data value of "accident, disease", and the data value of "521,1020", to obtain multiple augmented tabular data obtained by performing noise augmentation.
[0087] Refer to Figure 5 , Figure 5 is the step diagram of the mask augmentation process. For the first data augmentation process being the mask augmentation process, then performing mask augmentation on the tabular data includes:
[0088] Step S221, select at least one data value as the second target data;
[0089] Step S222: Perform masking processing on the second target data to obtain masked data;
[0090] Step S223: Perform data prediction on the masked data according to the unmasked data in the tabular data to obtain predicted data;
[0091] Step S224: Replace the masked data with the predicted data to obtain enhanced tabular data.
[0092] For example, perform masking enhancement processing on the original tabular data in Table 1. Select "N81.1" as the second target data, perform masking processing on "N81.1" to obtain masked data, that is, mark it as "Masked", and the tabular data with masked data is shown in Table 3.
[0093] Table 3 Example table of tabular data with masked data
[0094] Initial diagnosis disease Warning line Whether it is a designated hospital Department code Reason for application Insurance type code Masked Y Hospitals within the scope 17.0 Accident, disease 521,1020
[0095] Then, according to the unmasked data in the tabular data, that is, the index values of "primary diagnosis disease", the index value of "warning line", the index value of "whether it is a designated hospital", the index value of "department code", the index value of "application reason", the index value of "insurance type code", the data value "Y" corresponding to the "warning line" index value, the data value "hospital within the range" corresponding to the "whether it is a designated hospital" index value, the data value "17.0" corresponding to the "department code" index value, the data value "accident, disease" corresponding to the "application reason" index value, and the data values "521, 1020" corresponding to the "insurance type code" index value, perform data prediction on the masked data "Masked" through a data prediction model to obtain predicted data "N82.1". Replace the masked data "Masked" with the predicted data "N82.1" to obtain an example table of enhanced tabular data obtained by masking enhancement processing of the original tabular data as shown in Table 4.
[0096] Table 4 Example table of enhanced tabular data obtained by masking enhancement processing of the original tabular data
[0097] Initial diagnosis disease Warning line Whether it is a designated hospital Department code Reason for application Insurance type code N82.1 Y Hospitals within the scope 17.0 Accident, disease 521,1020
[0098] Of course, in practical applications, the tabular data includes a large number of index values and data values. Generally, the masked data ratio is set to 15%, and the unmasked data ratio is 85%. When the masked data ratio takes a value of 15%, for 7 cell data, one of the data is masked. Exactly the central word of the sliding window with a length of 7 in CBOW, which is beneficial to improving the prediction effect of the data prediction model.
[0099] In addition, mask enhancement processing includes the following, mainly word masking, full word masking, entity masking, N-gram masking and Span masking.
[0100] Among them, for full-word masking, full-word masking uses the word segmentation result as the minimum granularity to complete the masking task. For N-gram masking, N-gram masking uses the word segmentation result as the minimum granularity and uses n-gram to take words for masking. For entity masking, entity masking introduces named entity information and uses entities as the minimum granularity for masking. For Span masking, Span masking first randomly selects the length of a segment based on geometric distribution, then randomly selects the starting position of this segment based on uniform distribution, and finally masks according to the length.
[0101] Reference Figure 6 , Figure 6 , which is a step diagram of multi-view enhancement processing. If the first data enhancement processing is multi-view enhancement processing, multi-view enhancement processing is performed on the table data, including:
[0102] Step S231, randomly selecting a number of at least two index values as first index values, and the others as second index values;
[0103] Step S232: for the table data, retain the first index value and the data value corresponding to the first index value, and remove the second index value and the data value corresponding to the second index value to obtain enhanced table data.
[0104] For example, the original table data of Table 1 is enhanced from multiple perspectives. The index value of "first diagnosis disease", the index value of "department code", the index value of "application reason", and the index value of "insurance type code" are selected as the first index value; the index value of "warning line" and the index value of "whether it is a designated hospital" are selected as the second index value.
[0105] The index value of "first-visit disease", the index value of "department code", the index value of "application reason", the index value of "insurance type code", the data value "N81.1" corresponding to the index value of "first-visit disease", the data value "17.0" corresponding to the index value of "department code", the data value "accident, disease" corresponding to the index value of "application reason", and the data value "521,1020" corresponding to the index value of "insurance type code" are retained, and the index value of "warning line", the index value of "whether it is a designated hospital", the data value "Y" corresponding to the index value of "warning line", and the data value "hospital within the range" corresponding to the index value of "whether it is a designated hospital" are eliminated, and the enhanced table data example table obtained by multi-view enhancement processing of the original table data shown in Table 5 is obtained.
[0106] Table 5 Example of enhanced table data obtained by multi-view enhancement processing of original table data
[0107] Initial diagnosis disease Department code Reason for application Insurance type code N82.1 17.0 Accident, disease 521,1020
[0108] Of course, in other embodiments, other index values and data values can also be removed to form enhanced table data with multiple perspectives.
[0109] For step S300, the table data and multiple enhanced table data output by the first data enhancement unit 110 are input into the second data enhancement unit 120; in the second data enhancement unit 120, the table data and multiple enhanced table data are mapped through the embedding layer unit 130, and the second data enhancement process is performed by the dropout method during the mapping process of the embedding layer unit 130 to obtain multiple target enhanced data.
[0110] Refer to Figure 7 , specifically, the table data and multiple enhanced table data are mapped through the embedding layer unit 130, and the second data enhancement process is performed during the mapping process of the embedding layer unit 130 to obtain multiple target enhanced data, including but not limited to the following steps:
[0111] Step S310, select several from multiple hidden neurons 132 as target hidden neurons 134;
[0112] Step S320, fully connect the input neuron 131 with the target hidden neuron 134, and fully connect the target hidden neuron 134 with the output neuron 133 to obtain the target embedding layer 135;
[0113] Step S330, map the table data and multiple enhanced table data through the target embedding layer 135 to obtain the target enhanced data.
[0114] Among them, for the embedding layer unit 130, the embedding layer unit 130 adopts an embedding layer structure, which is composed of multiple neurons, namely an input neuron 131, hidden neurons 132, and an output neuron 133. The total number of neurons is embedding_size, and this parameter embedding_size can be adjusted according to actual needs. The input neuron 131 and the hidden neurons 132 adopt a fully connected connection method, and the hidden neurons 132 and the output neuron 133 adopt a fully connected connection method. Word embedding mapping is implemented through the embedding layer unit 130, which is used to map the table data and multiple enhanced table data output by the first data enhancement unit 110 into real number vectors according to the vocabulary, and the hidden neurons 132 are responsible for calculating the parameter matrix.
[0115] The embedding layer essentially performs matrix multiplication on the input data, reducing the dimension of the input data to obtain output data with a lower dimension. Since the original data has insufficient information density and there is a lack of correlation between features, the embedding layer re-encodes the previous information into information with a higher density and correlations between features.
[0116] For the embedding layer, the input data can be regarded as a sparse matrix, and the output data can be regarded as a dense matrix. The embedding layer transforms the sparse matrix into a dense matrix through some linear transformations; this dense matrix uses N features to represent all words. In this dense matrix, ostensibly representing a one-to-one correspondence between the dense matrix and a single word, it actually also contains a large number of internal relationships between words, between words and phrases, and even between sentences and sentences. This internal relationship is represented by the parameters learned by the embedding layer. More importantly, this relationship is continuously updated during the backpropagation process. Therefore, after multiple trainings, this relationship becomes relatively mature, that is, it can correctly express the entire semantics and the relationships between various sentences. This relationship can actually be represented by all the weight parameters of the embedding layer.
[0117] From the perspective of dimension, after defining the dimension of the embedding layer as n, by multiplying the input tabular data or enhanced tabular data with the parameter matrix, all words in the dictionary can be represented by n dimensions, which is equivalent to feature extraction. Therefore, the dimension of the target enhanced data is different from both the dimension of the tabular data and the first data enhancement unit 110. In this embodiment, the embedding layer unit 130 reduces the dimension of the tabular data and multiple enhanced tabular data, and the dimension of the target enhanced data is lower than the dimension of the tabular data and the first data enhancement unit 110.
[0118] During the mapping process of the embedding layer unit 130, second data enhancement processing is performed by the dropout method, which can also avoid overfitting. For dropout, by setting the node values of some hidden neurons 132 to 0, the interaction between the hidden neurons 132 is reduced; that is, during the forward propagation, the activation value of a certain neuron stops working with a certain probability p, which can make the embedding layer unit 130 more generalizable.
[0119] Data augmentation is achieved by performing dropout on the embedding layer unit 130, which has the effect of taking an average. Since different networks may produce different overfittings, taking an average may cause some "opposite" fittings to cancel each other out. Dropping out different hidden neurons 132 is similar to training different embedding layer units 130. Randomly deleting half of the hidden neurons 132 results in different network structures of the embedding layer units 130. The entire dropout process is equivalent to taking an average of many different neural networks. Different networks produce different overfittings, and some "opposite" fittings cancel each other out, which can reduce overfitting as a whole.
[0120] In addition, data augmentation is achieved by performing dropout on the embedding layer unit 130, which has the effect of reducing the complex co-adaptation relationship between neurons. Because dropout causes two hidden neurons 132 not to necessarily appear in a network of an embedding layer unit 130 every time. In this way, the update of the weights no longer depends on the joint action of hidden nodes with fixed relationships, preventing the situation where some features are only effective under other specific features, forcing the network to learn more robust features, and reducing the weights to improve the robustness of the network to the loss of specific neuron connections. This makes the embedding layer unit 130 not overly sensitive to some specific clue segments of tabular data. Even if specific clue segments of tabular data are lost, the embedding layer unit 130 can learn some common features from many other clues, improving the robustness of the embedding layer unit 130 to the loss of specific neuron connections.
[0121] Refer to Figure 2 , for example, in some embodiments, the embedding layer unit 130 includes two input neurons 131, four hidden neurons 132, and three output neurons 133, where the four hidden neurons 132 are respectively hidden neuron 132A, hidden neuron 132B, hidden neuron 132C, and hidden neuron 132D. Of course, in other embodiments, the embedding layer can also be other network structures with more neurons.
[0122] For the initial embedding layer unit 130, two input neurons 131 are fully connected to hidden neurons 132A, hidden neuron 132B, hidden neuron 132C, and hidden neuron 132D respectively, and hidden neurons 132A, hidden neuron 132B, hidden neuron 132C, and hidden neuron 132D are fully connected to three output neurons 133 respectively. During the dropout process, any several of hidden neurons 132A, hidden neuron 132B, hidden neuron 132C, and hidden neuron 132D are selected as target hidden neurons 134, such as hidden neurons 132A and hidden neuron 132C. Of course, in other embodiments, other hidden neurons 132 can also be selected. Two input neurons 131 are fully connected to hidden neurons 132A and hidden neuron 132C respectively, and hidden neurons 132A and hidden neuron 132C are fully connected to three output neurons 133 respectively to obtain a target embedding layer 135. The target enhanced data is obtained by mapping the tabular data and multiple enhanced tabular data through the target embedding layer 135.
[0123] Referring to Figure 8 , for step S400, self-supervised learning training is performed on the self-supervised learning model according to multiple target enhanced data, including but not limited to the following steps:
[0124] Step S410, input each target enhanced data into the backbone network layer 200, mine the supervision information from the target enhanced data, and obtain the representation information according to the supervision information for training;
[0125] Step S420, input multiple representation information into the loss function layer 300, calculate the target loss function value, and the target loss function value represents the correlation between multiple classification results;
[0126] Step S430, adjust the parameters of the self-supervised learning model according to the target loss function value.
[0127] In this embodiment, the main part of self-supervised learning is implemented through the backbone network layer 200. The backbone network layer 200 can adopt a CNN network model or a transformer network model.
[0128] For step S410, in the backbone network layer 200, the supervision information is mined from the target enhanced data through an auxiliary task related to the tabular data, and the representation information is obtained according to the supervision information for training.
[0129] For step S420, input multiple representation information into the loss function layer 300, and calculate the target loss function value. Specifically, the target loss function value is expressed as: where τ is a temperature hyperparameter used to control the difficulty of the target loss function; zi is the i-th characterization information, z j is the j-th characterization information, z k is the k-th characterization information; N is the total number of characterization information; II [k≠i] represents selecting the characterization information where k≠i.
[0130] Specifically, multiple table data are batch-processed to form a batch of data, and the self-supervised learning model processes the data in batches. The batchsize is the number of sample data processed in each batch. For an original table data, it undergoes various first data augmentation processes on the table data by the first data augmentation unit 110 to obtain multiple augmented table data, where two of the augmented table data are x i and x j , x i and x j are mapped through the embedding layer unit 130 and second data augmentation processing is performed during the mapping process of the embedding layer unit 130 to obtain the target augmented data x′ i and x′ j ; x′ i and x′ j correspondingly obtain the characterization information z i and z j . Then z k is the other characterization information in this batch of data. This batch of data is input into the loss function layer 300 to calculate the target loss function value.
[0131] For step S430, calculate the network gradient of the self-supervised learning model and backpropagate it in reverse. According to the target loss function value, adjust the parameters of the self-supervised learning model by the gradient descent method. Until the target loss function value is minimized, the training of the self-supervised learning model is completed.
[0132] To achieve the above object, an embodiment of the present application provides a self-supervised learning device for table data. Refer to Figure 9 , Figure 9 is the structural diagram of the self-supervised learning device for table data. The self-supervised learning device for table data includes an input module 510, a first data augmentation module 520, a second data augmentation module 530, and a training module 540.
[0133] Among them, the input module 510 is used to obtain tabular data and input the tabular data into the self-supervised learning model. The self-supervised learning model includes a data augmentation layer 100, and the data augmentation layer 100 includes a first data augmentation unit 110 and an embedding layer unit 130. The first data augmentation module 520 is used to perform a variety of first data augmentation processes on the tabular data through the first data augmentation unit 110 to obtain multiple augmented tabular data, where the dimension of the augmented tabular data is the same as that of the original tabular data. The second data augmentation module 530 is used to map the tabular data and multiple augmented tabular data through the embedding layer unit 130, and perform a second data augmentation process during the mapping process of the embedding layer unit 130 to obtain multiple target augmented data, where the dimension of the target augmented data is different from that of the augmented tabular data. The training module 540 is used to perform self-supervised learning training on the self-supervised learning model according to multiple target augmented data.
[0134] In this embodiment, the self-supervised learning device for tabular data uses the input module 510, the first data augmentation module 520, the second data augmentation module 530, and the training module 540. Through the first data augmentation unit 110, a variety of first data augmentation processes are performed on the tabular data to obtain multiple augmented tabular data; the tabular data and multiple augmented tabular data are mapped through the embedding layer unit 130, and a second data augmentation process is performed during the mapping process of the embedding layer unit 130 to obtain multiple target augmented data; self-supervised learning training is performed on the self-supervised learning model according to multiple target augmented data; by performing a variety of data augmentation processes on the tabular data, the augmented data is made more diverse, making up for the lack of association in time, space, and context of the tabular data, providing a powerful representation ability for improving tabular tasks, and being beneficial to improving the accuracy of downstream tabular tasks.
[0135] It can be understood that the content in the self-supervised learning method embodiment of tabular data is applicable to the self-supervised learning device embodiment of this tabular data. The functions specifically implemented by the self-supervised learning device embodiment of this tabular data are the same as those of the self-supervised learning method embodiment of tabular data, and the beneficial effects achieved are also the same as those of the self-supervised learning method embodiment of tabular data.
[0136] To achieve the above object, an embodiment of the present application also provides an electronic device. Refer to Figure 10 , Figure 10 is the structural diagram of the electronic device. The electronic device includes a memory 620, a processor 610, a program stored on the memory 620 and executable on the processor 610, and a data bus 630 for realizing the connection and communication between the processor 610 and the memory 620. When the program is executed by the processor 610, the above self-supervised learning method for tabular data is realized.
[0137] In this embodiment, by obtaining the original table data and inputting the original table data into a self-supervised learning model, the self-supervised learning model includes a data augmentation layer, and the data augmentation layer includes a first data augmentation unit and an embedding layer unit; performing various first data augmentation processes on the table data through the first data augmentation unit to obtain multiple augmented table data; mapping the table data and the multiple augmented table data through the embedding layer unit, and performing a second data augmentation process during the mapping process of the embedding layer unit to obtain multiple target augmented data, where the dimension of the target augmented data is lower than that of the augmented table data; performing self-supervised learning training on the self-supervised learning model according to the multiple target augmented data; by performing various data augmentation processes on the table data, the augmented data is made more diverse, compensating for the lack of association in time, space, and context of the table data, providing a powerful representation ability for improving table tasks, and being beneficial to improving the accuracy of downstream table tasks.
[0138] The memory 620, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the self-supervised learning method for table data in the above embodiments of the present invention. The processor 610 realizes the self-supervised learning method for table data in the above embodiments of the present invention by running the non-transitory software programs and programs stored in the memory 620.
[0139] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data required for executing the self-supervised learning method for table data in the above embodiments of the present invention, etc. In addition, the memory 620 may include a high-speed random access memory 620, and may also include a non-transitory memory 620, such as at least one disk memory 620 device, a flash memory device, or other non-transitory solid-state memory 620 devices. In some embodiments, the memory 620 may optionally include a memory 620 remotely disposed relative to the processor 610, and these remote memories 620 can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0140] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above self-supervised learning method for table data.
[0141] In this embodiment, by obtaining the original table data and inputting the original table data into a self-supervised learning model, the self-supervised learning model includes a data augmentation layer, and the data augmentation layer includes a first data augmentation unit and an embedding layer unit; performing a variety of first data augmentation processes on the table data through the first data augmentation unit to obtain multiple augmented table data; mapping the table data and the multiple augmented table data through the embedding layer unit, and performing a second data augmentation process during the mapping process of the embedding layer unit to obtain multiple target augmented data, where the dimension of the target augmented data is lower than the dimension of the augmented table data; performing self-supervised learning training on the self-supervised learning model according to the multiple target augmented data; by performing a variety of data augmentation processes on the table data, the augmented data is made more diverse, compensating for the lack of association in time, space, and context of the table data, providing a powerful representation ability for improving table tasks, and being beneficial to improving the accuracy of downstream table tasks.
[0142] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium. In the foregoing description of the present specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0143] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
[0144] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
[0145] In the description of this specification, the description referring to terms such as "one embodiment", "another embodiment", or "certain embodiments" means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
Claims
1. A self-supervised learning method for tabular data, characterized in that, Including: Obtain original tabular data, and input the original tabular data into a self-supervised learning model. The self-supervised learning model includes a data augmentation layer, and the data augmentation layer includes a first data augmentation unit and an embedding layer unit. The embedding layer unit includes an input neuron, a plurality of hidden neurons, and an output neuron; Perform multiple first data augmentation processes on the original tabular data through the first data augmentation unit to obtain a plurality of augmented tabular data, where the dimension of the augmented tabular data is the same as that of the original tabular data; wherein, the first data augmentation process includes at least one of the following: noise augmentation process, masking augmentation process, and multi-perspective augmentation process; Select several from the plurality of hidden neurons as target hidden neurons; Fully connect the input neuron and the target hidden neurons, and fully connect the target hidden neurons and the output neuron to obtain a target embedding layer; Map the original tabular data and the plurality of augmented tabular data through the target embedding layer to obtain a plurality of target augmented data, where the dimension of the target augmented data is different from that of the augmented tabular data; Perform self-supervised learning training on the self-supervised learning model according to the plurality of target augmented data.
2. The self-supervised learning method for tabular data according to claim 1, wherein The self-supervised learning model includes a backbone network layer and a loss function layer. The performing self-supervised learning training on the self-supervised learning model according to the plurality of target augmented data includes: Input each of the target augmented data into the backbone network layer, mine supervision information from the target augmented data, and perform training according to the supervision information to obtain representation information; Input the plurality of representation information into the loss function layer, calculate a target loss function value, and the target loss function value represents the correlation between a plurality of classification results; Adjust the parameters of the self-supervised learning model according to the target loss function value.
3. The self-supervised learning method for tabular data according to claim 2, characterized in that The target loss function value is expressed as: ; where, is a temperature hyperparameter for controlling the difficulty of the target loss function; is the i-th characterization information, is the j-th characterization information, is the k-th characterization information; N is the total number of characterization information; denotes the selection of the characterization information of i.
4. A self-supervised learning method for tabular data according to claim 1, wherein, The original tabular data includes index values and data values corresponding to the index values; The first data augmentation process includes the noise augmentation process; Performing a noise augmentation process on the original tabular data includes: Select at least one of the data values as first target data; Add noise information to the first target data to obtain the augmented tabular data.
5. A self-supervised learning method for tabular data according to claim 1, wherein The original tabular data includes index values and data values corresponding to the index values; The first data augmentation process includes the masking augmentation process; Performing a masking augmentation process on the original tabular data includes: Select at least one of the data values as second target data; Perform a masking process on the second target data to obtain masked data; Perform data prediction on the masked data according to the unmasked data in the original tabular data to obtain predicted data; Replace the masked data with the predicted data to obtain the augmented tabular data.
6. A self-supervised learning method for tabular data according to claim 1, characterized in that The original tabular data includes at least two index values and data values corresponding to the index values; The first data augmentation process includes the multi-perspective augmentation process; Performing a multi-perspective augmentation process on the original tabular data includes: Randomly select a number of at least two index values as first index values and the others as second index values; For the original table data, the first index value and the data value corresponding to the first index value are retained, and the second index value and the data value corresponding to the second index value are deleted to obtain the enhanced table data.
7. A self-supervised learning device for tabular data, characterized in that, include: An input module, used to obtain original table data, and input the original table data into a self-supervised learning model, wherein the self-supervised learning model includes a data enhancement layer, wherein the data enhancement layer includes a first data enhancement unit and an embedding layer unit, wherein the embedding layer unit includes an input neuron, a plurality of hidden neurons, and an output neuron; A first data enhancement module, configured to perform a plurality of first data enhancement processes on the original table data through the first data enhancement unit to obtain a plurality of enhanced table data, wherein the dimension of the enhanced table data is the same as the dimension of the original table data; wherein the first data enhancement process comprises at least one of the following: noise enhancement process, mask enhancement process and multi-view enhancement process; A second data enhancement module is used to select a number of hidden neurons from the plurality of hidden neurons as target hidden neurons, fully connect the input neurons with the target hidden neurons, and fully connect the target hidden neurons with the output neurons to obtain a target embedding layer, and map the original table data and the plurality of enhanced table data through the target embedding layer to obtain a plurality of target enhanced data, wherein the dimension of the target enhanced data is different from the dimension of the enhanced table data; A training module is used to perform self-supervised learning training on the self-supervised learning model according to the multiple target enhancement data.
8. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the self-supervised learning method for tabular data as described in any one of claims 1 to 6 is realized.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the self-supervised learning method for tabular data according to any one of claims 1 to 6.