Data processing method and device

By introducing instance set-level tags in the pre-training stage and using pseudo-negative pairs to establish matching relationships in the fine-tuning stage, the problem of improving the prediction performance of neural network models in tabular data analysis while protecting user privacy is solved, and more efficient personalized data processing and prediction are achieved.

CN120354907APending Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410080966.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

On the premise of protecting user privacy, how to obtain sufficient information from limited aggregation information to provide personalized data services and improve the prediction performance of neural network models, especially when processing table data, it is difficult for the existing technology to effectively use aggregation information for training.

Method used

By obtaining the first instance set and the second instance set, pre-training is performed using the instance set-level tags, combining the pseudo-oriented pair to introduce instance-level supervision information in the fine-tuning stage, establish a matching relationship between the first instance and the second instance, guide the training of the neural network model, and improve the robustness and prediction performance of the model.

Benefits of technology

Without augmenting the original instance set, the prediction performance and robustness of the neural network model in tabular data analysis are improved, especially under LLP tasks, which can predict instance-level information more accurately, enhancing the model's category perception ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354907A_ABST
    Figure CN120354907A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method, a data processing device, electronic equipment and a computer readable storage medium. The method comprises the steps that a first instance set and a second instance set are obtained, the first instance set comprises a plurality of first instances, the second instance set comprises a plurality of second instances, the first instance set is provided with first tags at the instance set level, and the second instance set is provided with second tags at the instance set level; pre-training a neural network model by taking the first label and the second label as supervision information; fine-tuning training is carried out on the pre-trained neural network model, in each time step of the fine-tuning training, a plurality of pseudo-pairs are determined, and each pseudo-pair comprises a first instance and a second instance matched with the first instance; and adjusting the parameters of the neural network model by taking the plurality of pseudo alignment as supervision information. The prediction performance of the neural network model under the LLP task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence services and cloud, and more particularly to a method for processing data, an apparatus for processing data, an electronic device, and a computer-readable storage medium. The present disclosure is particularly applicable to processing privacy-protected data. Background Art

[0002] Currently, various applications based on artificial intelligence technology have been deployed in the Internet environment to provide personalized data to a large number of users. At the same time, the number of users and network data on the Internet has reached an unprecedented scale. Although rich user data is valuable training data for driving and improving artificial intelligence technology, directly obtaining and using raw personal sensitive data poses a huge threat to user privacy.

[0003] To balance the utilization value of user data and the risk of privacy leakage, most enterprises and institutions currently only provide user data based on statistical aggregation and do not share any raw user records. Such an aggregated instance set usually only contains the overall behavior distribution or statistical information of a certain user group, thus ensuring that personal information data will not be misused.

[0004] In such a background, how an artificial intelligence system can obtain sufficient information from limited aggregated information to provide personalized data to users while protecting user privacy has become an urgent problem to be solved. Summary of the Invention

[0005] Embodiments of the present disclosure provide a method for processing data, an apparatus for processing data, an electronic device, and a computer-readable storage medium.

[0006] Embodiments of the present disclosure provide a method for processing data, the method including: obtaining a first instance set and a second instance set, where the first instance set includes a plurality of first instances, the second instance set includes a plurality of second instances, the first instance set has a first label at the instance set level, and the second instance set has a second label at the instance set level; pre-training a neural network model with the first label and the second label as supervision information; performing fine-tuning training on the pre-trained neural network model, where in each time step of the fine-tuning training, determining a plurality of pseudo-positive pairs, where each pseudo-positive pair includes: a first instance, and a second instance matching the first instance; and adjusting parameters of the neural network model with the plurality of pseudo-positive pairs as supervision information.

[0007] An embodiment of the present disclosure provides an apparatus for processing data, the apparatus including: a pre-training module configured to: obtain a first instance set and a second instance set, where the first instance set includes a plurality of first instances, the second instance set includes a plurality of second instances, the first instance set has a first label at the instance set level, and the second instance set has a second label at the instance set level; use the first label and the second label as supervision information to pre-train a neural network model; and a fine-tuning module configured to: perform fine-tuning training on the pre-trained neural network model, where in each time step of the fine-tuning training, determine a plurality of pseudo-positive pairs, where each pseudo-positive pair includes: a first instance, and a second instance that matches the first instance; and use the plurality of pseudo-positive pairs as supervision information to adjust the parameters of the neural network model.

[0008] An embodiment of the present disclosure discloses an electronic device, including: one or more processors; and one or more memories, where computer-executable programs are stored in the memories, and when the computer-executable programs are executed by the processors, the above-mentioned method is performed.

[0009] An embodiment of the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above-mentioned method is implemented.

[0010] According to another aspect of the present disclosure, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable medium, and the processor executes the computer instructions, so that the computer device executes the method provided in the above-mentioned aspects or various optional implementation manners of the above-mentioned aspects.

[0011] By generating pseudo-positive pairs, the embodiment of the present disclosure introduces instance-level supervision information for the training of a neural network model that only has instance set-level supervision information. This instance-level supervision information bridges the deficiency of instance set-level supervision information (such as aggregated information) and improves the prediction performance of the neural network model.

[0012] Especially in the field of tabular data analysis, according to the embodiment of the present disclosure, based on artificial intelligence, without augmenting the original instance set, a matching relationship between the first instance and the second instance is established through pseudo-positive pairs, and the matching relationship is used to guide the training of the neural network model, improving the robustness of the neural network model.

[0013] In addition, in some embodiments of the present disclosure, the performance of the neural network model is further improved based on a two-stage training scheme. For example, instance set-level supervision information is introduced for contrastive learning in the pre-training stage, while instance-level supervision information is introduced for contrastive learning in the fine-tuning stage. Thus, the feature representations obtained through pre-training match the label proportion distribution, and in the fine-tuning stage, the learning ability and analysis ability of the neural network model for instance data are enhanced, and the category perception ability of the neural network is enhanced. As a result, the performance of the neural network model is further improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for the description of the embodiments. The drawings in the following description are merely exemplary embodiments of the present disclosure.

[0015] Figure 1 It is an exemplary schematic diagram showing a scenario according to an embodiment of the present disclosure.

[0016] Figure 2 It shows a schematic diagram of an application scenario according to an embodiment of the present disclosure.

[0017] Figure 3 It shows a flowchart of a method for processing data according to an embodiment of the present disclosure.

[0018] Figure 4 It shows a schematic diagram of an instance set and example pairs according to an embodiment of the present disclosure.

[0019] Figure 5 It shows a schematic diagram of a pre-training process according to an embodiment of the present disclosure.

[0020] Figure 6 It shows a schematic diagram of a fine-tuning process according to an embodiment of the present disclosure.

[0021] Figure 7 It shows a test comparison diagram according to an embodiment of the present disclosure.

[0022] Figure 8 It shows a test comparison diagram according to an embodiment of the present disclosure.

[0023] Figure 9 It shows a test comparison diagram according to an embodiment of the present disclosure.

[0024] Figure 10 It shows a schematic diagram of an electronic device according to an embodiment of the present disclosure.

[0025] Figure 11 It shows a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure.

[0026] Figure 12 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown. Detailed implementation manners

[0027] In order to make the objectives, technical solutions, and advantages of the present disclosure more apparent, exemplary embodiments according to the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.

[0028] In this specification and the accompanying drawings, operations and elements having substantially the same or similar functions are denoted by the same or similar reference numerals, and repeated descriptions of these operations and elements will be omitted. At the same time, in the description of the present disclosure, terms such as "first" and "second" are only used for differential description and cannot be construed as indicating or implying relative importance or order.

[0029] To facilitate the description of the present disclosure, the following concepts related to the present disclosure are introduced.

[0030] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a manner similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0031] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interaction, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0032] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0033] After pre-training and fine-tuning, the artificial intelligence model of the present disclosure demonstrates a similar level of understanding to humans. In particular, the artificial intelligence model of the present disclosure has excellent performance in deeply mining aggregated tabular data in different instance sets or instance sets, can achieve personalized data analysis and processing, and at the same time can ensure the security of user privacy data.

[0034] Optionally, the models that can be used in the embodiments of the present disclosure hereinafter can all be artificial intelligence models, especially neural network models based on artificial intelligence. Generally, a neural network model based on artificial intelligence is implemented as an acyclic graph, where neurons are arranged in different layers. Generally, a neural network model includes an input layer and an output layer, and the input layer and the output layer are separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating an output in the output layer. Network nodes are fully connected to nodes in adjacent layers via edges, and there are no edges between nodes within each layer. The data received at the nodes of the input layer of the neural network is propagated to the nodes of the output layer via any one of a hidden layer, an activation layer, a pooling layer, a convolutional layer, etc. The input and output of the neural network model can take various forms, and the present disclosure places no restrictions thereon.

[0035] The solutions provided by the embodiments of the present disclosure relate to technologies such as artificial intelligence and / or machine learning, and are specifically described through the following embodiments.

[0036] First, refer to Figure 1 to describe the application scenarios of the method for processing data and the corresponding devices, etc. according to the embodiments of the present disclosure. Figure 1 FIG. 100 shows a schematic diagram of an application scenario 100 according to an embodiment of the present disclosure, in which a server 110 and a plurality of terminals 120 are schematically shown.

[0037] The neural network model of the embodiments of the present disclosure can specifically be integrated in various electronic devices. For example, Figure 1 any of the electronic devices in the server 110 and the plurality of terminals 120. For instance, the neural network model can be integrated in the terminal 120. The terminal 120 can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, a personal computer (PC), a smart speaker, or a smart watch, etc., but is not limited thereto. Also, for example, the neural network model can also be integrated in the server 110. The server 110 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present disclosure places no restrictions thereon.

[0038] It can be understood that the device for performing inference using the neural network model according to the embodiments of the present disclosure can be either a terminal, a server, or a system composed of a terminal and a server. The method for processing data according to the embodiments of the present disclosure can be executed on a terminal, on a server, or jointly executed by a terminal and a server.

[0039] The artificial intelligence model provided by the embodiments of the present disclosure may also be related to artificial intelligence cloud services in the field of cloud technology. Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used on demand, and be flexible and convenient. Cloud computing technology will become an important support. The background services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, drug research websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system backup support, which can only be achieved through cloud computing.

[0040] It should be noted that both the terminal 110 and the server 120 according to the embodiments of the present disclosure comply with the data protection principle, respect the user's data rights, and ensure the user's data security and privacy. The terminal 110 and the server 120 according to the embodiments of the present disclosure will clearly inform the user of the purpose, manner, and scope of collecting, using, storing, transmitting, and deleting the user's data, and obtain the user's consent. The terminal 110 and the server 120 according to the embodiments of the present disclosure will take reasonable technical and management measures to prevent the user's data from being leaked, tampered with, damaged, or lost. The providers of the terminal 110 and the server 120 according to the embodiments of the present disclosure will regularly review and update the user's data and timely delete expired or useless data. In addition, the cloud service provider adopting the embodiments of the present disclosure respects the user's rights such as data access, correction, deletion, revocation of consent, complaint, and claim, and provides convenient channels and procedures for the user to effectively exercise these rights.

[0041] Furthermore, the process of using artificial intelligence technology for data analysis in the terminal 120 or the server 110 is carried out based on the principles of legality, reasonableness, and transparency. The data collected and processed by the artificial intelligence model according to the embodiments of the present disclosure is relevant, necessary, and appropriate for the prediction purpose, and does not contain any personal identification information or sensitive information. The neural network model according to the embodiments of the present disclosure adopts appropriate technical and organizational measures to protect the security and integrity of the data and prevent the data from being accessed, used, or disclosed without authorization.

[0042] The artificial intelligence-based neural network model according to the embodiments of the present disclosure will comply with relevant data protection regulations and ethical principles. The neural network model is trained based on a large amount of anonymized and de-identified data and does not infringe on the privacy rights of any individual or group. The artificial intelligence model has also undergone rigorous testing and evaluation to ensure that its output results are accurate and reliable and will not cause any misinformation or discrimination. The artificial intelligence model is only designed to improve service quality and customer satisfaction and will not be used for any illegal or immoral purposes. In addition, the neural network model will be regularly reviewed and updated to adapt to changes in the data environment and legal norms.

[0043] In the current Internet ecosystem, due to privacy issues, user data is usually presented in the form of aggregated information. Specifically, in order to protect user privacy, enterprises adopt data aggregation techniques and principles and only share the statistical information (i.e., aggregated information) of user data rather than the original records. This method helps to reduce the risk of sensitive information leakage while meeting the needs of business analysis and decision-making.

[0044] Currently, the industry has proposed to use artificial intelligence technology to process aggregated information. However, due to the lack of specific user data, training a neural network model based only on aggregated information often results in low prediction performance.

[0045] In addition, the industry has also proposed using contrastive learning techniques to train neural network models to process aggregated information. The core idea of contrastive learning techniques is to use positive and negative sample pairs to make them close or far away in the embedding space. It is also proposed that since the amount of information in aggregated information is small, a data augmentation (Augmentation) scheme can be combined to perform various equivariance-based transformations on user data and aggregated information to increase the diversity of training data. It has now been proven that such a scheme has shown excellent performance in processing image data or text data.

[0046] However, such a solution often exhibits poor performance when dealing with user data with aggregated information. Usually, user data may be heterogeneous tabular data, which includes multimodal data forms (including data in forms such as text, numerical values, images, videos, etc.). At the same time, tabular data does not have spatial invariance in most cases. This means that data augmentation schemes applicable to most image data or text data cannot be applied to the processing of tabular data.

[0047] To this end, embodiments of the present disclosure provide a method for processing data, the method comprising: obtaining a first instance set and a second instance set, wherein the first instance set includes a plurality of first instances, the second instance set includes a plurality of second instances, the first instance set has a first label at the bag level, and the second instance set has a second label at the bag level; using the first label and the second label as supervision information to pre-train a neural network model; performing fine-tuning training on the pre-trained neural network model, wherein, in each time step of the fine-tuning training, determining a plurality of pseudo positive pairs, wherein each pseudo positive pair includes: a first instance, and a second instance that matches the first instance; and using the plurality of pseudo positive pairs as supervision information to adjust the parameters of the neural network model.

[0048] Embodiments of the present disclosure introduce instance-level supervision information for the training of the neural network model through pseudo positive pairs. This instance-level supervision information bridges the deficiency of instance set-level supervision information (such as aggregated information) and improves the prediction performance of the neural network model.

[0049] Especially in the field of tabular data analysis, according to the embodiments of the present disclosure based on artificial intelligence, without augmenting the original instance set, a matching relationship between the first instance and the second instance is established through pseudo positive pairs, and this matching relationship is used to guide the training of the neural network model, improving the robustness of the neural network model. It should be noted that although the present disclosure mainly uses the field of tabular data analysis as an example for illustration, the present disclosure is not limited thereto. Embodiments of the present disclosure can actually be applied to various different types of data processing schemes, such as image data or text data, etc.

[0050] In addition, in some embodiments of the present disclosure, the performance of the neural network model is further improved based on a two-stage training scheme. For example, in the pre-training stage, instance set-level supervision information is introduced for contrastive learning, while in the fine-tuning stage, instance-level supervision information is introduced for contrastive learning. Thereby, the feature representations obtained by pre-training match the label proportion distribution, and in the fine-tuning stage, the learning ability and analysis ability of the neural network model for instance data are enhanced, and the category perception ability of the neural network is enhanced. Thus, the performance of the neural network model is further improved.

[0051] The present disclosure can particularly improve the prediction performance of the neural network model in the LLP (Learning from Label Proportion) task. In the LLP task, although there are only instance set-level labels in the training stage, instance-level information needs to be predicted in the actual inference stage. Since the present disclosure refines the coarse-grained labels (instance set-level labels) into fine-grained labels (instance-level labels) by generating instance pairs during the training process, the neural network model obtained by the present disclosure is applicable to any LLP type of task.

[0052] Next, refer to Figures 2 to 10 Describe a method for processing data according to an embodiment of the present disclosure.

[0053] Figure 2 A schematic diagram of an application scenario according to an embodiment of the present disclosure is shown.

[0054] As Figure 2 shown, the embodiments of the present disclosure can be integrated into various types of databases based on deep learning as one of the tools for database analysis and processing. The embodiments of the present disclosure can merge the data in multiple instance sets into the database, especially the database storing the feature representations corresponding to certain instances.

[0055] Specifically, Figure 2 Although the first instance set and the second instance set shown in

[0056] For example, assume that both the first instance set and the second instance set are for data of similar users. For example, the first instance set and the second instance set include usage information of social network applications of some same / similar users. To ensure user privacy, the gender information of each user in the first instance set is deleted, and only the aggregated information that 70% of the users in the first instance set are female and 30% of the users are male is given. Similarly, the gender information of each user in the second instance set is also deleted, and only the aggregated information that 60% of the users in the second instance set are female and 40% of the users are male is given. Of course, the present disclosure is not limited thereto.

[0057] The trained neural network model can be integrated into the database to convert each instance in the first instance set and the second instance set into a specific representation vector and store it in the database for subsequent data analysis. For the convenience of analysis, the representation vectors corresponding to the same or similar user data should be stored in adjacent or similar positions for subsequent retrieval and subsequent downstream task calls. The neural network model of the embodiments of the present disclosure can learn instance-level class information from the aggregated information and generate feature representations with stronger representation capabilities, which is beneficial to subsequent task processing.

[0058] Figure 3 A flowchart of a method 30 for processing data according to an embodiment of the present disclosure is shown. Figure 4 A schematic diagram of an instance set and an example pair according to an embodiment of the present disclosure is shown.

[0059] The method 30 can be executed at a terminal device (such as Figure 1 the terminal 120 described). Among them, the method 30 includes operation S301 to operation S303. Of course, the method 30 may further include more or fewer operations, and the present disclosure is not limited thereto.

[0060] In operation S301, a first instance set and a second instance set are obtained, where the first instance set includes a plurality of first instances, the second instance set includes a plurality of second instances, the first instance set has a first label at the instance set level, and the second instance set has a second label at the instance set level.

[0061] Optionally, the instance set (such as the first instance set and the second instance set) according to the embodiments of the present disclosure may also be referred to as a "bag" or an "Instance Bag". An instance set contains a plurality of instances (for example, the first instance or the second instance). An instance (Instance) is also called a sample or a data point, which can be a single sample or data point with similar features or attributes, such as one or more images, one or more texts, a record in tabular data, etc.

[0062] Optionally, both the first label and the second label are bag-level labels, which also means that both are some kind of aggregated information. For example, the first label is a first proportion label indicating the proportion among various categories to which the multiple first instances belong. And the second label is a second proportion label indicating the proportion among various categories to which the multiple second instances belong. In addition, the first label and the second label do not include instance-level labels, so as to avoid the disclosure of user privacy.

[0063] As Figure 4 shown in the left figure of, assume that the first instance set has 100 instances, and the first proportion label is [0.7:0.3], which means that approximately 70% of the instances belong to category 0 and 30% belong to category 1. No instance-level label indicating whether each first instance belongs to category 0 or category 1 is given in the first instance set. Similarly, assume that the second instance set has 100 instances, and the second proportion label is [0.4:0.6], which means that approximately 40% of the instances belong to category 0 and 60% belong to category 1. No instance-level label indicating whether each second instance belongs to category 0 or category 1 is given in the second instance set.

[0064] Optionally, the first instance and the second instance are tabular data. For example, assume that the first instance set is 100 pieces of user data provided by an advertiser, and its first proportion label is [0.7:0.3], indicating that 70% of the users are female and 30% are male. Assume that the second instance set is 100 pieces of user data provided by a shopping platform provider, and its second proportion label is [0.4:0.6], indicating that 40% of the users are female and 60% are male. Each piece of user data in the first instance set and the second instance set includes: the timestamp for collecting the user data, the user behavior identifier, the user's location, the user's age, and so on. Of course, the present disclosure is not limited thereto.

[0065] As Figure 4 shown in the right figure of, the proportion label can also correspond to more than two categories. Assume that the first instance set has 100 instances, and the first proportion label is [0.5:0.3:0.2], which means that approximately 50% of the instances belong to category 0, 30% belong to category 1, and 20% of the instances belong to category 2. No instance-level label indicating which category each first instance belongs to is given in the first instance set. Similarly, assume that the second instance set has 100 instances, and the second proportion label is [0.3:0.4:0.3], which means that approximately 30% of the instances belong to category 0, 40% belong to category 1, and 30% of the instances belong to category 2. No instance-level label indicating which category each second instance belongs to is given in the second instance set. Of course, the present disclosure is not limited thereto.

[0066] In operation S302, the neural network model is pre-trained using the first label and the second label as supervision information.

[0067] Optionally, the pre-training process can be briefly described as follows: Using the encoder in the neural network model, each first instance is encoded to obtain a feature representation corresponding to each first instance, and based on the feature representation corresponding to each first instance, the instance set-level feature representation corresponding to the first instance set is predicted; using the encoder in the neural network model, each second instance is encoded to obtain a feature representation corresponding to each second instance, and based on the feature representation corresponding to each second instance, the instance set-level feature representation corresponding to the second instance set is predicted; determining the similarity between the instance set-level feature representation corresponding to the first instance set and the instance set-level feature representation corresponding to the second instance set; determining the metric distance between the first label and the second label; and adjusting the parameters of the neural network model so that the combination of the difference between the similarity and the metric distance reaches an extreme value.

[0068] Wherein, the combination of the difference between the first predicted label and the first label and the difference between the second predicted label and the second label is the instance set-level pre-training loss, and the pre-training loss is based on the label proportion intersection metric between the first instance set and the second instance set. The calculation method of the label proportion intersection metric between the first instance set and the second instance set will be described in detail later. Of course, the present disclosure is not limited thereto.

[0069] Specifically, using the shared encoder during training (i.e., this encoder can be shared by the first instance set and the second instance set), each first instance is encoded to obtain a feature representation corresponding to each first instance. Similarly, the feature representation corresponding to each second instance is obtained using this shared encoder during training. Then, based on the predicted feature representation corresponding to the first instance, the feature representation corresponding to the first instance set is determined using the bag representation aggregator. Furthermore, the first predicted label corresponding to the first instance set is determined based on the feature representation corresponding to the first instance set.

[0070] In a similar manner, the second predicted label corresponding to the second instance set is determined. Similarly, the first predicted label is also an instance set-level predicted label, which indicates the proportion between the categories to which the predicted first instances belong. Likewise, the second predicted label indicates the proportion between the categories to which the predicted second instances belong. Adjust the parameters of the neural network model so that the difference between the combination of the first label and the second label and the combination of the first predicted label and the second predicted label is minimized. Of course, the present disclosure is not limited thereto.

[0071] Optionally, the strategy of contrastive training can also be used during the pre-training process. Assume that the neural network model is represented by θ and the projection network is represented by δ. fθ (·) represents the corresponding feature representation output after any instance passes through this neural network. g δ (·) is the feature vector after projecting a certain feature vector through a projection network. The goal of this pre-training process is based on g δ (f θ (·)), to optimize the pre-training loss L Bag (·, ·) and the combination of the contrastive loss L com (·, ·), so that the neural network model can align the feature representations learned in the contrastive learning process with the ratio labels.

[0072] Optionally, the contrastive loss L com (·, ·) includes a self-contrastive loss or a denoising loss. The self-contrastive loss indicates the combination of the differences between the class tokens and the latent vectors in the feature representations corresponding to multiple instances. The denoising loss indicates the combination of the differences between the instance data decoded based on the feature representations corresponding to the multiple instances after a mixing operation and / or a noise addition operation and the real instance data. Among them, the mixing operation refers to replacing the vectors of specific dimensions in the feature representation corresponding to a certain instance with the vectors of specific dimensions in the feature representations corresponding to other instances in the instance set. Such a method is also called a same-position replacement operation. The noise addition operation refers to adding noise data (or interference data) to the feature representation corresponding to a certain instance. The feature representations corresponding to the multiple instances after the mixing operation and / or the noise addition operation can also be called the feature representations corresponding to the multiple instances in a noisy view.

[0073] Assume that after pre-training, the parameters of the obtained neural network are θ init , and the parameters of the projection network are δ * . Then, the pre-training process can be expressed by formula (1). Of course, the present disclosure is not limited thereto.

[0074]

[0075] Among them, B represents the first instance set / second instance set, is the first ratio label / second ratio label. The projection network is only a neural network used to assist in training during the pre-training process. Unless required by a specific downstream task, this projection network may not be used in the inference process of the downstream task. Some details of this pre-training process will be further described with reference to Figure 5 , and the present disclosure will not elaborate herein.

[0076] In operation S303, fine-tuning training is performed on the pre-trained neural network model. Wherein, in each time step of the fine-tuning training, a plurality of pseudo-positive pairs are determined. Each pseudo-positive pair includes: a first instance and a second instance that matches the first instance; and using the plurality of pseudo-positive pairs as supervision information, the parameters of the neural network model are adjusted.

[0077] Optionally, in each time step of the fine-tuning stage of the pre-trained neural network model, each instance can be encoded to obtain the feature representation of each first instance and the feature representation of each second instance. Based on the similarity between these feature representations, the second instance with the highest similarity to a certain first instance can be determined as the second instance that matches the first instance. For example, a linear programming method can be used to determine the instance pairs. The linear programming solutions include but are not limited to the linear sum assignment scheme. The similarity here can use the cosine similarity between two features. Of course, other similarity metrics can also be selected, such as Euclidean distance, Manhattan distance, Hamming distance, etc. Of course, the present disclosure is not limited thereto.

[0078] Optionally, as Figure 4 shown, each pseudo-positive pair indicates a one-to-one mapping relationship between the first instance and the second instance. Wherein, the solid line indicates that there is a corresponding matching relationship between the two. For example, the similarity between the two may be higher than the similarity between the first instance and other second instances. The dashed line indicates a lower matching degree between the two. For example, the first instance and the second embodiment connected by the dashed line have the lowest similarity. The two instances connected by the solid line are also called pseudo-positive pairs. The two instances connected by the dashed line are also called negative pairs. Of course, the present disclosure is not limited thereto.

[0079] In traditional contrastive learning tasks, it is usually required to use the already labeled positive pairs as input data. Specifically, a positive pair refers to a pair of instances belonging to the same true class (having the same true label) or a pair of instances that should be of the same class in terms of semantics / attributes. However, in the LLP scenario corresponding to the embodiments of the present disclosure, positive pairs in the traditional sense cannot be obtained, and only some other technical solutions can be used to generate pseudo-positive pairs and apply them. A pseudo-positive pair is a pair of instances that are very likely to belong to the same predicted class (labeled by linear programming) or are very likely to be of the same class in terms of semantics / attributes. The label is the true label, while the pseudo-label is indirectly inferred or generated by technical means to approximate the true positive label (i.e., the pseudo-label) in the absence of a true positive label. Optionally, continue Figure 4For the example in the left figure, in the optimal case, it can be determined that 70% of the instances in the first instance set and the second instance set may correspond to the same object, that is, 70 pseudo-positive pairs can be generated. At the same time, 30% of the instances may correspond to different objects, that is, 30 negative pairs can be generated. Continue Figure 4 For the example in the right figure, in the optimal case, approximately 80% of the data in the first instance set and the second instance set corresponds to similar objects, while 20% of the instances correspond to different objects. That is, in Figure 4 For the example in the right figure, 80 pseudo-positive pairs and 20 negative pairs can be generated. Of course, the present disclosure is not limited thereto.

[0080] Specifically, in the fine-tuning stage, the purpose of contrastive learning is to make the distance between the feature representations of the two instances in a pseudo-positive pair as close as possible, while making the distance between the feature representations of the two instances in a pseudo-negative pair as far as possible. At the same time, based on the assumption that the first instance set and the second instance set may have references to the same or similar objects, the second instance matched with each first instance should be unique. Then, a one-to-one matching scheme can be determined by means of linear programming based on the feature representations corresponding to each instance.

[0081] Therefore, the process of determining multiple pseudo-positive pairs in step S303 can be briefly described as: based on the neural network model in training, determining the feature representation corresponding to each first instance and the feature representation corresponding to each second instance; and based on the similarity between the feature representations corresponding to each first instance and each second instance, using linear programming to determine the multiple pseudo-positive pairs, where the first instance and the second instance in each pseudo-positive pair have a one-to-one matching relationship. That is, each first instance in the first instance set appears at most once in all pseudo-positive pairs, and each second instance in the second instance set also appears at most once in all pseudo-positive pairs. Of course, the present disclosure is not limited thereto.

[0082] Optionally, the neural network model in the fine-tuning process can be utilized to determine, at each time step, the feature representation corresponding to each first instance and the feature representation corresponding to each second instance based on the parameters of the neural network model at the current time step. Then, a linear programming scheme can be adopted to determine, based on the feature representation of the first instance and the feature representations of the respective second instances at the current time step, which feature representation of each first instance can be paired one-to-one with the feature representation of which second instance, while enabling such a pairing scheme to obtain multiple pseudo-positive pairs with optimal matching. Specifically, any instance can only appear in one instance pair and cannot appear in multiple instance pairs. That is to say, the entire linear programming process can be described as: adjusting the second instances matched with the respective first instances in the multiple pseudo-positive pairs to determine multiple pseudo-positive pairs that maximize or minimize the objective function. Optionally, as will be described in detail later, the objective function indicates one of the following: the sum of the cosine similarities between the first instance and the second instance in the pseudo-positive pairs; or the difference between the sum of the cosine similarities between the first instance and the second instance in the pseudo-positive pairs and the sum of the cosine similarities between the first instance and the second instance in the pseudo-negative pairs, where the pseudo-negative pairs include: a first instance and a second instance that does not match the first instance. For example, assume that the number of instances in the first instance set and the second instance set is b. If n pseudo-positive pairs can be obtained, then the number of pseudo-negative pairs is the remaining b - n instance pairs. That is, all instance pairs except the pseudo-positive pairs are pseudo-negative pairs. Similarly, an instance can only appear in one pseudo-negative pair. Of course, the present disclosure is not limited thereto.

[0083] Specifically, due to the matching relationship between the first instance and the second instance corresponding to each pseudo-positive pair, the two should belong to the same category. Thus, based on the assumption that the two belong to the same category, drawing on the idea of contrastive learning, the two can be used as a pseudo-positive pair, and when fine-tuning and training the neural network model, the feature representations of the two are made as close as possible in the embedding space. Since the two belong to the same category, using this pseudo-positive pair as supervision information has instance-level category-aware information content, thus being able to bridge the defect of only using aggregated information as supervision information. Of course, the present disclosure is not limited thereto.

[0084] In some other embodiments, it can also be assumed that only one second instance in the second instance set can form a pseudo-positive pair with the first instance, and the other second instances in the second instance set do not match the first instance. Thus, based on such an assumption, when fine-tuning and training the neural network model, the feature representation of the first instance is made as far away as possible from the feature representations of the other second instances except the second instance in its corresponding pseudo-positive pair. Of course, the present disclosure is not limited thereto.

[0085] In addition, in some other embodiments, it may be further assumed that there is only one second instance in the second instance set that can form a pseudo-negative pair with the first instance. Thus, based on such an assumption, when fine-tuning the neural network model, the feature representation of the first instance is made to be as far away as possible from the feature representation of the second instance in this pseudo-negative pair. Of course, the present disclosure is not limited thereto.

[0086] Optionally, the purpose of using the contrastive training strategy in the fine-tuning training process is to bridge the difference between the pre-training process with instance set-level supervision information and the instance-level prediction. Therefore, the design of the loss function in the fine-tuning training process should consider both the instance-level supervision information (i.e., the matching information provided by the pseudo-positive pairs) and the ratio labels. Of course, the present disclosure is not limited thereto.

[0087] Optionally, the instance-level supervision information can be reflected by the difference contrast loss. Among them, the difference contrast loss is positively correlated with the similarity between the feature representations of the first instance and the second instance in the plurality of pseudo-positive pairs, and negatively correlated with the similarity between the feature representation of the first instance and the feature representation of the second instance that does not match the first instance in the plurality of pseudo-positive pairs. Therefore, in the process of adjusting the parameters of the neural network model with the plurality of pseudo-positive pairs as the supervision information, it can be considered to adjust the parameters of the neural network model to make the difference contrast loss reach an extreme value. Of course, the present disclosure is not limited thereto.

[0088] Optionally, the ratio label loss can be used to bridge the difference between the instance-level supervision information and the instance set-level supervision information. The ratio label loss indicates the difference between the predicted value and the actual value of the first label or the second label, where the predicted value of the first label or the second label is statistically obtained based on the categories corresponding to each instance predicted by the neural network parameters at the current time step. Therefore, in the process of adjusting the parameters of the neural network model with the plurality of pseudo-positive pairs as the supervision information, it can be considered to adjust the parameters of the neural network model to make the ratio label loss reach an extreme value. Of course, the present disclosure is not limited thereto.

[0089] Optionally, the difference contrast loss and the ratio label loss outlined above can be jointly used as the loss function in the fine-tuning stage. Among them, as the time step increases, the weight of the ratio label loss gradually decreases, and the weight of the difference contrast loss gradually increases. Of course, the present disclosure is not limited thereto.

[0090] Specifically, it is assumed that the neural network model continues to be represented by θ, and the prediction network is represented by ω. f θ (·) represents the corresponding feature representation output after any instance passes through this neural network model. h ω(·) is the result after predicting a certain feature vector through a prediction network (e.g., the predicted category). The goal of this pre-training process is based on f θ (·), such that the instance-level differential contrast loss L Diff (·, ·) and / or the instance set-level ratio label loss L LLP (·, ·) reaches an extreme value. Assume that after pre-training, the parameters of the obtained neural network are θ*, and the parameters of the prediction network are ω * . Then, the fine-tuning training process can be expressed by Equation (2). Of course, the present disclosure is not limited thereto.

[0091]

[0092] Among them, B represents the first instance set / second instance set, is the first ratio label / second ratio label. Some details in this pre-training process will be further described with reference to Figure 6 , and the present disclosure will not elaborate herein.

[0093] Optionally, method 30 may further include operation S304 to perform inference using the fine-tuned neural network model. In operation S304, a data set is obtained, the data set includes multiple instances; and each instance in the multiple instances in the data set is predicted for the corresponding class label by using the fine-tuned neural network model.

[0094] Therefore, by introducing multiple pseudo-positive pairs as supervision information, the difference in the label ratio between the two instance sets can be adjusted. This makes the neural network trained in the fine-tuning stage have stronger class awareness. Through contrastive learning based on the correspondence relationship, both the instance-level supervision information and the class awareness information can be integrated, thus eliminating the need in the traditional scheme of neural networks trained based on aggregated information to generate multiple views to enable the neural network model to have sufficient training data. Thus, the embodiments of the present disclosure are particularly suitable for processing tabular data. During the entire training process, there is no need to utilize a data augmentation scheme based on spatial invariance. Therefore, when implementing the embodiments of the present disclosure, it is not required that the training samples (instances) have invariance.

[0095] Next, operation S302 will be further described with reference to Figure 5 . Figure 5 shows a schematic diagram of the pre-training process according to an embodiment of the present disclosure.

[0096] As Figure 5As shown, it is assumed that the neural network model of the embodiment of the present disclosure is an encoder with a Transformer structure based on the self-attention mechanism. The encoder includes an embedding layer and a Transformer based on the self-attention mechanism.

[0097] This encoder is shared by the first instance set and the second instance set during the training process. Optionally, the encoder uses a multi-head self-attention mechanism to process the input embeddings. The multi-head self-attention mechanism is a mechanism that emphasizes the importance of individual features and captures complex relationships in the data.

[0098] Specifically, after being processed by the encoder, each instance will correspondingly obtain a feature representation. The feature representation includes two parts: the class token located in the previous one or more dimensions of the feature representation, and the latent vector located after the class token. The class token, also known as the CLS token (Classification Token), indicates the predicted class for each instance. The CLS token is a special token. In the processing of tabular data, due to the use of the multi-head self-attention mechanism, the encoder can capture the context information of each instance and the input order of each field in a row of tabular data. Therefore, the generated CLS token can be used to represent the overall meaning of a row of tabular data. In addition, only the CLS token can also be used for subsequent prediction of the instance class. The present disclosure is not limited thereto.

[0099] Specifically, assume that the first instance set is represented by B A Select an arbitrary first instance x from the first instance set B A and represent it with i Then, the first instance x after being encoded by the encoder f i The encoded feature vector includes the CSL token z′ enc and its corresponding latent vector z i . That is, similarly, the second instance set B i can be defined, and the second instance x B ∈B j . The feature vector corresponding to the second instance x B includes the CLS token z′ j and its corresponding latent vector z j . j .

[0100] As described above, during the pre-training process, the neural network is pre-trained based on the pre-training loss L Bag (·, ·) at the instance set level. The pre-training loss at the instance set level can also be referred to as the bag contrast loss. Among them, the pre-training loss L Bag(·, ·)'s supervision information is the first label and the second label Optionally, the following defines an optional instance set-level pre-training loss by formula (3). Those skilled in the art should understand that the present disclosure is not limited thereto.

[0101]

[0102] b is the representation corresponding to the combination of the first instance set and the second instance set, b1 is the representation corresponding to the first instance set, and b2 is the representation corresponding to the first instance set. is the first label, is the second label, margin is a user-defined hyperparameter, and mPIoU is the label proportion intersection metric between two instance sets. where represents the proportion corresponding to the type indicator C in the first instance set. For example, assume that the first instance set has 100 instances and the first proportion label is [0.7:0.3], where approximately 70% of the instances belong to class 0 and 30% of the instances belong to class 1. Then, mPIoU can be used to evaluate the consistency between the label proportions across classes. Other metric schemes (such as L1) can also be used instead of the mPIoU metric. The present disclosure is not limited thereto.

[0103] In formula (3), cos(b1, b2) represents the cosine similarity between the feature representations of two instance sets. To better represent the instance set, the feature representations of multiple first instances can be combined through a bag aggregator to obtain a reasonable instance set representation. Specifically, the feature representation of an instance set can be solved by formulas (4) and (5). This scheme will be applicable to the solution of the feature vectors of the first instance set and the second instance set.

[0104]

[0105]

[0106] z i is the embedding of the i-th instance in an instance set, W is the weight matrix in the bag aggregator, and a i is the attention score representing the contribution of the i-th instance to its contribution to the instance set representation. b is the feature representation of the instance set. Through the above calculations, the contribution of each pseudo-positive pair to the feature representation of an instance set depends on the relationship between the instance and other instances in the instance set to which it belongs. Thus, the feature representation of the instance set implicitly reflects the proportion of the instance set label.

[0107] In some embodiments of the present disclosure, the packet contrast loss L Bag (·, ·) can also be combined with the contrast loss L com (·, ·), where the contrast loss L com (·, ·) can be a self-contrast loss L Self (Z), or can be a weighted sum of a self-contrast loss L Self (Z) and a noise reduction loss L Denoise (Z). Among them, the self-contrast loss can be defined by formula (6).

[0108]

[0109] Among them, as described above, z′ i is the CLS token corresponding to the first instance x i , and z i is the latent vector corresponding to the first instance x i . z′ k is the CLS token corresponding to the second instance x k . Thus, the self-contrast loss L Self (z) indicates a combination of the differences between the class tokens and the latent vectors in the representations corresponding to multiple instances. The self-contrast loss L (Z) is an optional loss in the pre-training stage, and other losses can also be used in the embodiments of the present disclosure to replace the self-contrast loss L Self (Z). Self (Z).

[0110] Optionally, the noise reduction loss L Denoisr (Z) can be defined using formula (7).

[0111]

[0112] Among them, L j is the cross-entropy loss of MLP j (z i ) and x i , and MLP j is the j-th non-linear single-hidden layer perceptron in the decoder, which is used to based on the feature representations r iRestore the specific data of the instance. This decoder is only used during the pre-training process and will not be used in subsequent downstream tasks unless the downstream task is directly related to this decoder. Thus, the denoising loss indicates a combination of the differences between the instance data decoded based on the feature representations corresponding to the multiple instances after the mixing operation and / or the noise addition operation and the true instance data. Among them, the mixing operation refers to replacing the vectors of specific dimensions in the feature representation corresponding to a certain instance with the vectors of specific dimensions in the feature representations corresponding to other instances in the instance set. Such a method is also called the same-position replacement operation. The noise addition operation is to add noise data (or interference data) to the feature representation corresponding to a certain instance. The feature representations corresponding to the multiple instances after the mixing operation and / or the noise addition operation can also be referred to as the feature representations corresponding to the multiple instances in the noisy view. Of course, the present disclosure is not limited thereto.

[0113] Next, refer to Figure 4 and Figure 6 to further illustrate operation S303. Figure 6 FIG. shows a schematic diagram of the fine-tuning training process according to an embodiment of the present disclosure.

[0114] The fine-tuning training aims to make up for the lack of instance-level supervision signals caused by only using instance set-level labels during pre-training.

[0115] As described above, operation S303 can determine the multiple pseudo-positive pairs through linear programming. For example, at each time step of the fine-tuning training, based on the neural network model in training, the feature representation corresponding to each first instance and the feature representation corresponding to each second instance can be determined; and based on the similarity between the feature representations corresponding to each first instance and the feature representations corresponding to each second instance, the multiple pseudo-positive pairs are determined by using linear programming, where the similarity between the feature representation of the first instance and the feature representation of the second instance in each pseudo-positive pair is greater than the similarity between the feature representation of the first instance and the feature representations of other second instances in the second instance set.

[0116] Assume the encoder is the optimal encoder obtained after fine-tuning training. For the first instance set b1 with the first label θ and the second instance set b2 with the second label , there exists a solution that can form pseudo-positive pairs between the representation vectors of the first instance and the second instance. Suppose after linear programming, pseudo-positive pairs (also pseudo-positive pairs) are generated, and the two instances in each pseudo-positive pair have the same category. At the same time, n is also generated, where neg = m - n posa negative pair, and the two instances in each negative pair have different categories.

[0117] In an embodiment according to the present disclosure, the cosine similarity can be used to measure the matching degree of the feature vectors of two instances. Then, the objective function based on the pseudo-positive pairs can be set as the difference between the sum of the cosine similarities between the pseudo-positive pairs and the sum of the cosine similarities between the negative pairs. By maximizing this difference, a matching scheme can be determined that increases the similarity of the embedding representations between the pseudo-positive pairs while decreasing the similarity between the embedding representations of the negative pairs. Therefore, for the first instance set b1 with the first label and the second instance set b2 with the second label , assuming that a similarity matrix S can be obtained after linear programming. Among them, both the first instance set b1 and the second instance set b2 have n instances, and the element S ij in the similarity matrix S represents the cosine similarity between the i-th element in the first instance set b1 and the j-th element in the second instance set b2. Then, after linear programming (also known as optimal matching), the first objective function as shown in formula (7) can be maximized.

[0118] O(S, n pos ) = max P,Q (∑ (i,j)∈P S i,j + ∑ (i,j)∈Q - S i,j ) (8)

[0119] where P is the set of pseudo-positive pairs (which includes n pos pseudo-positive pairs), Q is the set of negative pairs (which includes s - n pos pseudo-positive pairs, and s is the number of instances in the first instance set or the second instance set). The U formed by P and Q, U has the constraint: the two instances in each instance pair in U are in one-to-one mapping. By making the first objective function O(S, n pos ) obtain the maximum value, multiple pseudo-positive pairs of the best match can be obtained. The generated pseudo-positive pairs can then be used as pseudo-labels to train the neural network model with the cosine embedding loss, so that the distances between the corresponding first instance and the second instance in each pseudo-positive pair are close, and the distances between the two instances in the negative pairs are pulled apart, and vice versa.

[0120] Furthermore, in some embodiments according to the present disclosure, to simplify the calculation, it can be assumed that the similarity between the negative pairs is zero, so that the objective function is set as the sum of the cosine similarities between the pseudo-positive pairs. By maximizing this sum, a matching scheme that increases the similarity of the embedding representations between the pseudo-positive pairs can be determined. Therefore, for the first instance set b1 with the first label and the second instance set with the second label For the second instance set b2, it is assumed that a similarity matrix S can be obtained after linear programming. Among them, both the first instance set b1 and the second instance set b2 have n instances, and the element S ij in the similarity matrix S represents the cosine similarity between the i-th element in the first instance set b1 and the j-th element in the second instance set b2. Then, after linear programming (also known as optimal matching), the second objective function shown in formula (8) can be maximized.

[0121] O(S, n pos ) = max U (∑( i,j)∈P S i,j ) (9)

[0122] where P is a set of pseudo positive pairs (which includes n pos pseudo positive pairs), Q is a set of negative pairs (which includes s - n pos pseudo positive pairs, and s is the number of instances in the first instance set or the second instance set). U = P ∪ Q, which is a one-to-one mapping between instances.

[0123] If the second objective function is adopted, then the objective of the fine-tuning training in operation S304 can be written as minimizing the difference contrast loss shown in formula (9).

[0124]

[0125] Among them, a pseudo positive pair in P is shown as (i, j), where i is the first instance in this pseudo positive pair and j is the second instance in this pseudo positive pair. k is all the instances in the second instance set including the second instance j. Of course, it can also be the other way around, where i is the second instance in this pseudo positive pair and j is the first instance in this pseudo positive pair. k is all the instances in the first instance set including the first instance j. L Diff (Z, P) is called the difference contrast loss, which reflects the process of generating pseudo positive pairs based on the difference between the label ratios of two instance sets, bridging the difference between the pre-training process with instance set-level supervision information and the instance-level prediction.

[0126] Through the difference contrast loss, the embodiments of the present disclosure improve the category perception ability of the neural network model by determining that the two instances in the pseudo positive pair both belong to the same category. This difference contrast loss not only fuses the instance-level supervision information and the category-related supervision information into the training process of the neural network model, but also only needs to be trained under the constraint of a limited number of pseudo positive pairs (i.e., n pos ) during the training process of each time step of the neural network model, thereby saving computational effort.

[0127] As described above, in addition to the contrastive loss, the ratio label loss can be considered in the fine-tuning stage as the training objective. Specifically, the ratio label loss can be set as shown in Equation (11).

[0128]

[0129] where is the class of each instance predicted by the neural network during training. Thus, the ratio label loss can indicate the difference between the predicted class ratio and the actual first label or second label.

[0130] In an alternative embodiment of the present disclosure, the weight of the contrastive loss can be gradually increased during the fine-tuning stage, while the weight of the ratio label loss is gradually decreased. Specifically, the ratio label loss can be used as a starter in the fine-tuning stage to better determine multiple pseudo-positive pairs. As the training process progresses, the proportion of the ratio label loss can be gradually reduced to better learn the supervision information provided by the pseudo-positive pairs. Thus, the loss function in the fine-tuning stage can be set as shown in Equation (11).

[0131]

[0132] where where t represents the current time step and T represents the total number of time steps in the fine-tuning stage. Thus, as the time step increases, w(t) gradually increases and λ(t) gradually decreases. Expanding Equation (11) gives the loss function in the fine-tuning stage shown in Equation (12).

[0133]

[0134] Thus, through the processing in the fine-tuning stage, the embodiments of the present disclosure can solve two main problems faced by deep learning in processing tabular data: First, traditional solutions usually require complex equivalent transformations (such as view transformations) of the data for data augmentation, while the embodiments of the present disclosure do not require such processing, thus avoiding the poor performance of the neural network model caused by the inability to perform data augmentation on tabular data. Second, the embodiments of the present disclosure also enable the fine-tuned neural network model to better understand the ratio labels of the first instance set and the second instance set and their internal relationships through Equations (9) to (12). By using a linear programming scheme to construct pseudo-positive pairs, the embodiments of the present disclosure successfully enable the neural network model to learn the matching relationship between different instances in the instance set without specific instance labels. This greatly improves the ability of the neural network model to more accurately predict the classes of similar instances.

[0135] AfterFigures 7 to 9 From the shown verification experiments, it can be seen that the neural network model trained by the solution of the present disclosure embodiment has a significant improvement in the performance of predicting the categories of instances.

[0136] For common verification experiments, which include accessible instance-level verification (i.e., the instance-level labels can be viewed during the verification process, while hidden during the training and inference processes), 20 comparison experiments were conducted between the solution of the present disclosure embodiment and other solutions. In this case, Figure 7 The table shows the AUC scores for binary instance sets and the accuracies for multi-class instance sets. The embodiments of the present disclosure (shown as TabLLP-BDC) are consistently superior to other classifiers, highlighting the combined effectiveness of the pre-training task and the fine-tuning training task. It is worth noting that the neural network model trained by the training method of the embodiments of the present disclosure combines the instance-level weak supervision loss through category-aware signals, enabling the learning of similar representations for instances of the same category. This makes the embodiments of the present disclosure significantly different from traditional solutions such as SelfCLR-LLP or TabLLP-SELF in the table.

[0137] Figure 8 shows a more rigorous and practical LLP scenario where, due to user privacy reasons, instance-level verification is not performed. Figure 8 The analysis focuses on the label ratios of the instance sets, and the model is based on the L1 score or the mPIoU metric. Figure 8 The hidden instance-level test results are listed in the table. The embodiments of the present disclosure (shown as TabLLP-BDC) are consistently superior to other classifiers. However, although it does not always achieve the highest mPIoU or L1 score. This is attributed to the inherent limitations of tabular data, such as the lack of comprehensive semantics and the inclusion of noisy instances.

[0138] Figure 9 The influence of the data size in the instance set on the experimental results was studied, with a focus on larger instance set sizes as they are more common in practical applications. The trends of DLLP, SelfCLR-LLP, and the embodiments of the present disclosure (shown as BDC, the solid line in the figure) are as Figure 9 shown. Among them, the tabular data shows stronger resilience and adaptability to increasing instance set sizes, improving the practical applicability of this task.

[0139] An embodiment of the present disclosure also provides an apparatus for processing data, the apparatus including: a pre-training module configured to: obtain a first instance set and a second instance set, where the first instance set includes a plurality of first instances, the second instance set includes a plurality of second instances, the first instance set has a first label at the instance set level, and the second instance set has a second label at the instance set level; pre-train a neural network model using the first label and the second label as supervision information; and a fine-tuning module configured to: perform fine-tuning training on the pre-trained neural network model, where in each time step of the fine-tuning training, determine a plurality of pseudo positive pairs, where each pseudo positive pair includes: a first instance, and a second instance matching the first instance; and adjust the parameters of the neural network model using the plurality of pseudo positive pairs as supervision information.

[0140] According to another aspect of the present disclosure, there is also provided an electronic device for implementing the method according to the embodiments of the present disclosure. Figure 10 A schematic diagram of an electronic device 2000 according to an embodiment of the present disclosure is shown.

[0141] As Figure 10 shown, the electronic device 2000 may include one or more processors 2010 and one or more memories 2020. Among them, computer-readable code is stored in the memory 2020, and when the computer-readable code is run by the one or more processors 2010, the above-mentioned method can be executed.

[0142] The processor in the embodiment of the present disclosure may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The various methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.

[0143] Generally speaking, the various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor or other computing devices. When the aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts or using some other graphical representations, it will be understood that the blocks, devices, systems, technologies or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0144] For example, the method or apparatus according to an embodiment of the present disclosure may also be implemented by means of Figure 11 the architecture of the computing device 3000 shown. As Figure 11 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 11 the architecture shown is only exemplary, and when implementing different devices, one or more components shown in the Figure 11 illustrated computing device may be omitted according to actual needs.

[0145] According to another aspect of the present disclosure, a computer-readable storage medium is also provided. Figure 12 A schematic diagram of the storage medium 4000 according to the present disclosure is shown.

[0146] As Figure 12 shown, computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the method according to an embodiment of the present disclosure described with reference to the above drawings can be executed. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories for the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories. It should be noted that the memories for the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0147] Embodiments of the present disclosure also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to the embodiments of the present disclosure.

[0148] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0149] In general, the various example embodiments of the present disclosure can be implemented in hardware or a dedicated circuit, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, a dedicated circuit or logic, general hardware or a controller or other computing devices, or some combination thereof.

[0150] The example embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for processing data, the method comprising: Obtaining a first instance set and a second instance set, wherein the first instance set includes a plurality of first instances, the second instance set includes a plurality of second instances, the first instance set has a first label at the instance set level, and the second instance set has a second label at the instance set level; Pre-training a neural network model using the first label and the second label as supervision information; Performing fine-tuning training on the pre-trained neural network model, wherein, at each time step of the fine-tuning training, Determining a plurality of pseudo-positive pairs, where each pseudo-positive pair includes: a first instance, and a second instance that matches the first instance; and Adjusting the parameters of the neural network model using the plurality of pseudo-positive pairs as supervision information.

2. The method for processing data according to claim 1, wherein, The first label is a first ratio label indicating the ratio between the various categories to which the plurality of first instances belong; And The second label is a second ratio label indicating the ratio between the various categories to which the plurality of second instances belong.

3. The method for processing data according to claim 1, wherein, The first instance set and the second instance set do not have instance-level labels; and The first instance and the second instance are a row record in different tabular data.

4. The method for processing data according to claim 1, wherein, The determining of the plurality of pseudo-positive pairs includes: Using the neural network model during training to determine the feature representation corresponding to each first instance and the feature representation corresponding to each second instance; and Based on the similarity between the feature representations corresponding to each first instance and the feature representations corresponding to each second instance, using linear programming to determine the plurality of pseudo-positive pairs, wherein the first instance and the second instance in each pseudo-positive pair have a one-to-one matching relationship.

5. The method for processing data according to claim 4, wherein, The using of linear programming to determine the plurality of pseudo-positive pairs includes: Adjusting the second instances that match each first instance in the plurality of pseudo-positive pairs to determine a plurality of pseudo-positive pairs that make the objective function reach an extreme value, wherein the objective function indicates one of the following: The sum of the cosine similarities between the first instance and the second instance in the pseudo-positive pairs; or The difference between the sum of the cosine similarities between the first instance and the second instance in the pseudo-positive pairs and the sum of the cosine similarities between the first instance and the second instance in the pseudo-negative pairs, wherein the pseudo-negative pairs include: a first instance, and a second instance that does not match the first instance.

6. The method for processing data according to claim 1, wherein, The adjusting of the parameters of the neural network model using the plurality of pseudo-positive pairs as supervision information includes: Adjusting the parameters of the neural network model to make the differential contrast loss reach an extreme value, wherein the differential contrast loss is positively correlated with the similarity between the feature representations of the first instance and the second instance in the plurality of pseudo-positive pairs, and negatively correlated with the similarity between the feature representations of the first instance and the second instance that does not match the first instance in the plurality of pseudo-positive pairs.

7. The method for processing data according to claim 6, wherein, The adjusting of the parameters of the neural network model using the plurality of pseudo-positive pairs as supervision information includes: Adjust the parameters of the neural network model to make the ratio label loss reach an extreme value, where the ratio label loss indicates the difference between the predicted value and the actual value of the first label or the second label, and the predicted value of the first label or the second label is obtained based on the class statistics corresponding to each instance predicted by the neural network parameters at the current time step.

8. The method for processing data according to claim 7, wherein, Adjusting the parameters of the neural network model with the multiple pseudo-positive pairs as the supervision information includes: Adjust the parameters of the neural network model to make the combination of the difference contrast loss and the ratio label loss reach an extreme value, wherein, as the time step increases, the weight of the ratio label loss gradually decreases, and the weight of the difference contrast loss gradually increases.

9. The method for processing data according to claim 1, wherein Pre-training the neural network model with the first label and the second label as the supervision information includes: Using the encoder in the neural network model to encode each first instance to obtain the feature representation corresponding to each first instance, and predicting the instance set-level feature representation corresponding to the first instance set based on the feature representation corresponding to each first instance; Using the encoder in the neural network model to encode each second instance to obtain the feature representation corresponding to each second instance, and predicting the instance set-level feature representation corresponding to the second instance set based on the feature representation corresponding to each second instance; Determine the similarity between the instance set-level feature representation corresponding to the first instance set and the instance set-level feature representation corresponding to the second instance set; Determine the metric distance between the first label and the second label; and Adjust the parameters of the neural network model to make the combination of the difference between the similarity and the metric distance reach an extreme value.

10. The method for processing data according to claim 9, wherein, The combination of the difference between the similarity and the metric distance is the pre-training loss at the instance set level, and the pre-training loss is based on the label ratio intersection metric between the first instance set and the second instance set.

11. The method for processing data according to claim 10, wherein, Pre-training the neural network model with the first label and the second label as the supervision information includes: Adjust the neural network model to make the combination of the pre-training loss and the contrast loss reach an extreme value, where the contrast loss includes a self-contrast loss or a denoising loss, the self-contrast loss indicates the combination of the differences between the class token and the latent vector in the feature representations corresponding to multiple instances, and the denoising loss indicates the combination of the differences between the instance data decoded based on the feature representations corresponding to multiple instances after a mixing operation or a noise addition operation and the true instance data.

12. The method according to claim 1, further comprising: Obtain a data set, the data set including multiple instances; and Use the fine-tuned and trained neural network model to predict the class label corresponding to each instance in the multiple instances in the data set.

13. An apparatus for processing data, the apparatus comprising: A pre-training module, configured to: obtain a first instance set and a second instance set, where the first instance set includes a plurality of first instances, the second instance set includes a plurality of second instances, the first instance set has a first label at the instance set level, and the second instance set has a second label at the instance set level; use the first label and the second label as supervision information to pre-train a neural network model; and A fine-tuning module, configured to: perform fine-tuning training on the pre-trained neural network model, where in each time step of the fine-tuning training, determine a plurality of pseudo-positive pairs, where each pseudo-positive pair includes: a first instance and a second instance that matches the first instance; and use the plurality of pseudo-positive pairs as supervision information to adjust the parameters of the neural network model.

14. An electronic device, comprising: one or more processors; and one or more memories, where computer-executable programs are stored in the memories, and when the computer-executable programs are executed by the processors, the method according to any one of claims 1-12 is performed.

15. A computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the method according to any one of claims 1-12 is implemented.