Training method and device of contrast learning model, electronic equipment and storage medium

CN116051919BActive Publication Date: 2026-10-09瀚依科技(杭州)有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211460697.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2026-10-09
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

其主要目的在于解决对比学习中模型易坍缩,进而造成对比学习模型训练效果不佳的问题

Benefits of technology

[0056]本公开提供的对比学习模型的训练方法、装置、电子设备和存储介质,主要技术方案包括:首先,将样本图像输入预设对比学习模型,所述样本图像包含图像标识,其次,基于预设对比学习模型中的G网络,对至少两个样本图像进行特征提取,得到样本特征对,将所述G网络提取的样本特征对输入预设对比学习模型中的D网络,基于所述D网络根据所述图像标识确定所述样本特征对之间的相似度,最后,根据所述G网络提取的样本特征对及所述D网络确定对所述样本特征对之间的相似度,进行对抗训练。与相关技术相比,本申请实施例通过将生成对抗网络中的G网络与D网络引入对比学习中,通过G网络与D网络的对抗训练,互相进行监督,有效的解决了传统对比学习模型坍缩的问题,且通过G网络与D网络的对抗训练能够使G网络与D网络的性能都达到最优,提高了对比训练模型的训练效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051919B_ABST
    Figure CN116051919B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a training method and device of a contrast learning model, an electronic device and a storage medium, relates to the technical field of artificial intelligence, and introduces a G network and a D network in a generative adversarial network into contrast learning, performs mutual supervision training through adversarial training of the G network and the D network, effectively solves the problem of collapse of a traditional contrast learning model, and can make the performance of the G network and the D network optimal through adversarial training of the G network and the D network, thereby improving the training effect of the contrast training model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method and apparatus for training a contrastive learning model, an electronic device, and a storage medium. Background Technology

[0002] Pre-trained models are widely used in the field of deep learning. In deep learning, using pre-trained models to fine-tune them on the target task often achieves better results than training models from scratch, thereby improving the performance of the model in subsequent tasks.

[0003] Contrastive learning is a popular self-supervised pre-training model method in the field of image processing in recent years. It learns the core features in an image to obtain a pre-trained model that can extract key features of the image, and then applies the pre-trained model to tasks such as image classification.

[0004] However, the learning model trained by the contrastive learning method usually has the following problem: if all samples are mapped to the same point in the feature space, then the similarity between sample pairs is at its maximum value. This leads to the problem of learning model collapse, and the learning model has no practical meaning. In this case, the training effect of the contrastive learning model will be poor. Summary of the Invention

[0005] This disclosure provides a training method, apparatus, electronic device, and storage medium for a contrastive learning model. Its main purpose is to address the problem of model collapse in contrastive learning, which leads to poor training results.

[0006] According to a first aspect of this disclosure, a method for training a contrastive learning model is provided, comprising:

[0007] The sample images are input into a preset contrastive learning model, and the sample images contain image labels;

[0008] Based on the G network in the pre-defined contrastive learning model, feature extraction is performed on at least two sample images to obtain sample feature pairs;

[0009] The sample features extracted by the G network are input into the D network in the preset contrastive learning model;

[0010] Based on the D network, the similarity between the sample feature pairs is determined according to the image identifier;

[0011] Adversarial training is performed based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network.

[0012] Optionally, the step of performing adversarial training based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network includes:

[0013] When the number of training iterations is greater than or equal to the first preset round threshold, the network parameters configured for the sample feature pairs extracted by the G network are fixed, and the network parameters configured for the similarity determination by the D network are trained.

[0014] After the number of training iterations is greater than or equal to the second preset threshold, the network parameters configured for the D network are fixed to determine the similarity, and the network parameters configured for the G network are trained by extracting sample features.

[0015] Optionally, the adversarial training based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network includes:

[0016] Calculate the loss function of the pre-defined comparative training model;

[0017] Determine whether the G network and the D network simultaneously satisfy the convergence condition based on the loss function.

[0018] When both the G network and the D network meet the convergence condition, the training of the preset contrastive learning model is completed.

[0019] Optionally, before inputting the sample images into the preset contrastive learning model, the method further includes:

[0020] Each sample image is processed separately without superposition using at least two preset image enhancement methods.

[0021] Optionally, determining the similarity between the sample feature pairs based on the image identifier using the D network includes:

[0022] If the sample feature pair is a sample image feature pair enhanced by at least two different image enhancement methods for the same image identifier, then the negative correlation similarity between the sample feature pairs is determined based on the D network;

[0023] If the sample feature pair is a sample image feature pair enhanced by a random image enhancement method for different image identifiers, then the positive correlation similarity between the sample feature pairs is determined based on the D network.

[0024] Optionally, the method further includes:

[0025] Input the image to be detected into the pre-trained contrastive learning model;

[0026] Image features of the image to be detected are extracted based on the G network in the preset comparison model.

[0027] According to a second aspect of this disclosure, a method apparatus for training a contrastive learning model is provided, comprising:

[0028] The first input unit is used to input a sample image into a preset contrastive learning model, wherein the sample image contains an image identifier;

[0029] The first extraction unit is used to extract features from at least two sample images based on the G network in the preset contrastive learning model to obtain sample feature pairs.

[0030] The second input unit is used to input the sample features extracted by the G network into the D network of the preset contrastive learning model;

[0031] The determining unit is configured to determine the similarity between the sample feature pairs based on the image identifiers using the D network;

[0032] The training unit is used to perform adversarial training based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network.

[0033] Optionally, the training unit includes:

[0034] The first fixed module is used to fix the network parameters configured for the sample feature pairs extracted by the G network after the number of training cycles is greater than or equal to the first preset round threshold.

[0035] The first training module is used to train the network parameters configured for determining similarity in the D network;

[0036] The second fixed module is used to fix the network parameters configured by the D network to determine the similarity after the number of training rounds is greater than or equal to the second preset round threshold.

[0037] The second training module is used to extract sample features from the G network and train the configured network parameters.

[0038] Optionally, the training unit includes:

[0039] The calculation module is used to calculate the loss function of the pre-defined comparative training model;

[0040] The judgment module is used to determine whether the G network and the D network simultaneously satisfy the convergence condition based on the loss function.

[0041] The completion module is used to complete the training of the preset contrastive learning model when both the G network and the D network simultaneously meet the convergence condition.

[0042] Optionally, before the first input unit, the device further includes:

[0043] The processing unit is used to perform non-overlapping image enhancement processing on each sample image according to at least two preset image enhancement methods.

[0044] Optionally, the determining unit includes:

[0045] The first determining module is used to determine the negative correlation similarity between the sample feature pairs based on the D network when the sample feature pairs are sample image feature pairs enhanced by at least two different image enhancement methods for the same image identifier.

[0046] The second determining module is used to determine the positive correlation similarity between the sample feature pairs based on the D network when the sample feature pairs are sample image feature pairs enhanced by different image identifiers using a random image enhancement method.

[0047] Optionally, the device further includes:

[0048] The third input unit is used to input the image to be detected into the pre-trained contrast learning model;

[0049] The second extraction unit is used to extract image features of the image to be detected based on the G network in the preset comparison model.

[0050] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0051] At least one processor; and

[0052] A memory communicatively connected to the at least one processor; wherein,

[0053] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0054] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0055] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0056] The training method, apparatus, electronic device, and storage medium for the contrastive learning model disclosed herein mainly include the following technical solutions: First, inputting sample images into a preset contrastive learning model, wherein the sample images contain image identifiers; second, based on the G network in the preset contrastive learning model, extracting features from at least two sample images to obtain sample feature pairs; inputting the sample feature pairs extracted by the G network into the D network in the preset contrastive learning model; determining the similarity between the sample feature pairs based on the image identifiers using the D network; and finally, performing adversarial training based on the sample feature pairs extracted by the G network and the similarity determined by the D network. Compared with related technologies, the embodiments of this application introduce the G network and D network from generative adversarial networks into contrastive learning, and through adversarial training of the G network and D network, mutual supervision is achieved, effectively solving the problem of collapse in traditional contrastive learning models. Furthermore, adversarial training of the G network and D network can optimize the performance of both the G network and D network, improving the training effect of the contrastive learning model.

[0057] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0058] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0059] Figure 1 This is a flowchart illustrating a training method for a contrastive learning model provided in an embodiment of the present disclosure.

[0060] Figure 2 A flowchart illustrating a method for training G-networks and D-networks provided in an embodiment of this application;

[0061] Figure 3 A schematic diagram of the structure of a training device for a contrastive learning model provided in an embodiment of this disclosure;

[0062] Figure 4 A schematic diagram of the structure of a training device for another contrastive learning model provided in an embodiment of this disclosure;

[0063] Figure 5 A schematic block diagram of an example electronic device 400 provided for embodiments of this disclosure. Detailed Implementation

[0064] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0065] The following description, with reference to the accompanying drawings, outlines a method, apparatus, electronic device, and storage medium for training a contrastive learning model according to embodiments of the present disclosure.

[0066] Figure 1 This is a flowchart illustrating a training method for a contrastive learning model provided in an embodiment of the present disclosure.

[0067] like Figure 1 As shown, the method includes the following steps:

[0068] Step 101: Input the sample image into the preset contrastive learning model. The sample image contains image labels.

[0069] The preset contrastive learning model is trained based on sample images. The number and content of the sample images can be determined according to the actual research question. For example, when studying the classification of animals and plants, sample images containing animal and plant content can be selected to train the preset contrastive learning model. The more sample images there are, the better the training effect of the preset contrastive learning model will be, and the more accurate the image features extracted will be when applied to subsequent actual research. The embodiments of this application do not limit the content and number of sample images.

[0070] Image tags are used to identify the relationships between different images. For example, two images may be identical in content, but they may have been flipped, rotated, or otherwise altered.

[0071] Step 102: Based on the G network in the preset contrastive learning model, extract features from at least two sample images to obtain sample feature pairs.

[0072] Since the G network and D network need to be trained by comparison based on multiple sample features, if only one feature is extracted in a single training session, the training cannot be completed. Therefore, the pre-set contrastive learning model randomly extracts at least two sample images from the sample images, and extracts the sample features of the two sample images based on the G network to form a sample feature pair, and trains the G network and D network based on the sample feature pair.

[0073] Step 103: Input the sample features extracted by the G network into the D network of the preset contrastive learning model.

[0074] The sample feature pairs are input into the D network, which is used to determine the similarity between the sample feature pairs extracted by the G network.

[0075] Step 104: Based on the D network, determine the similarity between the sample feature pairs according to the image identifier.

[0076] Step 105: Based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network, adversarial training is performed.

[0077] During adversarial training, the G network and the D network are in a game-like relationship. The recognition ability of the D network is trained based on the sample feature pairs extracted by the G network; and the ability of the G network to extract sample features is trained based on the similarity identified by the D network, thereby achieving the goal of training the model's ability to extract image features.

[0078] The training method for the contrastive learning model disclosed herein mainly includes the following technical solutions: First, inputting sample images into a preset contrastive learning model, wherein the sample images contain image identifiers; second, based on the G network in the preset contrastive learning model, extracting features from at least two sample images to obtain sample feature pairs; inputting the sample feature pairs extracted by the G network into the D network in the preset contrastive learning model; determining the similarity between the sample feature pairs based on the image identifiers using the D network; and finally, performing adversarial training based on the sample feature pairs extracted by the G network and the similarity determined by the D network. Compared with related technologies, this application embodiment introduces the G network and D network from generative adversarial networks into contrastive learning, and through adversarial training of the G network and D network, mutual supervision is achieved, effectively solving the problem of collapse in traditional contrastive learning models. Furthermore, adversarial training of the G network and D network can optimize the performance of both the G network and D network, improving the training effect of the contrastive learning model.

[0079] As an extension of the above-mentioned embodiments, before inputting the sample images into the preset contrastive learning model, each sample image needs to be processed separately using at least two preset image enhancement methods without being superimposed, and image labels are added. For example, different sample images are labeled as 1, 2, and 3, and different image enhancement methods are labeled as A and B. Then, the image enhancement results for the sample images are 1A, 1B, 2A, 2B, 3A, and 3B. The image enhancement methods include rotation, flipping, cropping, changing image brightness, etc. This embodiment does not limit the image enhancement methods.

[0080] As an extension of the above-mentioned embodiments, when determining the similarity between the sample feature pairs based on the D network, different similarities need to be extracted according to the different sample images; if the sample feature pair is a sample image feature pair enhanced by at least two different image enhancement methods for the same image identifier, then the negative correlation similarity between the sample feature pairs is determined based on the D network; if the sample feature pair is a sample image feature pair enhanced by random image enhancement methods for different image identifiers, then the positive correlation similarity between the sample feature pairs is determined based on the D network.

[0081] In this embodiment of the application, an iterative training method is used when training the G network and the D network; such as Figure 2 As shown, Figure 2 A flowchart illustrating a method for training G-networks and D-networks provided in this application embodiment includes:

[0082] Step 201: When the number of training iterations is greater than or equal to the first preset round threshold, fix the network parameters configured for the sample feature pairs extracted by the G network, and train the network parameters configured for the similarity determination of the D network.

[0083] At the start of training, no preset contrastive learning model is configured. When the preset contrastive learning model reaches the preset threshold of training epochs, that is, when the G network in the preset contrastive learning model has the ability to extract some image features and the G network also has the ability to determine the similarity between some sample feature pairs, the network parameters configured by the G network to extract sample feature pairs are fixed. That is, based on the current ability of the G network to extract sample features, sample features are extracted, and the extracted sample feature pairs are input into the D network. The D network determines the similarity between sample feature pairs, and the network parameters configured by the D network to determine the similarity are trained.

[0084] The first preset round threshold is an empirical value, which can be set according to the training situation in actual training. This application embodiment does not limit this.

[0085] Step 202: After the number of training iterations is greater than or equal to the second preset round threshold, fix the network parameters configured for similarity in the D network, and train the G network to extract sample features based on the configured network parameters.

[0086] When the D network can easily extract the similarity between the feature pairs extracted by the G network, the training effect will be greatly reduced. Therefore, after the number of training times reaches the preset second preset threshold, the training of the network parameters configured by the D network to determine the similarity is stopped, and the network parameters configured by the D network to determine the similarity are opened. The network parameters of the D network are fixed, and the network parameters configured by the G network to extract the sample feature pairs are trained. The D network enables the G network to extract more accurate sample features, and the G network enables the D network to identify sample feature pairs that are more difficult to distinguish.

[0087] As an extension of the above-mentioned embodiments, when training the G network and D network, it is necessary to constrain the training process and the training direction of the network parameters according to the loss function. The following method can be used: calculate the loss function of a preset contrastive training model; determine whether the G network and D network simultaneously meet the convergence condition based on the loss function; when the G network and D network simultaneously meet the convergence condition, the training of the preset contrastive learning model is completed; when the number of sample images is b, the preset contrastive learning loss function is:

[0088] The numerator is the loss function used when training images with the same image identifier processed by two different image enhancement methods, and the denominator is the loss function used when training images with different image identifiers processed by random image enhancement methods. Training is complete when both the G network and the D network simultaneously meet the convergence condition. The determination of the G network and D network based on the loss function can refer to any implementation method in the prior art; this embodiment will not elaborate on these details further.

[0089] As an extension of the embodiments of this application, after training the preset contrastive learning model, it can be used in other image recognition and classification tasks. When performing other image processing tasks, the following method can be used: input the image to be detected into the trained preset contrastive learning model, and extract the image features of the image to be detected based on the G network in the preset contrastive model; thus, the image feature extraction of the image to be detected can be completed. The extracted image features can be used for tasks such as image classification and image content detection. This application embodiment does not limit the application scenarios of the preset contrastive learning model.

[0090] Corresponding to the training method of the contrastive learning model described above, this invention also proposes a training device for the contrastive learning model. Since the device embodiments of this invention correspond to the method embodiments described above, details not disclosed in the device embodiments can be referred to in the method embodiments described above, and will not be repeated here.

[0091] Figure 3 This is a schematic diagram of the structure of a training device for a contrastive learning model provided in an embodiment of this disclosure, as shown below. Figure 3As shown, it includes:

[0092] The first input unit 31 is used to input a sample image into a preset contrast learning model, wherein the sample image contains an image identifier;

[0093] The first extraction unit 32 is used to extract features from at least two sample images based on the G network in the preset contrastive learning model to obtain sample feature pairs.

[0094] The second input unit 33 is used to input the sample features extracted by the G network into the D network in the preset contrastive learning model;

[0095] Determining unit 34 is used to determine the similarity between the sample feature pairs based on the image identifiers according to the D network;

[0096] Training unit 35 is used to perform adversarial training based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network.

[0097] The training apparatus for the contrastive learning model disclosed herein mainly includes the following technical solutions: First, inputting sample images into a preset contrastive learning model, wherein the sample images contain image identifiers; second, based on the G network in the preset contrastive learning model, extracting features from at least two sample images to obtain sample feature pairs; inputting the sample feature pairs extracted by the G network into the D network in the preset contrastive learning model; determining the similarity between the sample feature pairs based on the image identifiers using the D network; and finally, performing adversarial training based on the sample feature pairs extracted by the G network and the similarity determined by the D network. Compared with related technologies, the embodiments of this application introduce the G network and D network from generative adversarial networks into contrastive learning, and through adversarial training of the G network and D network, mutual supervision is achieved, effectively solving the problem of collapse in traditional contrastive learning models. Furthermore, adversarial training of the G network and D network can optimize the performance of both the G network and D network, improving the training effect of the contrastive learning model.

[0098] Furthermore, in one possible implementation of this embodiment, such as Figure 4 As shown, the training unit 35 includes:

[0099] The first fixed module 351 is used to fix the network parameters configured for the sample feature pairs extracted by the G network after the number of training cycles is greater than or equal to the first preset round threshold.

[0100] The first training module 352 is used to train the network parameters configured for determining similarity in the D network;

[0101] The second fixed module 353 is used to fix the network parameters configured by the D network to determine the similarity after the number of training times is greater than or equal to the second preset round threshold.

[0102] The second training module 354 is used to extract sample features from the G network and train the configured network parameters.

[0103] Furthermore, in one possible implementation of this embodiment, such as Figure 4 As shown, the training unit 35 includes:

[0104] The calculation module 355 is used to calculate the loss function of the preset contrastive training model;

[0105] The judgment module 356 is used to determine whether the G network and the D network simultaneously satisfy the convergence condition based on the loss function.

[0106] Module 357 is used to complete the training of the preset contrastive learning model when both the G network and the D network simultaneously meet the convergence condition.

[0107] Furthermore, in one possible implementation of this embodiment, such as Figure 4 As shown, the device further includes:

[0108] The processing unit 36 ​​is used to perform non-overlapping image enhancement processing on each sample image according to at least two preset image enhancement methods before the first input unit 31 inputs the sample image into the preset contrast learning model.

[0109] Furthermore, in one possible implementation of this embodiment, such as Figure 4 As shown, the determining unit 34 includes:

[0110] The first determining module 341 is used to determine the negative correlation similarity between the sample feature pairs based on the D network when the sample feature pairs are sample image feature pairs enhanced by at least two different image enhancement methods for the same image identifier.

[0111] The second determining module 342 is used to determine the positive correlation similarity between the sample feature pairs based on the D network when the sample feature pairs are sample image feature pairs enhanced by a random image enhancement method for different image identifiers.

[0112] Furthermore, in one possible implementation of this embodiment, such as Figure 4 As shown, the device further includes:

[0113] The third input unit 37 is used to input the image to be detected into the trained preset contrast learning model;

[0114] The second extraction unit 38 is used to extract image features of the image to be detected based on the G network in the preset comparison model.

[0115] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] Figure 5 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] like Figure 5 As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 402 or a computer program loaded from storage unit 408 into RAM (Random Access Memory) 403. RAM 403 may also store various programs and data required for the operation of device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. I / O (Input / Output) interface 405 is also connected to bus 404.

[0118] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0119] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as methods for training contrastive learning models. For example, in some embodiments, methods for training contrastive learning models can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, computing unit 401 may be configured to perform the aforementioned training method of the contrastive learning model by any other suitable means (e.g., by means of firmware).

[0120] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0124] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0125] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0126] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0127] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0128] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A training method for a contrastive learning model, characterized in that, include: The sample images are input into a preset contrastive learning model, and the sample images contain image labels; Based on the G network in the pre-defined contrastive learning model, feature extraction is performed on at least two sample images to obtain sample feature pairs; The sample features extracted by the G network are input into the D network in the preset contrastive learning model; Based on the D network, the similarity between the sample feature pairs is determined according to the image identifier; Based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network, adversarial training is performed; wherein, during adversarial training, the G network and the D network are in a game-like relationship. The step of determining the similarity between sample feature pairs based on the image identifier using the D network includes: If the sample feature pair is a sample image feature pair enhanced by at least two different image enhancement methods for the same image identifier, then the negative correlation similarity between the sample feature pairs is determined based on the D network; If the sample feature pair is a sample image feature pair enhanced by a random image enhancement method for different image identifiers, then the positive correlation similarity between the sample feature pairs is determined based on the D network.

2. The method according to claim 1, characterized in that, The adversarial training based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network includes: When the number of training iterations is greater than or equal to the first preset round threshold, the network parameters configured for the sample feature pairs extracted by the G network are fixed, and the network parameters configured for the similarity determination by the D network are trained. After the number of training iterations is greater than or equal to the second preset threshold, the network parameters configured for the D network are fixed to determine the similarity, and the network parameters configured for the G network are trained by extracting sample features.

3. The method according to claim 2, characterized in that, The adversarial training based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network includes: Determine whether the G network and D network simultaneously satisfy the convergence condition based on the loss function; When both the G network and the D network simultaneously meet the convergence condition, the training of the preset contrastive learning model is completed.

4. The method according to claim 1, characterized in that, Before inputting the sample images into the preset contrastive learning model, the method further includes: Each sample image is processed separately without superposition using at least two preset image enhancement methods.

5. The method according to claim 3, characterized in that, The method further includes: Input the image to be detected into the pre-trained contrastive learning model; The image features of the image to be detected are extracted based on the G network in the preset contrastive learning model.

6. A training device for a contrastive learning model, characterized in that, include: The first input unit is used to input a sample image into a preset contrastive learning model, wherein the sample image contains an image identifier; The first extraction unit is used to extract features from at least two sample images based on the G network in the preset contrastive learning model to obtain sample feature pairs. The second input unit is used to input the sample features extracted by the G network into the D network of the preset contrastive learning model; The determining unit is configured to determine the similarity between the sample feature pairs based on the image identifiers using the D network; The training unit is used to perform adversarial training based on the sample feature pairs extracted by the G network and the similarity between the sample feature pairs determined by the D network; wherein, during adversarial training, the G network and the D network are in a game-like relationship. The determining unit includes: The first determining module is used to determine the negative correlation similarity between the sample feature pairs based on the D network when the sample feature pairs are sample image feature pairs enhanced by at least two different image enhancement methods for the same image identifier. The second determining module is used to determine the positive correlation similarity between the sample feature pairs based on the D network when the sample feature pairs are sample image feature pairs enhanced by different image identifiers using a random image enhancement method.

7. The apparatus according to claim 6, characterized in that, The training unit includes: The first fixed module is used to fix the network parameters configured for the sample feature pairs extracted by the G network after the number of training cycles is greater than or equal to the first preset round threshold. The first training module is used to train the network parameters configured for determining similarity in the D network; The second fixed module is used to fix the network parameters configured by the D network to determine the similarity after the number of training rounds is greater than or equal to the second preset round threshold. The second training module is used to extract sample features from the G network and train the configured network parameters.

8. The apparatus according to claim 7, characterized in that, The training unit includes: The judgment module is used to determine whether the G network and the D network simultaneously satisfy the convergence condition based on the loss function; The completion module is used to complete the training of the preset contrastive learning model when both the G network and the D network simultaneously meet the convergence condition.

9. The apparatus according to claim 6, characterized in that, Before the first input unit, the device further includes: The processing unit is used to perform non-overlapping image enhancement processing on each sample image according to at least two preset image enhancement methods.

10. The apparatus according to claim 8, characterized in that, The device further includes: The third input unit is used to input the image to be detected into the pre-trained contrast learning model; The second extraction unit is used to extract image features of the image to be detected based on the G network in the preset contrastive learning model.

11. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

13. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Model training method and device and electronic equipment

    CN113344089A