Training methods, feature extraction methods and devices for feature extraction models

By training the feature extraction model in two stages and combining inter-domain and intra-domain comparative learning, the problem of extracting only overall features while ignoring detailed features in existing technologies is solved. This achieves comprehensive capabilities of the feature extraction model in both overall and detailed features, thus expanding its application scope.

CN116129210BActive Publication Date: 2026-05-26MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MASHANG CONSUMER FINANCE CO LTD
Filing Date
2022-08-03
Publication Date
2026-05-26

Smart Images

  • Figure CN116129210B_ABST
    Figure CN116129210B_ABST
Patent Text Reader

Abstract

This application discloses a training method, feature extraction method, and apparatus for a feature extraction model, belonging to the field of computer science. The training method provided by this application includes: obtaining first-stage training samples; performing comparative learning training on the feature extraction model to be trained using the first-stage training samples to obtain a first target feature extraction model; obtaining second-stage training samples; and performing comparative learning training on the first target feature extraction model using the second-stage training samples to obtain a second target feature extraction model; wherein the first-stage training samples are one of training samples used for inter-domain comparative learning and training samples used for intra-domain comparative learning, and the second-stage training samples are the other of training samples used for inter-domain comparative learning and training samples used for intra-domain comparative learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer science, and specifically relates to a training method for a feature extraction model, a feature extraction method, and an apparatus. Background Technology

[0002] Currently, in many cases of processing objects such as images or text (e.g., face recognition or text-to-speech conversion), the feature information of the object is often extracted first.

[0003] In the process of extracting feature information of an object, related technologies often focus on extracting information about one aspect of the object (e.g., the overall information of the object) while ignoring information about another aspect of the object (e.g., the detailed feature information of the object). As a result, the application scope of this feature information extraction method is relatively limited. Summary of the Invention

[0004] This application provides a training method for a feature extraction model, a feature extraction method, and an apparatus to address the problem that the application scope of feature information extraction methods in related technologies is relatively limited.

[0005] In a first aspect, embodiments of this application provide a method for training a feature extraction model, the method comprising:

[0006] Obtain the first phase of training samples;

[0007] The first target feature extraction model is obtained by comparing and learning the feature extraction model to be trained using the first stage training samples.

[0008] Obtain the second phase training samples;

[0009] The first target feature extraction model is trained by comparative learning using the second stage training samples to obtain the second target feature extraction model.

[0010] The first-stage training samples are one of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning, and the second-stage training samples are the other of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning.

[0011] The training samples used for inter-domain contrastive learning include: a first anchor sample, a positive sample within the domain corresponding to the first anchor sample, and a negative sample outside the domain corresponding to the first anchor sample; the training samples used for intra-domain contrastive learning include: a second anchor sample, a positive sample within the domain corresponding to the second anchor sample, and a negative sample within the domain corresponding to the second anchor sample.

[0012] Secondly, embodiments of this application provide a feature extraction method, the method comprising:

[0013] Obtain the target data;

[0014] The target data is input into the second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data;

[0015] The second target feature extraction model is trained according to the training method provided in the first aspect.

[0016] Thirdly, embodiments of this application provide a feature extraction device, including an acquisition module and a processing module;

[0017] The acquisition module is used to acquire target data;

[0018] The processing module is used to input the target data into the second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data;

[0019] The second target feature extraction model is obtained based on the training method provided in the first aspect.

[0020] Fourthly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first or second aspect.

[0021] Fifthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first or second aspect.

[0022] In this embodiment, a first-stage training sample is obtained; a first target feature extraction model is obtained by performing comparative learning training on the feature extraction model to be trained using the first-stage training sample; a second-stage training sample is obtained; and a second target feature extraction model is obtained by performing comparative learning training on the first target feature extraction model using the second-stage training sample; wherein, the first-stage training sample is one of training samples for inter-domain comparative learning and training samples for intra-domain comparative learning, and the second-stage training sample is the other of training samples for inter-domain comparative learning and training samples for intra-domain comparative learning; wherein, the training sample for inter-domain comparative learning includes: a first anchor sample, an intra-domain positive sample corresponding to the first anchor sample, and an extra-domain negative sample corresponding to the first anchor sample; the training sample for intra-domain comparative learning includes: a second anchor sample, an intra-domain positive sample corresponding to the second anchor sample, and an intra-domain negative sample corresponding to the second anchor sample. In this way, the model is trained in two stages using first-stage training samples and second-stage training samples. One stage focuses on inter-domain comparative learning, and the other stage focuses on intra-domain comparative learning. The resulting second target feature extraction model not only has the ability to distinguish the overall features of the data, but also has the ability to further distinguish the detailed features of the data. This greatly expands the application scope of the model and solves the problem that the application scope of feature information extraction methods in related technologies is relatively limited. Attached Figure Description

[0023] Figure 1 A schematic diagram illustrating the training process of a feature extraction model provided in an embodiment of this application;

[0024] Figure 2 A schematic diagram illustrating the training process of a feature extraction model provided in an embodiment of this application;

[0025] Figure 3 A schematic flowchart illustrating a training method for a feature extraction model provided in an embodiment of this application;

[0026] Figure 4-1 A schematic flowchart illustrating a training method for a feature extraction model provided in an embodiment of this application;

[0027] Figure 4-2 A schematic flowchart illustrating a training method for a second target feature extraction model provided in an embodiment of this application;

[0028] Figure 4-3 A schematic diagram illustrating the training process of a second target feature extraction model provided in an embodiment of this application;

[0029] Figure 5A schematic flowchart illustrating a feature extraction method provided in an embodiment of this application;

[0030] Figure 6 A schematic structural diagram of a training device for a feature extraction model provided in an embodiment of this application;

[0031] Figure 7 A schematic structural diagram of an electronic device provided in an embodiment of this application;

[0032] Figure 8 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0034] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0035] In the process of extracting feature information, feature extraction models can be used to extract features from input data (such as text or images). Since the quality of the feature extraction model directly affects the quality of the extracted feature information, it is crucial to train the model to ensure that the generated feature information accurately reflects as many features as possible of the input data. In related technologies, during the training of the feature extraction model, a specified sample is typically used as an anchor sample. Positive samples generated within the specified domain after enhancement processing are used as positive samples, and samples from different domains are used as negative samples. The feature extraction model is then trained through comparative learning, enabling it to distinguish the overall features of the data, thus ensuring that the feature information output by the model matches the features of the input data.

[0036] In this context, positive samples within the domain corresponding to anchor samples can be derived from anchor samples. Anchor samples and positive samples within the domain are similar in both overall and detailed features. Negative samples within the domain corresponding to anchor samples can be derived from anchor samples. Anchor samples and negative samples within the domain are similar in overall features but not so similar in detailed features. Anchor samples and their corresponding negative samples outside the domain are dissimilar in both overall and detailed features. Overall features represent the overall information of the sample, while detailed features represent the local features of the sample in specific aspects. For example, if the sample is an image of a cartoon-style cat, the overall feature of the sample could be the category "cat," while the detailed features could be the proportion of a certain part of the cat (e.g., a cat's paw) in the sample, the color of a certain part of the cat (e.g., a cat's paw), etc. Of course, the sample can also be a text sample, an audio sample, etc., and this application does not impose specific limitations here.

[0037] However, the applicant noted that in the relevant technologies, by using data from different domains as negative samples, the feature extraction model can only distinguish the overall features of the input data, but it is difficult to distinguish the detailed features of the input data. This results in the feature information generated by the feature extraction model being incomplete, which in turn limits its application scope.

[0038] Based on this, this application provides a training method for a feature extraction model. The model is trained in two stages using first-stage training samples and second-stage training samples. One stage focuses on inter-domain comparative learning, while the other stage focuses on intra-domain comparative learning. The resulting second target feature extraction model not only has the ability to distinguish the overall features of the data, but also has the ability to further distinguish the detailed features of the data. This greatly expands the application scope of the model and solves the problem in related technologies where it is difficult to extract the detailed feature information of objects using feature extraction models, resulting in a limited application scope.

[0039] For example, such as Figure 1As shown, the training process of the feature extraction model provided in this application embodiment may include: using specified samples in the first-stage dataset as first anchor samples; using positive example enhancement to enhance the first anchor samples to generate intra-domain positive samples corresponding to the first anchor samples; using samples in the first-stage dataset other than the specified samples as out-of-domain negative samples corresponding to the first anchor samples; using the first anchor samples, the intra-domain positive samples corresponding to the first anchor samples, and the out-of-domain negative samples corresponding to the first anchor samples as first-stage training samples, and performing first-stage comparative learning training on the feature extraction model to be trained to obtain a first target feature extraction model, so that the first target feature extraction model has the ability to distinguish the overall features of the data; then, using the second-stage dataset... The specified sample is used as the second anchor sample; positive example enhancement is used to enhance the second anchor sample to generate positive samples within the domain; negative example enhancement is used to enhance the second anchor sample to generate negative samples within the domain; the second anchor sample, the corresponding positive samples within the domain, and the corresponding negative samples within the domain are used as the second-stage training samples to perform the second-stage comparative learning training on the first target feature extraction model, thus obtaining the second target feature extraction model. This enables the second target feature extraction model to not only distinguish the overall features of the data, but also to further distinguish the detailed features of the data, greatly expanding the application scope of the model and thus solving the problem that the application scope of feature information extraction methods in related technologies is relatively limited.

[0040] Furthermore, to prevent the ability to distinguish overall data features obtained from the first stage of training from degrading during the second stage of training, in practical applications, out-of-domain negative samples corresponding to the second anchor point samples can be added to the second stage of training samples. Figure 1 (Not shown), a small number of out-of-domain negative samples corresponding to the second anchor point samples are used for contrastive learning training to maintain the overall ability of the discriminative data obtained from the first stage of training. In this way, the effectiveness of the second stage of training can be improved while ensuring the effectiveness of the first stage of training.

[0041] It should be pointed out that, Figure 1 The training process of the above-mentioned feature extraction model uses inter-domain learning samples in the first stage of training to train the model's ability to distinguish the overall features of the data; and uses intra-domain learning samples in the second stage of training to train the model's ability to distinguish the detailed features of the data. The above embodiments are merely examples and do not imply limitation. For example, ... Figure 2As shown in the embodiments of this application, a training method for a feature extraction model is also provided. In the first stage of training, intra-domain learning samples are used to train the feature extraction model to distinguish the detailed features of the data, and in the second stage of training, inter-domain learning samples are used to train the feature extraction model to distinguish the overall features of the data.

[0042] like Figure 2 As shown, the training process of the feature extraction model provided in this application embodiment may include: using the second anchor sample, the positive sample in the domain corresponding to the second anchor sample, and the negative sample in the domain corresponding to the second anchor sample as the first-stage training samples, performing the first-stage comparative learning training on the feature extraction model to be trained, to obtain the first target feature extraction model, so that the first target feature extraction model has the ability to distinguish the detailed features of the data; then, using the first anchor sample, the positive sample in the domain corresponding to the first anchor sample, and the negative sample out of domain corresponding to the first anchor sample as the second-stage training samples, performing the second-stage comparative learning training on the first target feature extraction model, to obtain the second target feature extraction model, so that the second target feature extraction model, in addition to having the ability to distinguish the detailed features of the data, also has the ability to further distinguish the overall features of the data, greatly expanding the application scope of the model, thereby solving the problem that the application scope of feature extraction methods in related technologies is relatively limited.

[0043] This application also provides a feature extraction method. A pre-trained second target feature extraction model is used to perform feature extraction processing on input objects such as images or text, generating feature information such as feature vectors or feature matrices. The generated feature information conforms to the overall and detailed features of the input object, greatly expanding the application scope of the model. For example, in practical applications, when the input object is an image, the pre-trained feature extraction model provided in this application can be used as a feature extraction layer for image-based neural network models such as image classification models, image recognition models, image conversion models, or image segmentation models. When the input object is a text-based object, the pre-trained feature extraction model provided in this application can be used as a feature extraction layer for text-based neural network models such as text classification models, text recognition models, text retrieval models, or text tag extraction models. Of course, the above examples are merely illustrations and do not imply limitation. When the target model to be trained includes a feature extraction layer for feature extraction processing, the training method provided in this application can be used to pre-train the feature extraction layer of the target model.

[0044] The training method, feature extraction method, and apparatus of the feature extraction model provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0045] Figure 3 This is a schematic flowchart illustrating a training method for a feature extraction model provided in an embodiment of this application.

[0046] like Figure 3 As shown, the training method for the feature extraction model provided in this application embodiment may include:

[0047] Step 310: Obtain the first-stage training samples; wherein, the first-stage training samples are one of the training samples used for inter-domain contrastive learning and the training samples used for intra-domain contrastive learning.

[0048] Step 320: Use the training samples from the first stage to perform comparative learning training on the feature extraction model to be trained, and obtain the first target feature extraction model;

[0049] Step 330: Obtain the second-stage training samples; the second-stage training samples are the other of the training samples used for inter-domain contrastive learning and the training samples used for intra-domain contrastive learning.

[0050] Step 340: The first target feature extraction model is trained by comparative learning using the training samples from the second stage to obtain the second target feature extraction model;

[0051] The training samples used for inter-domain contrastive learning include: a first anchor sample, a positive sample within the domain corresponding to the first anchor sample, and a negative sample outside the domain corresponding to the first anchor sample; the training samples used for intra-domain contrastive learning include: a second anchor sample, a positive sample within the domain corresponding to the second anchor sample, and a negative sample within the domain corresponding to the second anchor sample.

[0052] The first and second phase training samples can be in the form of text, images, etc. It should be noted that the concept of an in-domain sample refers to a sample whose overall features are highly similar to those of the anchor sample, while the concept of an out-of-domain sample refers to a sample whose overall features are not very similar to those of the anchor sample.

[0053] For example, a positive sample within the domain corresponding to the first anchor sample can be derived from the first anchor sample. The first anchor sample and the positive sample within the domain are similar in both overall and detailed features. A negative sample within the domain corresponding to the first anchor sample can be derived from the first anchor sample. The first anchor sample and the negative sample within the domain are similar in overall features but not so similar in detailed features. The first anchor sample and the negative sample outside the domain corresponding to the first anchor sample are dissimilar in both overall and detailed features.

[0054] It is understood that steps 310 and 320 above perform the first stage of comparative learning training for the feature extraction model, while steps 330 and 340 perform the second stage of comparative learning training for the feature extraction model. The training samples for the two learning stages can be interchanged. In other words, if the first stage training samples are used for inter-domain comparative learning, the second stage training samples can be used for intra-domain comparative learning; conversely, if the first stage training samples are used for intra-domain comparative learning, the second stage training samples can be used for inter-domain comparative learning. The following sections will explain each of these two cases in detail.

[0055] The first scenario: In the first stage of training, training samples for inter-domain contrastive learning are used to train the feature extraction model's ability to distinguish the overall features of the data; in the second stage of training, training samples for intra-domain contrastive learning are used to further train the feature extraction model's ability to distinguish the detailed features of the data.

[0056] In this case, in step 310, the first-stage training samples can be training samples used for inter-domain contrastive learning. The training samples used for inter-domain contrastive learning may include: a first anchor sample, positive samples within the domain corresponding to the first anchor sample, and negative samples outside the domain corresponding to the first anchor sample.

[0057] The first anchor sample can be obtained from a pre-prepared specified dataset according to actual needs. The positive samples within the domain corresponding to the first anchor sample can be obtained by performing data augmentation on the first anchor sample using positive example data augmentation. The negative samples outside the domain corresponding to the first anchor sample can be obtained from the aforementioned specified dataset. This application does not impose any specific restrictions on these.

[0058] In step 320, the first anchor sample, the positive sample within the domain corresponding to the first anchor sample, and the negative sample outside the domain corresponding to the first anchor sample are used to perform comparative learning training on the feature extraction model to be trained, so as to obtain the first target feature extraction model. Since the first anchor sample and the positive sample within the domain corresponding to the first anchor sample are highly similar in terms of overall features, and the first anchor sample and the negative sample outside the domain corresponding to the first anchor sample are dissimilar in terms of overall features, the first target feature extraction model has the ability to distinguish the overall features of the data.

[0059] In step 330, the training samples for the second stage can be training samples used for inter-domain contrastive learning. The training samples used for inter-domain contrastive learning include: a second anchor sample, positive samples within the domain corresponding to the second anchor sample, and negative samples within the domain corresponding to the second anchor sample. The second anchor sample can be obtained from a pre-prepared specified dataset according to actual needs. The positive samples within the domain corresponding to the second anchor sample can be obtained by augmenting the second anchor sample with positive examples, and the negative samples within the domain corresponding to the second anchor sample can be obtained by augmenting the second anchor sample with negative examples. This application does not impose specific limitations on these methods.

[0060] The first anchor sample and the second anchor sample can belong to the same dataset or different datasets. When the first anchor sample and the second anchor sample belong to the same dataset, the first anchor sample and the second anchor sample can be the same or different; when the first anchor sample and the second anchor sample belong to different datasets, the first anchor sample and the second anchor sample can be different. This application does not impose specific restrictions here.

[0061] In step 340, the first target feature extraction model is trained by comparative learning using the second anchor sample, the positive sample in the domain corresponding to the second anchor sample, and the negative sample in the domain corresponding to the second anchor sample, to obtain the second target feature extraction model. Since the second anchor sample and the positive sample in the domain corresponding to the second anchor sample are highly similar in terms of both overall features and detailed features, and the second anchor sample and the negative sample in the domain corresponding to the second anchor sample are similar in terms of overall features but dissimilar in terms of detailed features, the second target feature extraction model can learn to distinguish the detailed features of the data. Thus, the second target feature extraction model has the ability to distinguish the detailed features of the data in addition to the ability to distinguish the overall features of the data.

[0062] The second scenario: In the first stage of training, training samples used for inter-domain contrastive learning are used to train the feature extraction model's ability to distinguish the overall features of the data; in the second stage of training, training samples used for intra-domain contrastive learning are used to further train the feature extraction model's ability to distinguish the detailed features of the data.

[0063] In this case, in step 310, the first-stage training samples can be training samples used for intra-domain contrastive learning. The training samples used for intra-domain contrastive learning include: the second anchor sample, the intra-domain positive sample corresponding to the second anchor sample, and the intra-domain negative sample corresponding to the second anchor sample.

[0064] In step 320, the second anchor sample, the positive sample in the domain corresponding to the second anchor sample, and the negative sample in the domain corresponding to the second anchor sample are used to perform comparative learning training on the feature extraction model to be trained, so as to obtain the first target feature extraction model. Since the second anchor sample and the positive sample in the domain corresponding to the second anchor sample are similar in terms of both overall features and detailed features, and the second anchor sample and the negative sample in the domain corresponding to the second anchor sample are similar in terms of overall features but dissimilar in terms of detailed features, the first target feature extraction model learns the ability to distinguish the detailed features of the data.

[0065] In step 330, the training samples for the second stage can be training samples used for inter-domain contrastive learning. The training samples used for inter-domain contrastive learning may include: a first anchor sample, positive samples within the domain corresponding to the first anchor sample, and negative samples outside the domain corresponding to the first anchor sample.

[0066] In step 340, the first target feature extraction model is trained by comparative learning using the first anchor sample, the positive sample within the domain corresponding to the first anchor sample, and the negative sample outside the domain corresponding to the first anchor sample, to obtain the second target feature extraction model. Since the first anchor sample and the positive sample within the domain corresponding to the first anchor sample are highly similar in terms of overall features, and the first anchor sample and the negative sample outside the domain corresponding to the first anchor sample are dissimilar in terms of overall features, the second target feature extraction model learns the ability to distinguish the overall features of the data. Thus, the second target feature extraction model has the ability to distinguish the overall features of the data in addition to the ability to distinguish the detailed features of the data.

[0067] According to the training method of the feature extraction model provided in the embodiments of this application, a first-stage training sample is obtained; a first target feature extraction model is obtained by comparative learning training the feature extraction model to be trained using the first-stage training sample; a second-stage training sample is obtained; and a second target feature extraction model is obtained by comparative learning training the first target feature extraction model using the second-stage training sample; wherein, the first-stage training sample is one of training samples for inter-domain comparative learning and training samples for intra-domain comparative learning, and the second-stage training sample is the other of training samples for inter-domain comparative learning and training samples for intra-domain comparative learning; wherein, the training sample for inter-domain comparative learning includes: a first anchor sample, an intra-domain positive sample corresponding to the first anchor sample, and an extra-domain negative sample corresponding to the first anchor sample; the training sample for intra-domain comparative learning includes: a second anchor sample, an intra-domain positive sample corresponding to the second anchor sample, and an intra-domain negative sample corresponding to the second anchor sample. In this way, the model is trained in two stages using first-stage training samples and second-stage training samples. One stage focuses on inter-domain comparative learning, and the other stage focuses on intra-domain comparative learning. The resulting second target feature extraction model not only has the ability to distinguish the overall features of the data, but also has the ability to further distinguish the detailed features of the data. This greatly expands the application scope of the model and solves the problem that the application scope of feature information extraction methods in related technologies is relatively limited.

[0068] In one specific embodiment, to prevent the feature extraction capability obtained from the first stage of training from degrading during the second stage of training, in practical applications, a small number of negative samples similar to those in the first stage of training can be added to the second stage of training samples to maintain the feature extraction capability obtained from the first stage of training. An example is given below.

[0069] For example, if the first-stage training samples are training samples for inter-domain contrastive learning and the second-stage training samples are training samples for intra-domain contrastive learning, the second-stage training samples may include the second anchor sample, the intra-domain positive sample corresponding to the second anchor sample, and the intra-domain negative sample corresponding to the second anchor sample. In this case, the second-stage training samples may also include the out-of-domain negative sample corresponding to the second anchor sample.

[0070] In this way, by adding a small number of out-of-domain negative samples to the training samples in the second stage, the model is prevented from focusing too much on details during the second stage training process, which would affect the model's ability to distinguish overall features. This allows the second target feature extraction model to maintain the ability to distinguish the overall features of the data obtained in the first stage training, while also having the ability to distinguish the detailed features of the data.

[0071] In practical applications, when the second-stage training samples include the second anchor sample, positive samples within the domain corresponding to the second anchor sample, negative samples within the domain corresponding to the second anchor sample, and negative samples outside the domain corresponding to the second anchor sample, the number of negative samples within the domain corresponding to the second anchor sample is greater than or equal to the number of negative samples outside the domain corresponding to the second anchor sample. For example, the ratio of the number of negative samples within the domain to the number of negative samples outside the domain is 1:1 or 2:1, etc. Thus, during the second-stage training, a small number of negative samples outside the domain can be used for comparative learning training to maintain the overall ability to distinguish data obtained in the first-stage training, while a sufficient number of negative samples within the domain can be used for comparative learning training, enabling the second target feature extraction model to also distinguish the detailed features of the data. Of course, the number of negative samples within the domain corresponding to the second anchor sample can also be slightly less than the number of negative samples outside the domain corresponding to the second anchor sample; this application does not impose specific restrictions on this.

[0072] For example, if the first-stage training samples are training samples used for intra-domain contrastive learning and the second-stage training samples are training samples used for inter-domain contrastive learning, the second-stage training samples may include the first anchor sample, the intra-domain positive sample corresponding to the first anchor sample, and the extra-domain negative sample corresponding to the first anchor sample. In this case, the second-stage training samples may also include the intra-domain negative sample corresponding to the first anchor sample.

[0073] In this way, by adding a small number of negative samples within the domain to the training samples in the second stage, the model is prevented from focusing too much on the overall features during the second stage training process, which would affect the model's ability to distinguish detailed features. This allows the second target feature extraction model to maintain the ability to distinguish the detailed features of the data obtained in the first stage training, while also having the ability to distinguish the overall features of the data.

[0074] Similarly, in practical applications, when the second-stage training samples include the first anchor sample, positive samples within the domain corresponding to the first anchor sample, negative samples outside the domain corresponding to the first anchor sample, and negative samples within the domain corresponding to the first anchor sample, the number of negative samples outside the domain corresponding to the first anchor sample is greater than or equal to the number of negative samples within the domain corresponding to the first anchor sample. For example, the ratio of the number of negative samples outside the domain to the number of negative samples within the domain is 1:1 or 2:1, etc. Thus, during the second-stage training, a small number of negative samples within the domain can be used for contrastive learning training to maintain the ability to distinguish the detailed features of the data obtained in the first-stage training, while a sufficient number of negative samples outside the domain can be used for contrastive learning training, enabling the second target feature extraction model to distinguish the overall features of the data. Of course, the number of negative samples outside the domain corresponding to the first anchor sample can also be slightly less than the number of negative samples within the domain corresponding to the first anchor sample; this application does not impose specific restrictions on this.

[0075] The following example uses the first-stage training samples as training samples for inter-domain contrastive learning and the second-stage training samples as training samples for intra-domain contrastive learning to illustrate in detail the training method of the feature extraction model provided in this application.

[0076] Figure 4-1 A schematic flowchart illustrating another training method for a feature extraction model provided in an embodiment of this application.

[0077] like Figure 4-1 As shown, the training method for the feature extraction model provided in this application embodiment may include:

[0078] Step 410: Obtain the first anchor point sample, the positive sample within the domain corresponding to the first anchor point sample, and the negative sample outside the domain corresponding to the first anchor point sample;

[0079] Step 410 can be a sub-step of step 310.

[0080] In step 410, this embodiment of the application can obtain a first anchor sample from the first-stage dataset; samples in the first-stage dataset other than the first anchor sample can be used as out-of-domain negative samples corresponding to the first anchor sample. The first-stage dataset can be a specified dataset determined according to actual needs, and the first anchor sample can be any sample in the first-stage dataset.

[0081] In step 410, this embodiment of the application can perform data augmentation processing on the first anchor sample using a positive example data augmentation method to obtain a domain-specific positive sample corresponding to the first anchor sample. The first anchor sample can be a text sample or an image sample.

[0082] In the case where the first anchor sample is a text sample, the positive example data augmentation method can be at least one of the following: random deletion, masking, back translation, repetition, position transformation, and synonym replacement. For example, a positive example data augmentation method can be random deletion, which involves deleting at least a portion of the content of the first anchor sample to obtain a domain-specific positive sample corresponding to the first anchor sample. Another example is masking, which involves masking at least a portion of the content of the first anchor sample to obtain a domain-specific positive sample corresponding to the first anchor sample. Yet another example is back translation, which involves translating at least a portion of the content of the first anchor sample into a specified language to obtain a domain-specific positive sample corresponding to the first anchor sample. Yet another example is repetition, which involves repeating at least a portion of the content of the first anchor sample to obtain a domain-specific positive sample corresponding to the first anchor sample. Finally, a positive example data augmentation method can be position transformation, which involves changing the positions of various phrases in the first anchor sample to obtain a domain-specific positive sample corresponding to the first anchor sample. For example, positive data augmentation can also be done by synonym replacement. By replacing several random word groups in the first anchor sample with synonyms, positive samples within the domain corresponding to the first anchor sample can be obtained.

[0083] For example, the first anchor point sample in text form is "In this rapidly developing modern era, in order not to be overwhelmed by the torrent of time, people can only work even harder to improve themselves and make a living." The overall features of this first anchor point sample can include the meaning expressed by the sentence, such as "people in modern times improve themselves to make a living." Detailed features can include the specific modifying content of any of the sentence's subject, predicate, object, or adverbial phrase. For example, detailed features could include the modifying content of "era" ("rapidly developing"), the modifying content of "improve" ("working even harder"), and so on. No specific restrictions are imposed here.

[0084] After the first anchor sample is processed using positive example data augmentation methods (such as synonym replacement), the resulting domain-specific positive sample corresponding to the first anchor sample can be "In the rapidly developing modern era, in order not to be submerged by the torrent of time, people can only work tirelessly to improve themselves and be busy making a living." The overall features of this domain-specific positive sample can include the meaning expressed by the sentence, such as "people improve themselves in modern times to make a living." Detailed features can include the specific modifying content of any one of the sentence's subject, predicate, object, or adverbial phrase. For example, detailed features could include the modifying content of "era" as "rapidly developing," the modifying content of "improve" as "work tirelessly," and so on. No specific restrictions are imposed here.

[0085] The negative sample from outside the domain corresponding to the first anchor point sample could be "Mashang Consumer Finance is a technology-driven financial institution approved by the China Banking and Insurance Regulatory Commission and holding a consumer finance license." The overall features of this negative sample can include the meaning expressed in the sentence, such as "Mashang Consumer Finance is a financial institution." Detailed features can include the specific modifications of any one of the sentence's subject, predicate, object, or adverbial phrases. For example, detailed features could include the modifications to "financial institution" such as "technology-driven" and "approved by the China Banking and Insurance Regulatory Commission and holding a consumer finance license," etc., without specific limitations here.

[0086] In the case where the first anchor sample is an image, the positive example data augmentation method can be at least one of the following: image rotation, blurring, random image resizing, style transfer, etc. For example, a positive example data augmentation method could be image rotation, which involves rotating the first anchor sample by a specified angle to obtain a domain-specific positive sample corresponding to the first anchor sample. Another example is blurring, which involves blurring the first anchor sample to obtain a domain-specific positive sample corresponding to the first anchor sample. Yet another example is random image resizing, which involves enlarging or reducing the first anchor sample to obtain a domain-specific positive sample corresponding to the first anchor sample. Finally, a positive example data augmentation method could be style transfer, which involves transferring the original style of the first anchor sample to a specified style to obtain a domain-specific positive sample corresponding to the first anchor sample.

[0087] Step 420: Using the first anchor sample, the positive samples within the domain corresponding to the first anchor sample, and the negative samples outside the domain corresponding to the first anchor sample, perform comparative learning training on the feature extraction model to be trained to obtain the first target feature extraction model.

[0088] Step 420 can be a sub-step of step 320.

[0089] In step 420, F data pairs can be input into the feature extraction model to be trained, wherein the data pairs include a first anchor sample and a first pairing sample; the first pairing sample can be a positive sample within the domain corresponding to the first anchor sample, or a negative sample outside the domain corresponding to the first anchor sample.

[0090] Among them, the F data pairs include at least one of the fourth type data pairs and the fifth type data pairs; the fourth type data pair includes the first anchor sample and the positive in-domain sample corresponding to the first anchor sample; the fifth type data pair includes the first anchor sample and the negative out-of-domain sample corresponding to the first anchor sample.

[0091] Specifically, for one target data pair among the F data pairs: the target similarity between the first anchor sample and the first paired sample in the target data pair is determined by the feature extraction model to be trained; the contrast loss of the target data pair is determined based on the target similarity and the reference similarity of the target data pair; and the parameters of the feature extraction model to be trained are adjusted based on the contrast loss of each data pair among the F data pairs to obtain the first target feature extraction model.

[0092] Specifically, when the first paired sample in the target data pair is a positive sample within the domain corresponding to the first anchor sample, the reference similarity of the target data pair is 1; when the first paired sample in the target data pair is a negative sample outside the domain corresponding to the first anchor sample, the reference similarity of the target data pair is 0.

[0093] Step 430: Obtain the second anchor point sample from the second-stage dataset;

[0094] Step 430 can be a sub-step of step 330;

[0095] In step 430, the second-stage dataset can be a specified dataset determined according to actual needs, and the second anchor sample can be any sample in the second-stage dataset.

[0096] The second-stage dataset can be the same as or different from the first-stage dataset. If the second-stage dataset and the first-stage dataset are the same, the first anchor sample and the second anchor sample can be the same or different; if the second-stage dataset and the first-stage dataset are different, the first anchor sample and the second anchor sample can be different. This application does not impose specific restrictions here.

[0097] Step 440: Perform data augmentation on the second anchor sample using positive example data augmentation to obtain positive samples within the domain corresponding to the second anchor sample.

[0098] Step 440 can be a sub-step of step 330.

[0099] In step 440, the second anchor sample can be a text sample or an image sample. When the second anchor sample is a text sample, the positive example data augmentation method can be at least one of the following: random deletion, masking, back translation, repetition, position transformation, synonym replacement, etc. When the second anchor sample is an image sample, the positive example data augmentation method can be at least one of the following: image rotation, blurring, random image resizing, style transfer, etc. The specific details are similar to step 410; refer to the specific content of data augmentation processing of the first anchor sample using positive example data augmentation methods in step 410, which will not be repeated here. In this way, multiple positive example data augmentation methods can be used to augment the second anchor sample, resulting in multiple domain-specific positive samples corresponding to the second anchor sample, thus improving the anti-interference capability of the feature extraction model.

[0100] Step 450: Perform data augmentation on the second anchor sample using negative example data augmentation to obtain the domain negative sample corresponding to the second anchor sample.

[0101] Step 450 can be a sub-step of step 330.

[0102] In step 450, the second anchor sample can be a text sample or an image sample.

[0103] In step 450, when the second anchor sample is a text sample, the negative example data augmentation method can be at least one of the following: adding negative words, replacing antonyms, etc. Examples are given below.

[0104] For example, if the second anchor sample is a text sample and the negative sample data augmentation method is to add a negative word, the above step 450 may include: splitting the text sample into word groups to obtain multiple independent word groups; and obtaining the domain negative sample corresponding to the second anchor sample by adding a negative word before at least one of the multiple independent word groups.

[0105] For example, if the second anchor sample is a text sample and the negative example data augmentation method includes antonym replacement, step 450 above may include:

[0106] The second anchor point sample in text form is split into word groups to obtain P independent word groups;

[0107] By replacing Q independent word groups out of P independent word groups with their corresponding antonyms, we obtain the domain negative sample corresponding to the second anchor sample.

[0108] Where Q and P are positive integers, the quotient of Q and P lies between the first threshold and the second threshold, the first threshold is greater than 0, and the second threshold is less than 1.

[0109] For example, if the value of P is twice the value of Q, half of the word groups in the second anchor sample in text form are replaced with their corresponding antonyms to obtain the domain negative sample corresponding to the second anchor sample.

[0110] For example, the second anchor sample could be: "In this rapidly developing modern era, in order not to be swept away by the tide of time, people can only work even harder to improve themselves and make a living." The overall characteristics of this second anchor sample can include the meaning expressed in the sentence, such as "people in modern times improve themselves to make a living." Detailed characteristics can include the specific modifying content of any of the sentence's subject, predicate, object, or adverbial phrase. For example, detailed characteristics could include the modifying content of "the era" ("rapidly developing"), the modifying content of "improve" ("working even harder"), and so on. No specific restrictions are imposed here.

[0111] The negative sample obtained after the second anchor point sample is processed by antonym replacement can be "In this modern, slowly developing era, in order not to be submerged by the torrent of time, people can only eat more and spend their days improving themselves, in order to make a living." The overall features of this negative sample can include the meaning expressed in the sentence, such as "people in modern times improve themselves in order to make a living." Detailed features can include the specific modifying content of any one of the sentence's subject, predicate, object, or adverbial phrase. For example, detailed features could include the modifying content of "era" as "slowly developing," the modifying content of "improve" as "eat more and spend their days," etc., without specific limitations here.

[0112] In this way, compared with adding negation words, generating negative samples within the domain by replacing negative words can avoid the model only learning the existence and differences of negation words between data pairs during subsequent training. It can change the original semantics when the overall features of the original text data of the second anchor sample in text form do not change much. This allows the model to not only focus on sample features, but also to gain a deeper understanding of text semantics, especially the different semantics of the replaced antonyms, thus improving the model's semantic understanding ability.

[0113] In step 450, when the second anchor sample is an image, the negative example data augmentation method can be at least one of image rotation, blurring, random image resizing, style transfer, etc. Examples are given below.

[0114] For example, if the second anchor sample is an image and the negative example data augmentation method includes color adjustment, step 450 above may include:

[0115] Determine the target object in the second anchor point sample in image form;

[0116] By adjusting the color of the target object to a specified color, a negative sample within the domain corresponding to the second anchor point sample is obtained.

[0117] For example, if the target object in the second anchor point sample in the form of an image is determined to be the eye area, the negative sample in the domain corresponding to the second anchor point sample is obtained by adjusting the color of the eye area to blue.

[0118] For example, if the second anchor sample is an image and the negative example data augmentation method includes content replacement, step 450 above may include:

[0119] Determine the target object in the second anchor point sample in image form;

[0120] By replacing the target object with a specified object, a domain-specific negative sample corresponding to the second anchor point sample is obtained.

[0121] For example, if the target object in the second anchor point sample in the form of an image is determined to be the ear, the negative sample in the domain corresponding to the second anchor point sample is obtained by replacing the ear with a rectangle.

[0122] For example, if the second anchor sample is an image and the negative example data augmentation method includes size adjustment, step 450 above may include:

[0123] Determine the target object in the second anchor point sample in image form;

[0124] By adjusting the size of the target object to a specified size, a negative sample within the domain corresponding to the second anchor point sample is obtained.

[0125] For example, if the target object in the second anchor point sample in the form of an image is determined to be the lips, then by enlarging or shrinking the size of the lips, a negative sample in the domain corresponding to the second anchor point sample can be obtained.

[0126] In this way, various negative sample data augmentation methods can be used to augment the data of the second anchor sample, resulting in multiple domain-specific negative samples corresponding to the second anchor sample, thereby improving the anti-interference ability of the feature extraction model.

[0127] Step 460: Take the samples in the second stage dataset other than the second anchor sample as the out-of-domain negative samples corresponding to the second anchor sample.

[0128] Step 460 can be a sub-step of step 330. Step 430 can be executed first, followed by steps 440, 450 and 460, and there is no specific restriction on the execution order of steps 440, 450 and 460.

[0129] Step 470: Use the second anchor sample, the positive sample within the domain corresponding to the second anchor sample, the negative sample within the domain corresponding to the second anchor sample, and the negative sample outside the domain corresponding to the second anchor sample as training samples for the second stage, and perform comparative learning training on the first target feature extraction model to obtain the second target feature extraction model.

[0130] In the second-stage training samples, the number of negative samples within the domain corresponding to the second anchor sample can be greater than or equal to the number of negative samples outside the domain corresponding to the second anchor sample.

[0131] Step 470 can be a sub-step of step 340.

[0132] In one specific embodiment, step 470 above may include:

[0133] Step 4701: Obtain N data pairs from the second-stage training samples. Each data pair includes a second anchor sample and a second pairing sample corresponding to the second anchor sample. The second pairing sample is either a positive sample within the domain corresponding to the second anchor sample, a negative sample outside the domain corresponding to the second anchor sample, or a negative sample within the domain corresponding to the second anchor sample.

[0134] Where N is an integer greater than or equal to 2;

[0135] Step 4702: Input the N data pairs into the first target feature extraction model;

[0136] Where N is an integer greater than or equal to 3; the N data pairs include at least one of the first type of data pairs, the second type of data pairs, and the third type of data pairs; the first type of data pair includes the second anchor sample and the positive sample within the domain corresponding to the second anchor sample; the second type of data pair includes the second anchor sample and the negative sample outside the domain corresponding to the second anchor sample; the third type of data pair includes the second anchor sample and the negative sample within the domain corresponding to the second anchor sample.

[0137] Step 4703: For one target data pair among N data pairs: Determine the target similarity between the second anchor sample and the second paired sample in the target data pair using the first target feature extraction model; Determine the contrast loss of the target data pair based on the target similarity and the reference similarity of the target data pair.

[0138] Among them, such as Figure 4-3As shown, the first target feature extraction model may include: a feature extraction layer for performing feature extraction processing on the input data and a similarity calculation layer for performing similarity calculation processing.

[0139] Furthermore, such as Figure 4-3 As shown, for one target data pair among the N data pairs: a feature extraction layer is used to extract features of the target data pair to obtain the first feature information corresponding to the second anchor sample and the second feature information corresponding to the second pairing sample; a similarity calculation layer is used to calculate the similarity between the first feature information and the second feature information, which is used as the target similarity between the second anchor sample and the second pairing sample.

[0140] The first target feature extraction model can be a neural network learning model, such as... Figure 4-3 The layer shown consists of a feature extraction layer for feature extraction processing of input data and a similarity calculation layer for similarity calculation processing. Alternatively, it may also consist of a BERT (Bidirectional Encoder Representation from Transformers) structure, a Roberta (A Robustly Optimized BERT) structure, an XLNet (Transformer-XL based neural network) structure, a VGGNet (Visual Geometry Group Network) structure, a ResNet (Residual Network) structure, a DeBERTa (Decoding-enhanced BERT with disentangled attention) structure, etc. This application does not impose specific limitations.

[0141] Specifically, the reference similarity of the target data pair is 1 when the second paired sample in the target data pair is a positive sample within the domain corresponding to the second anchor sample; the reference similarity of the target data pair is 0 when the second paired sample in the target data pair is a negative sample outside the domain corresponding to the second anchor sample; and the reference similarity of the target data pair is 0 when the second paired sample in the target data pair is a negative sample within the domain corresponding to the second anchor sample.

[0142] In this way, the first feature information corresponding to the second anchor sample and the second feature information corresponding to the second paired sample are extracted by the first target feature extraction model, and the similarity between the first feature information and the second feature information is calculated as the target similarity between the second anchor sample and the second paired sample.

[0143] In this embodiment, the second anchor sample and the second pairing sample in the target data pair can be respectively input to the feature extraction layer to extract the first feature information corresponding to the second anchor sample and the second feature information corresponding to the second pairing sample; or, in this embodiment, the second anchor sample and the second pairing sample in the target data pair can be simultaneously input to the feature extraction layer to simultaneously extract the first feature information corresponding to the second anchor sample and the second feature information corresponding to the second pairing sample.

[0144] The first and second feature information can be data in the form of feature vectors or feature matrices. The feature extraction layer can include one or more feature extraction sub-layers. When the feature extraction layer includes multiple feature extraction sub-layers, the output of one sub-layer can be used as the input to the next sub-layer. Therefore, when the second anchor sample from the target data pair is input to the feature extraction layer, the output of the last layer of multiple feature extraction sub-layers can be used as the first feature information; similarly, when the second pairing sample from the target data pair is input to the feature extraction layer, the output of the last layer of multiple feature extraction sub-layers can be used as the second feature information.

[0145] In the case where the second anchor sample in the target data pair is a text sample, the first feature information corresponding to the second anchor sample can include the overall features and detailed features of the second anchor sample. The overall features of the second anchor sample can include the overall meaning of the sentence, and the detailed features of the second anchor sample can include the specific modifying content of any one of the subject, predicate, object, or adverbial of the sentence, etc., without specific limitations in this application. For example, the second anchor sample is "In this rapidly developing modern era, in order not to be submerged by the torrent of time, people can only work harder and harder to improve themselves and make a living." Here, the overall features of this second anchor sample can include the meaning expressed by this sentence, such as "people in modern times improve themselves to make a living," and the detailed features of this second anchor sample can include the specific modifying content of any one of the subject, predicate, object, or adverbial of the sentence, for example, the detailed features can include the modifying content of "era" "rapidly developing," the modifying content of "improve" "work harder and harder," etc., without specific limitations in this application. Furthermore, the first feature information corresponding to the second anchor sample can include the above-mentioned overall features and detailed features of the second anchor sample.

[0146] When the second anchor sample in the target data pair is a text sample, the second collocation sample in the target data pair is also a text sample. The second feature information corresponding to the second collocation sample can include the overall feature and detailed feature of the second collocation sample. The overall feature of the second collocation sample can include the overall meaning of the sentence, and the detailed feature of the second collocation sample can include the specific modification content of any one of the subject, predicate, object, or adverbial of the sentence, etc., which are not specifically limited in this application. For example, the second collocation sample can be the domain negative sample corresponding to the second anchor sample: "In this modern, slowly developing era, in order not to be submerged by the torrent of time, people can only eat more and improve themselves all day long to make a living." Here, the overall feature of this second collocation sample can include the meaning expressed by the sentence, such as "people in modern times improve themselves to make a living," and the detailed feature can include the specific modification content of any one of the subject, predicate, object, or adverbial of the sentence. For example, the detailed feature can include the modification content of "era" as "slowly developing," the modification content of "improve" as "eat more and improve themselves all day long," etc., which are not specifically limited here. Furthermore, the second feature information corresponding to the second pairing sample may include the aforementioned overall features and the aforementioned detailed features of the second pairing sample.

[0147] In the case where the second anchor sample in the target data pair is an image, the first feature information corresponding to the second anchor sample can include the overall features and detailed features of the second anchor sample. The overall features of the second anchor sample can include the image category, the image theme, etc., while the detailed features can include at least one of the following: image style, the proportion of objects in the image, the size of objects in the image, the shape of objects in the image, the color of objects in the image, the color of the background area in the image, etc., without specific limitations. For example, if the second anchor sample is a photo of a group dining together, the overall features of the second anchor sample can include the image theme "food," and the detailed features can include the type of food in the image, the quantity of food, the shape of the table, the number of people, etc., without specific limitations. The first feature information corresponding to the second anchor sample can include the aforementioned overall features and detailed features of the second anchor sample.

[0148] When the second anchor sample in the target data pair is an image, the second paired sample in the target data pair is also an image. The second feature information corresponding to the second paired sample can include the overall features and detailed features of the second paired sample. The overall features of the second paired sample can include at least one of the following: image category, image theme, etc. The detailed features of the second anchor sample can include at least one of the following: image style, the proportion of objects in the image, the size of objects in the image, the shape of objects in the image, the color of objects in the image, the color of the background in the image, etc., without specific limitations. For example, if the second paired sample is an out-of-domain negative sample corresponding to the second anchor sample, such as a photo of a basketball game, the overall features of the second paired sample can include the image theme "sports," and the detailed features of the second paired sample can include the number of people in the image, the background in the image, the position of the basketball in the image, the jersey number in the image, etc., without specific limitations. The second feature information corresponding to the second paired sample can include the above-mentioned overall features and detailed features of the second paired sample.

[0149] Step 4704: Based on the contrast loss of each data pair in the N data pairs, adjust the parameters of the first target feature extraction model to obtain the second target feature extraction model.

[0150] In this embodiment, the parameters of the first target feature extraction model can be adjusted N times based on the contrast loss of each of the N data pairs until the model converges, thereby obtaining the second target feature extraction model and improving the accuracy of model training.

[0151] In this embodiment, the first target feature extraction model can also be adjusted N / M times based on the mean of the contrast loss of M data pairs out of N data pairs, until the model converges, thus obtaining the second target feature extraction model, which improves the training efficiency of the model. Here, N is a positive integer greater than or equal to 3, M is a positive integer greater than or equal to 2, and N / M is a positive integer.

[0152] Specifically, when adjusting the parameters of the first target feature extraction model based on the mean of the contrast loss for every M data pairs, the second anchor sample of every two data pairs in the M data pairs can be different; or, if two data pairs in the M data pairs have the same second anchor sample, the second pairing samples in these two data pairs are different. This allows the first target feature extraction model to be trained with N data pairs randomly shuffled, improving the model's robustness against interference.

[0153] According to the training method of the feature extraction model provided in the embodiments of this application, a first target feature extraction model is obtained by acquiring a first anchor sample, a positive sample within the domain corresponding to the first anchor sample, and a negative sample outside the domain corresponding to the first anchor sample; the first anchor sample, the positive sample within the domain corresponding to the first anchor sample, and the negative sample outside the domain corresponding to the first anchor sample are used to perform comparative learning training on the feature extraction model to be trained, thereby obtaining a first target feature extraction model; a second anchor sample is obtained from the second stage dataset; the second anchor sample is augmented using positive example data augmentation to obtain a positive sample within the domain corresponding to the second anchor sample; the second anchor sample is augmented using negative example data augmentation to obtain a negative sample within the domain corresponding to the second anchor sample; samples in the second stage dataset other than the second anchor sample are used as negative samples outside the domain corresponding to the second anchor sample; the second anchor sample, the positive sample within the domain corresponding to the second anchor sample, the negative sample within the domain corresponding to the second anchor sample, and the negative sample outside the domain corresponding to the second anchor sample are used as training samples in the second stage to perform comparative learning training on the first target feature extraction model, thereby obtaining a second target feature extraction model. Thus, the first target feature extraction model, obtained through inter-domain contrastive learning training in the first stage, has the ability to distinguish the overall features of the data. Based on this, the second target feature extraction model, obtained through intra-domain contrastive learning training in the second stage, not only has the ability to distinguish the overall features of the data, but also has the ability to further distinguish the detailed features of the data, which greatly expands the application scope of the model. This solves the problem in related technologies where feature extraction models are difficult to extract detailed feature information of objects, resulting in a limited application scope.

[0154] It should be noted that, since the order of the training samples in the two stages can be interchanged, in another embodiment provided in this application, the first-stage training samples can be training samples used for intra-domain contrastive learning, and the second-stage training samples can be training samples used for inter-domain contrastive learning. In this case, in the training method of the feature extraction model provided in this application embodiment:

[0155] In step 310 above, obtaining the first-stage training samples may include:

[0156] Obtain the second anchor point sample from the first stage dataset;

[0157] The second anchor sample is augmented using positive example data augmentation to obtain positive samples within the domain corresponding to the second anchor sample.

[0158] The second anchor sample is augmented using negative example data augmentation to obtain a domain-specific negative sample corresponding to the second anchor sample.

[0159] The specific details of step 310 can be found in steps 430 to 450, and will not be repeated here.

[0160] In step 320 above, the first target feature extraction model is obtained by comparative learning training using the first-stage training samples to the feature extraction model to be trained, which may include:

[0161] The second anchor sample, the positive sample in the domain corresponding to the second anchor sample, and the negative sample in the domain corresponding to the second anchor sample are used as the first-stage training samples. The feature extraction model to be trained is compared and trained to obtain the first target feature extraction model.

[0162] Specifically, the second anchor point sample, the positive sample in the domain corresponding to the second anchor point sample, and the negative sample in the domain corresponding to the second anchor point sample are used as the first-stage training samples. These samples are then used to perform comparative learning training on the feature extraction model to be trained, resulting in the first target feature extraction model, which may include:

[0163] N data pairs are obtained from the first-stage training samples. Each data pair includes a second anchor sample and a second pairing sample corresponding to the second anchor sample. The second pairing sample is either a positive sample in the domain corresponding to the second anchor sample or a negative sample in the domain corresponding to the second anchor sample. N is an integer greater than or equal to 2.

[0164] Input the N data pairs into the feature extraction model to be trained;

[0165] Among them, the N data pairs include at least one of the first type data pairs and the second type data pairs; the first type data pair includes the second anchor sample and the positive sample in the domain corresponding to the second anchor sample; the second type data pair includes the second anchor sample and the negative sample in the domain corresponding to the second anchor sample.

[0166] For a target data pair among N data pairs: determine the target similarity between the second anchor sample and the second paired sample in the target data pair using the feature extraction model to be trained; determine the contrast loss of the target data pair based on the target similarity and the reference similarity of the target data pair.

[0167] The feature extraction model to be trained may include a feature extraction layer for performing feature extraction processing on the input data and a similarity calculation layer for performing similarity calculation processing.

[0168] Furthermore, for one target data pair among the N data pairs: a feature extraction layer is used to extract features of the target data pair to obtain the first feature information corresponding to the second anchor sample and the second feature information corresponding to the second pairing sample; a similarity calculation layer is used to calculate the similarity between the first feature information and the second feature information, which is used as the target similarity between the second anchor sample and the second pairing sample.

[0169] Specifically, when the second matching sample in the target data pair is a positive sample within the domain corresponding to the second anchor sample, the reference similarity of the target data pair is 1; when the second matching sample in the target data pair is a negative sample within the domain corresponding to the second anchor sample, the reference similarity of the target data pair is 0.

[0170] Based on the contrast loss of each of the N data pairs, the parameters of the feature extraction model to be trained are adjusted to obtain the first target feature extraction model.

[0171] The specific adjustment process in step 320 can be referred to in step 470, and will not be repeated here.

[0172] In step 330 above, obtaining the second-stage training samples may include:

[0173] Obtain the first anchor sample, the positive sample within the domain corresponding to the first anchor sample, the negative sample outside the domain corresponding to the first anchor sample, and the negative sample within the domain corresponding to the first anchor sample.

[0174] Wherein, the number of negative samples outside the domain corresponding to the first anchor sample is greater than or equal to the number of negative samples within the domain corresponding to the first anchor sample.

[0175] The specific details of obtaining the first anchor point sample, the positive sample within the domain corresponding to the first anchor point sample, and the negative sample outside the domain corresponding to the first anchor point sample can be found in step 410, and will not be repeated here.

[0176] Obtaining the domain-specific negative sample corresponding to the first anchor sample may include: performing data augmentation processing on the first anchor sample using negative example data augmentation methods to obtain the domain-specific negative sample corresponding to the first anchor sample. Where the first anchor sample is a text sample, the negative example data augmentation method may be at least one of adding negation words, replacing antonyms, etc. Where the first anchor sample is an image sample, the negative example data augmentation method may be at least one of image rotation, blurring, randomly adjusting image size, style transfer, etc. Specific details can be found similarly in step 450, and will not be repeated here.

[0177] In step 340 above, the first target feature extraction model is trained through comparative learning using the second-stage training samples to obtain the second target feature extraction model, which may include:

[0178] The first anchor sample, the positive sample within the domain corresponding to the first anchor sample, the negative sample outside the domain corresponding to the first anchor sample, and the negative sample within the domain corresponding to the first anchor sample are used as training samples for the second stage to perform comparative learning training on the first target feature extraction model, thereby obtaining the second target feature extraction model.

[0179] Specifically, the first anchor sample, the positive in-domain sample corresponding to the first anchor sample, the negative out-of-domain sample corresponding to the first anchor sample, and the negative in-domain sample corresponding to the first anchor sample are used as training samples for the second stage to perform comparative learning training on the first target feature extraction model, thereby obtaining the second target feature extraction model, which may include:

[0180] F data pairs are obtained from the first-stage training samples. Each data pair includes a first anchor sample and a first pairing sample corresponding to the first anchor sample. The first pairing sample is either a positive sample within the domain corresponding to the first anchor sample, a negative sample outside the domain corresponding to the first anchor sample, or a negative sample within the domain corresponding to the first anchor sample. Here, F is an integer greater than or equal to 3.

[0181] Input the F data pairs into the first target feature extraction model;

[0182] Among them, the F data pairs include at least one of the fourth type data pair, the fifth type data pair, and the sixth type data pair; the fourth type data pair includes the first anchor sample and the positive sample within the domain corresponding to the first anchor sample; the fifth type data pair includes the first anchor sample and the negative sample outside the domain corresponding to the first anchor sample; the sixth type data pair includes the first anchor sample and the negative sample within the domain corresponding to the first anchor sample.

[0183] For a target data pair among F data pairs: the target similarity between the first anchor sample and the first paired sample in the target data pair is determined by the first target feature extraction model; based on the target similarity and the reference similarity of the target data pair, the contrast loss of the target data pair is determined.

[0184] The first target feature extraction model may include a feature extraction layer for performing feature extraction processing on the input data and a similarity calculation layer for performing similarity calculation processing.

[0185] Furthermore, for one target data pair among the F data pairs: a feature extraction layer is used to extract features of the target data pair to obtain first feature information corresponding to the first anchor sample and second feature information corresponding to the first pairing sample; a similarity calculation layer is used to calculate the similarity between the first feature information and the second feature information, which is used as the target similarity between the first anchor sample and the first pairing sample.

[0186] Specifically, the reference similarity of the target data pair is 1 when the first paired sample in the target data pair is a positive sample within the domain corresponding to the first anchor sample; the reference similarity of the target data pair is 0 when the first paired sample in the target data pair is a negative sample within the domain corresponding to the first anchor sample; and the reference similarity of the target data pair is 0 when the first paired sample in the target data pair is a negative sample outside the domain corresponding to the first anchor sample.

[0187] Based on the contrast loss of each of the N data pairs, the parameters of the first target feature extraction model are adjusted to obtain the second target feature extraction model.

[0188] According to the training method of the feature extraction model provided in the embodiments of this application, a second anchor point sample, a positive sample in the domain corresponding to the second anchor point sample, and a negative sample in the domain corresponding to the second anchor point sample are obtained; the second anchor point sample, the positive sample in the domain corresponding to the second anchor point sample, and the negative sample in the domain corresponding to the second anchor point sample are used as first-stage training samples to perform comparative learning training on the feature extraction model to be trained, thereby obtaining a first target feature extraction model; a first anchor point sample is obtained from the second-stage dataset; and the first anchor point sample is subjected to data augmentation processing through positive example data augmentation to obtain a first target feature extraction model. The first target feature extraction model is obtained by taking positive samples within the domain corresponding to the first anchor sample, and negative samples outside the domain corresponding to the first anchor sample from the second-stage dataset. The first anchor sample is then augmented using negative example data augmentation to obtain negative samples within the domain corresponding to the first anchor sample. The first anchor sample, the positive samples within the domain corresponding to the first anchor sample, the negative samples outside the domain corresponding to the first anchor sample, and the negative samples within the domain corresponding to the first anchor sample are used as training samples for the second stage. This allows for comparative learning training of the first target feature extraction model, resulting in a second target feature extraction model. Thus, the first target feature extraction model, trained through intra-domain comparative learning in the first stage, has the ability to distinguish detailed features of the data. Based on this, the second target feature extraction model, trained through inter-domain comparative learning in the second stage, not only has the ability to distinguish detailed features of the data but also the ability to further distinguish overall features of the data, greatly expanding the model's application scope and thus solving the problem of limited application scope of feature extraction methods in related technologies.

[0189] It is understood that the second target feature extraction model is a pre-trained feature extraction model. This application embodiment also provides a method for extracting features from input objects such as images or text using the second target feature extraction model.

[0190] Figure 5 This is a schematic flowchart illustrating a feature extraction method provided in an embodiment of this application.

[0191] like Figure 5 As shown, the feature extraction method provided in this application embodiment may include:

[0192] Step 510: Obtain target data;

[0193] The target data can be in the form of text, images, or other similar data.

[0194] Step 520: Input the target data into the second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data;

[0195] The second target feature extraction model is trained using any of the training methods provided in the above embodiments.

[0196] The feature information corresponding to the target data may include overall features and detailed features corresponding to the target data.

[0197] In cases where the target data is in textual form, the overall characteristics of the target data can include the overall meaning of the sentence, while the detailed characteristics can include the specific modifications of any one of the sentence's subject, predicate, object, or adverbial phrase, etc., without specific limitations. For example, the target data is "In this rapidly developing modern era, in order not to be overwhelmed by the torrent of time, people can only work harder and harder to improve themselves and make a living." The overall characteristics of this target data can include the meaning expressed by the sentence, such as "people in modern times improve themselves to make a living," while the detailed characteristics can include the specific modifications of any one of the sentence's subject, predicate, object, or adverbial phrase, such as the modification of "the era" ("rapidly developing"), the modification of "improve" ("working harder and harder"), etc., without specific limitations.

[0198] When the target data is in the form of an image, the overall features of the target data may include at least one of the following: image category, image theme, etc. The detailed features of the target data may include at least one of the following: image style, the proportion of objects in the image, the size of objects in the image, the shape of objects in the image, the color of objects in the image, the color of the background area of ​​the image, etc. This application does not impose specific limitations. For example, if the target data is a photograph of a group of people dining together, the overall features may include the image's theme "food," and the detailed features may include the type of food in the image, the quantity of food, the shape of the table, the number of people, etc. This application does not impose specific limitations.

[0199] It is understandable that, since the second target feature extraction model not only has the ability to distinguish the overall features of the data, but also has the ability to further distinguish the detailed features of the data, the feature information corresponding to the target data extracted by the second target feature extraction model can accurately reflect the overall features and many detailed features of the target data, so that the feature information corresponding to the target data can be applied to various subsequent tasks (such as classification tasks, recognition tasks, transformation tasks, etc.).

[0200] According to the feature extraction method provided in this application embodiment, target data is acquired; the target data is then input into a second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data; wherein, the second target feature extraction model is trained according to any of the training methods provided in the above embodiments. Thus, the feature information corresponding to the target data extracted by the second target feature extraction model can accurately reflect the overall features and many detailed features of the target data, enabling the feature information corresponding to the target data to be applied to various subsequent tasks, greatly expanding the application scope of the model.

[0201] The training method for a feature extraction model based on contrastive learning provided in this application can be executed by a training device for the feature extraction model. This application uses an example of a training device for the feature extraction model executing the training method to illustrate the training device for the feature extraction model provided in this application.

[0202] This application provides a training apparatus for a feature extraction model, which may include: a sample acquisition module and a contrastive learning training module;

[0203] The sample acquisition module is used to acquire the first-stage training samples;

[0204] The contrastive learning training module is used to perform contrastive learning training on the feature extraction model to be trained using the first stage training samples to obtain the first target feature extraction model.

[0205] The sample acquisition module is also used to acquire training samples for the second stage;

[0206] The contrastive learning training module is further used to perform contrastive learning training on the first target feature extraction model using the second-stage training samples to obtain a second target feature extraction model.

[0207] The first-stage training samples are one of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning, and the second-stage training samples are the other of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning.

[0208] The training samples used for inter-domain contrastive learning include: a first anchor sample, a positive sample within the domain corresponding to the first anchor sample, and a negative sample outside the domain corresponding to the first anchor sample; the training samples used for intra-domain contrastive learning include: a second anchor sample, a positive sample within the domain corresponding to the second anchor sample, and a negative sample within the domain corresponding to the second anchor sample.

[0209] The training apparatus for the feature extraction model provided in the embodiments of this application may include a sample acquisition module and a contrastive learning training module. The sample acquisition module is used to acquire first-stage training samples. The contrastive learning training module is used to perform contrastive learning training on the feature extraction model to be trained using the first-stage training samples to obtain a first target feature extraction model. The sample acquisition module is also used to acquire second-stage training samples. The contrastive learning training module is also used to perform contrastive learning training on the first target feature extraction model using the second-stage training samples to obtain a second target feature extraction model. The first-stage training samples are one of training samples for inter-domain contrastive learning and training samples for intra-domain contrastive learning, and the second-stage training samples are the other of training samples for inter-domain contrastive learning and training samples for intra-domain contrastive learning. The training samples for inter-domain contrastive learning include: a first anchor sample, an intra-domain positive sample corresponding to the first anchor sample, and an extra-domain negative sample corresponding to the first anchor sample. The training samples for intra-domain contrastive learning include: a second anchor sample, an intra-domain positive sample corresponding to the second anchor sample, and an intra-domain negative sample corresponding to the second anchor sample. In this way, the model is trained in two stages using first-stage training samples and second-stage training samples. One stage focuses on inter-domain comparative learning, and the other stage focuses on intra-domain comparative learning. The resulting second target feature extraction model not only has the ability to distinguish the overall features of the data, but also has the ability to further distinguish the detailed features of the data. This greatly expands the application scope of the model and solves the problem in related technologies where feature extraction models are difficult to extract detailed feature information of objects, resulting in a limited application scope.

[0210] Optionally, in the training apparatus for the feature extraction model provided in the embodiments of this application,

[0211] In the case where the training samples in the second stage are training samples for intra-domain contrastive learning, the training samples for intra-domain contrastive learning further include: out-of-domain negative samples corresponding to the second anchor sample.

[0212] Alternatively, if the training samples in the second stage are training samples for inter-domain contrastive learning, the training samples for inter-domain contrastive learning may further include: intra-domain negative samples corresponding to the first anchor sample.

[0213] In this way, when the training samples in the second stage are used for intra-domain contrastive learning, by adding a small number of out-of-domain negative samples to the training samples in the second stage, the model is prevented from focusing too much on details during the second stage training process, which would affect the model's ability to distinguish overall features. This allows the second target feature extraction model to maintain the ability to distinguish the overall features of the data obtained in the first stage training, while also having the ability to distinguish the detailed features of the data.

[0214] Alternatively, if the training samples in the second stage are used for inter-domain contrastive learning, by adding a small number of intra-domain negative samples to the training samples in the second stage, the model can be prevented from focusing too much on the overall features during the second stage training process, which would affect the model's ability to distinguish detailed features. This allows the second target feature extraction model to maintain the ability to distinguish the detailed features of the data obtained in the first stage training, while also having the ability to distinguish the overall features of the data.

[0215] Optionally, in the training apparatus for the feature extraction model provided in the embodiments of this application,

[0216] When the training samples in the second stage are training samples used for intra-domain contrastive learning, the number of intra-domain negative samples corresponding to the second anchor sample is greater than or equal to the number of out-of-domain negative samples corresponding to the second anchor sample.

[0217] When the training samples in the second stage are training samples used for inter-domain contrast learning, the number of out-of-domain negative samples corresponding to the first anchor sample is greater than or equal to the number of in-domain negative samples corresponding to the first anchor sample.

[0218] Thus, when the training samples in the second stage are used for intra-domain contrastive learning, a small number of out-of-domain negative samples can be used for contrastive learning training during the second stage training process to maintain the overall ability of distinguishing data obtained in the first stage training. At the same time, a sufficient number of intra-domain negative samples can be used for contrastive learning training, so that the second target feature extraction model also has the ability to distinguish the detailed features of the data.

[0219] Alternatively, if the training samples in the second stage are used for inter-domain contrastive learning, a small number of intra-domain negative samples can be used for contrastive learning training during the second stage training process to maintain the ability to distinguish the detailed features of the data obtained in the first stage training. At the same time, a sufficient number of out-of-domain negative samples can be used for contrastive learning training to enable the second target feature extraction model to have the ability to distinguish the overall features of the data.

[0220] Optionally, in the training apparatus for the feature extraction model provided in the embodiments of this application,

[0221] The first-stage training samples are used for inter-domain contrastive learning, and the second-stage training samples are used for intra-domain contrastive learning.

[0222] Regarding the acquisition of training samples for the second stage, the sample acquisition module is specifically used for:

[0223] Obtain the second anchor point sample from the second phase dataset;

[0224] The second anchor sample is augmented using positive example data augmentation to obtain positive samples within the domain corresponding to the second anchor sample.

[0225] The second anchor sample is augmented using negative example data augmentation to obtain a domain-specific negative sample corresponding to the second anchor sample.

[0226] The samples in the second-stage dataset other than the second anchor sample are taken as the out-of-domain negative samples corresponding to the second anchor sample.

[0227] In this way, we can obtain the second anchor sample, the positive sample within the domain corresponding to the second anchor sample, the negative sample within the domain corresponding to the second anchor sample, and the negative sample outside the domain corresponding to the second anchor sample as training samples for the second stage. By comparing and learning with the negative sample outside the domain corresponding to the second anchor sample, we can prevent the model from focusing too much on details during the second stage training process, which would affect the model's ability to distinguish overall features. This allows the second target feature extraction model to maintain the ability to distinguish the overall features of the data obtained in the first stage training, while also having the ability to distinguish the detailed features of the data.

[0228] Optionally, in the training apparatus for the feature extraction model provided in the embodiments of this application,

[0229] The second anchor sample in the second-stage training samples is a text sample, and the negative example data augmentation method includes antonym replacement; in terms of obtaining the domain-specific negative samples corresponding to the second anchor sample, the sample acquisition module is specifically used for:

[0230] The second anchor point sample in text form is split into word groups to obtain P independent word groups;

[0231] By replacing Q independent word groups out of the P independent word groups with their corresponding antonyms, the domain negative sample corresponding to the second anchor point sample is obtained.

[0232] Where Q and P are positive integers, the quotient of Q and P lies between the first threshold and the second threshold, the first threshold is greater than 0, and the second threshold is less than 1.

[0233] In this way, generating negative samples within the domain by replacing antonyms can change the original semantics of the second anchor sample in text form without significantly altering the overall features of the original text data. This allows the model to not only focus on sample features but also gain a deeper understanding of text semantics, especially the different semantics of the replaced antonyms, thereby improving the model's semantic understanding ability.

[0234] Optionally, in the training apparatus for the feature extraction model provided in the embodiments of this application,

[0235] The second anchor sample in the second-stage training samples is an image sample, and the negative example data augmentation method includes at least one of color adjustment, content replacement, and size adjustment; in terms of obtaining the domain-specific negative samples corresponding to the second anchor sample, the sample acquisition module is specifically used for:

[0236] Determine the target object in the second anchor point sample in image form;

[0237] When the negative sample data augmentation method includes color adjustment, the negative sample in the domain corresponding to the second anchor point sample is obtained by adjusting the color of the target object to a specified color.

[0238] When the negative sample data augmentation method includes content replacement, the domain negative sample corresponding to the second anchor sample is obtained by replacing the target object with a specified object.

[0239] When the negative sample data augmentation method includes size adjustment, the domain negative sample corresponding to the second anchor point sample is obtained by adjusting the size of the target object to a specified size.

[0240] In this way, various negative sample data augmentation methods can be used to perform data augmentation processing on the second anchor point sample in image form, resulting in multiple domain negative samples corresponding to the second anchor point sample, thereby improving the anti-interference ability of the feature extraction model.

[0241] Optionally, in the training apparatus for the feature extraction model provided in the embodiments of this application,

[0242] In obtaining a second target feature extraction model by comparative learning training the first target feature extraction model using the second-stage training samples, the comparative learning training module is specifically used for:

[0243] N data pairs are obtained from the second stage training samples. Each data pair includes a second anchor sample and a second pairing sample corresponding to the second anchor sample. The second pairing sample is either a positive sample within the domain corresponding to the second anchor sample, a negative sample outside the domain corresponding to the second anchor sample, or a negative sample within the domain corresponding to the second anchor sample; where N is an integer greater than or equal to 2.

[0244] Input the N data pairs into the first target feature extraction model;

[0245] For a target data pair among the N data pairs: the target similarity between the second anchor sample and the second paired sample in the target data pair is determined by the first target feature extraction model; based on the target similarity and the reference similarity of the target data pair, the contrast loss of the target data pair is determined.

[0246] Based on the contrast loss of each of the N data pairs, the parameters of the first target feature extraction model are adjusted to obtain the second target feature extraction model.

[0247] Thus, based on the ability of the first target feature extraction model to distinguish the overall features of data obtained through inter-domain contrastive learning training in the first stage, the second target feature extraction model obtained through intra-domain contrastive learning training in the second stage not only has the ability to distinguish the overall features of data, but also has the ability to further distinguish the detailed features of data, which greatly expands the application scope of the model. This solves the problem in related technologies where it is difficult to extract detailed feature information of objects using feature extraction models, resulting in a relatively limited application scope.

[0248] The training device for the feature extraction model in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0249] The training device for the feature extraction model in this embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this embodiment does not specifically limit the specific operating system used.

[0250] The training device for the feature extraction model provided in this application embodiment can achieve... Figures 1 to 4-3 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0251] Figure 6 This is a schematic structural diagram of a feature extraction device provided in an embodiment of this application.

[0252] like Figure 6 As shown, this application embodiment also provides a feature extraction device 600, which may include: an acquisition module 601 and a processing module 602;

[0253] Module 601 is used to acquire target data;

[0254] The target data can be in the form of text, images, or other similar data.

[0255] Processing module 602 is used to input the target data into the second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data;

[0256] The second target feature extraction model is trained using any of the training methods provided in the embodiments of this application.

[0257] The feature extraction apparatus provided according to the embodiments of this application includes an acquisition module and a processing module. The acquisition module is used to acquire target data; the processing module is used to input the target data into a second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data. The second target feature extraction model is trained using any of the training methods provided in the embodiments of this application. Thus, the feature information corresponding to the target data extracted using the second target feature extraction model can accurately reflect the overall features and many detailed features of the target data, enabling the feature information corresponding to the target data to be applied to various subsequent tasks, greatly expanding the application scope of the model.

[0258] The feature extraction device provided in this application embodiment can achieve... Figure 5 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0259] Optionally, such as Figure 7 As shown, this application embodiment also provides an electronic device 700, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they implement the various steps of the above method embodiments and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0260] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0261] Figure 8 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0262] The electronic device 800 includes, but is not limited to, components such as: radio frequency unit 801, network module 802, audio output unit 803, input unit 804, sensor 805, display unit 806, user input unit 807, interface unit 808, memory 809, and processor 810.

[0263] Those skilled in the art will understand that the electronic device 800 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 810 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0264] The input unit 804 is used to acquire the first-stage training samples;

[0265] The processor 810 is used to perform comparative learning training on the feature extraction model to be trained using the first stage training samples to obtain the first target feature extraction model.

[0266] The input unit 804 is also used to acquire the second-stage training samples;

[0267] The processor 810 is further configured to perform comparative learning training on the first target feature extraction model using the second-stage training samples to obtain a second target feature extraction model;

[0268] The first-stage training samples are one of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning, and the second-stage training samples are the other of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning.

[0269] The training samples used for inter-domain contrastive learning include: a first anchor sample, a positive sample within the domain corresponding to the first anchor sample, and a negative sample outside the domain corresponding to the first anchor sample; the training samples used for intra-domain contrastive learning include: a second anchor sample, a positive sample within the domain corresponding to the second anchor sample, and a negative sample within the domain corresponding to the second anchor sample.

[0270] The electronic device provided according to embodiments of this application may include an input unit and a processor; the input unit is configured to acquire first-stage training samples; the processor is configured to perform comparative learning training on a feature extraction model to be trained using the first-stage training samples to obtain a first target feature extraction model; the input unit is further configured to acquire second-stage training samples; the processor is further configured to perform comparative learning training on the first target feature extraction model using the second-stage training samples to obtain a second target feature extraction model; wherein, the first-stage training samples are one of training samples for inter-domain comparative learning and training samples for intra-domain comparative learning, and the second-stage training samples are the other of training samples for inter-domain comparative learning and training samples for intra-domain comparative learning; wherein, the training samples for inter-domain comparative learning include: a first anchor sample, an intra-domain positive sample corresponding to the first anchor sample, and an extra-domain negative sample corresponding to the first anchor sample; the training samples for intra-domain comparative learning include: a second anchor sample, an intra-domain positive sample corresponding to the second anchor sample, and an intra-domain negative sample corresponding to the second anchor sample. In this way, the model is trained in two stages using first-stage training samples and second-stage training samples. One stage focuses on inter-domain comparative learning, and the other stage focuses on intra-domain comparative learning. The resulting second target feature extraction model not only has the ability to distinguish the overall features of the data, but also has the ability to further distinguish the detailed features of the data. This greatly expands the application scope of the model and solves the problem in related technologies where feature extraction models are difficult to extract detailed feature information of objects, resulting in a limited application scope.

[0271] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0272] It should be understood that, in this embodiment, the input unit 804 may include a graphics processing unit (GPU) 8041 and a microphone 8042. The GPU 8041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 806 may include a display panel 8061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 807 includes at least one of a touch panel 8071 and other input devices 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 may include a touch detection device and a touch controller. Other input devices 8072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0273] The memory 809 can be used to store software programs and various data. The memory 809 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 809 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 809 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0274] Processor 810 may include one or more processing units; optionally, processor 810 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 810.

[0275] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0276] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0277] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0278] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0279] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.

[0280] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0281] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0282] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A training method for a feature extraction model, characterized in that, include: Obtain the first phase of training samples; The first target feature extraction model is obtained by comparing and learning the feature extraction model to be trained using the first stage training samples. Obtain the second phase training samples; The first target feature extraction model is trained by comparative learning using the second stage training samples to obtain the second target feature extraction model. The first-stage training samples are one of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning, and the second-stage training samples are the other of training samples used for inter-domain contrast learning and training samples used for intra-domain contrast learning. The training samples used for inter-domain contrastive learning include: a first anchor sample, a positive intra-domain sample corresponding to the first anchor sample, and a negative out-of-domain sample corresponding to the first anchor sample; the training samples used for intra-domain contrastive learning include: a second anchor sample, a positive intra-domain sample corresponding to the second anchor sample, and a negative intra-domain sample corresponding to the second anchor sample; the first anchor sample and the second anchor sample include samples in text form or samples in image form.

2. The training method according to claim 1, characterized in that, In the case where the training samples in the second stage are training samples for intra-domain contrastive learning, the training samples for intra-domain contrastive learning further include: out-of-domain negative samples corresponding to the second anchor sample. Alternatively, if the training samples in the second stage are training samples for inter-domain contrastive learning, the training samples for inter-domain contrastive learning may further include: intra-domain negative samples corresponding to the first anchor sample.

3. The method according to claim 2, characterized in that, When the training samples in the second stage are training samples used for intra-domain contrastive learning, the number of intra-domain negative samples corresponding to the second anchor sample is greater than or equal to the number of out-of-domain negative samples corresponding to the second anchor sample. Alternatively, if the training samples in the second stage are training samples used for inter-domain contrast learning, the number of out-of-domain negative samples corresponding to the first anchor sample is greater than or equal to the number of in-domain negative samples corresponding to the first anchor sample.

4. The training method according to any one of claims 1-3, characterized in that, The first-stage training samples are used for inter-domain contrastive learning, and the second-stage training samples are used for intra-domain contrastive learning. The process of obtaining the second-stage training samples includes: Obtain the second anchor point sample from the second phase dataset; The second anchor sample is augmented using positive example data augmentation to obtain positive samples within the domain corresponding to the second anchor sample. The second anchor sample is augmented using negative example data augmentation to obtain a domain-specific negative sample corresponding to the second anchor sample. The samples in the second-stage dataset other than the second anchor sample are taken as the out-of-domain negative samples corresponding to the second anchor sample.

5. The method according to claim 4, characterized in that, The second anchor sample in the second stage training samples is a text sample, and the negative example data augmentation method includes antonym replacement; The step of performing data augmentation on the second anchor sample using negative example data augmentation to obtain in-domain negative samples corresponding to the second anchor sample includes: The second anchor point sample in text form is split into word groups to obtain P independent word groups; By replacing Q independent word groups out of the P independent word groups with their corresponding antonyms, the domain negative sample corresponding to the second anchor point sample is obtained. Where Q and P are positive integers, the quotient of Q and P lies between the first threshold and the second threshold, the first threshold is greater than 0, and the second threshold is less than 1.

6. The method according to claim 4, characterized in that, The second anchor sample in the second stage training samples is an image sample, and the negative example data augmentation method includes at least one of color adjustment, content replacement and size adjustment; The step of performing data augmentation on the second anchor sample using negative example data augmentation to obtain in-domain negative samples corresponding to the second anchor sample includes: Determine the target object in the second anchor point sample in image form; When the negative sample data augmentation method includes color adjustment, the negative sample in the domain corresponding to the second anchor point sample is obtained by adjusting the color of the target object to a specified color. When the negative sample data augmentation method includes content replacement, the domain negative sample corresponding to the second anchor sample is obtained by replacing the target object with a specified object. When the negative sample data augmentation method includes size adjustment, the domain negative sample corresponding to the second anchor point sample is obtained by adjusting the size of the target object to a specified size.

7. The method according to claim 4, characterized in that, The step of comparing and training the first target feature extraction model with the second-stage training samples to obtain the second target feature extraction model includes: N data pairs are obtained from the second stage training samples. Each data pair includes a second anchor sample and a second pairing sample corresponding to the second anchor sample. The second pairing sample is either a positive sample within the domain corresponding to the second anchor sample, a negative sample outside the domain corresponding to the second anchor sample, or a negative sample within the domain corresponding to the second anchor sample; where N is an integer greater than or equal to 2. Input the N data pairs into the first target feature extraction model; For a target data pair among the N data pairs: the target similarity between the second anchor sample and the second paired sample in the target data pair is determined by the first target feature extraction model; based on the target similarity and the reference similarity of the target data pair, the contrast loss of the target data pair is determined; Based on the contrast loss of each of the N data pairs, the parameters of the first target feature extraction model are adjusted to obtain the second target feature extraction model.

8. The method according to claim 7, characterized in that, N is an integer greater than or equal to 3; the N data pairs include at least one of the first type of data pairs, the second type of data pairs, and the third type of data pairs; The first type of data pair includes a second anchor sample and a positive sample within the domain corresponding to the second anchor sample; the second type of data pair includes a second anchor sample and a negative sample outside the domain corresponding to the second anchor sample; the third type of data pair includes a second anchor sample and a negative sample within the domain corresponding to the second anchor sample. When the second matching sample in the target data pair is a positive sample within the domain corresponding to the second anchor sample, the reference similarity of the target data pair is 1; If the second matching sample in the target data pair is an out-of-domain negative sample corresponding to the second anchor sample, the reference similarity of the target data pair is 0. If the second matching sample in the target data pair is a negative sample within the domain corresponding to the second anchor sample, the reference similarity of the target data pair is 0.

9. The method according to claim 7, characterized in that, The first target feature extraction model includes: a feature extraction layer for performing feature extraction processing on the input data and a similarity calculation layer for performing similarity calculation processing; For one of the N data pairs: The feature extraction layer is used to extract features from the target data pair to obtain the first feature information corresponding to the second anchor sample and the second feature information corresponding to the pairing sample. The similarity calculation layer is used to calculate the similarity between the first feature information and the second feature information, which serves as the target similarity between the second anchor sample and the paired sample.

10. A feature extraction method, characterized in that, include: Obtain the target data; The target data is input into the second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data; The second target feature extraction model is trained using the training method according to any one of claims 1-9.

11. A feature extraction device, characterized in that, include: Acquisition module and processing module; The acquisition module is used to acquire target data; The processing module is used to input the target data into the second target feature extraction model for feature extraction processing to obtain feature information corresponding to the target data; The second target feature extraction model is obtained by the method according to any one of claims 1-9.

12. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1-10.

13. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-10.