A machine learning model construction method, device, electronic device and storage medium
By clustering and feature enhancement of the original samples, constructing sample pairs and feature extraction and fusion, the problem that traditional unsupervised single-source domain generalization technology is difficult to adapt to multiple target domains, and the applicability and accuracy of the model are improved.
Patent Information
- Application Number
- CN202510053970.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Traditional unsupervised single-source domain generalization technology is difficult to adapt to multiple unknown target domains, resulting in a significant decline in model performance.
By clustering the original samples, cluster enhancement samples are generated, sample pairs are constructed, and feature extraction and fusion are used for convolutional neural networks, and finally, machine learning models are built based on positive and negative sample sets.
Improve the applicability and accuracy of machine learning models, and better learn similarities and differences between different cluster features, and overcome false correlations between domains.
Smart Images

Figure CN119476407B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device and storage medium for constructing a machine learning model. Background Art
[0002] In recent years, machine learning, as a particularly critical branch in the field of artificial intelligence, is committed to enabling models and algorithms to learn from data. Unsupervised single-source domain generalization technology refers to training a model to generalize to an unknown target domain using only unlabeled data from a single source domain. Therefore, it has been widely studied and applied in machine learning scenarios.
[0003] However, when using traditional unsupervised single-source domain generalization technology for model training, only a limited amount of data is often available, and this data may only come from a specific domain. When the model needs to be applied to a new, unlabeled target domain, the domain shift problem makes it difficult for the learned model to adapt to multiple unknown target domains, and the performance of the model may drop significantly. Summary of the invention
[0004] The present application provides a method, device, electronic device and storage medium for constructing a machine learning model to improve the applicability and accuracy of the machine learning model.
[0005] In a first aspect, a method for building a machine learning model is provided, comprising:
[0006] Performing a clustering operation on a plurality of original samples contained in the original sample set to obtain N cluster samples; wherein N is an integer greater than 1;
[0007] Performing feature enhancement on the N cluster samples respectively to obtain cluster enhanced samples corresponding to the N cluster samples respectively, and forming N sample pairs based on the N cluster samples and the cluster enhanced samples corresponding to the N cluster samples respectively;
[0008] Using a convolutional neural network to extract features from the N sample pairs respectively; wherein the convolutional neural network includes a plurality of convolutional layers connected in sequence;
[0009] For each sample pair, the features output by each of the multiple convolutional layers are fused to obtain a fused feature of each sample pair;
[0010] A machine learning model that meets the requirements of the loss function is constructed based on a positive sample set, a negative sample set, and N fusion features; wherein the positive sample pairs contained in the positive sample set are constructed based on original samples and enhanced samples in the same cluster, and the negative sample pairs contained in the negative sample set are constructed based on original samples and enhanced samples in different clusters.
[0011] In an embodiment of the present application, first, a clustering operation is performed on multiple original samples included in the original sample set to obtain N cluster samples, and feature enhancement is performed on the N cluster samples to obtain cluster enhanced samples corresponding to each of the N cluster samples, and N sample pairs are respectively formed based on the N cluster samples and the cluster enhanced samples corresponding to each of the N cluster samples, so that respective enhancement strategies suitable for different clusters can be executed during the training process; then, a convolutional neural network is used to perform feature extraction on the N sample pairs, and for each sample pair, the features output by each of the multiple convolutional layers connected in sequence in the convolutional neural network are fused to obtain the fused features of each sample pair, which can better balance the global information and detail information of the image during the training process; finally, a machine learning model that meets the requirements of the loss function is constructed based on the positive sample set, the negative sample set, and the N fused features, so that the similarities and differences between different cluster features can be learned during the training process; the above-mentioned entire training process improves the applicability and accuracy of the machine learning model.
[0012] In some embodiments, the performing feature enhancement on the N cluster samples respectively to obtain cluster enhanced samples corresponding to the N cluster samples respectively includes:
[0013] A generative adversarial network is used to perform feature enhancement on the N cluster samples respectively to obtain enhanced samples corresponding to the N cluster samples respectively; wherein the generative adversarial network includes a generator and N cluster discriminators, one cluster discriminator corresponds to one cluster sample, and is used to discriminate the difference between the enhanced sample generated by the generator and the original sample.
[0014] In some embodiments, the adopting a generative adversarial network to perform feature enhancement on the N cluster samples respectively to obtain enhanced samples corresponding to the N cluster samples respectively includes:
[0015] Input any first original sample in the first cluster of samples into the generator to obtain a first enhanced sample of the first original sample; wherein the first cluster of samples is any one of the N cluster samples;
[0016] The first enhanced sample is input into a first cluster discriminator associated with the first cluster sample. If the first cluster discriminator determines that the probability that the first enhanced sample is the first original sample does not meet a probability threshold, the respective parameters of the first cluster discriminator and the generator are adjusted until the probability meets the probability threshold.
[0017] Through the above method, a generative adversarial network is used to perform adversarial training on the original samples in each cluster sample, which can generate enhanced samples that are suitable for the original samples in different cluster samples. Compared with the traditional unsupervised single-source domain generalization method, it takes into account the inconsistency of the enhancement strategies required for original samples of different categories, effectively analyzes and learns the differences between different cluster samples, thereby enhancing the effectiveness of different original samples and laying the foundation for subsequent feature extraction.
[0018] In some embodiments, the convolutional neural network further includes an attention mechanism layer; for each sample pair, the features output by each of the multiple convolutional layers are fused to obtain the fused features of each sample pair, including:
[0019] For each sample pair, the features output by each of the multiple convolutional layers are sequentially input into the attention mechanism layer to determine the weight parameters corresponding to each of the features;
[0020] The features are fused according to the weight parameters corresponding to the features to obtain the fused features of each sample pair.
[0021] In the above manner, for each sample pair, an attention mechanism layer is introduced to determine the weight parameters corresponding to the features output by different convolutional layers, and then the features are fused based on the weight parameters of each feature, making full use of the respective characteristics of shallow features and deep features, so that the model can obtain effective and comprehensive feature information, thereby improving the applicability of the model. Moreover, through the weight parameters of features in different layers, the model can pay more attention to domain-invariant features and ignore domain-related features, thereby overcoming the false correlation between domains.
[0022] In some embodiments, the loss function satisfies the following expression:
[0023]
[0024] Among them, the To characterize the fusion features of any one of the N sample pairs, Characterization and description The features of the i-th positive sample pair belonging to the same cluster, Characterization and description The features of the j-th negative sample pair belonging to different clusters, the τ represents a temperature parameter, which is used to adjust the sensitivity of the machine learning model to similarity.
[0025] In some embodiments, performing a clustering operation on the original samples included in the original sample set to obtain N cluster samples includes:
[0026] The K-Means++ clustering algorithm is used to perform a clustering operation on the original samples contained in the original sample set to obtain the N cluster samples.
[0027] In the above way, the K-Means++ clustering algorithm is used to ensure that the initial centroids have a more applicable distribution, avoid the centroids from gathering in a specific part of the data distribution, reduce the risk of converging to the local minimum, and is suitable for single-source domain generalization scenarios.
[0028] In a second aspect, a machine learning model construction device is provided, comprising:
[0029] A clustering module, used to perform a clustering operation on a plurality of original samples contained in the original sample set to obtain N cluster samples; wherein N is an integer greater than 1;
[0030] A feature enhancement module, configured to perform feature enhancement on the N cluster samples respectively, obtain cluster enhanced samples corresponding to the N cluster samples respectively, and form N sample pairs based on the N cluster samples and the cluster enhanced samples corresponding to the N cluster samples respectively;
[0031] A feature extraction module, used to extract features from the N sample pairs respectively using a convolutional neural network; wherein the convolutional neural network includes a plurality of convolutional layers connected in sequence;
[0032] A feature fusion module, used for fusing the features output by each of the multiple convolutional layers for each sample pair to obtain a fused feature of each sample pair;
[0033] A model construction module is used to construct a machine learning model that meets the requirements of the loss function based on a positive sample set, a negative sample set, and N fusion features; wherein the positive sample pairs contained in the positive sample set are constructed based on original samples and enhanced samples in the same cluster, and the negative sample pairs contained in the negative sample set are constructed based on original samples and enhanced samples in different clusters.
[0034] In some embodiments, the convolutional neural network further includes an attention mechanism layer; the feature fusion module is specifically used to:
[0035] For each sample pair, the features output by each of the multiple convolutional layers are sequentially input into the attention mechanism layer to determine the weight parameters corresponding to each of the features;
[0036] The features are fused according to the weight parameters corresponding to the features to obtain the fused features of each sample pair.
[0037] According to a third aspect, an electronic device is provided, including:
[0038] A memory for storing a computer program; a processor for implementing any one of the methods described in the first aspect when executing the computer program stored in the memory.
[0039] According to a fourth aspect, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of the first aspects is implemented.
[0040] For each aspect from the second to the fourth aspect and the technical effects that may be achieved by each aspect, please refer to the above description of the technical effects that can be achieved by the first aspect or various possible schemes in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A flowchart of a method for constructing a machine learning model provided in an embodiment of the present application;
[0042] Figure 2 A schematic diagram of the structure of a compression and activation network provided in an embodiment of the present application;
[0043] Figure 3 A logical schematic diagram of feature fusion provided in an embodiment of the present application;
[0044] Figure 4 A schematic diagram of a logical framework for training a machine learning model provided in an embodiment of the present application;
[0045] Figure 5 A schematic diagram of the structure of a machine learning model building device provided in an embodiment of the present application;
[0046] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily combined with each other. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.
[0048] The terms "first" and "second" in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of their variations are intended to cover non-exclusive protection. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in the present application can mean at least two, for example, two, three or more, and the embodiments of the present application are not limited.
[0049] The following is a description of exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, which should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description. It should be noted that in the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned, which should be considered as exemplary, and their purpose is only to illustrate the feasibility of the implementation of the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.
[0050] Figure 1 A flowchart of a method for constructing a machine learning model provided in an embodiment of the present application. The process can be executed by a machine learning model construction device to improve the applicability and accuracy of the machine learning model. Figure 1 As shown, the process includes the following steps:
[0051] 101: Perform a clustering operation on multiple original samples included in the original sample set to obtain N cluster samples.
[0052] In some embodiments, a K-Means++ clustering algorithm may be used to perform clustering operations on multiple original samples contained in the original sample set to obtain N cluster samples (one cluster sample may represent an original sample of one category attribute) to improve the quality and efficiency of clustering.
[0053] For example, an original sample is randomly selected in the original sample set as the first centroid. For other original samples in the original sample set, their distances to the selected centroid are calculated in turn, and the point where the original sample with the largest distance is located is selected as the next centroid. The process is repeated until all K centroids are selected; then, the distance between each original sample and the K centroids is calculated, and each original sample is assigned to the cluster to which the nearest centroid belongs, thereby forming multiple clusters. The mean of each cluster is calculated, and the mean is used as the new centroid. The process is repeated until the centroid no longer changes significantly or the predetermined number of iterations is reached, thereby obtaining N (N is an integer greater than 1) cluster samples to ensure that there are similarities between original samples in the same cluster, and at the same time, there are distinctions between original samples in different clusters.
[0054] Through the above steps, the K-Means++ clustering algorithm is used to ensure that the initial centroids have a more applicable distribution, avoid the centroids from gathering in a specific part of the data distribution, reduce the risk of converging to the local minimum, and is suitable for single-source domain generalization scenarios.
[0055] 102: Perform feature enhancement on the N cluster samples to obtain cluster enhanced samples corresponding to the N cluster samples, and form N sample pairs based on the N cluster samples and the cluster enhanced samples corresponding to the N cluster samples.
[0056] In some embodiments, a generative adversarial network can be used to perform feature enhancement on the N cluster samples. The generative adversarial network includes a generator and N cluster discriminators, and one cluster discriminator corresponds to one cluster sample, which is used to discriminate the difference between the enhanced sample generated by the generator and the original sample. The generator is used to adaptively generate different enhanced samples based on different input original samples. Through the structure of the generative adversarial network, it is possible to implement respective enhancement strategies for different cluster samples, thereby enhancing the effectiveness and diversity of the single source domain.
[0057] Further, taking any first cluster sample among the N cluster samples as an example, first, any first original sample in the first cluster sample is input into the generator to obtain the first enhanced sample of the first original sample; then, the first enhanced sample is input into the first cluster discriminator associated with the first cluster sample. If the probability that the first cluster discriminator discriminates that the first enhanced sample is the first original sample does not meet the probability threshold, it indicates that the enhanced sample generated by the generator is not realistic enough, and the parameters of the first cluster discriminator and the generator need to be adjusted until the probability meets the probability threshold (for example, 0 and 1 are used to distinguish whether the generated enhanced sample is realistic). By repeating this process, the parameters of the generator and the cluster discriminator are continuously adjusted and optimized, so that the generator can adaptively generate more realistic enhanced samples of different original samples, and the cluster discriminator can accurately identify the enhanced samples generated by the generator that are related to the cluster.
[0058] In some embodiments, the parameters of the cluster discriminator and the generator are adjusted by first updating the parameters of the cluster discriminator and then updating the parameters of the generator, and the two are performed alternately, which can specifically satisfy the following function:
[0059] … (1)
[0060] … (2)
[0061] Among them, N represents the number of cluster samples, Characterizes the discrimination of the i-th cluster discriminator for the i-th category of enhanced samples, Represents the original sample of the i-th category in N cluster samples, The representation generator generates enhanced samples based on the original samples of the i-th category, Characterize the original sample in The logarithm of the probability of being judged as the i-th category in (which can measure the credibility of the cluster discriminator to the original sample), The logarithm of the probability that the generated enhanced sample is judged by the cluster discriminator as not belonging to the i-th category (which can measure the authenticity discrimination ability of the i-th cluster discriminator for the generated enhanced samples).
[0062] In the above method, in the process of adjusting the parameters of the cluster discriminator, in order to ensure that the i-th cluster discriminator It can accurately identify the original samples of the i-th category, so it maximizes At the same time, the cluster discriminator also needs to be able to effectively distinguish the enhanced samples generated using the i-th class of original samples, so The value of is as close to 0 as possible, so that the i-th cluster discriminator can correctly distinguish true and false data, and then Also needs to be maximized; in the process of adjusting the parameters of the generator, in order to make the generator generate more realistic enhanced samples to deceive cluster discriminators of different categories, it should be maximized , which can be converted to minimize .
[0063] In some embodiments, after the N cluster samples are feature enhanced, N sample pairs can be formed based on the N cluster samples and the cluster enhanced samples corresponding to the N cluster samples, so as to extract effective and comprehensive feature information from them, thereby improving the model's representation ability and understanding depth of the input samples. For example, assuming that the first cluster sample includes original sample A1, original sample A2, and original sample A3, and the corresponding enhanced samples are enhanced sample A1_aug, enhanced sample A2_aug, and enhanced sample A3_aug, respectively, which constitute a sample pair; similarly, assuming that the second cluster sample includes original sample B1, original sample B2, and original sample B3, and the corresponding enhanced samples are enhanced sample B1_aug, enhanced sample B2_aug, and enhanced sample B3_aug, respectively, which constitute another sample pair.
[0064] Through the above steps, the generative adversarial network is used to perform adversarial training on the original samples in each cluster sample, which can generate enhanced samples that are suitable for the original samples in different cluster samples. Compared with the traditional unsupervised single-source domain generalization method, it takes into account the inconsistency of the enhancement strategies required for original samples of different categories, effectively analyzes and learns the differences between different cluster samples, thereby enhancing the effectiveness of different original samples and laying the foundation for subsequent feature extraction.
[0065] 103: Use a convolutional neural network to extract features from the N sample pairs respectively, where the convolutional neural network includes a plurality of convolutional layers connected in sequence.
[0066] In some embodiments, taking four convolutional layers as an example, the convolutional neural network can be a network structure of ResNet18. The first convolutional layer and the second convolutional layer among the four convolutional layers can be called shallow convolutional layers, which focus on capturing the basic feature information of sample pairs (such as texture, color, edges, etc.), which helps the model learn and understand the basic structure of the samples; and the third convolutional layer and the fourth convolutional layer among the four convolutional layers can be called deep convolutional layers, which have a wider receptive field, thereby extracting higher-level and more global deep features (such as tissue structure, shape, etc.).
[0067] 104: For each sample pair, the features output by each of the multiple convolutional layers are fused to obtain a fused feature of each sample pair.
[0068] In some embodiments, the convolutional neural network may further include an attention mechanism layer to facilitate the construction of adaptive weight parameters based on the features extracted by different convolutional layers, so that the model pays more attention to important features; for each sample pair, the features output by multiple convolutional layers are fused to obtain the fused features of each sample pair, which can be: for each sample pair, the features output by multiple convolutional layers are input into the attention mechanism layer in turn, the weight parameters corresponding to each feature are determined, and then the features are fused according to the weight parameters corresponding to each feature, so as to obtain the fused features of each sample pair.
[0069] Furthermore, the attention mechanism layer can be a Squeeze and Excitation Network (SENet), in which the compression part compresses the spatial dimension of the feature into a channel description through global average pooling, and the excitation part reallocates the weight parameters for each channel through the learned weight parameters, thereby enhancing the effective features and suppressing the invalid features. The specific structure is as follows: Figure 2 shown.
[0070] exist Figure 2 In the example, X represents the input features of dimension (H, W, C). First, The function performs compression operation, performs global average pooling on the input features, obtains compressed features and obtains global information on each channel. The specific formula is as follows:
[0071] … (3)
[0072] in, represents the channel of the input feature map X, Represents the feature vector after input feature compression, with dimension (1, 1, C), and then uses To perform activation operations on the feature vector, two fully connected layers can be introduced and activation functions can be applied to learn the weight parameters of each channel. The specific formula is as follows:
[0073] …… (4)
[0074] in, represents the weight matrices of the two fully connected layers, is the ReLU function, Represents the Sigmod function, s represents the weight parameter of the importance of each channel, and the dimension is (1, 1, C).
[0075] Finally, the weight vector s is weighted and multiplied with the input feature X to achieve the weighting of each channel. The specific formula is as follows:
[0076] … (5)
[0077] in, The multiplication function representing the eigenvector and the weight matrix, The output features with representation dimensions (H, W, C) ) channel.
[0078] For example, taking any one of the N sample pairs as an example, Figure 3 A logical diagram of feature fusion provided in an embodiment of the present application. Figure 3 In the example, the first convolution layer, the second convolution layer, the third convolution layer, and the fourth convolution layer sequentially input the features extracted from the sample pair into the SEnet network, thereby sequentially obtaining the first weight parameter of the features of the first convolution layer, the second weight parameter of the features of the second convolution layer, the third weight parameter of the features of the third convolution layer, and the fourth weight parameter of the features of the fourth convolution layer. Then, based on the first weight parameter, the second weight parameter, the third weight parameter, and the fourth weight parameter, the features output by each of the four convolution layers are fused to obtain the fused features of the sample pair.
[0079] Through the above steps, for each sample pair, the attention mechanism layer is introduced to determine the weight parameters corresponding to the features output by different convolutional layers. Then, based on the weight parameters of each feature, the features are fused to make full use of the respective characteristics of shallow features and deep features, so that the model can obtain effective and comprehensive feature information, thereby improving the applicability of the model. Moreover, through the weight parameters of features in different layers, the model can pay more attention to domain-invariant features and ignore domain-related features, thereby overcoming the false correlation between domains.
[0080] 105: Based on the positive sample set, negative sample set, and N fusion features, build a machine learning model that meets the requirements of the loss function.
[0081] In this step, the positive sample pairs contained in the positive sample set are constructed based on the original samples and enhanced samples in the same cluster, and the negative sample pairs contained in the negative sample set are constructed based on the original samples and enhanced samples in different clusters.
[0082] For example, assuming that the first cluster of samples includes original samples A1, original samples A2, and original samples A3, and the corresponding enhanced samples are enhanced samples A1_aug, enhanced samples A2_aug, and enhanced samples A3_aug, respectively, which constitute a sample pair; assuming that the second cluster of samples includes original samples B1, original samples B2, and original samples B3, and the corresponding enhanced samples are enhanced samples B1_aug, enhanced samples B2_aug, and enhanced samples B3_aug, respectively; wherein the positive sample pairs in the positive sample set are based on the original samples A1, original samples A2, original samples A3, enhanced samples A1_aug, enhanced samples A2_aug, and enhanced samples A The feature combination of any two of the original samples B1, B2, B3, enhanced samples B1_aug, enhanced samples B2_aug, and enhanced samples B3_aug is constructed. Similarly, the feature combination of any two of the original samples B1, B2, B3, enhanced samples B1_aug, enhanced samples B2_aug, and enhanced samples B3_aug is constructed. The negative sample pairs in the negative sample set are constructed based on any one of the original samples A1, A2, A3, enhanced samples A1_aug, enhanced samples A2_aug, and enhanced samples A3_aug, and any one of the original samples B1, B2, B3, enhanced samples B1_aug, enhanced samples B2_aug, and enhanced samples B3_aug.
[0083] In some embodiments, the loss function satisfies the following expression:
[0084] …… (6)
[0085] Among them, the To characterize the fusion features of any one of the N sample pairs, Characterization and description The features of the i-th positive sample pair belonging to the same cluster, Characterization and description The features of the j-th negative sample pair belonging to different clusters, the τ represents a temperature parameter, which is used to adjust the sensitivity of the machine learning model to similarity.
[0086] Through the above steps, by constructing the positive sample set and the negative sample set, during the model training process, the distance between the positive sample pairs can be shortened, and the distance between the negative sample pairs can be widened, so that the representations of the positive sample pairs tend to be similar and can be distinguished from the negative sample pairs. The loss function is set as the optimization goal during model training, and the model parameters are continuously updated to achieve the minimized L, which helps the model to better understand the complex relationship between features, and then learn more consistent and discriminative feature representations, solve the problems of different changes in unknown domains and inconsistent feature distributions, and further improve the adaptability of the model.
[0087] In an embodiment of the present application, first, a clustering operation is performed on multiple original samples included in the original sample set to obtain N cluster samples, and feature enhancement is performed on the N cluster samples to obtain cluster enhanced samples corresponding to each of the N cluster samples, and N sample pairs are respectively formed based on the N cluster samples and the cluster enhanced samples corresponding to each of the N cluster samples, so that respective enhancement strategies suitable for different clusters can be executed during the training process; then, a convolutional neural network is used to perform feature extraction on the N sample pairs, and for each sample pair, the features output by each of the multiple convolutional layers connected in sequence in the convolutional neural network are fused to obtain the fused features of each sample pair, which can better balance the global information and detail information of the image during the training process; finally, a machine learning model that meets the requirements of the loss function is constructed based on the positive sample set, the negative sample set, and the N fused features, so that the similarities and differences between different cluster features can be learned during the training process; the above-mentioned entire training process improves the applicability and accuracy of the machine learning model.
[0088] Based on the above Figure 2 The method shown, Figure 4 A schematic diagram of a logical framework for training a machine learning model provided in an embodiment of the present application.
[0089] exist Figure 4 In the method, first, the K-Means++ clustering algorithm is used to perform clustering operations on multiple original samples contained in the original sample set to obtain N cluster samples; then, the N cluster samples are input into the generator in the generative adversarial network to obtain the cluster enhancement samples corresponding to the N cluster samples, and the N cluster enhancement samples are respectively input into the corresponding cluster discriminators, and the process is repeated until the enhanced samples generated by the generator can deceive the cluster discriminator; then, each cluster sample and its corresponding cluster enhancement sample are respectively formed into N sample pairs, and the convolutional neural network is used to extract features from the N sample pairs respectively. For each sample pair, the features output by each convolutional layer in the network are fused to obtain the fused features of each sample pair; finally, based on the N fused features and the constructed positive sample set and negative sample set, comparative learning is performed to continuously learn the similarities and differences between different cluster features until the loss function is met, thereby generating an accurate and generalized machine learning model.
[0090] Based on the same technical concept, a machine learning model construction device is also provided in an embodiment of the present application, which can implement the above-mentioned machine learning model construction method process in the embodiment of the present application.
[0091] Figure 5 This is a schematic diagram of the structure of a machine learning model construction device provided in an embodiment of the present application. Figure 5As shown, the device includes a clustering module 501, a feature enhancement module 502, a feature extraction module 503, a feature fusion module 504, and a model construction module 505.
[0092] The clustering module 501 is used to perform a clustering operation on a plurality of original samples included in the original sample set to obtain N cluster samples; wherein N is an integer greater than 1.
[0093] The feature enhancement module 502 is used to perform feature enhancement on the N cluster samples respectively to obtain cluster enhanced samples corresponding to the N cluster samples respectively, and to form N sample pairs based on the N cluster samples and the cluster enhanced samples corresponding to the N cluster samples respectively.
[0094] The feature extraction module 503 is used to use a convolutional neural network to extract features from the N sample pairs respectively; wherein the convolutional neural network includes a plurality of convolutional layers connected in sequence.
[0095] The feature fusion module 504 is used to fuse the features output by each of the multiple convolutional layers for each sample pair to obtain a fused feature of each sample pair.
[0096] The model construction module 505 is used to construct a machine learning model that meets the requirements of the loss function based on the positive sample set, the negative sample set, and N fusion features; wherein the positive sample pairs contained in the positive sample set are constructed based on the original samples and enhanced samples in the same cluster, and the negative sample pairs contained in the negative sample set are constructed based on the original samples and enhanced samples in different clusters.
[0097] In some embodiments, the convolutional neural network further includes an attention mechanism layer; the feature fusion module 504 is specifically used to:
[0098] For each sample pair, the features output by each of the multiple convolutional layers are sequentially input into the attention mechanism layer to determine the weight parameters corresponding to each of the features;
[0099] The features are fused according to the weight parameters corresponding to the features to obtain the fused features of each sample pair.
[0100] It should be noted here that the above-mentioned device provided in the embodiment of the present application can implement all the method steps in the above-mentioned method embodiment and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0101] Based on the same technical concept, an electronic device is also provided in an embodiment of the present application, which can realize the functions of the aforementioned machine learning model building device.
[0102] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0103] At least one processor 601, and a memory 602 connected to the at least one processor 601. The specific connection medium between the processor 601 and the memory 602 is not limited in the embodiment of the present application. Figure 6 In the example, the processor 601 and the memory 602 are connected via a bus 600. Figure 6 The connection between other components is shown by bold lines, and is not intended to be limiting. The bus 600 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 601 can also be called a controller, and there is no limitation on the name.
[0104] In the embodiment of the present application, the memory 602 stores instructions that can be executed by at least one processor 601. The at least one processor 601 can execute a machine learning model construction method discussed above by executing the instructions stored in the memory 602. The processor 601 can implement Figure 5 The functions of each module in the device shown.
[0105] Among them, the processor 601 is the control center of the device, and can use various interfaces and lines to connect the various parts of the entire control device. By running or executing instructions stored in the memory 602 and calling the data stored in the memory 602, the various functions of the device and processing data, the device can be monitored as a whole.
[0106] In the embodiment of the present application, the processor 601 may include one or more processing units, and the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the modem processor may not be integrated into the processor 601. In some embodiments, the processor 601 and the memory 602 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on separate chips.
[0107] Processor 601 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of a machine learning model construction method disclosed in the embodiments of the present application can be directly embodied as a hardware processor execution, or a combination of hardware and software modules in the processor.
[0108] The memory 602 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 602 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 602 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 602 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0109] By programming the processor 601, the code corresponding to the machine learning model construction method described in the above embodiment can be fixed into the chip, so that the chip can execute the code when running. Figure 1 A method for constructing a machine learning model in the embodiment shown. How to design and program the processor 601 is a technique known to those skilled in the art and will not be described in detail here.
[0110] It should be noted here that the above-mentioned electronic device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0111] Based on the same technical concept, an embodiment of the present application provides a computer storage medium, which includes: a computer program code, which, when executed on a computer, enables the computer to execute any of the methods for constructing a machine learning model discussed above. Since the principle of solving the problem by the above-mentioned computer storage medium is similar to that of constructing a machine learning model, the implementation of the above-mentioned computer storage medium can refer to the implementation of the method, and the repeated parts will not be repeated.
[0112] In a specific implementation process, computer storage media may include: Universal Serial Bus Flash Drive (USB), mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other storage media that can store program codes.
[0113] Based on the same technical concept, the embodiment of the present application also provides a computer program product, which includes: computer program code, when the computer program code is run on a computer, the computer executes any of the methods for building a machine learning model as discussed above. Since the principle of solving the problem by the above-mentioned computer program product is similar to that of building a machine learning model, the implementation of the above-mentioned computer program product can refer to the implementation of the method, and the repeated parts will not be repeated.
[0114] The computer program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0115] The method in the present application can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in the present application is executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM or other programmable device.
[0116] The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0117] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.
[0118] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that has the functions specified in one or more boxes.
[0119] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0121] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for constructing a machine learning model, characterized in that: include: Performing a clustering operation on a plurality of original samples contained in the original sample set to obtain N cluster samples; wherein N is an integer greater than 1; A generative adversarial network is used to perform feature enhancement on the N cluster samples respectively to obtain enhanced samples corresponding to the N cluster samples respectively, and N sample pairs are respectively formed based on the N cluster samples and the cluster enhanced samples corresponding to the N cluster samples respectively; wherein the generative adversarial network includes a generator and N cluster discriminators, one cluster discriminator corresponds to one cluster sample, and is used to discriminate the difference between the enhanced sample generated by the generator and the original sample; Using a convolutional neural network to extract features from the N sample pairs respectively; wherein the convolutional neural network includes a plurality of convolutional layers connected in sequence; For each sample pair, the features output by each of the multiple convolutional layers are fused to obtain a fused feature of each sample pair, so as to balance the global information and detail information of the image during the training process; A machine learning model that conforms to the loss function value is constructed based on the positive sample set, the negative sample set, and N fusion features; wherein the positive sample pairs contained in the positive sample set are constructed based on the original samples and enhanced samples in the same cluster, and the negative sample pairs contained in the negative sample set are constructed based on the original samples and enhanced samples in different clusters; The adopting of a generative adversarial network to perform feature enhancement on the N cluster samples respectively to obtain enhanced samples corresponding to the N cluster samples respectively includes: Input any first original sample in the first cluster of samples into the generator to obtain a first enhanced sample of the first original sample; wherein the first cluster of samples is any one of the N cluster samples; The first enhanced sample is input into a first cluster discriminator associated with the first cluster sample. If the first cluster discriminator determines that the probability that the first enhanced sample is the first original sample does not meet a probability threshold, the respective parameters of the first cluster discriminator and the generator are adjusted until the probability meets the probability threshold.
2. The method according to claim 1, characterized in that The convolutional neural network also includes an attention mechanism layer; For each sample pair, the features output by each of the multiple convolutional layers are fused to obtain the fused features of each sample pair, including: For each sample pair, the features output by each of the multiple convolutional layers are sequentially input into the attention mechanism layer to determine the weight parameters corresponding to each feature; The features are fused according to the weight parameters corresponding to the features to obtain the fused features of each sample pair.
3. The method according to claim 1, characterized in that The loss function satisfies the following expression: Among them, the To characterize the fusion features of any one of the N sample pairs, Characterization and description The features of the i-th positive sample pair belonging to the same cluster, Characterization and description The features of the j-th negative sample pair belonging to different clusters, Characterizing a temperature parameter for adjusting the sensitivity of the machine learning model to similarity.
4. The method according to claim 1, characterized in that The clustering operation is performed on the original samples contained in the original sample set to obtain N cluster samples, including: The K-Means++ clustering algorithm is used to perform a clustering operation on the original samples contained in the original sample set to obtain the N cluster samples.
5. A machine learning model construction device, characterized in that: include: A clustering module, used to perform a clustering operation on a plurality of original samples contained in the original sample set to obtain N cluster samples; wherein N is an integer greater than 1; A feature enhancement module, used to perform feature enhancement on the N cluster samples respectively by using a generative adversarial network to obtain enhanced samples corresponding to the N cluster samples respectively, and to form N sample pairs based on the N cluster samples and the cluster enhanced samples corresponding to the N cluster samples respectively; wherein the generative adversarial network includes a generator and N cluster discriminators, one cluster discriminator corresponds to one cluster sample, and is used to discriminate the difference between the enhanced sample generated by the generator and the original sample; A feature extraction module, used to extract features from the N sample pairs respectively using a convolutional neural network; wherein the convolutional neural network includes a plurality of convolutional layers connected in sequence; A feature fusion module, used for fusing the features output by each of the multiple convolutional layers for each sample pair to obtain a fused feature of each sample pair, so as to balance the global information and detail information of the image during the training process; A model construction module, used to generate a machine learning model that meets the requirements of the loss function based on a positive sample set, a negative sample set, and N fusion features; wherein the positive sample pairs contained in the positive sample set are constructed based on original samples and enhanced samples in the same cluster, and the negative sample pairs contained in the negative sample set are constructed based on original samples and enhanced samples in different clusters; Wherein, the feature enhancement module is specifically used for: Input any first original sample in the first cluster of samples into the generator to obtain a first enhanced sample of the first original sample; wherein the first cluster of samples is any one of the N cluster samples; The first enhanced sample is input into a first cluster discriminator associated with the first cluster sample. If the first cluster discriminator determines that the probability that the first enhanced sample is the first original sample does not meet a probability threshold, the respective parameters of the first cluster discriminator and the generator are adjusted until the probability meets the probability threshold.
6. The device according to claim 5, characterized in that The convolutional neural network also includes an attention mechanism layer; the feature fusion module is specifically used for: For each sample pair, the features output by each of the multiple convolutional layers are sequentially input into the attention mechanism layer to determine the weight parameters corresponding to each feature; The features are fused according to the weight parameters corresponding to the features to obtain the fused features of each sample pair.
7. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to implement the method according to any one of claims 1 to 4 when executing a computer program stored in the memory.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Method for identifying wrong actions in playing music
CN116740815A
Target model training method and device
CN118350448A