Model training method, device and storage medium
By determining the same samples and different samples in the sample set during the training process and using similarity to train the classification model, the problem that noise samples affect the accuracy of model prediction is solved, and efficient utilization of noise samples and improved the accuracy of model prediction is achieved.
Patent Information
- Application Number
- CN202210785492.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The prior art is difficult to effectively utilize noise samples when training algorithm models, resulting in low model prediction accuracy.
By obtaining the multimedia samples in the sample set, determining the same and different samples, training the classification model with the first and second similarities increases the first similarity and decreases the second similarity to ensure that the target sample is closer to the same samples and away from the different samples.
It improves the accuracy of model prediction, can efficiently use noise samples for training, and ensures the prediction effect of the model.
Smart Images

Figure CN115130597B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and specifically to a model training method, device, and storage medium. Background Art
[0002] In recent years, artificial intelligence technology has achieved unprecedented development in various fields. Machine learning models (algorithm models) trained based on data have been applied in many fields, such as common virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc.
[0003] However, it is currently difficult to obtain a sample set consisting entirely of clean samples before training the algorithm model. The sample set currently used to train the algorithm model includes noise samples. Clean samples refer to samples that are correctly labeled, and noise samples refer to samples that are incorrectly labeled. When the algorithm model is trained using noisy samples in the sample set, the accuracy of the algorithm model's prediction is low. Summary of the Invention
[0004] The embodiments of the present application provide a model training method, device and storage medium, which can efficiently utilize noise samples to train the model and ensure the accuracy of model prediction.
[0005] The present invention provides a model training method, including:
[0006] Acquire a sample set, the sample set including multiple multimedia samples, the multimedia samples carrying pending labels and pending label types, the pending label types including correct types and incorrect types;
[0007] According to the pending label and pending label type carried by the target sample, the same sample and different sample are determined from the sample set. The same sample is a multimedia sample with the same actual label as the target sample, and the different sample is a multimedia sample with a different actual label from the target sample. The target sample is any multimedia sample in the sample set;
[0008] Determining a first similarity and a second similarity, the first similarity being a similarity between the target sample and the same sample, and the second similarity being a similarity between the target sample and the different sample;
[0009] The classification model is trained based on the first similarity and the second similarity to obtain a trained classification model.
[0010] The present application also provides a model training device, including:
[0011] An acquisition unit is used to acquire a sample set, the sample set including a plurality of multimedia samples, the multimedia samples carrying a pending label and a pending label type, the pending label type including a correct type and an incorrect type;
[0012] a sample determination unit, configured to determine identical samples and different samples from a sample set based on the pending label and the pending label type carried by the target sample, wherein the identical sample is a multimedia sample with the same actual label as the target sample, and the different sample is a multimedia sample with a different actual label from the target sample, and the target sample is any multimedia sample in the sample set;
[0013] a similarity determination unit, configured to determine a first similarity and a second similarity, the first similarity being the similarity between the target sample and the same sample, and the second similarity being the similarity between the target sample and the different sample;
[0014] The model training unit is used to train the classification model based on the first similarity and the second similarity to obtain a trained classification model.
[0015] In some embodiments, the pending label type includes a correct type and an incorrect type. Determining identical samples and different samples from a sample set based on the pending label and the pending label type carried by the target sample includes:
[0016] Screening the multimedia samples in the sample set according to the pending tag type carried by the multimedia samples to obtain a first set and a second set, wherein the first set includes multimedia samples carrying the pending tag type of a correct type, and the second set includes multimedia samples carrying the pending tag type of an incorrect type;
[0017] According to the pending label and the pending label type carried by the target sample, the same samples and different samples are determined from the first set and the second set.
[0018] In some embodiments, determining identical samples and different samples from a sample set based on the pending label and the pending label type carried by the target sample includes:
[0019] When the type of the pending label carried by the target sample is the correct type, the same sample is determined from the first set, and the same sample has the same pending label as the target sample;
[0020] Different samples are determined from the first set and the second set, where the different samples include multimedia samples in the first set with different pending labels from the target sample, and multimedia samples in the second set with the same pending labels as the target sample.
[0021] In some embodiments, the sample set includes multiple sample pairs, the target sample is a multimedia sample in any sample pair, and determining identical samples and different samples from the sample set based on the pending label and the pending label type carried by the target sample includes:
[0022] When the pending label type carried by the target sample is an incorrect type, the multimedia samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located;
[0023] And different samples are determined from the first set, where the different samples are multimedia samples in the first set that have the same to-be-determined label as the target sample.
[0024] In some embodiments, the sample set includes multiple sample pairs, the target sample is a sample in any sample pair, and determining the same sample from the first set further includes:
[0025] The multimedia samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located.
[0026] In some embodiments, determining the first similarity and the second similarity includes:
[0027] Extracting a target vector, a first vector, and a second vector, wherein the target vector is the vector of the target sample, the first vector is the vector of the same sample, and the second vector is the vector of the different samples;
[0028] A first similarity is determined based on the target vector and the first vector, and a second similarity is determined based on the target vector and the second vector.
[0029] In some embodiments, training a classification model based on the first similarity and the second similarity to obtain a trained classification model includes:
[0030] Adding all first similarities and all second similarities to obtain a total similarity;
[0031] determining a quotient of each first similarity divided by the total similarity;
[0032] Determine the logarithm of each quotient;
[0033] Divide the sum of all logarithms by the determinant of the first matrix to obtain the loss value corresponding to the target sample. The first matrix includes the vectors of all identical samples.
[0034] According to the loss value corresponding to each target sample, the classification model is trained to obtain a trained classification model.
[0035] In some embodiments, after training the classification model based on the first similarity and the second similarity to obtain the trained classification model, the method further includes:
[0036] Using the trained classification model, obtain the multimedia clips to be identified;
[0037] Extract features from multimedia clips to obtain features corresponding to the multimedia clips;
[0038] The multimedia segments are classified according to the features to obtain the types of the multimedia segments.
[0039] An embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the steps in any model training method provided in the embodiment of the present application.
[0040] An embodiment of the present application can obtain a sample set, which includes multiple multimedia samples. The multimedia samples carry pending labels and pending label types, and the pending label types include correct types and incorrect types. According to the pending labels and pending label types carried by the target samples, identical samples and different samples are determined from the sample set, where the identical samples are multimedia samples with the same actual labels as the target samples, and the different samples are multimedia samples with different actual labels from the target samples. The target sample is any multimedia sample in the sample set. A first similarity and a second similarity are determined, where the first similarity is the similarity between the target sample and the identical sample, and the second similarity is the similarity between the target sample and the different sample. Based on the first similarity and the second similarity, a classification model is trained to obtain a trained classification model.
[0041] In the present application, the sample set may include noise samples, which carry pending labels of the wrong type. By using any multimedia sample in the sample set as the target sample, identical samples with the same actual labels as the target sample and different samples with different actual labels from the target sample can be determined from the sample set. By using identical samples and different samples, the noise samples will not affect the accuracy of the first similarity and the second similarity. When training the classification model, the first similarity can be increased and the second similarity can be decreased, so that the target sample is closer to the identical sample and away from the different sample to ensure the accuracy of the model prediction. Therefore, this solution can efficiently utilize the noise sample training model. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0043] Figure 1a This is a schematic diagram of a scenario of the model training method provided in an embodiment of the present application;
[0044] Figure 1b This is a flow chart of the model training method provided in the embodiment of the present application;
[0045] Figure 2a This is a schematic diagram of classification of identical samples and different samples provided by an embodiment of the present application;
[0046] Figure 2b This is another schematic diagram of classification of identical samples and different samples provided by the embodiment of the present application;
[0047] Figure 3 Schematic diagram of the structure of the model training device provided in the embodiment of the present application;
[0048] Figure 4 It is a structural diagram of the server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0050] The embodiments of the present application provide a model training method, device and storage medium.
[0051] The model training device can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC); the server can be a single server or a server cluster consisting of multiple servers.
[0052] In some embodiments, the model training device can also be integrated into multiple electronic devices. For example, the model training device can be integrated into multiple servers, and the model training method of the present application can be implemented by multiple servers.
[0053] In some embodiments, the server may also be implemented in the form of a terminal.
[0054] For example, reference Figure 1a The electronic device can obtain a sample set, which includes multiple multimedia samples, and the multimedia samples carry pending labels and pending label types; determine identical samples and different samples from the sample set according to the pending labels and pending label types carried by the target samples, where the identical samples are multimedia samples with the same actual labels as the target samples, and the different samples are samples with different actual labels from the target samples, and the target sample is any multimedia sample in the sample set; determine a first similarity and a second similarity, where the first similarity is the similarity between the target sample and the identical sample, and the second similarity is the similarity between the target sample and the different sample; train a classification model based on the first similarity and the second similarity to obtain a trained classification model.
[0055] The sample set may include noise samples, which carry pending labels of the wrong type. During model training, any multimedia sample in the sample set is used as a target sample. The target sample can be a clean sample or a noise sample. Before model training, it is necessary to determine from the sample set multimedia samples with the same actual label as the target sample as the same sample, and multimedia samples with different actual labels from the target sample as the different samples. Due to the same samples and different samples, the noise samples do not affect the accuracy of the first similarity and the second similarity. When training the classification model, the first similarity can be increased and the second similarity can be decreased, so that the target sample is closer to the same sample and away from the different sample to ensure the accuracy of the model prediction. Therefore, this solution can efficiently utilize noise samples to train the model.
[0056] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0057] Artificial intelligence (AI) is a technology that uses digital computers to simulate human-like perception of the environment, acquisition, and application of knowledge. This technology enables machines to possess human-like perception, reasoning, and decision-making capabilities. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.
[0058] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0059] Autonomous driving technology usually includes high-precision maps, environmental perception, behavioral decision-making, path planning, motion control and other technologies. Autonomous driving technology has broad application prospects. With the research and advancement of artificial intelligence technology, artificial intelligence technology has been researched and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, intelligent customer service, Internet of Vehicles, autonomous driving, smart transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0060] In this embodiment, a model training method involving artificial intelligence is provided, such as Figure 1b As shown, the specific process of the model training method can be as follows:
[0061] 110. Obtain a sample set, where the sample set includes multiple multimedia samples, and the multimedia samples carry pending tags and pending tag types.
[0062] The sample set consists of the multimedia samples used to train the model. It includes clean samples and noisy samples. Clean samples carry the correct type of pending labels, while noisy samples carry the incorrect type of pending labels. Correct labels are true labels, while incorrect labels are labels that do not match reality. For example, the multimedia samples in the sample set can be multiple frames of images, multiple audio clips, multiple video clips, multiple text fields, and so on.
[0063] For example, suppose a sample set D = ,in, , Refers to the i-th multimedia sample in the sample set, R is the value range of the sample vector, and d is The dimension is The number of vectors with M is the number of samples.
[0064] In some embodiments, there are multiple ways to obtain multimedia samples in the sample set, for example, they can be collected offline, crawled on the network, obtained from a server, and so on.
[0065] A pending label refers to the annotated label of a multimedia sample. For example, if the multimedia sample is an image, and the object to be identified in the image is "cat," then the annotated label for the image could be "cat," "dog," "pig," and so on. If the annotated label is "cat," then the pending label is the true label, and the sample is a clean sample. If the annotated label is "dog," "pig," and so on, then the pending label is an untrue label, and the sample is a noisy sample.
[0066] The pending label type is used to indicate whether the label is true. For example, if the pending label matches the target to be identified in the multimedia sample, the pending label is a true label, and the pending label type can use the correct type to indicate that the pending label is a true label. If the pending label does not match the target to be identified in the multimedia sample, the pending label is an untrue label, and the pending label type can use the error type to indicate that the pending label is an untrue label.
[0067] In some embodiments, there are multiple ways to determine the type of tag to be determined, for example, it can be determined manually, or obtained through model screening, etc.
[0068] For example, in some embodiments, a method for model classification of pending label types may include the following steps:
[0069] 1. Use multimedia samples in the sample set to train the model and obtain the trained initial model;
[0070] 2. Use the trained initial model to predict the sample set and obtain the prediction results;
[0071] 3. Determine the pending label type of the pending label based on the absolute error between the predicted result and the pending label.
[0072] The model used for training may be a clustering model, for example, a tree model (a tree-like clustering model), a system clustering model, and the like.
[0073] For example, if the absolute error between the predicted result and the pending label is less than a preset value, the pending label of the multimedia sample is considered to be the correct type. If the absolute error between the predicted result and the pending label is greater than a preset value, the pending label of the multimedia sample is considered to be the wrong type.
[0074] The preset value can be set according to the actual application scenario.
[0075] In some embodiments, the method for determining the type of the undetermined tag may also be based on a noise sample detection method of feature clustering, an algorithm classification model (using K-means clustering), and the like.
[0076] 120. According to the pending label and the pending label type carried by the target sample, the same sample and the different sample are determined from the sample set. The same sample is a multimedia sample with the same actual label as the target sample, and the different sample is a multimedia sample with a different actual label from the target sample. The target sample is any multimedia sample in the sample set.
[0077] The target sample is the multimedia sample used to train the model at the current moment, and the target sample is any multimedia sample in the sample set. For example, if any sample in the sample set is used as the target sample to train the model, the pending label carried by the target sample can be a correct type of label or an incorrect type of label.
[0078] The actual label of the target sample refers to the true label of the target sample. For example, if the target sample is a clean sample, the actual label of the target sample is the same as the pending label. If the target sample is a noise sample, the actual label of the target sample is different from the pending label.
[0079] An identical sample is a multimedia sample with the same actual label as the target sample. For example, if the target sample is an image and its actual label is "cat," then an identical sample is a multimedia sample with the actual label "cat." If the target sample is audio and its actual label is "lyrical," then an identical sample is a multimedia sample with the actual label "lyrical." If the target sample is a video and its actual label is "comedy," then an identical sample is a multimedia sample with the actual label "comedy."
[0080] Dissimilar samples are multimedia samples with different actual labels than the target sample. For example, if the target sample is an image and its actual label is "cat," then a dissimilar sample is a multimedia sample whose actual label is definitely not "cat." If the target sample is audio and its actual label is "lyrical," then a dissimilar sample is a multimedia sample whose actual label is definitely not "lyrical." If the target sample is a video and its actual label is "comedy," then a dissimilar sample is a multimedia sample whose actual label is definitely not "comedy."
[0081] In some embodiments, the method for determining identical samples and different samples can be applied locally or remotely, specifically in a server or a cloud server.
[0082] In some embodiments, considering that the sample set includes clean samples and noise samples, in order to facilitate determining multimedia samples with the same actual label as the target sample and multimedia samples with different labels from the sample set, the pending label type includes a correct type and an incorrect type. According to the pending label and the pending label type carried by the target sample, determining the same sample and the different sample from the sample set includes:
[0083] 121. Filter samples in the sample set according to the pending tag types carried by the multimedia samples to obtain a first set and a second set, wherein the first set includes multimedia samples carrying the pending tag type of a correct type, and the second set includes multimedia samples carrying the pending tag type of an incorrect type.
[0084] 122. Determine identical samples and different samples from the first set and the second set based on the pending label and the pending label type carried by the target sample.
[0085] The first set includes multimedia samples carrying the correct type of pending tags.
[0086] The second set includes multimedia samples carrying pending labels of the wrong type.
[0087] For example, if the multimedia samples are images, the first set consists of images with pending labels of the correct type (the clean sample set), and the second set consists of images with pending labels of the incorrect type (the noisy sample set). If the multimedia samples are audio, the first set can also consist of audio with pending labels of the correct type (the clean sample set), and the second set can consist of audio with pending labels of the incorrect type (the noisy sample set). If the multimedia samples are videos, the first set can also consist of videos with pending labels of the correct type (the clean sample set), and the second set can consist of videos with pending labels of the incorrect type (the noisy sample set), and so on.
[0088] In some embodiments, considering that the undetermined label type carried by the target sample may be the correct type, in order to determine the same samples and different samples matching the target sample, step 122 includes:
[0089] When the type of the pending label carried by the target sample is the correct type, the same sample is determined from the first set, and the same sample has the same pending label as the target sample;
[0090] Different samples are determined from the first set and the second set, where the different samples include multimedia samples in the first set with different pending labels from the target sample, and multimedia samples in the second set with the same pending labels as the target sample.
[0091] For example, if the target sample's pending label type is the correct type, that is, the target sample belongs to the multimedia sample in the first set (clean sample set), the first set (clean sample set) is , the second set (noise sample set) is .
[0092] The same sample includes ,in, are the same sample set, and Indicates that the target sample comes from the first set. j It means taking the multimedia samples numbered j in the first set as the same samples, refers to the label of the target sample in the sample set, refers to the label of the same sample in the first set, It means that the target sample is the pending label of the i-th sample in the sample set. It means that the same sample is the pending label of the j-th sample in the sample set.
[0093] Different samples include ,in, is a set of different samples, and Indicates that the target sample comes from the first set. It means that the second set is labeled Multimedia samples of are used as different samples, It means that the first set is labeled Multimedia samples of are used as different samples, refers to the label of the target sample in the sample set, refers to the labels of different samples in the second set, refers to the labels of different samples in the first set, It means that the target sample is the pending label of the i-th multimedia sample in the sample set. It means that the different samples are the first The pending labels of multimedia samples, It means that the different samples are the first The pending labels of multimedia samples.
[0094] In some embodiments, considering that the type of the undetermined label carried by the target sample may be an incorrect type, in order to ensure that there is a multimedia sample in the sample set with the same actual label as the target sample, the sample set includes multiple sample pairs, the target sample is a multimedia sample in any sample pair, and determining the same sample from the first set further includes:
[0095] The multimedia samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located.
[0096] The sample pair includes at least two multimedia samples from the same target. For example, the target is image A, and an image processing is performed on image A to obtain image , perform another image processing on image A to obtain image , then the image and images A sample pair is formed, where image processing includes image compression, image angle adjustment, color adjustment, etc. For another example, the target is video S, which is composed of video segments S1, S2, S3...S10 in time sequence. For video S, at least two video segments from S1, S2, S3...S10 can be selected. Then, at least two video segments form a sample pair.
[0097] The target sample pair refers to the sample pair where the target sample is located. For example, the target sample pair includes and ,in, and Used to distinguish two multimedia samples in a sample pair, is the multimedia sample with number G in the i-th sample pair in the sample set, The number of the i-th sample pair in the sample set is The target sample pairs can be image sample pairs, audio sample pairs, video sample pairs, text sample pairs, etc.
[0098] For example, the target sample pair includes and , the target sample is ,but for The same sample of .
[0099] In some embodiments, considering that the undetermined label type carried by the target sample may be an erroneous type, in order to ensure that there is a multimedia sample in the sample set with the same actual label as the target sample, the sample set includes multiple sample pairs, and the target sample is a multimedia sample in any sample pair. Step 122 includes:
[0100] When the type of the pending label carried by the target sample is the wrong type, the samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located;
[0101] And different samples are determined from the first set, where the different samples are multimedia samples in the first set that have the same to-be-determined label as the target sample.
[0102] For example, when the target sample carries an incorrect tag type, that is, the target sample belongs to a multimedia sample in the second set (noise sample set), the first set (clean sample set) is , the second set (noise sample set) is , the target sample pairs include and ,in, is the multimedia sample with number G in the i-th sample pair in the sample set, The number of the i-th sample pair in the sample set is Multimedia samples.
[0103] The same sample includes ,in, are the same sample set, and Indicates that the target sample comes from the second sample set.
[0104] Different samples include ,in, is a set of different samples, and Indicates that the target sample comes from the second sample set. It means taking the multimedia samples numbered k in the first set as different samples, refers to the label of the target sample in the sample set, refers to the labels of different samples in the first set, It means that the target sample is the pending label of the i-th multimedia sample in the sample set. It means that the different samples are the first The pending labels of multimedia samples.
[0105] 130. Determine a first similarity and a second similarity, the first similarity being the similarity between the target sample and the same sample, and the second similarity being the similarity between the target sample and the different sample.
[0106] The first similarity is the similarity between the target sample and the same sample.
[0107] The second similarity is the similarity between the target sample and the different samples.
[0108] For example, similarity can refer to the Euclidean distance between the vector of the target sample and the vector of different samples / different samples, the cosine distance between the vector of the target sample and the vector of different samples / different samples, the Jaccard distance between the vector of the target sample and the vector of different samples / different samples, and so on.
[0109] In some embodiments, in order to train a model using noise samples without reducing the accuracy of the model in predicting the target, determining the first similarity and the second similarity includes:
[0110] Extracting a target vector, a first vector, and a second vector, wherein the target vector is the vector of the target sample, the first vector is the vector of the same sample, and the second vector is the vector of the different samples;
[0111] A first similarity is determined based on the target vector and the first vector, and a second similarity is determined based on the target vector and the second vector.
[0112] Among them, the target vector is used to represent the target sample.
[0113] The first vector is used to represent the same sample.
[0114] The second vector is used to represent the vectors of different samples.
[0115] Determine the first similarity based on the target vector and the first vector: .
[0116] Determine the second similarity based on the target vector and the second vector: .
[0117] in, is the vector of target samples, is a vector of the same sample, is the vector of different samples, is a hyperparameter for the temperature scale.
[0118] 140. Based on the first similarity and the second similarity, train the classification model to obtain a trained classification model.
[0119] The classification model is used to identify the type of the object to be detected, which can be a target to be identified in an image, a multimedia clip, a text field, etc.
[0120] For example, a classification model can identify cats, dogs, etc. in an image. It can also identify the genre of a video clip, which could be comedy, food, fashion, travel, etc. It can also identify the genre of a text field, which could be martial arts, suspense, science fiction, etc.
[0121] In some embodiments, in order to train a classification model by similarity to maximize the similarity between similar samples and minimize the similarity between dissimilar samples, the classification model is trained based on the first similarity and the second similarity to obtain a trained classification model, including:
[0122] Adding all first similarities and all second similarities to obtain a total similarity;
[0123] determining a quotient of each first similarity divided by the total similarity;
[0124] Determine the logarithm of each quotient;
[0125] Divide the sum of all logarithms by the determinant of the first matrix to obtain the loss value corresponding to the target sample. The first matrix includes the vectors of all identical samples.
[0126] According to the loss value corresponding to each target sample, the classification model is trained to obtain a trained classification model.
[0127] The total similarity includes a first similarity between the target object and all identical samples, and a second similarity between the target object and all different samples.
[0128] Total similarity = = + ,in, are multimedia samples labeled y in the same sample set P and different sample sets N, for vector, is the multimedia sample numbered p in the same sample set, is the multimedia sample numbered n in the different sample set.
[0129] Business = .
[0130] Logarithm = .
[0131] The first matrix contains all vectors of the same sample.
[0132] Loss value = ,in, = , is the determinant of the first matrix.
[0133] In some embodiments, in order to improve the similarity between the target sample and the same sample and reduce the similarity between the target sample and the different samples, the classification model is trained according to the loss value corresponding to each target sample to obtain the trained classification model, including:
[0134] Dividing the sum of the loss values corresponding to all target samples by the determinant of the second matrix to obtain a target loss value corresponding to all target samples, where the second matrix includes vectors of all multimedia samples in the sample set;
[0135] Perform gradient descent on the target loss value in the classification model to obtain the trained classification model.
[0136] The second matrix includes vectors of all samples in the sample set.
[0137] The target loss value refers to a regularization term corresponding to the classification model trained on all target samples in the sample set. The target loss value is , , is the determinant of the second matrix.
[0138] In some embodiments, considering that classification of videos, audios, etc. composed of multiple frames of images can be achieved, after training the classification model based on the first similarity and the second similarity to obtain the trained classification model, the method further includes:
[0139] Using the trained classification model, obtain the multimedia clips to be identified;
[0140] Extract features from multimedia clips to obtain features corresponding to the multimedia clips;
[0141] The multimedia segments are classified according to the features to obtain the types of the multimedia segments.
[0142] The multimedia segments to be identified may be multimedia segments that need to be classified, for example, multimedia segments may be images, video segments, audio segments, text fields, and the like.
[0143] Features are used to represent multimedia segments. For example, features can be vectors of multimedia segments.
[0144] The type of a multimedia segment refers to the type to which the multimedia segment is classified. For example, if the multimedia segment is a video, the type of the multimedia segment may be comedy, food, fashion, travel, etc.
[0145] Among them, the trained classification model classifies the multimedia clips according to the features through the normalized exponential function (softmax function) to obtain the type of the multimedia clips. The multimedia segment is represented by the mapping header The normalized features (vectors) of the mapping, through Z, we can know the type of multimedia segment, , and there are , It refers to the dimension of the multimedia clip after normalization. The normalization can be done using a normalization model such as softmax or sigmoid.
[0146] The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.
[0147] As can be seen from the above, the embodiment of the present application can obtain a sample set, which includes multiple multimedia samples, and the samples carry pending labels and pending label types, and the pending label types include correct types and incorrect types; according to the pending labels and pending label types carried by the target samples, the same samples and different samples are determined from the sample set, the same samples are multimedia samples with the same actual labels as the target samples, and the different samples are multimedia samples with different actual labels from the target samples, and the target sample is any multimedia sample in the sample set; determine a first similarity and a second similarity, the first similarity is the similarity between the target sample and the same sample, and the second similarity is the similarity between the target sample and the different samples; based on the first similarity and the second similarity, train the classification model to obtain the trained classification model.
[0148] In an embodiment of the present invention, the sample set may include noise samples, which carry a pending label of the wrong type. By using any sample in the sample set as a target sample, identical samples with the same actual label as the target sample and different samples with different actual labels can be determined from the sample set. By using identical samples and different samples, the noise samples will not affect the accuracy of the first similarity and the second similarity. When training the classification model, the first similarity can be increased and the second similarity can be decreased, so that the target sample is closer to the identical sample and away from the different sample to ensure the accuracy of the model prediction. Therefore, this solution can efficiently utilize noise samples to train the model.
[0149] The method described in the above embodiment will be further described below.
[0150] In this embodiment, the method of the embodiment of the present application will be described in detail by taking the application in a video classification model as an example.
[0151] The specific process of a model training method is as follows:
[0152] 210. Obtain a sample set, where the sample set includes multiple sample pairs, and each multimedia sample in the sample pair carries a to-be-determined label.
[0153] For example, the sample set D= ,in , where M represents the number of multimedia samples, Refers to the i-th multimedia sample in the sample set.
[0154] 220. Classify the pending label carried by each multimedia sample in the sample set according to the multimedia samples and the pending label carried by the multimedia samples to obtain a pending label type, where the pending label type includes a correct type and an incorrect type.
[0155] 230. Screen the multimedia samples in the sample set according to the pending tag type carried by the sample to obtain a first set and a second set, wherein the first set includes samples carrying the pending tag type of a correct type, and the second set includes multimedia samples carrying the pending tag type of an incorrect type.
[0156] For example, the first set is a clean sample set , the second set is the noise sample set .
[0157] 240. Determine identical samples and different samples from the first set and the second set based on the pending label and the pending label type carried by the target sample. Identical samples are multimedia samples with the same actual label as the target sample, and different samples are multimedia samples with different actual labels from the target sample. The target sample is any multimedia sample in the sample set.
[0158] For example, randomly sampling a segment pair from a video ( , ) to represent two different perspectives of the video, and both perspective segments are composed of T frames. For multimedia samples that are images, a sample pair refers to the results of two different enhancements to the original image. The samples of these two perspectives are defined as ( , ).
[0159] In noise contrastive learning, one of the samples in the pair is as the target sample.
[0160] when When from the first set, the same sample includes The samples in , different samples include Multimedia samples in .
[0161] like Figure 2a As shown, the Y interface divides the sample set into the first set and the second set, and the X interface divides the multimedia samples in the first set into the same samples and different samples, and divides the multimedia samples in the second set into different samples and other samples, where the other samples are multimedia samples whose actual labels are uncertain to be the same or different from those of the target samples. When coming from the first set, identical samples include multimedia samples in the same sample pair as the target sample, as well as multimedia samples and sample pairs in the first set that share the same pending label as the target sample. Dissimilar samples include sample pairs in the first set whose pending labels differ from those of the target sample, as well as multimedia samples and sample pairs in the second set whose pending labels are the same as those of the target sample. Thus, identical samples share the same actual label as the target sample, while dissimilar samples differ from the target sample in actual label.
[0162] when When from the second set, the same sample includes , different samples include The samples in .
[0163] The set of identical samples is P and the set of different samples is N.
[0164] like Figure 2b As shown, the Y interface divides the sample set into the first set and the second set, and the X interface divides the multimedia samples in the first set into different samples and other samples, and divides the multimedia samples in the second set into the same samples and other samples. The other samples are multimedia samples that are uncertain whether they are the same or different from the actual labels of the target samples. When the target sample is from the second set, identical samples include multimedia samples in the same sample pair as the target sample, and distinct samples include multimedia samples and sample pairs from the first set whose pending labels are the same as the target sample's pending labels. Thus, identical samples have the same actual labels as the target sample, while distinct samples have different actual labels from the target sample.
[0165] 250. Determine a first similarity and a second similarity, where the first similarity is the similarity between the target sample and the same sample, and the second similarity is the similarity between the target sample and the different sample.
[0166] 260. Determine a loss value corresponding to the target sample according to the first similarity and the second similarity.
[0167] Loss value = ,in, is the vector of a multimedia sample in the same sample set P and different sample set N, for vector, is a multimedia sample in the same sample set, is the nth sample in the set of different samples, = , is the determinant of the first matrix, which includes vectors of all identical samples.
[0168] 270. Determine a target loss value according to the loss value corresponding to each target sample.
[0169] The target loss value is , , is the determinant of the second matrix.
[0170] 280. Perform gradient descent on the target loss value in the classification model to obtain the trained classification model.
[0171] Experimental verification of the model training method:
[0172] Experiments were conducted on three common video classification datasets: Kinetics (K400), Mini-Kinetics (K200), and Something-Something-V1 (SthV1), as well as two common image classification datasets: CIFAR10 and CIFAR100. For video tasks, TSM-ResNet50 was used as the base architecture for all experiments. For image tasks, ResNet32 was used as the base model.
[0173] Symmetrical noise and equivalent noise were constructed. Symmetrical noise randomly and independently assigns a label to each sample in the sample set. This label will never be the same as its actual label, and the probability of different labels is equal. In this experiment, the noise ratios are 20%, 40%, 60%, and 80%. Equivalent noise, on the other hand, replaces the labels of samples in one category with the labels of another category in a targeted manner. In this experiment, the noise ratios are 10%, 20%, and 40%.
[0174] The experiment used a noise sample detection method (CT) based on feature clustering. This detection method (CT) has a high accuracy rate when detecting noise samples in a sample set, but it cannot guarantee that the detection results are completely correct.
[0175] Based on the CT method, the first set (clean sample set) and the second set (noise sample set) were obtained. On this basis, the model training method of this application was compared with several common methods using noisy data. The pseudo-label method (PL) uses the model's predicted value to label the sample with a pseudo label and uses the pseudo label to train the sample. Instance contrastive learning (CL) enhances the model's ability to capture sample features. Supervised contrastive learning (SCL) takes category information into account when performing contrastive learning. DivideMix is one of the best methods for solving noisy learning problems on previous image datasets. For the sake of fairness, all methods did not use Mix-up data enhancement.
[0176] The experimental results on videos and images are shown in Tables 1-5. The tables record the best accuracy achieved by all methods on the sample set during training.
[0177] Table 1: Test accuracy of the Mini-kinetics sample set (%)
[0178]
[0179] Table 2: Test accuracy (%) of Kinetics sample set
[0180]
[0181] Table 3: Test accuracy of the Something-Something-V1 sample set (%)
[0182]
[0183] Table 4: Test accuracy (%) of CIFAR10 sample set
[0184]
[0185] Table 5: Test accuracy of CIFAR100 dataset (%)
[0186]
[0187] As can be seen from the above, the model training method of this application can stably improve the model accuracy in all groups and is superior to previous methods. In addition, the model training method of this application abandons the previous practice of using multiple enhancements of the same sample to construct consistency to improve the credibility of pseudo-labels. It does not require exponentially augmented data. Compared with semi-supervised methods such as DivideMix, the model training method of this application speeds up computational efficiency and effectively utilizes noise samples.
[0188] To better implement the above method, the present application also provides a model training device. The model training device can be integrated into an electronic device, such as a terminal or server. The terminal includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, and other devices. The server can be a single server or a server cluster consisting of multiple servers.
[0189] For example, in this embodiment, the method of the embodiment of the present application will be described in detail by taking the specific integration of the model training device into an electronic device as an example.
[0190] For example, Figure 3As shown, the model training device may include an acquisition unit 310, a sample determination unit 320, a similarity determination unit 330, and a model training unit 340, as follows:
[0191] (1) Acquisition unit 310.
[0192] The acquisition unit is used to acquire a sample set, where the sample set includes multiple multimedia samples, and the multimedia samples carry a pending label and a pending label type.
[0193] (2) Sample determination unit 320.
[0194] The sample determination unit 320 is used to determine the same sample and different samples from the sample set based on the pending label and the pending label type carried by the target sample. The same sample is a multimedia sample with the same actual label as the target sample, and the different sample is a multimedia sample with a different actual label from the target sample. The target sample is any multimedia sample in the sample set.
[0195] In some embodiments, the pending label type includes a correct type and an incorrect type. Determining identical samples and different samples from a sample set based on the pending label and the pending label type carried by the target sample includes:
[0196] Screening samples in the sample set according to the pending label type carried by the multimedia samples to obtain a first set and a second set, wherein the first set includes multimedia samples carrying the pending label type of a correct type, and the second set includes multimedia samples carrying the pending label type of an incorrect type;
[0197] According to the pending label and the pending label type carried by the target sample, the same samples and different samples are determined from the first set and the second set.
[0198] In some embodiments, determining identical samples and different samples from a sample set based on the pending label and the pending label type carried by the target sample includes:
[0199] When the type of the pending label carried by the target sample is the correct type, the same sample is determined from the first set, and the same sample has the same pending label as the target sample;
[0200] Different samples are determined from the first set and the second set, where the different samples include multimedia samples in the first set with different pending labels from the target sample, and multimedia samples in the second set with the same pending labels as the target sample.
[0201] In some embodiments, the sample set includes multiple sample pairs, the target sample is a sample in any sample pair, and according to the pending label and the pending label type carried by the target sample, determining the same samples and different samples from the sample set includes:
[0202] When the pending label type carried by the target sample is an incorrect type, the multimedia samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located;
[0203] And different samples are determined from the first set, where the different samples are multimedia samples in the first set that have the same to-be-determined label as the target sample.
[0204] In some embodiments, the sample set includes multiple sample pairs, the target sample is a sample in any sample pair, and determining the same sample from the first set further includes:
[0205] The multimedia samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located.
[0206] (3) Similarity determination unit 330.
[0207] The similarity determination unit 330 is configured to determine a first similarity and a second similarity, wherein the first similarity is the similarity between the target sample and the same sample, and the second similarity is the similarity between the target sample and the different sample.
[0208] In some embodiments, determining the first similarity and the second similarity includes:
[0209] Extracting a target vector, a first vector, and a second vector, wherein the target vector is the vector of the target sample, the first vector is the vector of the same sample, and the second vector is the vector of the different samples;
[0210] A first similarity is determined based on the target vector and the first vector, and a second similarity is determined based on the target vector and the second vector.
[0211] (4) Model training unit 340.
[0212] The model training unit 340 is configured to train the classification model based on the first similarity and the second similarity to obtain a trained classification model.
[0213] In some embodiments, training a classification model based on the first similarity and the second similarity to obtain a trained classification model includes:
[0214] Adding all first similarities and all second similarities to obtain a total similarity;
[0215] determining a quotient of each first similarity divided by the total similarity;
[0216] Determine the logarithm of each quotient;
[0217] Divide the sum of all logarithms by the determinant of the first matrix to obtain the loss value corresponding to the target sample. The first matrix includes the vectors of all identical samples.
[0218] According to the loss value corresponding to each target sample, the classification model is trained to obtain a trained classification model.
[0219] In some embodiments, after training the classification model based on the first similarity and the second similarity to obtain the trained classification model, the method further includes:
[0220] Using the trained classification model, obtain the multimedia clips to be identified;
[0221] Extract features from multimedia clips to obtain features corresponding to the multimedia clips;
[0222] The multimedia segments are classified according to the features to obtain the types of the multimedia segments.
[0223] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0224] As can be seen from the above, the model training device of this embodiment obtains a sample set by the acquisition unit, and the sample set includes multiple multimedia samples, and the multimedia samples carry pending labels and pending label types; the sample determination unit determines the same samples and different samples from the sample set according to the pending labels and pending label types carried by the target samples, the same samples are multimedia samples with the same actual labels as the target samples, and the different samples are multimedia samples with different actual labels from the target samples, and the target sample is any multimedia sample in the sample set; the similarity determination unit determines the first similarity and the second similarity, the first similarity is the similarity between the target sample and the same sample, and the second similarity is the similarity between the target sample and the different samples; the model training unit trains the classification model based on the first similarity and the second similarity to obtain the trained classification model.
[0225] Therefore, the embodiments of the present application can efficiently utilize noise samples to train the model and ensure the accuracy of the model prediction.
[0226] The present application also provides an electronic device, which may be a terminal, a server, or the like. The terminal may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or the like; the server may be a single server or a server cluster consisting of multiple servers, or the like.
[0227] In some embodiments, the model training device can also be integrated into multiple electronic devices. For example, the model training device can be integrated into multiple servers, and the model training method of the present application can be implemented by multiple servers.
[0228] In this embodiment, the electronic device of this embodiment is a server as an example for detailed description, for example, Figure 4 As shown, it shows a schematic diagram of the structure of the server involved in the embodiment of the present application, specifically:
[0229] The server may include one or more processing core processors 410, one or more computer-readable storage media memories 420, a power supply 430, an input module 440, and a communication module 450. It will be understood by those skilled in the art that Figure 4 The server structure shown in the figure does not constitute a limitation on the server, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0230] Processor 410 is the server's control center, connecting various components of the server using various interfaces and circuits. It executes software programs and / or modules stored in memory 420 and accesses data stored in memory 420 to perform various server functions and process data. In some embodiments, processor 410 may include one or more processing cores. In some embodiments, processor 410 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 410.
[0231] Memory 420 can be used to store software programs and modules. Processor 410 executes various functional applications and data processing by running the software programs and modules stored in memory 420. Memory 420 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as sound playback or image playback). The data storage area may store data generated based on server usage. Memory 420 may also include high-speed random access memory (RAM) and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state memory device. Accordingly, memory 420 may also include a memory controller to provide processor 410 with access to memory 420.
[0232] The server also includes a power supply 430 that supplies power to various components. In some embodiments, the power supply 430 can be logically connected to the processor 410 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 430 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0233] The server may further include an input module 440, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0234] The server may also include a communication module 450. In some embodiments, the communication module 450 may include a wireless module. The server may use the wireless module of the communication module 450 to perform short-range wireless transmission, thereby providing users with wireless broadband Internet access. For example, the communication module 450 may be used to help users send and receive emails, browse web pages, and access streaming media.
[0235] Although not shown, the server may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 410 in the server will load the executable files corresponding to one or more application processes into the memory 420 according to the following instructions, and the processor 410 will run the application stored in the memory 420 to implement various functions as follows:
[0236] Acquire a sample set, the sample set including multiple multimedia samples, the multimedia samples carrying pending labels and pending label types;
[0237] According to the pending label and pending label type carried by the target sample, the same sample and different sample are determined from the sample set. The same sample is a multimedia sample with the same actual label as the target sample, and the different sample is a multimedia sample with a different actual label from the target sample. The target sample is any multimedia sample in the sample set;
[0238] Determining a first similarity and a second similarity, the first similarity being a similarity between the target sample and the same sample, and the second similarity being a similarity between the target sample and the different sample;
[0239] The classification model is trained based on the first similarity and the second similarity to obtain a trained classification model.
[0240] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0241] From the above, we can see that noise samples can be used efficiently to train the model and ensure the accuracy of model prediction.
[0242] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0243] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any model training method provided in the embodiment of the present application. For example, the instructions can execute the following steps:
[0244] Acquire a sample set, the sample set including multiple multimedia samples, the multimedia samples carrying pending labels and pending label types;
[0245] According to the pending label and pending label type carried by the target sample, the same sample and different sample are determined from the sample set. The same sample is a multimedia sample with the same actual label as the target sample, and the different sample is a multimedia sample with a different actual label from the target sample. The target sample is any multimedia sample in the sample set;
[0246] Determining a first similarity and a second similarity, the first similarity being a similarity between the target sample and the same sample, and the second similarity being a similarity between the target sample and the different sample;
[0247] The classification model is trained based on the first similarity and the second similarity to obtain a trained classification model.
[0248] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0249] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the model training aspects provided in the above embodiments.
[0250] Since the instructions stored in the storage medium can execute the steps in any model training method provided in the embodiments of the present application, the beneficial effects that can be achieved by any model training method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0251] The above is a detailed introduction to a model training method, device and storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A model training method, characterized in that: include: Obtaining a sample set, the sample set comprising a plurality of multimedia samples, the multimedia samples carrying pending labels and pending label types, the multimedia samples in the sample set being any one of a plurality of image frames, a plurality of audio clips, a plurality of video clips, and a plurality of text fields, and the pending label types comprising a correct type and an incorrect type; Determine identical samples and different samples from the sample set according to the pending label and the pending label type carried by the target sample, wherein the identical samples are multimedia samples having the same actual label as the target sample, and the different samples are multimedia samples having different actual labels from the target sample, and the target sample is any multimedia sample in the sample set; Determining a first similarity and a second similarity, the first similarity being a similarity between the target sample and the same sample, and the second similarity being a similarity between the target sample and the different sample; Based on the first similarity and the second similarity, a classification model is trained to obtain a trained classification model.
2. The model training method according to claim 1, wherein: The determining, from the sample set, identical samples and different samples based on the pending label and the pending label type carried by the target sample, includes: Filtering the multimedia samples in the sample set according to the pending label type carried by the multimedia samples to obtain a first set and a second set, wherein the first set includes multimedia samples carrying the pending label type as the correct type, and the second set includes multimedia samples carrying the pending label type as the incorrect type; According to the pending label and the pending label type carried by the target sample, identical samples and different samples are determined from the first set and the second set.
3. The model training method according to claim 2, wherein: The determining, from the sample set, identical samples and different samples based on the pending label and the pending label type carried by the target sample, includes: When the type of the pending label carried by the target sample is the correct type, determining an identical sample from the first set, where the identical sample has the same pending label as the target sample; And different samples are determined from the first set and the second set, where the different samples include multimedia samples in the first set that have different labels from those carried by the target sample, and multimedia samples in the second set that have the same labels as those carried by the target sample.
4. The model training method according to claim 2, wherein: The sample set includes a plurality of sample pairs, the target sample is a multimedia sample in any one of the sample pairs, and determining identical samples and different samples from the sample set according to the pending label and the pending label type carried by the target sample includes: When the undetermined label type carried by the target sample is the error type, the multimedia samples other than the target sample in the target sample pair are regarded as identical samples, and the target sample pair is the sample pair in which the target sample is located; And different samples are determined from the first set, where the different samples are multimedia samples in the first set that have the same undetermined label as the target sample.
5. The model training method according to claim 3, wherein: The sample set includes a plurality of sample pairs, the target sample is a multimedia sample in any one of the sample pairs, and the determining of the same sample from the first set further includes: The multimedia samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located.
6. The model training method according to claim 1, wherein: Determining the first similarity and the second similarity includes: Extracting a target vector, a first vector, and a second vector, wherein the target vector is a vector of the target sample, the first vector is a vector of the same sample, and the second vector is a vector of the different samples; A first similarity is determined based on the target vector and the first vector, and a second similarity is determined based on the target vector and the second vector.
7. The model training method according to claim 1, wherein: The step of training a classification model based on the first similarity and the second similarity to obtain a trained classification model includes: Adding all the first similarities and all the second similarities to obtain a total similarity; determining a quotient of each first similarity divided by the total similarity; determining the logarithm of each of said quotients; Dividing the sum of all the logarithms by the determinant of a first matrix to obtain a loss value corresponding to the target sample, where the first matrix includes vectors of all the same samples; The classification model is trained according to the loss value corresponding to each target sample to obtain a trained classification model.
8. The model training method according to claim 1, wherein: After training the classification model based on the first similarity and the second similarity to obtain the trained classification model, the method further includes: Using the trained classification model, obtaining a multimedia segment to be identified; Performing feature extraction on the multimedia segment to obtain features corresponding to the multimedia segment; The multimedia segment is classified according to the feature to obtain the type of the multimedia segment.
9. A model training device, characterized in that: include: an acquisition unit, configured to acquire a sample set, the sample set comprising a plurality of multimedia samples, the multimedia samples carrying pending labels and pending label types, the multimedia samples in the sample set being any one of a plurality of image frames, a plurality of audio clips, a plurality of video clips, and a plurality of text fields, and the pending label types comprising a correct type and an incorrect type; a sample determination unit, configured to determine identical samples and different samples from the sample set based on the pending label and the pending label type carried by the target sample, wherein the identical sample is a multimedia sample having the same actual label as the target sample, and the different sample is a multimedia sample having a different actual label from the target sample, and the target sample is any multimedia sample in the sample set; a similarity determination unit, configured to determine a first similarity and a second similarity, wherein the first similarity is the similarity between the target sample and the same sample, and the second similarity is the similarity between the target sample and the different sample; The model training unit is used to train the classification model based on the first similarity and the second similarity to obtain a trained classification model.
10. The model training device according to claim 9, characterized in that: The determining, from the sample set, identical samples and different samples based on the pending label and the pending label type carried by the target sample, includes: Filtering the multimedia samples in the sample set according to the pending label type carried by the multimedia samples to obtain a first set and a second set, wherein the first set includes multimedia samples carrying the pending label type as the correct type, and the second set includes multimedia samples carrying the pending label type as the incorrect type; According to the pending label and the pending label type carried by the target sample, identical samples and different samples are determined from the first set and the second set.
11. The model training device according to claim 10, wherein: The determining, from the sample set, identical samples and different samples based on the pending label and the pending label type carried by the target sample, includes: When the type of the pending label carried by the target sample is the correct type, determining an identical sample from the first set, where the identical sample has the same pending label as the target sample; And different samples are determined from the first set and the second set, where the different samples include multimedia samples in the first set that have different labels from those carried by the target sample, and multimedia samples in the second set that have the same labels as those carried by the target sample.
12. The model training device according to claim 10, wherein: The sample set includes a plurality of sample pairs, the target sample is a multimedia sample in any one of the sample pairs, and determining identical samples and different samples from the sample set according to the pending label and the pending label type carried by the target sample includes: When the undetermined label type carried by the target sample is the error type, the multimedia samples other than the target sample in the target sample pair are regarded as identical samples, and the target sample pair is the sample pair in which the target sample is located; And different samples are determined from the first set, where the different samples are multimedia samples in the first set that have the same undetermined label as the target sample.
13. The model training device according to claim 11, wherein: The sample set includes a plurality of sample pairs, the target sample is a multimedia sample in any one of the sample pairs, and the determining of the same sample from the first set further includes: The multimedia samples other than the target sample in the target sample pair are regarded as the same samples, and the target sample pair is the sample pair in which the target sample is located.
14. The model training device according to claim 9, wherein: Determining the first similarity and the second similarity includes: Extracting a target vector, a first vector, and a second vector, wherein the target vector is a vector of the target sample, the first vector is a vector of the same sample, and the second vector is a vector of the different samples; A first similarity is determined based on the target vector and the first vector, and a second similarity is determined based on the target vector and the second vector.
15. The model training device according to claim 9, wherein: The step of training a classification model based on the first similarity and the second similarity to obtain a trained classification model includes: Adding all the first similarities and all the second similarities to obtain a total similarity; determining a quotient of each first similarity divided by the total similarity; determining the logarithm of each of said quotients; Dividing the sum of all the logarithms by the determinant of a first matrix to obtain a loss value corresponding to the target sample, where the first matrix includes vectors of all the same samples; The classification model is trained according to the loss value corresponding to each target sample to obtain a trained classification model.
16. The model training device according to claim 9, wherein: After training the classification model based on the first similarity and the second similarity to obtain the trained classification model, the method further includes: Using the trained classification model, obtaining a multimedia segment to be identified; Performing feature extraction on the multimedia segment to obtain features corresponding to the multimedia segment; The multimedia segment is classified according to the feature to obtain the type of the multimedia segment.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores multiple instructions, which are suitable for loading by a processor to execute the steps in the model training method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Data classification method and device and electronic equipment
CN113705598A
Garbage picture classification method and device
CN113989567A