Classification model training method and device, equipment, storage medium and program product
By using different loss determination methods for samples to be classified for different categories of labels, the classification model is trained, which solves the problem of inaccurate loss values in the existing technology and improves the performance of the classification model.
Patent Information
- Application Number
- CN202410018422.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, a single loss calculation method is used during the training of classification model, resulting in inaccurate loss values, resulting in poor classification performance.
By obtaining multiple samples to be classified and their different categories of labels, the target loss determination method is determined, and the classification model is trained using different loss determination methods, including the loss determination method for the target label and reference label.
The accuracy of the loss value is improved, thereby improving the classification performance of the target classification model obtained by training.
Smart Images

Figure CN120257033A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, device, equipment, storage medium, and program product for training a classification model. Background Art
[0002] Artificial Intelligence (AI) is a comprehensive discipline with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0003] In related technologies, for the training of a classification model, usually a single loss calculation method is used for various different samples to determine the loss value for training the classification model, and then the classification model is trained through the loss value. Due to the single loss calculation method, the loss values determined by different samples are inaccurate, resulting in poor classification performance of the trained classification model. Summary of the Invention
[0004] The embodiments of this application provide a method, object classification method, device, electronic device, computer-readable storage medium, and computer program product for training a classification model, which can effectively improve the classification performance of the trained target classification model.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] The embodiments of this application provide a method for training a classification model, including:
[0007] Obtain a plurality of samples to be classified and the category labels carried by each of the samples to be classified, where the category labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different;
[0008] Call the classification model to respectively perform category prediction on each of the samples to be classified, and obtain the predicted category corresponding to each of the samples to be classified;
[0009] For each of the samples to be classified, based on the class label carried by the sample to be classified, determine the target loss determination method corresponding to the sample to be classified, and based on the target loss determination method and the predicted class corresponding to the sample to be classified, determine the loss value corresponding to the sample to be classified;
[0010] Based on the loss values respectively corresponding to each of the samples to be classified, train the classification model to obtain a target classification model.
[0011] An embodiment of the present application provides an object classification method, including:
[0012] Obtain an object to be classified, and object description information for describing the object to be classified, and perform feature extraction on the object description information to obtain the object features of the object to be classified;
[0013] Call the target classification model, and based on the object features of the object to be classified, perform class prediction on the object to be classified to obtain the object class corresponding to the object to be classified;
[0014] Wherein, the object class is used to indicate whether the object to be classified is a normal object, the target classification model is obtained by training a classification model based on samples of objects to be classified carrying different class labels, the class labels include a target label and a reference label, and the loss determination methods respectively corresponding to the target label and the reference label are different.
[0015] An embodiment of the present application provides a training device for a classification model, including:
[0016] An acquisition module, configured to acquire a plurality of samples to be classified, and class labels respectively carried by each of the samples to be classified, the class labels include a target label and a reference label, and the loss determination methods respectively corresponding to the target label and the reference label are different;
[0017] A prediction module, configured to call the classification model to perform class prediction on each of the samples to be classified respectively, to obtain the predicted classes respectively corresponding to each of the samples to be classified;
[0018] A loss module, configured to, for each of the samples to be classified, based on the class label carried by the sample to be classified, determine the target loss determination method corresponding to the sample to be classified, and based on the target loss determination method and the predicted class corresponding to the sample to be classified, determine the loss value corresponding to the sample to be classified;
[0019] A training module, configured to train the classification model based on the loss values respectively corresponding to each of the samples to be classified, to obtain a target classification model.
[0020] In the above solution, the above loss module is further configured to obtain a target mapping relationship between multiple preset class labels and corresponding loss determination methods, and compare each of the class labels with each of the preset class labels in the target mapping relationship to obtain a label comparison result corresponding to the class label; when the label comparison result indicates that the class label exists among the multiple preset class labels, determine the loss determination method corresponding to the class label in the target mapping relationship as the target loss determination method.
[0021] In the above solution, the above target label includes multiple sub-target labels, and the above loss module is further configured to obtain prediction probabilities of the sample to be classified corresponding to each of the sub-target labels, and the predicted class corresponding to the sample to be classified is the class indicated by the sub-target label with the highest prediction probability; when the target loss determination method is the loss determination method corresponding to the target label, determine a target prediction probability corresponding to the sample to be classified based on each of the prediction probabilities, and determine the negative value of the logarithm of the target prediction probability as the loss value corresponding to the sample to be classified; when the target loss determination method is the loss determination method corresponding to the reference label, determine a reference prediction probability corresponding to the sample to be classified based on each of the prediction probabilities, and determine the loss value corresponding to the sample to be classified based on the reference prediction probability; wherein, the sum of the reference prediction probability and the target prediction probability is equal to 1.
[0022] In the above solution, the above sub-target labels include a positive target label and a negative target label, and the prediction probabilities include a first prediction probability corresponding to the positive target label and a second prediction probability corresponding to the negative target label; the above loss module is further configured to sum the first prediction probability and the second prediction probability to obtain a sum probability; divide the exponential value of the second prediction probability by the exponential value of the sum probability to obtain the reference prediction probability corresponding to the sample to be classified.
[0023] In the above solution, the above loss module is further configured to compare the reference prediction probability with a reference probability threshold to obtain a probability comparison result; when the probability comparison result indicates that the reference prediction probability is less than the reference probability threshold, determine the loss value corresponding to the sample to be classified as zero; when the probability comparison result indicates that the reference prediction probability is greater than or equal to the reference probability threshold, determine the loss value corresponding to the sample to be classified in combination with the reference prediction probability and the reference probability threshold.
[0024] In the above solution, the above loss module is further configured to subtract the reference probability threshold from the reference prediction probability to obtain the loss value corresponding to the sample to be classified; or, determine the nth power of the difference between the reference prediction probability and the reference probability threshold as the loss value corresponding to the sample to be classified; where n is a positive integer greater than or equal to 2.
[0025] In the above solution, the classification model includes a conversion layer, a feature extraction layer, and a classification layer. The above prediction module is further configured to perform the following processing for each of the samples to be classified: call the conversion layer to perform enhancement conversion on the sample to be classified to obtain an enhanced sample corresponding to the sample to be classified; call the feature extraction layer to perform feature extraction on the enhanced sample to obtain a sample feature corresponding to the sample to be classified; call the classification layer to perform class prediction on the sample to be classified based on the sample feature to obtain a predicted class corresponding to the sample to be classified.
[0026] In the above solution, the enhancement conversion includes a partitioning process and a fusion process. The above prediction module is further configured to perform a partitioning process on the sample to be classified to obtain a plurality of sub-samples in the sample to be classified, and obtain enhanced sub-samples corresponding to each of the sub-samples; perform a fusion process on each of the enhanced sub-samples and each of the sub-samples to obtain an enhanced sample corresponding to the sample to be classified.
[0027] In the above solution, the above prediction module is further configured to obtain a mapping relationship between a plurality of preset samples and corresponding preset enhanced samples, and perform the following processing for each of the sub-samples: compare each of the sub-samples with each of the preset samples to obtain a comparison result corresponding to the sub-sample; when the comparison result indicates that the sub-sample exists in the preset samples, determine the preset enhanced sample corresponding to the sub-sample in the mapping relationship as the enhanced sub-sample of the sub-sample.
[0028] In the above solution, the above prediction module is further configured to call a first feature extraction layer to perform feature extraction on the enhanced sample to obtain a first sample feature; traverse i and perform the following processing: call an ith feature extraction layer to perform feature extraction on the enhanced sample based on the (i - 1)th sample feature to obtain an ith sample feature, where 2 ≤ i ≤ N; determine the Nth sample feature as the sample feature corresponding to the sample to be classified.
[0029] An object classification device provided by an embodiment of the present application includes:
[0030] A feature extraction module, configured to obtain an object to be classified, and object description information for describing the object to be classified, and perform feature extraction on the object description information to obtain an object feature of the object to be classified;
[0031] A classification module, configured to call a target classification model, and based on the object features of the object to be classified, predict the category of the object to be classified, so as to obtain the object category corresponding to the object to be classified; wherein, the object category is used to indicate whether the object to be classified is a normal object. The target classification model is obtained by training a classification model based on samples of objects to be classified carrying different category labels, and the category labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different.
[0032] An embodiment of the present application provides an electronic device, including:
[0033] A memory, configured to store computer-executable instructions or a computer program;
[0034] A processor, configured to implement the training method of the classification model provided by the embodiment of the present application when executing the computer-executable instructions or the computer program stored in the memory.
[0035] An embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, and when executed by a processor, implementing the training method of the classification model provided by the embodiment of the present application.
[0036] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the training method of the classification model described above in the embodiment of the present application.
[0037] The embodiment of the present application has the following beneficial effects:
[0038] By obtaining multiple samples to be classified carrying different category labels, calling a classification model, predicting the category of each sample to be classified respectively, obtaining the predicted category corresponding to each sample to be classified, for each sample to be classified, based on the category label carried by the sample to be classified, determining the target loss determination method corresponding to the sample to be classified, and based on the target loss calculation method and the predicted category corresponding to the sample to be classified, determining the loss value of the sample to be classified, and based on the loss values corresponding to each sample to be classified respectively, training the classification model to obtain a target classification model. In this way, for each sample to be classified, based on the category label carried by the sample to be classified, determining the target loss determination method corresponding to the sample to be classified, and based on the target loss calculation method and the predicted category corresponding to the sample to be classified, determining the loss value of the sample to be classified, so that for samples to be classified with different category labels, different loss determination methods are used to determine the loss values of the corresponding samples to be classified, thereby effectively improving the accuracy of the loss values, and training the classification model with accurate loss values, thereby effectively improving the classification performance of the obtained target classification model. Description of the Drawings
[0039] Figure 1 is a schematic structural diagram of a training system for a classification model provided by an embodiment of the present application;
[0040] Figure 2 is a schematic structural diagram of an electronic device for training a classification model provided by an embodiment of the present application;
[0041] Figure 3 is a schematic structural diagram of an electronic device for object classification provided by an embodiment of the present application;
[0042] Figure 4 is a schematic flowchart of a method for training a classification model provided by an embodiment of the present application Figure 1 ;
[0043] Figure 5 is a schematic flowchart of an object classification method provided by an embodiment of the present application Figure 2 ;
[0044] Figure 6 is a schematic principle diagram of a method for training a classification model provided by an embodiment of the present application;
[0045] Figure 7 is a schematic flowchart of a method for training a classification model provided by an embodiment of the present application Figure 3 ;
[0046] Figure 8 is a schematic flowchart of a method for training a classification model provided by an embodiment of the present application Figure 4 ;
[0047] Figure 9Schematic flowchart of the training method of the classification model provided by the embodiments of the present application Figure 5 ;
[0048] Figure 10 Schematic flowchart of the training method of the classification model provided by the embodiments of the present application Figure 6 。 Detailed implementation manners
[0049] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0050] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0051] In the following description, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0053] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0054] 1) Artificial Intelligence (AI): It is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include, for example, sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, pre-trained model technologies, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0055] 2) Machine Learning (ML): It is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.
[0056] 3) Convolutional Neural Networks (CNN): It is a type of feed-forward neural network (FNN) with convolutional calculations and a deep structure, and it is one of the representative algorithms of deep learning. Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on input images according to their hierarchical structure.
[0057] 4) In response to: It is used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or have a set delay; without special instructions, there is no restriction on the execution order of multiple executed operations.
[0058] 5) Convolutional Neural Networks (CNN): It is a type of feed-forward neural network (FNN) with convolutional calculations and a deep structure, and it is one of the representative algorithms of deep learning. Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on input images according to their hierarchical structure.
[0059] 6) Convolutional Layer: Each convolutional layer in a convolutional neural network consists of several convolutional units, and the parameters of each convolutional unit are optimized through the backpropagation algorithm. The purpose of the convolution operation is to extract different features of the input. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners. More layers of the network can iteratively extract more complex features from the low-level features.
[0060] 7) Pooling Layer: After feature extraction in the convolutional layer, the output feature map is passed to the pooling layer for feature selection and information filtering. The pooling layer contains a pre-set pooling function, whose function is to replace the result of a single point in the feature map with the statistical quantity of the feature map in its adjacent area. The pooling layer selects the pooling area in the same way as the convolutional kernel scans the feature map, which is controlled by the pooling size, stride, and padding.
[0061] 8) Fully-Connected Layer: The fully-connected layer in a convolutional neural network is equivalent to the hidden layer in a traditional feedforward neural network. The fully-connected layer is located at the last part of the hidden layer of the convolutional neural network and only transmits signals to other fully-connected layers. The spatial topology structure of the feature map is lost in the fully-connected layer, and it is unfolded into a vector and passed through an activation function.
[0062] 9) Mutual Information: It is a concept in information theory used to measure the mutual dependence or correlation between two random variables. Mutual Information: It is a concept in information theory used to measure the mutual dependence or correlation between two random variables. It is usually used to measure the amount of information provided by the value of one random variable about the value of another random variable.
[0063] During the implementation process of the embodiments of the present application, the applicant found the following problems in the related art:
[0064] In the related art, for the training of a classification model, usually a single loss calculation method is used for various different samples to determine the loss value for training the classification model, and then the classification model is trained through the loss value. Due to the single loss calculation method, the loss values determined by different samples are inaccurate, resulting in poor classification performance of the trained classification model.
[0065] The embodiments of the present application provide a training method, device, electronic device, computer-readable storage medium, and computer program product for a classification model, which can effectively improve the classification performance of the trained target classification model. The following describes an exemplary application of the classification model training system provided by the embodiments of the present application.
[0066] See Figure 1 ,Figure 1 It is a schematic architecture diagram of a training system 100 for a classification model provided by an embodiment of the present application. A terminal (exemplarily showing terminal 400) is connected to a server 200 through a network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0067] The terminal 400 is used for a user to use a client 410 to display a classification result on a graphical interface 410-1 (exemplarily showing graphical interface 410-1). The terminal 400 and the server 200 are connected to each other through a wired or wireless network.
[0068] In some embodiments, the server 200 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The electronic device provided by the embodiment of the present application can be implemented as a terminal or as a server. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0069] In some embodiments, the terminal 400 invokes a classification model to respectively perform class prediction on each sample to be classified, obtains the predicted class corresponding to each sample to be classified, and for each sample to be classified, determines a target loss determination method based on the class label carried by the sample to be classified, and determines the loss value corresponding to the sample to be classified based on the target loss determination method and the predicted class, and sends the loss value to the server 200. The server 200 trains the classification model based on the loss value to obtain a target classification model.
[0070] In other embodiments, the server 200 invokes a classification model to respectively perform class prediction on each sample to be classified, obtains the predicted class corresponding to each sample to be classified, and for each sample to be classified, determines a target loss determination method based on the class label carried by the sample to be classified, and determines the loss value corresponding to the sample to be classified based on the target loss determination method and the predicted class, and sends the loss value to the terminal 400. The terminal 400 trains the classification model based on the loss value to obtain a target classification model.
[0071] In some other embodiments, the embodiments of the present application can be implemented by means of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing.
[0072] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources.
[0073] See Figure 2 , Figure 2 which is a schematic structural diagram of an electronic device 500 provided by the embodiments of the present application. Among them, Figure 2 the shown electronic device 500 can be Figure 1 the server 200 or the terminal 400 in Figure 2 The shown electronic device 500 includes: at least one processor 430, a memory 450, and at least one network interface 420. Each component in the electronic device 500 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.
[0074] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0075] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located far from the processor 430.
[0076] The memory 450 includes volatile memory, non-volatile memory, or both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0077] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.
[0078] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0079] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.
[0080] In some embodiments, the training device of the classification model provided in the embodiments of the present application can be implemented in software. Figure 2 Shown is the training device 455 of the classification model stored in the memory 450, which can be software in the form of programs and plugins, etc., and includes the following software modules: an acquisition module 4551, a prediction module 4552, a loss module 4553, and a training module 4554. These modules are logical, and thus can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.
[0081] See Figure 3 , Figure 3 is a schematic structural diagram of the electronic device 600 for object classification provided by the embodiments of the present application, where Figure 3 the shown electronic device 600 can be Figure 1 the server 200 or the terminal 400 in Figure 3The electronic device 600 shown includes: at least one processor 530, a memory 550, and at least one network interface 520. Each component in the electronic device 600 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3 all kinds of buses are labeled as the bus system 540.
[0082] The processor 530 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0083] The memory 550 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. Optionally, the memory 550 includes one or more storage devices that are physically located far from the processor 530.
[0084] The memory 550 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM, Read Only Memory), and the volatile memory can be a random access memory (RAM, Random Access Memory). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0085] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.
[0086] An operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0087] A network communication module 552, for reaching other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.
[0088] In some embodiments, the object classification device provided by the embodiments of the present application may be implemented in software. Figure 3 FIG. shows an object classification device 555 stored in a memory 550, which may be software in the form of a program and a plug-in, etc., including the following software modules: a feature extraction module 5551 and a classification module 5552. These modules are logical, and thus can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.
[0089] In other embodiments, the training device of the classification model provided by the embodiments of the present application may be implemented in hardware. As an example, the training device of the classification model provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the classification model provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0090] In some embodiments, a terminal or a server may implement the training method of the classification model provided by the embodiments of the present application by running a computer program or computer executable instructions. For example, the computer program may be a native program in an operating system (e.g., a dedicated training program for a classification model) or a software module. For example, it may be a training module for a classification model that can be embedded in any program (such as an instant messaging client, a photo album program, an electronic map client, a navigation client); for example, it may be a native application (APP), that is, a program that needs to be installed in an operating system to run. In short, the above computer program may be any form of application program, module, or plug-in.
[0091] The training method of the classification model provided by the embodiments of the present application will be described in conjunction with the exemplary applications and implementations of the server or terminal provided by the embodiments of the present application.
[0092] See Figure 4 , Figure 4 is a flowchart of the training method of the classification model provided by the embodiments of the present application Figure 1 will be described in conjunction with Figure 4The steps 101 to 104 shown will be described. The training method of the classification model provided by the embodiments of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in cooperation. Hereinafter, the case where the server implements it alone will be taken as an example for description.
[0093] In step 101, a plurality of samples to be classified are obtained, as well as the category labels carried by each sample to be classified respectively.
[0094] In some embodiments, the category label includes a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different.
[0095] In some embodiments, the above samples to be classified can be samples in various expression forms such as text samples to be classified, video samples to be classified, audio samples to be classified, object samples to be classified, etc.
[0096] As an example, in the application scenario of text classification, the above samples to be classified can be text samples to be classified. Then, the above classification model can be a text classification model. Thus, by obtaining a plurality of text samples to be classified and the category labels carried by each text sample to be classified respectively, and calling the text classification model to perform category prediction on each text sample to be classified respectively, the predicted category corresponding to each text sample to be classified is obtained. For each text sample to be classified, based on the category label carried by the text sample to be classified, the target loss determination method corresponding to the text sample to be classified is determined, and based on the target loss determination method and the predicted category corresponding to the text sample to be classified, the loss value corresponding to the text sample to be classified is determined, and based on the loss value, the text classification model is trained to obtain the target text classification model.
[0097] As an example, in the application scenario of video classification, the above samples to be classified can be video samples to be classified. Then, the above classification model can be a video classification model. Thus, by obtaining a plurality of video samples to be classified and the category labels carried by each video sample to be classified respectively, and calling the video classification model to perform category prediction on each video sample to be classified respectively, the predicted category corresponding to each video sample to be classified is obtained. For each video sample to be classified, based on the category label carried by the video sample to be classified, the target loss determination method corresponding to the video sample to be classified is determined, and based on the target loss determination method and the predicted category corresponding to the video sample to be classified, the loss value corresponding to the video sample to be classified is determined, and based on the loss value, the video classification model is trained to obtain the target video classification model.
[0098] As an example, in the application scenario of object classification, the above-mentioned samples to be classified can be object samples to be classified, and then the above-mentioned classification model can be an object classification model. Thus, by obtaining multiple object samples to be classified and the class labels carried by each object sample to be classified, and calling the object classification model to respectively perform class prediction on the object samples to be classified, the predicted classes corresponding to each object sample to be classified are obtained. For each object sample to be classified, based on the class label carried by the object sample to be classified, a target loss determination method corresponding to the object sample to be classified is determined, and based on the target loss determination method and the predicted class corresponding to the object sample to be classified, the loss value corresponding to the object sample to be classified is determined, and based on the loss value, the object classification model is trained to obtain a target object classification model.
[0099] In some embodiments, the above-mentioned class labels include a target label and a reference label. The classes indicated by the target label and the reference label are different respectively. The target label includes multiple sub-target labels, and the classes indicated by each sub-target label are different respectively.
[0100] As an example, when the class indicated by the above-mentioned target label is the positive class and the negative class in the binary classification scenario, then the sub-target labels include a positive target label and a negative target label. Specifically, in the application scenario of anomaly classification, the class indicated by the positive target label can be the normal class, and the class indicated by the negative target label can be the abnormal class. In the application scenario of object classification, the class indicated by the reference label can be a fuzzy class, and the objects of the fuzzy class can be objects in the normal class that tend to evolve into the abnormal class, or objects in the abnormal class that tend to evolve into the normal class.
[0101] In some embodiments, the classes indicated by the target label and the reference label are different respectively. The target label includes multiple sub-target labels, and the reference label includes at least one sub-reference label, and the classes indicated by each sub-reference label are different respectively.
[0102] In step 102, the classification model is called to respectively perform class prediction on each sample to be classified, and the predicted classes corresponding to each sample to be classified are obtained.
[0103] In some embodiments, the above-mentioned classification model includes a conversion layer, a feature extraction layer, and a classification layer. Refer to Figure 6 , Figure 6 which is a schematic diagram of the principle of the training method of the classification model provided by the embodiments of the present application. Figure 6 The classification model shown includes a conversion layer 41, a feature extraction layer 42, and a classification layer 43.
[0104] In some embodiments, refer to Figure 7 , Figure 7Schematic flowchart of the training method of the classification model provided by the embodiments of the present application Figure 3 , Figure 4 The step 102 shown can be executed for each sample to be classified Figure 7 and is implemented by the steps 1021 to 1023 shown
[0105] In step 1021, a conversion layer is called to perform enhancement conversion on the sample to be classified, and an enhanced sample corresponding to the sample to be classified is obtained
[0106] In some embodiments, the above enhancement conversion is used to enhance the information volume of the object to be processed, and the information volume of the enhanced sample of the sample to be classified is greater than that of the sample to be classified. The above enhancement conversion includes partitioning processing and fusion processing
[0107] In some embodiments, the above step 1021 can be implemented in the following manner: the sample to be classified is subjected to partitioning processing to obtain multiple sub-samples in the sample to be classified, and enhanced sub-samples corresponding to the sub-samples are obtained; the enhanced sub-samples and the sub-samples are subjected to fusion processing to obtain an enhanced sample corresponding to the sample to be classified
[0108] As an example, the expression of the sample to be classified can be: A1A2A3A4A5, the expressions of multiple sub-samples in the sample to be classified can be: sub-sample A1, sub-sample A2, sub-sample A3, sub-sample A4, and sub-sample A5, the enhanced sub-sample corresponding to sub-sample A1 is enhanced sub-sample B1, the enhanced sub-sample corresponding to sub-sample A2 is enhanced sub-sample B2, the enhanced sub-sample corresponding to sub-sample A3 is enhanced sub-sample B3, the enhanced sub-sample corresponding to sub-sample A4 is enhanced sub-sample B4, and the enhanced sub-sample corresponding to sub-sample A5 is enhanced sub-sample B5
[0109] Continuing with the above example, the expression of the enhanced sample corresponding to the above sample to be classified can be
[0110] Z = {A1, A2, A3, A4, A5, B1, B2, B3, B4, B5} (1)
[0111] In some embodiments, the above obtaining the enhanced sub-samples corresponding to the sub-samples can be implemented in the following manner: obtaining the mapping relationship between multiple preset samples and the corresponding preset enhanced samples, and performing the following processing for each sub-sample respectively: comparing the sub-sample with each preset sample to obtain the comparison result corresponding to the sub-sample; when the comparison result indicates that the preset sample exists in the sub-sample, the preset enhanced sample corresponding to the sub-sample in the mapping relationship is determined as the enhanced sub-sample of the sub-sample
[0112] As an example, for sub-sample A1, sub-sample A1 is compared with preset sample A1, preset sample A2, and preset sample A3 respectively to obtain the comparison results corresponding to sub-sample A1; when the comparison results indicate that sub-sample A1 exists in preset sample A1, preset sample A2, and preset sample A3, the preset enhanced sample corresponding to preset sample A1 in the mapping relationship is determined as the enhanced sub-sample of the sub-sample.
[0113] In step 1022, the feature extraction layer is called to extract features from the enhanced sample to obtain the sample features corresponding to the sample to be classified.
[0114] In some embodiments, the above step 1022 can be implemented in the following manner: call the first feature extraction layer to extract features from the enhanced sample to obtain the first sample features; traverse i and perform the following processing: call the i-th feature extraction layer to extract features from the enhanced sample based on the (i - 1)-th sample features to obtain the i-th sample features, where 2 ≤ i ≤ N; determine the N-th sample features as the sample features corresponding to the sample to be classified.
[0115] As an example, when N is equal to 3, call the first feature extraction layer to extract features from the enhanced sample to obtain the first sample features, call the second feature extraction layer to extract features from the enhanced sample based on the first sample features to obtain the second sample features, call the third feature extraction layer to extract features from the enhanced sample based on the second sample features to obtain the third sample features, and determine the third sample features as the sample features corresponding to the sample to be classified.
[0116] In some embodiments, the above i-th feature extraction layer includes a convolutional layer. Each convolutional layer in the convolutional neural network is composed of several convolutional units, and the parameters of each convolutional unit are optimized through the backpropagation algorithm. The purpose of the convolutional operation is to extract different features of the input. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners, etc. More layers of the network can iteratively extract more complex features from the low-level features.
[0117] In step 1023, the classification layer is called to predict the category of the sample to be classified based on the sample features to obtain the predicted category corresponding to the sample to be classified.
[0118] In some embodiments, the above calling the classification layer to predict the category of the sample to be classified based on the sample features to obtain the predicted category corresponding to the sample to be classified can be implemented in the following manner: call the classification layer to predict the category of the sample to be classified based on the sample features to obtain the predicted probabilities of each category corresponding to the sample to be classified, and determine the category with the maximum predicted probability as the predicted category corresponding to the sample to be classified.
[0119] As an example, the classification layer is called, and based on the sample features, class prediction is performed on the sample to be classified, obtaining a prediction probability of 0.2 for class A corresponding to the sample to be classified and a prediction probability of 0.3 for class B. The class B with the maximum prediction probability is determined as the predicted class corresponding to the sample to be classified.
[0120] In this way, by calling the conversion layer, enhanced conversion is performed on the sample to be classified to obtain an enhanced sample with more information than the sample to be classified. The feature extraction layer is called to extract features from the enhanced sample to obtain the sample features corresponding to the sample to be classified. The classification layer is called, and based on the sample features, class prediction is performed on the sample to be classified to obtain the predicted class corresponding to the sample to be classified. Thus, through enhanced conversion, the information content of the sample to be classified is significantly enhanced, and based on the enhanced sample with more information, class prediction is performed, thereby significantly improving the accuracy of the predicted class obtained by the prediction.
[0121] In step 103, for each sample to be classified, based on the class label carried by the sample to be classified, the target loss determination method corresponding to the sample to be classified is determined.
[0122] In some embodiments, referring to Figure 8 , Figure 8 is the flowchart of the training method of the classification model provided by the embodiments of the present application Figure 4 , Figure 4 The step 103 shown can be executed separately for each sample to be classified Figure 8 The steps 1031 to 1033 shown are implemented.
[0123] In step 1031, the target mapping relationship between multiple preset class labels and the corresponding loss determination methods is obtained.
[0124] As an example, the above target mapping relationship can be seen in Table 1 shown below. Table 1 is a schematic table of the target mapping relationship provided by the embodiments of the present application.
[0125] Table 1 Schematic Table of Target Mapping Relationship
[0126] Preset category label Loss determination method Category label A Loss determination method A Category label B Loss determination method B
[0127] In step 1032, the class label is respectively compared with each preset class label in the target mapping relationship to obtain the label comparison result corresponding to the class label.
[0128] In some embodiments, the above label comparison result is used to indicate whether there is a class label among multiple preset class labels.
[0129] In step 1033, when the label comparison result indicates that there is a category label among multiple preset category labels, the loss determination method corresponding to the category label in the target mapping relationship is determined as the target loss determination method.
[0130] Continuing with the above example, when the label comparison result indicates that there is category label A among multiple preset category labels, the loss determination method A corresponding to category label A in the target mapping relationship is determined as the target loss determination method.
[0131] Continuing with the above example, when the label comparison result indicates that there is category label B among multiple preset category labels, the loss determination method B corresponding to category label B in the target mapping relationship is determined as the target loss determination method.
[0132] In this way, for each sample to be classified, based on the category label carried by the sample to be classified, the target loss determination method corresponding to the sample to be classified is determined, and based on the target loss calculation method and the predicted category corresponding to the sample to be classified, the loss value of the sample to be classified is determined. Thus, for samples to be classified with different category labels, different loss determination methods are used to determine the loss values of the corresponding samples to be classified, effectively improving the accuracy of the loss value. By training the classification model with the accurate loss value, the classification performance of the obtained target classification model is effectively improved.
[0133] In step 104, based on the target loss determination method and the predicted category corresponding to the sample to be classified, the loss value corresponding to the sample to be classified is determined.
[0134] In some embodiments, referring to Figure 9 , Figure 9 is the flowchart of the training method of the classification model provided by the embodiments of the present application Figure 5 , the above target label includes multiple sub-target labels, Figure 4 the step 104 shown can be executed for each sample to be classified separately Figure 8 The steps 1041 to 1043 shown are implemented.
[0135] In step 1041, the prediction probabilities of the sample to be classified corresponding to each sub-target label are obtained.
[0136] In some embodiments, the predicted category corresponding to the sample to be classified is the category indicated by the sub-target label with the highest prediction probability.
[0137] As an example, the target label includes sub-target label A and sub-target label B. The prediction probability of the sample to be classified corresponding to sub-target label A is 0.2, and the prediction probability of the sample to be classified corresponding to sub-target label B is 0.5. The predicted category corresponding to the sample to be classified is the category indicated by sub-target label B.
[0138] In step 1042, when the target loss determination method is the loss determination method corresponding to the target label, based on each prediction probability, determine the target prediction probability corresponding to the sample to be classified, and determine the negative value of the logarithm of the target prediction probability as the loss value corresponding to the sample to be classified.
[0139] As an example, the expression of the loss value corresponding to the sample to be classified above can be:
[0140] L1 = -log(p i ) (2)
[0141] where L1 is used to indicate the loss value corresponding to the sample to be classified, and p i is used to indicate the target prediction probability.
[0142] In some embodiments, the sub-target labels include a positive target label and a negative target label, and the prediction probabilities include a first prediction probability corresponding to the positive target label and a second prediction probability corresponding to the negative target label.
[0143] In some embodiments, the above-mentioned determining the target prediction probability corresponding to the sample to be classified based on each prediction probability can be implemented by the following method: sum the first prediction probability and the second prediction probability to obtain a sum probability; divide the exponential value of the first prediction probability by the exponential value of the sum probability to obtain the target prediction probability corresponding to the sample to be classified.
[0144] As an example, the expression of the above-mentioned target prediction probability can be:
[0145]
[0146] where p1 is used to indicate the first prediction probability, p2 is used to indicate the second prediction probability, p1 + p2 is used to indicate the sum probability, and p i is used to indicate the target prediction probability corresponding to the sample to be classified.
[0147] In step 1043, when the target loss determination method is the loss determination method corresponding to the reference label, based on each prediction probability, determine the reference prediction probability corresponding to the sample to be classified, and based on the reference prediction probability, determine the loss value corresponding to the sample to be classified.
[0148] In some embodiments, the sum of the reference prediction probability and the target prediction probability is equal to 1.
[0149] In some embodiments, the sub-target labels include a positive target label and a negative target label, and the prediction probabilities include a first prediction probability corresponding to the positive target label and a second prediction probability corresponding to the negative target label.
[0150] In some embodiments, based on the respective prediction probabilities, to determine the reference prediction probability corresponding to the sample to be classified, it can be achieved in the following manner: sum the first prediction probability and the second prediction probability to obtain a sum probability; divide the exponential value of the second prediction probability by the exponential value of the sum probability to obtain the reference prediction probability corresponding to the sample to be classified.
[0151] As an example, the expression of the above reference prediction probability can be:
[0152]
[0153] where p1 is used to indicate the first prediction probability, p2 is used to indicate the second prediction probability, p1 + p2 is used to indicate the sum probability, and p f is used to indicate the reference prediction probability corresponding to the sample to be classified.
[0154] In some embodiments, based on the above reference prediction probability, to determine the loss value corresponding to the sample to be classified, it can be achieved in the following manner: compare the reference prediction probability with a reference probability threshold to obtain a probability comparison result; when the probability comparison result indicates that the reference prediction probability is less than the reference probability threshold, determine the loss value corresponding to the sample to be classified as zero; when the probability comparison result indicates that the reference prediction probability is greater than or equal to the reference probability threshold, combine the reference prediction probability and the reference probability threshold to determine the loss value corresponding to the sample to be classified.
[0155] In some embodiments, the above probability comparison result is used to indicate whether the reference prediction probability is less than the reference probability threshold.
[0156] As an example, when the target loss determination method is the loss determination method corresponding to the reference label, the expression of the loss value corresponding to the above sample to be classified can be:
[0157]
[0158] where L2 is used to indicate the loss value corresponding to the above sample to be classified when the target loss determination method is the loss determination method corresponding to the reference label, p f is used to indicate the reference prediction probability corresponding to the sample to be classified, and m is used to indicate the reference probability threshold.
[0159] In some embodiments, combining the reference prediction probability and the reference probability threshold to determine the loss value corresponding to the sample to be classified can be achieved in the following manner: subtract the reference probability threshold from the reference prediction probability to obtain the loss value corresponding to the sample to be classified; or, determine the nth power of the difference between the reference prediction probability and the reference probability threshold as the loss value corresponding to the sample to be classified.
[0160] In some embodiments, n is a positive integer greater than or equal to 2.
[0161] As an example, when the target loss determination method is the loss determination method corresponding to the reference label, the expression of the loss value corresponding to the to-be-classified sample can also be:
[0162]
[0163] where L2 is used to indicate the loss value corresponding to the to-be-classified sample when the target loss determination method is the loss determination method corresponding to the reference label, p f is used to indicate the reference prediction probability corresponding to the to-be-classified sample, m is used to indicate the reference probability threshold, and n is a positive integer greater than or equal to 2.
[0164] In this way, when the target loss determination method is the loss determination method corresponding to the target label, based on each prediction probability, the target prediction probability corresponding to the to-be-classified sample is determined, and the negative value of the logarithm of the target prediction probability is determined as the loss value corresponding to the to-be-classified sample; when the target loss determination method is the loss determination method corresponding to the reference label, based on each prediction probability, the reference prediction probability corresponding to the to-be-classified sample is determined, and based on the reference prediction probability, the loss value corresponding to the to-be-classified sample is determined, so as to adopt different loss determination methods for the target label and the reference label to determine the corresponding loss values, thereby effectively improving the accuracy of the loss value, and training the classification model through the accurate loss value, thereby effectively improving the classification performance of the obtained target classification model.
[0165] In step 105, based on the loss values respectively corresponding to each to-be-classified sample, the classification model is trained to obtain the target classification model.
[0166] In some embodiments, the above step 105 can be implemented in the following manner: the loss values respectively corresponding to each to-be-classified sample are summed to obtain the total loss value, and based on the total loss value, the classification model is trained to obtain the target classification model.
[0167] In some embodiments, the model parameters of the above target classification model are different from those of the classification model, and the model structures of the above target classification model and the classification model are the same.
[0168] In some embodiments, after training the classification model based on the loss values respectively corresponding to each to-be-classified sample to obtain the target classification model, the to-be-classified object can also be classified in the following manner: the to-be-classified object is obtained, the target classification model is called, and the category prediction is performed on the to-be-classified object to obtain the object category corresponding to the to-be-classified object.
[0169] In some embodiments, after training a classification model based on the loss values respectively corresponding to each sample to be classified to obtain a target classification model, the video to be classified can also be classified in the following manner: obtaining the video to be classified, invoking the target classification model, and performing class prediction on the video to be classified to obtain the video class corresponding to the video to be classified.
[0170] In some embodiments, after training a classification model based on the loss values respectively corresponding to each sample to be classified to obtain a target classification model, the text to be classified can also be classified in the following manner: obtaining the text to be classified, invoking the target classification model, and performing class prediction on the text to be classified to obtain the text class corresponding to the text to be classified.
[0171] In this way, by obtaining multiple samples to be classified carrying different class labels, invoking the classification model, performing class prediction on each sample to be classified respectively to obtain the predicted classes respectively corresponding to each sample to be classified, for each sample to be classified, based on the class label carried by the sample to be classified, determining the target loss determination method corresponding to the sample to be classified, and based on the target loss calculation method and the predicted class corresponding to the sample to be classified, determining the loss value of the sample to be classified, and training the classification model based on the loss values respectively corresponding to each sample to be classified to obtain a target classification model. In this way, for each sample to be classified, based on the class label carried by the sample to be classified, determining the target loss determination method corresponding to the sample to be classified, and based on the target loss calculation method and the predicted class corresponding to the sample to be classified, determining the loss value of the sample to be classified, so as to adopt different loss determination methods for samples to be classified with different class labels to determine the loss values of the corresponding samples to be classified, thereby effectively improving the accuracy of the loss values, and training the classification model with accurate loss values, thereby effectively improving the classification performance of the obtained target classification model.
[0172] See Figure 5 , Figure 5 is a flowchart illustration of the object classification method provided by an embodiment of the present application Figure 2 , which will be described in conjunction with Figure 5 Steps 301 to 304 shown. The object classification method provided by an embodiment of the present application can be implemented independently by a server or a terminal, or jointly implemented by a server and a terminal. Hereinafter, the case where the server implements independently will be taken as an example for description.
[0173] In step 301, an object to be classified and object description information for describing the object to be classified are obtained, and feature extraction is performed on the object description information to obtain object features of the object to be classified.
[0174] In some embodiments, the object description information of the object to be classified may refer to information such as the age and historical behavior of the object to be classified, which can objectively describe the object to be classified.
[0175] In some embodiments, the above feature extraction is used to convert the object description information of the object to be classified into object features in vector form. The above feature extraction of the object description information to obtain the object features of the object to be classified can be achieved in the following manner: call a feature extraction network to perform feature extraction on the object to be classified to obtain the object features of the object to be classified.
[0176] In some embodiments, the above feature extraction network includes a convolutional layer, and the convolutional layer is composed of several convolutional units, and the parameters of each convolutional unit are optimized through the backpropagation algorithm. The purpose of the convolutional operation is to extract different features of the input. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners, etc. More layers of the network can iteratively extract more complex features from the low-level features.
[0177] In step 302, call a target classification model, and based on the object features of the object to be classified, perform class prediction on the object to be classified to obtain the object class corresponding to the object to be classified.
[0178] In some embodiments, the object class is used to indicate whether the object to be classified is a normal object. The target classification model is obtained by training a classification model based on samples of objects to be classified with different class labels. The class labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different.
[0179] In this way, by calling the target classification model, based on the object features of the object to be classified, performing class prediction on the object to be classified to obtain the object class corresponding to the object to be classified. Since the target classification model is obtained by training a classification model based on samples of objects to be classified with different class labels, and the loss determination methods corresponding to the target label and the reference label are different, the accuracy of the loss value used for training is effectively improved. By training the classification model with an accurate loss value, the classification performance of the trained target classification model is effectively improved. Call the target classification model, based on the object features of the object to be classified, perform class prediction on the object to be classified, so that the obtained object class corresponding to the object to be classified is more accurate.
[0180] Next, an exemplary application of the embodiments of the present application in an actual object classification application scenario will be described.
[0181] The embodiment of the present application proposes a Rectify Loss for the problem of abnormal detection in the e-commerce scenario, which is used to solve the problem that fuzzy merchant samples mislead the model learning.
[0182] The embodiment of the present application makes a fine-grained division of merchants. On the basis of dividing merchants into normal merchants and abnormal merchants, the normal merchants are further divided into normal merchants and fuzzy merchants. Fuzzy merchants are those that have been manually verified to have no problems after consumer complaints. Some behaviors of these merchants are already on the verge of violation and can be regarded as ambiguous samples. On the basis of the fine-grained division, the embodiment of the present application proposes Rectify Loss to provide looser supervision for fuzzy merchant samples, which can avoid fuzzy merchant samples misleading the model learning and improve the classification performance of the model for abnormal merchant samples.
[0183] In some embodiments, refer to Figure 10 , Figure 10 is the flow diagram of the training method of the classification model provided by the embodiment of the present application Figure 6 , the training method of the classification model provided by the embodiment of the present application can be achieved through Figure 10 the steps 201 to 205 shown.
[0184] In step 201, normal merchants, abnormal merchants and fuzzy merchants are determined as merchant samples.
[0185] In some embodiments, through the fine-grained division of merchants, the uncomplained merchants are divided into normal merchants, the complained and manually verified abnormal merchants are divided into abnormal merchants, and the complained but manually verified normal merchants are divided into fuzzy merchants. Through the fine-grained division, key prior knowledge is retained.
[0186] In step 202, a feature tokenizer is called.
[0187] In some embodiments, the feature tokenizer converts the input feature x into an embedding matrix Given the feature x j The embedding calculation is as follows:
[0188]
[0189] where b j is the j-th feature deviation, is achieved by multiplying each element with the vector one by one, is achieved through the lookup table of categorical features Overall:
[0190]
[0191]
[0192]
[0193] in, is the one-hot encoded vector corresponding to the categorical feature.
[0194] In step 203, the feature extractor is called.
[0195] In some embodiments, the feature extractor is a Transformer. At this stage, the embeddings of the [CLS] tokens are appended to the embedding matrix T, and then L Transformer layers F1,...,F L Extract features:
[0196]
[0197] In step 204, the classification head is called.
[0198] In some embodiments, the final representation of the [CLS] token is used to predict:
[0199]
[0200] In step 205, the cross loss is determined in combination with the normal / abnormal merchant probability distribution, the correction loss is determined based on the fuzzy merchant probability distribution, and the target label is determined by combining the correction loss and the cross loss.
[0201] Traditional abnormal detection simply divides merchant types into normal classes (normal classes and fuzzy classes) and abnormal classes, and uses classification loss functions such as Cross Entropy Loss for training. Taking Cross Entropy Loss as an example, the model prediction logits in the fourth step are specifically expressed as: The embodiment of the present application uses the Softmax function on logits to calculate the probability. When the embodiment of the present application wants to classify the fuzzy merchant sample into the normal class, the cross loss of the fuzzy merchant sample is:
[0202] CrossEntropyLoss = -log(p n ) (13)
[0203]
[0204] When using Cross Entropy Loss as the optimization objective, it means that the embodiments of this application hope that the feature vectors are as close as possible to the prototypes of the normal classes. However, as mentioned in the embodiments of this application before, the features of the ambiguous samples are somewhat similar to those of the abnormal samples. Therefore, forcing the ambiguous samples to be classified as normal classes will mislead the model's classification of abnormal samples. Therefore, the embodiments of this application believe that it is unreasonable to directly use Cross Entropy Loss to train the ambiguous samples. According to the prior knowledge, the ambiguous samples belong to the normal classes, but their features are somewhat similar to those of the abnormal samples. To relax the supervision and avoid the problems of the above strategy, the embodiments of this application propose Rectify Loss to train the ambiguous samples, and continue to use Cross Entropy Loss to train the normal samples and abnormal samples.
[0205] In some embodiments, a threshold can be set for the predicted probability. When the predicted probability of the normal class is greater than the threshold, the loss value should be 0. Based on this idea, the embodiments of this application can easily write such a loss function:
[0206]
[0207]
[0208] where m is the Margin between the optimization objective and the prototype of the normal class, and 0 < m < 1. p f can be regarded as the probability of the abnormal class predicted by the model. Minimizing p f is equal to maximizing p n , that is, pushing the sample towards the normal class prototype. In addition, the embodiments of this application manually set m. When p f is small enough, the loss is 0. The Margin is very important because it can prevent the ambiguous samples from getting too close to the normal class prototype.
[0209] However, the gradient of the above loss function with respect to p f is:
[0210]
[0211] Regardless of the value of p f , the gradient of this loss function is a constant, which will lead to unstable model training. Finally, the embodiments of this application modify Rectify Loss to:
[0212]
[0213] Generally speaking, Rectify Loss provides a smaller penalty for ambiguous samples than Cross Entropy Loss.
[0214] Thus, based on the original coarse-grained division, the embodiments of the present application perform a fine-grained division of normal merchants into normal merchants and fuzzy merchants. The key prior knowledge of fuzzy merchants is retained, which is more suitable for the abnormal detection problem in the e-commerce scenario. The embodiments of the present application propose Rectify Loss to provide loose supervision for fuzzy samples, avoiding fuzzy samples from misleading the model learning and making up for the deficiencies considered in the previous solutions. The embodiments of the present application have good effects when experiments are carried out on multiple large real table datasets using various neural networks.
[0215] Thus, by obtaining multiple to-be-classified samples carrying different category labels, invoking a classification model, respectively predicting the categories of each to-be-classified sample, obtaining the predicted categories corresponding to each to-be-classified sample, for each to-be-classified sample, based on the category label carried by the to-be-classified sample, determining the target loss determination method corresponding to the to-be-classified sample, and based on the target loss calculation method and the predicted category corresponding to the to-be-classified sample, determining the loss value of the to-be-classified sample, and based on the loss values respectively corresponding to each to-be-classified sample, training the classification model to obtain a target classification model. Thus, for each to-be-classified sample, based on the category label carried by the to-be-classified sample, determining the target loss determination method corresponding to the to-be-classified sample, and based on the target loss calculation method and the predicted category corresponding to the to-be-classified sample, determining the loss value of the to-be-classified sample, so as to adopt different loss determination methods for to-be-classified samples with different category labels to determine the loss values of the corresponding to-be-classified samples, thereby effectively improving the accuracy of the loss values, and training the classification model with accurate loss values, thereby effectively improving the classification performance of the obtained target classification model.
[0216] It can be understood that in the embodiments of the present application, when it comes to data related to to-be-classified samples, etc., when the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0217] Next, continue to describe the exemplary structure of the training device 455 of the classification model provided by the embodiments of the present application as a software module. In some embodiments, such as Figure 2As shown, the software modules in the training device 455 of the classification model stored in the memory 450 may include: an acquisition module 4551, configured to acquire a plurality of samples to be classified and the class labels carried by each of the samples to be classified, where the class labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different; a prediction module 4552, configured to call the classification model to perform class prediction on each of the samples to be classified, and obtain the predicted classes corresponding to each of the samples to be classified; a loss module 4553, configured to, for each of the samples to be classified, determine the target loss determination method corresponding to the sample to be classified based on the class label carried by the sample to be classified, and determine the loss value corresponding to the sample to be classified based on the target loss determination method and the predicted class corresponding to the sample to be classified; a training module 4554, configured to train the classification model based on the loss values corresponding to each of the samples to be classified to obtain a target classification model.
[0218] In some embodiments, the above loss module is further configured to obtain a target mapping relationship between a plurality of preset class labels and corresponding loss determination methods, and compare each of the class labels with each of the preset class labels in the target mapping relationship to obtain a label comparison result corresponding to the class label; when the label comparison result indicates that the class label exists in the plurality of preset class labels, determine the loss determination method corresponding to the class label in the target mapping relationship as the target loss determination method.
[0219] In some embodiments, the above target label includes a plurality of sub-target labels, and the above loss module is further configured to obtain the prediction probabilities of the sample to be classified corresponding to each of the sub-target labels, and the predicted class corresponding to the sample to be classified is the class indicated by the sub-target label with the highest prediction probability; when the target loss determination method is the loss determination method corresponding to the target label, determine the target prediction probability corresponding to the sample to be classified based on each of the prediction probabilities, and determine the opposite number of the logarithm value of the target prediction probability as the loss value corresponding to the sample to be classified; when the target loss determination method is the loss determination method corresponding to the reference label, determine the reference prediction probability corresponding to the sample to be classified based on each of the prediction probabilities, and determine the loss value corresponding to the sample to be classified based on the reference prediction probability; where the sum of the reference prediction probability and the target prediction probability is equal to 1.
[0220] In some embodiments, the above-mentioned sub-goal labels include positive goal labels and negative goal labels, and the prediction probabilities include a first prediction probability corresponding to the positive goal label and a second prediction probability corresponding to the negative goal label; the above-mentioned loss module is further configured to sum the first prediction probability and the second prediction probability to obtain a sum probability; divide the exponential value of the second prediction probability by the exponential value of the sum probability to obtain a reference prediction probability corresponding to the sample to be classified.
[0221] In some embodiments, the above-mentioned loss module is further configured to compare the reference prediction probability with a reference probability threshold to obtain a probability comparison result; when the probability comparison result indicates that the reference prediction probability is less than the reference probability threshold, determine the loss value corresponding to the sample to be classified as zero; when the probability comparison result indicates that the reference prediction probability is greater than or equal to the reference probability threshold, determine the loss value corresponding to the sample to be classified by combining the reference prediction probability and the reference probability threshold.
[0222] In some embodiments, the above-mentioned loss module is further configured to subtract the reference probability threshold from the reference prediction probability to obtain the loss value corresponding to the sample to be classified; or, determine the nth power of the difference between the reference prediction probability and the reference probability threshold as the loss value corresponding to the sample to be classified; where n is a positive integer greater than or equal to 2.
[0223] In some embodiments, the classification model includes a conversion layer, a feature extraction layer, and a classification layer. The above-mentioned prediction module is further configured to perform the following processing for each of the samples to be classified: call the conversion layer to perform enhancement conversion on the sample to be classified to obtain an enhanced sample corresponding to the sample to be classified; call the feature extraction layer to perform feature extraction on the enhanced sample to obtain a sample feature corresponding to the sample to be classified; call the classification layer to perform class prediction on the sample to be classified based on the sample feature to obtain a prediction class corresponding to the sample to be classified.
[0224] In some embodiments, the enhancement conversion includes a partitioning process and a fusion process. The above-mentioned prediction module is further configured to perform a partitioning process on the sample to be classified to obtain multiple sub-samples in the sample to be classified, and obtain enhanced sub-samples corresponding to each of the sub-samples; perform a fusion process on each of the enhanced sub-samples and each of the sub-samples to obtain an enhanced sample corresponding to the sample to be classified.
[0225] In some embodiments, the above prediction module is further configured to obtain the mapping relationship between multiple preset samples and the corresponding preset enhanced samples, and perform the following processing for each of the sub-samples: compare each of the sub-samples with each of the preset samples to obtain the comparison result corresponding to the sub-sample; when the comparison result indicates that the sub-sample exists in the preset samples, determine the preset enhanced sample corresponding to the sub-sample in the mapping relationship as the enhanced sub-sample of the sub-sample.
[0226] In some embodiments, the above prediction module is further configured to call the first feature extraction layer to extract features from the enhanced sample to obtain the first sample feature; traverse i and perform the following processing: call the i-th feature extraction layer to extract features from the enhanced sample based on the (i - 1)-th sample feature to obtain the i-th sample feature, where 2 ≤ i ≤ N; determine the N-th sample feature as the sample feature corresponding to the sample to be classified.
[0227] An object classification device provided by an embodiment of the present application includes:
[0228] A feature extraction module, configured to obtain an object to be classified and object description information for describing the object to be classified, and extract features from the object description information to obtain the object feature of the object to be classified;
[0229] A classification module, configured to call a target classification model to predict the category of the object to be classified based on the object feature of the object to be classified to obtain the object category corresponding to the object to be classified; where the object category is used to indicate whether the object to be classified is a normal object. Wherein, the target classification model is obtained by training a classification model based on samples of objects to be classified carrying different category labels, and the category labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different.
[0230] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the training method of the classification model in the embodiment of the present application.
[0231] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, where the computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, the processor will be caused to execute the training method of the classification model provided by the embodiment of the present application, for example, Figure 4 the training method of the classification model shown.
[0232] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the object classification method described above in the embodiment of the present application.
[0233] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, where the computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, the processor will be caused to execute the training method of the classification model provided in the embodiment of the present application. For example, Figure 5 the object classification method shown.
[0234] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various electronic devices including one or any combination of the above memories.
[0235] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0236] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (such as files storing one or more modules, subroutines, or code portions).
[0237] In the embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0238] As an example, the computer-executable instructions can be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected through a communication network.
[0239] In summary, the embodiments of the present application have the following beneficial effects:
[0240] (1) By obtaining multiple to-be-classified samples carrying different category labels, invoking a classification model, respectively predicting the categories of each to-be-classified sample to obtain the predicted categories corresponding to each to-be-classified sample, for each to-be-classified sample, based on the category label carried by the to-be-classified sample, determining the target loss determination method corresponding to the to-be-classified sample, and based on the target loss calculation method and the predicted category corresponding to the to-be-classified sample, determining the loss value of the to-be-classified sample, and based on the loss values corresponding to each to-be-classified sample, training the classification model to obtain a target classification model. In this way, for each to-be-classified sample, based on the category label carried by the to-be-classified sample, determining the target loss determination method corresponding to the to-be-classified sample, and based on the target loss calculation method and the predicted category corresponding to the to-be-classified sample, determining the loss value of the to-be-classified sample, so as to adopt different loss determination methods for to-be-classified samples with different category labels, determine the loss values of the corresponding to-be-classified samples, thereby effectively improving the accuracy of the loss values, and training the classification model with accurate loss values, thereby effectively improving the classification performance of the trained target classification model.
[0241] (2) When the target loss determination method is the loss determination method corresponding to the target label, based on each prediction probability, determining the target prediction probability corresponding to the to-be-classified sample, and taking the opposite number of the logarithm value of the target prediction probability as the loss value corresponding to the to-be-classified sample; when the target loss determination method is the loss determination method corresponding to the reference label, based on each prediction probability, determining the reference prediction probability corresponding to the to-be-classified sample, and based on the reference prediction probability, determining the loss value corresponding to the to-be-classified sample, so as to adopt different loss determination methods for the target label and the reference label, determine the corresponding loss values, thereby effectively improving the accuracy of the loss values, and training the classification model with accurate loss values, thereby effectively improving the classification performance of the trained target classification model.
[0242] (3) For each sample to be classified, based on the class label carried by the sample to be classified, determine the target loss determination method corresponding to the sample to be classified, and based on the target loss calculation method and the predicted class corresponding to the sample to be classified, determine the loss value of the sample to be classified. Thus, for samples to be classified with different class labels, different loss determination methods are used to determine the loss values of the corresponding samples to be classified, effectively improving the accuracy of the loss value. By training the classification model with the accurate loss value, the classification performance of the obtained target classification model is effectively improved.
[0243] (4) By calling the conversion layer, perform enhanced conversion on the sample to be classified to obtain an enhanced sample with more information than the sample to be classified. Call the feature extraction layer to extract features from the enhanced sample to obtain the sample features corresponding to the sample to be classified. Call the classification layer to perform class prediction on the sample to be classified based on the sample features to obtain the predicted class corresponding to the sample to be classified. Thus, through enhanced conversion, the information volume of the sample to be classified is significantly enhanced, and class prediction is performed based on the enhanced sample with a larger information volume, significantly improving the accuracy of the obtained predicted class.
[0244] (5) In the embodiment of the present application, on the basis of the original coarse-grained division, normal merchants are finely divided into normal merchants and fuzzy merchants. The key prior knowledge of fuzzy merchants is retained, which is more suitable for the abnormal detection problem in the e-commerce scenario. The embodiment of the present application proposes Rectify Loss to provide loose supervision for fuzzy samples, avoiding fuzzy samples from misleading model learning and making up for the deficiencies considered in the previous solutions. The embodiment of the present application conducts experiments on multiple large real table datasets using various neural networks and achieves good results.
[0245] (6) By calling the target classification model, based on the object features of the object to be classified, perform class prediction on the object to be classified to obtain the object class corresponding to the object to be classified. Since the target classification model is obtained by training the classification model based on samples of objects to be classified with different class labels, and the loss determination methods corresponding to the target label and the reference label are different, the accuracy of the loss value used for training is effectively improved. By training the classification model with the accurate loss value, the classification performance of the obtained target classification model is effectively improved. Call the target classification model, based on the object features of the object to be classified, perform class prediction on the object to be classified, making the obtained object class corresponding to the object to be classified more accurate.
[0246] The above is only the embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A training method for a classification model, characterized in that, The method includes: Obtaining a plurality of samples to be classified, and class labels respectively carried by each of the samples to be classified, where the class labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different; Invoking the classification model to respectively perform class prediction on each of the samples to be classified, and obtaining the predicted classes respectively corresponding to each of the samples to be classified; For each of the samples to be classified, based on the class label carried by the sample to be classified, determining the target loss determination method corresponding to the sample to be classified, and based on the target loss determination method and the predicted class corresponding to the sample to be classified, determining the loss value corresponding to the sample to be classified; Training the classification model based on the loss values respectively corresponding to each of the samples to be classified to obtain a target classification model.
2. The method according to claim 1, wherein The determining the target loss determination method corresponding to the sample to be classified based on the class label carried by the sample to be classified includes: Obtaining a target mapping relationship between a plurality of preset class labels and corresponding loss determination methods, and comparing the class labels with each of the preset class labels in the target mapping relationship respectively to obtain a label comparison result corresponding to the class label; When the label comparison result indicates that the class label exists in the plurality of preset class labels, determining the loss determination method corresponding to the class label in the target mapping relationship as the target loss determination method.
3. The method according to claim 1, characterized in that The target label includes a plurality of sub-target labels, and the determining the loss value corresponding to the sample to be classified based on the target loss determination method and the predicted class corresponding to the sample to be classified includes: Obtaining the prediction probabilities corresponding to each of the sub-target labels of the sample to be classified, and the predicted class corresponding to the sample to be classified is the class indicated by the sub-target label with the largest prediction probability; When the target loss determination method is the loss determination method corresponding to the target label, determining the target prediction probability corresponding to the sample to be classified based on each of the prediction probabilities, and determining the opposite number of the logarithm value of the target prediction probability as the loss value corresponding to the sample to be classified; When the target loss determination method is the loss determination method corresponding to the reference label, determining the reference prediction probability corresponding to the sample to be classified based on each of the prediction probabilities, and determining the loss value corresponding to the sample to be classified based on the reference prediction probability; Wherein, the sum of the reference prediction probability and the target prediction probability is equal to 1.
4. The method according to claim 3, wherein The sub-target label includes a positive target label and a negative target label, and the prediction probability includes a first prediction probability corresponding to the positive target label and a second prediction probability corresponding to the negative target label; The determining the target prediction probability corresponding to the sample to be classified based on each of the prediction probabilities includes: Adding the first prediction probability and the second prediction probability to obtain a sum probability; Dividing the exponential value of the first prediction probability by the exponential value of the sum probability to obtain the target prediction probability corresponding to the sample to be classified.
5. The method according to claim 3, wherein The sub-goal labels include a positive goal label and a negative goal label, and the prediction probabilities include a first prediction probability corresponding to the positive goal label and a second prediction probability corresponding to the negative goal label; Determining the reference prediction probability corresponding to the sample to be classified based on each of the prediction probabilities includes: Adding the first prediction probability and the second prediction probability to obtain a sum probability; Dividing the exponential value of the second prediction probability by the exponential value of the sum probability to obtain the reference prediction probability corresponding to the sample to be classified.
6. The method according to claim 3, wherein Determining the loss value corresponding to the sample to be classified based on the reference prediction probability includes: Comparing the reference prediction probability with a reference probability threshold to obtain a probability comparison result; When the probability comparison result indicates that the reference prediction probability is less than the reference probability threshold, determining the loss value corresponding to the sample to be classified as zero; When the probability comparison result indicates that the reference prediction probability is greater than or equal to the reference probability threshold, combining the reference prediction probability and the reference probability threshold to determine the loss value corresponding to the sample to be classified.
7. The method according to claim 6, characterized in that, Combining the reference prediction probability and the reference probability threshold to determine the loss value corresponding to the sample to be classified includes: Subtracting the reference probability threshold from the reference prediction probability to obtain the loss value corresponding to the sample to be classified; or, Taking the nth power of the difference between the reference prediction probability and the reference probability threshold as the loss value corresponding to the sample to be classified; where n is a positive integer greater than or equal to 2.
8. The method according to claim 1, characterized in that, The classification model includes a conversion layer, a feature extraction layer, and a classification layer. Invoking the classification model to perform class prediction on each of the samples to be classified to obtain the predicted class corresponding to each of the samples to be classified includes: Performing the following processing for each of the samples to be classified: Invoking the conversion layer to perform enhancement conversion on the sample to be classified to obtain an enhanced sample corresponding to the sample to be classified; Invoking the feature extraction layer to perform feature extraction on the enhanced sample to obtain sample features corresponding to the sample to be classified; Invoking the classification layer to perform class prediction on the sample to be classified based on the sample features to obtain the predicted class corresponding to the sample to be classified.
9. The method according to claim 8, wherein The enhancement conversion includes a partitioning process and a fusion process. Performing the enhancement conversion on the sample to be classified to obtain an enhanced sample corresponding to the sample to be classified includes: Performing a partitioning process on the sample to be classified to obtain multiple sub-samples in the sample to be classified, and obtaining enhanced sub-samples corresponding to each of the sub-samples; Performing a fusion process on each of the enhanced sub-samples and each of the sub-samples to obtain an enhanced sample corresponding to the sample to be classified.
10. The method according to claim 9, wherein Obtaining the enhanced sub-samples corresponding to each of the sub-samples includes: Obtaining the mapping relationship between multiple preset samples and corresponding preset enhanced samples, and performing the following processing for each of the sub-samples: Comparing each of the sub-samples with each of the preset samples to obtain a comparison result corresponding to the sub-sample; When the comparison result indicates the existence of the sub-sample in the preset sample, determine the preset enhanced sample corresponding to the sub-sample in the mapping relationship as the enhanced sub-sample of the sub-sample.
11. The method according to claim 8, characterized in that The calling the feature extraction layer to extract features from the enhanced sample to obtain the sample features corresponding to the sample to be classified includes: Call the first feature extraction layer to extract features from the enhanced sample to obtain the first sample features; Traverse i and perform the following processing: Call the i-th feature extraction layer to extract features from the enhanced sample based on the (i - 1)-th sample features to obtain the i-th sample features, where 2 ≤ i ≤ N; Determine the N-th sample features as the sample features corresponding to the sample to be classified.
12. A method for object classification, characterized in that, The method includes: Obtain an object to be classified and object description information for describing the object to be classified, and extract features from the object description information to obtain the object features of the object to be classified; Call a target classification model to perform a class prediction on the object to be classified based on the object features of the object to be classified to obtain the object class corresponding to the object to be classified; Wherein, the object class is used to indicate whether the object to be classified is a normal object, and the target classification model is obtained by training a classification model based on samples of objects to be classified with different class labels, and the class labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different.
13. A training device for a classification model, characterized in that, The apparatus includes: An acquisition module, configured to acquire a plurality of samples to be classified and class labels carried by each of the samples to be classified, where the class labels include a target label and a reference label, and the loss determination methods corresponding to the target label and the reference label are different; A prediction module, configured to call the classification model to perform a class prediction on each of the samples to be classified to obtain the prediction classes corresponding to each of the samples to be classified; A loss module, configured to, for each of the samples to be classified, determine the target loss determination method corresponding to the sample to be classified based on the class label carried by the sample to be classified, and determine the loss value corresponding to the sample to be classified based on the target loss determination method and the prediction class corresponding to the sample to be classified; A training module, configured to train the classification model based on the loss values corresponding to each of the samples to be classified to obtain a target classification model.
14. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions or a computer program; A processor, configured to implement the method according to any one of claims 1 to 12 when executing the computer-executable instructions or the computer program stored in the memory.
15. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 12.
16. A computer program product, comprising a computer program or computer-executable instructions, characterized in that, The computer program or the computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 12.