Image anomaly detection model training method, image anomaly detection method and device
By iteratively adjusting model parameters and mapping labels using feedback data, the method improves the accuracy of image anomaly detection models, addressing the issue of noisy training labels and enhancing their predictive capabilities.
Patent Information
- Application Number
- CN202111079651.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-15
AI Technical Summary
In the prior art, training of image anomaly detection model based on simple binary classification has the problem of low accuracy, mainly due to the diversity of abnormal images, the training label is noisy.
By obtaining the current mapping label and target prediction label of the training image, the model feedback data is generated, the label loss is generated based on the data change reference information, the current mapping label and model parameters are adjusted until the training is completed, and the target image abnormality detection model is obtained.
Effectively filtering the impact of noise data on model performance improves the accuracy of the model and ensures that image abnormalities can be accurately detected during application.
Smart Images

Figure CN114332578B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to a method for training an image anomaly detection model, an image anomaly detection method, an apparatus, a computer device, and a storage medium. Background Art
[0002] With the progress of computer technologies, images are widely used in various industries, and the requirements for image quality are also getting higher and higher. Machine learning models can be trained to detect whether an image is abnormal, so as to screen out abnormal and low-quality images.
[0003] In traditional technologies, the machine learning model is usually trained based on training images. The training images are usually simple binary classification images, that is, the training images are divided into normal images and abnormal images. However, abnormal images usually correspond to various degrees of abnormal situations, and the simple binary classification labels may be very subjective, which may lead to noisy training labels of the training images. The model trained based on such training images has the problem of low accuracy. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide an image anomaly detection model training method, an image anomaly detection method, an apparatus, a computer device, and a storage medium that can improve the prediction accuracy of the model.
[0005] An image anomaly detection model training method, the method comprising:
[0006] Obtaining a current mapping label corresponding to a training image and a training label of the training image;
[0007] Inputting the training image into an initial image anomaly detection model to obtain a target prediction label corresponding to the training image;
[0008] Generating model feedback data based on the current mapping label corresponding to the training label of the training image and the target prediction label;
[0009] Generating a label loss based on data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label;
[0010] Adjusting model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain a target image anomaly detection model.
[0011] An image anomaly detection model training apparatus, the apparatus comprising:
[0012] A training data acquisition module, configured to acquire a training image and a current mapping label corresponding to the training label of the training image;
[0013] A model prediction module, configured to input the training image into an initial image anomaly detection model to obtain a target prediction label corresponding to the training image;
[0014] A model feedback data determination module, configured to generate model feedback data based on the current mapping label corresponding to the training label of the training image and the target prediction label;
[0015] A label adjustment module, configured to generate a label loss based on the data change reference information corresponding to the model feedback data, adjust the current mapping label based on the label loss to obtain an updated mapping label, and use the updated mapping label as the current mapping label;
[0016] A model adjustment module, configured to adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain a target image anomaly detection model.
[0017] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Acquire a training image and a current mapping label corresponding to the training label of the training image;
[0019] Input the training image into an initial image anomaly detection model to obtain a target prediction label corresponding to the training image;
[0020] Generate model feedback data based on the current mapping label corresponding to the training label of the training image and the target prediction label;
[0021] Generate a label loss based on the data change reference information corresponding to the model feedback data, adjust the current mapping label based on the label loss to obtain an updated mapping label, and use the updated mapping label as the current mapping label;
[0022] Adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain a target image anomaly detection model.
[0023] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0024] Obtain the current mapping label corresponding to the training image and the training label of the training image;
[0025] Input the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image;
[0026] Generate model feedback data based on the current mapping label corresponding to the training label of the training image and the target prediction label;
[0027] Generate a label loss based on the data change reference information corresponding to the model feedback data, adjust the current mapping label based on the label loss to obtain an updated mapping label, and use the updated mapping label as the current mapping label;
[0028] Adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0029] An image anomaly detection method, the method includes:
[0030] Obtain the image to be detected;
[0031] Input the image to be detected into the target image anomaly detection model to obtain the model prediction label corresponding to the image to be detected;
[0032] Determine the image anomaly detection result corresponding to the image to be detected based on the model prediction label;
[0033] Wherein, the training process of the target image anomaly detection model includes: obtaining the current mapping label corresponding to the training image and the training label of the training image; inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image; generating model feedback data based on the current mapping label corresponding to the training label of the training image and the target prediction label; generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0034] An image anomaly detection device, the device includes:
[0035] An image acquisition module, configured to acquire the image to be detected;
[0036] A label prediction module, configured to input the image to be detected into a target image anomaly detection model to obtain a model prediction label corresponding to the image to be detected;
[0037] A detection result determination module, configured to determine an image anomaly detection result corresponding to the image to be detected based on the model prediction label;
[0038] Wherein, the training process of the target image anomaly detection model includes: obtaining a training image and a current mapping label corresponding to the training label of the training image; inputting the training image into an initial image anomaly detection model to obtain a target prediction label corresponding to the training image; generating model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image; generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0039] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0040] Obtain an image to be detected;
[0041] Input the image to be detected into a target image anomaly detection model to obtain a model prediction label corresponding to the image to be detected;
[0042] Determine an image anomaly detection result corresponding to the image to be detected based on the model prediction label;
[0043] Wherein, the training process of the target image anomaly detection model includes: obtaining a training image and a current mapping label corresponding to the training label of the training image; inputting the training image into an initial image anomaly detection model to obtain a target prediction label corresponding to the training image; generating model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image; generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0044] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the following steps are implemented:
[0045] Obtain an image to be detected;
[0046] Input the image to be detected into a target image anomaly detection model to obtain a model prediction label corresponding to the image to be detected;
[0047] Determine an image anomaly detection result corresponding to the image to be detected based on the model prediction label;
[0048] Wherein, the training process of the target image anomaly detection model includes: obtaining a training image and a current mapping label corresponding to the training label of the training image; inputting the training image into an initial image anomaly detection model to obtain a target prediction label corresponding to the training image; generating model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image; generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0049] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0050] The above-mentioned method for training an image anomaly detection model, image anomaly detection method, device, computer device, and storage medium obtain the current mapping label corresponding to the training image and the training label of the training image, input the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image, generate model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image, generate a label loss based on the data change reference information corresponding to the model feedback data, adjust the current mapping label to obtain an updated mapping label, use the updated mapping label as the current mapping label, adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model. In this way, when training the model, the current mapping label and the model parameters are adjusted synchronously, and the self-correction of the current mapping label corresponding to the training label with noise is performed using the feature learning ability of the model, which can effectively filter the performance impact of noise data on the model, greatly improve the performance of the model, and finally train an image anomaly detection model with high accuracy. Subsequently, when applying the model, a high-accuracy image anomaly detection result can be obtained based on the image anomaly detection model with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 FIG. is an application environment diagram of the method for training an image anomaly detection model and the image anomaly detection method in an embodiment;
[0052] Figure 2 FIG. is a schematic flowchart of the method for training an image anomaly detection model in an embodiment;
[0053] Figure 3 FIG. is a schematic flowchart of the method for training an image anomaly detection model in another embodiment;
[0054] Figure 4 FIG. is a schematic flowchart of the method for training an image anomaly detection model in yet another embodiment;
[0055] Figure 5 FIG. is a schematic flowchart of the image anomaly detection method in an embodiment;
[0056] Figure 6 FIG. is a schematic diagram of a screen-frozen image in an embodiment;
[0057] Figure 7 FIG. is a schematic flowchart of the method for training an image screen-frozen detection model in an embodiment;
[0058] Figure 8 FIG. is a schematic diagram of the image screen-frozen detection result in an embodiment;
[0059] Figure 9 It is a structural block diagram of an image anomaly detection model training device in an embodiment;
[0060] Figure 10 It is a structural block diagram of an image anomaly detection device in an embodiment;
[0061] Figure 11 It is an internal structure diagram of a computer device in an embodiment;
[0062] Figure 12 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0063] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0064] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.
[0065] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0066] Computer Vision Technology (CV) is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphics processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0067] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, knowledge graph, etc.
[0068] Machine Learning (ML) is a multi-disciplinary cross-cutting discipline that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0069] The solution provided in the embodiments of this application involves technologies such as computer vision technology, natural language processing, and machine learning in artificial intelligence, and is specifically described through the following embodiments:
[0070] The image anomaly detection model training method and image anomaly detection method provided in this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through a network. The terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, vehicle-mounted terminals, roadside terminals, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster or cloud server composed of multiple servers.
[0071] Both the terminal 102 and the server 104 can be separately used to execute the image anomaly detection model training method and the image anomaly detection method provided in the embodiments of the present application.
[0072] For example, the server obtains the current mapping label corresponding to the training image and the training label of the training image, inputs the training image into the initial image anomaly detection model, and obtains the target prediction label corresponding to the training image. The server generates model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image, generates a label loss based on the data change reference information corresponding to the model feedback data, adjusts the current mapping label based on the label loss to obtain an updated mapping label, and uses the updated mapping label as the current mapping label. The server adjusts the model parameters of the initial image anomaly detection model based on the model feedback data, obtains an updated initial image anomaly detection model, and returns to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0073] The server obtains the image to be detected, inputs the image to be detected into the target image anomaly detection model, obtains the model prediction label corresponding to the image to be detected, and determines the image anomaly detection result corresponding to the image to be detected based on the model prediction label.
[0074] The terminal 102 and the server 104 can also be used in cooperation to execute the image anomaly detection model training method and the image anomaly detection method provided in the embodiments of the present application.
[0075] For example, the server obtains the training image and the current mapping label corresponding to the training label of the training image from the terminal. The server inputs the training image into the initial image anomaly detection model and obtains the target prediction label corresponding to the training image. The server generates model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image, generates a label loss based on the data change reference information corresponding to the model feedback data, adjusts the current mapping label based on the label loss to obtain an updated mapping label, and uses the updated mapping label as the current mapping label. The server adjusts the model parameters of the initial image anomaly detection model based on the model feedback data, obtains an updated initial image anomaly detection model, and returns to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model. The server sends the target image anomaly detection model to the terminal.
[0076] The terminal obtains the image to be detected, inputs the image to be detected into the target image anomaly detection model, obtains the model prediction label corresponding to the image to be detected, and the terminal determines the image anomaly detection result corresponding to the image to be detected based on the model prediction label.
[0077] In one embodiment, as Figure 2 shown, an image anomaly detection model training method is provided. Taking the computer in Figure 1 as an example for illustration, it can be understood that the computer device can be the terminal 102 or the server 104. In this embodiment, the image anomaly detection model training method includes the following steps:
[0078] Step S202, obtain the training image and the current mapping label corresponding to the training label of the training image.
[0079] Among them, the training image refers to the image used for model training. The image can be a picture or a video frame in a video. The training label is a manually marked label used to identify the image anomaly determination result of the training image. The training label can be a binary classification label. For example, the training label corresponding to the image with anomalies is a negative label, represented by 1, and the training label corresponding to the image without anomalies is a positive label, represented by 0. The training label can also be a multi-classification label. For example, the training label corresponding to the image without anomalies is the first label, represented by 0, the training label corresponding to the image with obvious anomalies is the second label, represented by 1, and the training label corresponding to the image with slight anomalies is the third label, represented by 2.
[0080] Furthermore, different model training samples can be set for different types of image anomalies, and each model training sample is used to specifically train the corresponding image anomaly detection model. For example, image anomalies include types such as image screen freeze, image blur, image mosaic, and image ghosting. Taking image screen freeze and image blur as examples, for image screen freeze, an image screen freeze detection model can be specifically trained, and the training label of the training image in the training sample corresponding to the image screen freeze detection model is used to identify whether there is a screen freeze in the image. For image blur, an image blur detection model can be specifically trained, and the training label of the training image in the training sample corresponding to the image blur detection model is used to identify whether there is blur in the image. It can be understood that there can be multiple training images participating in model training each time.
[0081] The initial current mapping label is obtained by mapping and converting the training label, and this current mapping label is used to convert the discrete training label into data that is easy for the model to learn and calculate. A training label is usually represented by a constant value. After mapping and conversion, the current mapping label corresponding to this training label can be represented by a vector. In one embodiment, the distances between the respective initial current mapping labels are the same.
[0082] Specifically, the computer device can obtain training samples of the image anomaly detection model locally or from other terminals or servers. The training samples include training images and training labels corresponding to the training images. The computer device can perform mapping transformation on the training labels to obtain the current mapping labels. It can be understood that the computer device can use a custom algorithm or formula for the mapping transformation. Of course, the training samples can also directly include the training images, the corresponding training labels, and the current mapping labels.
[0083] Step S204: Input the training images into the initial image anomaly detection model to obtain the target prediction labels corresponding to the training images.
[0084] Among them, the initial image anomaly detection model refers to the image anomaly detection model to be trained. The image anomaly detection model is a machine learning model, specifically, it can be a model of types such as a convolutional neural network, a deep neural network, etc. The input data of the image anomaly detection model is an image, and the output data is a prediction label. The prediction label is used to identify the probability that the input image is abnormal. The prediction label can also be represented by a vector.
[0085] Specifically, the computer device can input the training images into the initial image anomaly detection model. After the initial image anomaly detection model processes the input data, it can output the target prediction labels corresponding to the training images.
[0086] In one embodiment, the initial image anomaly detection model can be an original image anomaly detection model, and the model parameters in the original image anomaly detection model are randomly initialized. That is to say, the initial image anomaly detection model can be an original model that has not undergone any model training. The initial image anomaly detection model can also be a pre-trained image anomaly detection model. That is to say, the initial image anomaly detection model can be obtained by pre-training the original image anomaly detection model.
[0087] Step S206: Generate model feedback data based on the current mapping labels corresponding to the training labels of the training images and the target prediction labels.
[0088] Specifically, the computer device can calculate the model feedback data based on the data distribution difference between the current mapping label corresponding to the training label of the training image and the target prediction label. For example, the classification loss can be calculated by computing the distance between the current mapping label corresponding to the training label of the training image and the target prediction label, and the classification loss can be used as the model feedback data. The prior loss can also be obtained based on the training label and the target prediction label, and the model feedback data can be obtained based on the classification loss and the prior loss. The self-entropy loss can also be obtained based on the target prediction label, and the model feedback data can be obtained based on the classification loss, the prior loss, and the self-entropy loss, or the model feedback data can be obtained based on the classification loss and the self-entropy loss. Among them, the classification loss, the prior loss, and the self-entropy loss can all be obtained by calculating the divergence of the data, or by calculating the cross-entropy of the data, or by using custom formulas, algorithms, etc.
[0089] Step S208: Generate a label loss based on the data change reference information corresponding to the model feedback data, adjust the current mapping label based on the label loss to obtain an updated mapping label, and use the updated mapping label as the current mapping label.
[0090] Specifically, the data change reference information is used to measure the data change speed of the model feedback data along the direction of the current mapping label. It can be understood that the model feedback data is obtained based on the current mapping label and the target prediction label. For the model feedback data, both the current mapping label and the target prediction label are variables. Since the label loss is used to adjust the current mapping label corresponding to the training label, the computer device can generate the data change reference information based on the data change speed of the model feedback data in the direction of the current mapping label, and then generate the label loss based on the data change reference information, adjust the current mapping label based on the label loss to obtain an updated mapping label, and use the updated mapping label as the new current mapping label. For example, the gradient can be used to measure the change speed of a function along a certain direction. Therefore, the computer device can calculate the gradient of the model feedback data with respect to the current mapping label to obtain a loss gradient, use the loss gradient as the data change reference information, obtain the label loss based on the loss gradient, and adjust the current mapping label based on the label loss to obtain an updated current mapping label. Of course, the computer device can also use a custom formula or algorithm to calculate the data change reference information.
[0091] Among them, generating the label loss based on the data change reference information can be using the data change reference information as the label loss, or fusing the data change reference information and hyperparameters and using the fused data as the label loss. The hyperparameters can be preset fixed values, or the hyperparameters can also be used as a type of model parameter and be adjusted and learned during the model training process.
[0092] Step S210: Adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0093] Among them, the target image anomaly detection model refers to the image anomaly detection model that has completed training.
[0094] Specifically, in addition to adjusting the current mapping label based on the model feedback data, the computer device also needs to synchronously adjust the model parameters based on the model feedback data. The computer device can perform backpropagation based on the model feedback data to update the model parameters of the initial image anomaly detection model, thereby obtaining a new initial image anomaly detection model, that is, the updated initial image anomaly detection model. The computer device can input the training image into the updated initial image anomaly detection model to obtain an updated target prediction label, generate updated model feedback data based on the updated current mapping label and the updated target prediction label, and adjust the current mapping label and the model parameters of the initial image anomaly detection model again based on the updated model feedback data. Repeat the above steps until the model convergence condition is met, stop the training, and obtain the target image anomaly detection model. The model convergence condition can be that the number of model iterations reaches the preset number of iterations, the value of the model feedback data is less than the preset target value, the change rate of the model feedback data is less than the preset change rate, the model feedback data is minimized, etc.
[0095] In one embodiment, the computer device can adjust the model parameters of the initial image anomaly detection model based on the model feedback data through the gradient descent algorithm.
[0096] In one embodiment, the computer device can train the current mapping label and the initial image anomaly detection model based on the model feedback data until the model convergence condition is met to obtain an intermediate image anomaly detection model and a target mapping label, and directly use the intermediate image anomaly detection model as the target image anomaly detection model. Further, the computer device can also keep the target mapping label unchanged, input the training image into the intermediate image anomaly detection model to obtain an updated prediction label, calculate an updated loss based on the updated prediction label and the target mapping label, and train the intermediate image anomaly detection model based on the updated loss until the model convergence condition is met to obtain the target image anomaly detection model.
[0097] In one embodiment, the target image anomaly detection model is any one of an image screen freeze detection model, an image blur detection model, and an image mosaic detection model.
[0098] Among them, the image screen distortion detection model is used to detect whether there is a screen distortion phenomenon in the input image. The screen distortion phenomenon is an image abnormality caused by problems in the image encoding and decoding process. The image blurring detection model is used to detect whether there is a blurring phenomenon in the input image. The blurring phenomenon is an image abnormality caused by the shooting parameters or shooting angles of the shooting device during image shooting. The image mosaic detection model is used to detect whether there is a mosaic phenomenon in the input image. The mosaic phenomenon is an image abnormality caused by the deterioration of the color level details in a local area of the image.
[0099] In the above method for training the image abnormality detection model, by obtaining the training image and the current mapping label corresponding to the training label of the training image, inputting the training image into the initial image abnormality detection model to obtain the target prediction label corresponding to the training image, generating model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image, generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, using the updated mapping label as the current mapping label, adjusting the model parameters of the initial image abnormality detection model based on the model feedback data to obtain an updated initial image abnormality detection model, and returning to the step of inputting the training image into the initial image abnormality detection model until the training is completed to obtain the target image abnormality detection model. In this way, during model training, the current mapping label and the model parameters are adjusted synchronously, and the self-correction of the current mapping label corresponding to the training label with noise is performed using the feature learning ability of the model, which can effectively filter the performance impact of noise data on the model, greatly improve the performance of the model, and finally train an image abnormality detection model with higher accuracy. Subsequently, during model application, a higher-accuracy image abnormality detection result can be obtained based on the image abnormality detection model with higher accuracy.
[0100] In one embodiment, obtaining the training image and the current mapping label corresponding to the training label of the training image includes:
[0101] Performing label encoding on the training label based on the number of label categories corresponding to the training label to obtain the current mapping label.
[0102] Among them, the number of label categories refers to the number of label categories of the training label. For example, if the training label is a binary classification label, then the number of label categories is 2; if the training label is a four-classification label, then the number of label categories is 4. Label encoding is used to convert the training label into data represented by binary.
[0103] Specifically, the computer device can perform label encoding on the training labels based on the number of label categories corresponding to the training labels, and map and transform the training labels into initial current mapped labels. Specifically, the computer device can determine the number of vector dimensions corresponding to the initial current mapped labels based on the number of label categories. The initial vector values on each vector dimension are defaulted to the first preset value, and the initial vector values on each vector dimension are sequentially converted into the second preset value, so as to obtain the current mapped labels corresponding to the training labels of each label category. One vector dimension corresponds to one label category, and the vector values on each vector dimension represent the probability that the image belongs to the corresponding label category. For example, if the training label is a binary classification label and the number of label categories is 2, then the number of vector dimensions corresponding to the initial current mapped label is 2, that is, the current mapped label is a two-dimensional vector. The initial form of the two-dimensional vector can be [0, 0], that is, the first preset value is 0 and the second preset value is 1. Sequentially convert the initial values on each vector dimension into the second preset value, then the current mapped label corresponding to the positive label in the binary classification label is [0, 1], and the current mapped label corresponding to the negative label is [1, 0]. Taking [0, 1] as an example, 0 represents the probability that the image belongs to the negative label, and 1 represents the probability that the image belongs to the positive label. If the training label is a three-classification label and the number of label categories is 3, then the number of vector dimensions corresponding to the initial current mapped label is 3, that is, the current mapped label is a three-dimensional vector. The initial form of the three-dimensional vector can be [0, 0, 0], the current mapped label corresponding to the first label in the three-classification label is [0, 0, 1], the current mapped label corresponding to the second label is [0, 1, 0], and the current mapped label corresponding to the third label is [1, 0, 0]. The distances between the initial current mapped labels are the same. Of course, the first preset value can also be 1 and the second preset value can be 0. It can be understood that as the model and network are updated subsequently, the current mapped labels will be gradually softened. For example, the current mapped label [1, 0] corresponding to the negative label will be adjusted to [0.8, 0.2].
[0104] In this embodiment, label encoding is performed on the training labels based on the number of label categories corresponding to the training labels to obtain the initial current mapped labels. The initial current mapped labels are composed of binary data and the distances between them are the same, which are easy to be learned and calculated by the model.
[0105] In one embodiment, as Figure 3 shown, before inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image, the method further includes:
[0106] Step S302, input the training image into the candidate image anomaly detection model to obtain the initial prediction label corresponding to the training image.
[0107] Step S304: Based on the label difference between the training label and the initial prediction label, adjust the model parameters of the candidate image anomaly detection model until the first convergence condition is met, and obtain the initial image anomaly detection model.
[0108] Among them, the candidate image anomaly detection model is the original image anomaly detection model. The model parameters in the candidate image anomaly detection model can be randomly initialized or set to initial values manually.
[0109] Specifically, in addition to directly obtaining the candidate image anomaly detection model as the initial image anomaly detection model, the computer device can also pre-train the candidate image anomaly detection model and use the pre-trained candidate image anomaly detection model as the initial image anomaly detection model. When training the candidate image anomaly detection model, only adjust the model parameters and do not adjust the training labels, so that the subsequent model training for the initial image anomaly detection model can converge faster than training from scratch.
[0110] During training, the computer device inputs the training images into the candidate image anomaly detection model. After the candidate image anomaly detection model processes the input data, it can output the initial prediction labels corresponding to the training images. The computer device can generate a training loss based on the label difference between the training label and the initial prediction label, and perform backpropagation based on the training loss to update the model parameters of the candidate image anomaly detection model until the first convergence condition is met, and obtain the initial image anomaly detection model. Among them, the training loss can be calculated by calculating the distance between the training label and the initial prediction label. For example, the divergence between the initial prediction label and the training label is calculated to obtain the training loss, or the cross-entropy between the initial prediction label and the training label is calculated to obtain the training loss. The first convergence condition can be that the number of model iterations reaches the first preset threshold, the value of the training loss is less than the preset training value, etc.
[0111] In this embodiment, pre-training the candidate image anomaly detection model to obtain the initial image anomaly detection model can reduce the training difficulty of the initial image anomaly detection model and improve the training speed. It can be understood that if the candidate image anomaly detection model is directly obtained as the initial image anomaly detection model, due to the random initialization of the model parameters, it is difficult to update both the model parameters and the current mapping label at the same time, and the training may be unstable. Therefore, first fix the current mapping label at the beginning, only update the network parameters, and train the candidate image anomaly detection model to obtain the initial image anomaly detection model. Subsequently, when training the initial image anomaly detection model, update both the model parameters and the current mapping label at the same time, which can effectively reduce the training difficulty.
[0112] In one embodiment, generating model feedback data based on the current mapping label and the target prediction label includes:
[0113] Calculate the divergence between the current mapping label and the target prediction label to obtain the first loss; obtain the label distribution ratio corresponding to the training label, and obtain the second loss based on the label distribution ratio and the target prediction label; the label distribution ratio is determined based on the number of training images of each label category corresponding to the training label; obtain the model feedback data based on the first loss and the second loss.
[0114] Among them, divergence calculation is a way to measure the distribution difference between two data. The label distribution ratio is used to represent the distribution of training images among each label category. The label distribution ratio is determined based on the number of training images of each label category corresponding to the training label. Specifically, the ratio between the number of training images of each label category can be used as the label distribution ratio. For example, if the training label is a binary classification label, there are 30 images with a positive label and 30 images with a negative label, then the label distribution ratio is 30:30 = 1:1. If the training label is a three-classification label, there are 30 images with the first label, 30 images with the second label, and 60 images with the third label, then the label distribution ratio is 30:30:60 = 1:1:2.
[0115] Specifically, the computer device can calculate the divergence between the current mapping label and the target prediction label to obtain the first loss. The first loss is used to measure the difference between the prediction probability output by the current model and the manually annotated data. The first loss can also be called the classification loss. In addition to the first loss, the computer device can also obtain the label distribution ratio corresponding to the training label and calculate the second loss based on the label distribution ratio and the target prediction label. The second loss is used to measure the distance between the prediction probability output by the current model and the prior label distribution ratio. The second loss can also be called the prior loss. The purpose of the prior loss is to make the prediction probability not deviate too much from the label distribution ratio. The computer device can fuse the calculated first loss and second loss to obtain the model feedback data. For example, add the first loss and the second loss to obtain the model feedback data, or use the weighted sum of the first loss and the second loss as the model feedback data.
[0116] It can be understood that although the manually annotated labels are noisy, generally, the accuracy of most manually annotated labels is relatively high, that is, most manually annotated labels are correct, and only a small number of manually annotated labels are noisy. Therefore, by introducing the second loss, it is hoped that the target prediction label and the label distribution ratio will not differ too much, so as to improve the accuracy of the model feedback data.
[0117] In this embodiment, by calculating the divergence between the current mapping label and the target prediction label, a first loss is obtained. A second loss is obtained based on the label distribution ratio and the target prediction label. Model feedback data is obtained based on the first loss and the second loss. The model feedback data consists of different types of losses. By adjusting the model parameters and the current mapping label based on such model feedback data to train the model, the accuracy of model training can be improved.
[0118] In one embodiment, calculating the divergence between the current mapping label and the target prediction label to obtain the first loss includes:
[0119] Performing a logarithmic transformation on the ratio of the current mapping label and the target prediction label to obtain a label transformation ratio; fusing the label transformation ratio and the target prediction label to obtain the first loss.
[0120] Specifically, the computer device can calculate the ratio of the current mapping label and the target prediction label and perform a logarithmic transformation on this ratio to obtain the label transformation ratio for calculating the first loss. Specifically, the preset value can be used as the base, and the ratio of the current mapping label and the target prediction label can be used as the true number for logarithmic transformation to obtain the label transformation ratio. Alternatively, the preset value can be used as the base, and the fusion value of this ratio and the constant value can be used as the true number for logarithmic transformation to obtain the label transformation ratio. Among them, the preset value and the constant value can be set as needed. For example, the preset value and the constant value are data greater than 1. The fusion value can be the sum of the ratio of the current mapping label and the target prediction label and the constant value, or the product of the ratio of the current mapping label and the target prediction label and the constant value, etc. Furthermore, the computer device can fuse the target prediction label and the label transformation ratio to obtain the first loss. For example, the product of the target prediction label and the label transformation ratio is used as the first loss, and the weighted product of the target prediction label and the label transformation ratio is used as the first loss.
[0121] It can be understood that the closer the current mapping label and the target prediction label are, the closer the label transformation ratio is to zero. The optimization direction of the first loss is that the closer the current mapping label and the target prediction label are, the better.
[0122] In one embodiment, the calculation formula of the first loss is as follows:
[0123]
[0124] where L c represents the first loss, p represents the target prediction label, represents the current mapping label. represents the label transformation ratio. p and are vectors. When calculating L c , p and The vector values of the same vector dimension are calculated based on the above formula, and finally L is obtained based on the calculation results of each vector dimension. c 。
[0125] In this embodiment, by performing a logarithmic transformation on the ratio of the current mapping label and the target prediction label, a label transformation ratio is obtained. Fusing the label transformation ratio and the target prediction label can quickly obtain the first loss. The closer and more similar the current mapping label and the target prediction label are, the closer the first loss is to zero.
[0126] In one embodiment, calculating the divergence between the current mapping label and the target prediction label to obtain the first loss includes:
[0127] Statistical analysis is performed on the current mapping label and the target prediction label to obtain label statistical information; a logarithmic transformation is performed on the ratio of the label statistical information and the target prediction label to obtain a first transformation ratio, and a logarithmic transformation is performed on the ratio of the label statistical information and the current mapping label to obtain a second transformation ratio; the target prediction label and the first transformation ratio are fused to obtain a first sub-loss, and the second transformation ratio and the current mapping label are fused to obtain a second sub-loss; the first loss is obtained based on the first sub-loss and the second sub-loss.
[0128] Specifically, the computer device can perform statistical analysis on the current mapping label and the target prediction label, and use the statistical result as the label statistical information. For example, the mean value of the current mapping label and the target prediction label is used as the label statistical information, or the sum of the current mapping label and the target prediction label is used as the label statistical information. Furthermore, the computer device performs a logarithmic transformation on the ratio of the label statistical information and the target prediction label to obtain a first transformation ratio, and performs a logarithmic transformation on the ratio of the label statistical information and the current mapping label to obtain a second transformation ratio. For example, a preset value is used as the base, and the ratio of the label statistical information and the target prediction label is used as the true number for logarithmic transformation to obtain the first transformation ratio, and a preset value is used as the base, and the ratio of the label statistical information and the current mapping label is used as the true number for logarithmic transformation to obtain the second transformation ratio. The computer device fuses the target prediction label and the first transformation ratio to obtain a first sub-loss, and fuses the current mapping label and the second transformation ratio to obtain a second sub-loss. For example, the product of the target prediction label and the first transformation ratio is used as the first sub-loss, and the product of the current mapping label and the second transformation ratio is used as the second sub-loss. Finally, the computer device can obtain the first loss based on the first sub-loss and the second sub-loss. For example, the first loss is obtained by adding the first sub-loss and the second sub-loss, or the weighted sum of the first sub-loss and the second sub-loss is used as the first loss.
[0129] It can be understood that the closer the current mapping label and the target prediction label are, the closer the first transformation ratio and the second transformation ratio are to zero. The closer the current mapping label and the target prediction label are, the smaller the first loss is.
[0130] In one embodiment, the calculation formula of the first loss is as follows:
[0131]
[0132] where, L c represents the first loss, p represents the target prediction label, represents the current mapping label. represents the label statistical information, represents the first transformation ratio, represents the second transformation ratio. represents the first sub-loss, represents the second sub-loss.
[0133] In this embodiment, by statistically analyzing the current mapping label and the target prediction label, the label statistical information is obtained. The logarithm transformation is performed on the ratio of the label statistical information and the target prediction label to obtain the first transformation ratio. The logarithm transformation is performed on the ratio of the label statistical information and the current mapping label to obtain the second transformation ratio. The target prediction label and the first transformation ratio are fused to obtain the first sub-loss. The second transformation ratio and the current mapping label are fused to obtain the second sub-loss. Based on the first sub-loss and the second sub-loss, the first loss can be quickly obtained. The closer and more similar the current mapping label and the target prediction label are, the closer the first sub-loss and the second sub-loss are to zero, and further, the closer the first loss is to zero.
[0134] In one embodiment, obtaining the label distribution ratio corresponding to the training label and obtaining the second loss based on the label distribution ratio and the target prediction label includes:
[0135] Performing vectorization processing on the label distribution ratio to obtain a label distribution vector; performing cross-entropy calculation on the label distribution vector and the target prediction label to obtain the second loss.
[0136] Among them, the cross-entropy calculation is a calculation method for measuring the distribution difference between two data. The closer the label distribution vector and the target prediction label are, the smaller the second loss is.
[0137] Specifically, the label distribution ratio is used to represent the comparison relationship between the numbers of training images of each label category. For the convenience of calculation with the target predicted label, it is necessary to vectorize the label distribution ratio and convert it into a label distribution vector. The label distribution vector has the same vector dimension as the target predicted label, and the number of vector dimension quantities of the label distribution vector is the number of label categories. The computer device can use the ratio of the number of training images corresponding to each label category to the total number of training images as the vector value of each vector dimension in the label distribution vector. For example, if the training label is a binary classification label and the label distribution ratio is 30:30 = 1:1, then the label distribution vector can be [1 / 2, 1 / 2], that is, [0.5, 0.5]. Furthermore, the computer device can calculate the cross entropy between the label distribution vector and the target predicted label to obtain the second loss. Specifically, it can use a preset value as the base, perform a logarithmic transformation with the target predicted label as the true number, fuse the label distribution vector and the logarithmically transformed target predicted label, and take the opposite of the fusion result to obtain the second loss.
[0138] In one embodiment, the calculation formula of the second loss is as follows;
[0139] L p =-p prior *logp
[0140] where, L p represents the second loss, p prior represents the label distribution vector, and p represents the target predicted label.
[0141] In this embodiment, by vectorizing the label distribution ratio to obtain the label distribution vector and calculating the cross entropy between the label distribution vector and the target predicted label, the second loss can be quickly obtained.
[0142] In one embodiment, the model feedback data is obtained based on the first loss and the second loss, including:
[0143] Calculating the information entropy of the target predicted label to obtain the third loss; obtaining the model feedback data based on the first loss, the second loss, and the third loss.
[0144] Among them, the information entropy calculation is a calculation method for measuring the amount of information in data. Calculating the information entropy of a data is equivalent to calculating the cross entropy between the data and itself. The closer the target predicted label is to the current mapped label obtained by label encoding the training label, the smaller the third loss.
[0145] Specifically, the third loss is used to constrain the first loss to prevent the model from falling into local optima. It can be understood that during the network update process, since the information of the manually annotated labels is also updated synchronously (i.e., the current mapped labels are updated synchronously), and the purpose of network learning is to make the model output closer to the label information, then obviously, as long as the manually labeled information is updated to be exactly the same as the model output, in this case, the first loss is equal to 0. However, the model output at this time may be out of order because the model parameters are randomly initialized at the beginning of training. Therefore, to avoid this situation, the third loss is introduced, and the third loss is used as a positive term to constrain the first loss.
[0146] The computer device can calculate the third loss by calculating the information entropy of the target prediction label. Specifically, it can use a preset value as the base, take the target prediction label as the true number for logarithmic transformation, fuse the target prediction label and the logarithmically transformed target prediction label, and take the opposite of the fusion result to obtain the third loss. The optimization direction of the third loss is to make the model output as close as possible to the current mapped label obtained by label encoding the training label.
[0147] In one embodiment, the calculation formula of the third loss is as follows:
[0148] L e =-p*logp
[0149] where L e represents the third loss, and p represents the target prediction label.
[0150] In one embodiment, the model feedback data is obtained based on the first loss, the second loss, and the third loss, including:
[0151] Obtain the loss weights corresponding to the first loss, the second loss, and the third loss respectively; the loss weight corresponding to the second loss decreases as the proportion of noise images in the training images increases; fuse the first loss, the second loss, and the third loss based on the loss weights to obtain the model feedback data.
[0152] Among them, a noisy image refers to a training image with noise. A noisy image is an image for which it is difficult to determine the corresponding accurate label category from multiple label categories. That is, the label accuracy of the training label corresponding to the noisy image is lower than the preset accuracy. For example, the training label is a binary classification label, where the positive label represents a normal image and the negative training label represents a screen-frozen image. For an image with only slight or partial screen freezing, it is difficult to determine its corresponding accurate label. Some people will label it as the corresponding positive label, while some people will label it as the corresponding negative label. Therefore, the label annotation information of multiple users for the same training image can be collected, statistically analyzed for each label annotation information, and the label accuracy corresponding to the training image can be determined based on the statistical analysis results. For example, calculate the annotation ratio corresponding to each label category and select the largest annotation ratio as the label accuracy. The preset accuracy can be set as needed. For example, it can be set to 0.8. For example, for training image A, the label annotation information given by 3 users is the positive label, and the label annotation information given by 3 users is the negative label. Then the label accuracy is 3 / 6 = 0.5 < 0.8, so training image A is a noisy image.
[0153] The noisy image ratio refers to the ratio of the number of noisy images to the total number of training images. For example, if there are 100 training images and 20 of them are noisy images, then the noisy image ratio is 20 / 100 = 0.2.
[0154] Specifically, when obtaining the model feedback data based on the first loss, the second loss, and the third loss, the computer device can obtain the loss weights corresponding to the first loss, the second loss, and the third loss respectively, and perform weighted fusion on each loss based on the loss weights corresponding to each loss, so as to obtain the model feedback data. Among them, each loss weight can be a preset fixed value or can be used as a model parameter and adjusted and learned during the model training process.
[0155] Furthermore, since the purpose of the second loss is to make the model output not deviate too much from the label distribution ratio, if the proportion of noisy images in the training images is relatively large, then the loss weight corresponding to the second loss can be appropriately reduced to avoid the adverse impact of noisy images on model training. Therefore, the loss weight corresponding to the second loss can decrease as the noisy image ratio corresponding to the noisy images in the training images increases.
[0156] In one embodiment, the calculation formula for the model feedback data is as follows:
[0157] L = L c + αL e + βL p
[0158] Among them, L represents the model feedback data (which can also be called the loss function), L cRepresents the first loss, L p Represents the second loss, L e Represents the third loss. The loss weight corresponding to the first loss is 1, the loss weight corresponding to the second loss is β, and the loss weight corresponding to the third loss is α.
[0159] In this embodiment, by obtaining the loss weights corresponding to the first loss, the second loss, and the third loss respectively, and fusing the first loss, the second loss, and the third loss based on the loss weights to obtain model feedback data. Among them, the loss weight corresponding to the second loss decreases as the proportion of the noise images corresponding to the training images increases, so as to avoid the second loss affecting the training effect of model training in the case of more noise data.
[0160] In one embodiment, generating a label loss based on the data change reference information corresponding to the model feedback data, and adjusting the current mapping label based on the label loss to obtain an updated mapping label, including:
[0161] Calculating the gradient of the current mapping label based on the model feedback data to obtain data change reference information; obtaining the model learning rate, and adjusting the data change reference information based on the model learning rate to obtain a label loss; obtaining the updated mapping label based on the distance between the current mapping label and the label loss.
[0162] Among them, the model learning rate is a hyperparameter, which can be a preset fixed value or can be used as a model parameter to be adjusted and learned during the model training process.
[0163] Specifically, the computer device can calculate the gradient of the model feedback data with respect to the current mapping label, and use the calculation result as the data change reference information, that is, calculate the data change reference information by calculating the gradient of the model feedback data with respect to the current mapping label. It can be understood that since the self-entropy loss is obtained based on the target prediction label and has nothing to do with the current mapping label, even if the model feedback data includes the self-entropy loss, the data change reference information obtained by calculating the gradient of the model feedback data with respect to the current mapping label will not be affected by the self-entropy loss. Furthermore, the computer device can obtain the model learning rate, and fuse the model learning rate and the data change reference information to obtain a label loss. For example, the product of the model learning rate and the data change reference information is used as the label loss, or the weighted product of the model learning rate and the data change reference information is used as the label loss, etc. Then, the computer device calculates the distance between the current mapping label and the label loss, and uses the calculated distance as the updated mapping label. For example, the difference between the current mapping label and the label loss is used as the updated mapping label. Subsequently, the updated mapping label is used as the new current mapping label to participate in the next round of model iterative training.
[0164] In one embodiment, the current mapping label can be adjusted by the following formula:
[0165]
[0166] Wherein, represents the data change reference information, which is obtained by calculating the gradient of the loss function L with respect to , L represents the model feedback data (which can also be called the loss function), represents the current mapping label corresponding to the t-th moment (the t-th model iteration round), represents the current mapping label corresponding to the (t + 1)-th moment ((t + 1)-th model iteration round), that is, the updated mapping label obtained by adjusting the current mapping label, l r represents the model learning rate.
[0167] In this embodiment, the data change reference information is obtained by calculating the gradient of the current mapping label based on the model feedback data, the label loss is obtained by adjusting the data change reference information based on the model learning rate, and the updated mapping label is obtained based on the distance between the current mapping label and the label loss. In this way, by adjusting the current mapping label according to the loss gradient of the model feedback data with respect to the current mapping label, the current mapping label can be updated along the gradient descent direction, so that the target mapping label can be quickly adjusted.
[0168] In one embodiment, as Figure 4 shown, based on the model feedback data, the model parameters of the initial image anomaly detection model are adjusted to obtain an updated initial image anomaly detection model, and the step of inputting the training image into the initial image anomaly detection model is returned until the training is completed to obtain the target image anomaly detection model, including:
[0169] Step S402, based on the model feedback data, adjust the model parameters of the initial image anomaly detection model to obtain an updated initial image anomaly detection model, and return the step of inputting the training image into the initial image anomaly detection model until the second convergence condition is satisfied to obtain the intermediate image anomaly detection model and the target mapping label.
[0170] Among them, the intermediate image anomaly detection model is the model obtained by synchronously adjusting the model parameters and the current mapping label during model training. The target mapping label is the mapping label obtained after the initial image anomaly detection model converges.
[0171] Specifically, when training the initial image anomaly detection model, the computer device can update the model parameters and the current mapping labels simultaneously. The computer device can perform backpropagation based on the model feedback data to adjust the model parameters of the initial image anomaly detection model, obtain a new initial image anomaly detection model, input the training images into the new initial image anomaly detection model, obtain new target prediction labels, obtain new model feedback data based on the new current mapping labels and the new target prediction labels, and adjust the current mapping labels and the model parameters again based on the new model feedback data, and so on, and cycle through multiple training iterations until the second convergence condition is met, obtaining an intermediate image anomaly detection model and target mapping labels. Among them, the second convergence condition can be that the number of model iterations reaches the second iteration number, the value of the model feedback data is less than a preset target value, the change rate of the model feedback data is less than a preset change rate, the model feedback data is minimized, etc.
[0172] Step S404: Input the training images into the intermediate image anomaly detection model to obtain updated prediction labels corresponding to the training images.
[0173] Step S406: Generate an updated loss based on the updated prediction labels and the target mapping labels, and adjust the model parameters of the intermediate image anomaly detection model based on the updated loss until the third convergence condition is met, obtaining a target image anomaly detection model.
[0174] Specifically, after obtaining the intermediate image anomaly detection model, the computer device can fix the target mapping labels and perform further model training on the intermediate image anomaly detection model to fine-tune the model parameters and obtain a target image anomaly detection model. When training the intermediate image anomaly detection model, the computer device can input the training images into the intermediate image anomaly detection model to obtain updated prediction labels corresponding to the training images, generate an updated loss based on the data distribution difference between the updated prediction labels and the target mapping labels, and perform backpropagation based on the updated loss to adjust the model parameters of the intermediate image anomaly detection model until the third convergence condition is met, obtaining a target image anomaly detection model. Among them, the third convergence condition can be that the number of model iterations reaches the third iteration number, the value of the updated loss is less than a preset updated value, etc.
[0175] It can be understood that similar to the model feedback data, the updated loss can be obtained based on at least one loss. For example, the updated loss includes a classification loss calculated based on the target mapping labels and the updated prediction labels and a prior loss calculated based on the training labels and the updated prediction labels, and the updated loss includes a classification loss calculated based on the target mapping labels and the updated prediction labels, a prior loss calculated based on the training labels and the updated prediction labels, and a self-entropy loss calculated based on the updated prediction labels. The calculation of various losses can refer to the methods described in the foregoing various related embodiments and will not be elaborated here.
[0176] In this embodiment, when training the initial image anomaly detection model, the current mapping label and the model parameters are updated synchronously until the second convergence condition is met, obtaining an intermediate image anomaly detection model and a target mapping label. Then, the target mapping label is fixed, and the model parameters of the intermediate image anomaly detection model are further updated until the third convergence condition is reached, obtaining the target image anomaly detection model. The above training process can further improve the accuracy of model training.
[0177] In one embodiment, generating an update loss based on the updated prediction label and the target mapping label includes:
[0178] Calculating the divergence between the updated prediction label and the target mapping label to obtain a fourth loss; calculating the information entropy of the updated prediction label to obtain a fifth loss; calculating the cross-entropy based on the training label and the updated prediction label to obtain a sixth loss; and obtaining the update loss based on the fourth loss, the fifth loss, and the sixth loss.
[0179] Specifically, the update loss can be obtained based on the classification loss, the prior loss, and the self-entropy loss. The classification loss can be obtained by calculating the divergence between the updated prediction label and the target mapping label. That is, the computer device can calculate the divergence between the updated prediction label and the target mapping label to obtain the fourth loss. The prior loss can be obtained by calculating the information entropy of the updated prediction label. That is, the computer device can calculate the information entropy of the updated prediction label to obtain the fifth loss. The prior loss can be obtained by calculating the cross-entropy between the training label and the updated prediction label. That is, the computer device can calculate the cross-entropy based on the training label and the updated prediction label to obtain the sixth loss. Finally, the computer device can obtain the update loss based on the fourth loss, the fifth loss, and the sixth loss. For example, adding the fourth loss, the fifth loss, and the sixth loss to obtain the update loss, or performing weighted fusion on the fourth loss, the fifth loss, and the sixth loss to obtain the update loss.
[0180] In this embodiment, the fourth loss is obtained by calculating the divergence between the updated prediction label and the target mapping label, the fifth loss is obtained by calculating the information entropy of the updated prediction label, the sixth loss is obtained by calculating the cross-entropy based on the training label and the updated prediction label, and the update loss is obtained based on the fourth loss, the fifth loss, and the sixth loss. The update loss consists of multiple different types of losses. Adjusting the model parameters based on such an update loss to train the model can improve the accuracy of model training.
[0181] In one embodiment, as Figure 5 shown, an image anomaly detection method is provided, and this method is applied to Figure 1Taking the computer in it as an example for illustration, it can be understood that the computer device can be the terminal 102 or the server 104. In this embodiment, the image anomaly detection method includes the following steps:
[0182] Step S502, obtain the image to be detected.
[0183] Step S504, input the image to be detected into the target image anomaly detection model to obtain the model prediction label corresponding to the image to be detected.
[0184] Among them, the image to be detected refers to the image to be detected for anomalies. The image to be detected can be a picture or a video frame in a video. The model prediction label refers to the prediction label corresponding to the image to be detected.
[0185] Specifically, the computer device can obtain the image to be detected and the target image anomaly detection model locally, or from other terminals or servers, input the image to be detected into the target image anomaly detection model, and obtain the model prediction label corresponding to the image to be detected.
[0186] Among them, the training process of the target image anomaly detection model includes: obtaining the training image and the current mapping label corresponding to the training label of the training image; inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image; generating model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image; generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0187] It can be understood that the specific training process of the target image anomaly detection model can refer to the methods described in the relevant embodiments of the image anomaly detection model training method, which will not be elaborated here.
[0188] Step S506, determine the image anomaly detection result corresponding to the image to be detected based on the model prediction label.
[0189] Specifically, after obtaining the model prediction label, the computer device can determine the image anomaly detection result corresponding to the image to be detected based on the model prediction label. The computer device can set the target image anomaly detection model to output complete data, that is, the model prediction label is a prediction vector, and each vector value in the prediction vector represents the probability that the image to be detected belongs to the corresponding label category. The computer device can take the probability corresponding to the normal label category from the prediction vector as the target confidence level. If the target confidence level is greater than the preset confidence level, it is determined that the image to be detected is a normal image, and the image anomaly detection result of the image to be detected is that the image has no anomaly. If the target confidence level is less than or equal to the preset confidence level, it is determined that the image to be detected is an abnormal image, and the image anomaly detection result of the image to be detected is that the image is abnormal. If the model prediction label is a prediction vector, the computer device can also take the vector values greater than the preset vector value in the prediction vector as the target confidence level. If the target confidence level is greater than the preset confidence level, it is determined that the image anomaly detection result of the image to be detected is the label category corresponding to the target confidence level.
[0190] Of course, the computer device can also set the target image anomaly detection model to only output the probability corresponding to the normal label category as the model prediction label. If the model prediction label is greater than the preset confidence level, it is determined that the image to be detected is a normal image. If the model prediction label is less than or equal to the preset confidence level, it is determined that the image to be detected is an abnormal image. Among them, the preset confidence level can be set as needed. For example, it can be set to 0.5.
[0191] The image anomaly detection method of the present application can be applied to the quality analysis tasks of images or videos. For example, in a social application, a computer device obtains an image to be shared uploaded by a user, and performs image anomaly detection on the image to be shared based on a target image anomaly detection model. If the image anomaly detection result is that the image is normal, it is determined that the image to be shared meets the sharing conditions, and then the image to be shared is published in the social application for other users to view. If the image anomaly detection result is that the image is abnormal, the user can be prompted that the image quality is poor and the user is prompted to re-upload the image. In a video application, a computer device obtains a video to be shared uploaded by a user, and performs image anomaly detection on the video frames in the video to be shared based on a target image anomaly detection model. If the proportion of video frames with normal image anomaly detection results is greater than a preset proportion, it is determined that the video to be shared meets the sharing conditions, and then the video to be shared is published in the video application for other users to view. In a vehicle surrounding environment monitoring application, a computer device can obtain a monitoring video collected by an in-vehicle terminal, and perform image anomaly detection on the video frames in the monitoring video based on a target image anomaly detection model. If the proportion of video frames with normal image anomaly detection results is greater than a preset proportion, the monitoring video is stored or further data analysis is performed on the monitoring video to determine the environmental state of the vehicle. If the proportion of video frames with normal image anomaly detection results is less than or equal to the preset proportion, the in-vehicle terminal is instructed to re-collect the monitoring video.
[0192] In the above image anomaly detection method, by obtaining the training image and the current mapping label corresponding to the training label of the training image, inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image, generating model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image, generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label to obtain an updated mapping label based on the label loss, adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model. In this way, during model training, the current mapping label and the model parameters are adjusted synchronously, and the self-correction of the current mapping label corresponding to the noisy training label is performed using the feature learning ability of the model, which can effectively filter the performance impact of the noisy data on the model, greatly improve the performance of the model, and finally train an image anomaly detection model with higher accuracy. Therefore, when the model is applied, image anomaly detection is performed on the image to be detected based on the image anomaly detection model with higher accuracy, and an image anomaly detection result with higher accuracy can be obtained.
[0193] In a specific embodiment, the above-mentioned image anomaly detection model training method and image anomaly detection method can be applied to the task of detecting image screen tearing. Image screen tearing detection refers to detecting whether there is a screen tearing phenomenon in a picture. The task of detecting image screen tearing is an essential step in the quality analysis of pictures and videos and can be used to evaluate the quality of the current picture or video.
[0194] Reference Figure 6 , the screen-tearing images are not a simple binary classification. Many screen-tearing images are only slightly or partially screen-tearing. Therefore, simple binary labels may be subjective and, as a result, the manually labeled information corresponding to such screen-tearing images is noisy. Training a model with such noisy labels will reduce the performance of the model. Through the image anomaly detection model training method of the present application, the feature learning ability of the model can be used to identify and correct the artificial label data containing noise, thereby significantly improving the performance of the model. Finally, the trained model can output more accurate prediction results in the screen-tearing detection task.
[0195] Reference Figure 7 , during model training, the image is input into the image screen-tearing detection model, and through the data calculation of the deep learning model, the prediction probability (i.e., the prediction label) of whether the input image is screen-tearing can be obtained. At the same time, the artificial label corresponding to the input image is converted into a mapped label through label encoding. Among them, the artificial label refers to the manually labeled label, which can be a 0 / 1 binary label representing the human judgment result of whether the image has screen tearing or not. The loss function of the image screen-tearing detection model includes three losses, specifically, self-entropy loss, classification loss, and prior loss. The self-entropy loss is calculated based on the prediction probability and is the result of the prediction probability and its own calculated entropy value. This loss is a regularization term, and its purpose is to prevent the model from falling into a local optimum. The classification loss is obtained based on the prediction probability and the mapped label and is used to measure the distance between the prediction probability of the model and the artificial label. The prior loss is obtained based on the proportion of positive and negative label distributions in the prediction probability and the artificial label. This loss is also a regularization term. After the three losses are superimposed, they simultaneously guide the parameter update of the deep learning model and the label update of the mapped label corresponding to the artificial label. The updated mapped label is used as the calculation data corresponding to the loss function for the next round of model iteration. After multiple rounds of model iteration, the finally trained image screen-tearing detection model and the corrected artificial label result (i.e., the target mapped label) are obtained.
[0196] It can be understood that Figure 7 the solid arrow in Figure 6 represents the forward propagation process of the model, that is, the inference process,
[0197] Furthermore, the training process of the model can be divided into three steps. In Step 1, the model is trained using the original manually annotated labels without updating the labels. The purpose of this step is that at the initial stage of model training, the model parameters are randomly initialized, and it is difficult to update both the model parameters and the label distribution simultaneously, which may lead to unstable training. Therefore, at the beginning, the mapping labels are fixed, and only the model parameters are updated. The loss function at this stage can include only the classification loss, and this stage can be trained for K1 epochs. In Step 2, both the model parameters and the mapping labels are updated. The loss function at this stage can include three types of losses, and this stage can be trained for K2 epochs. In Step 3, the mapping labels obtained from the iteration in Step 2 are fixed, and the mapping labels are no longer updated. Only the model parameters are updated, and the model is updated for K3 more epochs. The loss function at this stage includes three types of losses. Finally, the entire training process requires K1 + K2 + K3 epochs. After training is completed, a trained image screen freeze detection model is obtained.
[0198] It can be understood that the training process of the model may not adopt the three-step training method of K1 + K2 + K3. For example, a training method of K1 + K2 can be adopted.
[0199] The trained image screen freeze detection model can be used to predict the screen freeze probability (i.e., screen freeze confidence) of the input image. Referring to Figure 8 , when the screen freeze confidence predicted by the model is greater than the preset confidence, it is determined that the input image has a screen freeze and is a screen freeze image. When the screen freeze confidence predicted by the model is less than or equal to the preset confidence, it is determined that the input image does not have a screen freeze and is a normal image. Among them, the preset threshold can be set as needed. For example, it can be set to 0.5.
[0200] In this embodiment, based on the trained image screen freeze detection model, the screen freeze degree of the image can be accurately detected. During model training, by self-correcting the artificial labels with noise, the performance impact of noise data on the model is effectively filtered, so as to train an accurate image screen freeze detection model. When the model is applied, based on the accurate image screen freeze detection model, stable and reliable screen freeze detection results can be output, thus providing reliable technical support for video quality assessment.
[0201] It can be understood that the above model training method can be applied not only to the image screen freeze detection model, but also to image anomaly detection models such as image blur detection models and image mosaic detection models.
[0202] It should be understood that although Figures 2 - 5The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 - 5 At least a part of the steps in
[0203] In one embodiment, as Figure 9 shown, an image anomaly detection model training device is provided. This device can be a software module, a hardware module, or a combination of both to form a part of a computer device. Specifically, the device includes: a training data acquisition module 902, a model prediction module 904, a model feedback data determination module 906, a label adjustment module 908, and a model adjustment module 910, where:
[0204] The training data acquisition module 902 is used to obtain the current mapped label corresponding to the training image and the training label of the training image.
[0205] The model prediction module 904 is used to input the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image.
[0206] The model feedback data determination module 906 is used to generate model feedback data based on the current mapped label corresponding to the training label of the training image and the target prediction label.
[0207] The label adjustment module 908 is used to generate a label loss based on the data change reference information corresponding to the model feedback data, adjust the current mapped label based on the label loss to obtain an updated mapped label, and use the updated mapped label as the current mapped label.
[0208] The model adjustment module 910 is used to adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0209] In one embodiment, the training data acquisition module is further used to perform label encoding on the training label based on the number of label categories corresponding to the training label to obtain the current mapped label.
[0210] In one embodiment, the image anomaly detection model training device further includes:
[0211] A model pre-training module for inputting training images into a candidate image anomaly detection model to obtain initial prediction labels corresponding to the training images; and adjusting model parameters of the candidate image anomaly detection model based on the label differences between the training labels and the initial prediction labels until a first convergence condition is met, thereby obtaining an initial image anomaly detection model.
[0212] In one embodiment, the model feedback data determination module includes:
[0213] A first loss determination unit for calculating the divergence between the current mapped label and the target prediction label to obtain a first loss.
[0214] A second loss determination unit for obtaining the label distribution ratio corresponding to the training labels, and obtaining a second loss based on the label distribution ratio and the target prediction label; the label distribution ratio is determined based on the number of training images in each label category corresponding to the training labels.
[0215] A model feedback data determination unit for obtaining model feedback data based on the first loss and the second loss.
[0216] In one embodiment, the first loss determination unit is further configured to perform a logarithmic transformation on the ratio of the current mapped label to the target prediction label to obtain a label transformation ratio; and fuse the label transformation ratio and the target prediction label to obtain a first loss.
[0217] In one embodiment, the first loss determination unit is further configured to perform statistics on the current mapped label and the target prediction label to obtain label statistical information; perform a logarithmic transformation on the ratio of the label statistical information to the target prediction label to obtain a first transformation ratio, and perform a logarithmic transformation on the ratio of the label statistical information to the current mapped label to obtain a second transformation ratio; fuse the target prediction label and the first transformation ratio to obtain a first sub-loss, and fuse the current mapped label and the second transformation ratio to obtain a second sub-loss; and obtain a first loss based on the first sub-loss and the second sub-loss.
[0218] In one embodiment, the second loss determination unit is further configured to perform a vectorization process on the label distribution ratio to obtain a label distribution vector; and perform a cross-entropy calculation on the label distribution vector and the target prediction label to obtain a second loss.
[0219] In one embodiment, the model feedback data determination unit is further configured to perform an information entropy calculation on the target prediction label to obtain a third loss; and obtain model feedback data based on the first loss, the second loss, and the third loss.
[0220] In one embodiment, the model feedback data determination unit is further configured to obtain loss weights corresponding to a first loss, a second loss, and a third loss respectively; the loss weight corresponding to the second loss decreases as the proportion of noise images corresponding to the noise images in the training images increases; and fuse the first loss, the second loss, and the third loss based on the loss weights to obtain model feedback data.
[0221] In one embodiment, the label adjustment module is further configured to calculate a gradient of the current mapped label based on the model feedback data to obtain data change reference information; obtain a model learning rate, adjust the data change reference information based on the model learning rate to obtain a label loss; and obtain an updated current mapped label based on the distance between the current mapped label and the label loss.
[0222] In one embodiment, the model adjustment module includes:
[0223] A first adjustment unit, configured to adjust model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training images into the initial image anomaly detection model until a second convergence condition is satisfied, to obtain an intermediate image anomaly detection model and a target mapped label.
[0224] A prediction unit, configured to input the training images into the intermediate image anomaly detection model to obtain updated prediction labels corresponding to the training images.
[0225] A second adjustment unit, configured to generate an updated loss based on the updated prediction labels and the target mapped label, and adjust model parameters of the intermediate image anomaly detection model based on the updated loss until a third convergence condition is satisfied, to obtain a target image anomaly detection model.
[0226] In one embodiment, the second adjustment unit is further configured to calculate a divergence between the updated prediction labels and the target mapped label to obtain a fourth loss; calculate an information entropy of the updated prediction labels to obtain a fifth loss; calculate a cross entropy based on the training labels and the updated prediction labels to obtain a sixth loss; and obtain the updated loss based on the fourth loss, the fifth loss, and the sixth loss.
[0227] In one embodiment, the target image anomaly detection model is any one of an image screen freeze detection model, an image blur detection model, and an image mosaic detection model.
[0228] The above-mentioned image anomaly detection model training device, when training the model, synchronously adjusts the training labels and the current mapping labels, and uses the feature learning ability of the model to self-correct the current mapping labels corresponding to the noisy training labels, which can effectively filter the performance impact of noisy data on the model, greatly improve the performance of the model, and finally train an image anomaly detection model with relatively high accuracy. Subsequently, when the model is applied, relatively high-accuracy image anomaly detection results can be obtained based on the image anomaly detection model with relatively high accuracy.
[0229] In one embodiment, as Figure 10 shown, an image anomaly detection device is provided. This device can be a software module, a hardware module, or a combination of both to become part of a computer device. Specifically, the device includes: an image acquisition module 1002, a label prediction module 1004, and a detection result determination module 1006, where:
[0230] The image acquisition module 1002 is used to acquire the image to be detected.
[0231] The label prediction module 1004 is used to input the image to be detected into the target image anomaly detection model to obtain the model prediction label corresponding to the image to be detected.
[0232] The detection result determination module 1006 is used to determine the image anomaly detection result corresponding to the image to be detected based on the model prediction label.
[0233] Among them, the training process of the target image anomaly detection model includes: acquiring the training image and the current mapping label corresponding to the training label of the training image; inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image; generating model feedback data based on the current mapping label and the target prediction label corresponding to the training label of the training image; generating a label loss based on the data change reference information corresponding to the model feedback data, and adjusting the current mapping label to obtain an updated mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
[0234] The above-mentioned image anomaly detection device, when training the model, synchronously adjusts the training labels and the current mapping labels, and uses the feature learning ability of the model to self-correct the current mapping labels corresponding to the noisy training labels, which can effectively filter the performance impact of noisy data on the model, greatly improve the performance of the model, and finally train an image anomaly detection model with relatively high accuracy. Thus, when the model is applied, relatively high-accuracy image anomaly detection results can be obtained based on the image anomaly detection model with relatively high accuracy.
[0235] For the specific limitations of the image anomaly detection model training device and the image anomaly detection device, reference may be made to the limitations of the image anomaly detection model training method and the image anomaly detection method in the foregoing text, which will not be elaborated here. Each module in the above-mentioned image anomaly detection model training device and image anomaly detection device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0236] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as training images, candidate image anomaly detection models, initial image anomaly detection models, intermediate image anomaly detection models, and target image anomaly detection models. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an image anomaly detection model training method and an image anomaly detection method.
[0237] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 12 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an image anomaly detection model training method and an image anomaly detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0238] Those skilled in the art can understand that Figure 11 , 12 the structures shown in are only block diagrams of some structures related to the solution of this application, and do not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0239] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0240] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0241] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the steps in the above method embodiments.
[0242] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0243] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0244] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for training an image anomaly detection model, characterized in that The method includes: Obtaining a training image and a current mapping label corresponding to the training label of the training image; Inputting the training image into an initial image anomaly detection model to obtain a target prediction label corresponding to the training image; Calculating the divergence between the current mapping label and the target prediction label to obtain a first loss, including: performing statistics on the current mapping label and the target prediction label to obtain label statistical information, performing a logarithmic transformation on the ratio of the label statistical information to the target prediction label to obtain a first transformation ratio, performing a logarithmic transformation on the ratio of the label statistical information to the current mapping label to obtain a second transformation ratio, fusing the first transformation ratio and the target prediction label to obtain a first sub-loss, fusing the second transformation ratio and the current mapping label to obtain a second sub-loss, and obtaining the first loss based on the first sub-loss and the second sub-loss; Obtaining model feedback data based on the first loss; Generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; Adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain a target image anomaly detection model.
2. The method according to claim 1, wherein The obtaining of the training image and the current mapping label corresponding to the training label of the training image includes: Performing label encoding on the training label based on the number of label categories corresponding to the training label to obtain the current mapping label.
3. The method according to claim 1, characterized in that, Before the step of inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image, the method further includes: Inputting the training image into a candidate image anomaly detection model to obtain an initial prediction label corresponding to the training image; Adjusting the model parameters of the candidate image anomaly detection model based on the label difference between the training label and the initial prediction label until a first convergence condition is met to obtain an initial image anomaly detection model.
4. The method according to claim 1, wherein The obtaining of the model feedback data based on the first loss includes: Obtaining the label distribution ratio corresponding to the training label, and obtaining a second loss based on the label distribution ratio and the target prediction label; the label distribution ratio is determined based on the number of training images of each label category corresponding to the training label; Obtaining the model feedback data based on the first loss and the second loss.
5. The method according to claim 1, wherein The calculating of the divergence between the current mapping label and the target prediction label to obtain a first loss further includes: Performing a logarithmic transformation on the ratio of the current mapping label and the target prediction label to obtain a label transformation ratio; Fusing the label transformation ratio and the target prediction label to obtain the first loss.
6. The method according to claim 4, wherein The obtaining of the label distribution ratio corresponding to the training label and obtaining a second loss based on the label distribution ratio and the target prediction label includes: Performing a vectorization process on the label distribution ratio to obtain a label distribution vector; Calculate the cross entropy of the label distribution vector and the target prediction label to obtain the second loss.
7. The method according to claim 4, characterized in that, Obtaining the model feedback data based on the first loss and the second loss includes: Calculate the information entropy of the target prediction label to obtain the third loss; Obtain the model feedback data based on the first loss, the second loss, and the third loss.
8. The method according to claim 7, wherein Obtaining the model feedback data based on the first loss, the second loss, and the third loss includes: Obtain the loss weights corresponding to the first loss, the second loss, and the third loss respectively; the loss weight corresponding to the second loss decreases as the proportion of noise images corresponding to the noise images in the training images increases; Fuse the first loss, the second loss, and the third loss based on the loss weights to obtain the model feedback data.
9. The method according to claim 1, wherein Generating a label loss based on the data change reference information corresponding to the model feedback data, and adjusting the current mapped label based on the label loss to obtain an updated mapped label, includes: Calculate the gradient of the current mapped label based on the model feedback data to obtain the data change reference information; Obtain the model learning rate, and adjust the data change reference information based on the model learning rate to obtain the label loss; Obtain the updated mapped label based on the distance between the current mapped label and the label loss.
10. The method according to claim 1, characterized in that Adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model, includes: Adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the second convergence condition is met to obtain the intermediate image anomaly detection model and the target mapped label; Input the training image into the intermediate image anomaly detection model to obtain the updated prediction label corresponding to the training image; Generate an updated loss based on the updated prediction label and the target mapped label, and adjust the model parameters of the intermediate image anomaly detection model based on the updated loss until the third convergence condition is met to obtain the target image anomaly detection model.
11. The method according to claim 10, wherein Generating the updated loss based on the updated prediction label and the target mapped label includes: Calculate the divergence of the updated prediction label and the target mapped label to obtain the fourth loss; Calculate the information entropy of the updated prediction label to obtain the fifth loss; Calculate the cross entropy based on the training label and the updated prediction label to obtain the sixth loss; Obtain the updated loss based on the fourth loss, the fifth loss, and the sixth loss.
12. The method according to any one of claims 1 to 11, characterized in that The target image anomaly detection model is any one of an image screen freeze detection model, an image blur detection model, and an image mosaic detection model.
13. An image anomaly detection method, characterized in that, The method includes: Obtain the image to be detected; Input the image to be detected into the target image anomaly detection model to obtain the model prediction label corresponding to the image to be detected; Determine the image anomaly detection result corresponding to the image to be detected based on the model prediction label; Among them, the training process of the target image anomaly detection model includes: obtaining the training image and the current mapping label corresponding to the training label of the training image; inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image; calculating the divergence between the current mapping label and the target prediction label to obtain the first loss, including: counting the current mapping label and the target prediction label to obtain label statistical information, performing a logarithmic transformation on the ratio of the label statistical information to the target prediction label to obtain the first transformation ratio, performing a logarithmic transformation on the ratio of the label statistical information to the current mapping label to obtain the second transformation ratio, fusing the first transformation ratio and the target prediction label to obtain the first sub-loss, fusing the second transformation ratio and the current mapping label to obtain the second sub-loss, and obtaining the first loss based on the first sub-loss and the second sub-loss; obtaining the label distribution ratio corresponding to the training label, and obtaining the second loss based on the label distribution ratio and the target prediction label; the label distribution ratio is determined based on the number of training images of each label category corresponding to the training label; obtaining model feedback data based on the first loss and the second loss; generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
14. The method according to claim 13, wherein The image to be detected is an image to be shared uploaded by the user, and the method further includes: When the image anomaly detection result is that the image is normal, determine that the image to be shared meets the sharing conditions, and publish the image to be shared to the social application.
15. An image anomaly detection model training device, characterized in that, The device includes: A training data acquisition module, configured to acquire a training image and a current mapping label corresponding to the training label of the training image; A model prediction module, configured to input the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image; A model feedback data determination module, configured to calculate the divergence between the current mapping label and the target prediction label to obtain a first loss, including: performing statistics on the current mapping label and the target prediction label to obtain label statistical information, performing a logarithmic transformation on the ratio of the label statistical information to the target prediction label to obtain a first transformation ratio, performing a logarithmic transformation on the ratio of the label statistical information to the current mapping label to obtain a second transformation ratio, fusing the first transformation ratio and the target prediction label to obtain a first sub-loss, fusing the second transformation ratio and the current mapping label to obtain a second sub-loss, and obtaining the first loss based on the first sub-loss and the second sub-loss; obtaining model feedback data based on the first loss; A label adjustment module, configured to generate a label loss based on the data change reference information corresponding to the model feedback data, adjust the current mapping label based on the label loss to obtain an updated mapping label, and use the updated mapping label as the current mapping label; A model adjustment module, configured to adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain a target image anomaly detection model.
16. The device according to claim 15, characterized in that, The training data acquisition module is further configured to: Perform label encoding on the training label based on the number of label categories corresponding to the training label to obtain a current mapping label.
17. The device according to claim 15, characterized in that The apparatus further includes: A model pre-training module, configured to input the training image into a candidate image anomaly detection model to obtain an initial prediction label corresponding to the training image; adjust the model parameters of the candidate image anomaly detection model based on the label difference between the training label and the initial prediction label until a first convergence condition is met to obtain an initial image anomaly detection model.
18. The device according to claim 15, characterized in that, The model feedback data determination module is further configured to: Obtain the label distribution ratio corresponding to the training label, and obtain a second loss based on the label distribution ratio and the target prediction label; the label distribution ratio is determined based on the number of training images of each label category corresponding to the training label; Obtain the model feedback data based on the first loss and the second loss.
19. The device according to claim 15, characterized in that, The model feedback data determination module is further configured to: Perform a logarithmic transformation on the ratio of the current mapping label to the target prediction label to obtain a label transformation ratio; Fuse the label transformation ratio and the target prediction label to obtain the first loss.
20. The device according to claim 18, characterized in that, The model feedback data determination module is further configured to: Perform a vectorization process on the label distribution ratio to obtain a label distribution vector; Perform a cross-entropy calculation on the label distribution vector and the target prediction label to obtain the second loss.
21. The device according to claim 18, wherein The model feedback data determination module is further configured to: Perform an information entropy calculation on the target prediction label to obtain a third loss; Obtain the model feedback data based on the first loss, the second loss, and the third loss.
22. The device according to claim 21, wherein, The model feedback data determination module is further configured to: Obtain the loss weights corresponding to the first loss, the second loss, and the third loss respectively; the loss weight corresponding to the second loss decreases as the proportion of the noise image corresponding to the noise image in the training image increases; Fuse the first loss, the second loss, and the third loss based on the loss weights to obtain the model feedback data.
23. The device according to claim 15, characterized in that, The label adjustment module is further configured to: Calculate the gradient of the current mapping label based on the model feedback data to obtain the data change reference information; Obtain the model learning rate, and adjust the data change reference information based on the model learning rate to obtain the label loss; Obtain the updated mapping label based on the distance between the current mapping label and the label loss.
24. The device according to claim 15, characterized in that, The model adjustment module includes: A first adjustment unit, configured to adjust the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and return to the step of inputting the training image into the initial image anomaly detection model until the second convergence condition is satisfied, to obtain an intermediate image anomaly detection model and a target mapping label; A prediction unit, configured to input the training image into the intermediate image anomaly detection model to obtain an updated prediction label corresponding to the training image; A second adjustment unit, configured to generate an updated loss based on the updated prediction label and the target mapping label, and adjust the model parameters of the intermediate image anomaly detection model based on the updated loss until the third convergence condition is satisfied, to obtain the target image anomaly detection model.
25. The device according to claim 24, wherein The second adjustment unit is further configured to: Calculate the divergence between the updated prediction label and the target mapping label to obtain a fourth loss; Calculate the information entropy of the updated prediction label to obtain a fifth loss; Calculate the cross entropy based on the training label and the updated prediction label to obtain a sixth loss; Obtain the updated loss based on the fourth loss, the fifth loss, and the sixth loss.
26. The device according to any one of claims 15 to 25, characterized in that The target image anomaly detection model is any one of an image screen freeze detection model, an image blur detection model, and an image mosaic detection model.
27. An image anomaly detection device, characterized in that, The device includes: An image acquisition module, configured to acquire an image to be detected; A label prediction module, configured to input the image to be detected into the target image anomaly detection model to obtain a model prediction label corresponding to the image to be detected; A detection result determination module, configured to determine the image anomaly detection result corresponding to the image to be detected based on the model prediction label; Among them, the training process of the target image anomaly detection model includes: obtaining the training image and the current mapping label corresponding to the training label of the training image; inputting the training image into the initial image anomaly detection model to obtain the target prediction label corresponding to the training image; calculating the divergence between the current mapping label and the target prediction label to obtain the first loss, including: performing statistics on the current mapping label and the target prediction label to obtain label statistical information, performing a logarithmic transformation on the ratio of the label statistical information to the target prediction label to obtain the first transformation ratio, performing a logarithmic transformation on the ratio of the label statistical information to the current mapping label to obtain the second transformation ratio, fusing the first transformation ratio and the target prediction label to obtain the first sub-loss, fusing the second transformation ratio and the current mapping label to obtain the second sub-loss, and obtaining the first loss based on the first sub-loss and the second sub-loss; obtaining model feedback data based on the first loss; generating a label loss based on the data change reference information corresponding to the model feedback data, adjusting the current mapping label based on the label loss to obtain an updated mapping label, and using the updated mapping label as the current mapping label; adjusting the model parameters of the initial image anomaly detection model based on the model feedback data to obtain an updated initial image anomaly detection model, and returning to the step of inputting the training image into the initial image anomaly detection model until the training is completed to obtain the target image anomaly detection model.
28. The device according to claim 27, characterized in that, The image to be detected is the image to be shared uploaded by the user, and the device is further configured to: When the image anomaly detection result is that the image is normal, determine that the image to be shared meets the sharing condition, and publish the image to be shared to the social application.
29. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 14.
30. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 14.
31. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Image processing method and device, server and storage medium
CN110866908A
Target object detection method and device, computer equipment and storage medium
CN112766244A