Image Processing Method, Apparatus, Storage Medium, and Computer Device

By adopting unsupervised pre-training and supervised training methods in image detection, combined with texture enhancement processing and the use of classifiers, the problem of overfitting simple distinctive feature detection in the prior art is solved, and the effect of image detection and live resolution capabilities are improved.

CN115100706BActive Publication Date: 2025-07-01ALIBABA (CHINA) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202210676305.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-07-01
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

In the prior art, image detection methods overfit the detection of simple prominent features, resulting in poor detection effects.

Method used

By using the first image sample set for unsupervised pre-training, a pre-trained classifier was obtained; then using the second image sample set for supervised training to obtain the target classifier; then using the target encoder to perform texture enhancement processing on the target image to generate a clue map; superimpose the target image and the clue map to obtain a texture enhancement map; finally using the target classifier to classify the texture enhancement map to obtain the image detection result.

Benefits of technology

The dependence of the facial detection model on simple features is reduced, the underlying texture features is enhanced, the living resolution ability of image detection is improved, and the detection effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100706B_ABST
    Figure CN115100706B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method, apparatus, storage medium, and computer device. Among them, the method includes: performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; performing texture enhancement processing on a target image using a target encoder to obtain a clue map of the target image; superimposing the target image and the clue map to obtain a texture-enhanced image; and classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image. The present invention solves the technical problem in the prior art that the image detection method has an overfitting situation for the detection of simple and prominent features, resulting in poor detection effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and in particular, to an image processing method, device, storage medium, and computer device. Background Art

[0002] In the technical field of facial image detection, in order to achieve a better prevention and control effect, a facial image detection algorithm is adopted to ensure that the authenticated person in remote identity authentication is a real living person, rather than a photo, video, 3D model, etc., to avoid false authentication. The detection algorithm needs to make good use of the significant features from the image, such as the screen / paper edge, reflection, etc.; at the same time, it also needs to collect the detail information such as texture / material from the local part of the real face or attack sample.

[0003] However, the facial image detection model obtained through data training is very likely to overfit to the simple significant features in the image, resulting in a significant decrease in the recall ability of the algorithm when encountering an attacker who conceals the attack clues (such as paper edges, reflections, etc.) well; and when encountering normal glasses / mask reflections, or similar objects such as door frames and billboards appear behind a real person, it will mistake them for facial attacks, resulting in false interception; seriously affecting the image detection effect.

[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide an image processing method, device, storage medium, and computer device to at least solve the technical problem that the existing image detection method has overfitting in the detection of simple significant features, resulting in poor detection effects.

[0006] According to one aspect of the embodiments of the present invention, an image processing method is provided, including: performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; performing texture enhancement processing on a target image using a target encoder to obtain a clue map of the target image; superimposing the target image and the clue map to obtain a texture enhancement map; classifying the texture enhancement map using the target classifier to obtain an image detection result of the target image.

[0007] Optionally, before performing texture enhancement processing on the target image using the target encoder to obtain the clue map of the target image, the method further includes: constructing an initial image detection model based on the initial encoder and the target classifier; locking the parameters of the target classifier, and training the initial image detection model using a third image sample set to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

[0008] Optionally, the unsupervised pre-training using the first image sample set to obtain a pre-trained classifier includes: performing partial masking processing on the first image samples in the first image sample set, and reconstructing the masked parts to obtain processed image samples; performing the unsupervised pre-training using the sample set formed by the processed image samples to obtain the pre-trained classifier.

[0009] Optionally, the supervised training of the pre-trained classifier using the second image sample set to obtain the target classifier includes: when the pre-trained classifier is a transformer model, deleting the decoder of the transformer model to obtain the encoder of the transformer model; performing supervised training on the encoder of the transformer model using the second image sample set to obtain the target classifier.

[0010] According to another aspect of the embodiments of the present invention, there is provided an image processing method, including: performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; obtaining an image detection sample image set and the target classifier; constructing an initial image detection model based on the initial encoder and the target classifier; locking the parameters of the target classifier, performing texture enhancement processing on the sample images in the image detection sample set using the initial encoder to obtain clue maps of the sample images, and superimposing the sample images and the clue maps to obtain a texture-enhanced map, and inputting the texture-enhanced map into the target classifier, training the initial image detection model based on the image detection sample set to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

[0011] According to another aspect of the embodiments of the present invention, there is provided an image processing method, including: displaying an image input control on an interaction interface; receiving a target image in response to an operation on the image input control and displaying the target image on the interaction interface; receiving an image detection instruction; obtaining a pre-trained classifier by performing unsupervised pre-training using a first image sample set in response to the image detection instruction; obtaining a target classifier by performing supervised training on the pre-trained classifier using a second image sample set; performing texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image, and superimposing the target image and the clue map to obtain a texture-enhanced image; and displaying an image detection result of the target image on the interaction interface, where the image detection result is obtained by classifying the texture-enhanced image using the target classifier.

[0012] According to another aspect of the embodiments of the present invention, there is provided an image processing method, including: obtaining a target image displayed on a virtual reality (VR) device or an augmented reality (AR) device; obtaining a pre-trained classifier by performing unsupervised pre-training using a first image sample set; obtaining a target classifier by performing supervised training on the pre-trained classifier using a second image sample set; performing texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimposing the target image and the clue map to obtain a texture-enhanced image; classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image; in the case where the image detection result indicates that the target image includes a living target, labeling the living target in the target image to obtain a target image with the living target labeled; and driving the VR device or the AR device to display the target image with the living target labeled.

[0013] According to another aspect of the embodiments of the present invention, there is provided an image processing method, including: receiving a payment instruction; collecting a target video and intercepting a target image from the target video in response to the payment instruction; obtaining a pre-trained classifier by performing unsupervised pre-training using a first image sample set; obtaining a target classifier by performing supervised training on the pre-trained classifier using a second image sample set; performing texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimposing the target image and the clue map to obtain a texture-enhanced image; classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image; and verifying the payment instruction in the case where the image detection result indicates that the target video includes a living target.

[0014] According to another aspect of the embodiments of the present invention, an image processing apparatus is provided, including: a first training module for performing unsupervised pre-training using a first set of image samples to obtain a pre-trained classifier; a second training module for performing supervised training on the pre-trained classifier using a second set of image samples to obtain a target classifier; an enhancement module for performing texture enhancement processing on a target image using a target encoder to obtain a clue map of the target image; a superimposing processing module for superimposing the target image and the clue map to obtain a texture-enhanced image; and a classification module for classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image.

[0015] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the above-mentioned image processing methods.

[0016] According to another aspect of the embodiments of the present invention, a computer device is provided, including: a memory and a processor. The memory stores a computer program; the processor is configured to execute the computer program stored in the memory, and when the computer program runs, the processor executes any one of the above-mentioned image processing methods.

[0017] In the embodiments of the present invention, through an image processing system, unsupervised pre-training is performed using a first set of image samples to obtain a pre-trained classifier; supervised training is performed on the pre-trained classifier using a second set of image samples to obtain a target classifier; texture enhancement processing is performed on a target image using a target encoder to obtain a clue map of the target image; the target image and the clue map are superimposed to obtain a texture-enhanced image; and the texture-enhanced image is classified using the target classifier to obtain an image detection result of the target image, achieving the purpose of reducing the dependence of the face detection model on simple features, thereby realizing the technical effect of enhancing the underlying texture features and enabling the underlying texture-enhanced image to have the ability to distinguish between real and fake, and further solving the technical problem in the prior art that the image detection method has overfitting in detecting simple and prominent features, resulting in poor detection effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0019] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method is shown;

[0020] Figure 2 is a flowchart of the image processing method according to Embodiment 1 of the present invention;

[0021] Figure 3 is a schematic diagram of the pre-trained classifier generation process according to Embodiment 1 of the present invention;

[0022] Figure 4 is a schematic diagram of the target image detection model generation process according to Embodiment 1 of the present invention;

[0023] Figure 5 is a flowchart of the image processing method according to Embodiment 2 of the present invention;

[0024] Figure 6 is a flowchart of the image processing method according to Embodiment 3 of the present invention;

[0025] Figure 7 is a flowchart of the image processing method according to Embodiment 4 of the present invention;

[0026] Figure 8 is a flowchart of the image processing method according to Embodiment 5 of the present invention;

[0027] Figure 9 is a schematic diagram of the device structure for implementing the above image processing method according to Embodiment 6 of the present invention;

[0028] Figure 10 is a block diagram of the structure of a computer terminal according to an embodiment of the present invention. Detailed implementation manners

[0029] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0030] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0032] Transformer model: The Transformer model uses a Self-Attention structure to replace the RNN network structure commonly used in NLP tasks. Compared with the RNN network structure, its greatest advantage is that it can perform parallel computing. The Transformer model includes an encoding component and a decoding component.

[0033] Supervised training: Also known as supervised learning, it is a method in machine learning that can learn or establish a pattern (function / learning model) from training data and infer new instances based on this pattern. The training data consists of input objects (usually vectors) and expected outputs. The output of the function can be a continuous value (referred to as regression analysis) or predict a classification label (referred to as classification).

[0034] Unsupervised training: Also known as unsupervised learning, different from supervised learning, unsupervised learning does not have any training samples in advance and needs to directly model the data, such as clustering algorithms.

[0035] Embodiment 1

[0036] According to an embodiment of the present invention, an embodiment of an image processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0037] In the recommendation scenario of related technologies, the following solution can be adopted to obtain the image processing result, that is, the classification result. An autoencoder is used to convert the input image into a clue map for live attack, followed by a classifier. The above classifier is only used for auxiliary training, and the classification result is finally determined based on the mean value of the clue map. The overall idea is based on the idea of anomaly detection, that is, the expected result is that the response in the clue map of a real person is all 0, while there will be more responses in the clue map of an attack. The design background of this algorithm is relatively ideal, which is easy to overfit the simple and significant features in the image, and it is easy to misjudge when detecting anomalies based on the clue map in the predetermined scenario. The ability of the trained classifier is not well utilized in the final classification decision.

[0038] In the recommendation scenario of related technologies, the following solution can also be adopted to alleviate the overfitting problem of downstream tasks. The Transformer architecture network dedicated to the converter model MAE is designed to cover the original image with a large-area mask, and then the image is predicted. The whole process is carried out in a self-supervised manner. This method can, to a certain extent, guide the model to pre-extract high-level features irrelevant to downstream supervised tasks, and can effectively alleviate the overfitting problem of downstream tasks. However, due to the task characteristics of large-area masking and the defects in the architecture design of the Transformer model itself, the underlying texture features of the image are not fully mined, and the strong transferability of the underlying texture features of the image cannot be fully utilized.

[0039] To solve the problems existing in the above algorithm, in the embodiments of the present application, the classifier is taken as the dominant in the image decision-making process, and the clue map is only used as an auxiliary. The classifier pre-trained by the converter model MAE already has more general and generalized features than ordinary classifiers. Therefore, during the process of learning the autoencoder, the relevant parameters of the classifier will be locked, and idealized constraints such as setting the clue map of a real person to 0 and the existence of amplitude in the clue map of an attack will not be added to the generated underlying texture map. Nor will the clue map be used for classification decision-making. Instead, amplitude constraints will be unified, so that the clue map is only used for auxiliary classification, thus ensuring that the obtained classification result can be steadily improved. At the same time, the underlying texture map generated by the autoencoder will also respond to the classification supervision signal fed back by the classifier, having a certain image detection ability, which can further enhance the mining of the underlying texture features by the classification algorithm and obtain better generalization.

[0040] In the embodiments of the present application, compared with the technical solution of simply pre-training with the converter model MAE and then fine-tuning, by adding the underlying texture map before the input in the present application, the underlying features required by the classification network are enhanced. Even if the classification network remains unchanged, a stable improvement can be obtained in the results. It alleviates the problems that the Transformer network architecture lacks in mining texture features within patches and that the features extracted by MAE are too high-level, enabling the self-supervised training method to have discriminative ability in tasks that focus on detailed textures such as image detection.

[0041] The method embodiment provided by Embodiment 1 of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image processing method. As Figure 1 shown, the computer terminal 10 (or mobile device) may include one or more processors (shown as 102a, 102b,..., 102n in the figure, and the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0042] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is used for processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0043] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image processing method in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the vulnerability detection method of the above application program. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0044] The transmission device is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0045] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0046] Under the above operating environment, the present application provides an Figure 2 image processing method as shown. Figure 2 It is a flowchart of the image processing method according to Embodiment 1 of the present invention.

[0047] Step S202, perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier;

[0048] Step S204, perform supervised training on the above pre-trained classifier using the second image sample set to obtain a target classifier;

[0049] Step S206, perform texture enhancement processing on the target image using the target encoder to obtain a clue map of the above target image;

[0050] Step S208, superimpose the above target image and the above clue map to obtain a texture-enhanced image;

[0051] Step S210: Classify the above texture-enhanced image using the above target classifier to obtain the image detection result of the above target image.

[0052] In the embodiment of the present invention, the execution subject of the image processing method provided in the above steps S202 to S208 is an image processing system. The above system is used to obtain the above target image, perform texture enhancement processing on the above target image to obtain a clue map of the target image; then superimpose the target image and the obtained clue map to obtain a texture-enhanced image, and finally use a target classifier to perform classification processing on the generated texture-enhanced image, and determine the image detection result of the target image based on the classification result.

[0053] It should be noted that the above target image may be a real facial image obtained by using a detection device or other mobile devices through a camera or other shooting devices, or may also be a virtual facial image obtained by using a detection device or other mobile devices through a camera or other shooting devices, such as: a photo image of a certain user, a facial mask, etc.

[0054] It should also be noted that the above target encoder is obtained by training an initial encoder. After pre-training, the above initial encoder can be used to perform texture enhancement processing on the above target image; the clue map of the above target image is a bottom-layer texture map obtained by the above target encoder performing feature extraction processing on the target image, and the above system then performs superimposition processing on the above bottom-layer texture map (i.e., the clue map of the above target image) and the original image (i.e., the above target image) to obtain the above texture-enhanced image.

[0055] As an optional embodiment, the above image processing method includes a model training method and an image detection method. The model training method is divided into three stages: First, perform unsupervised pre-training on real human samples and attack samples together to obtain a pre-trained classifier; second, perform supervised training on the above pre-trained classifier to obtain the above target classifier; finally, lock the parameters of the trained classifier and use it as the classifier for image detection. In the image detection method, use the above target classifier to classify the above texture-enhanced image, and use the classification result as the image detection result of the above target image.

[0056] Through the embodiments of the present invention, a self-supervised training scheme based on MAE is adopted in the image detection task, enabling the target classifier to learn higher-level semantic features in advance, so that when the detection model performs supervised live classification training, it is not easily trapped in the overfitting problem of simple and prominent features, and better generalization performance is obtained; by introducing a target encoder trained from the initial encoder before the target classifier, a texture enhancement map that enhances the underlying texture features of the image is generated. After being superimposed with the target image, the subsequent classifier can better utilize the detailed texture in the image, making up for the problem that Transformer under-exploits the underlying layer caused by MAE pre-training, and taking advantage of the better transferability of detailed texture features to significantly enhance the generalization performance of the model.

[0057] In the embodiments of the present invention, before the classifier classifies the above-mentioned texture enhancement map, the detection model needs to be trained first. First, an unsupervised pre-training is performed using a first image sample set to obtain a pre-trained classifier; secondly, a supervised training is performed on the above-mentioned pre-trained classifier using a second image sample set to obtain the above-mentioned target classifier. Using the above-mentioned pre-trained classifier can fully mine the underlying texture information of the target image, and using the above-mentioned target classifier to fine-tune the classification task in a supervised manner can fully discover the top-level texture information of the target image. Superimposing the above-mentioned underlying texture information and the above-mentioned top-level texture information makes the obtained texture enhancement map have the ability to distinguish between real and fake.

[0058] It should be noted that the above-mentioned first image sample set includes pre-set real human samples and attack samples without labels. Among them, the above-mentioned real human samples can be real facial images obtained by using a detection device or other mobile devices through shooting devices such as cameras, for example: video data when a certain user performs a face recognition operation, etc.; the above-mentioned attack samples can be virtual facial images obtained by using a detection device or other mobile devices through shooting devices such as cameras, for example: photo images of a certain user, face masks, etc. The above-mentioned second image sample set includes pre-set real human samples and attack samples with labels, that is, the images in the samples are labeled as real human samples or attack samples through labels.

[0059] As an optional embodiment, the above-mentioned unsupervised pre-training using the first image sample set to obtain a pre-trained classifier includes: performing partial masking processing on the first image samples in the first image sample set, and performing image reconstruction on the masked parts to obtain processed image samples; using the sample set formed by the processed image samples to perform the above-mentioned unsupervised pre-training to obtain the above-mentioned pre-trained classifier.

[0060] Optionally, as Figure 3The schematic diagram of the pre-trained classifier generation process shown below performs unsupervised pre-training of the Transformer model MAE on all real human samples and attack samples. That is, after masking a certain proportion of the sample images, the masked parts are reconstructed, and the least squares loss function is used for constraint to obtain the above-mentioned pre-trained classifier.

[0061] As an optional embodiment, the above-mentioned supervised training of the above-mentioned pre-trained classifier using the second image sample set to obtain the above-mentioned target classifier includes: when the above-mentioned pre-trained classifier is a Transformer model, deleting the decoder of the above-mentioned Transformer model to obtain the encoder of the above-mentioned Transformer model; using the above-mentioned second image sample set to perform supervised training on the encoder of the above-mentioned Transformer model to obtain the above-mentioned target classifier.

[0062] Optionally, delete the decoder part of the Transformer model MAE, use the pre-trained encoder as the feature extraction classifier, perform fine-tuning on the classification task in a supervised manner, and use the cross-entropy loss function for constraint to obtain the above-mentioned target classifier.

[0063] In an optional embodiment, before the above-mentioned texture enhancement processing of the above-mentioned target image using the target encoder to obtain the clue map of the above-mentioned target image, it further includes: constructing an initial image detection model based on the initial encoder and the above-mentioned target classifier; locking the parameters of the above-mentioned target classifier, and using the third image sample set to train the initial image detection model to obtain a target image detection model, where the above-mentioned target image detection model includes the above-mentioned target encoder obtained by training the above-mentioned initial encoder.

[0064] In the embodiment of the present invention, an autoencoder, that is, the above-mentioned initial encoder, is trained before inputting the target image. An initial image detection model is constructed using the above-mentioned initial encoder and the above-mentioned target classifier. The parameters of the trained above-mentioned target classifier are locked, and the initial image detection model is trained using the third image sample set to obtain a target image detection model.

[0065] Optionally, as Figure 4The schematic diagram of the target image detection model generation process shown locks the parameters of the above-mentioned trained target classifier, and uses the above-mentioned autoencoder to generate an underlying feature enhancement map, that is, the above-mentioned texture enhancement map; the generated features are constrained using the least squares function to minimize their amplitudes and prevent affecting the classification performance of the subsequent classifier; then, the triple loss function TripletLoss is used to constrain and widen the distance between the features of the real person samples and the attack samples in the autoencoder. The original image and the reconstructed clue map are directly added together, and after passing through the classifier, cross-entropy is used for constraint. Only the parameters of the autoencoder are trained to obtain the target image detection model, so that the learned underlying texture map has the ability to distinguish between live bodies.

[0066] In the embodiment of the present invention, during the final forward inference of the above system, the original image passes through the autoencoder to generate an underlying texture enhancement map, which is added to the original image and then passes through the classifier (i.e., the encoder) to obtain the classification result.

[0067] It should be noted that the above-mentioned third image sample set includes pre-set real person samples and attack samples. Among them, the above-mentioned real person samples can be real facial images obtained by using a detection device or other mobile devices through a shooting device such as a camera, for example: video data when a certain user performs a facial recognition operation, etc.; the above-mentioned attack samples can be virtual facial images obtained by using a detection device or other mobile devices through a shooting device such as a camera, for example: a photo image of a certain user, a facial mask, etc.

[0068] In the embodiment of the present invention, through self-supervised pre-training, the dependence of the main detection model on simple features is reduced. Then, through an autoencoder, the texture features beneficial to image detection in the underlying features of the input image are enhanced, so as to utilize the better transfer and generalization performance of the underlying texture features, and solve the problems of the decline in the recall ability of attackers with well-hidden attack clues and misinterception caused by mistaking normal reflections, etc. for attacks.

[0069] Embodiment 2

[0070] Under the above operating environment, the present application provides an Figure 5 image processing method as shown. Figure 5 It is a flowchart of the image processing method according to Embodiment 2 of the present invention.

[0071] Step S302: Perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier;

[0072] Step S304: Perform supervised training on the above-mentioned pre-trained classifier using the second image sample set to obtain a target classifier;

[0073] Step S306: Obtain an image detection sample image set and the above-mentioned target classifier;

[0074] Step S308: Construct an initial image detection model based on the initial encoder and the above-mentioned target classifier;

[0075] Step S310: Lock the parameters of the above-mentioned target classifier, use the above-mentioned initial encoder to perform texture enhancement processing on the sample images in the above-mentioned image detection sample set to obtain clue maps of the above-mentioned sample images, and superimpose the above-mentioned sample images and the above-mentioned clue maps to obtain texture-enhanced images, and input the above-mentioned texture-enhanced images into the above-mentioned target classifier. Based on the above-mentioned image detection sample set, train the above-mentioned initial image detection model to obtain a target image detection model, where the above-mentioned target image detection model includes a target encoder obtained by training the above-mentioned initial encoder.

[0076] In the embodiment of the present invention, the execution subject of the image processing method provided in the above steps S302 to S306 is a model training system. Use the above model training system to obtain an image detection sample image set and a target classifier; construct an initial image detection model based on the initial encoder and the above-mentioned target classifier, lock the parameters of the above-mentioned target classifier, use the above-mentioned initial encoder to perform texture enhancement processing on the sample images in the above-mentioned image detection sample set to obtain clue maps of the above-mentioned sample images, and superimpose the above-mentioned sample images and the above-mentioned clue maps to obtain texture-enhanced images, and input the above-mentioned texture-enhanced images into the above-mentioned target classifier. Based on the above-mentioned image detection sample set, train the above-mentioned initial image detection model to obtain a target image detection model, and finally send the above-mentioned target image detection model to an image detection system for the above image detection system to use the above target image detection model to perform image detection on a target image.

[0077] In an optional embodiment, the obtaining of the above-mentioned target classifier includes: performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; performing supervised training on the above-mentioned pre-trained classifier using a second image sample set to obtain the above-mentioned target classifier.

[0078] It should be noted that the above-mentioned first image sample set includes pre-set real-person samples and attack samples without labels, where the above-mentioned real-person samples can be real facial images obtained by using a detection device or other mobile devices through a shooting device such as a camera, for example: video data when a certain user performs a face recognition operation, etc.; the above-mentioned attack samples can be virtual facial images obtained by using a detection device or other mobile devices through a shooting device such as a camera, for example: a photo image of a certain user, a face mask, etc. The above-mentioned second image sample set includes pre-set real-person samples and attack samples with labels, that is, the images in the samples are labeled as real-person samples or attack samples through labels.

[0079] As an alternative embodiment, the above-mentioned unsupervised pre-training using the first image sample set to obtain the pre-trained classifier includes: performing partial masking on the first image samples in the first image sample set, and reconstructing the masked parts to obtain processed image samples; using the sample set formed by the processed image samples to perform the above-mentioned unsupervised pre-training to obtain the above-mentioned pre-trained classifier.

[0080] Optionally, all real person samples and attack samples are jointly subjected to unsupervised pre-training of the Transformer model MAE, that is, after masking a certain proportion of the sample images, reconstructing the masked parts, and using the least squares loss function for constraint to obtain the above-mentioned pre-trained classifier.

[0081] As an alternative embodiment, the above-mentioned supervised training of the pre-trained classifier using the second image sample set to obtain the target classifier includes: in the case where the pre-trained classifier is a Transformer model, deleting the decoder of the Transformer model to obtain the encoder of the Transformer model; using the second image sample set to perform supervised training on the encoder of the Transformer model to obtain the above-mentioned target classifier.

[0082] Optionally, delete the decoder part of the Transformer model MAE, use the pre-trained encoder as the feature extraction classifier, perform fine-tuning on the classification task in a supervised manner, and use the cross-entropy loss function for constraint to obtain the above-mentioned target classifier.

[0083] It should be noted that the above-mentioned target image detection model includes the above-mentioned target encoder obtained by training the above-mentioned initial encoder.

[0084] Embodiment 3

[0085] Under the above-mentioned operating environment, the present application provides an Figure 6 image processing method as shown in Figure 6 is a flowchart of the image processing method according to Embodiment 3 of the present invention.

[0086] Step S402, display an image input control on the interaction interface;

[0087] Step S404, in response to an operation on the above-mentioned image input control, receive a target image and display the target image on the above-mentioned interaction interface;

[0088] Step S406, receive an image detection instruction;

[0089] Step S408: In response to the above image detection instruction, perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier; perform supervised training on the above pre-trained classifier using the second image sample set to obtain a target classifier; use the target encoder to perform texture enhancement processing on the above target image to obtain a clue map of the above target image, and superimpose the above target image and the above clue map to obtain a texture-enhanced image.

[0090] Step S410: Display the image detection result of the above target image on the above interaction interface, where the above image detection result is obtained by classifying the above texture-enhanced image using the above target classifier.

[0091] In an embodiment of the present invention, the execution subject of the image processing method provided in the above steps S402 to S410 is an image detection system. After receiving an image detection application instruction, the image detection system displays an image input control on an interaction interface and receives a target image to display the above target image on the above interaction interface; after receiving an image detection instruction, the image detection system uses a target encoder to perform texture enhancement processing on the above target image to obtain a clue map of the above target image, and superimpose the above target image and the above clue map to obtain a texture-enhanced image; display the image detection result of the above target image on the above interaction interface.

[0092] It should be noted that the above image detection result is obtained by classifying the above texture-enhanced image using a target classifier; the above image detection system includes but is not limited to a detection device or other mobile devices.

[0093] Embodiment 4

[0094] Under the above operating environment, the present application provides an Figure 7 image processing method as shown. Figure 7 is a flowchart of an image processing method according to Embodiment 4 of the present invention.

[0095] Step S502: Obtain a target image displayed on a virtual reality (VR) device or an augmented reality (AR) device.

[0096] Step S504: Perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier.

[0097] Step S506: Perform supervised training on the above pre-trained classifier using the second image sample set to obtain a target classifier.

[0098] Step S508: Use the target encoder to perform texture enhancement processing on the above target image to obtain a clue map of the above target image.

[0099] Step S510: Superimpose the above-mentioned target image and the above-mentioned clue image to obtain a texture-enhanced image;

[0100] Step S512: Classify the above-mentioned texture-enhanced image using a target classifier to obtain an image detection result of the above-mentioned target image;

[0101] Step S514: In the case where the above-mentioned image detection result is that the above-mentioned target image includes a living target, label the living target in the above-mentioned target image to obtain a target image with the above-mentioned living target labeled;

[0102] Step S516: Drive the above-mentioned VR device or the above-mentioned AR device to display the target image with the above-mentioned living target labeled.

[0103] In the embodiment of the present invention, the execution subject of the image processing method provided in the above steps S502 to S512 is an image detection system running on a virtual reality VR device or an augmented reality AR device. After receiving an image detection application instruction, the image detection system acquires a target image displayed on the virtual reality VR device or the augmented reality AR device; performs texture enhancement processing on the above-mentioned target image using a target encoder to obtain a clue image of the above-mentioned target image; superimposes the above-mentioned target image and the above-mentioned clue image to obtain a texture-enhanced image; classifies the above-mentioned texture-enhanced image using a target classifier to obtain an image detection result of the above-mentioned target image; in the case where the above-mentioned image detection result is that the above-mentioned target image includes a living target, label the living target in the above-mentioned target image to obtain a target image with the above-mentioned living target labeled; finally, drive the above-mentioned VR device or the above-mentioned AR device to display the target image with the above-mentioned living target labeled.

[0104] Embodiment 5

[0105] In the above operating environment, the present application provides an Figure 8 image processing method as shown. Figure 8 It is a flowchart of the image processing method according to Embodiment 5 of the present invention.

[0106] Step S602: Receive a payment instruction;

[0107] Step S604: In response to the above-mentioned payment instruction, collect a target video and intercept a target image from the above-mentioned target video;

[0108] Step S606: Perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier;

[0109] Step S608: Perform supervised training on the above-mentioned pre-trained classifier using a second image sample set to obtain a target classifier;

[0110] Step S610: Use the target encoder to perform texture enhancement processing on the above-mentioned target image to obtain a clue map of the above-mentioned target image;

[0111] Step S612: Superimpose the above-mentioned target image and the above-mentioned clue map to obtain a texture-enhanced image;

[0112] Step S614: Use the target classifier to classify the above-mentioned texture-enhanced image to obtain an image detection result of the above-mentioned target image;

[0113] Step S616: In the case where the above-mentioned image detection result is that the above-mentioned target video includes a live target, verify the above-mentioned payment instruction.

[0114] In the embodiment of the present invention, the execution subject of the image processing method provided in the above steps S602 to S612 is an image detection system applied to the payment scenario of face recognition. After the payment device triggers a payment instruction, the above image detection system receives the above payment instruction, and after responding to the instruction, acquires a target video and intercepts a target image from the above target video; uses the target encoder to perform texture enhancement processing on the above target image to obtain a clue map of the above target image; superimposes the above target image and the above clue map to obtain a texture-enhanced image; uses the target classifier to classify the above texture-enhanced image to obtain an image detection result of the above target image; in the case where the above image detection result is that the above target video includes a live target, verify the above payment instruction.

[0115] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0117] Embodiment 6

[0118] According to an embodiment of the present invention, there is also provided an apparatus for implementing the above image processing method. As Figure 9 shown, the apparatus includes: a first training module 90, a second training module 92, an enhancement module 94, an overlay processing module 96, and a classification module 98, where:

[0119] The first training module 90 is configured to perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier;

[0120] The second training module 92 is configured to perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier;

[0121] The enhancement module 94 is configured to perform texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image;

[0122] The overlay processing module 96 is configured to overlay the target image and the clue map to obtain a texture-enhanced image;

[0123] The classification module 98 is configured to classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image.

[0124] It should be noted here that the above first training module 90, second training module 92, enhancement module 94, overlay processing module 96, and classification module 98 correspond to steps S202 to S210 in Embodiment 1. The functions realized by the four modules and the corresponding steps are the same in terms of examples and application scenarios, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the apparatus, can run in the computer terminal 10 provided in Embodiment 1.

[0125] Embodiment 7

[0126] An embodiment of the present invention may provide a computer terminal, and the computer terminal may be any one of the computer terminal devices in a computer terminal group. Optionally, in this embodiment, the above computer terminal may also be replaced with a terminal device such as a mobile terminal.

[0127] Optionally, in this embodiment, the above computer terminal may be located in at least one of multiple network devices in a computer network.

[0128] In this embodiment, the computer terminal can execute the program code for the following steps in the image processing method of the application program: perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on a target image using the target encoder to obtain a clue map of the target image; superimpose the target image and the clue map to obtain a texture-enhanced image; classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image.

[0129] In this embodiment, the computer terminal can execute the program code for the following steps in the image processing method of the application program: perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; obtain an image detection sample image set and the target classifier; construct an initial image detection model based on an initial encoder and the target classifier; lock the parameters of the target classifier, perform texture enhancement processing on the sample images in the image detection sample set using the initial encoder to obtain clue maps of the sample images, superimpose the sample images and the clue maps to obtain texture-enhanced images, and input the texture-enhanced images into the target classifier, and based on the image detection sample set, train the initial image detection model to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

[0130] In this embodiment, the computer terminal can execute the program code for the following steps in the image processing method of the application program: display an image input control on an interaction interface; in response to an operation on the image input control, receive a target image and display the target image on the interaction interface; receive an image detection instruction; in response to the image detection instruction, perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using the target encoder to obtain a clue map of the target image, and superimpose the target image and the clue map to obtain a texture-enhanced image; display an image detection result of the target image on the interaction interface, where the image detection result is obtained by classifying the texture-enhanced image using the target classifier.

[0131] In this embodiment, the computer terminal can execute the program code for the following steps in the image processing method of the application program: obtaining a target image displayed on a virtual reality (VR) device or an augmented reality (AR) device; performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; performing texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimposing the target image and the clue map to obtain a texture-enhanced image; classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image; in the case where the image detection result indicates that the target image includes a living target, annotating the living target in the target image to obtain a target image with the living target annotated; driving the VR device or the AR device to display the target image with the living target annotated.

[0132] In this embodiment, the computer terminal can execute the program code for the following steps in the image processing method of the application program: receiving a payment instruction; in response to the payment instruction, collecting a target video and intercepting a target image from the target video; performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; performing texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimposing the target image and the clue map to obtain a texture-enhanced image; classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image; in the case where the image detection result indicates that the target video includes a living target, verifying the payment instruction.

[0133] Optionally, Figure 10 is a structural block diagram of a computer terminal according to an embodiment of the present invention. As Figure 10 shown, the computer terminal may include: one or more (only one is shown in the figure) processors, a memory, etc.

[0134] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image processing method and device in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned image processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories can be connected to the computer terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0135] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using the target encoder to obtain a clue map of the target image; superimpose the target image and the clue map to obtain a texture-enhanced image; classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image.

[0136] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; obtain an image detection sample image set and the target classifier; construct an initial image detection model based on the initial encoder and the target classifier; lock the parameters of the target classifier, perform texture enhancement processing on the sample images in the image detection sample set using the initial encoder to obtain clue maps of the sample images, and superimpose the sample images and the clue maps to obtain texture-enhanced images, and input the texture-enhanced images into the target classifier, and based on the image detection sample set, train the initial image detection model to obtain a target image detection model, wherein the target image detection model includes the target encoder obtained by training the initial encoder.

[0137] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: display an image input control on an interaction interface; in response to an operation on the image input control, receive a target image and display the target image on the interaction interface; receive an image detection instruction; in response to the image detection instruction, perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image, and superimpose the target image and the clue map to obtain a texture-enhanced image; display an image detection result of the target image on the interaction interface, where the image detection result is obtained by classifying the texture-enhanced image using the target classifier.

[0138] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: obtain a target image displayed on a virtual reality (VR) device or an augmented reality (AR) device; perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimpose the target image and the clue map to obtain a texture-enhanced image; classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image; in the case where the image detection result is that the target image includes a live target, label the live target in the target image to obtain a target image with the live target labeled; drive the VR device or the AR device to display the target image with the live target labeled.

[0139] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: receive a payment instruction; in response to the payment instruction, collect a target video and intercept a target image from the target video; perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimpose the target image and the clue map to obtain a texture-enhanced image; classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image; in the case where the image detection result is that the target video includes a live target, verify the payment instruction.

[0140] Optionally, the above-mentioned processor may also execute the program code of the following steps: perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using the second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using the target encoder to obtain a clue map of the target image; superimpose the target image and the clue map to obtain a texture-enhanced image; classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image.

[0141] Optionally, the above-mentioned processor may also execute the program code of the following steps: perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using the second image sample set to obtain a target classifier; obtain an image detection sample image set and the target classifier; construct an initial image detection model based on the initial encoder and the target classifier; lock the parameters of the target classifier, perform texture enhancement processing on the sample images in the image detection sample set using the initial encoder to obtain clue maps of the sample images, superimpose the sample images and the clue maps to obtain texture-enhanced images, and input the texture-enhanced images into the target classifier, and based on the image detection sample set, train the initial image detection model to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

[0142] Optionally, the above-mentioned processor may also execute the program code of the following steps: display an image input control on the interaction interface; in response to an operation on the image input control, receive a target image and display the target image on the interaction interface; receive an image detection instruction; in response to the image detection instruction, perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using the second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using the target encoder to obtain a clue map of the target image, and superimpose the target image and the clue map to obtain a texture-enhanced image; display an image detection result of the target image on the interaction interface, where the image detection result is obtained by classifying the texture-enhanced image using the target classifier.

[0143] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtain a target image displayed on a virtual reality (VR) device or an augmented reality (AR) device; perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimpose the target image and the clue map to obtain a texture-enhanced image; classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image; in the case where the image detection result indicates that the target image includes a live target, label the live target in the target image to obtain a target image with the live target labeled; drive the VR device or the AR device to display the target image with the live target labeled.

[0144] Optionally, the above-mentioned processor may also execute the program code of the following steps: receive a payment instruction; in response to the payment instruction, collect a target video and intercept a target image from the target video; perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; perform texture enhancement processing on the target image using a target encoder to obtain a clue map of the target image; superimpose the target image and the clue map to obtain a texture-enhanced image; classify the texture-enhanced image using the target classifier to obtain an image detection result of the target image; in the case where the image detection result indicates that the target video includes a live target, verify the payment instruction.

[0145] By using the embodiment of the present invention, a scheme for image processing is provided. By performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier, performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier, performing texture enhancement processing on a target image using a target encoder to obtain a clue map of the target image, superimposing the target image and the clue map to obtain a texture-enhanced image, and classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image, the purpose of reducing the dependence of the face detection model on simple features is achieved, thereby realizing the technical effect of enhancing the underlying texture features and enabling the underlying texture-enhanced image to have the ability to distinguish live bodies, and further solving the technical problem that the image detection method in the prior art has an overfitting situation in detecting simple and significant features, resulting in poor detection effects.

[0146] Those of ordinary skill in the art can understand that Figure 10The structure shown is only illustrative. The computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 10 It does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 10 in the figure, or have a different configuration from that shown Figure 10 in the figure.

[0147] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program. The program can be stored in a computer-readable storage medium. The computer-readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.

[0148] The embodiment of the present invention also provides a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium can be used to store the program code executed by the image processing method provided in the above Embodiment 1.

[0149] Optionally, in this embodiment, the above computer-readable storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0150] Optionally, in this embodiment, the computer-readable storage medium is set to store program code for performing the following steps: performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; performing texture enhancement processing on a target image using the target encoder to obtain a clue map of the target image; superimposing the target image and the clue map to obtain a texture-enhanced image; classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image.

[0151] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining an image detection sample image set and a target classifier; constructing an initial image detection model based on an initial encoder and the above-mentioned target classifier; locking the parameters of the above-mentioned target classifier, using the above-mentioned initial encoder to perform texture enhancement processing on the sample images in the above-mentioned image detection sample set to obtain clue maps of the above-mentioned sample images, and superimposing the above-mentioned sample images and the above-mentioned clue maps to obtain a texture-enhanced image, and inputting the above-mentioned texture-enhanced image into the above-mentioned target classifier, training the above-mentioned initial image detection model based on the above-mentioned image detection sample set to obtain a target image detection model, where the above-mentioned target image detection model includes the above-mentioned target encoder obtained by training the above-mentioned initial encoder.

[0152] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: displaying an image input control on an interaction interface; in response to an operation on the above-mentioned image input control, receiving a target image and displaying the above-mentioned target image on the above-mentioned interaction interface; receiving an image detection instruction; in response to the above-mentioned image detection instruction, using a target encoder to perform texture enhancement processing on the above-mentioned target image to obtain a clue map of the above-mentioned target image, and superimposing the above-mentioned target image and the above-mentioned clue map to obtain a texture-enhanced image; displaying an image detection result of the above-mentioned target image on the above-mentioned interaction interface, where the above-mentioned image detection result is obtained by classifying the above-mentioned texture-enhanced image using a target classifier.

[0153] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining a target image displayed on a virtual reality (VR) device or an augmented reality (AR) device; using a target encoder to perform texture enhancement processing on the above-mentioned target image to obtain a clue map of the above-mentioned target image; superimposing the above-mentioned target image and the above-mentioned clue map to obtain a texture-enhanced image; classifying the above-mentioned texture-enhanced image using a target classifier to obtain an image detection result of the above-mentioned target image; in the case where the above-mentioned image detection result indicates that the above-mentioned target image includes a live target, labeling the live target in the above-mentioned target image to obtain a target image with the above-mentioned live target labeled; driving the above-mentioned VR device or the above-mentioned AR device to display the target image with the above-mentioned live target labeled.

[0154] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: receiving a payment instruction; in response to the above payment instruction, collecting a target video and intercepting a target image from the above target video; using a target encoder to perform texture enhancement processing on the above target image to obtain a clue map of the above target image; superimposing the above target image and the above clue map to obtain a texture-enhanced map; using a target classifier to classify the above texture-enhanced map to obtain an image detection result of the above target image; in the case where the above image detection result is that the above target video includes a live target, verifying the above payment instruction.

[0155] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0156] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0157] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.

[0158] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0159] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0160] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned computer-readable storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0161] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An image processing method, characterized in that, Including: Performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; Performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; Performing texture enhancement processing on a target image using a target encoder to obtain a clue map of the target image; Overlaying the target image and the clue map to obtain a texture-enhanced image; Classifying the texture-enhanced image using the target classifier to obtain an image detection result of the target image, wherein the pre-trained classifier obtained by unsupervised training in the target classifier is used to mine the underlying texture information of the target image, and the one obtained by supervised training in the target classifier is used to discover the top-level texture information of the target image; Wherein, before performing texture enhancement processing on the target image using the target encoder to obtain a clue map of the target image, it further includes: constructing an initial image detection model based on an initial encoder and the target classifier; locking the parameters of the target classifier, and training the initial image detection model using a third image sample set to obtain a target image detection model, wherein the target image detection model includes the target encoder obtained by training the initial encoder.

2. The method according to claim 1, wherein The performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier includes: Performing partial masking processing on a first image sample in the first image sample set, and performing image reconstruction on the masked part to obtain a processed image sample; Performing the unsupervised pre-training using the sample set formed by the processed image samples to obtain the pre-trained classifier.

3. The method according to claim 1, wherein The performing supervised training on the pre-trained classifier using a second image sample set to obtain the target classifier includes: In the case where the pre-trained classifier is a transformer model, deleting the decoder of the transformer model to obtain the encoder of the transformer model; Performing supervised training on the encoder of the transformer model using the second image sample set to obtain the target classifier.

4. An image processing method, characterized in that, Including: Performing unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; Performing supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; Obtaining an image detection sample image set and the target classifier; Constructing an initial image detection model based on an initial encoder and the target classifier; Lock the parameters of the target classifier, use the initial encoder to perform texture enhancement processing on the sample images in the image detection sample set to obtain clue maps of the sample images, and superimpose the sample images and the clue maps to obtain texture-enhanced images, and input the texture-enhanced images into the target classifier. Based on the image detection sample set, train the initial image detection model to obtain a target image detection model, where the target image detection model includes a target encoder obtained by training the initial encoder. The pre-trained classifier obtained by unsupervised training in the target classifier is used to mine the underlying texture information of the target image, and the classifier obtained by supervised training in the target classifier is used to discover the top-level texture information of the target image.

5. An image processing method, characterized in that, Comprising: Display an image input control on the interaction interface; In response to an operation on the image input control, receive a target image and display the target image on the interaction interface; Receive an image detection instruction; In response to the image detection instruction, perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; Perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; Use the target encoder to perform texture enhancement processing on the target image to obtain a clue map of the target image, and superimpose the target image and the clue map to obtain a texture-enhanced image; Display the image detection result of the target image on the interaction interface, where the image detection result is obtained by classifying the texture-enhanced image using the target classifier. The pre-trained classifier obtained by unsupervised training in the target classifier is used to mine the underlying texture information of the target image, and the classifier obtained by supervised training in the target classifier is used to discover the top-level texture information of the target image; Wherein, before using the target encoder to perform texture enhancement processing on the target image to obtain a clue map of the target image, it further includes: constructing an initial image detection model based on the initial encoder and the target classifier; locking the parameters of the target classifier, and using a third image sample set to train the initial image detection model to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

6. An image processing method, characterized in that, Comprising: Obtain a target image displayed on a virtual reality (VR) device or an augmented reality (AR) device; Perform unsupervised pre-training using a first image sample set to obtain a pre-trained classifier; Perform supervised training on the pre-trained classifier using a second image sample set to obtain a target classifier; Use the target encoder to perform texture enhancement processing on the target image to obtain a clue map of the target image; Superimpose the target image and the clue map to obtain a texture-enhanced image; Classify the texture-enhanced image using the target classifier to obtain the image detection result of the target image. Among them, the pre-trained classifier obtained by unsupervised training in the target classifier is used to mine the underlying texture information of the target image, and the one obtained by supervised training in the target classifier is used to discover the top-level texture information of the target image; In the case where the image detection result is that the target image includes a live target, label the live target in the target image to obtain a target image with the live target labeled; Drive the VR device or the AR device to display the target image with the live target labeled; Among them, before using the target encoder to perform texture enhancement processing on the target image to obtain the clue map of the target image, it further includes: constructing an initial image detection model based on the initial encoder and the target classifier; locking the parameters of the target classifier, and using the third image sample set to train the initial image detection model to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

7. An image processing method, characterized in that, Include: Receive a payment instruction; In response to the payment instruction, collect a target video and intercept a target image from the target video; Perform unsupervised pre-training using the first image sample set to obtain a pre-trained classifier; Perform supervised training on the pre-trained classifier using the second image sample set to obtain a target classifier; Perform texture enhancement processing on the target image using the target encoder to obtain the clue map of the target image; Overlay the target image and the clue map to obtain a texture-enhanced image; Classify the texture-enhanced image using the target classifier to obtain the image detection result of the target image. Among them, the pre-trained classifier obtained by unsupervised training in the target classifier is used to mine the underlying texture information of the target image, and the one obtained by supervised training in the target classifier is used to discover the top-level texture information of the target image; In the case where the image detection result is that the target video includes a live target, verify the payment instruction; Among them, before using the target encoder to perform texture enhancement processing on the target image to obtain the clue map of the target image, it further includes: constructing an initial image detection model based on the initial encoder and the target classifier; locking the parameters of the target classifier, and using the third image sample set to train the initial image detection model to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

8. An image processing apparatus, characterized in that, Include: A first training module for performing unsupervised pre-training using the first image sample set to obtain a pre-trained classifier; A second training module for performing supervised training on the pre-trained classifier using the second image sample set to obtain a target classifier; An enhancement module for performing texture enhancement processing on a target image using a target encoder to obtain the clue map of the target image; An overlay processing module for overlaying the target image and the clue map to obtain a texture-enhanced map; A classification module for classifying the texture-enhanced map using a target classifier to obtain an image detection result of the target image. Among them, the pre-trained classifier obtained by unsupervised training in the target classifier is used to mine the underlying texture information of the target image, and the one obtained by supervised training in the target classifier is used to discover the top-level texture information of the target image; Among them, the enhancement module is further configured to construct an initial image detection model based on the initial encoder and the target classifier; lock the parameters of the target classifier, and use the third image sample set to train the initial image detection model to obtain a target image detection model, where the target image detection model includes the target encoder obtained by training the initial encoder.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the image processing method according to any one of claims 1 to 7.

10. A computer device, characterized in that, Including: A memory and a processor, The memory stores a computer program; The processor is configured to execute the computer program stored in the memory, and when the computer program runs, the processor executes the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Transfer learning method and device, computer device and storage medium

    CN108805160A

  • Image enhancement method and device and storage medium

    CN112348747A

  • Multi-person behavior recognition method based on Transform network

    CN113033657A

  • Object detection method and device

    CN114140427A

  • Image recognition method and device, computer equipment, storage medium and product

    CN114359564A