An image recognition method, device, storage medium and electronic device

CN115620111BActive Publication Date: 2026-09-18ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210991167.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2026-09-18
Estimated Expiration
2042-08-18

AI Technical Summary

Benefits of technology

在本说明书一个或多个实施例中,电子设备可以基于图像识别任务构建包括第一主网络和第二元网络的初始图像识别模型,采用图像样本数据对第一主网络进行主网络训练以确定至少一个监督信号识别结果,然后基于各监督信号识别结果对第二元网络进行元网络训练以确定损失调整参数,再基于损失调整参数对第一主网络进行模型调整,直至得到针对初始图像识别模型的目标图像识别模型。通过若干监督信号从多个维度进行图像识别训练并结合元网络训练可实现基于损失调整参数的自适应监督,在模型训练过程中可准确高效的动态对模型网络结构以及参数分配进行调整,达到较好的资源利用率,在保证模型性能的前提下通过模型自适应监督调整可降低对模型资源的消耗,以及可大幅确保模型上线后的模型鲁棒性和模型适应能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620111B_ABST
    Figure CN115620111B_ABST
Patent Text Reader

Abstract

The specification discloses an image recognition method, device, storage medium and electronic equipment, wherein the method comprises: constructing an initial image recognition model based on an image recognition task, the image recognition model comprising a first main network and a second meta network; performing main network training on the first main network using image sample data to determine a supervised signal recognition result; then performing meta network training on the second meta network based on the supervised signal recognition result to determine a loss adjustment parameter, so as to perform model adjustment on the first main network and obtain a target image recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image recognition method, apparatus, storage medium and electronic device. Background Technology

[0002] With the widespread use of electronic devices, visual image data such as images and videos are increasing daily. Visual content perception and understanding has become a cutting-edge research direction in scientific research fields such as visual computing, computer vision, and computational photography, as well as their interdisciplinary areas. Among them, image recognition, such as liveness detection, object recognition, and scene recognition, is a recent research hotspot in the field of visual content perception and understanding. Summary of the Invention

[0003] This specification provides an image recognition method, apparatus, storage medium, and electronic device, the technical solution of which is as follows: Firstly, this specification provides an image recognition method, the method comprising: An initial image recognition model is constructed based on the image recognition task. The image recognition model includes a first master network and a second meta-network. The first main network is trained using image sample data to determine at least one supervisory signal recognition result. Based on the identification results of each of the supervision signals, the second meta-network is trained to determine the loss adjustment parameters. The first master network is adjusted based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model.

[0004] Secondly, this specification provides an image recognition device, the device comprising: The model building module is used to build an initial image recognition model based on the image recognition task. The image recognition model includes a first master network and a second meta-network. The model training module is used to train the first main network using image sample data and determine at least one supervision signal recognition result. The model training module is used to train the second meta-network based on the recognition results of each supervision signal and to determine the loss adjustment parameters. The model training module is used to adjust the first main network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model.

[0005] Thirdly, this specification provides a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.

[0006] Fourthly, this specification provides an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.

[0007] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following: In one or more embodiments of this specification, an electronic device can construct an initial image recognition model, including a first main network and a second meta-network, based on an image recognition task. The first main network is trained using image sample data to determine at least one supervisory signal recognition result. Then, based on the recognition results of each supervisory signal, the second meta-network is trained to determine loss adjustment parameters. Finally, the first main network is adjusted based on the loss adjustment parameters until a target image recognition model for the initial image recognition model is obtained. By training image recognition from multiple dimensions using several supervisory signals and combining this with meta-network training, adaptive supervision based on loss adjustment parameters can be achieved. During model training, the model network structure and parameter allocation can be dynamically adjusted accurately and efficiently, achieving better resource utilization. While ensuring model performance, adaptive supervision adjustment can reduce the consumption of model resources and significantly ensure the robustness and adaptability of the model after deployment. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a scene diagram of an image recognition system provided in this manual; Figure 2 This is a flowchart illustrating an image recognition method provided in this specification; Figure 3 This is a flowchart illustrating another image recognition method provided in this manual; Figure 4 This is a flowchart illustrating another image recognition method provided in this manual; Figure 5 This is a schematic diagram of the structure of an image recognition device provided in this specification; Figure 6 This is a schematic diagram of the structure of a model training module provided in this manual; Figure 7 This is a schematic diagram of the structure of a network training unit provided in this specification; Figure 8 This is a schematic diagram of another image recognition device provided in this specification; Figure 9 This is a schematic diagram of the structure of an electronic device provided in this specification; Figure 10 This is a schematic diagram of the operating system and user space provided in this manual; Figure 11 yes Figure 10 Architecture diagram of the Android operating system in China; Figure 12 yes Figure 10 Architecture diagram of the iOS operating system. Detailed Implementation

[0010] The technical solutions in this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0012] In related technologies, image recognition often relies on machine learning methods to build an initial image recognition model for training. After the model converges, a trained image recognition model is obtained and applied to the corresponding image recognition scenario. However, current image recognition methods for training image recognition models suffer from poor recognition performance, especially in harsh application scenarios where the models exhibit low robustness. For example, in the application of image recognition models in liveness detection scenarios, the trained models require highly cooperative user actions such as shaking the head or blinking to accurately identify the subject. However, the application environment for image recognition models is often not ideal (as mentioned above), leading to a significant decrease in recognition accuracy. Therefore, further improvements to image recognition methods for training image recognition models are needed.

[0013] The present application will now be described in detail with reference to specific embodiments.

[0014] Please see Figure 1 This is a scene diagram of an image recognition system provided in this specification. Figure 1 As shown, the image recognition system may include at least a client cluster and a service platform 100.

[0015] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.

[0016] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.

[0017] The service platform 100 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of other servers.

[0018] In one or more embodiments of this specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and complete data interaction during the image recognition process based on the communication connection, such as online transaction data interaction. For example, the service platform 100 can deploy the target image recognition model obtained by the image recognition method of this specification to several clients online, and the clients can perform image recognition based on the target image recognition model. Or, the service platform 100 can obtain the target detection image to be detected in the corresponding transaction scenario (such as a liveness detection transaction scenario) from the client, then input the target detection image into the target image recognition model, output at least one target supervision recognition result for the target detection image, and determine the image detection type corresponding to the target detection image based on each target supervision recognition result, and can send the image detection type to the client, etc.

[0019] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0020] The image recognition system embodiments provided in this specification and the image recognition methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the image recognition method involved in one or more embodiments of this specification can be the aforementioned service platform 100; the execution entity corresponding to the image recognition method involved in one or more embodiments of this specification can also be a client, depending on the actual application environment. The implementation process of the image recognition system embodiments can be detailed in the following method embodiments, and will not be repeated here.

[0021] based on Figure 1 The following is a detailed description of the image recognition method provided by one or more embodiments of this specification, as illustrated in the scene diagram.

[0022] Please see Figure 2 This document provides a flowchart illustrating an image recognition method according to one or more embodiments. This method can be implemented using a computer program and can run on an image recognition device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The image recognition device can be a service platform.

[0023] Specifically, the image recognition method includes: S102: Construct an initial image recognition model based on the image recognition task, wherein the image recognition model includes a first main network and a second meta-network; In real-world scenarios, machine learning-based image recognition models are often used to further process and identify image data generated during these scenarios. Image recognition models can be applied to different image recognition scenarios depending on the specific task. In other words, image recognition models can be neural networks suitable for various machine vision image recognition tasks, such as liveness detection, object recognition in autonomous driving scenarios, and interaction recognition in human-computer interaction scenarios. Furthermore, initial image recognition models can be pre-built based on different image recognition tasks.

[0024] In one or more embodiments of this specification, the initial image recognition model includes at least a first main network and a second sub-network. During the training process of the initial image recognition model, the first main network is mainly used for image recognition, and the second sub-network is used to adjust the network structure and computational resource allocation of the first main network during model training to achieve better resource utilization for model computation. Understandably, the meta-learning approach for training meta-networks based on second-order networks can adaptively adjust the weights of multiple supervisory signals and the parameter allocation of the network during model training, thereby achieving high performance and reducing resource consumption in less time.

[0025] Furthermore, the initial image recognition model can be created by fitting one or more machine learning models such as Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Networks (RNN), embedding models, Gradient Boosting Decision Tree (GBDT) models, and Logistic Regression (LR) models.

[0026] Indicatively, the electronic device pre-builds an initial image recognition model for the image recognition task. The initial image recognition model is set based on the corresponding image recognition task and includes a first main network and a second meta-network. S104: Use image sample data to train the first main network and determine at least one supervisory signal recognition result; Image sample data can be publicly available image data obtained from relevant databases, such as one or more of CIFAR-10, CIFAR-100, Tiny ImageNet, etc., or it can be user-defined image sample data collected for specific scenarios in actual image recognition tasks, such as image classification datasets created by labeling image data collected from the Internet. Image sample data can be one or more image sample training sets, each of which includes several sample images.

[0027] To illustrate, taking image recognition as an example of a liveness detection task, the data acquisition process for image sample data can be as follows: Data on live users can be collected using an image acquisition device, for example, 20 images of each of 500 different users under different lighting conditions and facial angles. Users should cover different factors, such as age, weight, etc. Simultaneously, photos of different attack materials can be collected, such as mobile phone screens, printed paper, etc., with 20 images of each attack material collected under different lighting conditions and angles. Furthermore, after acquiring image sample data, data filtering and preprocessing can be performed on the image sample data for the liveness detection task. For example, face detection and face quality judgment can be performed on the images in the acquired image sample data, and images that do not detect faces or have poor quality can be discarded.

[0028] In one or more embodiments of this specification, each sample image may carry a label. This label is used for training the initial image recognition model and corresponds to the output of the first main network. The label may be a label representing the recognition result of the supervision signal for the sample image. In one or more embodiments of this specification, each sample image may also be unlabeled, and the first main network may be trained using unlabeled sample images. The supervision signal loss function set by the first main network to calculate the supervision signal loss may be a loss function that is not based on the label in related technologies.

[0029] In one or more embodiments of this specification, the training image sample data may correspond to only one image recognition task.

[0030] Optionally, image sample data is used as the input master training data of the first master network, and the output recognition data of the first master network is used as the input meta training data of the second meta network.

[0031] In a schematic manner, image sample data is used to train the first master network for one or more rounds, and the corresponding supervision signal recognition results are output for one or more rounds.

[0032] Understandably, during or after the initial image recognition model is built, at least one image recognition supervision signal for a recognition dimension can be configured on the first master network based on the image recognition task. The first master network is then instructed to perform image recognition in the recognition dimension based on the image recognition supervision signal, and the output of the first master network is the recognition result corresponding to the image recognition supervision signal of the corresponding recognition dimension.

[0033] To illustrate, taking the liveness detection task as an example, it is usually necessary to combine the recognition results of multiple detection and recognition dimensions to comprehensively determine whether a target image is a live type or an attack type. The first main network included in the initial image recognition image performs recognition of the recognition dimensions corresponding to the image recognition supervision signals on the image sample data to obtain the results of the image recognition supervision signals, that is, the first main network outputs several supervision signal recognition results.

[0034] Taking a liveness detection task as an example, the liveness recognition supervision signal includes at least one of depth estimation supervision signal, image material classification supervision signal, reflectance spectrum prediction supervision signal, and liveness classification supervision signal. It can be understood that image sample data is input into the first main network, which performs image recognition on the image sample data in recognition dimensions such as the depth estimation dimension corresponding to the depth estimation supervision signal and the image material classification dimension corresponding to the image material classification supervision signal, and outputs the corresponding supervision signal recognition results, such as the recognition results corresponding to the depth estimation supervision signal, the recognition results corresponding to the image material classification supervision signal, and the reflectance spectrum prediction results corresponding to the reflectance spectrum prediction supervision signal, etc.

[0035] S106: Based on the identification results of each of the supervision signals, perform meta-network training on the second meta-network to determine the loss adjustment parameters; The loss adjustment parameter is used to adjust the loss function of the first main network in the next round of training of the first main network, so as to achieve a balance between model training effect, model recognition performance and resource consumption based on the updated loss function.

[0036] Understandably, the output recognition data of the first master network serves as the input meta-training data for the second meta-network. That is, during one or more rounds of model training of the first master network, the output recognition data (i.e., the recognition results of several supervision signals) of each round or more can be input into the second meta-network in real time for meta-network training. Meta-network training can realize the adaptive adjustment of model parameters of the first master network from multiple supervision signal recognition dimensions based on the output of the second meta-network, and the dynamic adjustment of model resource allocation for multiple supervision signal recognition dimensions during model training.

[0037] In one or more embodiments of this specification, the output loss adjustment parameter of the second-order network may be a loss adjustment weight for multiple monitoring signals, a sparsity intensity for the sparsity of multiple monitoring signals, etc.

[0038] Optionally, the second-order network included in the initial image recognition model can be composed of a fully connected layer MLP. The fully connected layer can also be regarded as a multilayer perceptron, which belongs to the multilayer fully connected neural network model. The second-order network assists or instructs the main network of the first main network to perform iterative updates of model parameters during the training process.

[0039] S108: Adjust the first main network model based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model.

[0040] In one feasible implementation, the initial image recognition model can be trained using an alternating network training method, and the number of the first training rounds for the first main network and the number of the second training rounds for the second meta-network can be determined based on the alternating network training method. Indicatively, the process could involve first training the first main network for a first number of training rounds, then training the second meta-network for a second number of training rounds after the first number of training rounds of main network training is completed (i.e., executing S106); then, after the second number of training rounds of meta-network training is completed, simultaneously executing S108 "adjusting the model of the first main network based on the loss adjustment parameters" and training the first main network for the next first number of training rounds (i.e., S104)... and so on, until the initial image recognition model meets the model termination conditions, thus obtaining the target image recognition model.

[0041] In one feasible implementation, the initial image recognition model can be trained using a synchronous network training method. First, the first main network can be trained for at least one round to accumulate input meta-training data for the second meta-network. Then, the first main network and the second meta-network can be trained synchronously, i.e., S104 and S106 are executed synchronously. During the synchronous training process, the first main network is adjusted based on the loss adjustment parameters output by the second meta-network until the model training termination condition corresponding to the initial image recognition model is met, thus obtaining the target image recognition model.

[0042] In one or more embodiments of this specification, the model termination condition may include, for example, the loss value of the loss function being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific model termination conditions can be determined based on actual circumstances and will not be elaborated here.

[0043] In one or more embodiments of this specification, the image recognition task is a liveness detection task, and the liveness detection supervision signal includes at least one of depth estimation supervision signal, image material classification supervision signal, reflectance spectrum prediction supervision signal, and liveness classification supervision signal.

[0044] Optionally, the loss function of the first main network can be adjusted based on the loss adjustment parameters output by the second-order network, so that model parameters can be adjusted based on the adjusted loss function during the subsequent model training of the first main network. For example, the connection weights and / or thresholds between neurons in each layer of the network can be adjusted by backpropagation based on the loss function.

[0045] In this specification, an electronic device can construct an initial image recognition model, including a first main network and a second meta-network, based on an image recognition task. The first main network is trained using image sample data to determine at least one supervisory signal recognition result. Then, based on the recognition results of each supervisory signal, the second meta-network is trained to determine loss adjustment parameters. Finally, the first main network is adjusted based on the loss adjustment parameters until a target image recognition model is obtained, corresponding to the initial image recognition model. By training image recognition from multiple dimensions using several supervisory signals and combining this with meta-network training, adaptive supervision based on loss adjustment parameters can be achieved. During model training, the model network structure and parameter allocation can be dynamically adjusted accurately and efficiently, achieving good resource utilization. While ensuring model performance, adaptive supervision and adjustment can reduce model resource consumption and significantly ensure the model's robustness and adaptability after deployment.

[0046] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating another embodiment of an image recognition method proposed in one or more embodiments of this specification. Specifically: S202: Construct an initial image recognition model based on the image recognition task, wherein the image recognition model includes a first master network and a second meta-network; For details, please refer to one or more embodiments of this specification for the method steps, which will not be repeated here.

[0047] S204: Configure at least one image recognition supervision signal for the first main network based on the image recognition task; The image recognition supervision signal can be immediately used as a type of model supervision signal to instruct or set the machine learning model to perform model recognition in the supervision signal dimension. The image recognition supervision signal is associated with the data type of the output of the image recognition model. For example, if the output of the image recognition model is text, then the image recognition supervision signal can be a text recognition supervision signal. If the output of the image recognition model is data for various dimensions of liveness detection, then the image recognition supervision signal can be at least one of depth estimation supervision signal, image material classification supervision signal, reflectance spectrum prediction supervision signal, and liveness classification supervision signal.

[0048] In one or more embodiments of this specification, at least one image recognition supervision signal for a recognition dimension can be configured on the first master network based on the image recognition task, so as to instruct the first master network to perform image recognition in the recognition dimension based on the image recognition supervision signal, and the output of the first master network is the result corresponding to the image recognition supervision signal of the corresponding recognition dimension.

[0049] In one or more embodiments of this specification, taking the liveness detection task as an example, it is usually necessary to combine the recognition results of multiple detection and recognition dimensions to comprehensively determine whether a target image is a live type or an attack type. The first main network included in the initial image recognition image performs recognition of the recognition dimensions corresponding to the image recognition supervision signals on the image sample data to obtain the results of the image recognition supervision signals. That is, the first main network outputs several supervision signal recognition results, such as at least one of the following supervision signals: depth estimation supervision signal, image material classification supervision signal, reflectance spectrum prediction supervision signal, liveness classification supervision signal, etc.

[0050] S206: Input the image sample data into the first main network for main network training, and output at least one supervision signal recognition result indicated by the image recognition supervision signal; The specific process is explained below: A2: The electronic device can input the image sample data into the first main network for at least one round of main network recognition; A4: In each round of the main network recognition process, the electronic device can output at least one of the image recognition supervision signals indicating the supervision signal recognition result, determine the supervision signal loss corresponding to each of the image recognition supervision signals, and adjust the first main network based on each supervision signal loss.

[0051] Understandably, image sample data is input into the first main network, which performs image recognition on the image sample data in recognition dimensions such as the depth estimation dimension corresponding to the depth estimation supervision signal and the image material classification dimension corresponding to the image material classification supervision signal, so as to output the corresponding supervision signal recognition results, such as the results corresponding to the depth estimation supervision signal, the results of the image material classification supervision signal, the reflection spectrum prediction results of the reflection spectrum prediction supervision signal, and so on.

[0052] Understandably, for each dimension of the image recognition process, a supervision signal loss function corresponding to the image recognition supervision signal can be pre-set for the first main network. That is, each image recognition supervision signal output by the model corresponds to a supervision signal loss function; for example, a depth estimation loss function can be set for the depth estimation dimension corresponding to the depth estimation supervision signal, and so on. In each round of the main network recognition process, the electronic device can output at least one supervision signal recognition result indicated by the image recognition supervision signal, and simultaneously obtain the supervision signal loss based on the determined supervision signal loss function corresponding to each image recognition supervision signal, so as to adjust the first main network based on each supervision signal loss.

[0053] Indicatively, the supervisory signal loss can typically be calculated using the supervisory signal loss function by applying the actual output value of the first master network in each round of image recognition and the labeled values ​​(i.e., theoretical output values) of the image sample data.

[0054] For example, taking the image recognition task as a liveness detection task, the set liveness detection supervision signal includes at least one of depth estimation supervision signal, image material classification supervision signal, reflectance spectrum prediction supervision signal, and liveness classification supervision signal. A depth estimation loss function is set for the depth estimation dimension corresponding to the depth estimation supervision signal, which can specifically be: The depth estimation supervision signal is used to instruct the depth estimation of the facial region based on the image. The depth estimation dimension can be set by the depth estimation loss function, for example, the loss function can be set to Euclidean distance loss in related techniques. Image material classification supervision signal is used to indicate the material of facial regions (e.g., normal facial material, paper material, screen material, etc.) based on modalities such as two-dimensional images and three-dimensional images. The image material classification supervision signal can be set with a corresponding image material classification loss function, for example, the loss function can be set as classification loss in related technologies. The reflectance spectrum prediction supervision signal is used to indicate the light reflectance characteristics of the facial region based on the image prediction of modalities such as two-dimensional images and three-dimensional images; the reflectance spectrum prediction supervision signal can be set with a corresponding reflectance spectrum prediction loss function, for example, it can be set to Euclidean distance loss in related techniques. The liveness classification supervision signal is used to instruct the classification of liveness / attack based on images of different modalities such as two-dimensional images and three-dimensional images; the liveness classification supervision signal can be configured with a corresponding liveness classification supervision loss function, for example, the loss function can be set as the classification loss in related technologies; It should be noted that the described "supervision signal loss function set for the result of image recognition supervision signal" is only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, other embodiments of the "supervision signal loss function set for the result of image recognition supervision signal" obtained by those skilled in the art based on related technologies without creative effort should fall within the scope of this application.

[0055] Optionally, the supervision signal loss function is set based on the corresponding supervision signal task. The output of the supervision signal loss function can be understood as a supervision loss for the supervision signal task specific to the corresponding supervision signal dimension. To further improve the model training effect, a sparse loss can also be set for each image recognition supervision signal. That is, a sparse loss function is introduced for each image recognition supervision signal, and the sparse loss function is combined with the supervision signal loss function. For each image recognition supervision signal, the supervision loss can be obtained by using the supervision signal loss function, and the sparse loss can be obtained by using the sparse loss function. The supervision loss and the sparse loss are used together as the supervision signal loss for the image recognition supervision signal. Furthermore, assuming the number of image recognition supervision signals is I, then the loss can be calculated using the Loss function. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal, which can be determined by the parameter. i Let represent the sparse loss corresponding to the i-th image recognition supervision signal, where i is a positive integer less than or equal to 1.

[0056] As an illustration, the sparse loss function can be set based on the algorithm for calculating sparse loss in related techniques. In one feasible implementation, the electronic device performing the master network adjustment of the first master network based on the losses of each of the monitoring signals may be: B2: The electronic device can acquire the target loss parameters for the first main network; The target loss parameter is associated with the loss function of the first main network (such as the first loss calculation formula). In one or more embodiments of this specification, the target loss parameter is the loss parameter or loss factor in the first loss calculation formula. In this specification, during the initial training of the image recognition model, the target loss parameter corresponding to the first loss calculation formula is updated based on the output loss adjustment parameter of the second network as the model training progresses, so as to achieve the effect of subsequently adjusting the first main network based on the loss adjustment parameter.

[0057] B4: Adjust the first master network based on the target loss parameters and the losses of each of the supervision signals.

[0058] Indicatively, in the initial stage of image training, before the training process of the second-level network is initiated, an initial value can be set for the target loss parameter in the loss function (first loss calculation formula). Subsequently, after the training process of the second-level network is initiated, the target loss parameter is adjusted based on the output loss adjustment parameter of the second-level network to obtain the adjusted target loss parameter. Then, based on the target loss parameter and the losses of each supervision signal, the model loss is output to the loss function. Based on this model loss, the main network parameters of the first main network are adjusted, such as adjusting the connection weights and / or thresholds between neurons in each layer of the network through backpropagation iterative adjustment based on the loss function, until the initial image recognition model meets the model termination conditions, such as the first loss being less than or equal to the loss threshold, or the total number of training rounds reaching the training round threshold, thus obtaining the trained target image recognition model.

[0059] Optionally, the supervision signal loss may be the supervision loss and sparsity loss included in some embodiments.

[0060] In a specific implementation scenario, taking the supervision signal loss as including supervision loss and sparse loss, and the target loss parameters as the supervision signal weights and signal sparsity intensity for the supervision loss, as an example, the details are as follows: The electronic device performing the main network adjustment of the first main network based on the target loss parameters and the losses of each of the supervision signals can be: The electronic device inputs the supervision loss, the sparsity loss, the supervision signal weight, and the signal sparsity intensity into the first loss calculation formula, calculates and determines the first loss based on the first loss calculation formula, and performs main network adjustment on the first main network based on the first loss. Furthermore, the first loss calculation formula can be regarded to some extent as the model loss function of the entire initial image recognition model. Based on the first loss calculation formula, it can be regarded as the model loss of the initial image recognition model. Usually, when the first loss calculated based on the first loss calculation formula converges, if the first loss is less than the loss threshold in the model end training condition, it is determined that the model has converged and the model training ends, thus obtaining the target image recognition model.

[0061] Furthermore, the first loss calculation formula satisfies the following formula:

[0062] Among them, Loss A The first loss is denoted by I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal, α i For the supervision signal weights of the i-th image recognition supervision signal, parameter i β is the sparse loss corresponding to the i-th image recognition supervision signal. i Let be the signal sparsity intensity for the i-th image recognition supervision signal.

[0063] S208: Based on the identification results of each of the supervision signals, perform meta-network training on the second meta-network to determine the loss adjustment parameters; In one or more embodiments of this specification, the loss adjustment parameter is used to adjust the target loss parameter corresponding to the first calculation formula. For example, the target loss parameter may be the supervision signal weight or the sparse signal strength.

[0064] In one feasible implementation, the electronic device performs at least one round of meta-network training by inputting the identification results of each of the supervision signals into the second meta-network; During each round of meta-network training, the electronic device outputs loss adjustment parameters for each image recognition supervision signal, obtains the supervision signal loss corresponding to each image recognition supervision signal based on the first main network, and performs meta-network adjustment on the second meta-network based on each supervision signal loss.

[0065] Understandably, during the training of the first main network, the supervisory signal loss corresponding to the image recognition supervisory signal is calculated at the same time as the output of the current round's image recognition supervisory signal. Based on this, the second sub-network can directly obtain the supervisory signal loss corresponding to the image recognition supervisory signal already calculated by the first main network. Then, when training the second sub-network based on the current image recognition supervisory signal, the second sub-network is simultaneously adjusted based on the supervisory signal loss, such as adjusting the connection weights and / or thresholds between neurons in each layer of the second sub-network through backpropagation based on the loss function.

[0066] In one feasible implementation, the electronic device performing the meta-network adjustment of the second meta-network based on the losses of each of the monitoring signals may be: The electronic device inputs the losses of each monitoring signal into the second loss calculation formula to determine the second loss, and performs meta-network adjustment on the second meta-network based on the second loss. The second loss calculation formula can be understood as the meta-network loss function of the second meta-network.

[0067] Optionally, the second loss calculation formula can satisfy the following formula:

[0068] Among them, Loss B The second loss is I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal.

[0069] Optionally, the weights of the image recognition supervision signal can be introduced into the second loss calculation formula to further accelerate model convergence. Furthermore, the second loss calculation formula can satisfy the following equation:

[0070] Among them, Loss B The second loss is I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal, α i The supervision signal weight is defined for the i-th image recognition supervision signal.

[0071] Schematic, in each round of meta-network training, after outputting the loss adjustment parameters for each of the image recognition supervision signals, the loss adjustment parameters typically include the supervision signal weights α for each of the image recognition supervision signals. iBased on this, the electronic device will change the original target loss parameter α i Update to α in the current loss adjustment parameter i .

[0072] S210: Based on the loss adjustment parameters, update the first loss calculation formula of the first main network to obtain the updated first loss calculation formula; Furthermore, after the electronic device completes at least one round of training on the second meta-network, it adjusts the first main network based on the output loss adjustment parameters until the initial image recognition model reaches the end-of-training condition, thereby obtaining the target image recognition model for the initial image recognition model.

[0073] Furthermore, the electronic device performing model adjustment on the first master network based on the output loss adjustment parameters can be as follows: the electronic device updates the loss parameters of the first loss calculation formula of the first master network based on the loss adjustment parameters to obtain the updated first loss calculation formula.

[0074] Specifically, the electronic device acquires the target loss parameter in the first loss calculation formula for the first master network, that is, the current supervision signal weights and / or signal sparsity intensity in the first loss calculation formula. Then, based on the loss adjustment parameters (new supervision signal weights and / or signal sparsity intensity), it updates the target loss parameter (original supervision signal weights and / or original signal sparsity intensity) to obtain the updated target loss parameter. In this way, during the next round of training the first master network, the first loss calculation formula can be updated based on the new target loss parameter.

[0075] S212: During the training of the first main network, the first main network is adjusted based on the first loss calculation formula until a target image recognition model for the initial image recognition model is obtained.

[0076] In one feasible implementation, the initial image recognition model can be trained using a synchronous network training method. First, the first master network can be trained for at least one round to accumulate input meta-training data for the second meta-network. Then, the first master network and the second meta-network are trained synchronously, i.e., S104 and S106 are executed synchronously. During synchronous training, the first loss calculation formula of the first master network is updated based on the loss adjustment parameters output by the second meta-network, resulting in an updated first loss calculation formula. When the first master network is trained for the next master network iteration, the first master network is adjusted based on the updated first loss calculation formula until the model training termination condition corresponding to the initial image recognition model is met, thus obtaining the target image recognition model.

[0077] In a specific implementation scenario, the initial image recognition model can be trained using an alternating network training method, as follows: C2: The electronic device can determine that the model training method for the initial image recognition model is the network alternation training method, and determine the first training round number for the first main network and the second training round number for the second meta-network based on the network alternation training method; As an illustration, the number of the first and second training rounds corresponding to the alternating training method of the network can be customized. For example, the number of the first training rounds for the first main network can be x rounds, and the number of the second training rounds for the second meta-network can be y rounds.

[0078] C4: Based on the first training round number, the first main network is trained using image sample data to determine at least one supervisory signal recognition result; Indicatively, the input to the first master network is image sample data; the output of the first master network is the recognition result of at least one supervisory signal.

[0079] Indicatively, during each training round of the main network in the first training round (e.g., 10 rounds), at least one image recognition supervision signal indicates the supervision signal recognition result, and the supervision signal loss corresponding to each image recognition supervision signal is determined. The supervision signal loss can be composed of supervision loss and sparse loss. Based on each supervision signal loss, the first main network is adjusted using a first calculation formula.

[0080] Schematic illustration: The target loss parameter in the first loss calculation formula for the first master network is obtained. When meta-network training is not started, the target loss parameter is a user-defined initial value. After meta-network training starts, the target loss parameter can be adjusted and updated based on the output loss adjustment parameter of the second meta-network. Further, the supervised loss, the sparse loss, the supervised signal weights, and the signal sparsity intensity are input into the first loss calculation formula to determine the first loss; the master network is then adjusted based on the first loss. C6: Based on the second training round number and the recognition results of each of the supervision signals, perform meta-network training on the second meta-network to determine the loss adjustment parameters.

[0081] After the main network training for the first training round (e.g., 10 rounds) is completed, several sets of supervision signal recognition results are accumulated according to the first training round (e.g., if the first training round is 10 rounds, there are 10 sets of supervision signal recognition results). Then, these sets of supervision signal recognition results are input into the second training round for allocation to determine the supervision signal recognition results to be input in each round. The supervision signal recognition results are then input into the second meta-network for training the meta-network corresponding to the second training round (e.g., 1 round). During each round of meta-network training, loss adjustment parameters for each image recognition supervision signal are output, and the supervision signal loss corresponding to each image recognition supervision signal is obtained based on the first main network. Meta-network adjustment is then performed on the second meta-network based on each supervision signal loss.

[0082] Optionally, the model termination training condition is usually set during the training of the meta-network. The initial image recognition model usually does not converge. The loss adjustment parameters of the output of each round during the training of the meta-network are updated with the loss parameters of the first loss calculation formula of the first main network.

[0083] C8: If the initial image recognition model does not meet the model training termination condition, then the loss parameter of the first loss calculation formula of the first main network is updated based on the loss adjustment parameter to obtain the updated first loss calculation formula. C10: When training the first main network in the next round, the first main network is adjusted based on the first loss calculation formula until the initial image recognition model meets the model training termination condition, thereby obtaining the target image recognition model for the initial image recognition model.

[0084] Furthermore, when training the first master network in the next round, that is, when performing the C4 step based on the next first training round number, the first master network is adjusted based on the first loss calculation formula.

[0085] Furthermore, during each round of main network training, it is checked whether the model training termination condition is met, such as the first loss being less than or equal to a loss threshold, or the total number of training rounds reaching a threshold. If the model training termination condition is not met, C6 is executed after step C4, and the first main network and the second sub-network are trained alternately based on the network alternating training method until the initial image recognition model meets the model training termination condition, thus obtaining the target image recognition model for the initial image recognition model.

[0086] Optionally, when the initial image recognition model meets the model training termination condition, the electronic device can use the initial image recognition model at this time as the target image recognition model. It should be noted that when the initial image recognition model at this time is used as the target image recognition model, in the subsequent model application stage, image recognition is usually only based on the first main network of the target image recognition model. After online deployment, since the target image recognition model also includes a second meta-network, the target image recognition model can be retrained based on actual online image data to enhance the robustness of the model after online deployment, effectively enhancing the stability of the model after online deployment and the image scene generalization ability.

[0087] Optionally, when the initial image recognition model meets the model training termination condition, the electronic device can use the first main network in the initial image recognition model as the target image recognition model, that is, discard the second network in the initial image recognition model and retain only the second network to obtain the target image recognition model, so as to perform lightweight processing on the model and facilitate its deployment to more implementation scenarios.

[0088] In one or more embodiments of this specification, after obtaining the target image recognition model, model pruning can be performed to obtain a lightweight model network.

[0089] As an illustration, the model pruning method in related technologies can be used to prune the target image recognition model. For example, the neuron parameters of the model are pruned to a specified value (such as 0). Through model pruning, a lightweight target image recognition model is obtained. At this time, the parameters and resources allocated to each supervision signal in the model are no longer consistent, and better resource allocation is achieved.

[0090] It is understood that the image recognition method of one or more embodiments of this specification realizes automatic adjustment of the weights of multiple supervision signals and the model parameters allocated to the multiple supervision signals through the model training method of the second network and the first network, thereby reducing the consumption of resources while ensuring performance.

[0091] In this specification, an electronic device can construct an initial image recognition model, including a first main network and a second meta-network, based on an image recognition task. The first main network is trained using image sample data to determine at least one supervisory signal recognition result. Then, based on the recognition results of each supervisory signal, the second meta-network is trained to determine loss adjustment parameters. Finally, the first main network is adjusted based on the loss adjustment parameters until a target image recognition model is obtained, corresponding to the initial image recognition model. By training image recognition from multiple dimensions using several supervisory signals and combining this with meta-network training, adaptive supervision based on loss adjustment parameters can be achieved. During model training, the model network structure and parameter allocation can be dynamically adjusted accurately and efficiently, achieving good resource utilization. While ensuring model performance, adaptive supervision and adjustment can reduce model resource consumption and significantly ensure the model's robustness and adaptability after deployment.

[0092] Please see Figure 3 , Figure 3 This is a schematic flowchart illustrating another embodiment of an image recognition method proposed in one or more embodiments of this specification. Specifically: S302: Construct an initial image recognition model based on the liveness detection task, wherein the image recognition model includes a first main network and a second-element network; S304: Use image sample data to train the first main network and determine at least one supervisory signal recognition result; S306: Based on the identification results of each of the supervision signals, perform meta-network training on the second meta-network to determine the loss adjustment parameters; S308: Adjust the model of the first main network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model; For details, please refer to one or more embodiments of this specification for the method steps, which will not be repeated here.

[0093] S310: Obtain the target detection image corresponding to the liveness detection task; In recent years, biometric technology has been widely applied to people's production and daily life. For example, facial recognition payment, facial access control, facial attendance, and facial recognition station entry all rely on biometrics. However, with the increasingly widespread application of biometric technology, the demand for liveness detection in biometric scenarios is becoming increasingly prominent. While biometrics provides convenience, it also brings new risks and challenges. The most common means of threatening the security of biometric systems is liveness attacks, which involve attempting to bypass image biometric verification through means such as device screens or printed photos. To detect liveness attacks, liveness prevention technology has become an essential component of biometric scenarios. The liveness detection task (also referred to as the liveness recognition task) in one or more embodiments of this specification is also a crucial part of biometric scenarios.

[0094] In related technologies, liveness detection is a method for determining the true physiological characteristics of an object in identity verification scenarios. In facial recognition applications, liveness detection often uses a combination of actions such as blinking, opening the mouth, shaking the head, and nodding to verify whether the user is a real, living person. Liveness detection tasks need to effectively resist common liveness attack methods such as photos, face swapping, masks, occlusion, and screen captures to help users identify fraudulent activities and protect their interests. However, using a combination of actions like blinking, opening the mouth, shaking the head, and nodding requires a high degree of user cooperation, which users often resist in practice. Furthermore, applications requiring such high user cooperation are somewhat unreasonable. Therefore, the target image recognition model obtained using the image recognition method described in this specification can be applied to the liveness detection task. Since the target image recognition model does not require high user cooperation to complete the combined actions, the liveness detection process is optimized, and the liveness detection effect is improved.

[0095] The target detection image can be a biological image to be detected in a biometrics scenario, such as a facial image, fingerprint image, etc.

[0096] S312: Input the target detection image into the target image recognition model, and output at least one target supervised recognition result for the target detection image; The target image recognition model performs image recognition from various recognition dimensions of multiple image recognition supervision signals based on the liveness detection task. It performs image recognition on the target detection image input to the model and outputs at least one target supervision recognition result for the target detection image.

[0097] Schematic, the liveness detection supervision signal includes at least one of the following types: depth estimation supervision signal, image material classification supervision signal, reflectance spectrum prediction supervision signal, and liveness classification supervision signal; Schematic, the target supervision signal recognition result includes at least one of the following: depth estimation recognition result, image material classification recognition result, reflectance spectrum prediction recognition result, and liveness classification recognition result.

[0098] S314: Based on the target supervision and recognition results, determine the image detection type corresponding to the target detection image, wherein the image detection type includes a liveness image type and an attack image type.

[0099] Understandably, electronic devices can further determine the image detection type based on the target supervision and recognition results of each liveness detection dimension, that is, whether the target detection image is a liveness image or an attack image.

[0100] In one feasible implementation, a liveness monitoring signal partitioning rule can be set in advance based on the liveness detection task requirements in the actual liveness detection scenario. The liveness monitoring signal partitioning rule is used to determine whether the target detection image is a liveness image type or an attack image type based on the target monitoring recognition results of each liveness detection dimension.

[0101] In specific implementation, the electronic device can determine the liveness supervision signal partitioning rule corresponding to the liveness supervision recognition result of each target, use the liveness supervision signal partitioning rule to perform liveness detection on each target supervision recognition, and determine the image detection type corresponding to the target detection image based on the liveness detection result.

[0102] Indicatively, the rules for classifying liveness monitoring signals can be: if the liveness classification result is a liveness category, then the target detection image can be determined to be a liveness image type; if the liveness classification result is an attack category, then the target detection image can be determined to be an attack image type. Indicatively, the rule for classifying liveness monitoring signals can be: if the average output value of the depth estimation recognition result is greater than a pre-set threshold, the target detection image is considered to be an attack image type; otherwise, the target detection image is considered to be a liveness image type. Indicatively, the rule for classifying liveness monitoring signals can be: if the image material classification and recognition result is not the object's face category, then the target detection image is considered to be an attack image type; conversely, if the image material classification and recognition result is the object's face category, then the target detection image is considered to be a liveness image type. Indicatively, the rule for classifying liveness monitoring signals can be: if the variance indicated by the reflectance spectrum prediction and recognition result is less than the set variance threshold, then the target detection image is considered to be an attack image type; otherwise, the target detection image is considered to be a liveness image type. It is understood that the image recognition method of one or more embodiments in this specification, based on image recognition and detection (such as liveness detection) based on multi-supervised signal dimensions (multi-angle), dynamically implements resource allocation for each supervision dimension of the model based on a second-order network, and adaptively adjusts the model network structure to achieve good resource utilization. It also balances user experience, computational resource consumption, and attack detection performance.

[0103] In one or more embodiments of this specification, image recognition training from multiple dimensions using several supervisory signals, combined with meta-network training, enables adaptive supervision based on loss-adjusted parameters. During model training, the model network structure and parameter allocation can be dynamically adjusted accurately and efficiently, achieving better resource utilization. While ensuring model performance, adaptive supervision reduces model resource consumption and significantly ensures model robustness and adaptability after deployment. Furthermore, in biometric scenarios related to liveness detection, relevant technologies require user interaction, such as head shaking or blinking upon prompting. After online deployment of the target image recognition model obtained in this specification, liveness detection does not require extensive user cooperation, improving the user experience.

[0104] The following will combine Figure 5 This manual provides a detailed description of the image recognition device provided. It should be noted that... Figure 5 The image recognition device shown is used to perform the present application. Figures 1-4 The methods of the embodiments shown are illustrated only in the parts relevant to this specification for ease of explanation. For specific technical details not disclosed, please refer to this application. Figures 1-4 The example shown.

[0105] Please see Figure 5 This diagram illustrates the structure of the image recognition device described in this specification. The image recognition device 1000 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the image recognition device 1000 includes a model building module 11 and a model training module 12, specifically used for: Model building module 11 is used to build an initial image recognition model based on the image recognition task. The image recognition model includes a first main network and a second meta-network. Model training module 12 is used to train the first main network using image sample data and determine at least one supervision signal recognition result; The model training module 12 is used to train the second meta-network based on the recognition results of each supervision signal and to determine the loss adjustment parameters. The model training module 12 is used to adjust the first main network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model. Optional, such as Figure 6 As shown, the model training module 12 includes: Network configuration unit 121 is used to configure at least one image recognition supervision signal for the first main network based on the image recognition task; The network training unit 122 is used to input the image sample data into the first main network for main network training and output at least one supervision signal recognition result indicated by the image recognition supervision signal.

[0106] Optional, such as Figure 7 As shown, the network training unit 122 includes: The network training subunit 1221 is used to input the image sample data into the first main network for at least one round of main network training. The network adjustment subunit 1222 is used to output at least one supervision signal recognition result indicated by the image recognition supervision signal during each round of main network training, and to determine the supervision signal loss corresponding to each image recognition supervision signal, and to adjust the first main network based on each supervision signal loss.

[0107] Optionally, the network adjustment subunit 1222 is used for: Obtain the target loss parameters for the first main network; The first master network is adjusted based on the target loss parameters and the losses of each of the supervision signals.

[0108] Optionally, the supervision signal loss includes supervision loss and sparse loss, the target loss parameters are the supervision signal weights and signal sparsity intensity for the supervision loss, and the network adjustment subunit 1222 is used for: The supervision loss, the sparse loss, the supervision signal weight, and the signal sparsity intensity are input into the first loss calculation formula to determine the first loss; Based on the first loss, adjust the first main network; The first loss calculation formula satisfies the following formula:

[0109] Among them, Loss A The first loss is denoted by I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal, αi For the supervision signal weights of the i-th image recognition supervision signal, parameter i β is the sparse loss corresponding to the i-th image recognition supervision signal. i Let be the signal sparsity intensity for the i-th image recognition supervision signal.

[0110] Optionally, the network training unit 122 is used for: The identification results of each of the aforementioned supervisory signals are input into the second meta-network for at least one round of meta-network training; In each round of meta-network training, loss adjustment parameters for each image recognition supervision signal are output, and the supervision signal loss corresponding to each image recognition supervision signal is obtained based on the first main network. Meta-network adjustment is performed on the second meta-network based on each supervision signal loss.

[0111] Optionally, the network adjustment subunit 1222 is used for: The losses of each of the aforementioned monitoring signals are input into the second loss calculation formula to determine the second loss; Based on the second loss, perform meta-network adjustments on the second meta-network; The second loss calculation formula satisfies the following formula:

[0112] Among them, Loss B The second loss is I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal.

[0113] Optionally, the model training module 12 is specifically used for: Based on the loss adjustment parameters, the first loss calculation formula of the first main network is updated to obtain the updated first loss calculation formula. During the training of the first main network, the first main network is adjusted based on the first loss calculation formula until a target image recognition model is obtained for the initial image recognition model.

[0114] Optionally, the model training module 12 is specifically used for: Obtain the target loss parameter in the first loss calculation formula for the first main network; The target loss parameter is updated based on the loss adjustment parameter to obtain the updated target loss parameter.

[0115] Optionally, the model building module 11 is specifically used for: The model training method for the initial image recognition model is determined to be the network alternation training method. Based on the network alternation training method, the first training round number for the first main network and the second training round number for the second meta-network are determined. Optionally, the model training module 12 is specifically used for: Based on the first training round, the first main network is trained using image sample data to determine at least one supervisory signal recognition result. Based on the second training round number and the identification results of each supervision signal, the second meta-network is trained to determine the loss adjustment parameters.

[0116] Optionally, the model training module 12 is specifically used for: If the initial image recognition model does not meet the model training termination condition, then the first loss calculation formula of the first main network is updated based on the loss adjustment parameter to obtain the updated first loss calculation formula. During the next round of training of the first main network, the first main network is adjusted based on the first loss calculation formula until the initial image recognition model meets the model training termination condition, thereby obtaining the target image recognition model for the initial image recognition model.

[0117] Optionally, the model training module 12 is specifically used for: Use the initial image recognition model as the target image recognition model; or, The first master network in the initial image recognition model is used as the target image recognition model.

[0118] Optionally, the image recognition task is a liveness detection task, and the supervision signal recognition result includes at least one of depth estimation recognition result, image material classification recognition result, reflectance spectrum prediction recognition result, and liveness classification recognition result.

[0119] Optional, such as Figure 8 As shown, the image recognition device 1000 further includes: The model recognition module 13 is used to acquire the target detection image corresponding to the liveness detection task, input the target detection image into the target image recognition model, and output at least one target supervised recognition result for the target detection image; Image detection module 14 is used to determine the image detection type corresponding to the target detection image based on the supervised recognition results of each target, wherein the image detection type includes liveness image type and attack image type.

[0120] Optionally, the image detection module 14 is used for: Determine the liveness monitoring signal segmentation rules corresponding to the monitoring and identification results of each target. The liveness detection is performed on each of the target supervision recognitions using the liveness supervision signal segmentation rules, and the image detection type corresponding to the target detection image is determined based on the liveness detection results.

[0121] It should be noted that the image recognition device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the image recognition method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the image recognition device and the image recognition method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0122] The serial numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0123] In one or more embodiments of this specification, image recognition training from multiple dimensions using several supervisory signals, combined with meta-network training, enables adaptive supervision based on loss-adjusted parameters. During model training, the model network structure and parameter allocation can be dynamically adjusted accurately and efficiently, achieving better resource utilization. While ensuring model performance, adaptive supervision reduces model resource consumption and significantly ensures model robustness and adaptability after deployment. Furthermore, in biometric scenarios related to liveness detection, relevant technologies require user interaction, such as head shaking or blinking upon prompting. After online deployment of the target image recognition model obtained in this specification, liveness detection does not require extensive user cooperation, improving the user experience.

[0124] This specification also provides a computer storage medium capable of storing multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-4 The image recognition method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-4 The specific details of the illustrated embodiments will not be elaborated here.

[0125] This application also provides a computer program product storing at least one instruction, which is loaded and executed by the processor as described above. Figures 1-4 The image recognition method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-4 The specific details of the illustrated embodiments will not be elaborated here.

[0126] Please refer to Figure 9 This diagram illustrates a structural block diagram of an electronic device provided in an exemplary embodiment of this application. The electronic device in this application may include one or more components such as a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 may be connected via the bus 150.

[0127] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device via various interfaces and lines, and performs various functions and processes data of electronic device 200 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of the following: central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.

[0128] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems. The data storage area may also store data created by the electronic device during use, such as phonebook data, audio and video data, chat log data, etc.

[0129] See Figure 10 As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, while native and third-party applications run in the user space. To ensure that different third-party applications can achieve good running performance, the operating system allocates corresponding system resources for each application. However, different application scenarios within the same third-party application have different requirements for system resources. For example, in local resource loading scenarios, third-party applications have high requirements for disk read speed; in animation rendering scenarios, third-party applications have high requirements for GPU performance. Since the operating system and third-party applications are independent of each other, the operating system often cannot promptly perceive the current application scenario of a third-party application, resulting in the operating system's inability to adapt system resources accordingly to the specific application scenario of the third-party application.

[0130] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0131] Taking the Android operating system as an example, the programs and data stored in memory 120 are as follows: Figure 11As shown, the memory 120 can store the Linux kernel layer 320, the system runtime library layer 340, the application framework layer 360, and the application layer 380. The Linux kernel layer 320, system runtime library layer 340, and application framework layer 360 belong to the operating system space, while the application layer 380 belongs to the user space. The Linux kernel layer 320 provides low-level drivers for various hardware components of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, and power management. The system runtime library layer 340 provides support for key features of the Android system through several C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D graphics support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library, which mainly provides core libraries that allow developers to write Android applications using the Java language. The Application Framework Layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application runs in the Application Layer 380. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera apps; or third-party applications developed by third-party developers, such as games, instant messaging, and photo editing apps.

[0132] Taking the operating system as an example (iOS), the programs and data stored in memory 120 are as follows: Figure 12As shown, the iOS system includes: Core OS layer 420, Core Services layer 440, Media layer 460, and Cocoa Touch layer 480. Core OS layer 420 includes the operating system kernel, drivers, and low-level program frameworks. These low-level program frameworks provide hardware-level functionality for use by the program frameworks located in Core Services layer 440. Core Services layer 440 provides system services and / or program frameworks required by applications, such as Foundation framework, account framework, advertising framework, data storage framework, network connectivity framework, geolocation framework, motion framework, etc. Media layer 460 provides applications with audiovisual interfaces, such as interfaces related to graphics and images, audio technology, video technology, and wireless playback (AirPlay) interfaces. Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development and is responsible for user touch interaction on electronic devices. Examples include local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (UI) frameworks, UIKit user interface frameworks, map frameworks, and so on.

[0133] exist Figure 12 The framework shown includes, but is not limited to, the base framework in the core service layer 440 and the UIKit framework in the touchable layer 480. The base framework provides many basic object classes and data types, offering the most basic system services to all applications, and is independent of the UI. The UIKit framework, on the other hand, provides a basic UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UI, thus providing the application's infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.

[0134] The methods and principles for implementing data communication between third-party applications and the operating system in the iOS system can be referenced from the Android system, and will not be elaborated here.

[0135] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined into a touch screen, which is used to receive touch operations from the user using a finger, stylus, or any suitable object on or near it, and to display the user interface of various applications. The touch screen is usually located on the front panel of the electronic device. The touch screen can be designed as a full-screen, curved screen, or irregularly shaped screen. The touch screen can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; this specification does not limit this.

[0136] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.

[0137] In this specification, the entity executing each step can be the electronic device described above. Optionally, the entity executing each step can be the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems; this specification does not limit this.

[0138] The electronic device described in this manual may also be equipped with a display device. This display device can be any device capable of displaying information, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an e-ink screen, a liquid crystal display (LCD), or a plasma display panel (PDP). Users can use the display device on electronic device 101 to view displayed text, images, videos, and other information. The electronic device may be a smartphone, tablet, gaming device, AR (Augmented Reality) device, automobile, data storage device, audio playback device, video playback device, laptop, desktop computing device, or wearable device such as an electronic watch, electronic glasses, electronic helmet, electronic bracelet, electronic necklace, or electronic clothing.

[0139] exist Figure 9 In the illustrated electronic device, the processor 110 can be used to call the application program stored in the memory 120 and specifically perform the following operations: An initial image recognition model is constructed based on the image recognition task. The image recognition model includes a first master network and a second meta-network. The first main network is trained using image sample data to determine at least one supervisory signal recognition result. Based on the identification results of each of the supervision signals, the second meta-network is trained to determine the loss adjustment parameters. The first master network is adjusted based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model.

[0140] In one embodiment, when the processor 110 performs master network training on the first master network using image sample data and determines at least one supervisory signal recognition result, it specifically performs the following operations: Based on the image recognition task, at least one image recognition supervision signal is configured for the first main network; The image sample data is input into the first main network for main network training, and at least one supervision signal recognition result indicated by the image recognition supervision signal is output.

[0141] In one embodiment, when the processor 110 performs the following operations to input the image sample data into the first main network for main network training and output at least one supervision signal recognition result indicated by the image recognition supervision signal: The image sample data is input into the first main network for at least one round of main network training. During each round of main network training, at least one image recognition supervision signal indicates the supervision signal recognition result, and the supervision signal loss corresponding to each image recognition supervision signal is determined. The first main network is then adjusted based on each supervision signal loss.

[0142] In one embodiment, when the processor 110 performs the master network adjustment based on the losses of each of the supervision signals, it specifically executes the following steps: Obtain the target loss parameters for the first main network; The first master network is adjusted based on the target loss parameters and the losses of each of the supervision signals.

[0143] In one embodiment, the supervision signal loss includes supervision loss and sparse loss, and the target loss parameter is the supervision signal weight and signal sparsity intensity for the supervision loss. When the processor 110 performs the main network adjustment based on the target loss parameter and each of the supervision signal losses, it specifically executes the following steps: The supervision loss, the sparse loss, the supervision signal weight, and the signal sparsity intensity are input into the first loss calculation formula to determine the first loss; Based on the first loss, adjust the first main network; The first loss calculation formula satisfies the following formula:

[0144] Among them, Loss A The first loss is denoted by I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal, α i For the supervision signal weights of the i-th image recognition supervision signal, parameter i β is the sparse loss corresponding to the i-th image recognition supervision signal. i Let be the signal sparsity intensity for the i-th image recognition supervision signal.

[0145] In one embodiment, when the processor 110 performs meta-network training on the second meta-network based on the identification results of each of the supervision signals and determines the loss adjustment parameters, it specifically performs the following steps: The identification results of each of the aforementioned supervisory signals are input into the second meta-network for at least one round of meta-network training; In each round of meta-network training, loss adjustment parameters for each image recognition supervision signal are output, and the supervision signal loss corresponding to each image recognition supervision signal is obtained based on the first main network. Meta-network adjustment is performed on the second meta-network based on each supervision signal loss.

[0146] In one embodiment, when the processor 110 performs the meta-network adjustment of the second meta-network based on the losses of each of the supervision signals, it specifically executes the following steps: The losses of each of the aforementioned monitoring signals are input into the second loss calculation formula to determine the second loss; Based on the second loss, perform meta-network adjustments on the second meta-network; The second loss calculation formula satisfies the following formula:

[0147] Among them, Loss B The second loss is I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal.

[0148] In one embodiment, when the processor 110 performs model adjustment on the first master network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model, it specifically executes the following steps: Based on the loss adjustment parameters, the first loss calculation formula of the first main network is updated to obtain the updated first loss calculation formula. During the training of the first main network, the first main network is adjusted based on the first loss calculation formula until a target image recognition model is obtained for the initial image recognition model.

[0149] In one embodiment, when the processor 110 updates the loss parameters of the first loss calculation formula of the first master network based on the loss adjustment parameters to obtain the updated first loss calculation formula, it specifically performs the following steps: Obtain the target loss parameter in the first loss calculation formula for the first main network; The target loss parameter is updated based on the loss adjustment parameter to obtain the updated target loss parameter.

[0150] In one embodiment, when executing the image recognition method, the processor 110 further performs the following steps: The model training method for the initial image recognition model is determined to be the network alternation training method. Based on the network alternation training method, the first training round number for the first main network and the second training round number for the second meta-network are determined. The process of training the first master network using image sample data to determine at least one supervisory signal recognition result, and training the second meta-network based on each supervisory signal recognition result to determine loss adjustment parameters includes: Based on the first training round, the first main network is trained using image sample data to determine at least one supervisory signal recognition result. Based on the second training round number and the identification results of each supervision signal, the second meta-network is trained to determine the loss adjustment parameters.

[0151] In one embodiment, when the processor 110 performs model adjustment on the first master network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model, it specifically executes the following steps: If the initial image recognition model does not meet the model training termination condition, then the first loss calculation formula of the first main network is updated based on the loss adjustment parameter to obtain the updated first loss calculation formula. During the next round of training of the first main network, the first main network is adjusted based on the first loss calculation formula until the initial image recognition model meets the model training termination condition, thereby obtaining the target image recognition model for the initial image recognition model.

[0152] In one embodiment, when the processor 110 executes the process of obtaining a target image recognition model for the initial image recognition model, it specifically performs the following steps: Use the initial image recognition model as the target image recognition model; or, The first master network in the initial image recognition model is used as the target image recognition model.

[0153] In one embodiment, the image recognition task is a liveness detection task, and the supervisory signal recognition result includes at least one of depth estimation recognition result, image material classification recognition result, reflectance spectrum prediction recognition result, and liveness classification recognition result.

[0154] In one embodiment, when executing the image recognition method, the processor 110 further performs the following steps: Obtain the target detection image corresponding to the liveness detection task; The target detection image is input into the target image recognition model, and at least one target supervised recognition result is output for the target detection image; Based on the target supervision and recognition results, the image detection type corresponding to the target detection image is determined, and the image detection type includes liveness image type and attack image type.

[0155] In one embodiment, when the processor 110 performs the step of determining the image detection type corresponding to the target detection image based on each of the target supervision recognition results, it specifically executes the following steps: Determine the liveness monitoring signal segmentation rules corresponding to the monitoring and identification results of each target. The liveness detection is performed on each of the target supervision recognitions using the liveness supervision signal segmentation rules, and the image detection type corresponding to the target detection image is determined based on the liveness detection results.

[0156] In one or more embodiments of this specification, image recognition training from multiple dimensions using several supervisory signals, combined with meta-network training, enables adaptive supervision based on loss-adjusted parameters. During model training, the model network structure and parameter allocation can be dynamically adjusted accurately and efficiently, achieving better resource utilization. While ensuring model performance, adaptive supervision reduces model resource consumption and significantly ensures model robustness and adaptability after deployment. Furthermore, in biometric scenarios related to liveness detection, relevant technologies require user interaction, such as head shaking or blinking upon prompting. After online deployment of the target image recognition model obtained in this specification, liveness detection does not require extensive user cooperation, improving the user experience.

[0157] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0158] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. An image recognition method, the method comprising: An initial image recognition model is constructed based on the image recognition task. The image recognition model includes a first master network and a second meta-network. The first main network is trained using image sample data to determine at least one supervisory signal recognition result. The step of training the first main network using image sample data to determine at least one supervisory signal recognition result includes: configuring at least one image recognition supervisory signal for the first main network based on the image recognition task, inputting image sample data into the first main network for main network training, and outputting at least one supervisory signal recognition result indicated by the image recognition supervisory signal. Based on the identification results of each of the supervision signals, the second meta-network is trained to determine the loss adjustment parameters. The first master network is adjusted based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model.

2. The method according to claim 1, wherein inputting image sample data into the first main network for main network training and outputting at least one supervision signal recognition result indicated by the image recognition supervision signal includes: Image sample data is input into the first main network for at least one round of main network training; During each round of main network training, at least one image recognition supervision signal indicates the supervision signal recognition result, and the supervision signal loss corresponding to each image recognition supervision signal is determined. The first main network is then adjusted based on each supervision signal loss.

3. The method according to claim 2, wherein adjusting the first master network based on the loss of each of the supervision signals comprises: Obtain the target loss parameters for the first main network; The first master network is adjusted based on the target loss parameters and the losses of each of the supervision signals.

4. The method according to claim 3, wherein the supervision signal loss includes supervision loss and sparsity loss, and the target loss parameter is the supervision signal weight and signal sparsity intensity for the supervision loss. The step of adjusting the first master network based on the target loss parameter and the losses of each of the supervision signals includes: The supervision loss, the sparse loss, the supervision signal weight, and the signal sparsity intensity are input into the first loss calculation formula to determine the first loss; Based on the first loss, adjust the first main network; The first loss calculation formula satisfies the following formula: Among them, Loss A The first loss is denoted by I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal, α i For the supervision signal weights of the i-th image recognition supervision signal, parameter i β is the sparse loss corresponding to the i-th image recognition supervision signal. i Let be the signal sparsity intensity for the i-th image recognition supervision signal.

5. The method according to claim 2, wherein the step of training the second meta-network based on the identification results of each of the supervision signals to determine the loss adjustment parameters includes: The identification results of each of the aforementioned supervisory signals are input into the second meta-network for at least one round of meta-network training; In each round of meta-network training, loss adjustment parameters for each image recognition supervision signal are output, and the supervision signal loss corresponding to each image recognition supervision signal is obtained based on the first main network. Meta-network adjustment is performed on the second meta-network based on each supervision signal loss.

6. The method according to claim 5, wherein adjusting the second meta-network based on each of the supervision signal losses comprises: The losses of each of the aforementioned monitoring signals are input into the second loss calculation formula to determine the second loss; Based on the second loss, perform meta-network adjustments on the second meta-network; The second loss calculation formula satisfies the following formula: Among them, Loss B The second loss is I, where I is the total number of image recognition supervision signals, and i is an integer. i (x) represents the supervision signal loss corresponding to the i-th image recognition supervision signal.

7. The method according to claim 1, wherein adjusting the first master network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model comprises: Based on the loss adjustment parameters, the first loss calculation formula of the first main network is updated to obtain the updated first loss calculation formula. During the training of the first main network, the first main network is adjusted based on the first loss calculation formula until a target image recognition model is obtained for the initial image recognition model.

8. The method according to claim 7, wherein updating the first loss calculation formula of the first main network based on the loss adjustment parameter to obtain the updated first loss calculation formula comprises: Obtain the target loss parameter in the first loss calculation formula for the first main network; The target loss parameter is updated based on the loss adjustment parameter to obtain the updated target loss parameter.

9. The method according to claim 1, further comprising: The model training method for the initial image recognition model is determined to be the network alternation training method. Based on the network alternation training method, the first training round number for the first main network and the second training round number for the second meta-network are determined. The process of training the first master network using image sample data to determine at least one supervisory signal recognition result, and training the second meta-network based on each supervisory signal recognition result to determine loss adjustment parameters includes: Based on the first training round, the first main network is trained using image sample data to determine at least one supervisory signal recognition result. Based on the second training round number and the identification results of each supervision signal, the second meta-network is trained to determine the loss adjustment parameters.

10. The method according to claim 9, wherein adjusting the first master network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model comprises: If the initial image recognition model does not meet the model training termination condition, then the first loss calculation formula of the first main network is updated based on the loss adjustment parameter to obtain the updated first loss calculation formula. During the next round of training of the first main network, the first main network is adjusted based on the first loss calculation formula until the initial image recognition model meets the model training termination condition, thereby obtaining the target image recognition model for the initial image recognition model.

11. The method according to claim 10, wherein obtaining the target image recognition model for the initial image recognition model comprises: The initial image recognition model is used as the target image recognition model; or, The first master network in the initial image recognition model is used as the target image recognition model.

12. The method according to any one of claims 1-11, wherein the image recognition task is a liveness detection task, and the supervision signal recognition result includes at least one of depth estimation recognition result, image material classification recognition result, reflectance spectrum prediction recognition result, and liveness classification recognition result.

13. The method according to claim 12, further comprising: Obtain the target detection image corresponding to the liveness detection task; The target detection image is input into the target image recognition model, and at least one target supervised recognition result is output for the target detection image; Based on the target supervision and recognition results, the image detection type corresponding to the target detection image is determined, and the image detection type includes liveness image type and attack image type.

14. The method according to claim 13, wherein determining the image detection type corresponding to the target detection image based on each of the target supervision recognition results includes: Determine the liveness monitoring signal segmentation rules corresponding to the monitoring and identification results of each target; The liveness detection is performed on each of the target supervision recognitions using the liveness supervision signal segmentation rules, and the image detection type corresponding to the target detection image is determined based on the liveness detection results.

15. An image recognition device, the device comprising: The model building module is used to build an initial image recognition model based on the image recognition task. The image recognition model includes a first master network and a second meta-network. The model training module is used to train the first main network using image sample data and determine at least one supervision signal recognition result. The step of training the first main network using image sample data to determine at least one supervisory signal recognition result includes: configuring at least one image recognition supervisory signal for the first main network based on the image recognition task, inputting image sample data into the first main network for main network training, and outputting at least one supervisory signal recognition result indicated by the image recognition supervisory signal. The model training module is used to train the second meta-network based on the recognition results of each supervision signal and to determine the loss adjustment parameters. The model training module is used to adjust the first main network based on the loss adjustment parameters to obtain a target image recognition model for the initial image recognition model.

16. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 14.

17. A computer program product storing at least one instruction, said at least one instruction being loaded by a processor and executing the method steps of any one of claims 1 to 14.

18. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 14.