Payment willingness identification method and device, electronic equipment, medium and program product
By obtaining the target face image and generating a mask image in face-swiping payment, combining gaze and body behavior characteristics, and using the payment willingness recognition model to identify the user's willingness, the problems of fraudulent and incorrect payment in face-swiping payment are solved, and security and user experience are improved.
Patent Information
- Application Number
- CN202210960359.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-08-11
AI Technical Summary
In the face-scanning payment scenario, there are problems of stolen and mistaken swipes, which reduces the security of face-scanning payment.
By obtaining the target face image and generating a target mask map, combined with gaze features and body behavior features, the willingness to pay recognition model is used for identification to determine the target face-scanning user's willingness to pay.
It improves the accuracy of payment intention recognition, prevents fraudulent and accidental swipes, and enhances the security and user experience of face-scanning payment.
Smart Images

Figure CN115375292B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a willingness-to-pay identification method, device, electronic device, medium, and program product. Background Art
[0002] Face payment is a new payment method based on artificial intelligence, machine vision, 3D sensing, big data and other technologies. It brings great convenience to users by using facial recognition as identity authentication.
[0003] Currently, in face-scanning payment scenarios, after the user activates face-scanning payment, they must stand in front of a device with face-scanning payment functionality for facial recognition. However, when using offline face-scanning devices for face-scanning payment, on the one hand, there is the possibility that unscrupulous individuals may use the face-scanning device to steal someone's face without their knowledge. On the other hand, there may be multiple users standing in front of the face-scanning device. That is, when the face image captured by the face-scanning device shows multiple users, it is possible that user A activates face-scanning but mistakenly scans user B, which can easily lead to public opinion about the security of face-scanning payment.
[0004] Based on this, face-scanning payment intention recognition is an important part of face-scanning security in the payment system, which helps to improve the face-scanning security experience. However, the above-mentioned theft and mistaken swipe situations will reduce the security of face-scanning payment. Therefore, a more secure payment intention recognition solution is needed for face-scanning payment. Summary of the Invention
[0005] The embodiments of this specification provide a method, device, electronic device, medium, and program product for identifying willingness to pay. These methods use the corresponding features of the target user's gaze and body movements in the target face-scanning image to identify the target user's willingness to pay. This improves the accuracy of willingness to pay identification, resolves the issues of fraudulent or misplaced payments, ensures the security of face-scanning payments, and enhances the user's experience of face-scanning payments. The technical solution is as follows:
[0006] In a first aspect, embodiments of this specification provide a method for identifying willingness to pay, including:
[0007] Obtain a target face recognition image; the target face recognition image includes the target face recognition user;
[0008] Generate a corresponding target mask image based on the target position of the target face-scanning user in the target face-scanning image; the target mask image is used to distinguish the facial area of the target face-scanning user from other areas except the facial area;
[0009] The target face-scanning image and the target mask image are input into a willingness-to-pay recognition model to obtain gaze features and body behavior features, and the recognition result corresponding to the target face-scanning user is output based on the gaze features and the body behavior features; the willingness-to-pay recognition model is trained based on face-scanning images of multiple face-scanning users with known willingness-to-pay information.
[0010] In one possible implementation, after obtaining the target face image, inputting the target face image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and before outputting the recognition result corresponding to the target face user based on the gaze features and the body behavior features, the method further includes:
[0011] Determining a target body region image of the target face-scanning user based on the face region of the target face-scanning user and the target face-scanning image; the target body region image includes the limbs of the target face-scanning user;
[0012] The target face image and the target mask image are input into the willingness to pay recognition model to obtain gaze features and body behavior features, and the recognition result corresponding to the target face user is output based on the gaze features and body behavior features, including:
[0013] The target human body area image and the target mask image are input into the willingness-to-pay recognition model to obtain gaze features and limb behavior features, and the recognition result corresponding to the target face-scanning user is output based on the gaze features and the limb behavior features.
[0014] In one possible implementation, the target human body region image and the target mask image are input into a willingness-to-pay recognition model to obtain gaze features and body behavior features, and based on the gaze features and body behavior features, an identification result corresponding to the target face-scanning user is output, including:
[0015] Extracting a first feature corresponding to the target human body region image;
[0016] Fusing the first feature with the target mask image to generate a second feature;
[0017] Determining a limb behavior feature corresponding to the target human body region image based on the first feature;
[0018] Determining a gaze feature corresponding to the target human body region image based on the second feature;
[0019] The recognition result corresponding to the target face scanning user is determined based on the gaze characteristics and the body behavior characteristics.
[0020] In one possible implementation, the above identification result includes a willingness-to-pay identification result;
[0021] After inputting the target face-scanning image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and outputting a recognition result corresponding to the target face-scanning user based on the gaze features and body behavior features, the method further includes:
[0022] Based on the above recognition results, it is determined whether the above target face-scanning user has the willingness to pay.
[0023] In a possible implementation, the above recognition result also includes a gaze recognition result and a payment behavior recognition result;
[0024] The above-mentioned determination of whether the target face-scanning user has a willingness to pay based on the above-mentioned recognition result includes:
[0025] If the payment willingness recognition result does not meet the preset conditions, whether the target face-scanning user has the payment willingness is determined based on the gaze recognition result and / or the payment behavior recognition result.
[0026] In a possible implementation, the target face-scanning image includes multiple face-scanning users, and the multiple face-scanning users include the target face-scanning user;
[0027] After obtaining the target face-scanning image, and before generating a corresponding target mask image based on the target position of the target face-scanning user in the target face-scanning image, the method further includes:
[0028] The target face-scanning user is determined from the multiple face-scanning users according to preset rules.
[0029] In one possible implementation, the recognition result includes a willingness-to-pay recognition result; the willingness-to-pay recognition model includes a first target convolutional network, a second target convolutional network, and a first fully connected layer;
[0030] The first target convolutional network is used to process the target human region image and the target mask image to obtain a gaze feature corresponding to the target human region image;
[0031] The second target convolutional network is used to process the target human body region image to obtain limb behavior features corresponding to the target human body region image;
[0032] The first fully connected layer is used to fuse the gaze features and the body behavior features to output the payment willingness recognition result corresponding to the target face-scanning user.
[0033] In a possible implementation, the first target convolutional network includes a first convolutional module, a second convolutional module, and a third convolutional module;
[0034] The first convolution module is used to extract the first feature corresponding to the target human body region image;
[0035] The second convolution module is used to fuse the first feature with the target mask image to generate a second feature;
[0036] The third convolution module is used to generate a gaze feature corresponding to the target human body area image based on the second feature.
[0037] In a possible implementation, the second target convolutional network includes a first convolutional module and a fourth convolutional module;
[0038] The first convolution module is used to extract the first feature corresponding to the target human body region image;
[0039] The fourth convolution module is used to generate a limb behavior feature corresponding to the target human body region image based on the first feature.
[0040] In a possible implementation, the recognition result further includes a gaze recognition result and a payment behavior recognition result; the willingness to pay recognition model further includes a second fully connected layer and a third connected layer;
[0041] The second connection layer is configured to output the gaze recognition result corresponding to the target face scanning user based on the gaze feature;
[0042] The third connection layer is used to output the payment behavior recognition result corresponding to the target face-scanning user based on the gaze feature.
[0043] In one possible implementation, the first target convolutional network is trained based on human body region images and mask images corresponding to the plurality of face-scanning users, each of which has known willingness-to-pay information and / or gaze information.
[0044] The second target convolutional network is trained based on human body area images corresponding to the multiple face-scanning users with known willingness to pay information and / or payment behavior information.
[0045] In a possible implementation, the target human region image and the target mask image have the same resolution.
[0046] In a second aspect, an embodiment of this specification provides a willingness-to-pay identification device, the device comprising:
[0047] An acquisition module is used to acquire a target face recognition image; the target face recognition image includes the target face recognition user;
[0048] A generating module, configured to generate a corresponding target mask image based on the target position of the target face-scanning user in the target face-scanning image; the target mask image is used to distinguish the facial area of the target face-scanning user from other areas except the facial area;
[0049] The willingness to pay recognition module is used to input the above-mentioned target face-scanning image and the above-mentioned target mask image into the willingness to pay recognition model, obtain gaze features and body behavior features, and output the recognition result corresponding to the above-mentioned target face-scanning user based on the above-mentioned gaze features and the above-mentioned body behavior features; the above-mentioned willingness to pay recognition model is trained based on the face-scanning images of multiple face-scanning users with known willingness to pay information.
[0050] In one possible implementation, the willingness-to-pay identification device further includes:
[0051] A first determining module is configured to determine a target body region image of the target face-scanning user based on the face region of the target face-scanning user and the target face-scanning image; the target body region image includes the limbs of the target face-scanning user;
[0052] The above-mentioned willingness to pay recognition module is specifically used to: input the above-mentioned target human body area image and the above-mentioned target mask image into the willingness to pay recognition model, obtain gaze features and limb behavior features, and output the recognition result corresponding to the above-mentioned target face-scanning user based on the above-mentioned gaze features and the above-mentioned limb behavior features.
[0053] In one possible implementation, the willingness-to-pay identification module includes:
[0054] an extraction unit, configured to extract a first feature corresponding to the target human body region image;
[0055] A fusion unit, configured to fuse the first feature with the target mask image to generate a second feature;
[0056] A first determining unit is configured to determine a limb behavior feature corresponding to the target human body region image based on the first feature;
[0057] A second determining unit is configured to determine a gaze feature corresponding to the target human body region image based on the second feature;
[0058] The third determination unit is used to determine the recognition result corresponding to the target face-scanning user based on the gaze characteristics and the body behavior characteristics.
[0059] In one possible implementation, the above identification result includes a willingness-to-pay identification result;
[0060] The willingness to pay identification device further includes:
[0061] The second determination module is used to determine whether the target face-scanning user has the willingness to pay based on the above recognition result.
[0062] In a possible implementation, the above recognition result also includes a gaze recognition result and a payment behavior recognition result;
[0063] The second determining module is specifically configured to:
[0064] If the payment willingness recognition result does not meet the preset conditions, whether the target face-scanning user has the payment willingness is determined based on the gaze recognition result and / or the payment behavior recognition result.
[0065] In a possible implementation, the target face-scanning image includes multiple face-scanning users, and the multiple face-scanning users include the target face-scanning user;
[0066] The willingness to pay identification device further includes:
[0067] The third determination module is used to determine the target face-scanning user from the multiple face-scanning users according to preset rules.
[0068] In one possible implementation, the recognition result includes a willingness-to-pay recognition result; the willingness-to-pay recognition model includes a first target convolutional network, a second target convolutional network, and a first fully connected layer;
[0069] The first target convolutional network is used to process the target human region image and the target mask image to obtain a gaze feature corresponding to the target human region image;
[0070] The second target convolutional network is used to process the target human body region image to obtain limb behavior features corresponding to the target human body region image;
[0071] The first fully connected layer is used to fuse the gaze features and the body behavior features to output the payment willingness recognition result corresponding to the target face-scanning user.
[0072] In a possible implementation, the first target convolutional network includes a first convolutional module, a second convolutional module, and a third convolutional module;
[0073] The first convolution module is used to extract the first feature corresponding to the target human body region image;
[0074] The second convolution module is used to fuse the first feature with the target mask image to generate a second feature;
[0075] The third convolution module is used to generate a gaze feature corresponding to the target human body area image based on the second feature.
[0076] In a possible implementation, the second target convolutional network includes a first convolutional module and a fourth convolutional module;
[0077] The first convolution module is used to extract the first feature corresponding to the target human body region image;
[0078] The fourth convolution module is used to generate a limb behavior feature corresponding to the target human body region image based on the first feature.
[0079] In a possible implementation, the recognition result further includes a gaze recognition result and a payment behavior recognition result; the willingness to pay recognition model further includes a second fully connected layer and a third connected layer;
[0080] The second connection layer is configured to output the gaze recognition result corresponding to the target face scanning user based on the gaze feature;
[0081] The third connection layer is used to output the payment behavior recognition result corresponding to the target face-scanning user based on the gaze feature.
[0082] In one possible implementation, the first target convolutional network is trained based on human body region images and mask images corresponding to the plurality of face-scanning users, each of which has known willingness-to-pay information and / or gaze information.
[0083] The second target convolutional network is trained based on human body area images corresponding to the multiple face-scanning users with known willingness to pay information and / or payment behavior information.
[0084] In a possible implementation, the target human region image and the target mask image have the same resolution.
[0085] In a third aspect, an embodiment of this specification provides an electronic device, including: a processor and a memory;
[0086] The processor is connected to the memory;
[0087] The aforementioned memory is used to store executable program code;
[0088] The processor runs the program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method provided by the first aspect of the embodiment of this specification or any possible implementation of the first aspect.
[0089] In a fourth aspect, an embodiment of this specification provides a computer storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor and executing the method provided by the first aspect of the embodiment of this specification or any possible implementation of the first aspect.
[0090] In a fifth aspect, an embodiment of this specification provides a computer program product comprising instructions, which, when the above-mentioned computer program product runs on a computer or a processor, enables the above-mentioned computer or the above-mentioned processor to execute the payment willingness identification method provided by the first aspect of the embodiment of this specification or any possible implementation method of the first aspect.
[0091] The embodiment of this specification obtains a target face-scanning image, which includes a target face-scanning user, and generates a corresponding target mask map based on the target position of the target face-scanning user in the target face-scanning image. The target mask map is used to distinguish the facial area of the target face-scanning user from other areas except the facial area. The target face-scanning image and the target mask map are then input into a willingness-to-pay recognition model to obtain gaze features and limb behavior features, and the recognition result corresponding to the target face-scanning user is output based on the gaze features and the limb behavior features. The willingness-to-pay recognition model is trained based on face-scanning images of multiple face-scanning users with known willingness-to-pay information. The embodiments of this specification adopt an end-to-end learning method of the willingness to pay recognition model. It not only determines the corresponding recognition result based on the gaze features of the target face-scanning user in the target face-scanning image, but also combines the gaze features and body behavior features of the target face-scanning user to jointly determine the recognition result corresponding to the target face-scanning user. Compared with only using the gaze features of the target face-scanning user in the target face-scanning image to determine the corresponding recognition result, the method of combining the gaze features and body behavior features of the target face-scanning user to jointly determine the recognition result corresponding to the target face-scanning user can improve the accuracy of the willingness to pay recognition. Therefore, in the face-scanning payment scenario, the more accurate willingness to pay can be identified, which solves the problem of stealing or mistakenly swiping the faces of other non-face-scanning users when using face-scanning devices for face-scanning payment in offline public places, ensures the security of face-scanning payment, and enhances the user's security experience of face-scanning payment. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0093] Figure 1A schematic diagram of the architecture of a willingness-to-pay identification system provided in an exemplary embodiment of this specification;
[0094] Figure 2A-2B A schematic diagram of an application scenario of willingness to pay identification provided by an exemplary embodiment of this specification;
[0095] Figure 3 A flowchart of a willingness-to-pay identification method provided in an exemplary embodiment of this specification;
[0096] Figure 4 A flowchart of another willingness-to-pay identification method provided as an exemplary embodiment of this specification;
[0097] Figure 5 A schematic diagram of an implementation process of determining a target human body region and a target mask image provided by an exemplary embodiment of this specification;
[0098] Figure 6 A schematic diagram of a process for implementing willingness to pay identification provided in an exemplary embodiment of this specification;
[0099] Figure 7 A schematic diagram of an implementation process of willingness-to-pay identification provided in an exemplary embodiment of this specification;
[0100] Figure 8 A schematic diagram of another implementation process of willingness-to-pay identification provided by an exemplary embodiment of this specification;
[0101] Figure 9 A schematic diagram of an implementation process for determining whether a target face-scanning user has a willingness to pay based on a recognition result, provided as an exemplary embodiment of this specification;
[0102] Figure 10 A schematic diagram of a willingness-to-pay identification model provided in an exemplary embodiment of this specification;
[0103] Figure 11 A schematic diagram of a willingness-to-pay identification device provided as an exemplary embodiment of this specification;
[0104] Figure 12 The present invention provides a structural diagram of an electronic device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0105] The technical solutions in the embodiments of this specification will be described clearly and completely below in conjunction with the drawings in the embodiments of this specification.
[0106] In this specification, claims, and the accompanying drawings, the terms "first," "second," "third," and so on are used to distinguish between different items, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0107] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the architecture of a willingness to pay identification system provided by an exemplary embodiment of this specification. Figure 1 As shown, the willingness to pay identification system may include: a face recognition device 110 and a server 120. Among them:
[0108] The face recognition device 110 can be a mobile phone, tablet computer, laptop computer or other device installed with user version software and a camera, or it can be other devices such as an Internet of Things (IoT) face recognition machine with a camera and payment function, etc. The embodiments of this specification are not limited to this.
[0109] Optionally, when the face-scanning device 110 collects the target face-scanning image, in order to avoid the situation where the payment is stolen or the face of another user is mistakenly swiped for payment, a corresponding target mask map can be generated based on the target position of the target face-scanning user appearing in the target face-scanning image. The target mask map is used to distinguish the facial area of the target face-scanning user from other areas except the facial area. The target face-scanning image and the target mask map are then input into the willingness to pay recognition model to obtain gaze features and body behavior features, and the recognition result corresponding to the target face-scanning user is output based on the gaze features and body behavior features. The willingness to pay recognition model is trained based on face-scanning images of known willingness to pay information corresponding to multiple face-scanning users. Finally, it can be learned from the recognition result corresponding to the target face-scanning user whether the target face-scanning user has the willingness to pay, thereby deciding whether the target face-scanning user will make the payment, and preventing the target face-scanning user's face from being stolen or mistakenly swiped by others.
[0110] It can be understood that after the face-scanning device 110 collects the target face-scanning image, the face-scanning device 110 directly performs payment willingness recognition based on the target face-scanning image without the need for additional transmission of the target face-scanning image. That is, when the face-scanning device 110 performs payment willingness recognition alone, it can avoid network limitations and ensure the efficiency and feasibility of payment willingness recognition.
[0111] Optionally, when the face-scanning device 110 collects the target face-scanning image, in order to avoid the situation where the payment is stolen or the face of another user is mistakenly swiped for payment, a data connection relationship can be established with the server 120 through the network, such as sending the target face-scanning image to the server 120, and then receiving the recognition result of the target face-scanning user in the target face-scanning image determined by the server 120 based on the target face-scanning image, and finally, according to the recognition result corresponding to the target face-scanning user, whether the target face-scanning user has the willingness to pay is determined, thereby deciding whether the target face-scanning user should make the payment, and preventing the target face-scanning user's face from being stolen or mistakenly swiped by others.
[0112] It can be understood that the server 120 is a high-performance computer with strong data processing capabilities and high stability and reliability. Therefore, when the face-scanning device 110 collects the target face-scanning image, compared with the method of directly performing payment willingness recognition based on the target face-scanning image through the face-scanning device 110, the face-scanning device 110 sends the collected target face-scanning image to the server 120 through the network for payment willingness recognition. That is, the method of jointly performing payment willingness recognition by the face-scanning device 110 and the server 120 can avoid problems such as inaccurate or slow payment willingness recognition due to the low configuration of the face-scanning device 110, and to a certain extent can also ensure the stability and accuracy of payment willingness recognition, providing stronger security protection for face-scanning payment.
[0113] Optionally, when it is learned from the recognition result corresponding to the target face-scanning user that the target face-scanning user is willing to pay, the face-scanning device 110 can determine that the target face-scanning user will make the payment, and after the face-scanning payment is completed, send the corresponding payment information to the terminal corresponding to the target face-scanning user through the network.
[0114] The server 120 can be a server that can provide multiple payment willingness identifications, and can receive data such as the target face-scanning image sent by the face-scanning device 110 through the network. The target face-scanning image includes the target face-scanning user, and generates a corresponding target mask map based on the target position of the target face-scanning user in the target face-scanning image. The target mask map is used to distinguish the facial area of the target face-scanning user from other areas except the facial area. The target face-scanning image and the target mask map are input into the payment willingness identification model to obtain gaze features and body behavior features, and the identification results corresponding to the target face-scanning user are output based on the gaze features and body behavior features. The payment willingness identification model is trained based on the face-scanning images of multiple face-scanning users with known payment willingness information.
[0115] Specifically, after the server 120 has identified the target face-scanning user's willingness to pay in the target face-scanning image, it can also send the target face-scanning user's corresponding recognition result to the face-scanning device 110 via the network, so that the face-scanning device 110 can determine whether the target face-scanning user should pay based on the recognition result. When the server 120 determines that the target face-scanning user has the willingness to pay, it can also send corresponding payment information or payment prompt information to the terminal corresponding to the target face-scanning user to prompt the target face-scanning user that the face-scanning payment has been made or to confirm whether the face-scanning payment is in progress.
[0116] Specifically, the server 120 may be, but is not limited to, a hardware server, a virtual server, a cloud server, etc.
[0117] The network can be a medium that provides a communication link between the server 120 and any of the face recognition devices 110, or can be the Internet including network devices and transmission media, but is not limited thereto. The transmission medium can be a wired link (such as, but not limited to, coaxial cable, optical fiber, and digital subscriber line (DSL)) or a wireless link (such as, but not limited to, wireless fidelity (WIFI), Bluetooth, and mobile device networks).
[0118] Understandably, Figure 1 The number of face-scanning devices 110 and servers 120 shown in the willingness-to-pay identification system is for illustrative purposes only. In a specific implementation, the willingness-to-pay identification system may include any number of face-scanning devices and servers. This specification does not impose any specific limitations on this. For example, but not limited to, the face-scanning device 110 may be a face-scanning device cluster consisting of multiple face-scanning devices, and the server 120 may be a server cluster consisting of multiple servers.
[0119] Please refer to Figure 2A-2B , Figure 2A-2B This is a schematic diagram of an application scenario of a willingness-to-pay identification method provided in an exemplary embodiment of this specification. Figure 1 The face recognition device 110 may be Figure 2A and Figure 2B After the user 220 has selected the goods to be purchased on the self-service payment machine 210, he can click Figure 2A The "Face Payment" control 212 displayed on the screen of the self-service payment machine 210 is used to pay by face. When the "Face Payment" control 212 is triggered, Figure 2BAs shown, the self-service payment machine 210 begins to collect a face image 230 through its installed camera 211, and performs facial recognition and payment intention recognition based on the collected face image 230, thereby determining the identity of the user using face payment and whether the user has the payment intention. However, there may often be multiple users standing in front of the self-service payment machine 210. That is, when multiple users appear in the face image collected by the self-service payment machine 210, it is possible that user 220 activates face payment and mistakenly swipes the face of another user near user 220 to make a payment, or user 220 steals the face of the user behind him while the user behind him is not paying attention. This can easily lead to public opinion about the security of face payment and make the security of face payment less effectively guaranteed.
[0120] In order to solve the above problems, we will combine Figures 1-2B , introduces the payment willingness identification method provided by the embodiment of this specification. For details, please refer to Figure 3 , which is a flow chart of a method for identifying willingness to pay provided by an exemplary embodiment of this specification. Figure 3 As shown in FIG, the willingness-to-pay identification method includes the following steps:
[0121] S302, obtaining a target face scan image.
[0122] Specifically, when a user triggers the face-scanning device 110 to perform face-scanning payment, the target face-scanning image can be collected through the camera installed on the face-scanning device 110, and the target face-scanning image includes the target face-scanning user. At this time, if the server 120 performs payment willingness recognition, the face-scanning device 110 also needs to send the collected target face-scanning image to the server 120 through the network. The server 120 can obtain the target face-scanning image sent by the face-scanning device 110 through the network, and thus perform payment willingness recognition based on the target face-scanning image. If the face-scanning device 110 performs payment willingness recognition, then after the face-scanning device 110 collects the target face-scanning image through the installed camera, it can directly start payment willingness recognition based on the target face-scanning image. The following embodiments are all explained by taking the face-scanning device 110 performing payment willingness recognition as an example.
[0123] It can be understood that when there are multiple users standing in front of the camera of the face-scanning device 110, the collected target face-scanning image may also include multiple users, and the target face-scanning user among the multiple users can be understood as the user who is most likely to perform facial recognition in the target face-scanning image, such as but not limited to the user who is closest to the face-scanning device 110 or occupies the largest area or is most centrally located in the target face-scanning image.
[0124] It is understandable that the target face-scanning user in the target face-scanning image may be the user who triggers the face-scanning device 110 or really needs to make a face-scanning payment, or may be other users standing in front of the camera of the face-scanning device 110 who have no intention to pay, etc. In order to prevent the target face-scanning user from being mistakenly swiped or stolen because he or she appears in front of the camera of the face-scanning device 110 without the intention to pay, when making a face-scanning payment, it is necessary to first determine whether the target face-scanning user has the intention to pay, and then decide whether to swipe the target face-scanning user's face to make the payment.
[0125] S304: Generate a corresponding target mask image based on the target position of the target face-scanning user in the target face-scanning image.
[0126] Specifically, after collecting the target face-scanning image, the face detection algorithm can be used to first detect the area where the target face-scanning user's face is located in the target face-scanning image, determine the target position of the target face-scanning user, and then generate a target mask map corresponding to the target face-scanning user based on the above target position, thereby distinguishing the target face-scanning user's facial area and other areas except the facial area, making the target face-scanning user's facial feature information more distinct.
[0127] For example, after collecting the target face-scanning image and determining the target position of the target face-scanning user, the filling value of the facial area of the target face-scanning user in the target face-scanning image can be set as 1, and the filling value of other areas can be set as 0 (other different filling values with higher discrimination can also be used), thereby generating a target mask map that can distinguish the facial area of the target face-scanning user and other areas except the facial area.
[0128] It can be understood that the facial area of the target face-scanning user represented by the above-mentioned target position can be a regular area such as a rectangle or a circle, or it can be any irregular area, and the embodiments of this specification do not limit this.
[0129] S306: Input the target face-scanning image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and output the recognition result corresponding to the target face-scanning user based on the gaze features and body behavior features.
[0130] Specifically, the willingness-to-pay recognition model is trained based on facial recognition images of multiple users with known willingness-to-pay information. This willingness-to-pay information includes a willingness-to-pay label corresponding to when the user has a willingness to pay and a non-willing-to-pay label corresponding to when the user has no willingness to pay.
[0131] Specifically, in order to enable the willingness to pay recognition model to more accurately obtain the gaze features of the target face-scanning image through the target mask map, the resolution and size of the target face-scanning image and the target mask map input into the willingness to pay recognition model should be consistent, so as to make it easier for the willingness to pay recognition model to obtain the corresponding gaze features from the facial area of the target face-scanning user in the target face-scanning image.
[0132] Specifically, the recognition result corresponding to the target face-scanning user may include the recognition result of the target face-scanning user's willingness to pay. The above-mentioned willingness to pay recognition result includes the probability of the target face-scanning user having willingness to pay and / or the probability of not having willingness to pay as identified by the willingness to pay recognition model.
[0133] Specifically, in the offline face-scanning payment scenario, when the user wants to make a face-scanning payment, he or she not only needs to look at the camera or screen of the face-scanning device 110, but also needs to click the face-scanning device 110 with his or her hand to trigger the face-scanning payment action, that is, the payment behavior. Therefore, the method of simply identifying the payment intention by the facial area of the target face-scanning user, that is, whether the target face-scanning user looks at the camera or screen of the face-scanning device 110, cannot effectively prevent theft or mistaken payment. The embodiment of this specification combines the gaze characteristics of the target face-scanning user (whether he or she looks at the camera or screen of the face-scanning device 110) and the body behavior characteristics (whether there is a payment behavior, such as but not limited to raising the hand to click The method of jointly determining the recognition result corresponding to the target face-scanning user by combining the gaze features and body behavior features of the target face-scanning user can improve the accuracy of recognition of willingness to pay. Therefore, in the face-scanning payment scenario, the more accurate willingness to pay can be identified, which solves the problem of stealing or mistakenly swiping the faces of other non-face-scanning users when using face-scanning devices for face-scanning payment in offline public places, ensures the security of face-scanning payment, and enhances the user's safety experience of face-scanning payment.
[0134] In an actual offline face-scanning payment scenario, the target face-scanning image collected by the face-scanning device 110 is likely to include multiple face-scanning users, and the target face-scanning image includes not only the face area of the target face-scanning user, but also the limb area of the target face-scanning user, such as upper limbs or lower limbs, etc., and may also include other objects that do not need to be identified, such as white clouds, tables and chairs in the environment where the target face-scanning user is located, and these additional objects will have a certain impact on the recognition of the payment willingness of the target face-scanning user in the target face-scanning image. At this time, in order to more accurately realize the recognition of payment willingness, better prevent theft or misuse of the brush, and further improve the security of offline face-scanning payment, the following is combined with Figure 4, introduces another method for identifying willingness to pay provided by the embodiments of this specification. Figure 4 As shown in FIG, the willingness-to-pay identification method includes the following steps:
[0135] S402, obtaining a target face scan image.
[0136] Specifically, the target face-scanning image may include multiple face-scanning users and / or objects that do not need to be identified other than the target face-scanning user, such as white clouds, tables and chairs in the environment where the target face-scanning user is located. This specification does not limit this.
[0137] Optionally, when the target face-scanning image includes multiple face-scanning users, after obtaining the target face-scanning image, the target face-scanning user may be determined from the multiple face-scanning users according to a preset rule. The preset rule may be to select the face-scanning user closest to the face-scanning device, or the face-scanning user with the largest area or the most central position in the target face-scanning image as the target face-scanning user. The target face-scanning user may also be determined by other methods, which are not limited in the embodiments of this specification.
[0138] S404: Generate a corresponding target mask image based on the target position of the target face-scanning user in the target face-scanning image.
[0139] Specifically, after acquiring the target face-scanning image, a face detection algorithm can be used to first detect the area where the target face-scanning user's face is located in the target face-scanning image, thereby determining the target position where the target face-scanning user's face area is located, and then generating a mask map corresponding to the target face-scanning user based on the above target position, thereby being able to distinguish the target face-scanning user's face area from other areas other than the face area in the target face-scanning image. Since the target face-scanning image may include multiple face-scanning users and / or objects other than the target face-scanning user that do not need to be identified, when identifying the target face-scanning user's willingness to pay, other areas other than the area where the target face-scanning user is located in the target face-scanning image may not need to be identified, that is, the target position where the target face-scanning user's face area is located in the mask map corresponding to the target face-scanning user can be used to determine the target human body area corresponding to the target face-scanning user in the mask map, and then crop it to obtain the target mask map.
[0140] It can be understood that when using an offline Internet of Things (IoT) face-scanning machine (face-scanning device) with face-scanning function set up in public consumption scenarios such as supermarkets / convenience stores / catering / hotels / campus education / medical care to make face-scanning payments, the collected target face-scanning image will not only include the face of the face-scanning user, but also their limbs. That is, in addition to the facial area of the target face-scanning user, the target human body area can also include the area where the target face-scanning user's limbs are located.
[0141] For example, Figure 5 As shown in the left figure, when the target face recognition user's facial area in the mask image corresponding to the target face recognition user is the inscribed circle area of the matrix frame composed of point A as the upper left vertex and point B as the lower right vertex, if the coordinates of point A are (x1, y1), the coordinates of point B are (x2, y2), and the radius of the facial area is R, the position coordinates of points A and B can be used to calculate the location of the target human body area. Figure 5 As shown in the left figure, the target human body area is a matrix area with point C as the upper left vertex and point D as the lower right vertex, where the coordinates of point C are (x3, y3), the coordinates of point D are (x4, y4), x3 = x1-2R, x4 = x2+2R, y3 = y1-R, y4 = y2+4R, that is, the rectangular box corresponding to the target user's face area is expanded in the four directions of up, down, left and right by a certain length to obtain the target human body area, and then the target human body area is cropped according to its position in the mask image to obtain Figure 5 (Right) The target mask shown.
[0142] It is understandable that the method of determining the target human body area of the target face-scanning user in the mask image of the target face-scanning user corresponding to the target face-scanning image is not limited to the above method. Figure 5 The method shown can also be determined directly through other methods such as key point detection, and the embodiments of this specification do not limit this.
[0143] It can be understood that the size of the target mask image in S304 is consistent with the size of the target face-scanning image, and the size of the target mask image in S404 is consistent with the size of the target human body area corresponding to the target face-scanning user in the target face-scanning image. Therefore, when multiple face-scanning users or other objects that do not need to be identified appear in the target face-scanning image, the willingness to pay recognition model can focus more on analyzing the corresponding features of the target face-scanning user, avoiding interference from other information, further improving the accuracy and efficiency of payment willingness recognition, and ensuring the security of offline face-scanning payment.
[0144] S406: Determine a target body region image of the target face-scanning user based on the face region of the target face-scanning user and the target face-scanning image.
[0145] Specifically, in order to avoid the influence of other information in the target face-scanning image other than the target face-scanning user on the recognition of the target face-scanning user's willingness to pay and to improve the efficiency and accuracy of the recognition of the willingness to pay, after the target face-scanning image is collected and the target face-scanning user for payment recognition is determined, the facial area of the target face-scanning user in the target face-scanning image can be expanded in all directions according to certain rules to obtain the target human body area of the target face-scanning user. Finally, the target human body area in the target face-scanning image is partially cropped and subjected to certain scaling transformations, etc., to obtain a target human body area image with the same resolution as the target mask image or the target face-scanning image. The above-mentioned target human body area image includes not only the face of the target face-scanning user, but also the limbs of the target face-scanning user.
[0146] S408: Input the target human body area image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and limb behavior features, and output the recognition result corresponding to the target face-scanning user based on the gaze features and limb behavior features.
[0147] Specifically, the willingness to pay recognition model is trained based on the human body area images corresponding to the face scan images of multiple face scan users with known willingness to pay information.
[0148] It can be understood that in order to better fuse the features in the target mask image with those in the target human body area image, obtain more accurate features that focus more on the face of the target face-scanning user, and thus determine the final gaze features, the resolution and size of the target human body area image and the target mask image should be the same.
[0149] Alternatively, as Figure 6 As shown, the implementation process of the willingness-to-pay recognition model outputting the recognition result corresponding to the target face-scanning user in S408 may include the following steps:
[0150] S602: Extract a first feature corresponding to the target human body region image.
[0151] Specifically, the target human body region image includes the face and limbs of the target face-scanning user, the limbs including at least the upper limbs, and the upper limbs may include the left hand and / or the right hand, although this specification does not limit this. The first feature corresponds to the basic feature corresponding to the target face-scanning user in the target human body region image.
[0152] S604: Fusing the first feature with the target mask image to generate a second feature.
[0153] Specifically, in order to be able to directly analyze (identify) whether the target face-scanning user is looking at the screen or camera of the face-scanning device based on the target human body area image, a first target mask image with the same resolution as the first feature can be generated by nearest neighbor sampling, and then the first feature and the first target mask image are connected together according to the channel dimension. Finally, the first target mask image is fused with the first feature corresponding to the target human body area image through a fusion convolutional network module to obtain a second feature that focuses more on the face of the target face-scanning user, avoiding the influence of features of other areas outside the face of the target face-scanning user on the analysis of whether the target face-scanning user is looking at the screen or camera of the face-scanning device, so that the gaze feature corresponding to the target human body area image determined based on the second feature is more accurate, so that in the process of payment willingness recognition, whether the target user is looking at the screen or camera of the face-scanning device can be more accurately analyzed (identified) to ensure the accuracy of the final recognition result.
[0154] It can be understood that the above-mentioned fused convolutional network module can be composed of one or more convolutional layers, and the embodiments of this specification are not limited to this.
[0155] S606: Determine a limb behavior feature corresponding to the target human body region image based on the first feature.
[0156] Specifically, the limb behavior features include features corresponding to the behavior of the upper limbs of the target face-scanning user in the target human body area image.
[0157] It can be understood that the behavior of the above-mentioned upper limbs can be a state in which the upper limbs (hands) are raised to show a clicking action when the target face-scanning user clicks the face-scanning device to trigger face-scanning payment, or the target face-scanning user does not click the face-scanning device at all, that is, the upper limbs of the target face-scanning user are not raised or are in other states that are not clicking actions, etc. The embodiments of this specification do not limit this.
[0158] S608: Determine a gaze feature corresponding to the target human body region image based on the second feature.
[0159] Specifically, the dimension of the second feature obtained after the fusion of the first feature and the target mask image is the same as the dimension of the first feature, thereby ensuring that the gaze feature corresponding to the target human body area image can be determined from the second feature that focuses more on the facial area of the target face-scanning user through one or more convolution modules of the same willingness to pay recognition model.
[0160] S610: Determine the recognition result corresponding to the target face scanning user based on the gaze characteristics and body behavior characteristics.
[0161] Specifically, after determining the gaze features and limb behavior features of the target face-scanning user in the target human area image, the above gaze features and limb behavior features can be connected in series to generate a fusion feature, and then the payment willingness of the target face-scanning user can be identified through the fully connected layer, that is, the identification result corresponding to the target face-scanning user can be obtained.
[0162] Alternatively, in order to improve the efficiency of willingness to pay identification, Figure 4 The willingness-to-pay identification process shown is Figure 7 The payment willingness recognition process shown in the figure determines the gaze features corresponding to the target face-scanning user directly on the basis of the target human body area image through the generated target mask map, and determines the limb behavior features corresponding to the target face-scanning user based on the target human body area image corresponding to the target face-scanning user in the target face-scanning image, and then determines the recognition result based on the gaze features and limb behavior features. This improves the accuracy of payment behavior recognition while reducing the amount of calculation in the recognition process and improving the efficiency of payment willingness recognition.
[0163] Optionally, in addition to following Figure 7 In addition to the willingness to pay identification process shown in the figure, you can also use Figure 8 The payment willingness recognition process shown in the figure is to directly capture the target facial image and target body area image of the target face-scanning user from the collected target face-scanning image, and then obtain the gaze features and body behavior features based on the target facial image and target body area image respectively, and finally determine the recognition result based on the gaze features and body behavior features. Figure 8 In the willingness-to-pay recognition process shown in FIG, since the target face image and the target body region image are both separate images, the corresponding gaze features can be directly extracted from the target face image through the pre-trained convolution module, and no additional target mask image is required to be fused with the features of the target body region image, i.e. Figure 8 The willingness-to-pay identification process shown is compared to Figure 7 The willingness-to-pay identification process shown in the figure is simpler and less difficult to implement.
[0164] Furthermore, the above-mentioned process of respectively obtaining gaze features and limb behavior features based on the target facial image and the target human body region image can be obtained by correspondingly using a gaze recognition model and a payment behavior recognition model.
[0165] It is understandable that the target facial image and target human body area image captured based on the target face-scanning image can also be directly input into the trained gaze recognition model and payment behavior recognition model respectively, so as to output the gaze recognition result and the payment behavior recognition result. Finally, the gaze recognition result and the payment behavior recognition result are used to jointly judge whether the target face-scanning user has the willingness to pay.
[0166] It can be understood that since it is necessary to understand whether the target face-scanning user is looking at the screen or camera of the face-scanning device from the facial area of the target face-scanning user in the target face-scanning image, the above-mentioned target face-scanning image can be directly replaced with the target eye image, and then the gaze direction or gaze state of the target face-scanning user is analyzed according to the target eye image, thereby realizing the gaze recognition of the target face-scanning user in the target face-scanning image.
[0167] Optionally, when there are multiple face-scanning users in the target face-scanning image, in order to maximize the effectiveness of face-scanning payment, it is possible to not only identify the payment willingness of the target face-scanning user among the multiple face-scanning users, but also identify the payment willingness of each face-scanning user among the multiple face-scanning users or face-scanning users that meet preset identification conditions, and then select the face-scanning user with the highest payment willingness for facial recognition and payment. The above-mentioned preset identification conditions may include, but are not limited to, the presence of a corresponding face-scanning user with a complete face in the target face-scanning image, the presence of a face-scanning user within a preset distance from the face-scanning device, or the presence of a face-scanning user with an area greater than a preset area in the target face-scanning image.
[0168] Optionally, the above recognition result includes a willingness-to-pay recognition result. After S306 or S408, i.e., after outputting the recognition result corresponding to the target face-scanning user, it is also possible to determine whether the target face-scanning user has a willingness to pay based on the above recognition result. The above willingness-to-pay recognition result includes the probability of having a willingness to pay and the probability of not having a willingness to pay (non-willingness to pay).
[0169] Optionally, determining whether the target face-scanning user has a willingness to pay based on the recognition result can be determining whether the target face-scanning user has a willingness to pay based on the willingness to pay recognition when the willingness to pay recognition result meets a preset condition. For example, but not limited to, when the probability of having a willingness to pay is greater than a first preset threshold and / or the probability corresponding to no willingness to pay is less than a second preset threshold, it is determined that the target face-scanning user has a willingness to pay; when the probability of having a willingness to pay is less than the first preset threshold and / or the probability corresponding to no willingness to pay is greater than the second preset threshold, it is determined that the target face-scanning user does not have a willingness to pay, etc. The above-mentioned first preset threshold can be 0.8, 0.9, etc., and the above-mentioned second preset threshold can be 0.1, 0.2, etc., which are values less than the first preset threshold, and the embodiments of this specification are not limited to this.
[0170] Optionally, in addition to the willingness to pay recognition result, the above-mentioned recognition results may also include gaze recognition results and payment behavior recognition results. The gaze recognition result includes the probability that the target face-scanning user is looking at the screen or camera of the face-scanning device and / or the probability that the target face-scanning user is not looking at the screen or camera of the face-scanning device (not looking at the screen or camera of the face-scanning device). The payment behavior recognition result includes the probability that the target face-scanning user has made a payment and / or the probability that the target face-scanning user has not made a payment (not making a payment). The above-mentioned determination of whether the target face-scanning user has a willingness to pay based on the recognition result may also be determined based on the gaze recognition result and / or the payment behavior recognition result when the willingness to pay recognition result does not meet a preset condition. For example, but not limited to, if the probability of having a willingness to pay in the willingness to pay recognition result is not within a preset range, if the gaze recognition result meets the preset gaze condition and / or the payment behavior recognition result meets the preset payment behavior condition, then the target face-scanning user is determined to have a willingness to pay; if the probability of having a willingness to pay in the willingness to pay recognition result is not within a preset range, if the gaze recognition result does not meet the preset gaze condition and the payment behavior recognition result does not meet the preset payment behavior condition, then the target face-scanning user is determined to have no willingness to pay. The above-mentioned preset range can be less than 0.1 and greater than 0.8, less than 0.2 and greater than 0.9, etc. The above-mentioned preset gaze condition can be that the probability that the target face-scanning user is looking at the screen or camera of the face-scanning device is greater than 0.99, 0.95, etc. The above-mentioned preset payment behavior condition can be that the probability that the target face-scanning user has payment behavior is greater than 0.99, 0.95, etc., and the embodiments of this specification do not limit this.
[0171] For example, when the recognition results include the payment willingness recognition results, the payment behavior recognition results and the gaze recognition results, the recognition results can be Figure 9 The process shown determines whether the target face-scanning user has the willingness to pay. Figure 9 The preset condition may be that the probability that the target face-scanning user has the willingness to pay in the payment willingness recognition result is greater than 0.8, 0.9, etc., and the preset requirement may be that the probability that the target face-scanning user is gazing at the screen or camera of the face-scanning device in the gaze recognition result is greater than 0.99, 0.95, etc., or the probability that the target face-scanning user has the payment behavior in the payment behavior recognition result is greater than 0.99, 0.95, etc., and the embodiments of this specification do not limit this.
[0172] In the embodiments of this specification, when the willingness to pay recognition result does not meet the preset conditions, that is, the willingness to pay in the willingness to pay recognition result is vague or weak, and it is difficult to accurately determine whether the target face-scanning user has the willingness to pay through the willingness to pay recognition result, it can also combine the gaze recognition result and / or the payment behavior recognition result to make a comprehensive judgment. For example, it is not limited to when the willingness to pay recognition result indicates that the willingness to pay of the target face-scanning user is 0.5, if the probability of the payment behavior recognition result indicating that the target face-scanning user has payment behavior is 1, it indicates that the target face-scanning user has a specific strong willingness to pay, that is, it can also be directly determined that the target face-scanning user has the willingness to pay, etc., thereby being able to more accurately realize the willingness to pay recognition when the willingness to pay in the willingness to pay recognition result is vague, thereby further improving the security of face-scanning payment.
[0173] Next, combine Figure 10 , introduces a willingness to pay identification model provided by the embodiment of this specification. Figure 10 As shown, when the recognition result in S408 includes a willingness-to-pay recognition result, its willingness-to-pay recognition model includes a first target convolutional network, a second target convolutional network, and a first fully connected layer.
[0174] The first target convolutional network is used to process the target human body area image and the target mask image to obtain the gaze feature corresponding to the target human body area image.
[0175] Specifically, the first target convolutional network includes a first convolutional module, a second convolutional module, and a third convolutional module. The first convolutional module is configured to extract a first feature corresponding to the target human region image. The second convolutional module is configured to fuse the first feature with the target mask image to generate a second feature. The third convolutional module is configured to generate a gaze feature corresponding to the target human region image based on the second feature.
[0176] Specifically, before fusing the first feature with the target mask map, the second convolution module needs to first adjust the resolution of the target mask map to the same as the resolution of the first feature, that is, to obtain the first target mask map, and then splice the first feature and the first target mask map according to the channel dimension. Since the splicing results in an additional channel dimension corresponding to the first target mask map, in order to ensure that the subsequent process can proceed normally, the spliced first feature and the first target mask map need to be fused through the second convolution module to generate a second feature with the same dimension as the first feature.
[0177] Optionally, the first convolution module and the third convolution module can be split from the same convolution network.
[0178] The second target convolutional network is used to process the target human body area image to obtain the limb behavior characteristics corresponding to the target human body area image.
[0179] Specifically, the second target convolutional network includes a first convolutional module and a fourth convolutional module. The first convolutional module is configured to extract a first feature corresponding to the target human body region image, and the fourth convolutional module is configured to generate a limb behavior feature corresponding to the target human body region image based on the first feature.
[0180] Optionally, the first convolution module and the fourth convolution module may be obtained by splitting the same convolution network. The third convolution module and the fourth convolution module may have the same corresponding structure but different parameters.
[0181] The first fully connected layer is used to fuse gaze features and body behavior features to output the payment willingness recognition result corresponding to the target face-scanning user.
[0182] Optionally, the recognition result in S408 may include not only the payment willingness recognition result, but also the gaze recognition result and the payment behavior recognition result. Figure 10 As shown, the willingness to pay recognition model also includes a second fully connected layer and a third connected layer. The second connected layer is used to output the gaze recognition result corresponding to the target face-scanning user based on the gaze features. The third connected layer is used to output the payment behavior recognition result corresponding to the target face-scanning user based on the gaze features.
[0183] Optionally, Figure 10 The first target convolutional network shown can be trained based on human body region images and mask images corresponding to multiple face-scanning users with known willingness to pay information and / or gaze information; the second target convolutional network is trained based on human body region images with known willingness to pay information and / or payment behavior information corresponding to multiple face-scanning users. The above-mentioned willingness to pay information includes a willingness to pay label corresponding to when the face-scanning user has a willingness to pay in the face-scanning image, and a willingness to pay label corresponding to when the face-scanning user does not have a willingness to pay; the above-mentioned gaze information includes a gaze label corresponding to when the face-scanning user is looking at the screen or camera of the face-scanning device in the face-scanning image, and a non-gaze label corresponding to when the face-scanning user is not looking at the screen or camera of the face-scanning device; the above-mentioned payment behavior information includes a payment behavior label corresponding to when the face-scanning user has a payment behavior such as clicking with body parts in the face-scanning image, and a non-payment behavior label corresponding to when the face-scanning user does not have a payment behavior such as clicking with body parts in the face-scanning image.
[0184] Understandably, Figure 10The shown willingness to pay recognition model can be obtained by directly training the face images of multiple face-scanning users with known willingness to pay information as a whole, or first training the first target convolutional network based on the human body area images and mask images with known willingness to pay information and / or gaze information corresponding to multiple face-scanning users, and then training the second target convolutional network based on the human body area images with known willingness to pay information and / or payment behavior information corresponding to multiple face-scanning users, or first training the second target convolutional network, and then training the first target convolutional network based on the trained second target convolutional network, so as to obtain a trained willingness to pay recognition model. The embodiments of this specification do not limit the training method of the willingness to pay recognition model.
[0185] Please refer to Figure 11 , Figure 11 A willingness-to-pay identification device is provided as an exemplary embodiment of this specification. The willingness-to-pay identification device 1100 includes:
[0186] An acquisition module 1110 is configured to acquire a target face recognition image; the target face recognition image includes the target face recognition user;
[0187] The generating module 1120 is configured to generate a corresponding target mask image based on the target position of the target face-scanning user in the target face-scanning image; the target mask image is configured to distinguish the facial region of the target face-scanning user from other regions except the facial region;
[0188] The willingness to pay recognition module 1130 is used to input the above-mentioned target face-scanning image and the above-mentioned target mask image into the willingness to pay recognition model, obtain gaze features and body behavior features, and output the recognition result corresponding to the above-mentioned target face-scanning user based on the above-mentioned gaze features and the above-mentioned body behavior features; the above-mentioned willingness to pay recognition model is trained based on the face-scanning images of multiple face-scanning users with known willingness to pay information.
[0189] In a possible implementation, the willingness-to-pay identification device 1120 further includes:
[0190] A first determining module is configured to determine a target body region image of the target face-scanning user based on the face region of the target face-scanning user and the target face-scanning image; the target body region image includes the limbs of the target face-scanning user;
[0191] The above-mentioned willingness to pay identification module 1130 is specifically used to: input the above-mentioned target human body area image and the above-mentioned target mask image into the willingness to pay identification model, obtain gaze features and limb behavior features, and output the identification result corresponding to the above-mentioned target face-scanning user based on the above-mentioned gaze features and the above-mentioned limb behavior features.
[0192] In one possible implementation, the willingness to pay identification module 1130 includes:
[0193] an extraction unit, configured to extract a first feature corresponding to the target human body region image;
[0194] A fusion unit, configured to fuse the first feature with the target mask image to generate a second feature;
[0195] A first determining unit is configured to determine a limb behavior feature corresponding to the target human body region image based on the first feature;
[0196] A second determining unit is configured to determine a gaze feature corresponding to the target human body region image based on the second feature;
[0197] The third determination unit is used to determine the recognition result corresponding to the target face-scanning user based on the gaze characteristics and the body behavior characteristics.
[0198] In one possible implementation, the above identification result includes a willingness-to-pay identification result;
[0199] The willingness to pay identification device 1120 further includes:
[0200] The second determination module is used to determine whether the target face-scanning user has the willingness to pay based on the above recognition result.
[0201] In a possible implementation, the above recognition result also includes a gaze recognition result and a payment behavior recognition result;
[0202] The second determining module is specifically configured to:
[0203] If the payment willingness recognition result does not meet the preset conditions, whether the target face-scanning user has the payment willingness is determined based on the gaze recognition result and / or the payment behavior recognition result.
[0204] In a possible implementation, the target face-scanning image includes multiple face-scanning users, and the multiple face-scanning users include the target face-scanning user;
[0205] The willingness to pay identification device 1120 further includes:
[0206] The third determination module is used to determine the target face-scanning user from the multiple face-scanning users according to preset rules.
[0207] In one possible implementation, the recognition result includes a willingness-to-pay recognition result; the willingness-to-pay recognition model includes a first target convolutional network, a second target convolutional network, and a first fully connected layer;
[0208] The first target convolutional network is used to process the target human region image and the target mask image to obtain a gaze feature corresponding to the target human region image;
[0209] The second target convolutional network is used to process the target human body region image to obtain limb behavior features corresponding to the target human body region image;
[0210] The first fully connected layer is used to fuse the gaze features and the body behavior features to output the payment willingness recognition result corresponding to the target face-scanning user.
[0211] In a possible implementation, the first target convolutional network includes a first convolutional module, a second convolutional module, and a third convolutional module;
[0212] The first convolution module is used to extract the first feature corresponding to the target human body region image;
[0213] The second convolution module is used to fuse the first feature with the target mask image to generate a second feature;
[0214] The third convolution module is used to generate a gaze feature corresponding to the target human body area image based on the second feature.
[0215] In a possible implementation, the second target convolutional network includes a first convolutional module and a fourth convolutional module;
[0216] The first convolution module is used to extract the first feature corresponding to the target human body region image;
[0217] The fourth convolution module is used to generate a limb behavior feature corresponding to the target human body region image based on the first feature.
[0218] In a possible implementation, the recognition result further includes a gaze recognition result and a payment behavior recognition result; the willingness to pay recognition model further includes a second fully connected layer and a third connected layer;
[0219] The second connection layer is configured to output the gaze recognition result corresponding to the target face scanning user based on the gaze feature;
[0220] The third connection layer is used to output the payment behavior recognition result corresponding to the target face-scanning user based on the gaze feature.
[0221] In one possible implementation, the first target convolutional network is trained based on human body region images and mask images corresponding to the plurality of face-scanning users, each of which has known willingness-to-pay information and / or gaze information.
[0222] The second target convolutional network is trained based on human body area images corresponding to the multiple face-scanning users with known willingness to pay information and / or payment behavior information.
[0223] In a possible implementation, the target human region image and the target mask image have the same resolution.
[0224] The division of the modules in the above-described willingness-to-pay identification device is for illustrative purposes only. In other embodiments, the willingness-to-pay identification device can be divided into different modules as needed to perform all or part of the functions of the above-described willingness-to-pay identification device. The various modules in the willingness-to-pay identification device provided in the embodiments of this specification can be implemented in the form of a computer program. This computer program can be executed on a terminal or server. The program modules comprising this computer program can be stored in the memory of the terminal or server. When executed by a processor, this computer program implements all or part of the steps of the willingness-to-pay identification method described in the embodiments of this specification.
[0225] See also Figure 12 , Figure 12 This is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. Figure 12 As shown, the electronic device 1200 may include: at least one processor 1210, at least one communication bus 1220, a user interface 1230, at least one network interface 1240, and a memory 1250. The communication bus 1220 may be used to implement connection and communication between the above components.
[0226] The user interface 1230 may include a display screen (Display) and a camera (Camera), and the optional user interface may also include a standard wired interface and a wireless interface.
[0227] The network interface 1240 may optionally include a Bluetooth module, a Near Field Communication (NFC) module, a Wireless Fidelity (Wi-Fi) module, and the like.
[0228] The processor 1210 may include one or more processing cores. The processor 1210 utilizes various interfaces and circuits to connect various components within the electronic device 1200. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1250, and accessing data stored in the memory 1250, the processor 1210 performs various functions and processes data for the routing electronic device 1200. Optionally, the processor 1210 may be implemented using at least one hardware form factor selected from the group consisting of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 1210 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 1210 and may be implemented as a separate chip.
[0229] Among them, the memory 1250 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 1250 includes a non-transitory computer-readable medium. The memory 1250 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 1250 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as an acquisition function, a generation function, a willingness to pay identification function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 1250 may also be optionally at least one storage device located away from the aforementioned processor 1210. As Figure 12 As shown, the memory 1250 as a computer storage medium may include an operating system, a network communication module, a user interface module, and program instructions.
[0230] Specifically, the processor 1210 may be configured to call program instructions stored in the memory 1250 and perform the following operations:
[0231] Obtain a target face-scanning image; the target face-scanning image includes the target face-scanning user.
[0232] A corresponding target mask image is generated based on the target position of the target face-scanning user in the target face-scanning image; the target mask image is used to distinguish the facial area of the target face-scanning user from other areas except the facial area.
[0233] The target face-scanning image and the target mask image are input into a willingness-to-pay recognition model to obtain gaze features and body behavior features, and the recognition result corresponding to the target face-scanning user is output based on the gaze features and the body behavior features; the willingness-to-pay recognition model is trained based on face-scanning images of multiple face-scanning users with known willingness-to-pay information.
[0234] In some possible embodiments, after the processor 1210 obtains the target face-scanning image, it inputs the target face-scanning image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and before outputting the recognition result corresponding to the target face-scanning user based on the gaze features and the body behavior features, it is further configured to execute:
[0235] The target body area image of the target face-scanning user is determined based on the facial area of the target face-scanning user and the target face-scanning image; the target body area image includes the limbs of the target face-scanning user.
[0236] The processor 1210 inputs the target face-scanning image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and outputs a recognition result corresponding to the target face-scanning user based on the gaze features and body behavior features, specifically for performing the following:
[0237] The target human body area image and the target mask image are input into the willingness-to-pay recognition model to obtain gaze features and limb behavior features, and the recognition result corresponding to the target face-scanning user is output based on the gaze features and the limb behavior features.
[0238] In some possible embodiments, the processor 1210 inputs the target human body region image and the target mask image into a willingness-to-pay recognition model to obtain gaze features and limb behavior features, and outputs a recognition result corresponding to the target face-scanning user based on the gaze features and limb behavior features, specifically for performing:
[0239] Extract the first feature corresponding to the target human body region image.
[0240] The first feature is fused with the target mask image to generate a second feature.
[0241] The limb behavior feature corresponding to the target human body region image is determined based on the first feature.
[0242] The gaze feature corresponding to the target human body region image is determined based on the second feature.
[0243] The recognition result corresponding to the target face scanning user is determined based on the gaze characteristics and the body behavior characteristics.
[0244] In some possible embodiments, the above-mentioned recognition result includes a willingness to pay recognition result; the above-mentioned processor 1210 inputs the above-mentioned target face-scanning image and the above-mentioned target mask image into the willingness to pay recognition model, obtains gaze features and body behavior features, and outputs the recognition result corresponding to the above-mentioned target face-scanning user based on the above-mentioned gaze features and the above-mentioned body behavior features, and is also used to execute: determining whether the above-mentioned target face-scanning user has a willingness to pay based on the above-mentioned recognition result.
[0245] In some possible embodiments, the above recognition results also include gaze recognition results and payment behavior recognition results;
[0246] When the processor 1210 determines whether the target face-scanning user has the willingness to pay based on the recognition result, it is specifically used to execute: if the recognition result of the willingness to pay does not meet the preset conditions, determine whether the target face-scanning user has the willingness to pay based on the gaze recognition result and / or the payment behavior recognition result.
[0247] In some possible embodiments, the target face-scanning image includes multiple face-scanning users, and the multiple face-scanning users include the target face-scanning user; after the processor 1210 obtains the target face-scanning image, before generating a corresponding target mask map based on the target position of the target face-scanning user in the target face-scanning image, the processor 1210 is further configured to perform:
[0248] The target face-scanning user is determined from the multiple face-scanning users according to preset rules.
[0249] In some possible embodiments, the recognition result includes a willingness-to-pay recognition result; the willingness-to-pay recognition model includes a first target convolutional network, a second target convolutional network, and a first fully connected layer;
[0250] The above-mentioned first target convolutional network is used to process the above-mentioned target human body area image and the above-mentioned target mask image to obtain the gaze features corresponding to the above-mentioned target human body area image; the above-mentioned second target convolutional network is used to process the above-mentioned target human body area image to obtain the limb behavior features corresponding to the above-mentioned target human body area image; the above-mentioned first fully connected layer is used to fuse the above-mentioned gaze features and the above-mentioned limb behavior features, and output the payment willingness recognition result corresponding to the above-mentioned target face-scanning user.
[0251] In some possible embodiments, the first target convolutional network includes a first convolutional module, a second convolutional module, and a third convolutional module;
[0252] The first convolution module is used to extract the first feature corresponding to the target human body region image;
[0253] The second convolution module is used to fuse the first feature with the target mask image to generate a second feature;
[0254] The third convolution module is used to generate a gaze feature corresponding to the target human body area image based on the second feature.
[0255] In some possible embodiments, the second target convolutional network includes a first convolutional module and a fourth convolutional module;
[0256] The first convolution module is used to extract the first feature corresponding to the target human body area image; the fourth convolution module is used to generate the limb behavior feature corresponding to the target human body area image based on the first feature.
[0257] In some possible embodiments, the above-mentioned recognition results also include gaze recognition results and payment behavior recognition results; the above-mentioned willingness to pay recognition model also includes a second fully connected layer and a third connected layer; the above-mentioned second connected layer is used to output the above-mentioned gaze recognition result corresponding to the above-mentioned target face-scanning user based on the above-mentioned gaze features; the above-mentioned third connected layer is used to output the above-mentioned payment behavior recognition result corresponding to the above-mentioned target face-scanning user based on the above-mentioned gaze features.
[0258] In one possible implementation, the first target convolutional network is trained based on human body region images and mask images corresponding to the plurality of face-scanning users, each of which has known willingness-to-pay information and / or gaze information.
[0259] The second target convolutional network is trained based on human body area images corresponding to the multiple face-scanning users with known willingness to pay information and / or payment behavior information.
[0260] In some possible embodiments, the target human body region image and the target mask image have the same resolution.
[0261] The embodiments of this specification also provide a computer-readable storage medium containing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps of the aforementioned embodiments. If the various components of the aforementioned willingness-to-pay identification device are implemented as software functional units and sold or used as independent products, they may be stored in the aforementioned computer-readable storage medium.
[0262] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The above-mentioned computer program product includes one or more computer instructions. When the above-mentioned computer program instructions are loaded and executed on a computer, the above-mentioned process or function according to the embodiment of this specification is generated in whole or in part. The above-mentioned computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above-mentioned computer instructions can be stored in a computer-readable storage medium or transmitted by the above-mentioned computer-readable storage medium. The above-mentioned computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The above-mentioned computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The above-mentioned available media can be magnetic media (for example, floppy disks, hard disks, tapes), optical media (for example, digital versatile discs (DVDs)), or semiconductor media (for example, solid state disks (SSDs)).
[0263] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. The technical features of this embodiment and the implementation scheme can be combined in any manner unless they conflict.
[0264] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Without departing from the design spirit of this specification, various modifications and improvements made to the technical solutions of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims.
[0265] The foregoing description of specific embodiments of this specification is intended to be a description of other embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims and specification can be performed in a different order than that described in the embodiments described in the specification and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A method for identifying willingness to pay, comprising: Get the target face scan image; The target face scanning image includes the target face scanning user; Generate a corresponding target mask image based on the target position of the target face-scanning user in the target face-scanning image; The target mask image is used to distinguish the facial area of the target face-scanning user from other areas except the facial area; The target face-scanning image and the target mask image are input into a willingness-to-pay recognition model to obtain gaze features and body behavior features, and the gaze features and body behavior features are fused, and an identification result corresponding to the target face-scanning user is output based on the fused features; the willingness-to-pay recognition model is trained based on face-scanning images of multiple face-scanning users each corresponding to known willingness-to-pay information; the gaze features represent whether the target face-scanning user is looking at the screen or camera of the face-scanning device, and the body behavior features represent whether the target face-scanning user has made a payment action; After acquiring the target face-scanning image, inputting the target face-scanning image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and before outputting the recognition result corresponding to the target face-scanning user based on the gaze features and the body behavior features, the method further includes: Determining a target body region image of the target face-scanning user based on the face region of the target face-scanning user and the target face-scanning image; the target body region image includes the limbs of the target face-scanning user; Inputting the target face-scanning image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and outputting a recognition result corresponding to the target face-scanning user based on the gaze features and the body behavior features, includes: Extracting a first feature corresponding to the target human body region image; Fusing the first feature with the target mask image to generate a second feature; Determining a limb behavior feature corresponding to the target human body region image based on the first feature; Determining a gaze feature corresponding to the target human body region image based on the second feature; The recognition result corresponding to the target face-scanning user is determined according to the gaze feature and the body behavior feature.
2. The method according to claim 1, wherein the identification result includes a willingness to pay identification result; After inputting the target face-scanning image and the target mask image into the willingness-to-pay recognition model to obtain gaze features and body behavior features, and outputting a recognition result corresponding to the target face-scanning user based on the gaze features and the body behavior features, the method further includes: Determine whether the target face-scanning user has a willingness to pay based on the recognition result.
3. The method according to claim 2, wherein the recognition result further comprises a gaze recognition result and a payment behavior recognition result; The determining, based on the recognition result, whether the target face-scanning user has a willingness to pay includes: If the payment willingness recognition result does not meet the preset conditions, whether the target face-scanning user has the payment willingness is determined according to the gaze recognition result and / or the payment behavior recognition result.
4. The method according to claim 1, wherein the target face-scanning image includes a plurality of face-scanning users, and the plurality of face-scanning users includes the target face-scanning user; After acquiring the target face-scanning image, and before generating a corresponding target mask map based on the target position of the target face-scanning user in the target face-scanning image, the method further includes: A target face-scanning user is determined from the multiple face-scanning users according to preset rules.
5. The method of claim 1, wherein the recognition result includes a willingness-to-pay recognition result; the willingness-to-pay recognition model includes a first target convolutional network, a second target convolutional network, and a first fully connected layer; The first target convolutional network is used to process the target human body region image and the target mask image to obtain a gaze feature corresponding to the target human body region image; The second target convolutional network is used to process the target human body region image to obtain limb behavior features corresponding to the target human body region image; The first fully connected layer is used to fuse the gaze features and the body behavior features to output the payment willingness recognition result corresponding to the target face-scanning user.
6. The method of claim 5, wherein the first target convolutional network comprises a first convolutional module, a second convolutional module, and a third convolutional module; The first convolution module is used to extract a first feature corresponding to the target human body area image; The second convolution module is used to fuse the first feature with the target mask image to generate a second feature; The third convolution module is used to generate a gaze feature corresponding to the target human area image based on the second feature.
7. The method of claim 5, wherein the second target convolutional network comprises a first convolutional module and a fourth convolutional module; The first convolution module is used to extract a first feature corresponding to the target human body area image; The fourth convolution module is used to generate a limb behavior feature corresponding to the target human body area image based on the first feature.
8. The method of claim 5, wherein the recognition result further comprises a gaze recognition result and a payment behavior recognition result; and the willingness to pay recognition model further comprises a second fully connected layer and a third connected layer; The second fully connected layer is configured to output the gaze recognition result corresponding to the target face scanning user based on the gaze feature; The third connection layer is used to output the payment behavior recognition result corresponding to the target face-scanning user based on the gaze feature.
9. The method of claim 5, wherein the first target convolutional network is trained based on body region images and mask images corresponding to the known willingness to pay and / or gaze information of the multiple face-scanning users; The second target convolutional network is trained based on human body area images corresponding to the multiple face-scanning users, each of which has known payment willingness information and / or payment behavior information.
10. The method according to any one of claims 1 or 5 to 9, wherein the target human body region image and the target mask image have the same resolution.
11. A willingness-to-pay identification device, comprising: An acquisition module is used to obtain the target face scan image; The target face scanning image includes the target face scanning user; A generating module, configured to generate a corresponding target mask image based on a target position of the target face-scanning user in the target face-scanning image; The target mask image is used to distinguish the facial area of the target face-scanning user from other areas except the facial area; A willingness-to-pay recognition module is configured to input the target face-scanning image and the target mask image into a willingness-to-pay recognition model, obtain gaze features and body behavior features, fuse the gaze features and body behavior features, and output a recognition result corresponding to the target face-scanning user based on the fused features; the willingness-to-pay recognition model is trained based on face-scanning images of multiple face-scanning users each corresponding to known willingness-to-pay information; the gaze features indicate whether the target face-scanning user is looking at the screen or camera of the face-scanning device, and the body behavior features indicate whether the target face-scanning user has made a payment action; The device further comprises: A first determining module is configured to determine a target body region image of the target face-scanning user based on the face region of the target face-scanning user and the target face-scanning image; the target body region image includes the limbs of the target face-scanning user; The willingness to pay identification module is specifically used to: Extract a first feature corresponding to the target human body area image; fuse the first feature with the target mask image to generate a second feature; determine the limb behavior feature corresponding to the target human body area image based on the first feature; determine the gaze feature corresponding to the target human body area image based on the second feature; determine the recognition result corresponding to the target face-scanning user according to the gaze feature and the limb behavior feature.
12. An electronic device comprising: processor and memory; The processor is connected to the memory; The memory is used to store executable program code; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method according to any one of claims 1 to 10.
13. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 10.
14. A computer program product comprising instructions, which, when executed on a computer or a processor, causes the computer or the processor to execute the willingness-to-pay identification method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Face-scanning payment willingness identification method, device and equipment
CN114511910A
Entry management device
JP2006251946A