A method and system for identifying a target object
By combining 2D and 3D image samples and using a willingness recognition model to extract and fuse behavioral and spatial distribution features, the problem of low accuracy in target object recognition in multi-user public places is solved, thereby improving the accuracy and security of facial recognition payment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2023-02-16
- Publication Date
- 2026-07-21
AI Technical Summary
In public spaces with multiple users, the accuracy of machine recognition of target objects is compromised, leading to frequent misidentification.
By combining 2D and 3D image samples, behavioral and spatial distribution features are extracted using a willingness recognition model, and the willingness attributes of the target object are determined by fusion, including willingness safety and willingness high risk.
It improves the accuracy of facial recognition payment in offline public places, avoids accidental swiping, and enhances the user's security experience.
Smart Images

Figure CN116071630B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a method and system for identifying target objects. Background Technology
[0002] In public places where multiple users are present, when using machines to identify target objects, the presence of multiple users often interferes with the accuracy of the machine's identification of the target object, resulting in misidentification. Summary of the Invention
[0003] In order to solve the problems existing in the prior art, the main purpose of this disclosure is to provide a method and system for identifying target objects.
[0004] This disclosure provides a method for identifying a target object, including:
[0005] Obtain 2D image samples and 3D image samples of the target scene, wherein each of the 2D image samples and the 3D image samples contains at least one target object;
[0006] Based on the intention recognition model, the 2D image samples and the 3D image samples are predicted to obtain the behavioral feature map and spatial distribution feature map of the at least one target object;
[0007] Based on the intention recognition model, the behavioral feature map and the spatial distribution feature map are predicted to determine the intention attributes of the at least one target object for the target transaction. The intention attributes include either intention to be safe or intention to be high-risk.
[0008] In some embodiments, the step of predicting the 2D image samples and the 3D image samples based on the intention recognition model to obtain the behavioral feature map and spatial distribution feature map of the at least one target object includes:
[0009] Based on the intent recognition model, the 2D image samples are predicted to determine the behavioral feature map of the at least one target object; and
[0010] Based on the intention recognition model, the spatial distribution feature map of the at least one target object is determined by predicting the 2D image samples and the 3D image samples.
[0011] In some embodiments, predicting the 2D image samples based on the intention recognition model to determine the behavioral feature map of the at least one target object includes:
[0012] Based on the intention recognition model, the 2D image sample is predicted to obtain a key point feature map, which contains key point information on at least one target object.
[0013] Based on the intent recognition model, feature extraction is performed on the 2D image samples to obtain a 2D modal feature map; and
[0014] The behavioral feature map is determined based on the key point feature map and the 2D modal feature map, and the behavioral feature map enhances the information of the key points in the 2D modal feature map.
[0015] In some embodiments, the intention recognition model includes a keypoint detection subnetwork; and
[0016] The step of predicting the 2D image samples based on the intent recognition model to obtain key point feature maps includes using the key point detection sub-network:
[0017] Human key points are extracted from the 2D image samples to obtain a human key point map, and
[0018] The key point map of the human body is predicted to obtain the key point feature map. The result of each position in the key point feature map corresponds to the probability value that each position of the 2D image sample is a key point of the at least one target object.
[0019] In some embodiments, determining the behavioral feature map based on the keypoint feature map and the 2D modal feature map includes:
[0020] The key point feature map and the 2D modal feature map are fused to obtain the behavior feature map, which emphasizes the behavior feature information brought by the key points in the 2D modal feature map.
[0021] In some embodiments, the step of predicting the spatial distribution feature map of the at least one target object based on the intention recognition model for the 2D image samples and the 3D image samples includes:
[0022] Based on the intention recognition model, the 2D image sample is predicted to obtain a target segmentation probability map, which contains the difference information between the background and the target object in the 2D image sample;
[0023] Based on the intent recognition model, feature extraction is performed on the 3D image samples to obtain a 3D modal feature map; and
[0024] The spatial distribution feature map is determined based on the target segmentation probability map and the 3D modal feature map, and the spatial distribution feature map enhances the spatial information of the target segmentation probability map.
[0025] In some embodiments, the intention recognition model further includes a target segmentation task subnetwork; and
[0026] The step of predicting the 2D image samples based on the intent recognition model to obtain a target segmentation probability map includes using the target segmentation task sub-network:
[0027] The 2D image sample is segmented to obtain a target segmentation feature map, and based on the target segmentation feature map, a target segmentation probability map is obtained. The result of each position in the target segmentation probability map corresponds to the probability value that each position of the 2D image sample is the at least one target object.
[0028] In some embodiments, the step of extracting features from the 3D image samples based on the intention recognition model to obtain a 3D modal feature map includes:
[0029] The intention recognition model is used to normalize the 3D image samples to obtain a normalized 3D depth image; and
[0030] Based on the intention recognition model, feature extraction is performed on the normalized 3D depth image to obtain the 3D modal feature map.
[0031] In some embodiments, determining the spatial distribution feature map based on the target segmentation probability map and the 3D modal feature map includes:
[0032] The target segmentation probability map and the 3D modal feature map are fused to obtain the spatial distribution feature map, which highlights the spatial location information of at least one target object in the 2D modal feature map.
[0033] In some embodiments, the step of predicting the behavioral feature map and the spatial distribution feature map based on the intention recognition model to determine the intention attribute of the at least one target object includes:
[0034] The behavioral feature map and the spatial distribution feature map are fused to obtain a multi-task fusion feature map, which integrates the behavioral features of the 2D attributes of the target object and the spatial distribution features of the 3D attributes in the target scene; and
[0035] Based on the intention recognition model, the multi-task fusion feature map is predicted to determine the intention attributes of at least one target object.
[0036] In some embodiments, the intention recognition model includes a multi-task prediction subnetwork; and
[0037] The step of predicting the multi-task fusion feature map based on the intention recognition model to determine the intention attribute of the at least one target object includes:
[0038] Obtain the initial intention object of the target scene, wherein the initial intention object is one of the at least one target object, and
[0039] The multi-task prediction sub-network is used to predict the multi-task fusion feature map to determine the intention attributes of the initial intention object.
[0040] In some embodiments, obtaining the initial intention object of the target scenario includes:
[0041] Determine the target scenario corresponding to the triggering of the target transaction, wherein the target scenario includes the at least one target object; and
[0042] Based on the initial identification model, a target object is determined from the at least one target object as the initial intention object.
[0043] In some embodiments, predicting the multi-task fusion feature map based on the multi-task prediction subnetwork to determine the intention attributes of the initial intention object includes:
[0044] Based on the multi-task prediction subnetwork, the multi-task fusion feature map is predicted to determine the initial intention classification probability value of the intention object; and
[0045] The intention classification probability value is compared with a preset intention classification probability threshold, and the intention attribute of the initial intention object is determined based on the comparison result.
[0046] In some embodiments, the training process of the intention recognition model includes:
[0047] Obtain 2D image samples and 3D image samples of the target scene, wherein each of the 2D image samples and the 3D image samples contains at least one target object;
[0048] Based on a preset intention recognition model, feature prediction is performed on the 2D image samples and the 3D image samples to obtain key point feature maps, target segmentation probability maps, and intention attributes of at least one target object; and
[0049] The preset intention recognition model is trained based on the key point feature map, the target segmentation probability map, and the multiple intention attributes to obtain the trained intention recognition model.
[0050] In some embodiments, the preset intention recognition model includes a preset keypoint detection subnetwork, a preset target segmentation task subnetwork, and a preset multi-task prediction subnetwork; and
[0051] The step of training the preset intention recognition model based on the key point feature map, the target segmentation probability map, and the intention attribute includes:
[0052] Based on the key point feature map, the first feature loss information of the preset key point detection sub-network is determined.
[0053] The second feature loss information of the preset target segmentation task sub-network is determined based on the target segmentation probability map.
[0054] The third feature loss information of the preset multi-task prediction sub-network is determined based on the intention attributes of the at least one target object, and
[0055] The preset intention recognition model is converged based on the first feature loss information, the second feature loss information, and the third feature loss information to obtain the trained intention recognition model.
[0056] In some embodiments, determining the first feature loss information of the preset keypoint detection sub-network based on the keypoint feature map includes:
[0057] Obtain the Gaussian heatmap of key points of the 2D image sample; and
[0058] The Gaussian heatmap of the key points is compared with the feature map of the key points to determine the first feature loss information.
[0059] In some embodiments, determining the second feature loss information of the preset target segmentation task sub-network based on the target segmentation probability map includes:
[0060] Obtain the target segmentation annotation map of the 2D image sample; and
[0061] The target segmentation annotation map is compared with the target segmentation probability map to determine the second feature loss information.
[0062] In some embodiments, determining the third feature loss information of the preset multi-task prediction sub-network based on the intention attributes of the at least one target object includes:
[0063] Obtain the original annotation attributes of the at least one target object; and
[0064] The original labeled attributes are compared with the intention attributes to determine the third feature loss information.
[0065] In some embodiments, the step of converging the preset intention recognition model based on the first feature loss information, the second feature loss information, and the third feature loss information to obtain the trained intention recognition model includes:
[0066] The first feature loss information, the second feature loss information, and the third feature loss information are fused to obtain comprehensive loss information; and
[0067] The preset intention recognition model is converged based on the comprehensive loss information to obtain the trained intention recognition model.
[0068] In some embodiments, fusing the first feature loss information, the second feature loss information, and the third feature loss information to obtain comprehensive loss information includes:
[0069] Obtain the weight attributes of the first feature loss information, the second feature loss information, and the third feature loss information respectively; and
[0070] The first feature loss information, the second feature loss information, and the third feature loss information are weighted based on the weight attributes to determine the comprehensive loss information.
[0071] This disclosure also provides a target object identification system, comprising: at least one storage medium including at least one instruction set for implementing and analyzing a target object identification method; and at least one processor communicatively connected to the at least one storage medium, wherein, when the system is running, the at least one processor reads the at least one instruction set and executes the target object identification method according to the instructions of the at least one instruction set.
[0072] As can be seen from the above technical solutions, the target object identification method and the system for executing this method provided in this disclosure. The method and system comprehensively process 2D and 3D image samples, fusing human keypoint detection with 2D modal feature maps to determine the behavioral feature map of the target object, and fusing human target segmentation with 3D modal feature maps to determine the spatial distribution feature map of the target object. Then, by fusing the behavioral feature map and the spatial distribution feature map, the intention attribute of the target object is predicted. This enables accurate prediction of the target object's intention to use facial recognition payment in offline public places, avoiding false scans, ensuring the security of the facial recognition system, and improving the user's security experience with facial recognition payment.
[0073] Other functions of the target object identification method and system provided in this disclosure will be partially listed in the following description. The figures and examples described below will be apparent to those skilled in the art. The inventive aspects of the target object identification method and system provided in this disclosure can be fully explained by practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0075] Figure 1 A schematic diagram of the terminal structure of the hardware operating environment involved in the embodiments of this disclosure is shown;
[0076] Figure 2 A server schematic diagram of a target object identification method according to some embodiments of the present disclosure is shown;
[0077] Figure 3 A flowchart of a method for identifying a target object according to some embodiments of the present disclosure is shown;
[0078] Figure 4 A schematic diagram of the structure of an intention recognition model provided according to some embodiments of the present disclosure is shown; and
[0079] Figure 5 A flowchart illustrating a method for training an intention recognition model according to some embodiments of the present disclosure is shown. Detailed Implementation
[0080] The following description provides specific application scenarios and requirements for this disclosure, intended to enable those skilled in the art to make and use the content of this disclosure. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this disclosure. Therefore, this disclosure is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.
[0081] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” as used herein may also include the plural forms. When used in this disclosure, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.
[0082] In view of the following description, these and other features of this disclosure, as well as the operation and function of the related elements of the structure, and the economy of assembly and manufacture of the components, can be significantly improved. All of these form part of this disclosure with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this disclosure. It should also be understood that the drawings are not drawn to scale.
[0083] The flowcharts used in this disclosure illustrate operations implemented according to some embodiments of this disclosure. It should be clearly understood that the operations in the flowcharts may not be implemented sequentially. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0084] Figure 1 A schematic diagram of a target object identification system 100 according to some embodiments of the present disclosure is shown. The target object identification system 100 may include a client 110, an integrated development platform 120, an integrated development platform server 130, and a database 140.
[0085] Client 110 can include various types of Internet of Things (IoT) devices, such as IoT devices in public places, including terminal devices with facial recognition capabilities deployed in public consumption scenarios such as supermarkets, convenience stores, restaurants, hotels, inns, and educational and medical facilities. These IoT devices can support payment for targeted transactions. IoT devices can collect various types of image data (such as 2D and 3D images) through cameras. Specifically, IoT devices can capture 2D images using a 2D camera and 3D images using a 3D camera. 3D cameras can capture images using structured light, Time-of-Flight (TOF), or binocular stereo imaging. Their working principle involves an infrared laser emitter emitting near-infrared light, which is reflected by the human body (mainly the face). The infrared information is received by an infrared CMOS image processor and aggregated into an image processing chip to obtain three-dimensional data of the human body (mainly the face), thereby achieving spatial positioning.
[0086] Generally, the image data collected by client 110 includes human (mainly facial) image data. IoT devices can directly process the collected image data to identify individuals in the image who have the willingness to pay for the target transaction; alternatively, the IoT device can send the collected image data to other devices in system 100 (such as the integrated development platform 120 or server 130) for processing to identify individuals in the image who have the willingness to pay for the target transaction. However, in actual offline IoT device facial recognition payment scenarios, there are often long queues of users. The images collected by the IoT device through the camera inevitably contain multiple target individuals. In this case, client 110 can determine the willingness attribute of at least one target individual for the target transaction (including either a safe or high-risk willingness), and decide whether to complete the payment for the target transaction based on the willingness attribute of at least one target individual.
[0087] Taking a scenario where offline IoT devices in public places use facial recognition payment for a target transaction as an example, when a user triggers a payment for the target transaction, the client 110 can capture an image within the camera's field of view (the target scene) (the image contains at least one human target object, and generally, the user triggering the target transaction is in the image as one of the human target objects). The client 110 can process the captured image to determine the willingness attribute of at least one human target object in the image for the target transaction. For example, if the image contains three target objects A, B, and C, the client 110 can process the captured image to determine the willingness attribute of target object A; the client 110 can also process the captured image to determine the willingness attribute of target object B; and the client 110 can also process the captured image to determine the willingness attribute of target object C.
[0088] It should be understood that the number of target objects contained in the image of the target scene collected by the client 110 is not limited to the above-mentioned number, and the image of the target scene collected by the client 110 may also contain any other number of target objects.
[0089] When a user triggers a target transaction to make a payment, the client 110 can also capture images of the target scene within the camera's field of view and send them to other devices in the system 100 (such as the integrated development platform 120 and the server 130). The other devices in the system 100 (such as the integrated development platform 120 and the server 130) can analyze and process the captured images to confirm the willingness attributes of at least one human target object in the image for the target transaction.
[0090] Integrated Development Platform 120, also known as an Integrated Development Environment (IDE), is an application that provides a program development environment, typically including tools such as a code editor, compiler, debugger, and graphical user interface. Developers can write program code (i.e., develop programs) on the IDE through client 110. The IDE server 130 (hereinafter referred to as server 130) can be a computing device on the IDE 120 specifically used to handle the identification of target objects in the program.
[0091] Server 130 may store data or instructions for performing the target object identification method described in this disclosure, and may execute or be used to execute said data and / or instructions. Server 130 may include hardware devices with data processing capabilities and the necessary programs required to drive the hardware devices. Of course, server 130 may also be merely a hardware device with data processing capabilities, or merely a program running on the hardware device. In some embodiments, server 130 may also be a plug-in and deployed on client 110.
[0092] Database 140 may store data and / or instructions. In some embodiments, database 140 may store data and / or instructions executed by server 130 or used to execute a method for identifying a target object in a program described in this disclosure. Client 110 and server 130 may have access to database 140, and client 110 and server 130 may access data or instructions stored in database 140 via a network. In some embodiments, database 140 may be directly connected to client 110 and server 130. In some embodiments, database 110 may be part of server 130. In some embodiments, database 140 may include mass storage, removable storage, volatile read-write memory, read-only memory (ROM), or similar content, or any combination thereof. Exemplary mass storage may include non-transitory storage media such as disks, optical discs, and solid-state drives. Exemplary removable storage may include flash drives, floppy disks, optical discs, memory cards, zip disks, magnetic tapes, etc. Typical volatile read-write memory may include random access memory (RAM). Example RAMs may include dynamic RAM (DRAM), dual date rate synchronous dynamic RAM (DDRSDRAM), static RAM (SRAM), thyristor RAM (T-RAM), and zero-capacitance RAM (Z-RAM), etc. Exemplary ROMs may include mask ROM (MROM), programmable ROM (PROM), virtual programmable ROM (PEROM), electronically programmable ROM (EEPROM), optical disc (CD-ROM), and digital multifunction disk ROM, etc.
[0093] It should be understood that Figure 1 The number of clients 110 and servers 130 shown is merely illustrative. Depending on implementation needs, there can be any number of clients 110 and servers 130.
[0094] It should be noted that the target object identification method can be executed entirely on the client 110, entirely on the server 130, or partially on the client 110 and partially on the server 130.
[0095] For ease of description, the following descriptions of this disclosure will use the execution of the target object identification method on server 130 as an example to describe the technical solutions involved in this disclosure.
[0096] Figure 2 This is a schematic diagram of the structure of a computing device 200 provided according to some embodiments of the present disclosure. The computing device 200 can be a general-purpose computer or a special-purpose computer. For example, the computing device 200 can be a server, a personal computer, a portable computer (such as a laptop computer, tablet computer, etc.), or other electronic devices with computing capabilities. Of course, the computing device can be... Figure 1 The server 130 can also be a terminal device used by multiple developers 110A, 110B, 110C (client 110) to develop programs on the integrated development platform.
[0097] like Figure 2 As shown, the computing device 200 may include a COM port 250, which can be connected to or from a network to facilitate data communication. The computing device 200 may also include a processor 220, such as a central processing unit (CPU), in the form of one or more processors for executing program instructions. The computing device 200 may also include an internal communication bus 210 and various forms of program storage media and data storage media, such as a disk 270 (non-transitory memory) and read-only memory (ROM) 230 or random access memory (RAM) 240, etc., for storing various data files to be processed and / or transmitted. The storage media may be local to the computing device 200 or shared by the computing device 200 (e.g., Figure 1 The computing device 200 may also include program instructions stored in ROM 230, RAM 240, and / or other types of non-transitory storage media to be executed by processor 220. The computing device 200 may also include I / O components 260 to support data communication with other computing devices in the distributed computing system 100. The computing device 200 may also receive programming and data via network communication.
[0098] For illustrative purposes only, only one processor 220 is described in the computing device 200. However, those skilled in the art will understand that the computing device 200 of this disclosure may also include multiple processors. Therefore, the methods / steps / operations performed by one processor as described in this disclosure may also be performed jointly or separately by multiple processors. For example, in this disclosure, the processors of the computing device 200 may simultaneously execute step A and step B. It should be understood that step A and step B may also be performed jointly by two different processors. For example, a first processor executes step A, a second processor executes step B, or a first processor and a second processor jointly execute steps A and B.
[0099] Figure 3 A flowchart 300 of a method for identifying a target object according to some embodiments of the present disclosure is shown. Figure 4 A schematic diagram of the structure of an intention recognition model provided according to some embodiments of the present disclosure is shown. The following will be combined with... Figure 3 , Figure 4 This disclosure describes the technical solution. The subject implementing the technical solution may be... Figure 1 The client 110, integrated development platform 120, and server 130 are selected from the above. Specifically, the client 110, integrated development platform 120, and / or server 130 may have the following characteristics: Figure 2 The aforementioned structure, namely, the client 110, the integrated development platform 120, and / or the server 130, can be a device for identifying target objects, comprising: at least one storage medium and at least one processor. The at least one storage medium includes at least one instruction set for a target object identification method in a program. The at least one processor is communicatively connected to the at least one storage medium. When the system is running, the at least one processor can read the at least one instruction set and execute instructions according to the at least one instruction set. Figure 3 The method 300. For illustrative purposes only, this application will describe the method 300 as being performed by server 130. The method 300 may include:
[0100] S310, obtain 2D image samples and 3D image samples of the target scene, wherein each of the 2D image samples and the 3D image samples contains at least one target object.
[0101] In this disclosure, the server 130 obtains 2D and 3D image samples of the target scene in the following way: Taking an offline payment scenario as an example, the client 110 can be an IoT payment device with facial recognition functionality. The client 110 can collect 2D image samples within the camera's field of view (target scene) using a 2D camera, and it can also collect 3D image samples within the camera's field of view (target scene) using a 3D camera. When a user triggers a payment for the target transaction, the client 110 receives or recognizes the payment task for the target transaction. The client 110 then begins collecting 2D and 3D image samples of the target scene and sends them to the server 130, which then obtains the 2D and 3D image samples of the target scene.
[0102] As mentioned earlier, in offline payment scenarios, multiple users often queue to pay by facial recognition, meaning there is at least one target object in the target scenario. Accordingly, both the 2D and 3D image samples contain at least one target object. Considering that in offline payment scenarios, when multiple users queue to pay by facial recognition, the randomness of their spatial positioning causes mutual occlusion, and the user triggering the target transaction may have their face partially or completely obscured for some reason (the 2D and 3D image samples do not capture the user's full facial information), in this disclosure, when the client 110 identifies a human body part (including the face and the torso excluding the face), it can be determined as a target object. The following example illustrates the scenario where three people, A, B, and C, appear in the target scenario. When client 110 can capture photos of A's entire face and part of his torso, client 110 can capture photos of B's part of his face and part of his torso, and client 110 can capture photos of C's part of his torso, then the 2D and 3D image samples collected by client 110 contain three target objects: target object A, target object B, and target object C.
[0103] When client 110 can capture images of A's entire face and part of his torso, and can capture images of B's part of his torso, but cannot capture images of C's face or torso, then the 2D and 3D image samples acquired by client 110 contain two target objects: target object A and target object B. C is not included in the scope of target objects. The method for determining target objects in other shooting scenarios follows the same logic and will not be listed here.
[0104] In this disclosure, since 3D image samples can reflect the depth information of the photographed object, that is, its three-dimensional position and size information, the server 130 can improve the accuracy of recognition and reduce the false recognition rate by obtaining 2D and 3D image samples and then identifying the target object based on the information contained in the 2D and 3D image samples.
[0105] S320, based on the intention recognition model, predict the 2D image samples and the 3D image samples to obtain the behavioral feature map and spatial distribution feature map of the at least one target object.
[0106] In this disclosure, for each of the at least one target object, the behavioral feature map contains behavioral feature information of the target object, which can reflect the payment behavior of the target object. The spatial distribution feature map contains spatial location information of the target object in the target scene, or the spatial distribution feature map contains distance information of the target object from the camera of the client 110.
[0107] In this disclosure, the intention recognition model can be understood as a neural network model that takes 2D and 3D image samples as input and can predict the behavioral feature map and spatial distribution feature map of at least one target object in the 2D and 3D image samples. Therefore, the server 130 can use the intention recognition model to make predictions using 2D and 3D image samples to obtain the behavioral feature map and spatial distribution feature map, both of which contain behavioral feature information and spatial distribution information of each of the at least one target object.
[0108] In some embodiments, S320 may include:
[0109] S321, based on the intention recognition model, predict the 2D image sample to determine the behavioral feature map of the at least one target object.
[0110] In this disclosure, server 130 can use a intent recognition model to predict 2D image samples, thereby obtaining a behavioral feature map of at least one target object, that is, the determination of the behavioral feature map can be independent of 3D image samples.
[0111] In some embodiments, S321 may include:
[0112] S3211, Based on the intention recognition model, predict the 2D image sample to obtain a key point feature map, wherein the key point feature map contains key point information on the at least one target object.
[0113] In this disclosure, after obtaining a 2D image sample, the server 130 can predict the 2D image sample based on a intent recognition model to obtain a key point feature map. The key point feature map contains key point information of at least one target object. The key point information can be understood as the key joints of the target object in the 2D image sample, such as the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. Among these, the nose, left eye, right eye, left ear, and right ear are facial key joints, while the others, such as the left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle, are non-facial key joints. In this disclosure, compared to non-facial key joints, facial key joints can provide more reliable information for the identification of the target object.
[0114] Furthermore, the facial information collected by client 110 may be a partial or complete face image, and the completeness of the face collected by client 110 will also affect the recognition of the target object. For example, in an offline payment scenario where multiple users are queuing to pay by facial recognition, for ease of understanding, assume there are three users (i.e., three target objects) queuing to pay by facial recognition, namely target object A, target object B, and target object C. In the 2D and 3D image samples collected by client 110, target object A shows its entire face, while target object B and target object C only show half of their faces. At this time, in the key point feature map obtained by server 130, the facial key joint information belonging to target object A is the most abundant. Since the facial key joint information belonging to target object A is the most abundant in the key point feature map, the server 130 is more likely to recognize that target object A's intention attribute for the current payment task is safe.
[0115] In some embodiments, the intention recognition model includes a keypoint detection subnetwork. In this disclosure, the keypoint detection subnetwork can be understood as a network structure in the intention recognition model that has keypoint detection functionality. The keypoint detection subnetwork can be combined with other functional network structures to form the intention recognition model.
[0116] In some implementations, predicting the 2D image samples based on the intent recognition model to obtain key point feature maps includes performing the following steps through the key point detection sub-network:
[0117] S3211-a, human keypoints are extracted from the 2D image sample to obtain a human keypoint map. Server 130 can input the 2D image sample into a keypoint detection subnetwork, and use the keypoint detection subnetwork to extract human keypoints from the 2D image sample, thereby obtaining a human keypoint map. The process of server 130 extracting human keypoints can be implemented by combining an encoding network module and a decoding network module. Specifically, server 130 can use the encoding network module to extract features from the 2D image sample. Feature extraction can be implemented through at least one convolutional neural network, and this process is downsampling processing; then, at least one convolutional network and a feature upsampling processor in the decoding network module are used to output the human keypoint map. In this disclosure, the resolution of the human keypoint map obtained after the 2D image sample passes through the encoding network module and the decoding network module is consistent with the resolution of the input 2D image, and the aspect ratio of the human keypoint map is also consistent with the aspect ratio of the input 2D image. During the key point extraction process of 2D image samples, server 130 ensures that the resolution and aspect ratio of the obtained human key point map are consistent with the original 2D image samples, which can lay the foundation for subsequent feature fusion operations.
[0118] S3211-b, predict the human body key point map to obtain the key point feature map, wherein the result of each position in the key point feature map corresponds to the probability value of each position of the 2D image sample being a key point of the at least one target object.
[0119] In this disclosure, the human keypoint map obtained by server 130 after extracting keypoints from 2D image samples contains various keypoint features of the target object. At this time, server 130 can also use a keypoint detection subnetwork to classify and predict the various keypoint features contained in the human keypoint map to determine which key joint the keypoint features contained in the human keypoint map belong to, thereby obtaining a keypoint feature map.
[0120] As previously mentioned, the resolution and aspect ratio of the human keypoint map are consistent with the original 2D image sample. In this disclosure, the keypoint feature map obtained by the server 130 predicting the human keypoint map also has the same resolution and aspect ratio as the original 2D image sample. At this point, each position in the keypoint feature map can be correlated with each position in the 2D image sample, thus making the result of each position in the keypoint feature map correspond to the probability value of each position in the 2D image sample being a keypoint of at least one target object. Here, each position can be understood as each pixel in the keypoint feature map or the 2D image sample, or as a feature region composed of multiple pixels. In the keypoint feature map, the result of each position can correspond to the probability value of the human keypoint with the highest probability at each position in the 2D image sample, that is, the probability value that best represents what kind of human keypoint each position corresponds to. For example, if server 130 predicts that a certain location in the keypoint feature map has a 78% probability of being a nose, a 55% probability of being an ear, and a 35% probability of being an eye, then in the keypoint feature map, the result for that location is: the human keypoint is a nose, with a probability of 78%. Correspondingly, in the 2D image sample, the human keypoint corresponding to that location is a nose, with a probability of 78%.
[0121] In this disclosure, server 130 obtains key point feature maps by predicting 2D image samples using a key point detection subnetwork. The key point feature maps can reflect the key point information of at least one target object contained in the 2D image samples, laying the foundation for subsequently determining the behavioral feature maps of at least one target object and improving the accuracy of the behavioral feature maps.
[0122] S3212, Based on the intention recognition model, feature extraction is performed on the 2D image sample to obtain a 2D modal feature map.
[0123] In this disclosure, server 130 can also use a intent recognition model to extract features from 2D image samples independently. The intent recognition model may include commonly used network structures (such as ResNet, VGG, MobileNet, ShuffleNetV2, etc.), and use these network structures to extract features from the 2D image samples to obtain 2D modal feature maps. The 2D modal feature maps contain various feature information (such as features reflecting key points of the human body, features reflecting the background, etc.).
[0124] S3213, the behavioral feature map is determined based on the key point feature map and the 2D modal feature map, wherein the behavioral feature map enhances the information of the key points in the 2D modal feature map.
[0125] In this disclosure, after obtaining the keypoint feature map and the 2D modal feature map, the server 130 can determine the behavioral feature map based on the keypoint feature map and the 2D modal feature map. The behavioral feature map integrates the feature information of both the keypoint feature map and the 2D modal feature map. The behavioral feature map can be understood as a feature map obtained by the server 130 after enhancing the human body keypoint information contained in the 2D modal feature map using the keypoint feature map.
[0126] In some embodiments, S3213 may include: fusing the keypoint feature map and the 2D modal feature map to obtain the behavioral feature map, wherein the behavioral feature map emphasizes the behavioral feature information brought by the keypoints in the 2D modal feature map.
[0127] In this disclosure, the server 130 can determine the behavioral feature map based on the keypoint feature map and the 2D modal feature map in various ways. For example, the server 130 can concatenate the keypoint feature map and the 2D modal feature map by channel and then perform feature fusion. The feature fusion method can be to use an attention mechanism to perform feature fusion, thereby obtaining the behavioral feature map.
[0128] As mentioned earlier, a behavioral feature map can be understood as a feature map obtained by server 130 after enhancing the human body key point information contained in the 2D modal feature map using the key point feature map. The human body key point information can include the target object's key joints, such as the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. Therefore, a behavioral feature map can also be understood as a feature map obtained by server 130 after enhancing the behavioral feature information brought about by the human body key points contained in the 2D modal feature map using the key point feature map; that is, the behavioral feature map can reflect the target object's payment behavior.
[0129] In this disclosure, the server 130 obtains a behavioral feature map by fusing human keypoint detection with 2D modal features. Compared with ordinary 2D modal feature maps, the behavioral attribute features of the target object can be reflected more comprehensively and accurately, thereby reflecting the payment behavior of the target object more comprehensively and accurately.
[0130] S322, Based on the intention recognition model, predict the 2D image samples and the 3D image samples to determine the spatial distribution feature map of the at least one target object.
[0131] Unlike the method used by server 130 to determine the behavioral feature map, in this disclosure, server 130 can use an intention recognition model to predict 2D and 3D image samples, thereby obtaining a spatial distribution feature map of at least one target object. The spatial distribution feature map can reflect the spatial location information of the target object or between target objects, thus helping server 130 to more accurately determine the intention object.
[0132] In some embodiments, S322 may include:
[0133] S3221, Based on the intention recognition model, predict the 2D image sample to obtain a target segmentation probability map, wherein the target segmentation probability map contains the difference information between the background and the target object in the 2D image sample.
[0134] In this disclosure, server 130 can extract or identify features from 2D image samples from multiple dimensions to obtain feature maps with different emphases. For example, as described above, server 130 uses an intention recognition model to predict 2D image samples to obtain a key point feature map containing key point information of at least one target object. Furthermore, since 2D image samples contain not only target objects but also background information other than target objects, and this background information may interfere with server 130's identification of intention objects within the target object, and the presence of background information also increases the computational load on server 130 during the identification of intention objects within the target object. Therefore, in this disclosure, after obtaining 2D image samples, server 130 can also use an intention recognition model to predict 2D image samples to obtain a target segmentation probability map that can distinguish between the background and the target object in the 2D image sample.
[0135] In some embodiments, the intention recognition model further includes a target segmentation task subnetwork. In this disclosure, the target segmentation task subnetwork can be understood as a network structure in the intention recognition model that has the function of distinguishing between background and target objects in 2D image samples. The target segmentation task subnetwork can be combined with other functional network structures (such as the aforementioned keypoint detection subnetwork) to form the intention recognition model.
[0136] In some embodiments, predicting the 2D image samples based on the intention recognition model to obtain a target segmentation probability map includes performing the following steps through the target segmentation task sub-network:
[0137] S3221-a, Perform target segmentation on the 2D image sample to obtain target segmentation feature map.
[0138] In this disclosure, server 130 can input 2D image samples into a target segmentation task sub-network, and use the target segmentation task sub-network to extract background features and / or human features from the 2D image samples, thereby obtaining a target segmentation feature map. The process of server 130 extracting background features and / or human features can be implemented by combining an encoding network module and a decoding network module. Specifically, server 130 can use the encoding network module to extract features from the 2D image samples. Feature extraction can be implemented by at least one convolutional neural network, and this process is downsampling processing; then, at least one convolutional network and a feature upsampling processor in the decoding network module are used to output the target segmentation feature map. In this disclosure, the resolution of the target segmentation feature map obtained after the 2D image sample passes through the encoding network module and the decoding network module is consistent with the resolution of the input 2D image, and the aspect ratio of the target segmentation feature map is also consistent with the aspect ratio of the input 2D image. During the process of extracting background features and / or human features from 2D image samples, server 130 ensures that the resolution and aspect ratio of the obtained target segmentation feature map are consistent with the original 2D image samples, which can lay the foundation for subsequent feature fusion operations.
[0139] S3221-b, Based on the target segmentation feature map, obtain the target segmentation probability map, wherein the result at each position in the target segmentation probability map corresponds to the probability value of each position of the 2D image sample being the at least one target object.
[0140] In this disclosure, the target segmentation feature map obtained by server 130 after extracting background features and / or human features from 2D image samples contains background-related features and human-related features from the 2D image samples. At this time, server 130 can use the target segmentation task sub-network to perform binary classification prediction on the aforementioned features contained in the target segmentation feature map to determine the probability that each position in the target segmentation probability map is either background or human, wherein the probability of human is the probability value of at least one target object.
[0141] Furthermore, as mentioned earlier, the resolution and aspect ratio of the target segmentation probability map are consistent with the original 2D image sample. In this disclosure, the target segmentation probability map obtained by the server 130 predicting the human keypoint map also has the same resolution and aspect ratio as the original 2D image sample. At this time, each position in the target segmentation probability map can correspond to each position in the 2D image sample, thereby making the result of each position in the target segmentation probability map correspond to the probability value of each position in the 2D image sample being at least one target object. Here, each position can be understood as each pixel in the target segmentation probability map or the 2D image sample, or as a feature region composed of multiple pixels.
[0142] In this disclosure, when the server 130 performs binary classification prediction on the background and human features contained in the target segmentation feature map using the target segmentation task sub-network, it can determine the human and background regions in the 2D image sample according to the respective weights of the background and human features. For example, if the probability value of a certain location in the keypoint feature map predicted by the server 130 is 80% for a human and 20% for background, then in the target segmentation probability map, the probability value of that location being at least one target object is 80%, which can also be understood as the probability of that location being a target object being greater than the probability of it being background. After performing binary classification prediction on each location in the keypoint feature map, the server 130 can distinguish between the human and background regions in the 2D image sample.
[0143] In some embodiments, the server 130 can remove the background region and retain only the feature information of the human body region. Specifically, the feature values of the background region can be set to 0, while the feature values of the human body region can remain unprocessed.
[0144] In this disclosure, server 130 obtains a target segmentation probability map by predicting 2D image samples using a target segmentation task sub-network. The target segmentation probability map can reflect the difference between the background and the target object in the 2D image sample, laying the foundation for subsequently determining the spatial position between each target object in at least one target object, and improving the accuracy of the spatial distribution feature map.
[0145] S3222, Based on the intention recognition model, feature extraction is performed on the 3D image sample to obtain a 3D modal feature map.
[0146] In this disclosure, after obtaining a 3D image sample, the server 130 can use a intention recognition model to extract features from the 3D image, extract the depth feature information contained in the 3D image sample, and thus obtain a 3D modal feature map.
[0147] In some embodiments, S3222 may include:
[0148] S3222-a, Based on the intention recognition model, the 3D image samples are normalized to obtain a normalized 3D depth image.
[0149] In this disclosure, server 130 can generate a normalized 3D depth image based on the depth value of the face region in a 3D image sample. The specific process is as follows: First, server 130 calculates the depth value of the face region in the 3D image sample, and takes the average of the depth measurement results in the face region as the depth value of the face; then, it normalizes the depth values of all regions in the 3D image sample based on the depth value of the face.
[0150] It is important to note that server 130 generates a normalized 3D depth image based on the depth values of the facial regions in the 3D image samples. This can be understood as server 130 focusing on the depth information between at least one facial region of the target object when normalizing the 3D image samples. In payment scenarios, client 110 typically performs facial recognition (focusing on the facial information of the target object). In this disclosure, server 130 also uses the facial depth information of the target object as a reference when generating the normalized 3D depth image.
[0151] S3222-b, Based on the intention recognition model, feature extraction is performed on the normalized 3D depth image to obtain the 3D modal feature map.
[0152] In this disclosure, after obtaining a normalized 3D depth image, the server 130 can further extract features from the normalized 3D depth image using a intent recognition model. The intent recognition model may include commonly used network structures (such as ResNet, VGG, MobileNet, ShuffleNetV2, etc.), and these network structures are used to extract features from the normalized 3D depth image to obtain a 3D modal feature map. The 3D modal feature map contains depth feature information corresponding to at least one facial region of the target object, and may also contain depth feature information corresponding to regions other than the facial region (such as at least one body or background region of the target user).
[0153] S3223, Based on the target segmentation probability map and the 3D modal feature map, the spatial distribution feature map is determined, and the spatial distribution feature map enhances the spatial information of the target segmentation probability map.
[0154] In this disclosure, after obtaining the target segmentation probability map and the 3D modal feature map, the server 130 can determine the spatial distribution feature map based on the target segmentation probability map and the 3D modal feature map. The spatial distribution feature map integrates the information contained in both the target segmentation probability map and the 3D modal feature map. The spatial distribution feature map can be understood as the server 130 using the 3D modal feature map to enhance the spatial information between target objects in the target segmentation probability map; that is, the spatial distribution feature map reflects the spatial location information of at least one target object (i.e., the distance information of the target object from the camera of the client 110). For example, when the 2D image sample collected by the client 110 contains two target objects, the server 130 can determine the location of the region in the 2D image sample containing only the two target objects after filtering out the background information through the target segmentation probability map. The server 130 can also determine the depth feature information of the two target objects and the background region through the 3D modal feature map. The server 130 can also obtain a spatial distribution feature map that reflects the spatial location information of the two target objects by comprehensively considering the target segmentation probability map and the 3D modal feature map.
[0155] In some embodiments, determining the spatial distribution feature map based on the target segmentation probability map and the 3D modal feature map may include: fusing the target segmentation probability map and the 3D modal feature map to obtain the spatial distribution feature map, wherein the spatial distribution feature map emphasizes the spatial location information of at least one target object in the 2D modal feature map.
[0156] In this disclosure, the server 130 can determine the spatial distribution feature map based on the target segmentation probability map and the 3D modal feature map in various ways. For example, the server 130 can concatenate the target segmentation probability map and the 3D modal feature map by channel and then perform feature fusion. The feature fusion method can be to use an attention mechanism to perform feature fusion, thereby obtaining the spatial distribution feature map.
[0157] In this disclosure, the server 130 obtains a spatial distribution feature map by fusing the target segmentation probability map and the 3D modal feature map. Compared with the ordinary 3D modal feature map, it can more specifically reflect the distance information of the target object from the camera of the client 110 (the closer the target object is to the camera of the client 110, the higher the probability that it is the intended object), thereby improving the accuracy of target object recognition.
[0158] S330, based on the intention recognition model, predict the behavioral feature map and the spatial distribution feature map to determine the intention attribute of the at least one target object for the target transaction, the intention attribute including one of intention safety and intention high risk.
[0159] In this disclosure, after obtaining the behavioral feature map and the spatial distribution feature map, the server 130 can predict the willingness attributes of at least one target object for the target transaction based on the behavioral feature map and the spatial distribution feature map. For example, when the 2D image sample and the 3D image sample contain three target objects A, B, and C, the server 130 can predict the behavioral feature map and the spatial distribution feature map based on the willingness recognition model and determine the willingness attributes of target object A for the target transaction; the server 130 can predict the behavioral feature map and the spatial distribution feature map based on the willingness recognition model and determine the willingness attributes of target object B for the target transaction; the server 130 can also predict the behavioral feature map and the spatial distribution feature map based on the willingness recognition model and determine the willingness attributes of target object C for the target transaction.
[0160] In this disclosure, the willingness attribute reflects the security of the target object for the target transaction. Server 130 can use a binary classification prediction method to predict the willingness attribute of the initial willingness object. The willingness attribute includes either willingness security or willingness high risk. Willingness security indicates that the target object is safe for the target transaction, while willingness high risk indicates that the target object is unsafe for the target transaction. For example, when server 130 determines that the willingness attribute of target object A is willingness security, it means that target object A is safe for the target transaction, and target object A can continue to complete the payment behavior for the target transaction; however, if server 130 determines that the willingness attribute of target object A is willingness high risk, it means that target object A is unsafe for the target transaction (there is a transaction risk). At this time, server 130 can issue an error reminder through client 110, or directly suspend the current target transaction until server 130 confirms that the willingness attribute of the target object is willingness security, and then continue to complete the payment behavior for the target transaction.
[0161] In some embodiments, S330 may include:
[0162] S331, the behavioral feature map and the spatial distribution feature map are fused to obtain a multi-task fusion feature map, wherein the multi-task fusion feature map fuses the behavioral features of the 2D attributes of the target object and the spatial distribution features of the 3D attributes in the target scene; and
[0163] S332, Based on the intention recognition model, predict the multi-task fusion feature map to determine the intention attribute of the at least one target object.
[0164] In this disclosure, the server 130 can determine the multi-task fusion feature map based on the behavioral feature map and the spatial distribution feature map in various ways. For example, the server 130 can concatenate the behavioral feature map and the spatial distribution feature map by channel and then perform feature fusion. The feature fusion method can be to use an attention mechanism to perform feature fusion, thereby obtaining the multi-task feature map. The multi-task fusion feature map integrates the behavioral features of the 2D attributes of the target object in the target scene and the spatial distribution features of the 3D attributes. The server 130 can combine the behavioral features of the 2D attributes (including the payment behavior of at least one target object, especially the payment behavior reflected by the face) and the distance information of each target object from the camera of the client 110 to determine the intention attribute of at least one target object.
[0165] In this disclosure, server 130 obtains a multi-task fusion feature map by fusing behavioral feature maps and spatial distribution feature maps, and predicts the intention attributes of the target object based on the multi-task fusion feature map. Compared with the method of predicting the intention attributes of the target object based solely on behavioral feature maps or spatial distribution feature maps, this method can more comprehensively and accurately reflect the behavioral attribute characteristics of the target object, thereby reducing the false recognition rate.
[0166] In some embodiments, the intention recognition model includes a multi-task prediction subnetwork. In this disclosure, the multi-task prediction subnetwork can be understood as a network structure in the intention recognition model that has the function of predicting the intention attributes of the target object. The multi-task prediction subnetwork can be combined with other functional network structures (such as the aforementioned keypoint detection subnetwork, target segmentation task subnetwork, etc.) to form the intention recognition model.
[0167] In some embodiments, S332 may include:
[0168] S3321, Obtain the initial intention object of the target scene, wherein the initial intention object is one of the at least one target object; and
[0169] S3322, Based on the multi-task prediction sub-network, predict the multi-task fusion feature map to determine the intention attribute of the initial intention object.
[0170] In this disclosure, when the server 130 determines the intention attribute of at least one target object, it can select one target object from the at least one target object as the initial intention object. The server 130 then predicts the multi-task fusion feature map based on the multi-task prediction sub-network to obtain the intention attribute prediction result for the initial intention object, thereby determining whether it is safe and reliable to use the initial intention object as the target transaction object.
[0171] The following example illustrates the concept of three target objects, A, B, and C, contained in 2D and 3D image samples. Server 130 can use any one of target object A, B, or C as the initial intention object. For instance, server 130 can use target object A as the initial intention object, predict the multi-task fusion feature map based on the multi-task prediction sub-network, obtain the prediction result of target object A's intention attribute for the target transaction, and determine whether target object A can continue to complete the payment behavior for the target transaction based on the prediction result of target object A's intention attribute for the target transaction. Similarly, server 130 can also use target object B as the initial intention object, predict the multi-task fusion feature map based on the multi-task prediction sub-network, obtain the prediction result of target object B's intention attribute for the target transaction, and determine whether target object B can continue to complete the payment behavior for the target transaction based on the prediction result of target object B's intention attribute for the target transaction. And so on, server 130 can also use target object C as the initial intention object and perform subsequent predictions.
[0172] In some embodiments, S3321 may include:
[0173] S3321-a, determine the target scenario corresponding to the triggering of the target transaction, wherein the target scenario includes the at least one target object; and
[0174] S3321-b, Based on the initial identification model, determine one target object from the at least one target object as the initial intention object.
[0175] In this disclosure, the server 130 can obtain the initial willingness object of the target scene in various ways. For example, the server 130 can predict the 2D image sample of the target scene through an initial recognition model, and predict a target object with potential payment willingness from at least one target object contained in the 2D image sample as the initial willingness object.
[0176] In this disclosure, the initial identification model is different from the intention identification model. The role of the initial identification model can be understood as making a preliminary prediction (judgment) on the 2D image sample of the target scene, and selecting a target object as the initial intention object from at least one target object contained in the 2D image sample (at this time, the initial intention object may be safe or unsafe for the target transaction), so that the intention identification model can make a secondary prediction (judgment) on the initial intention object, thereby determining whether the initial intention object is safe for the target transaction. The role of the intention recognition model can be understood as performing a secondary prediction (judgment) on the initial prediction (judgment) result of the initial recognition model. In the secondary prediction (judgment) process, information from 2D and 3D image samples is fused (the human key point detection task is fused with the 2D modal feature map to determine the behavioral feature map of the target object, and the human target segmentation task is fused with the 3D modal feature map to determine the spatial distribution feature map of the target object, and then the behavioral feature map and the spatial distribution feature map are fused). Then, the intention attributes of the initial intention object are predicted, which can improve the accuracy of the prediction results, avoid false scans, ensure the security of the facial recognition system, and thus improve the user's security experience of facial recognition payment.
[0177] In this disclosure, after obtaining the initial intention object of the target scene, the server 130 can generate a mask image based on the position of the initial intention object in the 2D image sample, and input the mask image into the intention recognition model for comparison and recognition of at least one target object. In this disclosure, the role of the mask image can be understood as the server 130 informing the intention recognition model of an initial intention object, so that the intention recognition model can perform binary classification prediction on the initial intention object, thereby determining the intention attribute of the initial intention object.
[0178] It should be understood that, for offline payment scenarios, especially those primarily using facial recognition as the main identification method, the initial intended object includes the target object with facial information collected by the client 110. In this disclosure, the background area of the mask image generated by the server 130 based on the position of the initial intended object in the 2D image sample can be filled with 0s, and the rectangular area contained in the position of the face bounding box of the initial intended object in the acquired image can be filled with 1s. It should be noted that the resolution and size of the mask image are the same as those of the 2D image sample.
[0179] In some embodiments, S3322 may include:
[0180] S3322-a, Based on the multi-task prediction sub-network, predict the multi-task fusion feature map to determine the intention classification probability value of the initial intention object; and
[0181] S3322-b, compare the intention classification probability value with the preset intention classification probability threshold, and determine the intention attribute of the initial intention object based on the comparison result.
[0182] In this disclosure, during the prediction of the multi-task fusion feature map based on the multi-task prediction sub-network, the server 130 employs a binary classification prediction method to predict the intention attributes of the initial intention object, thereby obtaining the intention classification probability value of the initial intention object. This intention classification probability value includes a safe intention probability value and a high-risk intention probability value, and the sum of the safe intention probability value and the high-risk intention probability value is 1. Therefore, the higher the safe intention probability value, the lower the high-risk intention probability value.
[0183] In this disclosure, server 130 may preset a willingness classification probability threshold. For example, server 130 may preset a willingness safety probability threshold and compare the willingness safety probability value with the willingness safety probability threshold. If the willingness safety probability value is greater than the willingness safety probability threshold, the willingness attribute of the initial willingness object is determined to be willingness safety; otherwise, if the willingness safety probability value is less than the willingness safety probability threshold, the willingness attribute of the initial willingness object is determined to be willingness high risk.
[0184] It should be understood that the willingness classification probability threshold can include a willingness safety probability threshold and a willingness high-risk probability threshold, and the sum of the willingness safety probability threshold and the willingness high-risk probability threshold is 1. When server 130 presets a willingness safety probability threshold, it is equivalent to server 130 also presets a willingness high-risk probability threshold. Therefore, server 130 can also compare the willingness high-risk probability value with the willingness high-risk probability threshold. If the willingness high-risk probability value is greater than the willingness high-risk probability threshold (equivalent to the willingness safety probability value being less than the willingness safety probability threshold), the willingness attribute of the corresponding initial willingness object is willingness high-risk; conversely, if the willingness high-risk probability value is less than the willingness high-risk probability threshold (equivalent to the willingness safety probability value being greater than the willingness safety probability threshold), the willingness attribute of the corresponding initial willingness object is willingness safe.
[0185] In this disclosure, server 130 performs comprehensive processing on 2D and 3D image samples, fuses human keypoint detection tasks with 2D modal feature maps to determine the behavioral feature map of the target object, and fuses human target segmentation tasks with 3D modal feature maps to determine the spatial distribution feature map of the target object. Then, it fuses the behavioral feature map and the spatial distribution feature map to predict at least one intention attribute of the target object. This enables accurate prediction of the target object's intention to use facial recognition payment in offline public places, avoids false scans, ensures the security of the facial recognition system, and improves the user's security experience of facial recognition payment.
[0186] In some embodiments, this disclosure also provides a training method 400 for an intention recognition model. Figure 5 A flowchart 400 of a method for training an intention recognition model according to some embodiments of the present disclosure is shown. The method 400 may include:
[0187] S410, obtain 2D image samples and 3D image samples of the target scene, wherein each of the 2D image samples and the 3D image samples contains at least one target object.
[0188] S420, based on a preset intention recognition model, perform feature prediction on the 2D image samples and the 3D image samples to obtain key point feature maps, target segmentation probability maps, and the intention attributes of at least one target object.
[0189] In this disclosure, server 130 can predict 2D image samples based on a preset intention recognition model to obtain a key point feature map, which contains key point information of multiple target objects in the 2D image sample.
[0190] Server 130 can predict 2D image samples based on a preset intention recognition model to obtain a target segmentation probability map, which contains information on the difference between the background and the target object in the 2D image sample.
[0191] Server 130 can fuse keypoint feature maps and 2D modal feature maps to obtain a behavior feature map, which emphasizes the behavior feature information brought by keypoints in the 2D modal feature map. Server 130 can also fuse target segmentation probability maps and 3D modal feature maps to obtain a spatial distribution feature map, which emphasizes the spatial location information of the missing target object in the 2D modal feature map. Furthermore, server 130 can fuse the behavior feature map and spatial distribution feature map based on a preset intention recognition model to obtain a multi-task fusion feature map, which integrates the behavior features of 2D attributes and the spatial distribution features of 3D attributes of target objects in the target scene. Further, server 130 can perform classification prediction on the multi-task fusion feature map based on the preset intention recognition model to obtain the intention attributes of each target object in at least one target object.
[0192] S430, the preset intention recognition model is trained based on the key point feature map, the target segmentation probability map and the intention attribute to obtain the trained intention recognition model.
[0193] In this disclosure, after obtaining the key point feature map, the target segmentation probability map, and the intention attribute of the at least one target object, the server 130 can train a preset intention recognition model based on the key point feature map, the target segmentation probability map, and the intention attribute of the at least one target object to obtain the trained intention recognition model.
[0194] In some embodiments, the preset intention recognition model includes a preset keypoint detection subnetwork, a preset target segmentation task subnetwork, and a preset multi-task prediction subnetwork. In this disclosure, the server 130 trains the preset intention recognition model, primarily by training the three subnetworks included in the preset intention recognition model (namely, the keypoint detection subnetwork, the target segmentation task subnetwork, and the multi-task prediction subnetwork). The keypoint detection subnetwork is related to a keypoint feature map, the target segmentation task subnetwork is related to a target segmentation probability map, and the multi-task prediction subnetwork is related to the intention attribute of at least one target object. At this time, S430 may include:
[0195] S431, determine the first feature loss information of the preset keypoint detection sub-network based on the keypoint feature map. In this disclosure, after obtaining the keypoint feature map, the server 130 can train the preset keypoint detection sub-network based on the keypoint feature map. The server 130 can determine the first feature loss information (L1) based on the keypoint feature map. The first feature loss information (L1) can be understood as the loss information formed by the difference between the predicted keypoint feature map and the keypoint feature information actually contained in the 2D image sample. The first feature loss information (L1) is used to constrain the keypoint feature map predicted by the keypoint detection sub-network to be consistent with the keypoint feature information actually contained in the 2D image sample.
[0196] In some embodiments, S431 may include:
[0197] S4311, Obtain the Gaussian heatmap of key points of the 2D image sample; and
[0198] S4312, compare the Gaussian heatmap of the key points with the feature map of the key points to determine the first feature loss information.
[0199] In this disclosure, there are multiple ways to obtain the key point feature information actually contained in the 2D image sample. For example, server 130 can draw a Gaussian heatmap of key points on the 2D image sample, which can reflect the key point feature information actually contained in the 2D image sample. After obtaining the Gaussian heatmap of key points of the 2D image sample, server 130 can compare the Gaussian heatmap with the key point feature map to determine the first feature loss information (L1). The first feature loss function can be of various types, such as the cross-entropy loss function or other loss functions that can be used to determine the first feature loss information (L1), etc.
[0200] S432, determine the second feature loss information of the preset target segmentation task sub-network based on the target segmentation probability map. In this disclosure, after obtaining the target segmentation probability map, the server 130 can train the preset target segmentation task sub-network based on the target segmentation probability map. The server 130 can determine the second feature loss information (L2) based on the target segmentation probability map. The second feature loss information (L2) can be understood as the loss information formed by the difference between the predicted target segmentation probability map and the actual human body region-background region division in the 2D image sample. The second feature loss information (L2) is used to constrain the target segmentation probability map predicted by the target segmentation task sub-network to be consistent with the actual human body region-background region division in the 2D image sample.
[0201] In some embodiments, S432 may include:
[0202] S4321, Obtain the target segmentation annotation map of the 2D image sample; and
[0203] S4322, compare the target segmentation annotation map with the target segmentation probability map to determine the second feature loss information.
[0204] In this disclosure, the target segmentation annotation map can be understood as an atlas obtained by annotating the actual human body region and background region in a 2D image sample. After obtaining the target segmentation annotation map of the 2D image sample, the server 130 can compare the target segmentation annotation map with the target segmentation probability map to determine the second feature loss information (L2). The second feature loss function can be of various types, such as the cross-entropy loss function or other loss functions that can be used to determine the second feature loss information (L2), etc.
[0205] S433, the third feature loss information of the preset multi-task prediction sub-network is determined based on the intention attributes of the at least one target object. In this disclosure, after obtaining the intention attributes of the target object contained in the 2D image sample or 3D image sample, the server 130 can train the preset multi-task prediction sub-network based on the intention attributes of the target object. The server 130 can determine the third feature loss information (L3) based on the intention attributes of the target object. The third feature loss information (L3) can be understood as the loss information formed by the difference between the predicted intention attributes of the target object and the actual intention attributes of the target object. The third feature loss information (L3) is used to constrain the intention attributes of the target object predicted by the multi-task prediction sub-network to be consistent with the actual intention attributes of the target object.
[0206] In some embodiments, S433 may include:
[0207] S4331, obtain the original annotation attributes of the at least one target object; and
[0208] S4332, compare the original labeled attributes with the intention attributes to determine the third feature loss information.
[0209] In this disclosure, the original labeled attributes can be understood as the actual intention attributes of the target object, which can include intention and non-intention. For ease of comparison, the server 130 uses the intention recognition model to perform binary classification of the target object's intention attributes according to the magnitude of the intention value, that is, classifying each target object into an intention object and a non-intention object based on the magnitude of the intention value. Specifically, when the 2D image sample and the 3D image sample contain two or more target objects, only one target object can be identified as an intention object, and the remaining target objects are non-intention objects. After obtaining the original labeled attributes, the server 130 can compare the original labeled attributes with the intention attributes to determine the third feature loss information (L3). The type of third feature loss function can be varied, such as including the cross-entropy loss function or other loss functions that can be used to determine the third feature loss information (L3), etc.
[0210] S434, converge the preset intention recognition model based on the first feature loss information, the second feature loss information and the third feature loss information to obtain the trained intention recognition model.
[0211] In this disclosure, after obtaining the first feature loss information (L1), the second feature loss information (L2), and the third feature loss information (L3), the server 130 can integrate the loss information of the first feature loss information (L1), the second feature loss information (L2), and the third feature loss information (L3) to converge the preset intention recognition model, thereby obtaining the trained intention recognition model.
[0212] In some embodiments, S434 may include:
[0213] S4341, the first feature loss information, the second feature loss information and the third feature loss information are fused to obtain comprehensive loss information.
[0214] In this disclosure, the server 130 can fuse the first feature loss information (L1), the second feature loss information (L2), and the third feature loss information (L3) in various ways. For example, the server 130 can directly add the first feature loss information (L1), the second feature loss information (L2), and the third feature loss information (L3) to obtain the comprehensive loss information (L), i.e., L = L1 + L2 + L3.
[0215] In some embodiments, S4341 may include:
[0216] S4341-a, obtain the weight attributes of the first feature loss information, the second feature loss information, and the third feature loss information respectively; and
[0217] S4341-b, the first feature loss information, the second feature loss information, and the third feature loss information are weighted based on the weight attributes to determine the comprehensive loss information.
[0218] In this disclosure, server 130 can also obtain the preset loss weights of the first feature loss information (L1), the second feature loss information (L2), and the third feature loss information (L3), and weight the first feature loss information (L1), the second feature loss information (L2), and the third feature loss information (L3) based on their respective preset loss weights. After adding the weighted first feature loss information (L1), the second feature loss information (L2), and the third feature loss information (L3), the comprehensive loss information L is obtained, that is, L = a*L1 + b*L2 + c*L3, where a is the preset loss weight of the first feature loss information (L1), b is the preset loss weight of the second feature loss information (L2), and c is the preset loss weight of the third feature loss information (L3).
[0219] In some embodiments, the server 130 may also obtain the comprehensive loss information L in the following manner, namely L=L3+λ(L1+L2), where λ is the balance coefficient. In practical applications, λ<1, so as to ensure that the training of the intention recognition network is more inclined to the learning of the multi-task prediction sub-network.
[0220] S4342, Based on the comprehensive loss information, the preset intention recognition model is converged to obtain the trained intention recognition model.
[0221] In this disclosure, after obtaining the comprehensive loss information, the server 130 can converge a preset intention recognition model based on the comprehensive loss information, thereby obtaining a trained intention recognition model. There are various ways to converge the intention recognition model based on the comprehensive loss information. For example, the server 130 can use a gradient descent algorithm to update the network parameters of the preset intention recognition model based on the comprehensive loss information until the preset intention recognition model converges, thereby obtaining a trained intention recognition model. Alternatively, other parameter update algorithms can be used to update the preset intention recognition model based on the comprehensive loss information until the preset intention recognition model converges, thereby obtaining a trained intention recognition model, and so on.
[0222] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this disclosure is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this disclosure and are within the spirit and scope of the exemplary embodiments of this disclosure.
[0223] Furthermore, certain terms used in this disclosure have been used to describe embodiments of this disclosure. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this disclosure. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this disclosure do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this disclosure.
[0224] It should be understood that in the foregoing description of the embodiments of this disclosure, various features are sometimes combined in a single embodiment, drawing, or description for the purpose of simplifying the disclosure and to aid in understanding a feature. Alternatively, various features may be distributed across multiple embodiments of this disclosure. However, this does not mean that the combination of these features is necessary, and those skilled in the art may extract some features as separate embodiments when reading this disclosure. That is, the embodiments in this disclosure can also be understood as an integration of multiple sub-embodiments. It is also possible for each sub-embodiment to contain fewer features than all of the features of a single foregoing disclosed embodiment.
[0225] Each patent, patent application, publication of a patent application, and other material such as articles, books, specifications, publications, documents, articles, etc., referenced in this disclosure, except for any related historical prosecution documents, any identical historical prosecution documents that may be inconsistent with or conflict with this disclosure, or any that may have a limiting effect on the widest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter in connection with this disclosure. Furthermore, in the event of any inconsistency or conflict between the description, definition, and / or use of terms related to any included material and those relating to this disclosure, the terms used in this disclosure shall prevail.
Claims
1. A method for identifying a target object, comprising: Obtain 2D image samples and 3D image samples of the target scene, wherein each of the 2D image samples and the 3D image samples contains at least one target object; Based on the intention recognition model, the 2D image samples are predicted to determine the behavioral feature map of the at least one target object; Based on the intention recognition model, the spatial distribution feature map of the at least one target object is determined by predicting the 2D image samples and the 3D image samples. Based on the intention recognition model, the 2D image sample is predicted to obtain a target segmentation probability map, which contains the difference information between the background and the target object in the 2D image sample; Based on the intention recognition model, feature extraction is performed on the 3D image samples to obtain a 3D modal feature map; The spatial distribution feature map is determined based on the target segmentation probability map and the 3D modal feature map, wherein the spatial distribution feature map enhances the spatial information of the target segmentation probability map; and Based on the intention recognition model, the behavioral feature map and the spatial distribution feature map are predicted to determine the intention attributes of the at least one target object for the target transaction. The intention attributes include either intention to be safe or intention to be high-risk.
2. The identification method as described in claim 1, wherein, The step of predicting the 2D image samples based on the intention recognition model to determine the behavioral feature map of the at least one target object includes: Based on the intention recognition model, the 2D image sample is predicted to obtain a key point feature map, which contains key point information on at least one target object. Based on the intent recognition model, feature extraction is performed on the 2D image samples to obtain a 2D modal feature map; and The behavioral feature map is determined based on the key point feature map and the 2D modal feature map, and the behavioral feature map enhances the information of the key points in the 2D modal feature map.
3. The identification method as described in claim 2, wherein, The intention recognition model includes a keypoint detection subnetwork; and The step of predicting the 2D image samples based on the intent recognition model to obtain key point feature maps includes using the key point detection sub-network: Human key points are extracted from the 2D image samples to obtain a human key point map, and The key point map of the human body is predicted to obtain the key point feature map. The result of each position in the key point feature map corresponds to the probability value that each position of the 2D image sample is a key point of the at least one target object.
4. The identification method as described in claim 2, wherein, Determining the behavioral feature map based on the keypoint feature map and the 2D modal feature map includes: The key point feature map and the 2D modal feature map are fused to obtain the behavior feature map, which emphasizes the behavior feature information brought by the key points in the 2D modal feature map.
5. The identification method as described in claim 1, wherein, The intention recognition model also includes a target segmentation task subnetwork; and The step of predicting the 2D image samples based on the intent recognition model to obtain a target segmentation probability map includes using the target segmentation task sub-network: The 2D image samples are segmented to obtain target segmentation feature maps, and Based on the target segmentation feature map, the target segmentation probability map is obtained, and the result at each position in the target segmentation probability map corresponds to the probability value that each position of the 2D image sample is the at least one target object.
6. The identification method as described in claim 1, wherein, The step of extracting features from the 3D image samples based on the intention recognition model to obtain a 3D modal feature map includes: The intention recognition model is used to normalize the 3D image samples to obtain a normalized 3D depth image; and Based on the intention recognition model, feature extraction is performed on the normalized 3D depth image to obtain the 3D modal feature map.
7. The identification method as described in claim 1, wherein, The step of determining the spatial distribution feature map based on the target segmentation probability map and the 3D modal feature map includes: The target segmentation probability map and the 3D modal feature map are fused to obtain the spatial distribution feature map, which highlights the spatial location information of at least one target object in the 2D modal feature map.
8. The identification method as described in claim 1, wherein, The step of predicting the behavioral feature map and the spatial distribution feature map based on the intention recognition model to determine the intention attribute of the at least one target object includes: The behavioral feature map and the spatial distribution feature map are fused to obtain a multi-task fusion feature map, which integrates the behavioral features of the 2D attributes of the target object and the spatial distribution features of the 3D attributes in the target scene; and Based on the intention recognition model, the multi-task fusion feature map is predicted to determine the intention attributes of at least one target object.
9. The identification method as described in claim 8, wherein, The intention recognition model includes a multi-task prediction subnetwork; as well as The step of predicting the multi-task fusion feature map based on the intention recognition model to determine the intention attribute of the at least one target object includes: Obtain the initial intention object of the target scene, wherein the initial intention object is one of the at least one target object, and The multi-task prediction sub-network is used to predict the multi-task fusion feature map to determine the intention attributes of the initial intention object.
10. The identification method as described in claim 9, wherein, The initial intention object for obtaining the target scenario includes: Determine the target scenario corresponding to the triggering of the target transaction, wherein the target scenario includes the at least one target object; and Based on the initial identification model, a target object is determined from the at least one target object as the initial intention object.
11. The identification method as described in claim 9, wherein, The step of predicting the multi-task fusion feature map based on the multi-task prediction sub-network to determine the intention attributes of the initial intention object includes: Based on the multi-task prediction subnetwork, the multi-task fusion feature map is predicted to determine the initial intention classification probability value of the intention object; and The intention classification probability value is compared with a preset intention classification probability threshold, and the intention attribute of the initial intention object is determined based on the comparison result.
12. The identification method as described in claim 1, wherein, The training process of the intention recognition model includes: Obtain 2D image samples and 3D image samples of the target scene, wherein each of the 2D image samples and the 3D image samples contains at least one target object; Based on a preset intention recognition model, feature prediction is performed on the 2D image samples and the 3D image samples to obtain key point feature maps, target segmentation probability maps, and intention attributes of at least one target object; and The preset intention recognition model is trained based on the key point feature map, the target segmentation probability map, and the intention attribute to obtain the trained intention recognition model.
13. The identification method as described in claim 12, wherein, The preset intention recognition model includes a preset key point detection subnetwork, a preset target segmentation task subnetwork, and a preset multi-task prediction subnetwork. as well as The step of training the preset intention recognition model based on the key point feature map, the target segmentation probability map, and the intention attribute includes: Based on the key point feature map, the first feature loss information of the preset key point detection sub-network is determined. The second feature loss information of the preset target segmentation task sub-network is determined based on the target segmentation probability map. The third feature loss information of the preset multi-task prediction sub-network is determined based on the intention attributes of the at least one target object, and The preset intention recognition model is converged based on the first feature loss information, the second feature loss information, and the third feature loss information to obtain the trained intention recognition model.
14. The identification method as described in claim 13, wherein, The step of determining the first feature loss information of the preset keypoint detection sub-network based on the keypoint feature map includes: Obtain the Gaussian heatmap of key points of the 2D image sample; and The Gaussian heatmap of the key points is compared with the feature map of the key points to determine the first feature loss information.
15. The identification method as described in claim 13, wherein, The step of determining the second feature loss information of the preset target segmentation task sub-network based on the target segmentation probability map includes: Obtain the target segmentation annotation map of the 2D image sample; and The target segmentation annotation map is compared with the target segmentation probability map to determine the second feature loss information.
16. The identification method as described in claim 13, wherein, The step of determining the third feature loss information of the preset multi-task prediction sub-network based on the intention attributes of the at least one target object includes: Obtain the original annotation attributes of the at least one target object; and The original labeled attributes are compared with the intention attributes to determine the third feature loss information.
17. The identification method as described in claim 13, wherein, The step of converging the preset intention recognition model based on the first feature loss information, the second feature loss information, and the third feature loss information to obtain the trained intention recognition model includes: The first feature loss information, the second feature loss information, and the third feature loss information are fused to obtain comprehensive loss information; and The preset intention recognition model is converged based on the comprehensive loss information to obtain the trained intention recognition model.
18. The identification method as described in claim 17, wherein, The step of fusing the first feature loss information, the second feature loss information, and the third feature loss information to obtain comprehensive loss information includes: Obtain the weight attributes of the first feature loss information, the second feature loss information, and the third feature loss information respectively; and The first feature loss information, the second feature loss information, and the third feature loss information are weighted based on the weight attributes to determine the comprehensive loss information.
19. A target object identification system, comprising: At least one storage medium, including at least one instruction set, for implementation analysis of the target object identification method; as well as At least one processor is communicatively connected to the at least one storage medium. When the system is running, the at least one processor reads the at least one instruction set and executes the method of any one of claims 1-18 according to the instructions of the at least one instruction set.