A test method for detecting living body and related device

By using embodied intelligence technology and VLA models, combined with image sensors and pose sensors, a robotic arm is used to grasp target objects to perform liveness detection testing tasks, solving the problems of low efficiency and high labor costs in existing technologies, and achieving efficient and accurate automated testing.

CN122493537APending Publication Date: 2026-07-31ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG DAHUA TECH CO LTD
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing liveness detection testing processes are inefficient and costly in terms of manpower, and embodied intelligence technology cannot accurately and efficiently implement liveness detection testing.

Method used

Employing embodied intelligence technology, this method utilizes a VLA model combined with image and pose sensors to perform liveness detection testing by using a robotic arm to grasp target objects. Grasping markers are set to improve accuracy and efficiency.

Benefits of technology

It improves the accuracy and efficiency of the liveness detection testing process, automates the testing process, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493537A_ABST
    Figure CN122493537A_ABST
Patent Text Reader

Abstract

This application provides a liveness detection testing method and related apparatus. The liveness detection testing method is applied to a task execution object, which includes a main body and an execution unit. The main body is equipped with a first-view image sensor, and the execution unit is equipped with a second-view image sensor. The task execution object also includes a pose sensor. The testing method includes: receiving text description information and sensor information corresponding to the liveness detection test task; the sensor information including image information and pose information; using a VLA model to determine the motion trajectory of the task execution object based on the text description information and sensor information; controlling the task execution object to grasp a target object based on the motion trajectory to perform the liveness detection test task; the target object includes a grasping marker. This application, combined with the characteristics of the liveness detection test task, sets a grasping marker on the target object, and grasps the target object based on the grasping marker, thus accurately and efficiently realizing the liveness detection test process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embodied intelligence technology, and in particular to a testing method and related apparatus for liveness detection. Background Technology

[0002] Liveness detection and anti-spoofing technologies are typically combined with biometric technologies such as facial recognition, palm print recognition, and fingerprint recognition, and deployed in access control systems, turnstiles, and various apps. This effectively prevents criminals from using other people's facial, palm print, or fingerprint photos to unlock access control systems, electronic devices, social media accounts, financial accounts, etc., thus avoiding various losses.

[0003] Typically, devices equipped with liveness detection and anti-spoofing technologies cannot be directly connected to the video stream for testing. Instead, manual testing in real-world scenarios is required. For example, testers might hold a piece of paper with biometric information, or use mobile phones, tablets, or other electronic devices in front of a turnstile (hereinafter, we'll use a turnstile as an example to represent electronic devices with liveness detection capabilities) and move them in various directions to simulate real-person behavior, attempting to pass liveness detection using photos or videos of others. This process requires constantly and mechanically replacing test materials and repeating actions, which is not only inefficient but also incurs high labor costs.

[0004] Currently, many scenarios, such as storage and retrieval tasks, have proposed embodied intelligence technology, which is used to achieve these tasks using robotic arms. However, in the testing of liveness detection, simple embodied intelligence technology cannot accurately and efficiently implement the liveness detection testing process. Summary of the Invention

[0005] This application mainly provides a liveness detection testing method and related apparatus, which can accurately and efficiently realize the liveness detection testing process by utilizing embodied intelligence technology.

[0006] To solve the above-mentioned technical problems, the first technical solution adopted in this application is: to provide a liveness detection testing method applied to a task execution object, the task execution object including a main body and an execution unit, the main body being provided with a first-view image sensor and the execution unit being provided with a second-view image sensor, and the task execution object also including a pose sensor, the testing method including: Receive text description information and sensor information corresponding to the liveness detection test task. The sensor information includes image information and pose information. The VLA model is used to determine the motion trajectory of the task execution object based on textual description information and sensor information; The task execution object is based on motion trajectory control to grasp the target object to perform a liveness detection test task; wherein, the target object includes a grasping marker.

[0007] In one embodiment, the target object is a piece of paper containing test material, and the gripping mark is set on the side of the paper opposite to the test material, and the gripping mark is a different color from the paper.

[0008] In one embodiment, the execution unit includes a suction cup, and the task execution object grasping the target object includes: the suction cup adsorbing and grasping the mark.

[0009] In one embodiment, the task execution object grasps the target object based on the motion trajectory control to perform a liveness detection test task, including: Control the task execution object to move to the corresponding position of the target object, and make the suction cup adsorb and grasp the mark; The control task execution object moves the target object to the test area of ​​the device under test and moves the target object according to a preset trajectory, so that the device under test can identify the test material in the target object from multiple different angles; The control task execution object places the target object at the target location.

[0010] In one embodiment, controlling the task execution object to move to the corresponding position of the target object and causing the suction cup to adhere to the grasping mark includes: Move the task execution object above the target object; Control the task execution object to move downwards so that the suction cup attaches to the gripping mark; Activate the suction cup to allow it to adhere to and grasp the marker; Controlling the task execution object to place the target object at the target location includes: The task execution object is controlled to move the target object to the target location and close the suction cup, placing the target object at the target location.

[0011] In one embodiment, the task execution object includes a first robotic arm and a second robotic arm, the first robotic arm including a main body and an execution part; Before using the VLA model to determine the motion trajectory of the task execution object based on textual description information and sensor information, the following steps are included: The first robotic arm is controlled based on the simulated trajectory information, and image sample data and pose sample data of the first robotic arm are collected to obtain target sample data; wherein, the simulated trajectory information is generated when the second robotic arm simulates a liveness detection test task; The VLA model is obtained by training a general task execution model using target sample data; the general task execution model is obtained by training the initial VLA model using mixed sample data corresponding to multiple types of tasks.

[0012] In one embodiment, the VLA model is used to determine the motion trajectory of the task execution object based on textual description information and sensor information, including: Textual description information and sensor information are processed to obtain multimodal fusion data; Multimodal fusion data is used as input to the VLA model. The VLA model is then used to process the multimodal fusion data to obtain encoded features. These encoded features are then further processed to obtain the action trajectory of the task execution object.

[0013] In one embodiment, textual description information and sensor information are processed to obtain multimodal fusion data, including: Feature transformation is performed on text description information and sensor information respectively to obtain text data features and sensor data features; Text data features and sensor data features are concatenated to obtain multimodal fusion data.

[0014] To solve the above-mentioned technical problems, the second technical solution adopted in this application is: to provide an electronic terminal, which includes a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory, and the processor being used to execute program data to implement the steps in the above-mentioned liveness detection test method.

[0015] To solve the above-mentioned technical problems, the third technical solution adopted in this application is: to provide a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, to implement the steps in the above-mentioned liveness detection test method.

[0016] The beneficial effects of this application are as follows: Unlike existing technologies, the provided liveness detection testing method is applied to a task execution object, which includes a main body and an execution unit. The main body is equipped with a first-view image sensor, and the execution unit is equipped with a second-view image sensor. The task execution object also includes a pose sensor. The testing method includes: receiving text description information and sensor information corresponding to the liveness detection test task; using a VLA model to determine the motion trajectory of the task execution object based on the text description information and sensor information; and controlling the task execution object to grasp a target object to perform the liveness detection test task based on the motion trajectory. The target object includes a grasping marker. This application, combining the characteristics of the liveness detection test task, sets a grasping marker on the target object, and then the task execution object grasps the target object based on this grasping marker to perform the liveness detection test task. This allows for accurate and efficient implementation of the liveness detection testing process. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic flowchart of an embodiment of the liveness detection test method provided in this application; Figure 2 yes Figure 1 A flowchart illustrating an embodiment of step S13; Figure 3 This is a schematic diagram of the structure of an embodiment of the task execution object provided in this application; Figure 4 This is a schematic diagram of another embodiment of the task execution object provided in this application; Figure 5 This is a schematic diagram of an embodiment of the motion trajectory of the liveness authentication task provided in this application; Figure 6 This is a schematic diagram of the framework of an embodiment of the electronic terminal provided in this application; Figure 7 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0019] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0020] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0021] In this article, the term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "more" in this article means two or more objects.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0023] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0024] The liveness detection testing method provided in this application embodiment can be implemented by a server or terminal alone, or by a server and terminal working together. In some embodiments, the terminal or server can implement the liveness detection testing method provided in this application embodiment by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as a client that supports virtual scenes, such as a game APP; it can also be a mini-program, that is, a program that only needs to be downloaded into a browser environment to run; or it can be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.

[0025] To enable those skilled in the art to better understand the technical solution of this application, the following describes in further detail a liveness detection test method provided by this application in conjunction with the accompanying drawings and specific embodiments.

[0026] Please see Figure 1 This is a flowchart illustrating the first embodiment of the liveness detection testing method of this application. Specifically, the liveness detection testing method of this application is applied to a task execution object, which is an embodied intelligent robot, including but not limited to a fixed-base robotic arm, a wheeled robot, and a humanoid robot. The task execution object includes a main body and an execution part. The main body refers to the main body of the task execution object, and the execution part refers to the end effector. Taking a fixed-base robotic arm as an example, combined with... Figure 3 The main body is shown as P1, and the execution unit is shown as P2. The main body P1 is equipped with a first-view image sensor, and the execution unit P2 is equipped with a second-view image sensor. Specifically, the first view is a top-down view; the top-down view image sensor in the main body P1 can acquire a global image of the target object in real time. The second view is multi-view; since the pose of the execution unit P2 changes during task execution, multi-view image sensors can acquire image data from various perspectives. Furthermore, the task execution object also includes a pose sensor, which can acquire various pose information of the task execution object during task execution.

[0027] Based on the aforementioned task execution objects, the liveness detection test method of this application specifically includes: Step S11: Receive the text description information and sensor information corresponding to the liveness detection test task. The sensor information includes image information and pose information.

[0028] Specifically, when performing a liveness detection test task, the system receives a text description of the task, such as "Grab the paper from the left basket using the red marker." Here, "red marker" is the grab marker, "paper" is the target object, and "left basket" indicates the target object's location. This accurate text description serves as guidance for the task execution.

[0029] To perform this task, the main body P1 of the task execution object is equipped with a top-view camera, and the execution unit P2 is equipped with a multi-view camera. During the execution of the task, images are acquired through the cameras to obtain image information; and / or, the pose information of the task execution object is detected through a pose sensor installed on the task execution object, and the image information and pose information are used as sensor data.

[0030] Step S12: Use the VLA model to determine the motion trajectory of the task execution object based on text description information and sensor information.

[0031] It should be noted that model training is required before using the VLA model. Specifically, in conjunction with... Figure 4 The task execution objects include the first robotic arm A and the second robotic arm B. The first robotic arm A includes a main body P1 and an execution part P2.

[0032] In one specific embodiment, the first robotic arm is controlled based on the simulated trajectory information, and image sample data and pose sample data of the first robotic arm are collected to obtain target sample data; wherein, the simulated trajectory information is generated when the second robotic arm simulates a liveness detection test task.

[0033] Specifically, when collecting target sample data, since the general task execution model does not yet have the ability to autonomously execute the target task, this application sets up a robot with a master-slave robotic arm structure, such as... Figure 4 As shown, the black robotic arm is the second robotic arm B, which acts as the master arm; the white robotic arm is the first robotic arm A, which acts as the slave arm. The black and white robotic arms are linked and have identical degrees of freedom. The second robotic arm B is manually dragged to perform a liveness detection test task. The angles of each joint of the second robotic arm B, i.e., the simulated trajectory information, are acquired in real time via serial communication. Then, the joint angles (simulated trajectory information) are written to the servos of each joint of the first robotic arm A via serial port. Based on the simulated trajectory information, the first robotic arm A is driven to move to the exact same pose as the second robotic arm B. During this process, images from a top-down camera and a wrist camera are simultaneously recorded. The camera images (i.e., image sample data) are time-stamped and aligned with the robotic arm joint angle information (i.e., pose sample data) to obtain the target sample data.

[0034] After obtaining the target sample data, the general task execution model is trained using the target sample data to obtain the VLA model; the general task execution model is obtained by training the initial VLA model using mixed sample data corresponding to multiple types of tasks.

[0035] It should be noted that the mixed sample data includes multiple first sample data sets, each corresponding to a type of task, and each first sample data set includes multiple sets of first sample data. The first sample data includes text data and sensor data corresponding to the type of task. Sensor data includes, for example, visual data acquired by a vision sensor and / or actuator pose information acquired by an encoder.

[0036] In one embodiment, various task types include sorting, folding, and storage tasks. Text data and sensor data are collected for the sorting task. For example, the text data in the sorting task might be "picking ping-pong balls out of the basket." Visual data from the sensor data might be a top-down view of the table where the ping-pong balls are placed, which can be captured by a camera mounted on the actuator body and / or end effector. The actuator's pose information from the sensor data includes the angles of each joint of the actuator during the sorting task, which can be detected using sensors mounted on the robot. Text data, visual data, and actuator pose information are collected for the folding task to obtain a second set of first sample data. For example, the text data in the folding task might be "folding the socks on the table and storing them in the storage box on the table." Visual data might be a top-down view of the table where the socks and storage box are placed, which can also be captured by a camera. The actuator's pose information includes the angles of each joint during the folding task. Text data, visual data, and actuator pose information are collected for the storage task to obtain a third set of first sample data. The Chinese data for the organization task is, for example, "Put the stationery on the desktop into the pencil case." The visual data is also a top-down view of the desktop, and the actuator's pose information includes the angles of each joint when performing the organization task. Understandably, the actuator's motion trajectory differs depending on the task being performed, and the angles of the joints corresponding to each movement during its motion are also different.

[0037] Through the above process, we can collect the first sample data sets corresponding to three different types of tasks; and use the mixed sample data set composed of these three first sample data sets to train the initial VLM model to obtain a general task execution model.

[0038] When training the initial VLM model, loss calculation is required, which outputs the continuous future actions A. tBy combining the actual action trajectories with the calculated flow matching loss function, the gradient is calculated and directly backpropagated to update the network parameters, ultimately resulting in a general task execution model. It should be noted that the actual action trajectories are the action trajectories represented by the first sample data in each group when collecting mixed sample data, and can be fed into the model through annotation.

[0039] This application obtains a general task execution model through pre-training. The purpose of pre-training is to expose the model to various tasks, enabling it to acquire widely applicable general physical capabilities. Therefore, the mixed sample data should cover robots with different structures and as many tasks as possible, and each task should cover various behaviors. The motion expert model unifies the motion dimension of all robot structures to D dimension. For the remaining vectors in the network input where the robot's degrees of freedom are less than D dimensions, zeros are used to fill them. Therefore, the D value should be greater than the full-body degrees of freedom of most commercially available robots, for example, D=32 dimensions. Similarly, the network output only calculates the loss of the robot's effective motion output dimension (for example, when a 6-DOF robotic arm is input, the first 6 dimensions are filled, and the remaining 26 dimensions are all filled with 0s; at the same time, the network predicts that only the first 6 dimensions are taken as the effective motion sequence).

[0040] The general task execution model is obtained through the above process. The target sample data of the liveness detection test task collected in this application is used to train the general task execution model, thereby obtaining the VLM model.

[0041] Specifically, general task execution models lack the ability to perform robustly and efficiently in specific tasks such as liveness detection testing. Traditionally trained, fixed-trajectory task execution models are limited by the availability of error correction data samples, failing to teach the model how to recover from errors or to generalize to more scenarios. This application employs a two-stage independent training process for the VLA model. This approach endows the model with the ability to execute the target task proficiently and smoothly, while also providing a degree of environmental generalization and error recovery capabilities. For example, in a grasping task, if the object is not grasped on the first attempt, the task execution model will replan its actions and attempt to grasp again.

[0042] This application adopts a two-stage training method. A large number of first sample datasets are collected from robots performing different tasks to form training data to train a general task execution model. Then, a small number of target sample data from a liveness detection test task are used to train the general task execution model to obtain a VLA model. This can improve the generalization of the model, enabling the model to transfer in different environments and tasks, and reduce costs.

[0043] After training the VLA model using the above method, the trained VLA model is used to determine the action trajectory of the task execution object based on textual description information and sensor information.

[0044] In one embodiment, textual description information and sensor information are processed to obtain multimodal fusion data.

[0045] Specifically, feature transformation is performed on text description information and sensor information respectively to obtain text data features and sensor data features. For example, the text description information is encoded into discrete token sequences. A token sequence, in natural language processing, refers to an ordered list formed by segmenting the original text into discrete, processable basic units (i.e., tokens) according to certain rules (such as spaces, punctuation marks, or sub-word units). Each token can be a word, sub-word, punctuation mark, or special marker. Token sequences are the fundamental form of input and text processing for large language models, and their structure directly affects the model's ability to understand and generate semantics, syntax, and context. In a specific embodiment, a token segmenter can be used to split the text description information to obtain token sequences, and then an embedding layer can be used to map the token sequences to obtain text data features. A visual encoder is used to process visual data, such as image information acquired by a camera, to extract high-level semantic features of the image. Then, a linear projection layer is used to map the extracted features to obtain visual data features. By using a multilayer perceptron to perform dimensional transformation and mapping on pose information such as joint angles and end effector posture, pose information features are obtained. It can be understood that pose information features and visual data features together constitute sensor data features.

[0046] It should be noted that the text data features, visual data features, and pose information features obtained through the above processing are all d_model-dimensional features, which facilitates subsequent cross-modal attention calculation and feature fusion.

[0047] Text data features and sensor data features are concatenated to obtain multimodal fusion data. This is understandable, as text data features, visual data features, and pose information features are all d_model-dimensional features; therefore, they can be concatenated to obtain multimodal fusion data.

[0048] Multimodal fusion data is used as VLA (Vision) Language The input to the Action (VLA, Visual Language Action Model) model is used to process the multimodal fusion data to obtain encoded features; and the encoded features are then processed to obtain the action trajectory of the task execution object.

[0049] It should be noted that the VLA model includes a Vision-Language Model (VLM) and an Action Expert Model (AEM), with the output of the VLM connected to the input of the AEM. Specifically, the VLA model of this application adopts an end-to-end structure of the VLM and AEM, forming a closed-loop multimodal intelligent system of perception-decision-execution. This system can map multimodal semantic features to continuous action features, directly converting the intent understood by the VLM into executable actions by the AEM.

[0050] In one embodiment, multimodal fusion data is used as input to a visual language model. The encoder in the visual language model performs cross-modal alignment and fusion on the multimodal fusion data to obtain encoded features. It should be noted that these encoded features include alignment information between text data features and visual data features, alignment information between text data features and pose information features, alignment information between visual data features and pose information features, and global task intent understanding information.

[0051] In one embodiment, when multimodal fusion data is input into a visual language model, the weights of the visual language model can be frozen, retaining only the multimodal semantic reasoning capability of the visual language model itself, thus reducing the computational load.

[0052] The encoded features are further processed using a motion expert model to obtain the motion trajectory of the task execution object.

[0053] In one specific embodiment, the motion expert model employs a flow matching architecture to simulate the continuous distribution of actions. By processing the encoded features using the motion expert model, it directly predicts consecutive future actions: A t = [a t ,a t+1 , a t+2 , …, a t+H-1 ], where H represents the size of the predicted consecutive future action block. For example, H=50 means predicting a sequence of fifty future actions, with each action a. t It is a vector containing the angles of all the robot's joints. This action can be directly used for robot control, such as joint velocities, end-effector pose increments, or gripper opening and closing commands. The motion expert model includes a cross-attention mechanism, which processes the encoded features to obtain conditional signals, processes these conditional signals, and ultimately predicts continuous future actions, i.e., motion trajectories.

[0054] Specifically, practical applications have shown that in liveness detection tasks, task execution models such as robotic arms / actuators have 6 degrees of freedom, and the output motion trajectory includes continuous joint angles in the [6×H] dimension.

[0055] Step S13: Based on the motion trajectory control task execution object, grasp the target object to perform a liveness detection test task; wherein, the target object includes a grasping marker.

[0056] Specifically, the target object is a piece of paper containing the test material. The gripping mark is set on the opposite side of the paper from the test material, and the gripping mark is a different color from the paper.

[0057] It should be noted that since the test materials in liveness detection tasks are mostly paper, and paper is breathable, directly using the robotic arm's suction cup to grab the paper during the collection process can easily result in multiple sheets being picked up simultaneously. Furthermore, because the paper materials may vary in shape and size, errors may occur, or the paper edges may be grabbed, thus affecting the progress of subsequent normal testing procedures. Therefore, gripping marks are made on the back of the paper, i.e., the side opposite to the test material, for example... Figure 4 The red markers not only solve the problem of breathable material gripping, but also provide clear visual information for the model, serving as target points for the robotic arm to accurately pick up each image. Of course, gripping markers can also be other non-breathable markers with clear visual characteristics.

[0058] In one embodiment, the execution unit includes a suction cup, and the task execution object grasping target object includes: the suction cup adsorbing and grasping a marker. Specifically, in order to display the test material flat, a suction cup is designed in the execution unit, and the suction cup is used to adsorb and grasp the marker to achieve the grasping of the target object.

[0059] Furthermore, in combination Figure 2 Step S13 includes: Step S131: Control the task execution object to move to the corresponding position of the target object, and make the suction cup adsorb the grab mark.

[0060] Specifically, in combination Figure 5 By using a top-down camera view to observe the initial position of the test material in the current scene, the task execution object is controlled to move from the initial position ( Figure 5 (as shown in (a)) Move above the target object ( Figure 5 (As shown in (b)). The control task object moves downward, causing the suction cup to attach to the gripping mark, and the suction cup is activated, causing the suction cup to adhere to the gripping mark ( Figure 5 (as shown in (c)).

[0061] Step S132: Control the task execution object to move the target object to the test area of ​​the device under test, and move the target object according to the preset trajectory so that the device under test can identify the test material in the target object from multiple different angles.

[0062] The control task execution object moves the test material with biometric characteristics, i.e., the target object, in front of the gate camera. Figure 5 As shown in (d); the task execution object randomly shakes the target object up, down, left, right, forward, and backward according to a preset trajectory (as shown in d); Figure 5 As shown in (e), the device under test can identify the test material in the target object from multiple different angles, thereby testing whether the gate will misidentify the target on the paper, and simulating a live attack in turn.

[0063] Step S133: Control the task execution object to place the target object at the target location.

[0064] After the simulation is complete, the control task execution object moves the target object to the target location, closes the suction cup, and places the target object at the target location. Figure 5 (as shown in f), then return to the initial position to enter the first step, and repeat the cycle.

[0065] This embodiment uses embodied intelligent robots (including but not limited to fixed-base robotic arms, wheeled robots, and humanoid robots) to replace manual labor in grasping test materials of different sizes and shapes, and simulates a live attack and a cyclical process of replacing test materials until all materials are used up and the test task ends. This achieves fully automated testing, freeing up manpower and reducing costs.

[0066] Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the electronic terminal provided in this application. The electronic terminal 80 includes a memory 81 and a processor 82 coupled to each other. The processor 82 is used to execute program instructions stored in the memory 81 to implement the steps of any of the above-described embodiments of the liveness detection test method. In a specific implementation scenario, the terminal 80 may include, but is not limited to, a microcomputer or a server. In addition, the terminal 80 may also include mobile devices such as laptops and tablets, which are not limited here.

[0067] Specifically, processor 82 controls itself and memory 81 to implement the steps of any of the above-described liveness detection test method embodiments. Processor 82 can also be referred to as a CPU (Central Processing Unit). Processor 82 may be an integrated circuit chip with signal processing capabilities. Processor 82 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 82 can be implemented using integrated circuit chips.

[0068] Please see Figure 7 , Figure 7 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 90 stores program instructions 901 that can be executed by a processor. The program instructions 901 are used to implement the steps of any of the above-described liveness detection test methods implementation embodiments.

[0069] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0070] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0071] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0072] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0073] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0074] The above are merely embodiments of this application and do not limit the scope of patent protection of this application. Any equivalent structural or procedural changes made using the content of this application’s specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.

Claims

1. A test method for detecting life, characterized by, The method is applied to a task execution object, which includes a main body and an execution unit. The main body is equipped with a first-view image sensor, and the execution unit is equipped with a second-view image sensor. The task execution object also includes a pose sensor. The testing method includes: Receive text description information and sensor information corresponding to the liveness detection test task, wherein the sensor information includes image information and pose information; The VLA model is used to determine the motion trajectory of the task execution object based on the text description information and the sensor information; Based on the motion trajectory, the task execution object is controlled to grasp the target object to perform the liveness detection test task; wherein, the target object includes a grasping marker.

2. The test method of claim 1, wherein, The target object is a piece of paper containing test material. The gripping mark is located on the side of the paper opposite to the test material, and the gripping mark is a different color from the paper.

3. The test method of claim 1, wherein, The execution unit includes a suction cup, and the task execution object grasping the target object includes: the suction cup adsorbing the grasping mark.

4. The test method according to claim 2, characterized in that, Based on the motion trajectory, the task execution object is controlled to grasp the target object to perform the liveness detection test task, including: Control the task execution object to move to the corresponding position of the target object, and cause the suction cup to adhere to the grasping mark; The task execution object is controlled to move the target object to the test area of ​​the device under test, and the target object is moved according to a preset trajectory so that the device under test can identify the test material in the target object from multiple different angles; Control the task execution object to place the target object at the target location.

5. The test method according to claim 4, characterized in that, Controlling the task execution object to move to the corresponding position of the target object, and causing the suction cup to adhere to the grasping mark, includes: Control the task execution object to move above the target object; Control the task execution object to move downwards, so that the suction cup attaches to the grasping mark; Activate the suction cup so that it adheres to the grasping mark; Controlling the task execution object to place the target object at the target location includes: The task execution object is controlled to move the target object to the target location, and the suction cup is turned off, thus placing the target object at the target location.

6. The test method according to claim 1, characterized in that, The task execution objects include a first robotic arm and a second robotic arm, wherein the first robotic arm includes the main body and the execution part; Before determining the motion trajectory of the task execution object using the VLA model based on the text description information and the sensor information, the process includes: The first robotic arm is controlled based on the simulated trajectory information, and image sample data and pose sample data of the first robotic arm are collected to obtain target sample data; wherein, the simulated trajectory information is generated by the second robotic arm when simulating the liveness detection test task; The general task execution model is trained using the target sample data to obtain the VLA model; wherein, the general task execution model is obtained by training the initial VLA model using mixed sample data corresponding to multiple types of tasks.

7. The test method according to claim 1, characterized in that, Determining the motion trajectory of the task execution object using a VLA model based on the text description information and the sensor information includes: The text description information and the sensor information are processed to obtain multimodal fusion data; The multimodal fusion data is used as input to the VLA model, and the VLA model is used to process the multimodal fusion data to obtain encoded features; then the encoded features are processed to obtain the action trajectory of the task execution object.

8. The test method according to claim 7, characterized in that, The text description information and the sensor information are processed to obtain multimodal fusion data, including: The text description information and the sensor information are respectively subjected to feature transformation to obtain text data features and sensor data features; The text data features and sensor data features are concatenated to obtain the multimodal fusion data.

9. An electronic terminal, characterized in that, The electronic terminal includes a memory and a processor coupled to each other. The processor is used to execute program instructions stored in the memory and to execute program data to implement the steps in the liveness detection test method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the liveness detection test method as described in any one of claims 1 to 8.