Transaction information evaluation method and device, equipment and medium

By combining video frames and touch data into a deep learning model to evaluate the attributes of target objects during the transaction process, the security and flexibility issues of existing identity verification are resolved, achieving higher accuracy in identity authentication and a better user interaction experience.

CN120931392APending Publication Date: 2025-11-11AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511031720.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing methods of evaluating transaction information through authentication methods such as usernames, passwords, and PIN codes are vulnerable to brute-force attacks and suffer from low flexibility, adaptability, and security.

Method used

By acquiring video frames and touch data during the transaction process, a pre-trained deep learning model is used to determine the object evaluation attributes, gaze confidence, and touch evaluation attributes of the target object. Combined with transaction scenario information, an interaction risk assessment function is used to determine the transaction attributes of the target object.

Benefits of technology

It enhances the accuracy and security of target object identity authentication and optimizes the customer's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931392A_ABST
    Figure CN120931392A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a transaction information evaluation method and device, equipment and a medium. The method comprises the following steps: acquiring a video frame and touch data in a transaction process; inputting the video frame into a first model obtained by pre-training, and determining an object evaluation attribute of a target object in the video frame; determining the watching confidence of the target object according to the video frame and the center coordinate of the key area in the transaction interface; for the touch data, inputting the touch data into a second model obtained by pre-training, and determining a touch evaluation attribute of the target object; and determining a target object transaction attribute according to the object evaluation attribute, the gaze confidence, the touch evaluation attribute, the transaction scene information and the interaction risk evaluation function. According to the technical scheme of the embodiment of the invention, the transaction attribute of the target object is determined based on the video frame and the touch data in the transaction process, and through combination of multiple biological recognition modes, the accuracy and safety of identity authentication of the target object are enhanced, and the interaction experience of customers is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of financial technology, and in particular to a method, apparatus, device and medium for evaluating transaction information. Background Technology

[0002] With the development of financial informatization, financial services are gradually transforming towards online and intelligent models. At this time, in order to improve the security, convenience, and user experience of financial services, it is necessary to evaluate transaction information when users interact with terminal devices.

[0003] Currently, the main method used for evaluating transaction information during user interaction is through authentication methods such as usernames, passwords, and PIN codes. However, this method suffers from vulnerabilities to brute-force attacks, as well as low flexibility, adaptability, and security. Summary of the Invention

[0004] This disclosure provides a method, apparatus, device, and medium for evaluating transaction information, thereby enhancing the accuracy and security of target object identity authentication and optimizing the customer's interactive experience.

[0005] In a first aspect, embodiments of this disclosure provide a method for evaluating transaction information, the method comprising:

[0006] Acquire at least one video frame and at least one touch data during the transaction process; wherein, the touch data includes touch pressure data when the target object contacts the device, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data;

[0007] The at least one video frame is input into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame;

[0008] The gaze confidence of the target object is determined based on the center coordinates of at least one video frame and at least one key area in the transaction interface.

[0009] For the at least one touch data, the touch data is input into a pre-trained second model to determine the touch evaluation attributes of the target object;

[0010] The target object's transaction attributes are determined based on the object evaluation attributes, the gaze confidence level, the touch evaluation attributes, the transaction scenario information, and the interaction risk assessment function; wherein, the transaction scenario information includes at least transaction time information, transaction location information, and transaction device information.

[0011] Secondly, embodiments of the present invention also provide a transaction information evaluation device, the device comprising:

[0012] The data acquisition module is used to acquire at least one video frame and at least one touch data during the transaction process; wherein, the touch data includes touch pressure data when the target object contacts the device, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data;

[0013] An object evaluation attribute determination module is used to input the at least one video frame into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame;

[0014] A gaze confidence determination module is used to determine the gaze confidence of the target object based on the center coordinates of at least one video frame and at least one key area in the transaction interface.

[0015] A touch evaluation attribute determination module is used to input the touch data into a pre-trained second model for the at least one touch data to determine the touch evaluation attributes of the target object;

[0016] The target object transaction attribute determination module is used to determine the target object transaction attributes based on the object evaluation attributes, the gaze confidence level, the touch evaluation attributes, transaction scenario information, and the interaction risk assessment function; wherein, the transaction scenario information includes at least transaction time information, transaction location information, and transaction device information.

[0017] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0018] One or more processors;

[0019] Storage device for storing one or more programs.

[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the transaction information evaluation method as described in any of the embodiments of the present invention.

[0021] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the transaction information evaluation method as described in any of the embodiments of the present invention.

[0022] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program, characterized in that, when executed by a processor, the computer program implements the transaction information evaluation method as described in any of the embodiments of the present invention.

[0023] The technical solution of this disclosure involves acquiring at least one video frame and at least one touch data point during a transaction. The touch data includes touch pressure data, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data when the target object contacts the device. Then, the at least one video frame is input into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame. Based on the center coordinates of the at least one video frame and at least one key area in the transaction interface, the gaze confidence of the target object is determined. Further, for the at least one touch data point, the touch data is input into a pre-trained second model to determine the touch evaluation attributes of the target object. Finally, based on the object evaluation attributes, gaze confidence, touch evaluation attributes, transaction scenario information, and an interaction risk assessment function, the transaction attributes of the target object are determined. The transaction scenario information includes at least transaction time information, transaction location information, and transaction device information. This solution addresses the problems of vulnerability to brute-force attacks, low flexibility, adaptability, and security when evaluating transaction information using authentication methods such as usernames, passwords, and PIN codes. This invention determines the transaction attributes of a target object based on video frames and touch data during the transaction process. By combining multiple biometric identification methods, it enhances the accuracy and security of target object identity authentication, thereby improving and optimizing the customer's interactive experience. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of exemplary embodiments of the present invention, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the drawings of the embodiments to be described in this invention, and not all of the drawings. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort.

[0025] Figure 1 This is a schematic flowchart of a transaction information evaluation method provided in an embodiment of this disclosure;

[0026] Figure 2 This is a schematic diagram of the first model provided in the embodiments of this disclosure;

[0027] Figure 3 This is a schematic diagram illustrating the determination of gaze confidence provided in an embodiment of this disclosure;

[0028] Figure 4 This is a schematic diagram of the second model provided in an embodiment of the present invention;

[0029] Figure 5 This is a schematic diagram of the execution decision results provided in an embodiment of the present invention;

[0030] Figure 6This is a schematic diagram of the structure of a transaction information evaluation device provided in an embodiment of the present invention;

[0031] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0032] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0033] Before introducing the technical solutions provided in the embodiments of this disclosure, the application scenarios can be illustrated first. The technical solutions provided in the embodiments of this disclosure can be applied to scenarios where multiple biometric identification methods are combined during the transaction process to determine the transaction attributes of a target object. Based on the technical solutions of the embodiments of this disclosure, the transaction attributes of the target object are determined based on video frames and touch data during the transaction process. By combining multiple biometric identification methods, the accuracy and security of the target object's identity authentication are enhanced, thereby optimizing the customer's interactive experience.

[0034] Example 1

[0035] Figure 1 This is a flowchart illustrating a transaction information evaluation method provided in this embodiment. This embodiment is applicable to situations where the transaction attributes of a target object are determined through a combination of multiple biometric identification methods during the transaction process. This method can be executed by a transaction information evaluation device, which can be implemented in the form of software and / or hardware. The hardware can be a mobile electronic device that can execute the transaction information evaluation method provided in this technical solution.

[0036] like Figure 1 As shown, the method includes:

[0037] S110. Acquire at least one video frame and at least one touch data during the transaction process.

[0038] The touch data includes touch pressure data when the target object comes into contact with the device, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data.

[0039] Optionally, the touch data includes: data related to device contact generated when the target object clicks on the device screen or buttons from the start to the end of the transaction; video frames include image data containing the target object captured by a camera device at a preset frame rate during the transaction.

[0040] It should be noted that the transaction process refers to the sum of all operational steps and related interactive behaviors performed by the target object in a customer self-service device, such as a self-service terminal, mobile APP interface, or ATM, in order to complete various transaction businesses. The entire transaction process refers to all activities that occur within the entire time period from when the target object initiates the transaction-related operations until the transaction is successfully completed or fails for various reasons. The preset frame rate refers to the number of image frames captured per second when the video frame of the transaction process is captured by the camera device. The target object refers to the subject that interacts with the device during the transaction process, usually the customer conducting the transaction. In some special cases, the target object may also include personnel assisting the customer in the transaction operations. The device can include the target object's self-service device, client device, or ATM, etc.

[0041] It's also important to note that pressure sensitivity data refers to the amount of pressure applied to the device's contact surface when the target object comes into contact with it, such as when clicking or swiping on a touchscreen. Pressure sensitivity data can be measured using the device's built-in pressure sensor. Different levels of pressure may reflect the target object's intentions and habits; for example, a forceful click might indicate certainty about an operation, while a light touch might be merely tentative or cursory. Contact area data refers to the actual contact area between the target object's fingers or other contact parts and the device surface. This data reveals the target object's contact method, such as whether they are using a single fingertip or the entire fingertip for a larger area of ​​contact. Different contact areas may correspond to different levels of accuracy and intention. Touch duration data refers to the length of time the target object remains in contact with the device. Timing begins when the target object's finger or other contact part touches the device surface and ends when they leave. Touch duration data reflects the target object's understanding and familiarity with a particular operation; for example, a shorter touch duration might be used for familiar users, while a longer duration might be used for unfamiliar operations or operations requiring careful confirmation. Touch interval duration data refers to the time interval between two consecutive touch operations during a complete operation of the target object. Touch pressure change rate data refers to the rate at which the touch pressure changes over time while the target object is in contact with the device. Touch pressure change rate data reflects how the pressure changes during target object operation, such as whether the pressure increases rapidly, decreases slowly, or remains relatively stable.

[0042] Specifically, during the entire transaction process, at least one video frame is captured using the device's camera. Furthermore, sensors are used to collect data on the contact pressure, contact area, touch duration, touch interval, and touch pressure change rate when the target object comes into contact with the device.

[0043] S120. Input at least one video frame into a pre-trained first model to determine the object evaluation attributes of the target object in at least one video frame.

[0044] The first model refers to the model used to determine the object evaluation attribute of a target object in an input video frame. As a pre-trained deep learning model, the first model extracts feature representations of the target object from the video frame and calculates its similarity score with a preset standard, such as an image in a human traffic database. It should be noted that the similarity score typically ranges from 0 to 1, where 1 indicates a perfect match between the target object and an image in the database, and 0 indicates a complete mismatch. For example, when a video frame is input into the pre-trained first model, and the object evaluation attribute of the target object in the video frame is determined to be 0.9, it means that the target object is highly similar to an image in the database. The object evaluation attribute refers to the quantitative evaluation result of the first model on the target object, used to determine the similarity or conformity between the target object and an image in the database.

[0045] It should be noted that, based on the input-output relationship, the first model includes, in order, a backbone network, a neck network, and a head network. The backbone network includes at least ghost convolutional layers, convolutional layers, cross-stage partial fusion layers, channel attention layers, spatial attention layers, and spatial pyramid pooling layers. The neck network includes at least a path aggregation network, a feature pyramid network, and a bidirectional feature pyramid network. The head network includes at least a classification branch, a bounding box regression branch, and a keypoint regression branch.

[0046] It's important to note that ghost convolution layers are a type of efficient convolutional operation designed to reduce computational complexity and the number of parameters by generating "ghost features." Ghost convolution layers generate features by performing a small number of convolution operations on the input feature map, and then use simple linear operations to generate additional features, thus reducing computational cost. Convolutional layers are primarily used to extract local features from the input data. A sliding filter is used to perform convolution operations on the input feature map to generate the feature map. The main function of convolutional layers is to capture spatial features. Cross-stage partial fusion layers are used to fuse features from different stages, typically to improve the expressive power of features. By partially connecting features from different stages, cross-stage partial fusion layers can achieve effective information transfer, thereby enhancing the first model's ability to learn multi-scale features. Channel attention layers are used to dynamically adjust the importance of each channel in the feature map. By calculating the channel weights, the first model can focus more on features that are more important to the task and suppress unimportant features, thereby improving the performance of the first model. Spatial attention layers focus on spatial information in the feature map. By calculating the weights at each location, the first model can focus on more important regions in the image, thereby enhancing the response to features in specific regions. Spatial pyramid pooling layers are used to process feature maps at different scales. They can pool features at different scales to generate a fixed-size output. Spatial pyramid pooling layers can effectively capture multi-scale information in images and are suitable for processing input images of different sizes. Path aggregation networks aim to enhance the network's learning of multi-scale features by aggregating features from different paths. Path aggregation networks integrate features through different paths to improve the expressive power of features. Feature pyramid networks are a network structure for multi-scale feature extraction. Feature pyramid networks achieve effective fusion of features at different scales by constructing feature pyramids, enhancing the first model's ability to detect targets of different sizes. Bidirectional feature pyramid networks are an extension of feature pyramid networks, capable of fusing features in bottom-up and top-down paths, further improving the expressive power of multi-scale features and detection performance. The classification branch is responsible for classifying detected targets and outputting the probability of the target's class. The classification branch typically uses a softmax function to map features to class probabilities for target recognition. The bounding box regression branch is used to predict the bounding box position of the target. The regression model outputs the target's position coordinates and size for accurate target localization. The keypoint regression branch is used to predict the location of keypoints of a target. Keypoints can represent specific facial features or body joints, and the regression model outputs the coordinates of the keypoints.

[0047] It should also be noted that in the backbone network, channel attention layers and spatial attention layers are inserted after the C2f layer. The C2f layer, as an efficient residual block, enhances the model's occlusion robustness when followed by these layers. The channel attention layer adaptively adjusts the weights of each channel, strengthening key features. The spatial attention layer focuses on salient regions of the target object, such as unoccluded parts, thus suppressing background interference. Furthermore, ghost convolutional layers are used to replace some convolutional layers. Using ghost convolutional layers allows for the generation of redundant features through inexpensive linear transformations, reducing the computational cost of standard convolutions and decreasing the size of the first model while maintaining feature expressiveness. In the head network, a keypoint regression branch is added to the existing detection head, in addition to the classification and regression branches. For example, it regresses five keypoints: the left and right eyes, the tip of the nose, and the left and right corners of the mouth, improving alignment accuracy and performance for subsequent tasks. In the first model, these components together constitute a complex deep learning model. Through different network structures and layers, features can be effectively extracted and fused, thereby achieving the task of efficiently determining object evaluation attributes.

[0048] Specifically, after acquiring at least one video frame during the transaction process, the video frame is input into a pre-trained first model. Based on the first model, the object evaluation attribute of the target object in the video frame is determined, that is, the similarity score between the target object in the video frame and the image in the People's Bank of China database is determined.

[0049] For example, see Figure 2 For the video frames acquired during the transaction process, before inputting the original video frame images into the model, the video frames can undergo image preprocessing steps such as distortion correction or high dynamic range (HDR) processing. After preprocessing, the video frames are input into the YOLOv8 model, which outputs the object evaluation attributes of the target object in the video frame. Furthermore, in the backbone network of the YOLOv8 model, channel attention layers and spatial attention layers are inserted after the C2f layer, and ghost convolutional layers are used to replace some convolutional layers. After acquiring continuous video frames, 3D structured light liveness detection can be used to determine whether the target object is alive. 3D structured light liveness detection is a technique used to distinguish between real faces and static images. After the video frames undergo 3D structured light liveness detection, the device projects a structured light pattern onto the target, and a camera captures the reflected light pattern to form a depth map. Then, the depth map and light pattern deformation are analyzed to extract 3D information. Finally, based on the acquired depth information, features are extracted to determine whether the target object is alive. Live objects typically exhibit dynamically changing characteristics, such as subtle facial expressions and blinking. The depth map of the target object is calculated using 3D structured light. The dynamic threshold is automatically adjusted based on the real-time ambient light intensity; for example, the threshold is widened to 5mm in low light and tightened to 2mm in strong light. The formula for adjusting the dynamic threshold is: T depth =Tbase +α*E light T base This refers to the base threshold, such as 3mm, corresponding to standard lighting conditions of 500-1000 lux; α refers to the illumination influence coefficient, which controls the sensitivity of the threshold to changes in illumination; E light This refers to the current light intensity, which is collected in real time by a light sensor. 3D structured light liveness detection can effectively distinguish between real target objects and static images, which is crucial for improving security and accuracy.

[0050] When training the first model, a first training image is determined based on at least one video frame, and each first training sample in the training sample set is obtained based on the object evaluation attributes of the target object in the first training image and the corresponding video frame.

[0051] To improve the accuracy of the first model training, video frames captured from different perspectives can be obtained and used as the first training images. The sum of all the first training images constitutes the training sample set. That is, the training sample set includes multiple first training images. The term "first training image" is relative and not a specific limitation. To improve the accuracy of the trained first model, as many and varied first training images as possible can be obtained. Each first training sample includes the first training image and the corresponding object evaluation attribute. Accordingly, the training sample set can include multiple first training samples.

[0052] For each first training sample, the first image to be trained in the current first training sample is input into the first model to be trained to obtain the actual object evaluation attributes corresponding to the current first training sample.

[0053] It should be noted that the above training method can be used to train each first training sample to obtain the desired target first model.

[0054] Here, the first model to be trained is a model whose model parameters are either initial parameters or default parameters. The actual object evaluation attribute is the object evaluation attribute output after the first training image from the current first training sample is input into the first model to be trained.

[0055] It should be noted that the model parameters in the first model to be trained do not meet the expected requirements. Therefore, there is a certain difference between the actual object evaluation attributes output based on the model parameters at this time and the theoretical object evaluation attributes. Therefore, the corresponding error loss value can be determined based on the actual object evaluation attributes and theoretical object evaluation attributes corresponding to each image to be processed.

[0056] Based on the first preset loss function in the first model to be trained, the actual object evaluation attributes and theoretical object evaluation attributes of the current first training sample are subjected to loss processing, so as to correct the model parameters in the first model to be trained according to the obtained loss value.

[0057] It should be noted that the training parameters can be set to default values ​​before training the first model to be trained. During the training of the first model to be trained, the training parameters can be adjusted based on its output. In other words, the target first model can be obtained by modifying the loss function in the first model to be trained. Each image to be processed has a corresponding loss value, which is determined based on the actual object evaluation attributes and theoretical object evaluation attributes of each image.

[0058] Specifically, after inputting the first training image from the first training sample into the first training model, the first training model can obtain the actual object evaluation attribute corresponding to the first training image. Based on this actual object evaluation attribute and the theoretical object evaluation attribute, the loss value corresponding to the first training image can be determined, and the model parameters in the first training model can be corrected using the backpropagation method.

[0059] The convergence of the first preset loss function is used as the training objective to obtain the first target model. This first target model is the final trained model used to determine the evaluation attributes of the target object in the image to be processed.

[0060] Specifically, the training error of the loss function, i.e., the loss parameter, can be used as a condition to detect whether the loss function has reached convergence. For example, this could be whether the training error is less than a preset error, whether the error trend is stable, or whether the current number of iterations equals a preset number. If convergence is detected, such as the training error of the loss function being less than the preset error or the error trend being stable, it indicates that the first model to be trained has completed training, and iterative training can be stopped. If convergence is not detected, the first training sample can be obtained again to train the first model to be trained until the training error of the loss function is within a preset range. When the training error of the loss function converges, the first model to be trained can be used as the target first model.

[0061] S130. Determine the gaze confidence of the target object based on the center coordinates of at least one video frame and at least one key area in the transaction interface.

[0062] Optionally, for at least one video frame, the video frame is input into the third model to determine the pupil coordinates corresponding to the eyes of the target object in the video frame; after smoothing the at least one pupil coordinate, the pupil motion trajectory is determined; based on the pupil motion trajectory and the center coordinates of at least one key region, at least one first distance is determined; based on at least one first distance and a preset threshold, the gaze confidence of the target object is determined.

[0063] The transaction interface refers to the graphical interface through which the target object performs operations, such as clicking the "Confirm" button. The transaction interface typically includes a key area and a coordinate system. The key area includes interactive controls such as buttons, input boxes, and checkboxes. The coordinate system can be based on the top-left corner of the screen corresponding to the transaction interface, with the x-axis horizontal and the y-axis vertical, and is expressed in pixels. The center coordinates of the key area refer to the coordinates of the geometric center point of the key area, such as the button. Gaze confidence is an indicator that quantifies whether the target object's gaze is stably focused on the key area. The gaze confidence value can range from 0 to 1. 0 represents that the target object's pupil coordinates continuously deviate from the key area, while 1 represents that the target object's pupil coordinates continuously lie within the key area. The third model is a pre-trained model used for pupil localization. The third model can be a lightweight pupil detection model based on deep learning, such as the MediaPipe Iris model. The input of the third model is a single video frame, and the output is the center coordinates of the pupil. Based on the pupil coordinates corresponding to the target object's eyes in the video frame determined by the third model, the coordinate system in the transaction interface can be used to determine the pupil location. Pupil motion trajectory refers to the path formed by the change of pupil coordinates of a target object over a period of time. The change in pupil position across multiple video frames can be represented by a series of coordinate points. To reduce noise and irregular changes, the pupil coordinates corresponding to the target object's eyes in the video frames can be smoothed, for example, using moving averages or Kalman filtering, to make the pupil motion trajectory more stable and consistent. The first distance refers to the distance from the current pupil position, i.e., the pupil coordinates, to the center coordinates of the key area. For each pupil coordinate (x... p ,y p The first distance can be expressed using the Euclidean distance formula, that is, when the center coordinates of a certain key region are (x... k ,y k When ), the first distance can be: The first distance reflects the spatial relationship between the object's gaze point and the key region. The preset threshold is a standard value used to determine the object's gaze confidence. The preset threshold is a value set according to the actual application scenario and can be determined based on experimental data or experience. For example, setting it to a specific preset threshold, such as 10 pixels or 20 pixels, indicates that within this range, the object is considered to be gazing at the key region. By comparing the first distance and the preset threshold, the object's gaze confidence can be determined. If the first distance is not greater than the preset threshold, it proves that the object's gaze confidence is high, indicating that the object is paying attention to the key region, and the corresponding video frame can be defined as a valid gaze frame; if the first distance is greater than the preset threshold, it is considered that the object's gaze confidence is low, indicating that the object is not paying attention to the key region, and the corresponding video frame can be defined as an invalid gaze frame.

[0064] It should be noted that, see Figure 3 After inputting video frames into the pre-trained third model, it can output the pupil coordinates corresponding to the target object's eyes. The training process of the third model is similar to that of the first model and will not be repeated here. The training samples for the third model during training are video frames and the theoretical pupil coordinates corresponding to the target object's eyes. When determining the gaze confidence of the target object based on at least one first distance and a preset threshold, the gaze confidence of the target object can be determined by dividing the number of effective gaze frames by the total number of analysis frames.

[0065] Specifically, after determining the pupil movement trajectory based on video frames, the fixation confidence is determined based on the distance relationship between the pupil movement trajectory and the center coordinates of the key area. Based on the fixation confidence, the stability of the target object's gaze is quantified as it remains focused on the key area of ​​the trading interface.

[0066] S140. For at least one touch data, input the touch data into a pre-trained second model to determine the touch evaluation attributes of the target object.

[0067] Optionally, for at least one touch data, touch pressure data, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data are input into the second model, and the touch data are processed based on the input layer, long short-term memory layer, random deactivation layer, fully connected layer, and output layer to output the touch evaluation attributes of the target object.

[0068] It should be noted that the second model is a pre-trained deep learning model specifically designed to process touch data and evaluate the touch assessment attributes of the target object. The design of the second model can include various neural network architectures, with the main purpose of extracting and analyzing features from the input touch data to output the probability of abnormal touch behavior of the target user. The training process of the second model is similar to that of the first model and will not be repeated here. The training samples for the second model are touch data and the touch assessment attributes of the target object. Touch assessment attributes refer to the features or indicators related to the touch behavior of the target object obtained by analyzing the touch data. Touch assessment attributes are the key indicators output by the second model, representing the probability of the detected touch behavior deviating from normal behavior. This probability value is typically between 0 and 1; a value closer to 1 indicates that the touch behavior is more likely to be abnormal, while a value closer to 0 indicates that the touch behavior is normal.

[0069] It should also be noted that, see Figure 4 The input layer is the first layer of the second model, responsible for receiving touch data. It passes the received touch data to subsequent layers for processing. Each input feature can be a separate node. The Long Short-Term Memory (LSTM) layer learns and remembers long-term dependencies in sequential data. When processing touch data, the LSM layer effectively captures the dynamic features of touch behavior over time, such as the touch patterns of a target object at different points in time. The LSM layer is particularly suitable for processing time-series data, helping to identify trends and anomalies in touch behavior. The random deactivation layer is a regularization technique used to prevent overfitting in the second model. During the training of the second model, the random deactivation layer can randomly "shut down" a certain proportion of neurons, making the second model more robust during learning and improving its generalization ability. This is crucial for processing touch data, as real-world touch data may contain noise and uncertainty. Each input node in the fully connected layer is connected to an output node. The fully connected layer is used to integrate and transform the features extracted from previous layers, typically used in the final stage of the second model to produce the final output. When determining touch evaluation attributes, the fully connected layer combines the features extracted from the Long Short-Term Memory layer into a higher-level representation for further analysis. The output layer, the last layer of the second model, is responsible for generating the final prediction result. In the second model, the output layer generates the touch evaluation attributes of the target object based on the preceding layers, i.e., the probability of abnormal touch behavior by the target user.

[0070] Specifically, touch pressure data, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data are input into a pre-trained second model. After processing the input touch data using the second model's long short-term memory layer, random deactivation layer, and fully connected layer, the touch evaluation attributes of the target object can be determined.

[0071] S150. Based on object evaluation attributes, gaze confidence, touch evaluation attributes, transaction scenario information, and interaction risk assessment function, determine the transaction attributes of the target object.

[0072] The transaction scenario information includes at least the transaction time information, transaction location information, and transaction device information.

[0073] Optionally, a first risk assessment attribute is determined based on the object assessment attribute and gaze confidence; a second risk assessment attribute is determined based on gaze confidence and touch assessment attribute; a third risk assessment attribute is determined based on transaction time information, transaction location information, and transaction device information; and the target object transaction attribute is determined based on the first risk assessment attribute, the second risk assessment attribute, the third risk assessment attribute, and the weights corresponding to each risk assessment attribute.

[0074] Transaction time information refers to the specific time when the transaction occurred, including the date, day of the week, and time of day, such as morning, afternoon, or evening. Transaction time information helps analyze transaction time patterns, such as the target's activity level or unusual transaction behavior during a specific time period. Transactions conducted late at night may carry higher risks. Transaction location information refers to the geographical location of the target when the transaction was conducted. Transaction location information can be a specific address or the device's IP address. By analyzing transaction location information, the target's behavioral patterns can be identified, and any anomalies can be determined. Transaction device information refers to the type of device used by the target, such as a mobile phone, tablet, or personal computer. Transaction device information can also include the device's operating system, browser information, and security status. Transaction device information helps assess transaction security and identify potential security risks. The interaction risk assessment function is an algorithm used to assess the risk level of a specific transaction or interaction. Target object transaction attributes refer to a transaction risk level index derived by comprehensively analyzing factors such as object assessment attributes, gaze confidence, touch assessment attributes, and transaction scenario information. Target object transaction attributes are usually represented as a risk score or level.

[0075] It should be noted that the first risk assessment attribute can be determined by applying linear and non-linear weights to the object evaluation attributes and gaze confidence. The second risk assessment attribute can be determined by applying linear and non-linear weights to the gaze confidence and touch evaluation attributes. When determining the third risk assessment attribute based on transaction time, transaction location, and transaction device information, firstly, a risk standard is defined for each transaction scenario. For example, for transaction time information, transactions during times such as late night or holidays can be assigned a higher risk value. For transaction location information, certain geographical locations can be assigned a higher risk value. For transaction device information, devices such as outdated devices or devices connected to public networks can be assigned a higher risk value. Then, the risk scores for the transaction time, transaction location, and transaction device are combined to derive the third risk assessment attribute. For example, addition or a weighted average can be used to calculate the third risk assessment attribute.

[0076] It should also be noted that when determining the transaction attributes of the target object, weights can be assigned to the first, second, and third risk assessment attributes. The weights corresponding to each risk assessment attribute can be determined based on their importance in the risk assessment.

[0077] In this embodiment, when the gaze confidence is lower than a preset threshold, the voiceprint data of the target object is collected, the voiceprint data is input into the pre-trained fourth model, and the voiceprint similarity evaluation attribute is output; based on the voiceprint similarity evaluation attribute, the object evaluation attribute and the gaze confidence, the first risk evaluation attribute is determined.

[0078] It's important to note that voiceprint data refers to biometric information obtained by analyzing the voice characteristics of a target object. Different target objects possess unique characteristics such as frequency, pitch, loudness, and timbre, which can be used to identify and verify the identity of the target object. Typically, voiceprint data can be collected through microphones, voice capture software, and mobile devices. The fourth model is a pre-trained deep learning model used to process voiceprint data and output voiceprint similarity evaluation attributes. The fourth model can employ convolutional neural networks, recurrent neural networks, or other structures suitable for processing audio signals. Through the fourth model, features are extracted from the input voiceprint data, and then voiceprint similarity is calculated by comparing it with existing voiceprints in the data. Finally, a voiceprint similarity evaluation attribute is output, representing the degree of similarity between the voiceprint data to be evaluated and the reference voiceprint. The voiceprint similarity evaluation attribute is the similarity value calculated by comparing the voiceprint data of the target object with reference voiceprint data, such as a stored voiceprint template. The voiceprint similarity evaluation attribute is typically represented as a value between 0 and 1. A value close to 1 indicates that the voiceprints are very similar, and the target objects may share the same identity. A value close to 0 in the voiceprint similarity evaluation attribute indicates that the voiceprints are dissimilar, and the identities of the target objects may not be consistent. The training process of the fourth model is similar to that of the first model, and will not be repeated here. The training samples for the fourth model are the voiceprint data of the target object and the voiceprint similarity evaluation attribute of the target object.

[0079] It should also be noted that when the output gaze confidence score is lower than a preset threshold, the voiceprint similarity assessment attribute is determined based on the voiceprint data. After determining the voiceprint similarity assessment attribute, the voiceprint similarity assessment attribute, the object assessment attribute, and the gaze confidence score are linearly and non-linearly weighted to determine the first risk assessment attribute.

[0080] In this embodiment, after determining the transaction attributes of the target object, a first response measure is executed when the transaction attributes of the target object fall within the first attribute range; a second response measure is executed when the transaction attributes of the target object fall within the second attribute range; and a third response measure is executed when the transaction attributes of the target object fall within the third attribute range.

[0081] It should be noted that the first attribute range typically refers to the target object's transaction attributes falling within a safe, low-risk range. For example, the first attribute range can be set to 0 to 0.3, indicating low transaction risk, and the target object's identity and transaction behavior are considered normal. Within the first attribute range, the transaction is considered safe. The second attribute range refers to the target object's transaction attributes falling within a medium-risk range. For example, the second attribute range can be set to 0.3 to 0.7, indicating medium transaction risk, with potential risk factors. Within the second attribute range, the system will take additional measures to verify the target object's identity or restrict certain functions. The third attribute range refers to the target object's transaction attributes falling within a high-risk range. For example, the third attribute range can be set to 0.7 to 1.0, indicating high transaction risk, and the target object's identity or transaction behavior may be abnormal. Within the third attribute range, the system will take mandatory measures to protect the security of the target object and the system. The first response measures may include logging or allowing normal access. The second response measures may include triggering two-factor authentication or restricting certain functions. When the target's transaction attributes fall within the second attribute range, requiring additional identity verification, such as entering a one-time verification code, SMS verification, or using biometrics, can enhance security and prevent unauthorized access. Third response measures may include blocking, freezing accounts, or sending alerts. By assessing the target's transaction attributes, appropriate response measures can be taken based on different risk levels to protect user and system security. This tiered response mechanism effectively addresses various risk scenarios, ensuring the security and reliability of transactions.

[0082] Specifically, based on object assessment attributes, gaze confidence, touch assessment attributes, and transaction scenario information, the first risk assessment attribute, the second risk assessment attribute, and the third risk assessment attribute can be determined. Based on each risk assessment attribute and its corresponding weight, the target object's transaction attributes can be determined.

[0083] The technical solution of this disclosure involves acquiring at least one video frame and at least one touch data point during a transaction. The touch data includes touch pressure data, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data when the target object contacts the device. Then, the at least one video frame is input into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame. Based on the center coordinates of the at least one video frame and at least one key area in the transaction interface, the gaze confidence of the target object is determined. Further, for the at least one touch data point, the touch data is input into a pre-trained second model to determine the touch evaluation attributes of the target object. Finally, based on the object evaluation attributes, gaze confidence, touch evaluation attributes, transaction scenario information, and an interaction risk assessment function, the transaction attributes of the target object are determined. The transaction scenario information includes at least transaction time information, transaction location information, and transaction device information. This solution addresses the problems of vulnerability to brute-force attacks, low flexibility, adaptability, and security when evaluating transaction information using authentication methods such as usernames, passwords, and PIN codes. This invention determines the transaction attributes of a target object based on video frames and touch data during the transaction process. By combining multiple biometric identification methods, it enhances the accuracy and security of target object identity authentication, thereby improving and optimizing the customer's interactive experience.

[0084] Example 2

[0085] As an optional embodiment of the present invention, an example is provided to further illustrate the invention.

[0086] See Figure 5 When collecting video frames, touch data, and voiceprint data of the target object, a dual-mode camera module can be used for video frame acquisition. This module can include a visible light camera and an infrared camera. The visible light camera is used for video frame acquisition under normal lighting conditions, while the infrared camera is used for video frame acquisition in low-light / high-light environments. A microphone array is used to collect voiceprint data of the target object. A four-microphone linear array supporting beamforming can be used. The microphone array's sound source localization accuracy is approximately 5 degrees horizontally, preventing recording attacks. A pressure sensor is used to collect touch data of the target object. The pressure sensor can be integrated into the touch layer of the ATM screen or device, with a linear error of approximately 1%. For video frames, touch data, and voiceprint data of the target object, the timestamps of the camera, microphone, and pressure sensor can be aligned using a precise time protocol to ensure an error within 1ms. A sliding window, with a window size of 200ms, can also be used to align multimodal data streams. Abnormal data can be filtered out, eliminating data frames with time deviations exceeding 10ms.

[0087] Among them, see Figure 5 When processing video frames, the YOLOv8 model is used. For the YOLOv8 model, backbone optimization is first performed by inserting CBAM after the C2f layer to enhance occlusion robustness. The Ghost module is used to replace some standard convolutions to reduce the model size. Improvements to the output head include adding keypoint regression branches for the left and right eyes, nose tip, and left and right corners of the mouth.

[0088] During 3D liveness detection, a facial depth map is calculated using 3D structured light to verify whether the attack is a planar attack, such as a photograph / screen capture. A dynamic threshold is set in the liveness detection process. The dynamic threshold T... depth The threshold is automatically adjusted based on real-time ambient light intensity; for example, the threshold is widened to 5mm in low light and tightened to 2mm in strong light. The formula for dynamic threshold adjustment is:

[0089] T depth =T base +αE light ;

[0090] Among them, T base This refers to the base threshold, such as 3mm, corresponding to standard illumination conditions of 500-1000 lux; α refers to the illumination influence coefficient, which controls the sensitivity of the threshold to changes in illumination; E light This refers to the current light intensity, which is collected in real time by a light sensor.

[0091] During eye-tracking, pupil localization is performed first using the MediaPipe Iris model, outputting the pupil center coordinates (x, y). Screen mapping introduces some accuracy error. For trajectory analysis, key areas of the user interface are first defined, such as the screen coordinate range of the "Confirm" button as x1, x2, y1, y2. Next, the Euclidean distance between the pupil coordinates and the key areas is calculated. If the Euclidean distance D is greater than a threshold, such as 50 pixels, it is marked as a gaze deviation. The formula for calculating the Euclidean distance D is as follows:

[0092]

[0093] Based on the above processing, biometric features are extracted from the raw data and their authenticity is verified, ensuring that attackers cannot bypass authentication by forging information. This addresses issues such as the ability of 3D structured light to resist traditional attacks like photos, videos, and silicone masks, and also improves environmental adaptability.

[0094] Click analysis begins with the collection of touch data, using pressure sensors to record the force, contact area, and duration of each click. Temporal features collected include click intervals and pressure gradients. An anomaly detection model is then constructed. Training data can include datasets of normal user actions (e.g., 10,000 clicks from different users across various age groups) and datasets of anomalous user actions (e.g., 2,000 simulated stress clicks involving rapid, continuous clicks and sudden changes in pressure). Performance metrics for the anomaly detection model can be evaluated using standard metrics such as precision, recall, and F1-Score.

[0095] For the collected data, edge-cloud collaborative computing can be employed, with tasks at the edge and tasks in the cloud. Edge tasks refer to lightweight computing tasks performed on local devices close to the data source, such as ATMs. At the edge, operations with high real-time requirements are prioritized. Tasks such as determining object evaluation attributes, gaze confidence, and touch evaluation attributes can all be categorized as edge tasks. These primarily involve data preprocessing and preliminary detection. A typical application scenario is as follows: during ATM operations, when a target enters their PIN at the ATM, the edge device performs real-time monitoring, completes authentication within 0.3 seconds, and outputs the final authentication result.

[0096] Risk response is the final execution layer of the system, responsible for taking tiered response measures based on the authentication results of multimodal decisions to balance security and operational smoothness. This module includes the following functions: Secondary authentication trigger: Initiating an enhanced authentication process based on the authentication risk score, such as SMS OTP or manual review; Real-time alarm reporting: Pushing high-risk events to the back-end risk control center to trigger manual intervention; Operation interceptor: Suspending or terminating suspicious transactions and freezing temporary account permissions; Log recorder: Storing multimodal data, decision-making processes, and response actions, supporting real-time auditing and model optimization.

[0097] The technical solution of this disclosure introduces an infrared-visible dual-mode camera, which uses a narrowband filter to suppress ambient light interference and combines it with 3D structured light dynamic threshold adjustment to adaptively adjust the depth difference threshold according to the light intensity. This improves the authentication rate in low-light scenes, thereby reducing the operation interruption rate of the target object, further optimizing the target object's experience, and reducing complaints. This invention also introduces spatiotemporal consistency verification, such as spatiotemporal alignment of pupil movement trajectory and touch data, which can improve security and increase trust.

[0098] Example 3

[0099] Figure 6 This is a schematic diagram of the structure of a transaction information evaluation device provided in an embodiment of this disclosure, as shown below. Figure 6As shown, the device includes: a data acquisition module 210, an object evaluation attribute determination module 220, a gaze confidence determination module 230, a touch evaluation attribute determination module 240, and a target object transaction attribute determination module 250.

[0100] A data acquisition module is used to acquire at least one video frame and at least one touch data during the transaction process; wherein, the touch data includes touch pressure data, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data when the target object contacts the device; an object evaluation attribute determination module is used to input the at least one video frame into a pre-trained first model to determine the object evaluation attribute of the target object in the at least one video frame; a gaze confidence determination module is used to determine the gaze confidence of the target object based on the at least one video frame and the center coordinates of at least one key area in the transaction interface; a touch evaluation attribute determination module is used to input the at least one touch data into a pre-trained second model to determine the touch evaluation attribute of the target object; a target object transaction attribute determination module is used to determine the target object transaction attribute based on the object evaluation attribute, the gaze confidence, the touch evaluation attribute, transaction scenario information, and an interaction risk assessment function; wherein, the transaction scenario information includes at least transaction time information, transaction location information, and transaction device information.

[0101] The technical solution of this disclosure involves acquiring at least one video frame and at least one touch data point during a transaction. The touch data includes touch pressure data, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data when the target object contacts the device. Then, the at least one video frame is input into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame. Based on the center coordinates of the at least one video frame and at least one key area in the transaction interface, the gaze confidence of the target object is determined. Further, for the at least one touch data point, the touch data is input into a pre-trained second model to determine the touch evaluation attributes of the target object. Finally, based on the object evaluation attributes, gaze confidence, touch evaluation attributes, transaction scenario information, and an interaction risk assessment function, the transaction attributes of the target object are determined. The transaction scenario information includes at least transaction time information, transaction location information, and transaction device information. This solution addresses the problems of vulnerability to brute-force attacks, low flexibility, adaptability, and security when evaluating transaction information using authentication methods such as usernames, passwords, and PIN codes. This invention determines the transaction attributes of a target object based on video frames and touch data during the transaction process. By combining multiple biometric identification methods, it enhances the accuracy and security of target object identity authentication, thereby improving and optimizing the customer's interactive experience.

[0102] Based on the above technical solutions, the touch data includes: data related to device contact generated when the target object clicks the device screen or button during the period from the start to the end of the transaction; the video frame includes image data containing the target object acquired by a camera device at a preset frame rate during the transaction process.

[0103] Based on the above technical solutions, the gaze confidence determination module 230 includes: a pupil coordinate determination submodule, a pupil movement trajectory determination submodule, a first distance determination submodule, and a gaze confidence determination submodule.

[0104] The pupil coordinate determination submodule is used to input the video frame into the third model for the at least one video frame and determine the pupil coordinates corresponding to the eyes of the target object in the video frame; the pupil motion trajectory determination submodule is used to smooth the at least one pupil coordinate and then determine the pupil motion trajectory; the first distance determination submodule is used to determine at least one first distance based on the pupil motion trajectory and the center coordinates of the at least one key region; and the gaze confidence determination submodule is used to determine the gaze confidence of the target object based on the at least one first distance and a preset threshold.

[0105] Based on the above technical solutions, the touch evaluation attribute determination module 240 includes a second model output submodule, which is used to input the touch pressure data, the contact area data, the touch duration data, the touch interval duration data, and the touch pressure change rate data into the second model for the at least one touch data, and to process the touch data based on the input layer, long short-term memory layer, random deactivation layer, fully connected layer, and output layer to output the touch evaluation attributes of the target object.

[0106] Based on the above technical solutions, the target object transaction attribute determination module 250 includes: a first risk assessment attribute determination submodule, a second risk assessment attribute determination submodule, a third risk assessment attribute determination submodule, and a transaction attribute determination submodule.

[0107] The first risk assessment attribute determination submodule is used to determine a first risk assessment attribute based on the object assessment attribute and the gaze confidence level; the second risk assessment attribute determination submodule is used to determine a second risk assessment attribute based on the gaze confidence level and the touch assessment attribute; the third risk assessment attribute determination submodule is used to determine a third risk assessment attribute based on the transaction time information, the transaction location information, and the transaction device information; the transaction attribute determination submodule is used to determine the target object transaction attribute based on the first risk assessment attribute, the second risk assessment attribute, the third risk assessment attribute, and the weights corresponding to each risk assessment attribute.

[0108] Based on the above technical solutions, when the gaze confidence is lower than a preset threshold, the device further includes: a voiceprint similarity assessment attribute output module and a first risk assessment attribute calculation module.

[0109] The voiceprint similarity assessment attribute output module is used to trigger the collection of voiceprint data of the target object, input the voiceprint data into the pre-trained fourth model, and output the voiceprint similarity assessment attribute; the first risk assessment attribute calculation module is used to determine the first risk assessment attribute based on the voiceprint similarity assessment attribute, the object assessment attribute, and the gaze confidence.

[0110] Based on the above technical solutions, after determining the transaction attributes of the target object, the device further includes: a first response measure execution module, a second response measure execution module, and a third response measure execution module.

[0111] The first response measure execution module is used to execute a first response measure when the transaction attribute of the target object is within the first attribute range; the second response measure execution module is used to execute a second response measure when the transaction attribute of the target object is within the second attribute range; and the third response measure execution module is used to execute a third response measure when the transaction attribute of the target object is within the third attribute range.

[0112] The transaction information evaluation device provided in this disclosure can execute the transaction information evaluation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0113] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0114] Example 4

[0115] Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Refer to the following... Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals). Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0116] like Figure 7 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.

[0117] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0118] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0119] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0120] The electronic device provided in this embodiment and the transaction information evaluation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0121] Example 5

[0122] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the transaction information evaluation method provided in the above embodiments.

[0123] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0124] In some implementations, the server may communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and may interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0125] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0126] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0127] Acquire at least one video frame and at least one touch data during the transaction process; wherein, the touch data includes touch pressure data when the target object contacts the device, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data;

[0128] The at least one video frame is input into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame;

[0129] The gaze confidence of the target object is determined based on the center coordinates of at least one video frame and at least one key area in the transaction interface.

[0130] For the at least one touch data, the touch data is input into a pre-trained second model to determine the touch evaluation attributes of the target object;

[0131] The target object's transaction attributes are determined based on the object evaluation attributes, the gaze confidence level, the touch evaluation attributes, the transaction scenario information, and the interaction risk assessment function; wherein, the transaction scenario information includes at least transaction time information, transaction location information, and transaction device information.

[0132] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0134] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0135] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0136] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0137] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0138] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0139] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for evaluating transaction information, characterized in that, The method includes: Acquire at least one video frame and at least one touch data during the transaction process; wherein, the touch data includes touch pressure data when the target object contacts the device, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data; The at least one video frame is input into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame; The gaze confidence of the target object is determined based on the center coordinates of at least one video frame and at least one key area in the transaction interface. For the at least one touch data, the touch data is input into a pre-trained second model to determine the touch evaluation attributes of the target object; The target object's transaction attributes are determined based on the object evaluation attributes, the gaze confidence level, the touch evaluation attributes, the transaction scenario information, and the interaction risk assessment function; wherein, the transaction scenario information includes at least transaction time information, transaction location information, and transaction device information.

2. The method according to claim 1, characterized in that, The touch data includes: Data related to device contact generated when the target object clicks on the device screen or buttons during the period from the start to the end of the transaction; the video frame includes image data containing the target object acquired by a camera device at a preset frame rate during the transaction process.

3. The method according to claim 1, characterized in that, Determining the gaze confidence of the target object based on the center coordinates of at least one video frame and at least one key region in the transaction interface includes: For the at least one video frame, the video frame is input into the third model to determine the pupil coordinates corresponding to the eyes of the target object in the video frame; After smoothing the at least one pupil coordinate, the pupil movement trajectory is determined; Based on the pupil movement trajectory and the center coordinates of the at least one key region, at least one first distance is determined; The gaze confidence of the target object is determined based on the at least one first distance and a preset threshold.

4. The method according to claim 1, characterized in that, For the at least one touch data, inputting the touch data into a pre-trained second model to determine the touch evaluation attributes of the target object includes: For the at least one touch data, the touch pressure data, the contact area data, the touch duration data, the touch interval duration data, and the touch pressure change rate data are input into the second model, and the touch data are processed based on the input layer, long short-term memory layer, random deactivation layer, fully connected layer, and output layer to output the touch evaluation attributes of the target object.

5. The method according to claim 1, characterized in that, The determination of the target object's transaction attributes based on the object evaluation attributes, the gaze confidence level, the touch evaluation attributes, transaction scenario information, and the interaction risk assessment function includes: Based on the object assessment attributes and the gaze confidence level, a first risk assessment attribute is determined; Based on the gaze confidence level and the touch assessment attribute, a second risk assessment attribute is determined; Based on the transaction time information, the transaction location information, and the transaction device information, a third risk assessment attribute is determined; The target object transaction attributes are determined based on the first risk assessment attribute, the second risk assessment attribute, the third risk assessment attribute, and the weights corresponding to each risk assessment attribute.

6. The method according to claim 5, characterized in that, When the gaze confidence is lower than a preset threshold, the method further includes: Trigger the collection of voiceprint data of the target object, input the voiceprint data into the pre-trained fourth model, and output the voiceprint similarity evaluation attribute; The first risk assessment attribute is determined based on the voiceprint similarity assessment attribute, the object assessment attribute, and the gaze confidence.

7. The method according to claim 1, characterized in that, After determining the transaction attributes of the target object, the method further includes: When the transaction attribute of the target object is within the first attribute range, the first response measure is executed; When the target object's transaction attribute falls within the range of the second attribute, the second response measure is executed; When the target object's transaction attribute falls within the third attribute range, a third response measure is executed.

8. A transaction information evaluation device, characterized in that, include: The data acquisition module is used to acquire at least one video frame and at least one touch data during the transaction process; wherein, the touch data includes touch pressure data when the target object contacts the device, contact area data, touch duration data, touch interval duration data, and touch pressure change rate data; An object evaluation attribute determination module is used to input the at least one video frame into a pre-trained first model to determine the object evaluation attributes of the target object in the at least one video frame; A gaze confidence determination module is used to determine the gaze confidence of the target object based on the center coordinates of at least one video frame and at least one key area in the transaction interface. A touch evaluation attribute determination module is used to input the touch data into a pre-trained second model for the at least one touch data to determine the touch evaluation attributes of the target object; The target object transaction attribute determination module is used to determine the target object transaction attributes based on the object evaluation attributes, the gaze confidence level, the touch evaluation attributes, transaction scenario information, and the interaction risk assessment function; wherein, the transaction scenario information includes at least transaction time information, transaction location information, and transaction device information.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the transaction information evaluation method as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the transaction information evaluation method as described in any one of claims 1-7.