Screen sliding behavior detection method and device, equipment and medium

By combining video classification technology and face detection matching technology, analyzing video frames and removing specific tags, the problem of inaccurate judgment of sliding screen action positions in the existing technology is solved, and high-precision sliding screen behavior detection is achieved to meet the operational authenticity needs in the financial and medical fields.

CN120198737APending Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349192.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately determine the specific location of the sliding screen action, and it is impossible to effectively distinguish the operating subjects in multi-character scenarios, resulting in high false alarm rates and missed alarm rates, which cannot meet the strict requirements for operational authenticity in the financial and medical fields.

Method used

By obtaining video data and target face images, detecting face feature information and determining the position of each face, analyzing video frames in combination with the sliding screen classification model, removing tags of unknown locations and target face positions, and the remaining tags are used to determine whether there is any screen-sliding behavior.

Benefits of technology

It realizes the accurate judgment of the occurrence location of the sliding screen action in multiple roles and scenarios, reduces false alarms and missed reports, improves detection accuracy and reliability, and meets the needs of the financial and medical fields for operational authenticity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198737A_ABST
    Figure CN120198737A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a screen sliding behavior detection method, which comprises the following steps: acquiring video data and a target face image, detecting face feature information from the video data, and determining the position of a target face in a video by combining the target face image, and extracting video frames from the video data according to a preset time interval, inputting the video frames into the screen sliding classification model to obtain classification tags, removing unknown position tags and target face position tags from the classification tags to obtain residual screen sliding action tags, and when the number of the residual screen sliding action tags is not null, determining that a screen sliding behavior exists. According to the method, the video classification technology and the face detection matching technology are combined, the specific occurrence position of the screen sliding action is accurately judged in the screen sliding action detection process, the operation main body in the multi-role scene is effectively distinguished, the false alarm problem caused by matching of the hand and the face is avoided, and the false alarm rate and the missing report rate can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and storage medium for detecting substitute screen sliding behavior. Background Art

[0002] In the dual-recording scenario in the financial field, customers usually need to complete business confirmation, risk disclosure or document signing through screen sliding operations to ensure the compliance of the business and the authenticity of the transaction. However, there are multiple technical problems in the verification of the authenticity of screen sliding operations in the prior art. On the one hand, many methods judge whether the screen sliding operation is completed by the customer himself through hand detection and face matching. However, due to the limited shooting range of the dual-recording video, usually only the area above the shoulders is included in the video, and the hand information is often missing or incomplete, resulting in a high false positive rate of screen sliding behavior. In addition, in a multi-person scenario, it is difficult for the prior art to accurately determine the specific occurrence position of the screen sliding action, making the verification effect of the authenticity of the screen sliding behavior unsatisfactory and unable to meet the strict requirements of financial operations for the authenticity of operations.

[0003] In the scenarios of telemedicine and electronic health record (EHR) management in the field of medical and health, patients or authorized persons usually need to complete key tasks through screen sliding operations, such as confirming medical records, signing surgical authorization forms or filling in health information. The prior art also faces technical bottlenecks in the detection of the authenticity of such remote operations. On the one hand, identity verification technology usually only relies on face detection and cannot accurately match the screen sliding behavior with the specific operator, and identity confusion is likely to occur in a multi-role scenario; on the other hand, the detection technology of screen sliding behavior lacks accurate determination of the position of the screen sliding action. Especially in the case of diverse device operation methods (such as tablets, mobile phones or dedicated medical terminals) and complex scenarios, the false positive rate and false negative rate of the existing methods are relatively high, affecting the verification effect of the authenticity of the operation.

[0004] Whether in the financial field or the medical and health field, the existing screen sliding behavior detection technologies show common deficiencies in the following technical problems: First, the prior art lacks multi-modal feature fusion of screen sliding actions and often only relies on single hand or face features, unable to accurately determine the specific occurrence position of the screen sliding action; second, the prior art fails to establish an effective verification mechanism for unknown labels. When an unknown state appears in the classification label, the method of directly removing the label is likely to lead to false positives or false negatives.

[0005] Therefore, there is an urgent need in the medical and health field and the financial field for a technology that can accurately determine the occurrence position of the screen sliding action in multi-role and multi-scenario situations, reduce false positives and false negatives, improve the detection accuracy, and at the same time ensure the applicability and reliability of the detection results, so as to meet the technical requirements of these two fields for the authenticity of interactive operations. Summary of the Invention

[0006] The main object of the present invention is to provide a method, device, equipment and storage medium for detecting proxy swiping behavior, aiming to solve the technical problems in the prior art that it is difficult to accurately determine the specific occurrence position of the swiping action, and it is impossible to effectively distinguish the operation subjects in a multi-role scenario, resulting in high false alarm rates and missed alarm rates, and the detection results are difficult to meet the actual application requirements.

[0007] To achieve the above object, the present invention provides a method for detecting proxy swiping behavior, including:

[0008] Obtain video data and a target face image;

[0009] Detect face feature information from the video data, and determine the face position of each face in the video based on the detected face feature information;

[0010] Match the target face image with the detected face feature information, and determine the target face position of the target face in the video based on the face position;

[0011] Extract video frames from the video data at a preset time interval, and input the video frames into a swiping classification model to obtain a classification label indicating the position of the swiping action;

[0012] Remove the label indicating an unknown position and the label corresponding to the target face position from the classification label to obtain the remaining swiping action labels in the classification label;

[0013] When the number of the remaining swiping action labels is not empty, it is determined that there is proxy swiping behavior.

[0014] Furthermore, to achieve the above object, the present invention provides a device for detecting proxy swiping behavior, including:

[0015] A data acquisition module for obtaining video data and a target face image;

[0016] A face detection module for detecting face feature information from the video data, and determining the face position of each face in the video based on the detected face feature information;

[0017] A target face matching module for matching the target face image with the detected face feature information, and determining the target face position of the target face in the video based on the face position;

[0018] A swiping classification module for extracting video frames from the video data at a preset time interval, and inputting the video frames into a swiping classification model to obtain a classification label indicating the position of the swiping action;

[0019] A label filtering module, configured to remove labels indicating unknown positions and labels corresponding to the target face position from the classification labels, so as to obtain the remaining swipe screen action labels in the classification labels;

[0020] A proxy swipe screen behavior detection module, configured to determine that there is a proxy swipe screen behavior when the number of the remaining swipe screen action labels is not empty.

[0021] Further, to achieve the above object, the present invention further provides a computer device, which includes a memory, a processor, and a proxy swipe screen behavior detection program stored in the memory and executable on the processor. When the proxy swipe screen behavior detection program is executed by the processor, the steps of the proxy swipe screen behavior detection method as described above are implemented.

[0022] Further, to achieve the above object, the present invention further provides a computer-readable storage medium, on which a proxy swipe screen behavior detection program is stored. When the proxy swipe screen behavior detection program is executed by a processor, the steps of the proxy swipe screen behavior detection method as described above are implemented.

[0023] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. A proxy swipe screen behavior detection method is disclosed, including: obtaining video data and a target face image, detecting face feature information from the video data and determining the position of each face in the video, matching the target face image with the detected face feature information and determining the position of the target face in the video, extracting video frames from the video data at a preset time interval, inputting the video frames into a swipe screen classification model to obtain classification labels, removing unknown position labels and target face position labels from the classification labels to obtain the remaining swipe screen action labels, and determining that there is a proxy swipe screen behavior when the number of the remaining swipe screen action labels is not empty. By combining video classification technology and face detection and matching technology, the present invention accurately determines the specific occurrence position of the swipe screen action during the swipe screen action detection process, effectively distinguishes the operation subjects in a multi-role scenario, avoids false alarms caused by hand-face matching, can reduce the false alarm rate and the missed alarm rate, and improves the applicability and reliability of the swipe screen behavior detection, so as to meet the verification requirements for the authenticity of operations in complex scenarios. Description of the Drawings

[0024] The following will further illustrate the present invention in conjunction with the drawings. In the drawings:

[0025] Figure 1 is a schematic diagram of an application environment of the proxy swipe screen behavior detection method in an embodiment of the present invention;

[0026] Figure 2 is a schematic flowchart of an embodiment of the proxy swipe screen behavior detection method of the present invention;

[0027] Figure 3 Schematic diagram of the functional modules of a preferred embodiment of the sliding screen behavior detection device of the present invention;

[0028] Figure 4 Schematic diagram of the structure of a computer device in an embodiment of the present invention;

[0029] Figure 5 Another schematic diagram of the structure of a computer device in an embodiment of the present invention. Detailed implementation manners

[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0031] The sliding screen behavior detection method provided by the embodiment of the present invention can be applied to an application environment such as Figure 1 . Among them, the client communicates with the server through the network. The server can obtain video data and target face images through the client, detect face feature information from the video data and determine the position of each face in the video, match the target face image with the detected face feature information and determine the position of the target face in the video, extract video frames from the video data at a preset time interval, input the video frames into the sliding screen classification model to obtain classification labels, remove the unknown position labels and target face position labels from the classification labels, and obtain the remaining sliding screen action labels. When the number of the remaining sliding screen action labels is not empty, it is determined that there is a sliding screen behavior on behalf of others. By combining video classification technology and face detection and matching technology, the present invention accurately determines the specific occurrence position of the sliding screen action during the sliding screen action detection process, effectively distinguishes the operation subjects in a multi-role scenario, avoids the false alarm problem caused by the matching of hands and faces, can reduce the false alarm rate and missed alarm rate, improves the applicability and reliability of the sliding screen behavior detection, and thus meets the verification requirements for the authenticity of operations in complex scenarios. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0032] Please refer to Figure 2 , Figure 2 which is a flowchart of an embodiment of the sliding screen behavior detection method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described herein may be executed in a different order.

[0033] As Figure 2 shown, the sliding screen behavior detection method proposed by the present invention includes the following steps:

[0034] S10. Obtain video data and a target face image;

[0035] In this embodiment, obtaining video data means continuously capturing dynamic images in a target scene through a collection device (such as a camera or other video recording device) to form raw video data. The video data may include images of single or multiple persons and records a video frame sequence and time information.

[0036] The target scene is recorded in real time through a video collection device, which can be a fixed device (such as a camera installed on a wall) or a portable device (such as a smartphone or tablet). During the video collection process, parameters such as resolution and frame rate need to be configured to ensure the picture quality and the accuracy of subsequent processing. The raw video data can be stored in a local storage device or transmitted to a central processing system through a network.

[0037] Obtaining a target face image means obtaining a static image of the target face that needs to be focused on through a specified method. The target face image can be sourced from a photo provided by the user or a static face region extracted from a video frame. This face image needs to include clear key features (such as eyes, nose, mouth, etc.).

[0038] The target face image can be obtained through the following methods:

[0039] Static upload method: The user manually uploads a photo or selects a target face image from an existing database.

[0040] Dynamic extraction method: The system automatically extracts the target face region from a certain frame of the video data, and after locating the face region through a face detection algorithm, crops and saves it as the target face image.

[0041] Automatic matching method: The system automatically extracts the target face image from the database according to known user data (such as ID photos or historical records).

[0042] To ensure the quality of the obtained face image, preprocessing is usually performed, such as clarity enhancement, denoising, and contrast adjustment.

[0043] Example illustration: In a remote medical consultation scenario, the system collects interactive video data of the patient and the doctor through a camera. At the same time, the patient uploads an ID photo or a face image during medical record registration as the source of the target face image before the consultation starts. Ensure that during the consultation process, the system can lock on to the patient himself / herself and accurately identify whether there is a proxy operation situation during the swipe screen behavior detection.

[0044] In the double-recording scenario of financial institutions, customers interact with sales staff. The system collects video data through a camera fixedly installed in the meeting room and extracts the target face image from the ID card photo uploaded by the customer. This target face image will be used in the subsequent face matching process to confirm whether the swiping operation is completed by the customer himself and avoid the risk of proxy operation.

[0045] Through the above steps, the efficient acquisition and accurate matching of video data and the target face image can be achieved. The dynamic images recorded by the video acquisition device provide the basic data source for the subsequent swiping behavior detection. By obtaining the target face image, the target object can be locked in a multi-role scenario, avoiding unnecessary misjudgment and confusion, and providing technical support for subsequent face matching and swiping classification. Using multiple methods such as dynamic extraction or static input to obtain the target face image can adapt to different application scenarios and improve the flexibility of the system and the applicability of detection.

[0046] S20, Detect face feature information from the video data, and determine the face position of each face in the video based on the detected face feature information;

[0047] In this embodiment, continuous video frames are extracted from the video data, and face detection operations are performed on each frame to extract the face regions and related feature information (such as key point features of eyes, nose, mouth, etc.) that appear in the video. The detected face feature information includes the bounding box coordinates and key point information of the face, which are used to represent the position and structural features of each face in the video frame.

[0048] Through the video frame decomposition technology, the video data is converted into a series of continuous static video frames, and the extraction frequency can be set according to requirements (such as extracting several frames per second). A deep learning model (such as the Haar cascade classifier based on CNN, YOLO, or MTCNN) is used to detect the face regions in each frame, and the bounding box position and size of each face are output.

[0049] The detected face regions are further analyzed, and the key point feature information of the face is extracted through a key point detection algorithm (such as Dlib or FacialLandmark Detection), including the specific coordinate positions of the eyes, nose, and mouth, for subsequent precise positioning and matching.

[0050] By analyzing the face feature information, the spatial position of each face in the video frame is determined. The face position is determined by calculating the center point coordinates of the face bounding box or other geometric center points.

[0051] According to the upper-left and lower-right coordinates of the face bounding box, calculate the center point coordinates of each face, or use the position of key points (such as the tip of the nose) as the center point of the face position. In a multi-person scenario, number the face positions in the order from left to right according to the center point coordinates of each face to distinguish the specific positions of multiple characters in the video.

[0052] In addition, tracking algorithms (such as Kalman filtering or SORT algorithm) can also be used to track the dynamic positions of each face in the video based on the face feature information between frames. Ensure that the face feature information of the same person can be associated in consecutive frames, and ensure consistent sorting of multiple characters by assigning a unified number.

[0053] Example: In a telemedicine scenario, patients need to interact with doctors through video to complete online consultations, medical record confirmations, and signing of surgical authorizations, etc. For example, in a typical scenario, the video screen includes the patient, the doctor, and the patient's family members. Through the face detection module, the system can identify all participants in the video and extract the feature information of each face, such as key point features like the bounding box, eyes, nose, and mouth.

[0054] In such a multi-character scenario, the face detection module sorts all participants by position, such as "the patient is on the left, the doctor is in the middle, and the family member is on the right", and numbers and labels each face position. In this way, the identity and position of the patient can be clarified, so as to accurately lock the executor of the screen sliding action during the screen sliding behavior detection, ensuring that key operations (such as medical record confirmation or surgical authorization) are completed by the patient himself, and avoiding the situation of proxy operation by the patient's family members or others.

[0055] Similarly, in the financial field, the dual recording (recording audio and video) technology is widely used in the whole process of selling high-risk financial products. For example, customers need to sign investment agreements or risk disclosure documents on the screen sliding device. In a typical scenario, the dual recording video usually includes the customer and the salesperson. In the video, the customer needs to slide the screen to complete relevant confirmation operations, while the salesperson can only explain beside. Any proxy operation behavior will lead to doubts about the compliance of the transaction.

[0056] In this scenario, the face detection module detects the feature information of all faces from the dual recording video. For example, the customer is on the left side of the screen and the salesperson is on the right side of the screen. The system labels and sorts the roles according to the face positions (such as "the customer is the left-side role and the salesperson is the right-side role"). Subsequently, by matching with the target face image (the customer's ID photo or the previously uploaded avatar), the system further confirms the specific position of the customer in the screen.

[0057] By identifying the customer's identity and the position of the screen swipe, the system can accurately detect whether the screen swipe action is completed by the customer himself. If the screen swipe behavior occurs outside the customer's position, for example, a screen swipe operation is detected at the position of the salesperson, the system can immediately mark it as a suspected proxy operation behavior, reminding the financial institution to conduct key audits and reducing transaction risks. This process reduces the workload of manual quality inspection while improving the detection accuracy and efficiency.

[0058] By detecting the facial feature information from the video data and determining the face position based on the detected information, it is possible to accurately identify the position of each face in the video, providing basic support for subsequent face matching and screen swipe behavior detection. In a multi-role scenario, by sorting and numbering the face positions, it helps to distinguish the operations of different roles and avoid misjudgments caused by confusion.

[0059] S30, match the target face image with the detected facial feature information, and determine the target face position of the target face in the video based on the face position;

[0060] In this embodiment, key point feature information in the target face image is extracted for matching with the facial feature information in the detected video frame. The key point feature information usually includes stable feature points such as the eyes, nose, and mouth of the face and their spatial coordinates.

[0061] Use a face key point detection algorithm (such as Dlib, Facial Landmark Detection, etc.) to process the target face image and extract its key point coordinates. During the extraction process, the following steps can be carried out: preprocess the target face image (such as grayscale conversion, normalization) to enhance the detection effect; detect the face bounding box and extract the key point coordinates from it; use key point encoding to generate a feature description vector for subsequent matching operations.

[0062] Extract the feature information of each face in all the face regions detected in the video frame, including key point features (such as the positions of the eyes, nose, and mouth) and face bounding box information.

[0063] Through a face detection algorithm, extract key point features for each face region in the video frame and generate corresponding feature vectors to form a feature information set. The feature information needs to contain the position identifier of each face so that it can be corresponding to the specific position during the subsequent matching process.

[0064] Calculate the similarity between the feature vector of the target face image and the feature vector of each detected face, and judge whether the target face matches a certain face in the video frame through the similarity score.

[0065] Use distance measurement methods (such as Euclidean distance, cosine similarity) to calculate the similarity of two sets of feature vectors. Set a similarity threshold to filter out the detected faces with similarity higher than the threshold as possible target faces. If there are multiple faces that meet the similarity condition, the best matching result can be determined by further comparison (such as weighted feature scores or multimodal fusion).

[0066] Determine the specific position of the target face in the video frame according to the matching result (such as the center point of the bounding box or the coordinates of key points), and mark the position of this face as the position of the target face.

[0067] Extract the bounding box coordinates or the center point position of the matching face from the matching result. Combine multi-frame dynamic tracking technology to ensure the consistency of the target face position. Even if the target face moves in multiple frames of video, its position can still be accurately located. Output the position of the target face for use by the subsequent swipe screen action detection module.

[0068] Example: In the electronic medical record management system of a hospital, patients may need to sign medical record confirmation letters, surgical consent forms or privacy authorization documents through electronic devices (such as tablets or self-service terminals) during the medical treatment process. To ensure that these critical operations are completed by the patients themselves, the system extracts the ID card photos or historical medical treatment photos from the patients' medical treatment records as the target face images.

[0069] During the operation, the system collects the facial images of patients in real time through the camera and detects all the face feature information that appears in the video data. By matching the key point feature information of the target face image with the face feature information in the video frame, the system can accurately locate the position of the patient in the video and track the identity change of the patient in real time.

[0070] For example, in the waiting area or the self-service terminal area, patients may temporarily leave or be accompanied by family members. In this case, the system can dynamically identify each face in the picture and confirm whether the patient is the executor of the swipe screen operation through the similarity matching algorithm. If the system detects that the swipe screen operation comes from a family member or someone else, a warning will be triggered to remind the staff to conduct further verification.

[0071] It can effectively prevent family members or others from signing medical documents on behalf of patients, ensure the authenticity and legality of electronic medical records, and avoid medical disputes caused by identity errors. In addition, the system can also dynamically track the movement of patients in front of the camera to ensure the continuity of patient identity verification. Even if the patient's position changes during the operation, their identity can be accurately recognized.

[0072] By matching the target face image with the facial feature information detected in the video and determining the position of the target face based on the matching result, the target operating subject can be accurately identified and located, providing reliable input data for subsequent sliding action detection.

[0073] S40, extracting video frames from the video data according to a preset time interval, and inputting the video frames into a screen sliding classification model to obtain a classification label representing a screen sliding action position;

[0074] In this embodiment, static video frames are extracted from the video data at a set time interval to form a video frame sequence. The selection of the time interval should be based on the action duration of the sliding operation and the computing power of the detection model, which can ensure the integrity of the action features and reduce the amount of redundant calculations.

[0075] Time interval setting: Set the time interval according to the average duration of the sliding action and the processing capacity of the model, for example, 0.1 second or 0.5 second.

[0076] Frame extraction: Traverse the video data and extract video frames in sequence according to the set time interval to form a frame sequence arranged in chronological order to ensure that the action trajectory is captured continuously.

[0077] Frame screening: The extracted video frames are screened for quality, and blurry or noisy frames are removed to ensure that the input frame sequence can represent the main features of the sliding action.

[0078] The extracted video frames are input into the trained screen sliding classification model frame by frame. The model recognizes and classifies the screen sliding action for each frame by extracting the action feature information in the video frames.

[0079] The extracted frames are normalized (e.g., resized, grayed) to meet the input requirements of the classification model. The convolutional layer of the sliding classification model is used to extract information such as sliding tracks, hand movements, and background features in the video frames. Based on the extracted features, the classification model outputs the classification results for each frame (e.g., the sliding position is "left", "right", or "uncertain").

[0080] The slide classification model analyzes the motion features of each frame and generates a corresponding classification label. The classification label indicates the location of the slide action, such as "left", "right", "center" or "uncertain".

[0081] According to the output of the classification model, the sliding action result of each frame is marked as a classification label, and the corresponding position (such as left sliding, right sliding, etc.) is recorded. The classification labels of all frames are aggregated to form the sliding action classification result for use in subsequent steps. The output results of the classification model can be further optimized through multi-frame voting or time series analysis to avoid the impact of single-frame misjudgment on the overall result.

[0082] Example illustration: In a telemedicine scenario, a patient may need to complete surgical authorization or signing of an electronic medical record by swiping the screen. The camera captures video data of the patient's screen-swiping operation in real time, and the system extracts video frames from the video at a fixed time interval (such as 0.2 seconds) and inputs them into the screen-swiping classification model. The classification model identifies the specific position of the patient's screen-swiping action (such as "left side" or "right side") and generates classification labels.

[0083] If the position detected by the screen-swiping action is consistent with the target face position of the patient, it can be preliminarily determined that the screen-swiping operation is completed by the patient himself; if the detection result shows that the position of the screen-swiping action does not match the patient's position, the risk of proxy operation can be further investigated to ensure the authenticity and legality of critical operations.

[0084] Similarly, in a financial double-recording scenario, a customer needs to sign relevant documents for high-risk products by swiping the screen. The system extracts frames from the double-recording video at a preset time interval (such as 0.1 seconds) and inputs them into the screen-swiping classification model for action classification. The model outputs the position label of the screen-swiping action for each frame (such as "right side" corresponding to the position of the salesperson), and integrates the classification results of all frames. If there are multiple screen-swiping actions in the position of the salesperson in the classification results, it is determined that there is a proxy operation behavior, and the financial institution is prompted to conduct manual verification to ensure the compliance of the transaction and the rights and interests of the customer.

[0085] By extracting video frames from video data at a preset time interval and inputting them into the screen-swiping classification model, key action features in the video can be obtained with relatively low computational overhead, reducing the interference of redundant data. The screen-swiping classification model can accurately identify the position of the screen-swiping action for each frame, and further improve the robustness of the classification results through the integration of frame-level classification labels. It can not only ensure the accuracy of identifying the screen-swiping action, but also significantly improve the real-time performance and efficiency of detection, laying a foundation for the subsequent judgment of the screen-swiping behavior.

[0086] S50, remove the labels indicating unknown positions and the labels corresponding to the target face position from the classification labels to obtain the remaining screen-swiping action labels in the classification labels;

[0087] In this embodiment, the label indicating an unknown position refers to a classification label for which the screen-swiping classification model cannot clearly locate the position of the screen-swiping action, and is usually used to label the uncertainty or fuzzy judgment of the classification model.

[0088] Traverse the classification label set and filter out the labels marked as "unknown position", which may be caused by fuzzy areas or interference factors (such as hand occlusion, irregular actions) of the screen-swiping action. Whether a label needs to be removed can be further confirmed through the labeling rules of the label or the confidence threshold of the classification model. For example, when the confidence is lower than a certain threshold, the label is classified as "unknown position".

[0089] The labels corresponding to the target face position refer to the labels where the swiping action occurs at the position of the target face. The removal of these labels aims to exclude the swiping actions of the target face itself, thus focusing on possible proxy swiping behaviors.

[0090] Filter and classify the labels in the classification labels that match the target face position according to the target face position (such as left, middle, or right). If the swiping action position of the classification label is consistent with the target face position, it is marked as a label corresponding to the target face position.

[0091] By removing the filtered unknown position labels and the labels corresponding to the target face position from the set of classification labels, the remaining swiping action labels are obtained, and these remaining labels are used for subsequent judgment of swiping behaviors.

[0092] Use set operations to remove the filtered "unknown position labels" and "labels corresponding to the target face position" from the set of classification labels. Update the set of classification labels to ensure that the remaining labels only include swiping action positions that have nothing to do with the target face position, reducing detection interference factors.

[0093] The remaining swiping action labels are the classification labels that are retained after screening and removal, and these labels are used to judge whether there is a proxy swiping behavior.

[0094] The remaining labels should include all swiping action positions where the swiping action may occur outside the target face position (such as left or right). Ensure the integrity and accuracy of the remaining labels to provide a reliable basis for subsequent judgment of proxy swiping behaviors.

[0095] Example illustration: In a telemedicine scenario, a patient needs to sign a surgical authorization document through a swiping device. The system detects the swiping actions in the video through a swiping classification model and outputs classification labels (such as "left", "right", or "unknown"). During the swiping action screening process, the system removes the labels indicating unknown positions to avoid misjudgment caused by model uncertainty. At the same time, the system removes the swiping action labels that match the patient's position to focus on possible proxy operations performed by doctors or family members. The screened swiping action labels are used to judge the occurrence of proxy swiping behaviors, thus ensuring the authenticity of the authorization process.

[0096] In a financial double-recording scenario, a customer needs to swipe the screen to confirm the contract content, and the video footage includes the customer and the salesperson. The system generates swiping action labels (such as "left", "middle", "right", or "unknown") through a swiping classification model. During the screening process, the system first removes the labels indicating unknown positions (caused by model uncertainty or hand occlusion), and then removes the swiping labels corresponding to the customer's position. If the remaining swiping action labels show that the swiping behavior occurs at the salesperson's position, the system preliminarily judges that there is a proxy operation behavior, providing a basis for subsequent detection.

[0097] By removing the tags representing unknown positions and the tags corresponding to the target face position from the classification tags, the interference factors in the swiping classification results can be effectively reduced, focusing on the swiping actions outside the target face position, thereby improving the accuracy of detecting proxy swiping behavior. The removal of the unknown position tags solves the influence of model uncertainty on the detection results, while the removal of the target face position tags eliminates the redundant information related to the target face in the swiping classification, ensuring the reliability and pertinence of the detection results.

[0098] S60. When the number of the remaining swiping action tags is not empty, it is determined that there is proxy swiping behavior.

[0099] In this embodiment, the remaining swiping action tags refer to the set of swiping action tags remaining after removing the tags representing unknown positions and the tags corresponding to the target face position from the classification tags. By judging the number of this set, it is decided whether there is a swiping action occurring in the area outside the target face position.

[0100] Count the filtered set of swiping action tags. If the number is greater than zero, it means that there is a swiping action outside the target face position. If the set is empty, it is determined that no swiping action occurs outside the target face position, and all swiping behaviors are completed by the target face.

[0101] When the number of the remaining swiping action tags is not empty, it means that the swiping action occurs in the area outside the target face position, which may be a swiping operation performed by someone else on behalf of the target face.

[0102] Based on the determination result of the remaining tag number, it is determined whether there is proxy swiping behavior. If the number is greater than zero, the determination logic of proxy swiping behavior is triggered; otherwise, a conclusion that there is no proxy swiping behavior is output.

[0103] The system generates a detection conclusion according to the determination result and transmits it to subsequent modules (such as a log recording module or a detection report generation module).

[0104] When it is determined that there is proxy swiping behavior, mark the current video segment as a video related to proxy swiping behavior for subsequent analysis and processing. If there is no proxy swiping behavior, record the normal swiping operation without further processing.

[0105] Example illustration: In a telemedicine scenario, a patient needs to confirm medical records or surgical authorization documents by swiping the screen. The system detects the position of the swiping action through a swiping screen classification model and removes the label corresponding to the target face position. If the number of remaining swiping action labels is not empty, it indicates that the swiping action may be completed by a doctor or family member, and the system will determine that there is a behavior of proxy swiping the screen. For example, if it is detected in some frames that the swiping action occurs outside the target face position, the system will trigger a warning to remind the doctor or medical institution to conduct further verification, so as to ensure the authenticity of the medical record confirmation operation.

[0106] In a dual-recording scenario, a customer needs to confirm contract content or sign an agreement through a swiping screen device. The system detects the swiping action through a swiping screen classification model and removes the label corresponding to the customer's position. If the remaining swiping action labels show that the swiping behavior occurs at the salesperson's position, the system determines that there is a behavior of proxy swiping the screen. For example, if the customer is in the left position and the swiping action label shows in the right (salesperson's position), the system automatically marks this operation as a proxy operation behavior and generates a warning to prompt the financial institution to conduct key audits to prevent non-compliant behaviors.

[0107] By determining the number of the remaining swiping action labels, it can be quickly judged whether the swiping behavior occurs in an area outside the target face, so as to effectively identify the existence of the proxy swiping behavior. Through the precise screening and quantity determination mechanism, the interference of redundant information on the detection result is reduced, and the accuracy and efficiency of proxy operation detection are improved. In addition, through the clear output of the determination result, the system can realize the automatic identification of the proxy swiping behavior, providing a reliable basis for subsequent recording and analysis.

[0108] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical health and fintech. It discloses a method for detecting proxy swiping behavior, including: obtaining video data and a target face image, detecting face feature information from the video data, and determining the position of the target face in the video in combination with the target face image, extracting video frames from the video data at a preset time interval, inputting the video frames into a swiping screen classification model to obtain classification labels, removing the unknown position labels and the target face position labels from the classification labels to obtain the remaining swiping action labels, and determining that there is a proxy swiping behavior when the number of the remaining swiping action labels is not empty. By combining video classification technology and face detection and matching technology, the present invention accurately judges the specific occurrence position of the swiping action during the swiping action detection process, effectively distinguishes the operation subjects in a multi-role scenario, avoids false alarms caused by hand-face matching, and can reduce the false alarm rate and the missed alarm rate.

[0109] In one embodiment, the above S20 includes:

[0110] S201, extracting a video frame sequence from the video data and preprocessing the video frame sequence;

[0111] S202. Perform face detection operations on each video frame in the pre - processed video frame sequence to obtain a set of face feature information;

[0112] S203. Extract key - point feature information including parts such as eyes, nose, and / or mouth from the set of face feature information;

[0113] S204. Based on the extracted key - point feature information, screen the set of face feature information to remove false face feature information and retain valid face feature information;

[0114] S205. According to the retained valid face feature information, determine the face positions of each face in the video and label the face positions.

[0115] In this embodiment, extracting the video frame sequence is to decompose the video data into a set of static frames so that subsequent face detection operations can be performed frame by frame. The pre - processing process includes adjusting the frame size, normalization processing, and noise reduction, aiming to improve the accuracy and efficiency of face detection. In the implementation process, first, video decoding technology is used to extract a continuous frame sequence from the video at a set frame rate. Then, the resolution of the frame is adjusted to match the input requirements of the classification model, and the pixel values of the frame are normalized, for example, scaling the pixel values to the interval [0, 1]. To eliminate background noise caused by the shooting environment, a filter (such as Gaussian filter) can be used to perform noise reduction processing on the frame.

[0116] In face detection operations, it is necessary to process video frames frame by frame to extract the bounding box positions and feature information of all faces. Detection methods usually adopt deep - learning algorithms, such as MTCNN, YOLO, or RetinaFace. These algorithms can quickly locate the positions of each face in the video frame. In the implementation process, the detection algorithm first scans each frame and identifies all face regions. Each detection result usually includes the coordinate information of the face bounding box (such as the coordinates of the upper - left and lower - right points). All detected face feature information will be stored in a set for subsequent screening and matching operations.

[0117] Key - point feature information refers to the core points of the face, such as the corners of the eyes, the tip of the nose, the corners of the mouth, etc. By extracting these stable features, the accuracy of face screening and position calculation can be enhanced. In the implementation process, key - point detection algorithms (such as Dlib or Facial Landmark Detection) can be used to process each face region. These algorithms usually return a list of two - dimensional coordinates of face key points. For example, the most representative points (such as eyes and nose) are selected from 68 feature points to reduce the computational burden and improve the screening efficiency.

[0118] By analyzing the distribution of key point features, the system can effectively distinguish real human faces from false detection results (such as interference objects in the background). False human faces are usually identified due to abnormal feature distributions (such as asymmetry or insufficient number of points). During the implementation process, the system first verifies whether the distribution of key points conforms to the geometric features of a normal human face. For example, by calculating the distance between the eyes and the mouth, it is determined whether it is within a reasonable range. If the distribution of key points of certain human faces is abnormal or incomplete, they are removed from the set of human face feature information. The remaining valid human face feature information will be used as the input for subsequent operations.

[0119] Based on the filtered valid human face feature information, the system can calculate the position of each human face in the video frame. For example, the position of each human face is determined by the center point coordinates of the bounding box. For a multi-role scenario, the system sorts the human faces according to the horizontal position of the center points of the human faces and labels the positions of each human face (such as "left", "center", "right"). During the implementation process, the system first calculates the center point coordinates of the bounding box of each human face, and then sorts them according to the horizontal distribution of these coordinates. If three human faces are detected in the picture, the system will label them as "left", "center", and "right" in sequence and record their position information for subsequent swipe screen action detection operations.

[0120] In this embodiment, through the detection, screening, and labeling of human face feature information in the video frame, different roles in the picture can be accurately distinguished and their specific positions marked, providing accurate input data for swipe screen behavior detection. The system can effectively remove the interference of false human faces and dynamically track the position changes in a multi-role scenario, adapting to diverse application scenarios and improving the reliability and efficiency of detection.

[0121] In one embodiment, the above S205 includes:

[0122] S2051, based on the valid human face feature information, locate the human face region in each video frame and obtain the bounding box coordinates of the human face region;

[0123] S2052, analyze the position center point of each human face region in the video frame according to the bounding box coordinates of the human face region;

[0124] S2053, based on the position center point of each human face region in the video frame, number and sort the positions of the human faces in each video frame.

[0125] In this embodiment, locating the human face region is to determine the position of each human face in the video frame according to the valid human face feature information (such as the detected bounding box coordinates or key point distribution). The bounding box coordinates are usually defined by four values, representing the pixel positions of the upper left corner and the lower right corner of the human face region respectively.

[0126] Process each frame using a deep learning face detection model (such as YOLO, MTCNN, or RetinaFace) to extract the coordinate information of the face bounding box. If there is noise in the bounding box information, the position of the bounding box in the face region can be further optimized through correction algorithms (such as IoU threshold or confidence screening) to ensure accurate positioning results.

[0127] By calculating the center point of the bounding box, the specific position of the face in the video frame can be determined, and this can be used as the basis for subsequent sorting. The center point coordinates are the median of the bounding box coordinates. For example:

[0128] Center point X = (Upper left corner X coordinate + Lower right corner X coordinate) / 2

[0129] Center point Y = (Upper left corner Y coordinate + Lower right corner Y coordinate) / 2

[0130] Use simple geometric formulas to calculate the center point coordinates of each face region in turn. Compare the calculated center point positions with the spatial distribution of the video frame (such as frame width and height) to mark their relative positions (such as left, centered, or right). If there is an overlap in the face bounding boxes, the distribution of key points can be further analyzed to adjust the calculation method of the center point to avoid incorrect labeling.

[0131] By analyzing the center point coordinates of each face, the system sorts the face positions in order from left to right and assigns a unique numbered label (such as "left", "center", "right"). This numbering and sorting method is particularly important in multi-person scenarios and can clearly mark the specific position of each face in the picture.

[0132] Extract the X coordinate values of the center points of all faces and sort them in ascending order. Assign a label or number to each face. For example, the first face is labeled "left" and the last face is labeled "right". If the center point coordinates of multiple faces are close or overlapping, the bounding box size or key point features can be combined for further differentiation to ensure that the sorting result is unique and reasonable.

[0133] In this embodiment, by locating the face regions in each video frame and calculating their center point positions, the system can accurately label the specific positions of each face in the picture and provide a clear sorting result for multi-person scenarios. It effectively avoids the risk of role confusion in multi-person scenarios and ensures that subsequent swipe screen action detection can focus on specific target face regions.

[0134] In one embodiment, the above S30 includes:

[0135] S301, extracting the key point feature information of the target face image;

[0136] S302. Obtain the key-point feature information of each face from the detected face feature information;

[0137] S303. Perform similarity matching between the key-point feature information of the target face image and the key-point feature information of each face, and determine the target face whose similarity exceeds the preset similarity threshold;

[0138] S304. Based on the face positions of each face in the video, determine the target face position corresponding to the target face in the video.

[0139] In this embodiment, extracting the key-point feature information in the target face image is to standardize the feature data of the target face, facilitating subsequent matching with the detected face features. These key points usually include facial core points such as the corners of the eyes, the tip of the nose, and the corners of the mouth.

[0140] Use a key-point detection algorithm (such as Dlib or Facial Landmark Detection) to analyze the target face image and extract the key-point information of the target face. The extracted key points are represented in the form of two-dimensional coordinates, covering the main feature points of the target face (such as eyes, nose, mouth, etc.). Normalize the extracted key-point features to adapt to subsequent similarity calculations.

[0141] The detected face feature information usually contains the key-point features of all faces in multiple video frames. To ensure accurate matching with the target face, it is necessary to extract the key-point feature information of each face.

[0142] For the face regions detected in each video frame, call the key-point detection algorithm to extract the corresponding key-point features. Store the key-point information for each face region. These key points usually include 68 or fewer important facial feature points. Generate a set containing the key-point features of all faces in all video frames for one-by-one matching with the target face.

[0143] The key-point feature information of the target face is matched one by one with the key-point feature information of each detected face, and the target face is identified through similarity calculation. The face whose similarity exceeds the preset threshold is considered the successfully matched target face.

[0144] Similarity calculation: Use the Euclidean distance, cosine similarity, or deep learning embedding representation method (such as the embedding vector generated by FaceNet) to calculate the matching degree between the key points of the target face and the detected face key points.

[0145] Threshold setting: Set a reasonable similarity threshold (such as 0.8) to distinguish between successful and failed matches. The selection of the threshold is based on the tolerance requirements of the system and the complexity of the detection scenario.

[0146] Matching result: If the similarity of a certain face exceeds the preset threshold, mark the face as the target face and save the key point features of the match at the same time.

[0147] After matching the target face, the system marks the spatial position information of the target face in the video as the position of the target face, such as "left", "middle", or "right" in the picture.

[0148] Position acquisition: Extract the bounding box coordinates of the successfully matched target face in the video frame.

[0149] Center point calculation: Determine the specific position of the target face in the picture through the center point coordinates of the bounding box.

[0150] Position annotation: Compare the position of the target face with the positions of other faces in the video frame and annotate the position of the target face, such as "left", "middle", "right".

[0151] Position tracking: If the target face moves in multiple consecutive frames, dynamically update its position annotation to ensure the temporal consistency of the matching result.

[0152] Example illustration: In a telemedicine scenario, the patient needs to confirm the surgery authorization or health data through a screen swiping operation. The system extracts the target face features from the historical photos provided by the patient and compares them with the faces detected in the real-time video to quickly confirm the position of the patient in the picture (such as the left side of the picture). This position confirmation provides the basic data support for the authenticity verification of the subsequent screen swiping operation behavior and prevents doctors or family members from operating on behalf of the patient.

[0153] In a dual-recording scenario, the customer confirms the contract terms through a screen swiping device. The system extracts the key point features of the customer from the previously provided customer identification photos and compares them one by one with the multiple faces detected in the dual-recording video, and finally confirms the position of the customer in the picture (such as the "middle" position). This confirmation step effectively ensures that the screen swiping behavior is completed by the customer himself and provides technical guarantee for the legality and compliance of the transaction.

[0154] In this embodiment, by matching the key point features of the target face with the faces detected in the video frame one by one, the system can accurately identify and locate the specific position of the target face in the video. By setting the similarity threshold, false matches can be effectively filtered, and the accuracy of target face recognition can be improved. In a multi-person scenario, different roles in the picture can be dynamically distinguished, and the spatial position of the target face can be accurately tracked, providing accurate data input for the detection of subsequent screen swiping actions.

[0155] In one embodiment, the above S40 includes:

[0156] S401, Sequentially extract video frames from the video data according to a preset time interval to form a video frame sequence;

[0157] S402, Input the video frames in the video frame sequence into the swiping screen classification model in chronological order;

[0158] S403, Generate a frame-level classification label for each video frame based on the swiping screen classification model, where the frame-level classification label is used to represent the occurrence position of the swiping screen action in a single video frame;

[0159] S404, Integrate the frame-level classification labels corresponding to each video frame to form a classification label representing the position of the swiping screen action.

[0160] In this embodiment, the system sequentially extracts frames from the video data at a set time interval to form a continuous frame sequence. The time interval can be set according to the detection requirements of the swiping screen action to ensure that the video data is reasonably decomposed into multiple static frames. In the implementation process, the system first uses a video decoding tool to parse the video data frame by frame and extracts specific frames according to the set time interval. For example, a frame can be extracted every certain time or a fixed number of frames, so as to generate a set of video frame sequences arranged in chronological order. After extraction, these frame sequences are stored as image data or memory objects for subsequent swiping screen action classification analysis.

[0161] To ensure that the swiping screen classification model can accurately process the frame data, the video frame sequence needs to be input into the model one by one in chronological order. The model can perform an independent analysis of the swiping screen action position for each frame. In the implementation process, each frame in the frame sequence will be loaded and formatted. For example, operations such as resolution adjustment and pixel normalization can be performed on the frame data to meet the input requirements of the swiping screen classification model. After processing, they are sequentially passed to the classification model according to the time order of the frames to ensure that the classification results can reflect the temporal continuity and spatial consistency of the swiping screen action.

[0162] The swiping screen classification model will perform an independent analysis on each video frame to generate a frame-level classification label, which is used to describe the specific occurrence position of the swiping screen action in a single frame. The frame-level classification label usually includes options such as "left", "middle", "right", or "unknown". In the implementation process, the swiping screen classification model first extracts spatial features (such as the swiping trajectory or finger position) from each input frame. Subsequently, these features are classified through a classifier module (such as a Softmax layer) in the model, and the position label of the swiping screen action is output. For example, if the swiping screen action occurs on the left side of the screen, the model will generate a "left" label for this frame. The system will bind these labels to the corresponding frame timestamps to form frame-level classification records.

[0163] The system integrates the frame-level classification labels of all video frames to generate a set of classification labels that describe the position distribution of the entire video's swiping action. These sets are used to represent the overall distribution characteristics of the swiping action in the video. During the implementation process, the system summarizes the frame-level classification labels and performs statistics based on the occurrence frequency or chronological order. For example, the occurrence frequency of a certain swiping position can be counted or the time distribution of the labels can be calculated. During the integration process, the system also screens and denoises the classification results, removing labels with low confidence or potential interference. Finally, the generated set of classification labels will serve as the core output of the swiping action detection result, providing a basis for subsequent analysis and judgment.

[0164] In this embodiment, by extracting video frames at preset time intervals and sequentially inputting them into the swiping classification model for analysis, the position where the swiping action occurs can be efficiently detected. The frame-level classification labels provide fine-grained swiping position information, while the integrated classification labels reflect the distribution characteristics of the swiping action in the entire video. Combining the classification method with dynamic inter-frame features, the system can adapt to the diverse manifestation forms of the swiping action while ensuring the accuracy and robustness of the detection results.

[0165] In one embodiment, before the above S50, it further includes:

[0166] S501, determining preliminary unknown labels indicating unknown positions from the classification labels;

[0167] S502, locating the video frames at the unknown positions corresponding to the preliminary unknown labels, and extracting swiping trajectory features, inter-frame change features, and time series features from the video frames at the unknown positions;

[0168] S503, based on the swiping trajectory features, inter-frame change features, and time series features, verifying the preliminary unknown labels through a label verification module to obtain a verification result;

[0169] S504, if the verification result shows that the preliminary unknown label indicates an unknown position, then determining the preliminary unknown label as a label indicating an unknown position;

[0170] S505, if the verification result shows that the preliminary unknown label indicates a known position, then determining the preliminary unknown label as a label indicating a known position.

[0171] In this embodiment, the preliminary unknown label refers to a label that cannot be clearly classified as the position of the swiping action in the classification labels, usually representing the uncertain part in the detection result. This step provides input data for subsequent verification and screening operations.

[0172] Traverse the classification labels, filter out the parts that are not clearly marked with the swipe screen position among all the labels, and mark them as preliminary unknown labels. The preliminary unknown labels can be determined by the confidence threshold of the model classification result. For example, when the confidence is lower than the set value (such as lower than 50%), it is marked as an unknown position. All preliminary unknown labels are stored in a list to be processed for subsequent operations.

[0173] Locating video frames with unknown positions is to obtain the video frame data related to the preliminary unknown labels and extract their features for the verification module to analyze.

[0174] Locating video frames: Locate the relevant video frames in the video data according to the timestamps corresponding to the preliminary unknown labels.

[0175] Extracting swipe screen trajectory features: Analyze the possible swipe screen action paths (such as swipe directions and trajectory shapes) in the video frames and extract the coordinates of feature points.

[0176] Extracting inter-frame change features: Compare the pixel changes between the video frames with unknown positions and their adjacent frames (such as by calculating using the optical flow method) to judge the coherence of the swipe screen action and the rationality of the trajectory.

[0177] Extracting time series features: Extract the time series information related to the swipe screen action, such as the start time, duration, and interval time of the action.

[0178] The label verification module is an independent analysis module used to further verify the authenticity of the preliminary unknown labels. By combining the swipe screen trajectory features, inter-frame change features, and time series features, the verification module can determine whether the preliminary unknown labels represent unknown positions.

[0179] Verifying the swipe screen trajectory: Perform rule matching on the extracted swipe screen trajectory features to judge whether the trajectory conforms to the standard form of the swipe screen action (such as straight-line swiping or curve swiping).

[0180] Verifying inter-frame changes: Check whether the pixel differences between the video frames with unknown positions and their adjacent frames match the characteristics of the swipe screen action.

[0181] Verifying time series: Analyze whether the time characteristics of the swipe screen action are reasonable, such as whether the action duration is within the normal range of human operations.

[0182] Comprehensive determination: Input the verification results of the above features into the decision-making module (such as logical rules or machine learning models) to output the verification results.

[0183] When the verification result indicates that the initially unknown tags cannot be classified into known swiping positions, the system will re-label them as tags representing unknown positions to retain their uncertainty information. Based on the output result of the verification module, the tags that meet the definition of unknown positions are filtered out. These tags are uniformly labeled as unknown position tags and stored in the classification tag set.

[0184] When the verification result indicates that the initially unknown tags can be classified into known swiping positions, the system will re-label them as specific known position tags. Based on the output result of the verification module, the tags that meet the definition of known positions are filtered out. According to the frame-level classification result, they are labeled with specific positions (such as "left", "middle", or "right"). The classification tag set is updated by replacing the initially unknown tags with the corresponding known position tags.

[0185] In this embodiment, by introducing the positioning and verification mechanism for initially unknown tags, the system can effectively solve the problem of handling the uncertain part in the classification tags. The verification module comprehensively verifies the initially unknown tags using swiping trajectory features, inter-frame change features, and time series features to ensure the accuracy and reliability of the classification results. The final classification tag set can more accurately reflect the position distribution of swiping actions and provide high-quality data for subsequent detection and analysis.

[0186] In one embodiment, after S60 above, it further includes:

[0187] S701, marking the swiping video frames corresponding to the remaining swiping action tags in the video data and recording the time points corresponding to the swiping video frames in the video data;

[0188] S702, storing the marked swiping video frames and the corresponding time points in the detection log;

[0189] S703, generating a detection report on the swiping behavior based on the detection log;

[0190] S704, inputting the detection report into the feedback module of the swiping classification model to update the training data of the swiping classification model.

[0191] In this embodiment, the video frames corresponding to the swiping action tags are located by timestamps, and the swiping action tag marks are added to these video frames, while recording the time points when the swiping behavior occurs. The implementation method includes locating the relevant video frames from the video data according to the timestamps of the swiping action tags, adding the swiping action tag marks to the metadata of the located video frames to distinguish specific swiping behaviors, and recording the time points of each video frame to ensure that these time points can support subsequent behavior traceability and analysis.

[0192] The marked swipe-screen video frames and their time-point information are sorted out and stored as detection logs for subsequent analysis and model optimization. The implementation methods include formatting the swipe-screen video frames and their time-point information to generate structured detection logs; storing the sorted logs in local files or databases to ensure convenience for subsequent retrieval and use; verifying the integrity and consistency of the log data before storage to avoid omissions or errors.

[0193] The content in the detection logs is used to generate a detection report describing the swipe-screen behavior. The implementation methods include extracting swipe-screen behavior-related information from the detection logs, counting the time range, frequency, and location of the occurrence of the swipe-screen behavior; generating the detection report in a standardized format according to the extraction and statistical results, such as including the time period when the swipe-screen behavior occurred, the swipe-screen location, and video frame examples; saving the generated report in a file format (such as PDF or HTML), or directly providing it for the user to view.

[0194] The data in the detection report is used to optimize the training of the swipe-screen classification model and improve the classification accuracy and robustness. The implementation methods include extracting swipe-screen-related video frames and action labels from the detection report to generate new training samples; performing data augmentation operations on the extracted training data, such as adjusting the frame brightness, changing the frame angle, adding noise, etc., to expand the applicability of the model; inputting the newly generated training data into the classification model for incremental training or retraining to optimize the classification performance of the model.

[0195] Example illustration: In a telemedicine scenario, patients need to confirm health data or authorization letters through swipe-screen operations. To ensure the authenticity of the swipe-screen operations, the system detects the patients' swipe-screen behaviors. When a surrogate swipe-screen behavior is detected, the system automatically marks the relevant swipe-screen video frames and records the time points when the swipe-screen behavior occurs. For example, when a family member or caregiver completes the swipe-screen operation on behalf of the patient, the system generates a detection report by analyzing the swipe-screen trajectory and the location of the swipe-screen action. The report details the occurrence time, location, and specific video frame examples of the swipe-screen behavior. Subsequently, this report is fed back to the swipe-screen classification model to further optimize the classification ability of the model and ensure more accurate and reliable future detections.

[0196] During the dual recording process of financial transactions, customers need to confirm transaction terms through a screen swiping operation. The system detects the customer's screen swiping behavior through a screen swiping classification model. When it detects that the screen swiping action occurs at a non-customer position (such as the salesperson's position), the system will mark these screen swiping behaviors as abnormal behaviors and generate a detection log. For example, during a certain transaction, the screen swiping action continuously occurs at the salesperson's position on the left side of the customer. The system records the relevant screen swiping video frames and time points and generates a detection report. This report includes the specific time period, location distribution of the screen swiping action, and an example of the screen swiping trajectory, which is used for transaction compliance review and fed back as training data to the screen swiping classification model to improve the accuracy of the model.

[0197] In this embodiment, by marking the video frames corresponding to the screen swiping action tags and recording their time points, the system can accurately track the specific location and time when the screen swiping behavior occurs. The generation and storage of the detection log provide a basis for traceability and verification, while the detection report systematically summarizes the specific characteristics of the proxy screen swiping behavior, providing support for review and further analysis. Inputting the detection report as feedback data into the training module of the classification model effectively improves the classification accuracy and adaptability of the model, providing a reliable guarantee for subsequent detections.

[0198] In one embodiment, a proxy screen swiping behavior detection device is provided, and this proxy screen swiping behavior detection device corresponds one-to-one with the proxy screen swiping behavior detection method in the above embodiment. Refer to Figure 3 , Figure 3 which is a schematic diagram of the functional modules of a preferred embodiment of the proxy screen swiping behavior detection device of the present invention. The data acquisition module 10, the face detection module 20, the target face matching module 30, the screen swiping classification module 40, the label filtering module 50, and the proxy screen swiping behavior detection module 60. The detailed description of each functional module is as follows:

[0199] The data acquisition module 10 is used to acquire video data and target face images;

[0200] The face detection module 20 is used to detect face feature information from the video data and determine the face position of each face in the video based on the detected face feature information;

[0201] The target face matching module 30 is used to match the target face image with the detected face feature information and determine the target face position of the target face in the video based on the face position;

[0202] The screen swiping classification module 40 is used to extract video frames from the video data at a preset time interval and input the video frames into the screen swiping classification model to obtain classification labels representing the screen swiping action positions;

[0203] A label filtering module 50, configured to remove labels representing unknown positions and labels corresponding to the target face positions from the classification labels, so as to obtain the remaining swipe screen action labels in the classification labels;

[0204] A proxy swipe behavior detection module 60, configured to determine that there is a proxy swipe behavior when the number of the remaining swipe screen action labels is not empty.

[0205] In one embodiment, the face detection module 20 is specifically configured to:

[0206] Extract a video frame sequence from the video data, and preprocess the video frame sequence;

[0207] Perform a face detection operation on each video frame in the preprocessed video frame sequence to obtain a set of face feature information;

[0208] Extract key point feature information including parts such as eyes, nose, and / or mouth from the set of face feature information;

[0209] Based on the extracted key point feature information, screen the set of face feature information to remove false face feature information and retain valid face feature information;

[0210] According to the retained valid face feature information, determine the face positions of each face in the video, and label the face positions.

[0211] In one embodiment, the face detection module 20 is specifically configured to:

[0212] Based on the valid face feature information, locate the face regions in each video frame to obtain the bounding box coordinates of the face regions;

[0213] Analyze the position center points of each face region in the video frame according to the bounding box coordinates of the face regions;

[0214] Based on the position center points of each face region in the video frame, number and sort the face positions in each video frame.

[0215] In one embodiment, the target face matching module 30 is specifically configured to:

[0216] Extract the key point feature information of the target face image;

[0217] Obtain the key point feature information of each face from the detected face feature information;

[0218] Perform a similarity match between the key point feature information of the target face image and the key point feature information of each face to determine the target face whose similarity exceeds a preset similarity threshold;

[0219] Determine the target face position corresponding to the target face in the video based on the face position of each face in the video.

[0220] In one embodiment, the screen swiping classification module 40 is specifically configured to:

[0221] Extract video frames sequentially from the video data according to a preset time interval to form a video frame sequence;

[0222] Input the video frames in the video frame sequence into the screen swiping classification model in chronological order;

[0223] Generate frame-level classification labels for each video frame based on the screen swiping classification model, where the frame-level classification labels are used to represent the occurrence position of the screen swiping action in a single video frame;

[0224] Integrate the frame-level classification labels corresponding to each video frame to form a classification label representing the position of the screen swiping action.

[0225] In one embodiment, the label filtering module 50 is specifically configured to:

[0226] Determine a preliminary unknown label representing an unknown position from the classification labels;

[0227] Locate the video frames at the unknown positions corresponding to the preliminary unknown label, and extract the screen swiping trajectory features, inter-frame change features, and time series features from the video frames at the unknown positions;

[0228] Based on the screen swiping trajectory features, inter-frame change features, and time series features, verify the preliminary unknown label through the label verification module to obtain a verification result;

[0229] If the verification result shows that the preliminary unknown label represents an unknown position, determine the preliminary unknown label as a label representing an unknown position;

[0230] If the verification result shows that the preliminary unknown label represents a known position, determine the preliminary unknown label as a label representing a known position.

[0231] In one embodiment, the surrogate screen swiping behavior detection module 60 is specifically configured to:

[0232] Mark the screen swiping video frames corresponding to the remaining screen swiping action labels in the video data, and record the time points corresponding to the screen swiping video frames in the video data;

[0233] Store the marked screen swiping video frames and the corresponding time points in the detection log;

[0234] Generate a detection report on the surrogate screen swiping behavior according to the detection log;

[0235] Input the detection report into the feedback module of the sliding screen classification model to update the training data of the sliding screen classification model.

[0236] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a method for detecting proxy sliding screen behavior.

[0237] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a method for detecting proxy sliding screen behavior

[0238] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0239] Obtain video data and a target face image;

[0240] Detect face feature information from the video data, and determine the face position of each face in the video based on the detected face feature information;

[0241] Match the target face image with the detected face feature information, and determine the target face position of the target face in the video based on the face position;

[0242] Extract video frames from the video data according to a preset time interval, and input the video frames into a swiping screen classification model to obtain classification labels indicating the positions of swiping screen actions;

[0243] Remove the labels indicating unknown positions and the labels corresponding to the target face positions from the classification labels to obtain the remaining swiping screen action labels in the classification labels;

[0244] When the number of the remaining swiping screen action labels is not empty, it is determined that there is a proxy swiping behavior.

[0245] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0246] Obtain video data and a target face image;

[0247] Detect face feature information from the video data, and determine the face positions of each face in the video based on the detected face feature information;

[0248] Match the target face image with the detected face feature information, and determine the target face position of the target face in the video based on the face positions;

[0249] Extract video frames from the video data according to a preset time interval, and input the video frames into a swiping screen classification model to obtain classification labels indicating the positions of swiping screen actions;

[0250] Remove the labels indicating unknown positions and the labels corresponding to the target face positions from the classification labels to obtain the remaining swiping screen action labels in the classification labels;

[0251] When the number of the remaining swiping screen action labels is not empty, it is determined that there is a proxy swiping behavior.

[0252] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0253] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0254] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0255] It should be noted that if non-company software tools or components appear in the embodiments of this application, they are only used for example introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for detecting screen sliding behavior, characterized in that: The following steps are involved: Obtain video data and target face image; Detecting facial feature information from the video data, and determining the facial position of each face in the video based on the detected facial feature information; Matching the target face image with the detected face feature information, and determining the target face position of the target face in the video based on the face position; Extracting video frames from the video data according to a preset time interval, and inputting the video frames into a screen sliding classification model to obtain a classification label representing a screen sliding action position; Remove the label indicating the unknown position and the label corresponding to the target face position from the classification label to obtain the remaining sliding action label in the classification label; When the number of the remaining sliding screen action tags is not empty, it is determined that there is a substitute sliding screen behavior.

2. The method for detecting screen sliding behavior according to claim 1, characterized in that: Detecting facial feature information from the video data, and determining the facial position of each face in the video based on the detected facial feature information, including: Extracting a video frame sequence from the video data, and preprocessing the video frame sequence; Performing a face detection operation on each video frame in the preprocessed video frame sequence to obtain a face feature information set; Extracting key point feature information including eyes, nose and / or mouth from the face feature information set; Based on the extracted key point feature information, the face feature information set is screened to remove false face feature information and retain valid face feature information; The face position of each face in the video is determined based on the retained effective face feature information, and the face position is marked.

3. The method for detecting screen sliding behavior according to claim 2, characterized in that: Determine the face position of each face in the video according to the retained valid face feature information, and mark the face position, including: Based on the effective facial feature information, locate the face area in each video frame to obtain the bounding box coordinates of the face area; Analyzing the position center point of each face region in the video frame according to the bounding box coordinates of the face region; Based on the position center point of each face area in the video frame, the face positions in each video frame are numbered and sorted.

4. The method for detecting screen sliding behavior according to claim 1, characterized in that: Matching the target face image with the detected face feature information, and determining the target face position of the target face in the video based on the face position, including: Extracting key point feature information of the target face image; Obtain key point feature information of each face from the detected face feature information; Performing similarity matching on the key point feature information of the target face image and the key point feature information of each face, and determining a target face whose similarity exceeds a preset similarity threshold; Based on the face position of each face in the video, a target face position corresponding to the target face in the video is determined.

5. The method for detecting screen sliding behavior according to claim 1, characterized in that: Extracting video frames from the video data according to a preset time interval, and inputting the video frames into a screen sliding classification model to obtain a classification label indicating a screen sliding action position, including: Sequentially extracting video frames from the video data according to a preset time interval to form a video frame sequence; Inputting the video frames in the video frame sequence into the screen sliding classification model in chronological order; Generating a frame-level classification label for each video frame based on the screen sliding classification model, wherein the frame-level classification label is used to indicate the occurrence position of the screen sliding action in a single video frame; The frame-level classification labels corresponding to each video frame are integrated to form a classification label representing the position of the screen sliding action.

6. The method for detecting screen sliding behavior according to claim 1, characterized in that: Before removing the label indicating the unknown position and the label corresponding to the target face position from the classification label to obtain the remaining sliding action label in the classification label, the method further includes: determining a preliminary unknown label representing an unknown position from the classification labels; Locating the unknown position video frame corresponding to the preliminary unknown tag, and extracting the sliding track feature, the inter-frame change feature and the time series feature from the unknown position video frame; Based on the sliding track features, inter-frame change features and time series features, the preliminary unknown tags are verified by a tag verification module to obtain a verification result; If the verification result shows that the preliminary unknown tag represents an unknown position, the preliminary unknown tag is determined as a tag representing an unknown position; If the verification result shows that the preliminary unknown tag indicates a known location, the preliminary unknown tag is determined to be a tag indicating a known location.

7. The method for detecting screen sliding behavior according to claim 1, characterized in that: When the number of the remaining sliding screen action tags is not empty, after determining that there is a substitute sliding screen behavior, the method further includes: Marking the sliding screen video frames corresponding to the remaining sliding screen action tags in the video data, and recording the time points corresponding to the sliding screen video frames in the video data; The marked sliding screen video frames and the corresponding time points are stored in the detection log; Generate a detection report of the screen sliding behavior according to the detection log; The detection report is input into a feedback module of the screen sliding classification model to update the training data of the screen sliding classification model.

8. A device for detecting screen sliding behavior, characterized in that: The device for detecting screen sliding behavior comprises: A data acquisition module, used to acquire video data and target face images; A face detection module, used to detect facial feature information from the video data, and determine the facial position of each face in the video based on the detected facial feature information; A target face matching module, used to match the target face image with the detected face feature information, and determine the target face position of the target face in the video based on the face position; A screen sliding classification module, used to extract video frames from the video data according to a preset time interval, and input the video frames into a screen sliding classification model to obtain a classification label representing a position of a screen sliding action; A label filtering module, used to remove labels indicating unknown positions and labels corresponding to the target face positions from the classification labels, and obtain the remaining sliding action labels in the classification labels; The proxy screen sliding behavior detection module is used to determine the presence of proxy screen sliding behavior when the number of the remaining screen sliding action tags is not empty.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a proxy screen sliding behavior detection program stored in the memory and executable on the processor. When the proxy screen sliding behavior detection program is executed by the processor, the steps of the proxy screen sliding behavior detection method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores a proxy screen sliding behavior detection program, and when the proxy screen sliding behavior detection program is executed by the processor, the steps of the proxy screen sliding behavior detection method according to any one of claims 1 to 7 are implemented.