Abnormal behavior monitoring method, apparatus, device, medium, and program product

CN122839014APending Publication Date: 2026-09-29INDUSTRIAL AND COMMERCIAL BANK OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610880318.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,该技术路线在实际应用中存在明显不足

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839014A_ABST
    Figure CN122839014A_ABST
Patent Text Reader

Abstract

The application provides an abnormal behavior monitoring method and device, equipment, a storage medium and a program product, which can be applied to the field of artificial intelligence technology. The method comprises the following steps: in response to the hand of a target object entering a target area, determining a space-time overlap index between the hand of the target object and a target screen; inputting voice text of the target object, log text associated with the target object and action features of the target object into a multi-modal large model in a continuous time window, performing cross-modal interaction reasoning through a self-attention mechanism of the multi-modal large model, and obtaining comprehensive feature representation of fused multi-modal semantic information; mapping the comprehensive feature representation through linear transformation and a normalization index function, and outputting an abnormal intention probability; weighting and fusing the space-time overlap index and the abnormal intention probability to generate an abnormal score; and triggering an abnormal behavior alarm in the case that the abnormal score is higher than a preset threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and specifically to an abnormal behavior monitoring method, device, equipment, medium, and program product. Background Technology

[0002] Currently, bank branches and wealth management rooms mainly rely on fixed surveillance cameras for passive recording, with some more advanced systems incorporating conventional computer vision algorithms. These solutions typically set a simple physical distance threshold, such as an intersection-union ratio (IU) greater than zero between the bounding boxes of a hand and a mobile phone. Once overlap is detected and the duration exceeds the set threshold, a violation alarm is triggered.

[0003] However, this technical approach has significant shortcomings in practical applications. First, the cameras in the financial planning room are mostly overhead or side-overhead angles. Normal and compliant physical interactions such as employees handing materials and guiding screens are easily overlapped with "directly touching a mobile phone" under two-dimensional visual projection, resulting in a very high false alarm rate and seriously affecting system availability. Second, the existing algorithm lacks semantic intent understanding capabilities, judging solely based on geometric position and motion trajectory. It cannot distinguish the subtle spatiotemporal differences between "employees pointing at the screen to guide customers" and "employees directly swiping to input," nor can it understand the business context when the behavior occurs, such as dialogue content and business process nodes, leading to numerous misjudgments. In addition, the time dimension modeling method is simplistic, relying only on simple sliding windows to accumulate time, ignoring the dynamic sequence characteristics of the occurrence, development, and decay of actions. This makes it unable to accurately capture sudden, instantaneous customer-assisted actions, resulting in a high risk of missed detections. Summary of the Invention

[0004] In view of the above problems, embodiments of this application provide an abnormal behavior monitoring method, apparatus, device, medium, and program product.

[0005] According to a first aspect of this application, an abnormal behavior monitoring method is provided, the method comprising: responding to a target object's hand entering a target area, determining a spatiotemporal overlap index between the target object's hand and a target screen, wherein the target area is a three-dimensional space of a preset size at the location of the target screen, and the spatiotemporal overlap index is a quantitative indicator characterizing the intensity of physical interaction between the target object's hand and the target screen within a continuous time window; inputting the target object's speech text, log text associated with the target object, and the target object's action features within the continuous time window into a multimodal large model, performing cross-modal interaction reasoning through the multimodal large model's self-attention mechanism to obtain a comprehensive feature representation incorporating multimodal semantic information, wherein the log text is log text from a business system related to the target object; mapping the comprehensive feature representation through a linear transformation and a normalized exponential function to output an abnormal intent probability; weighting and fusing the spatiotemporal overlap index and the abnormal intent probability to generate an abnormal score; and triggering an abnormal behavior alarm if the abnormal score is higher than a preset threshold.

[0006] According to an embodiment of this application, determining the spatiotemporal overlap index between the target object's hand and the target screen in response to the target object's hand entering the target area includes: acquiring video images within the target area within the continuous time window in response to the target object's hand entering the target area; calculating the Euclidean distance between the target object's hand and the target screen based on the video images, and calculating the three-dimensional volume intersection-union ratio between the bounding sphere of the target object's hand and the bounding box of the target screen; constructing an instantaneous association strength factor based on the Euclidean distance and the three-dimensional volume intersection-union ratio; determining a directional weight based on the approach direction of the target object's hand relative to the target screen; and performing time integration on the product of the instantaneous association strength factor and the directional weight within the continuous time window to obtain the spatiotemporal overlap index.

[0007] According to an embodiment of this application, the step of constructing the instantaneous correlation strength factor based on the Euclidean distance and the three-dimensional volume intersection-union ratio includes: performing an exponential decay transformation on the Euclidean distance and then weighting and summing it with the three-dimensional volume intersection-union ratio to obtain the instantaneous correlation strength factor.

[0008] According to an embodiment of this application, when the hand of the target object moves toward the target screen or remains stationary, the value of the directional weight approaches 1; when the speed at which the hand of the target object moves away from the target screen exceeds a preset threshold, the value of the directional weight approaches 0.

[0009] According to an embodiment of this application, the step of inputting the speech text of the target object, the log text associated with the target object, and the action features of the target object into a multimodal large model within the continuous time window, and performing cross-modal interactive reasoning through the self-attention mechanism of the multimodal large model to obtain a comprehensive feature representation that integrates multimodal semantic information includes: extracting the visual behavior tokens corresponding to the action features, the speech text tokens corresponding to the speech text, and the log text tokens corresponding to the log text; concatenating the cue tokens with the visual behavior tokens, the speech text tokens, and the log text tokens in sequence to form an input sequence, wherein the cue tokens are tokenized representations of preset instruction texts, and the instruction texts are used to define constraints on the probability of abnormal intent output by the multimodal large model; feeding the input sequence into the multimodal large model, calculating the attention score between all token pairs in the input sequence through the multi-layer self-attention mechanism of the multimodal large model; using the attention score to perform weighted fusion of the original representations of each token, and updating the representation of each token; and obtaining a comprehensive feature representation that integrates multimodal semantic information through the stacking of multiple self-attention mechanisms.

[0010] According to an embodiment of this application, the step of mapping the comprehensive feature representation through linear transformation and normalized exponential function to output the abnormal intent probability includes: mapping the comprehensive feature representation to the vector space corresponding to the abnormal operation to obtain the original logical value; and calculating the original logical value through a normalized exponential function to obtain the abnormal intent probability that the hand action of the target object belongs to the abnormal operation.

[0011] According to an embodiment of this application, the step of weightedly fusing the spatiotemporal overlap index and the abnormal intent probability to generate an anomaly score includes: determining a fusion weight coefficient based on the Euclidean distance between the target object's hand and the target screen; normalizing the spatiotemporal overlap index and multiplying it by the abnormal intent probability by the fusion weight coefficient and its complement, and then adding the products to obtain the anomaly score; wherein the fusion weight coefficient is negatively correlated with the Euclidean distance.

[0012] According to a second aspect of this application, an abnormal behavior monitoring device is provided, the device comprising: an overlap index calculation module, configured to determine a spatiotemporal overlap index between the hand of a target object and a target screen in response to the hand of a target object entering a target area, wherein the target area is a three-dimensional space of a preset size at the location of the target screen, and the spatiotemporal overlap index is a quantitative indicator characterizing the intensity of physical interaction between the hand of the target object and the target screen within a continuous time window; a comprehensive feature extraction module, configured to input the speech text of the target object, the log text associated with the target object, and the action features of the target object within the continuous time window into a multimodal large model, and perform cross-modal interaction reasoning through the self-attention mechanism of the multimodal large model to obtain a comprehensive feature representation that integrates multimodal semantic information, wherein the log text is the log text of a business system related to the target object; an abnormal probability calculation module, configured to output an abnormal intent probability by linear transformation and normalized exponential function mapping of the comprehensive feature representation; an abnormal scoring module, configured to weightedly fuse the spatiotemporal overlap index and the abnormal intent probability to generate an abnormal score; and an alarm triggering module, configured to trigger an abnormal behavior alarm when the abnormal score is higher than a preset threshold.

[0013] According to a third aspect of this application, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0014] According to a fourth aspect of this application, a computer-readable storage medium is also provided, on which a computer program or instructions are stored, wherein the computer program or instructions, when executed by a processor, implement the steps of the above-described method.

[0015] According to a fifth aspect of this application, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description

[0016] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0017] Figure 1 The illustrations depict application scenarios of the abnormal behavior monitoring method, apparatus, device, medium, and program products according to embodiments of this application.

[0018] Figure 2 A flowchart illustrating an abnormal behavior monitoring method according to an embodiment of this application is shown schematically.

[0019] Figure 3This schematic diagram illustrates the structural block diagram of an abnormal behavior monitoring device according to an embodiment of this application;

[0020] Figure 4 A block diagram of an electronic device suitable for implementing an abnormal behavior monitoring method according to an embodiment of this application is shown schematically. Detailed Implementation

[0021] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0024] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0025] It should be noted that the abnormal behavior monitoring method and device provided in this application can be used in business transaction scenarios in the fintech field, or in business transaction scenarios in any field other than fintech. The application field of the abnormal behavior monitoring method and device provided in this application is not limited.

[0026] For solutions involving personal information, if it cannot be avoided, please include the following statement in the instruction manual:

[0027] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0028] In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided in this application all provide users with corresponding operation entry points for users to choose to agree to or reject the automated decision results; if the user chooses to reject, the process enters the expert decision-making process.

[0029] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0030] It's important to note that the term "neural network" can refer to a machine learning network based on deep learning. A neural network processes input and provides corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between them. Neural networks used in deep learning applications often include many hidden layers, increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer serves as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output becomes the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each processing the input from the layer above.

[0031] It should be understood that machine learning generally includes three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values ​​to determine the corresponding output.

[0032] In one or more embodiments described herein, the term "large model" can refer to a deep learning model with a large number of model parameters, which can include hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. Large models can also be called foundational models or basic models. They are pre-trained using large-scale unlabeled corpora to produce pre-trained models with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability, such as large language models and multimodal pre-trained models. It should be understood that in practical applications, large models only require a small number of samples to fine-tune the pre-trained model before being applied to different tasks. Large models can be widely used in natural language processing, computer vision, and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering, image captioning, and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. Major application scenarios for large models can include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0033] In financial service scenarios such as bank wealth management, when wealth managers conduct business face-to-face with clients, clients often place their mobile phones on the table and open banking applications. To prevent wealth managers from operating the phone on behalf of clients without their knowledge or under duress, such as entering verification codes or transferring funds, it is necessary to monitor the wealth manager's hand movements in real time. This embodiment uses a bank wealth management office as a specific application environment. The office ceiling is equipped with a depth camera for capturing 3D spatial information and a microphone array for capturing audio. The system interfaces with the bank's intermediary business system in real time to obtain business node logs.

[0034] This application provides an abnormal behavior monitoring method that achieves accurate identification of abnormal hand touch screen behavior by integrating three-dimensional spatial physical interaction quantization and multimodal semantic reasoning.

[0035] Figure 1The illustration shows an application scenario diagram of the abnormal behavior monitoring method, apparatus, device, medium, and program product according to embodiments of this application.

[0036] In this example environment 100, a depth camera 110 and a microphone array 120 are installed on the ceiling of the financial advisory room. The depth camera 110 is used to capture three-dimensional spatial video images containing financial advisor A and customer B, and the microphone array 120 is used to capture on-site voice signals. A business terminal 130 is set up in the financial advisory room, which communicates with the bank's intermediary business system to obtain the operation log text of the current business handled by financial advisor A in real time (e.g., "Current node: Customer APP self-service login in progress"). Customer B places his mobile phone 140 on the table, and the mobile phone screen is used as the target screen in this application.

[0037] In some embodiments, the depth camera 110 can be any type of device capable of acquiring three-dimensional spatial information, including but not limited to a binocular camera, a structured light camera, a time-of-flight (ToF) camera, a lidar, or any combination thereof. The microphone array 120 can be an omnidirectional or directional microphone, supporting multi-channel audio acquisition and sound source localization. The business terminal 130 can be a desktop computer, a laptop computer, a tablet computer, a mobile handheld terminal, or any combination thereof, and it synchronizes data in real time with the bank's back-office business system via a network.

[0038] In the embodiments of this application, the system deploys an abnormal behavior monitoring engine 150. This engine can run locally on the business terminal 130, or on the server 160, or some components can be deployed in the cloud. The server 160 can be a mainframe, an edge computing node, a computing device in a cloud environment, etc. The 3D images, audio, and log data collected by the depth camera 110, microphone array 120, and business terminal 130 are transmitted to the abnormal behavior monitoring engine 150 in real time.

[0039] In some embodiments, the abnormal behavior monitoring engine 150 includes at least two core processing streams: a physical space quantization stream and a multimodal semantic inference stream. The physical space quantization stream calculates the spatiotemporal overlap index between the financial manager A's hand and the customer's mobile phone screen 140 based on the 3D images acquired by the depth camera 110. The multimodal semantic inference stream integrates a multimodal large model, which receives voice text acquired by the microphone array 120, log text acquired by the business terminal 130, and motion features extracted from the 3D skeletal sequence as input. It performs cross-modal interactive inference through a self-attention mechanism and outputs the probability of abnormal intent and a natural language explanation. The outputs of the two processing streams are weighted and fused to obtain an abnormal score. When the score exceeds a preset threshold, an alarm is triggered, and the alarm information can be pushed to the business terminal 130 or a remote monitoring center.

[0040] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and are not intended to limit the scope of this application in any way.

[0041] Figure 2 A flowchart illustrating an abnormal behavior monitoring method according to an embodiment of this application is shown schematically. Figure 2 As shown, the abnormal behavior monitoring method 200 according to the embodiments of this application may include steps S210 to S250.

[0042] In step S210, in response to the target object's hand entering the target area, the spatiotemporal overlap index between the target object's hand and the target screen is determined.

[0043] In this embodiment, the target object refers to the financial manager, i.e., the bank employee being monitored. The target area is a three-dimensional cube space of a preset size centered on the location of the target screen (e.g., a customer's mobile phone screen or a banking system screen). The spatiotemporal overlap index is a quantitative indicator used to characterize the intensity of physical interaction between the target object's hand and the target screen within a continuous time window (e.g., 2 seconds).

[0044] The system uses a depth camera to acquire the 3D coordinates of the hand and the screen in real time. Based on the Euclidean distance between the hand and the screen, the degree of 3D volume overlap, and the movement trend of the hand relative to the screen, it integrates within a time window to obtain a value between 0 and 1. The larger the value, the closer the physical contact between the hand and the screen and the longer the duration. For example, when the financial manager's fingertips are only 2 centimeters away from the screen and remain stationary, the spatiotemporal overlap index is approximately 0.92; if the hand is quickly retracted, the index drops rapidly.

[0045] In step S220, the speech text of the target object, the log text associated with the target object, and the action features of the target object within the continuous time window are input into the multimodal big model. Cross-modal interactive reasoning is performed through the self-attention mechanism of the multimodal big model to obtain a comprehensive feature representation that integrates multimodal semantic information.

[0046] In this embodiment, within the same continuous time window (e.g., 2 seconds) described in step S210, the system simultaneously acquires the target object's voice text, log text associated with the target object, and the target object's action features. The voice text is collected via a microphone array installed on the top of the financial advisory room, and an automatic speech recognition model converts the speech into text in real time. For example, the converted text might contain phrases like "Tell me the verification code, and I'll enter it for you." The log text is obtained from the bank's intermediary business system via a system interface, showing log information related to the target object's (financial advisor's) current business processing status. For example, the log text might read "Current business node: Obtaining customer SMS verification code" or "Client self-service login in progress." This log reflects the compliance boundaries and current stage of the business operation. The action features are extracted from the target object's hand movement features based on a 3D skeletal sequence acquired by a depth camera. Specifically, a spatiotemporal graph convolutional network can be used to encode a high-dimensional action feature vector from the skeletal key points. This feature vector characterizes the hand's movement trajectory, speed, dwell time, and relative positional changes with the phone screen.

[0047] The system takes the aforementioned speech text, log text, and action features as input and feeds them into a pre-trained multimodal large model. This large model contains multiple layers of self-attention mechanisms, which calculate the relevance score between each element in the input sequence (including words in the speech, fields in the log, action feature vectors, etc.) and all other elements. Through this cross-modal attention calculation, the scattered modal information is fused into a unified, context-rich comprehensive feature representation. This comprehensive feature representation is no longer merely an encoding of visual actions, but a high-order representation that integrates dialogue semantics, business stage, and action patterns. For example, when the speech contains "verification code," the log displays "get verification code," and the action feature is "finger continuously touching the screen," the large model's self-attention mechanism will closely correlate these three elements, causing the output comprehensive feature representation to strongly point to the abnormal behavior pattern of "entering the verification code on behalf of the customer." As another example, if the speech is "please click here," the log is "product display," and the action is "finger pointing to the top of the screen," the comprehensive feature representation will tend towards normal guidance behavior.

[0048] Through step S220, the system successfully aligned the originally heterogeneous speech, log and action data into a unified semantic space, providing a high-quality fusion feature foundation for the subsequent output of abnormal intent probabilities.

[0049] In step S230, the comprehensive feature representation is mapped by linear transformation and normalized exponential function to output the probability of abnormal intent.

[0050] In this embodiment, the comprehensive feature representation output in step S220 is a high-dimensional vector that integrates multimodal information such as hand physical interaction, speech semantics, and business logs. To transform this representation into a deducible probability value of anomalies, the system further performs a linear transformation and normalized exponential function mapping. The system pre-sets a linear transformation layer (i.e., a fully connected layer) with the same input dimension as the comprehensive feature representation and an output dimension of 2, corresponding to the two categories of "normal behavior" and "abnormal behavior," respectively. After inputting the comprehensive feature representation into this linear layer, two raw logical values ​​are obtained. Subsequently, these two logical values ​​are fed into a softmax function, which converts the logical values ​​into a probability distribution, where the sum of the probabilities of the two categories is 1, and each probability value is between 0 and 1. The system takes the probability value corresponding to the "abnormal behavior" category as the abnormal intent probability output.

[0051] Through this step, the system quantifies the fused high-level semantic features into an intuitive anomaly probability value, which facilitates joint decision-making with the spatiotemporal overlap index of the physical space and provides a clear semantic basis for the final alarm.

[0052] In step S240, the spatiotemporal overlap index and the probability of abnormal intent are weighted and fused to generate an anomaly score.

[0053] In this embodiment, step S210 calculates the spatiotemporal overlap index characterizing the intensity of physical interaction, and step S230 outputs the probability of abnormal intent characterizing the likelihood of semantic anomalies. Since the two indicators have different dimensions and meanings, they need to be fused into a unified anomaly score for final judgment. Through this dynamic weighted fusion, the system can adaptively balance the importance of physical and semantic evidence based on the actual physical distance, avoiding misjudgment based on a single indicator. The resulting anomaly score is a unified value that integrates the intensity of three-dimensional spatial interaction and multimodal semantic understanding, providing an accurate basis for triggering the alarm in step S250.

[0054] In step S250, if the abnormal score is higher than a preset threshold, an abnormal behavior alarm is triggered.

[0055] In this embodiment, the system pre-sets a high-confidence threshold. If the anomaly score calculated in step S240 is greater than or equal to this threshold, the current behavior is determined to constitute an abnormal operation risk, and the system immediately triggers a tiered alarm. The alarm information includes at least the anomaly score, the alarm level, and a natural language explanation generated by the multimodal large model during the inference process, such as "the voice contains content requesting a verification code, the log shows that the person is in the verification code acquisition stage, and the hand is continuously touching the mobile phone screen area." This alarm information can be pushed to the bank's monitoring center through a visual interface, and video clips and audio transcripts within the time period of the anomaly can be extracted for manual review. If the anomaly score is lower than the threshold, the system does not trigger an alarm and continues monitoring the next time window.

[0056] Through this step, the system achieves real-time, explainable, and automated alerts for abnormal behavior, effectively assisting bank compliance managers in promptly identifying and intervening in potential violations.

[0057] The abnormal behavior monitoring method provided in this application will be described in detail below through specific embodiments.

[0058] In S210, in response to the target object's hand entering the target area, the spatiotemporal overlap index between the target object's hand and the target screen is determined. S210 includes S211 to S215.

[0059] In S211, in response to the target object's hand entering the target area, video images within the target area within a continuous time window are acquired.

[0060] In this embodiment, when the financial manager's hand enters a three-dimensional cube target area with a side length of 50 centimeters centered on the customer's mobile phone screen, the system initiates continuous time window (e.g., 2 seconds) acquisition. The depth camera mounted on the top synchronously acquires depth and color images of this area at a frequency of 30 frames per second, ensuring the acquisition of three-dimensional spatial information of the hand and the mobile phone screen.

[0061] In S212, based on the video image, the Euclidean distance between the target object's hand and the target screen is calculated, and the three-dimensional volume intersection-union ratio between the bounding sphere of the target object's hand and the bounding box of the target screen is calculated.

[0062] The system performs 3D pose estimation on each frame of depth image, extracting multiple key points of the financial manager's hand, including the left and right wrists, palms, and fingertips, and calculating the centroid of all key points as the hand center. Simultaneously, it acquires the center and four corner points of the 3D bounding box of the client's mobile phone screen. The system calculates the Euclidean distance between the hand's centroid and the screen center. Furthermore, the system constructs the bounding sphere of the hand (centered on the hand center, with the distance from the farthest key point on the hand to the sphere center as the radius) and the bounding box of the mobile phone screen (determined by the length, width, height, and yaw angle of the bounding box), calculating the ratio of their intersection volume to their union volume in 3D space, obtaining the 3D volume intersection-union ratio. This ratio ranges from 0 to 1; a higher value indicates a greater degree of overlap between the hand and the screen.

[0063] In S213, the instantaneous correlation strength factor is constructed based on the Euclidean distance and the intersection-union ratio of the three-dimensional volume.

[0064] The instantaneous association strength factor is obtained by weighted summing the Euclidean distance after undergoing an exponential decay transformation and the intersection-union ratio of the 3D volume. This instantaneous association strength factor comprehensively reflects the spatial proximity and overlap between the hand and the screen at the current moment; a higher value indicates a stronger instantaneous physical interaction. The formula is as follows:

[0065]

[0066] in, Indicates the instantaneous correlation strength factor. To adjust the weights, This is the distance attenuation coefficient. Represents the intersection-union ratio of three-dimensional volumes. This represents the Euclidean distance.

[0067] In S214, the directional weight is determined based on the approach direction of the target object's hand relative to the target screen.

[0068] When the target object's hand moves towards the target screen or remains stationary, the directional weight value approaches 1; when the target object's hand moves away from the target screen at a speed exceeding a preset threshold, the directional weight value approaches 0. This weight is used to filter out the contribution of non-continuous contact actions such as hand withdrawal to the spatiotemporal overlap index.

[0069] In S215, the spatiotemporal overlap index is obtained by integrating the product of the instantaneous correlation strength factor and the directional weight within a continuous time window.

[0070] Defined in time window Continuous spatiotemporal overlap index :

[0071] in, The relative speed at which the hand approaches the phone. The step rectified function is used when the hand approaches the phone (negative velocity) or remains stationary (zero velocity). When the hand is quickly withdrawn (at a speed greater than a threshold) .

[0072] In S220, the speech text of the target object, the log text associated with the target object, and the action features of the target object within a continuous time window are input into the multimodal large model. Cross-modal interactive reasoning is performed through the self-attention mechanism of the multimodal large model to obtain a comprehensive feature representation that integrates multimodal semantic information. S220 includes S221~S224.

[0073] In S221, the visual behavior token corresponding to the action feature, the voice text token corresponding to the voice text, and the log text token corresponding to the log text are extracted. The prompt token is then concatenated with the visual behavior token, the voice text token, and the log text token in sequence to form the input sequence.

[0074] In this embodiment, the system maps the action feature vector obtained in step S220 to the token embedding space of the multimodal large model through linear projection to obtain a visual behavior token sequence, where each token corresponds to a sub-segment or the entire encoding of the action feature. Simultaneously, the large model's word segmenter segments the speech text and log text into several speech text tokens and log text tokens, respectively. The prompt token is a tokenized representation of a preset instruction text, which defines the constraints on the probability of abnormal intent output by the multimodal large model. For example, the instruction text could be set as "Please determine whether the behavior is an abnormal operation (abnormal operation refers to: a financial manager operating a customer's mobile phone without authorization, such as entering a password, transferring funds, or requesting a verification code) based on the input visual behavior token (describing the interaction between the hand and the screen), speech text token (on-site dialogue content), and log text token (current business node)." The above instruction text is converted into one or more prompt tokens by the multimodal large model's word segmenter and concatenated to the beginning of the input sequence.

[0075] The input sequence is represented as:

[0076]

[0077] in, A prompt token that represents the command text. Represents a log text token. A voice text token representing the target object. Represents a visual behavior token.

[0078] In S222, the input sequence is fed into the multimodal large model, and the attention score between all token pairs in the input sequence is calculated through the multi-layer self-attention mechanism of the multimodal large model.

[0079] The input sequence formed in step S221 is fed into the pre-trained multimodal large model. The model's first-layer self-attention mechanism generates a query vector (Q), a key vector (K), and a value vector (V) for each token in the sequence. The core multi-layer self-attention computation of the multimodal large model can be represented as:

[0080]

[0081] The process calculates the association strength between all token pairs in the sequence, including cue tokens and visual tokens, visual tokens and voice tokens, voice tokens and log tokens, etc.

[0082] In S223, the original representations of each token are weighted and fused using attention scores to update the representation of each token.

[0083] Based on the attention score matrix obtained in step S222, the value vector of each token is weighted and summed to obtain the updated representation of the token. For example, the updated representation of a visual behavior token will integrate information from all related voice tokens, log tokens, and cue tokens. In this way, the semantics of each token are fused with context from other modalities, achieving preliminary alignment and fusion of cross-modal information.

[0084] In S224, after stacking multiple layers of self-attention mechanisms, a comprehensive feature representation that integrates multimodal semantic information is obtained.

[0085] Multimodal large models typically contain multiple layers (e.g., 12 or 24 layers) of self-attention modules. Each layer repeats steps S222 and S223, using the output of the previous layer as the input of the next. As the number of layers increases, the interaction between tokens becomes more comprehensive, and high-level semantics gradually become more abstract. After stacking all layers, the model outputs an updated input sequence. The system extracts the final representation of the token at a preset position (usually the first position in the input sequence, i.e., the position corresponding to the cue token) from this sequence as a comprehensive feature representation that integrates multimodal semantic information. This representation vector contains the collaborative semantics of hand gestures, voice dialogue, and business logs throughout the entire time window, providing sufficient evidence for the subsequent output of abnormal intent probabilities.

[0086] In S230, the comprehensive feature representation is mapped by linear transformation and normalized exponential function to output the probability of abnormal intent, S231~S232.

[0087] In S231, the comprehensive feature representation is mapped to the vector space corresponding to the abnormal operation to obtain the original logic value.

[0088] The system pre-sets a linear transformation layer, i.e., a fully connected layer, whose input dimension is the same as the dimension of the comprehensive feature representation output in step S224, and whose output dimension is 2, corresponding to the two categories of "normal behavior" and "abnormal behavior" respectively. The weight matrix W and bias b of this linear layer are trainable parameters, which are fixed after pre-training. The comprehensive feature representation h is input into this linear layer to calculate the two-dimensional original logical value vector z=[z 正常 ,z 异常 In this application, abnormal operation refers to "the financial manager operating the client's mobile phone (such as entering passwords, transferring funds, or requesting verification codes) without the client's knowledge or authorization".

[0089] For example, suppose that within a certain time window, the comprehensive feature representation h is considered by a pre-trained model to strongly indicate anomalous behavior, then after linear transformation z 异常 It could be 2.5, z 正常 It could be -1.2; conversely, if the overall characteristics indicate normal guiding behavior, then z... 正常 Higher, z 异常 Lower.

[0090] In S232, the original logical value is calculated using a normalized exponential function to obtain the probability that the target object's hand gesture belongs to an abnormal intention.

[0091] The system will use the two-dimensional logic value vector z=[z] obtained in step S231. 正常 ,z 异常 Input the normalized exponential function, the formula for which is calculated is:

[0092]

[0093] Among them, P 异常 This represents the probability of abnormal intent, with a value between 0 and 1. Meanwhile, P... 正常 =1− P 异常 The system ultimately outputs P. 异常 This is the result of step S230.

[0094] In S240, the spatiotemporal overlap index and the probability of abnormal intent are weighted and fused to generate an anomaly score. S240 includes S241 to S242.

[0095] In S241, the fusion weight coefficients are determined based on the Euclidean distance between the target object's hand and the target screen.

[0096] The system dynamically calculates the fusion weight coefficient γ based on the Euclidean distance between the hand's centroid and the screen center within the current time window (or the current keyframe). This coefficient is negatively correlated with the Euclidean distance: when the hand is very close to the screen, γ takes a larger value, such as 0.8 to 0.9, indicating that the spatiotemporal overlap index in physical space is more reliable; when the hand is far from the screen, γ takes a smaller value, such as 0.2 to 0.3, in which case the probability of abnormal intent in semantic reasoning dominates.

[0097] In S242, the spatiotemporal overlap index is normalized and multiplied by the probability of abnormal intent by the fusion weight coefficient and its complement, and then the products are added together to obtain the anomaly score.

[0098] The spatiotemporal overlap index calculated in step S215 Normalize to the [0,1] interval. Let P be the probability of the abnormal intent output in step S232. 异常 The fusion weight coefficient is γ, and its complement is 1−γ. The anomaly score is calculated as follows:

[0099] Score=γ⋅ +(1−γ)⋅ P 异常

[0100] The abnormal behavior monitoring method provided in this application eliminates false alarms caused by projection parallax in a two-dimensional view by calculating the spatiotemporal overlap index between the hand and the screen in three-dimensional space. Simultaneously, it introduces a multimodal large model to fuse voice text, business logs, and action features, utilizing a self-attention mechanism for cross-modal reasoning. This enables it to understand behavioral semantics and distinguish between compliant guidance and unauthorized customer operations. It dynamically weights and fuses physical indicators and semantic probabilities based on real-time distance, balancing evidence from both sides to improve robustness. The large model simultaneously generates natural language explanations, making alarms interpretable and facilitating compliance auditing. The design of time window integration and directional weights accurately characterizes the complete dynamic process of hand approach, contact, and departure, effectively capturing instantaneous abnormal actions. In summary, this application achieves real-time abnormal behavior monitoring with low false alarm rate, high accuracy, and strong interpretability.

[0101] Based on the above-described abnormal behavior monitoring method, embodiments of this application also provide an abnormal behavior monitoring device. The following will be combined with... Figure 3 The device is described in detail.

[0102] Figure 3 A schematic block diagram of an abnormal behavior monitoring device according to an embodiment of this application is shown.

[0103] like Figure 3 As shown, the abnormal behavior monitoring device 300 in this embodiment includes an overlap index calculation module 310, a comprehensive feature extraction module 320, an abnormal probability calculation module 330, an abnormal scoring module 340, and an alarm triggering module 350.

[0104] The overlap index calculation module 310 is used to determine the spatiotemporal overlap index between the target object's hand and the target screen in response to the target object's hand entering the target area. The target area is a three-dimensional space of a preset size where the target screen is located, and the spatiotemporal overlap index is a quantitative indicator characterizing the intensity of physical interaction between the target object's hand and the target screen within a continuous time window. In one embodiment, the overlap index calculation module 310 can be used to execute step S210 described above, which will not be repeated here.

[0105] The comprehensive feature extraction module 320 is used to input the speech text of the target object within a continuous time window, the log text associated with the target object, and the action features of the target object into the multimodal large model. Cross-modal interactive reasoning is performed through the self-attention mechanism of the multimodal large model to obtain a comprehensive feature representation that integrates multimodal semantic information. The log text is the log text of the business system related to the target object. In one embodiment, the comprehensive feature extraction module 320 can be used to execute step S220 described above, which will not be repeated here.

[0106] The anomaly probability calculation module 330 is used to map the comprehensive feature representation through linear transformation and normalized exponential function to output the probability of abnormal intent. In one embodiment, the anomaly probability calculation module 330 can be used to perform step S230 described above, which will not be repeated here.

[0107] The anomaly scoring module 340 is used to weightedly fuse the spatiotemporal overlap index and the probability of anomalous intent to generate an anomaly score. In one embodiment, the anomaly scoring module 340 can be used to perform step S240 described above, which will not be repeated here.

[0108] The alarm triggering module 350 is used to trigger an abnormal behavior alarm when the abnormal score is higher than a preset threshold. In one embodiment, the alarm triggering module 350 can be used to execute step S250 described above, which will not be repeated here.

[0109] According to embodiments of this application, any multiple modules among the overlap index calculation module 310, comprehensive feature extraction module 320, anomaly probability calculation module 330, anomaly scoring module 340, and alarm triggering module 350 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the overlap index calculation module 310, comprehensive feature extraction module 320, anomaly probability calculation module 330, anomaly scoring module 340, and alarm triggering module 350 can be at least partially implemented as hardware circuits, such as field-programmable gate arrays, programmable logic arrays, systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits, or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, and firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the overlap index calculation module 310, comprehensive feature extraction module 320, anomaly probability calculation module 330, anomaly scoring module 340, and alarm triggering module 350 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0110] Figure 4 A block diagram of an electronic device suitable for implementing an abnormal behavior monitoring method according to an embodiment of this application is shown schematically.

[0111] like Figure 4 As shown, an electronic device 400 according to an embodiment of this application includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory 402 or a program loaded from a storage portion 408 into a random access memory 403. The processor 401 may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a dedicated microprocessor. The processor 401 may also include onboard memory for caching purposes. The processor 401 may include a single processing unit or multiple processing units for executing different steps of the method flow according to an embodiment of this application.

[0112] Random access memory 403 stores various programs and data required for the operation of electronic device 400. Processor 401, read-only memory 402, and random access memory 403 are interconnected via bus 404. Processor 401 executes various steps of the method flow according to embodiments of this application by executing programs stored in read-only memory 402 and / or random access memory 403. It should be noted that the programs may also be stored in one or more memories other than read-only memory 402 and random access memory 403. Processor 401 may also execute various steps of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0113] According to embodiments of this application, the electronic device 400 may further include an input / output interface 405, which is also connected to a bus 404. The electronic device 400 may also include one or more of the following components connected to the input / output interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube, liquid crystal display, etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card, such as a local area network card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 410 as needed so that computer programs read from it can be installed into the storage section 408 as needed.

[0114] Embodiments of this application also provide a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0115] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In embodiments of this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include the read-only memory 402 described above, and / or random access memory 403, and / or one or more memories other than read-only memory 402 and random access memory 403.

[0116] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.

[0117] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via communication section 409, and / or installed from removable medium 411. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0118] In embodiments of this application, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by processor 401, it performs the functions defined in the system of this application embodiment. According to embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0119] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0121] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. A method for monitoring abnormal behavior, characterized in that, The method includes: In response to the hand of a target object entering the target area, the spatiotemporal overlap index between the hand of the target object and the target screen is determined. The target area is a three-dimensional space of a preset size where the target screen is located. The spatiotemporal overlap index is a quantitative indicator that characterizes the intensity of physical interaction between the hand of the target object and the target screen within a continuous time window. The speech text of the target object within the continuous time window, the log text associated with the target object, and the action features of the target object are input into the multimodal big model. Cross-modal interactive reasoning is performed through the self-attention mechanism of the multimodal big model to obtain a comprehensive feature representation that integrates multimodal semantic information. The log text is the log text of the business system related to the target object. The comprehensive feature representation is mapped by linear transformation and normalized exponential function to output the probability of abnormal intent; The spatiotemporal overlap index and the probability of abnormal intent are weighted and fused to generate an anomaly score; If the abnormal score is higher than a preset threshold, an abnormal behavior alarm is triggered.

2. The method according to claim 1, characterized in that, The step of determining the spatiotemporal overlap index between the target object's hand and the target screen in response to the target object's hand entering the target area includes: In response to the target object's hand entering the target area, video images of the target area within the continuous time window are acquired; Based on the video image, calculate the Euclidean distance between the hand of the target object and the target screen, and calculate the three-dimensional volume intersection-union ratio between the bounding sphere of the hand of the target object and the bounding box of the target screen; Construct an instantaneous correlation strength factor based on the Euclidean distance and the three-dimensional volume intersection-union ratio; The directional weight is determined based on the approach direction of the target object's hand relative to the target screen; The spatiotemporal overlap index is obtained by integrating the product of the instantaneous correlation strength factor and the directional weight within the continuous time window.

3. The method according to claim 2, characterized in that, The step of constructing the instantaneous correlation strength factor based on the Euclidean distance and the three-dimensional volume intersection-union ratio includes: The instantaneous correlation strength factor is obtained by performing an exponential decay transformation on the Euclidean distance and then weighting it with the intersection-union ratio of the three-dimensional volume.

4. The method according to claim 2, characterized in that, When the target object's hand moves toward the target screen or remains stationary, the value of the directional weight approaches 1; when the target object's hand moves away from the target screen at a speed exceeding a preset threshold, the value of the directional weight approaches 0.

5. The method according to claim 1, characterized in that, The step of inputting the speech text of the target object within the continuous time window, the log text associated with the target object, and the action features of the target object into a multimodal large model, and performing cross-modal interactive reasoning through the self-attention mechanism of the multimodal large model to obtain a comprehensive feature representation that integrates multimodal semantic information includes: Extract the visual behavior token corresponding to the action feature, the voice text token corresponding to the voice text, and the log text token corresponding to the log text. Concatenate the prompt token with the visual behavior token, the voice text token, and the log text token in sequence to form an input sequence. The prompt token is a tokenized representation of a preset instruction text. The instruction text is used to define the constraint conditions for the probability of abnormal intent output by the multimodal large model. The input sequence is fed into the multimodal large model, and the attention score between all token pairs in the input sequence is calculated through the multi-layer self-attention mechanism of the multimodal large model. The original representations of each token are weighted and fused using the attention scores to update the representation of each token; By stacking multiple layers of self-attention mechanisms, a comprehensive feature representation that integrates multimodal semantic information is obtained.

6. The method according to claim 1, characterized in that, The step of mapping the comprehensive feature representation through linear transformation and normalized exponential function to output the probability of abnormal intent includes: The comprehensive feature representation is mapped to the vector space corresponding to the abnormal operation to obtain the original logical value; The original logical value is calculated using a normalized exponential function to obtain the probability that the target object's hand movements belong to an abnormal operation or an abnormal intent.

7. The method according to claim 1, characterized in that, The step of weightedly fusing the spatiotemporal overlap index and the abnormal intent probability to generate an anomaly score includes: The fusion weighting coefficients are determined based on the Euclidean distance between the target object's hand and the target screen; The spatiotemporal overlap index is normalized and multiplied by the probability of abnormal intent by the fusion weight coefficient and its complement, and then the products are added together to obtain the abnormal score. The fusion weight coefficient is negatively correlated with the Euclidean distance.

8. An abnormal behavior monitoring device, characterized in that, The device includes: The overlap index calculation module is used to determine the spatiotemporal overlap index between the hand of the target object and the target screen in response to the hand of the target object entering the target area. The target area is a three-dimensional space of a preset size where the target screen is located. The spatiotemporal overlap index is a quantitative index that characterizes the intensity of physical interaction between the hand of the target object and the target screen within a continuous time window. The comprehensive feature extraction module is used to input the speech text of the target object, the log text associated with the target object, and the action features of the target object into the multimodal big model within the continuous time window. Cross-modal interactive reasoning is performed through the self-attention mechanism of the multimodal big model to obtain a comprehensive feature representation that integrates multimodal semantic information. The log text is the log text of the business system related to the target object. The anomaly probability calculation module is used to map the comprehensive feature representation through linear transformation and normalized exponential function to output the probability of abnormal intent; Anomaly scoring module is used to weight and fuse the spatiotemporal overlap index and the probability of abnormal intent to generate anomaly score; The alarm triggering module is used to trigger an abnormal behavior alarm when the abnormal score is higher than a preset threshold.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.