Method and system for behavior semantic understanding based on spatio-temporal causal reasoning and confidence fusion

By fusing spatiotemporal causal reasoning with confidence levels, this method solves the problems of evidence fragmentation and decision-making rigidity in behavioral semantic understanding in existing technologies. It achieves highly robust and interpretable behavioral understanding, which is applicable to accurate risk assessment and rapid auditing in various scenarios such as retail and libraries.

CN122493364APending Publication Date: 2026-07-31NINGBO YUNJIN MICRO INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO YUNJIN MICRO INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing behavioral semantic understanding technologies suffer from fragmented evidence utilization, rigid decision-making logic, lack of causality in reasoning, and untraceable evidence chains in complex real-world scenarios, making it difficult to achieve high-precision, high-robustness, and interpretable behavioral understanding.

Method used

By employing a method based on spatiotemporal causal reasoning and confidence fusion, and combining multi-source heterogeneous visual data fusion and confidence calculation with a causal reasoning engine and behavioral language model, a multi-level evidence graph network is constructed to achieve structured representation and tracing of target object behavior.

Benefits of technology

It improves the robustness of status confirmation, reduces the false alarm rate, enhances the interpretability of the system and the structuring of the evidence chain, and is highly adaptable, enabling accurate risk assessment and rapid auditing in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493364A_ABST
    Figure CN122493364A_ABST
Patent Text Reader

Abstract

This invention provides a behavioral semantic understanding method and system based on spatiotemporal causal reasoning and confidence fusion. It simultaneously extracts three types of evidence from multiple video streams: interaction, trajectory, and credentials of the target object, and quantifies their confidence levels. A joint decision function with dynamically adjusted weights is used to fuse multi-source confidence levels, outputting highly robust state judgments and interpretable evidence. A causal reasoning engine and behavioral language model are used to perform causal violation detection, context anomaly detection, and intent inference on behavioral primitive sequences. A multi-level evidence graph network is constructed to achieve structured representation and tracing of interactive behavioral events. This addresses the problems of existing behavioral recognition technologies, such as the use of single evidence, rigid decision logic, and lack of causal understanding. It achieves automated, interpretable, and high-precision semantic understanding of the target object's "intent-behavior-result" link, making it particularly suitable for intelligent monitoring scenarios such as retail loss prevention and process compliance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method and system for behavioral semantic understanding based on spatiotemporal causal reasoning and confidence fusion. Background Technology

[0002] Behavioral semantic understanding is a core technology in the fields of computer vision, artificial intelligence, and intelligent monitoring. It aims to achieve automated parsing, status confirmation, and risk assessment of the entire chain of intent, behavior, and result of a target object through multi-source visual perception and intelligent analysis. It has become a key supporting technology for scenarios such as retail loss prevention, warehouse management, public self-service, park process compliance, library self-service borrowing and returning, and security duty. In the above practical applications, the system needs to accurately determine whether the target object has completed the preset compliant operation, identify abnormal behavior patterns, and output traceable judgment criteria to meet the core needs of scenario-based risk prevention and control, automated process supervision, and post-event auditing.

[0003] Current behavioral semantic understanding technologies still suffer from four core defects in practical applications: fragmented evidence utilization, rigid decision-making logic, lack of causal reasoning, and untraceable evidence chains. These shortcomings make it difficult to meet the high-precision, high-robustness, and interpretability requirements of complex real-world scenarios. Specific problems are as follows: ① Existing technologies mostly focus on a single visual analysis dimension, such as using only target tracking for trajectory analysis, only posture estimation for hand movement recognition, and only OCR for credential recognition. These various technical modules are isolated from each other. A unified framework has not been built to jointly model and quantify the uncertainty of the three core semantic information types: item interaction, movement trajectory, and credential verification. This makes it impossible to fully depict the complete behavioral chain of "picking up an item - moving along a path - verifying credential verification," resulting in a one-sided understanding of behavior and insufficient basis for decision-making.

[0004] ② Existing solutions generally use rigid "if-then" rules to determine the operation status, such as determining the operation is complete only when a payment voucher is detected. This model is prone to misjudgment and omission when faced with interference from real-world scenarios such as viewpoint obstruction, lighting changes, blurred vouchers, and non-standard operating procedures; moreover, it cannot output quantitative inference results with confidence when the evidence is insufficient, resulting in extremely poor decision stability and fault tolerance.

[0005] ③ The risk warning of the existing system is based only on simple surface rules such as "taking without payment", which cannot uncover the true intention behind the behavior and makes it difficult to distinguish between normal shopping, in-store browsing, employee stock management and suspicious theft. At the same time, it does not embed a lightweight causal reasoning engine, so it cannot identify hidden abnormal patterns that violate spatiotemporal causal logic, such as "staying for a long time after taking" and "touching items outside of business hours". The risk identification accuracy is low and the false alarm rate is high.

[0006] ④ The existing system simply piles up and associates raw data such as video clips, detection boxes, and recognized text, which is a "data dumping" evidence storage method. It does not perform logical abstraction and multi-level summary processing on the raw data, and does not form a structured and hierarchical evidence chain. When manually verifying, it is impossible to quickly grasp the whole picture of the event, the traceability of evidence is difficult, the audit efficiency is extremely low, and it is difficult to meet the actual needs of compliance supervision and post-event traceability. Summary of the Invention

[0007] In view of this, the present invention proposes a behavioral semantic understanding method and system based on spatiotemporal causal reasoning and confidence fusion, which can integrate multi-source heterogeneous visual data and perform spatiotemporal causal reasoning to achieve highly robust, quantifiable, and interpretable automated understanding and confirmation of the behavioral semantics of target objects.

[0008] The technical solution of this invention is implemented as follows: A behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion includes the following steps: Step S1: Simultaneously acquire video streams of the target scene from at least two monitoring viewpoints with spatiotemporal correlation, and extract interaction evidence, trajectory evidence, and credential evidence of the target object from the video streams; Step S2: Calculate the confidence level based on interaction evidence, trajectory evidence, and credential evidence to obtain the interaction confidence level, trajectory compliance confidence level, and credential validity confidence level. Step S3: Dynamically weight and fuse the interaction confidence, trajectory compliance confidence, and credential validity confidence through the constructed joint decision-making model. The dynamic weights are adaptively adjusted according to the prior knowledge of the target scenario and the real-time context, and the overall confidence is output and the state label is determined. Step S4: Parse the video stream into discrete behavioral primitive sequences, and use a causal inference engine to perform causal violation detection and context anomaly detection on the behavioral primitive sequences; Step S5: Infer the target object's intent based on the sequence of behavioral primitives using a behavioral language model. The behavioral language model is used to assist decision-making when the confidence level of the credentials is insufficient. Step S6: Construct a multi-layered evidence graph network containing a summary layer, a logic layer, and a data layer to structurally represent and trace the interactive behavior of the target object.

[0009] Preferably, the specific steps of step S1 are as follows: Simultaneously acquire video streams from at least two monitoring viewpoints that are spatially and temporally related; Extract video frames from the video stream that show the entire process of the target object's hand touching the object as evidence of interaction; Extract video frames from the video stream that contain the entire process of the target object moving from the interaction point to the preset operation confirmation area as trajectory evidence; In the operation confirmation area, the text is recognized by an OCR model to obtain credential evidence.

[0010] Preferably, the calculation steps for the interaction confidence level are as follows: Calculate the IoU sequence between the target object's hand bounding box and the item bounding box over N consecutive frames, and calculate the sustained contact score using exponential decay weighting. The calculation formula is as follows:

[0011] in To maintain continuous exposure to scores, As the attenuation factor, Let t be the IoU sequence of the t-th frame; The statistical IoU exceeds the threshold Number of consecutive frames The duration fraction is mapped using an S-shaped function, and the calculation formula is as follows:

[0012] in Let k be the duration fraction, and k be the slope parameter of the sigmoid function. The reference frame number for determining valid interactions; Based on the item's state before and after the interaction, if a state change occurs, the state change score is calculated. Otherwise, it is 0; The interaction confidence score is obtained by fusing the sustained contact score, duration score, and state change score. The calculation formula is as follows: ,in For interactive confidence, It is an adjustable weight, and .

[0013] Preferably, the calculation steps for the trajectory compliance confidence level are as follows: Based on the scene topology map of the target object, the theoretically optimal path from the interaction point to the operation confirmation area is generated using the A* algorithm. ; Using a dynamic time warping algorithm, the actual movement path point sequence is transformed. with the optimal path Align and calculate the minimum cumulative distance. ; Path similarity is calculated based on the minimum cumulative distance, and the formula is as follows:

[0014] in For path similarity, For scale parameters, The length of the optimal path; When it is determined that the target object has entered the geofence of the operation confirmation area within the time window, the arrival score is output. ,otherwise ; The trajectory compliance confidence score is obtained by fusing path similarity and arrival score, and the calculation formula is as follows:

[0015] in For the confidence level of trajectory compliance, These are the path similarity weight and the arrival score weight, respectively.

[0016] Preferably, the calculation steps for the confidence level of the certificate validity are as follows: An OCR model is used to obtain the recognition confidence of the text and each character, and the overall text confidence is calculated using the following formula:

[0017] in For the overall credibility of the text, Let i be the recognition confidence score of the i-th character output by the OCR model. To calculate the average value, To calculate the standard deviation; Using a pre-trained sentence embedding model, the cosine similarity between the text recognized by the OCR model and the expected voucher template is calculated. ; Based on the text recognition of the OCR model, the voucher time and voucher amount are verified to obtain a logical reward score. ; The confidence level of credential validity is calculated based on overall text credibility, cosine similarity, and logical reward score. The calculation formula is as follows: ,in Assuming confidence level for the validity of the document, These are text credibility weight and semantic similarity weight, respectively.

[0018] Preferably, the specific steps of step S3 are as follows: A joint decision-making model with dynamic weight adjustment is constructed. The inputs to the joint decision-making model are the interaction confidence, trajectory compliance confidence, and credential validity confidence. The output is the overall confidence of the target object completing the preset operation. The joint decision function of the joint decision-making model adopts the weighted geometric mean form, and the calculation formula is as follows:

[0019] in For the overall confidence level, Dynamic weights are dynamically adjusted based on prior knowledge of the target scenario and the real-time context. The overall confidence level is compared with a preset threshold range, and a status label representing the completion status of the operation is output.

[0020] Preferably, step S4 includes the following specific steps: The continuous video stream is parsed to obtain semantic abstract representations of discrete behaviors of several target objects in the target scene, and output as behavior primitives. A sequence of behavior primitives is formed based on the behavior primitives. The behavioral primitive sequence is fed into the causal reasoning engine to detect patterns in the behavioral primitive sequence that violate common sense causal relationships. At the same time, the abnormality of the behavioral primitive is judged by combining contextual information, and the detection results are obtained.

[0021] Preferably, in step S5, when the confidence level of the credential validity is lower than a preset threshold, the behavioral language model is triggered to infer intent, and the inference process includes: The sequence of behavioral primitives is converted into a vector sequence through an embedding layer, and then concatenated with the context feature vector or fused with attention to obtain the embedding sequence. The embedded sequence is encoded using a lightweight Transformer encoder or a bidirectional LSTM to obtain a context vector representing the overall semantics of the behavioral sequence. The context vector is input into a fully connected layer classifier, which outputs a probability distribution for a predefined set of intents, including purchase intent, browsing intent, suspicious intent, and work intent. The maximum probability value is determined from the probability distribution. When the maximum probability value is greater than the preset trigger threshold, the intent category corresponding to the maximum probability value is determined as the target object intent.

[0022] Preferably, step S6 includes the following specific steps: In the summary layer, the overall confidence level and status labels are combined to describe the complete interaction process of the target object using natural language. In the logic layer, several behavioral primitive nodes are created according to the time sequence of the behavioral primitive sequence and marked with corresponding timestamps. For each behavioral primitive, the interaction confidence, trajectory compliance confidence, and credential validity confidence are associated. Temporal sequence and causal trigger relationship edges are established between behavioral primitive nodes. Causal violation results and context abnormal results are marked on the corresponding behavioral primitive nodes. At the same time, the target object intent is attached to the entire behavioral primitive sequence. In the data layer, video evidence nodes and credential evidence nodes are created. The video evidence nodes include spatiotemporal indexes of multiple video streams, key video segments of behavior, key frames of interaction, video segments of trajectory, and credential presentation frame nodes. The credential evidence nodes include credential OCR text results, semantic similarity results, and logical verification record nodes. The nodes within the logic layer and between the logic layer and the data layer are connected by relational edges. The relational edges between the nodes within the logic layer are temporal or causal relationships, and the relational edges between the nodes in the logic layer and the data layer are proof relationships.

[0023] Systems applying behavioral semantic understanding methods based on spatiotemporal causal reasoning and confidence fusion include: At least two edge computing nodes are deployed in the target scene. Each edge computing node integrates an image sensor and a computing unit to perform video stream analysis in real time, extract interaction evidence and trajectory evidence of the target object and items, and generate a structured evidence data package with spatiotemporal stamps. The central server, deployed in the cloud, communicates with the edge computing nodes and is configured as follows: Receive and associate structured evidence data packets from different edge nodes, calculate interaction confidence based on interaction evidence, and calculate trajectory compliance confidence based on trajectory evidence; Run a dynamic fusion decision service, adjust weights based on real-time context, and fuse the interaction confidence and trajectory compliance confidence from edge computing nodes with the credential validity confidence from cloud-recognized data; Run the causal reasoning service, load the scenario-based causal rule library, and reason about the behavioral primitive sequences obtained by parsing the video stream reported from the edge computing node; Run the intent inference service. When the confidence level of the credentials is insufficient, load the behavioral language model to infer the intent from the sequence of behavioral primitives. Run the evidence graph management service to build and store a multi-layered evidence graph network containing a summary layer, a logic layer, and a data layer based on a graph database.

[0024] Compared with the prior art, the beneficial effects of the present invention are: The robustness of status confirmation is significantly improved: through the multi-source confidence fusion mechanism, a stable judgment can still be made using other evidence when a single source of evidence is unreliable.

[0025] The false alarm rate is significantly reduced, and the risk warning is more accurate: The spatiotemporal causal reasoning engine can effectively distinguish between "suspicious unpaid transactions" and "normal in-store browsing or employee stock management", reducing the false alarm rate in retail scenarios. At the same time, through intent inference, it can discover hidden suspicious patterns that traditional rules cannot cover.

[0026] The decision-making process is interpretable and the system has high credibility: the system outputs not only the results, but also a decomposed view that integrates confidence and a reasoning chain based on behavioral primitives, enabling security personnel to quickly understand the basis of the system's judgment, shortening the average event verification time, and greatly improving the efficiency of human-machine collaboration and the credibility of the system.

[0027] The evidence chain is highly structured and easy to trace: The multi-level evidence graph network abstracts the original data into logical events and provides a flexible access path from summary to details, which improves the efficiency of auditors in locating key events and extracting the complete evidence chain compared to the original video playback.

[0028] The framework is highly versatile and adaptable to various scenarios: The "evidence-fusion-reasoning" framework proposed in this invention can be decoupled from specific detection algorithms. By configuring different behavioral primitive libraries, causal rules and weight strategies, the same system core can be quickly adapted to various scenarios such as retail, library, warehouse, and office area, reducing the cost of secondary development. Existing behavioral semantic understanding technologies face a structural technical contradiction: if rigid rules are adopted (such as "detecting payment voucher = completion"), the system has an extremely high false alarm rate under occlusion, ambiguity, or non-standard processes (lacking fault tolerance); if a purely data-driven model (such as end-to-end deep learning) is used, the system has extremely poor interpretability, cannot provide the judgment criteria required for auditing, and the reliability of black-box models is uncontrollable in edge scenarios. This invention, through an architecture of "multi-source confidence quantification fusion + causal reasoning verification + intent inference supplementation + structured evidence tracing," simultaneously resolves the above contradictions within the same technical framework: it achieves flexible fault tolerance (reducing false alarms) through confidence fusion and intent inference, and achieves white-box interpretability (meeting auditing requirements) through a causal reasoning engine and a three-layer evidence graph network. Existing technologies do not provide technical inspiration for simultaneously achieving "flexible decision-making" and "interpretable tracing" within the same system. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only preferred embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart of the behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion of the present invention; Figure 2 This is a schematic diagram of the spatiotemporal causal reasoning process of the behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion of the present invention; Figure 3 This is a schematic diagram of the structure of the multi-level evidence graph network of the behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion of the present invention; Figure 4 This is a schematic diagram of the architecture of the behavioral semantic understanding system based on spatiotemporal causal reasoning and confidence fusion of the present invention; Figure 5This is a timeline diagram illustrating a specific application example of the behavioral semantic understanding method and system based on spatiotemporal causal reasoning and confidence fusion of the present invention in a retail scenario. Detailed Implementation

[0031] To better understand the technical content of this invention, a specific embodiment is provided below, and the invention will be further described in conjunction with the accompanying drawings.

[0032] In this embodiment, "behavioral language model" refers to a neural network model that semantically encodes behavioral primitive sequences based on a sequence encoder (including but not limited to a lightweight Transformer encoder or bidirectional LSTM) and outputs the probability distribution of intent categories; "dynamic weight adjustment" refers to the process by which the joint decision-making model adaptively updates the fusion weights based on prior knowledge of the target scenario (such as the baseline weight configuration of scenarios such as retail, warehousing, and library) and real-time context (such as time period, crowd density, and historical false alarm feedback); "behavioral primitive" refers to the semantic abstract representation of discrete behavioral actions of the target object in the target scenario, including but not limited to approaching, taking, moving, waiting, and showing.

[0033] See Figures 1 to 3 The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion provided by this invention includes the following steps: Step S1: Simultaneously acquire video streams of the target scene from at least two monitoring viewpoints with spatiotemporal correlation, and extract interaction evidence, trajectory evidence, and credential evidence of the target object from the video streams; Step S2: Calculate the confidence level based on interaction evidence, trajectory evidence, and credential evidence to obtain the interaction confidence level, trajectory compliance confidence level, and credential validity confidence level. Step S3: Dynamically weight and fuse the interaction confidence, trajectory compliance confidence, and credential validity confidence through the constructed joint decision-making model. The dynamic weights are adaptively adjusted according to the prior knowledge of the target scenario and the real-time context, and the overall confidence is output and the state label is determined. Step S4: Parse the video stream into discrete behavioral primitive sequences, and use a causal inference engine to perform causal violation detection and context anomaly detection on the behavioral primitive sequences; Step S5: Infer the target object's intent based on the sequence of behavioral primitives using a behavioral language model. The behavioral language model is used to assist decision-making when the confidence level of the credentials is insufficient. Step S6: Construct a multi-layered evidence graph network containing a summary layer, a logic layer, and a data layer to structurally represent and trace the interactive behavior of the target object.

[0034] The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion of the present invention can be applied to behavioral semantic recognition in various scenarios, such as retail loss prevention, warehouse management, and self-service. In the target scenario, multiple sets of monitoring viewpoints with spatiotemporal correlations are set up. The monitoring viewpoints can collect video streams containing the behavior of the target object in the target scenario. Through the video stream, interaction evidence between the target object and the item, trajectory data of the target object's own movement, and voucher data after the target object has completed a transaction can be extracted. The interaction data, trajectory data, and voucher data constitute multi-source heterogeneous visual data. After acquiring the multi-source heterogeneous visual data, the confidence of each type of evidence needs to be calculated. Then, the confidence of the multi-source evidence is fused to obtain the overall confidence. Then, based on the overall confidence, the state of the target object's entire interactive behavior can be evaluated and the state label can be determined. When a single source of evidence is unreliable, other evidence can still be used to make a stable judgment. Even when there is interference such as occlusion or blurring, a high state confirmation accuracy can still be maintained.

[0035] In addition, traditional behavioral semantic understanding methods may produce false alarms during execution. To reduce the false alarm rate, the acquired video stream is parsed into discrete behavioral primitive sequences. Then, a causal inference engine is used to perform causal violation detection and context anomaly detection on the behavioral primitive sequences. This can effectively distinguish between "suspicious unpaid transactions" and "normal in-store browsing or employee stock management". At the same time, when the confidence level of the target object's credential validity is lower than a preset threshold, it indicates that the target object's credential validity is weak and it is impossible to determine whether the payment was valid. In this case, the intent of the target object can be inferred through the set behavioral language model. If it is inferred that the target object is indeed only browsing or stocking, no alarm is triggered, further reducing the false alarm rate.

[0036] Finally, after completing the semantic understanding of the target object's behavior, a multi-level evidence graph network is constructed. The summary layer can provide a complete and concise description of the target object's behavior in a given instance, while the logic layer can display each behavioral primitive of the target object and show the causal or temporal relationship of the behavioral primitives. The data layer can provide corresponding evidence data to facilitate tracing. By abstracting the original data into logical events through the multi-level evidence graph network, and providing a flexible access path from summary to details, it can significantly improve the efficiency of auditors in locating key events and extracting a complete chain of evidence compared to the original video playback.

[0037] Preferably, the specific steps of step S1 are as follows: Simultaneously acquire video streams from at least two monitoring viewpoints that are spatially and temporally related; Extract video frames from the video stream that show the entire process of the target object's hand touching the object as evidence of interaction; Extract video frames from the video stream that contain the entire process of the target object moving from the interaction point to the preset operation confirmation area as trajectory evidence; In the operation confirmation area, the text is recognized by an OCR model to obtain credential evidence.

[0038] In the target scenario, at least two surveillance cameras are set up as monitoring viewpoints to collect multi-view video streams of the target object in the target scenario. Then, video frames in different states can be extracted from the video stream. In the interaction scenario, the target object will pick up or put back items. After capturing the video frames of the entire process before and after the target object's hand touches the item, it can be used as interaction evidence. In the target scenario, the target object will move frequently. After the target object interacts with the item, it needs to move to a designated area. The video frames of the entire process of the target object from the point of interaction to the preset operation confirmation area are captured as trajectory evidence. Finally, when the target object makes a payment in the operation confirmation area, the text recognition of the transaction process can be performed by the OCR model to obtain voucher evidence.

[0039] Preferably, the calculation steps for the interaction confidence level are as follows: Calculate the IoU sequence between the target object's hand bounding box and the item bounding box over N consecutive frames, and calculate the sustained contact score using exponential decay weighting. The calculation formula is as follows:

[0040] in To maintain continuous exposure to scores, As the attenuation factor, For example, 0.9 assigns a higher weight to recent contacts. Let t be the IoU sequence of the t-th frame; The statistical IoU exceeds the threshold Number of consecutive frames The duration fraction is mapped using an S-shaped function, and the calculation formula is as follows:

[0041] in Let k be the duration fraction, and k be the slope parameter of the sigmoid function. , The reference frame number for determining valid interactions, (Corresponds to 0.3-0.6 seconds of 25fps video); Based on the item's state before and after the interaction, if a state change occurs, such as changing from "on shelf" to "off shelf," the state change score is calculated. Otherwise, it is 0; The interaction confidence score is obtained by fusing the sustained contact score, duration score, and state change score. The calculation formula is as follows: ,in For interactive confidence, It is an adjustable weight, and It is obtained by optimizing on the validation set using the gradient descent method.

[0042] In the above formula, the attenuation factor ∈[0.8,0.95], preferably 0.9, assigning higher weight to recent contacts; S-shaped function slope parameter k∈[0.3,0.8]; number of reference frames for determining effective interactions. ∈[8,15] frames (corresponding to 0.3-0.6 seconds of 25fps video); adjustable weights α, β, and χ satisfy α+β+χ=1, and are obtained by optimization on the validation set through gradient descent. State mutation detection is achieved by comparing the confidence change and position shift of the item detection box before and after interaction. When the item changes from "on the shelf" to "off the shelf", it is determined that a state mutation has occurred.

[0043] The calculation of interaction confidence includes spatial contact quantification, contact persistence determination, item state change detection, and confidence fusion. By quantifying interaction confidence, the authenticity of the interaction between the target and the item can be accurately determined, avoiding misjudgment of core actions.

[0044] Preferably, the calculation steps for the trajectory compliance confidence level are as follows: Based on the scene topology map of the target object, the theoretically optimal path from the interaction point to the operation confirmation area is generated using the A* algorithm. The path cost takes into account both Euclidean distance and the type of travel area.

[0045] Using a dynamic time warping algorithm, the actual movement path point sequence is transformed. with the optimal path Align and calculate the minimum cumulative distance. ; Path similarity is calculated based on the minimum cumulative distance, and the formula is as follows:

[0046] in For path similarity, The length of the optimal path. For scale parameters, (meters), mapping the distance to similarity in (0,1]; When it is determined that the target object has entered the geofence of the operation confirmation area within the time window, the arrival score is output. ,otherwise ; The trajectory compliance confidence score is obtained by fusing path similarity and arrival score, and the calculation formula is as follows:

[0047] in For the confidence level of trajectory compliance, These are the path similarity weight and the arrival score weight, respectively.

[0048] In the above formula, the scale parameter λ∈[5,20] (meters) is used to map the DTW distance to path similarity in the interval (0,1); the optimal path length The theoretically optimal total path length output by the A* algorithm; the path similarity weight μ and the arrival score weight ν satisfy μ+ν=1, and are determined through validation set optimization based on the scenario. The actual movement path point sequence is obtained by sampling at a fixed frequency using a target tracking algorithm (such as DeepSORT), and the sampling points include timestamps and two-dimensional coordinates; DTW alignment uses Euclidean distance as a local distance metric, and the alignment window width is set according to the scene scale.

[0049] The calculation steps for trajectory compliance confidence include optimal path prediction, alignment of actual and predicted paths, path similarity calculation, destination arrival verification, and confidence fusion. By quantifying trajectory compliance confidence, the rationality of behavioral paths can be effectively verified, and it can be determined whether the process is followed to reach the designated area.

[0050] Preferably, the calculation steps for the confidence level of the certificate validity are as follows: An OCR model is used to obtain the recognition confidence of the text and each character, and the overall text confidence is calculated using the following formula:

[0051] in For the overall credibility of the text, Let i be the recognition confidence score of the i-th character output by the OCR model. To calculate the average value, To process the standard deviation, both the average confidence level and high variance results are considered; Using a pre-trained sentence embedding model (such as BERT), calculate the cosine similarity between the text recognized by the OCR model and the expected credential template (such as payment success). ; Based on the text recognition of the OCR model, the voucher time and voucher amount are verified to obtain a logical reward score. For each successful check, one logical reward score is added. The confidence score of the credential validity is calculated based on the overall text credibility, cosine similarity, and logical reward score, and truncated within the range of [0,1]. The calculation formula is as follows: ,in Assuming confidence level for the validity of the document, These are text credibility weight and semantic similarity weight, respectively.

[0052] In the above formula, the overall text credibility The cosine similarity is calculated by multiplying the mean character confidence score by (1 - standard deviation), which considers both average recognition quality and suppresses high variance results. The sentence embedding model is preferably a pre-trained model based on BERT, which encodes the OCR text and the expected voucher template into a 768-dimensional vector before calculating the cosine similarity. Logical reward score The reward is accumulated item by item through rule validation, with a fixed reward value added for each valid validation. The final confidence level of the credential validity is determined by this process. Cut off within the range [0,1].

[0053] The calculation steps for the confidence level of credential validity include text credibility calculation, semantic compliance calculation, logical consistency verification, and confidence level fusion. By quantifying the confidence level of credential validity, the compliance of operation results can be accurately verified, providing the final verification basis for the completion status of the behavior.

[0054] The key parameters involved in calculating the interaction confidence score, trajectory compliance confidence score, and voucher validity confidence score are all determined through validation set optimization based on the application scenario (retail loss prevention, warehouse management, self-service), including the attenuation factor. S-shaped function slope parameter k, number of reference frames Adjustable weights Scale parameters .

[0055] Preferably, the specific steps of step S3 are as follows: A joint decision-making model with dynamic weight adjustment is constructed. The inputs to the joint decision-making model are the interaction confidence, trajectory compliance confidence, and credential validity confidence. The output is the overall confidence of the target object completing the preset operation. The joint decision function of the joint decision-making model adopts the weighted geometric mean form, and the calculation formula is as follows:

[0056] in For the overall confidence level, Dynamic weights are dynamically adjusted based on prior knowledge of the target scenario and the real-time context. The dynamic weight adjustment mechanism includes: loading a baseline weight configuration based on the target scenario type (such as the default configuration for retail scenarios). The baseline weights are offset and corrected based on real-time context features, including the current time period (business hours / non-business hours), the population density in the monitored area, and the false alarm feedback statistics within the historical period. When the confidence of a single piece of evidence fluctuates abnormally (such as a sudden drop in interaction confidence while the trajectory confidence is normal), the weight of that piece of evidence is automatically reduced and the weights of the other evidence are increased to ensure the robustness of the fusion decision.

[0057] The overall confidence level is compared with a preset threshold range, and a status label representing the completion status of the operation is output. The status label includes strongly certain completion, possibly completion, uncertain, and possibly incomplete.

[0058] The joint decision function in the joint decision model uses a weighted geometric average to fuse the three types of confidence to obtain the overall confidence. Based on the threshold range into which the overall confidence falls, the final state label can be determined. The weights applied in the joint decision function can be dynamically adjusted according to the prior knowledge of the target scenario (such as retail or library) and the real-time context (such as crowd density or time period).

[0059] Preferably, step S4 includes the following specific steps: The continuous video stream is parsed to obtain semantic abstract representations of discrete behaviors of several target objects in the target scene, and output as behavior primitives. A sequence of behavior primitives is formed based on the behavior primitives. The behavioral primitive sequence is fed into the causal reasoning engine, which loads a scenario-based causal rule library to detect patterns in the behavioral primitive sequence that violate common sense causal relationships. At the same time, the abnormality of the behavioral primitive is judged by combining contextual information, and the detection results are obtained.

[0060] A continuous video stream can be parsed into several behavioral primitives. A behavioral primitive is a semantic abstraction of the discrete behavioral actions of a target object in a target scene, including but not limited to approach, pick up, move, wait, and present. Examples include Approach (shelf A), Grasp (product X), MoveTo (cashier), Wait (20 seconds), and Present (phone screen). Then, behavioral primitives can be combined into a behavioral primitive sequence according to the time sequence.

[0061] The causal rule base is constructed as follows: Based on the business process knowledge and compliance requirements of the target scenario, operational specifications are formalized into several causal constraint rules; each rule includes antecedent behavior primitives, consequent behavior primitives, the maximum allowed interval time, and anomaly severity level; the rule base supports incremental updates and version management through the scenario configuration management interface; when multiple rules are triggered simultaneously, the detection results are output first according to the anomaly severity level, and those with the same severity level are sorted by time. Example rules include, but are not limited to: Rule R1, if the behavior primitive sequence contains the GRASP behavior primitive, then it must be within the subsequent maximum interval time. If the MOVE_TO behavior primitive (destination is the operation confirmation area) appears within the timeout period, it is marked as a causal violation, and the severity level increases with the timeout duration; Rule R2, if the GRASP behavior primitive is detected outside of business hours, it is directly marked as a context exception.

[0062] Causal violation detection and context anomaly detection can be performed based on behavioral primitive sequences. Causal violation detection checks for patterns that violate common-sense causal relationships within the behavioral primitive sequence. This can be achieved by constructing a causal relation library containing several detection rules. One such rule is: if a behavioral primitive sequence contains the GRASP behavioral primitive, then the MOVE_TO behavioral primitive (with the destination being the operation confirmation area) must appear within a specified interval. The detection algorithm is as follows: traverse the behavioral primitive sequence; when a GRASP behavioral primitive is detected, start a timer; search for the MOVE_TO behavioral primitive within the subsequent maximum interval; if not found, mark a causal violation, indicating its severity. Where f is an increasing function, This refers to the timeout duration.

[0063] Similarly, context anomaly detection is also implemented through a context database containing multiple detection rules. One of the detection rules in the context database is: when GRASP behavioral primitives exist outside of business hours, they are marked as context anomalies. The detection algorithm is: when parsing each behavioral primitive, the business hours and authorized areas in the context database are queried simultaneously before making a judgment.

[0064] Preferably, in step S5, when the confidence level of the credential validity is lower than a preset threshold, the behavioral language model is triggered to infer intent, and the inference process includes: The sequence of behavioral primitives is converted into a vector sequence through an embedding layer, and then concatenated with the context feature vector or fused with attention to obtain the embedding sequence. The embedded sequence is encoded using a lightweight Transformer encoder or a bidirectional LSTM to obtain a context vector representing the overall semantics of the behavioral sequence. The context vector is input into a fully connected layer classifier, which outputs a probability distribution for a predefined set of intents, including purchase intent, browsing intent, suspicious intent, and work intent. The maximum probability value is determined from the probability distribution. When the maximum probability value is greater than the preset trigger threshold, the intent category corresponding to the maximum probability value is determined as the target object intent.

[0065] The training process of the behavioral language model includes: collecting historical behavioral primitive sequences and corresponding manually labeled intent tags from the target scene to construct a training sample set; converting the behavioral primitive sequences into vector sequences through an embedding layer, and concatenating or fusion with context feature vectors to obtain embedding sequences; encoding the embedding sequences using a lightweight Transformer encoder (preferably 4 layers, 256-dimensional hidden layers, and a 4-head attention mechanism) or a bidirectional LSTM to obtain context vectors representing the overall semantics of the behavioral sequences; inputting the context vectors into a fully connected layer classifier, which outputs a probability distribution for a predefined intent set; performing end-to-end training using a cross-entropy loss function, and preventing overfitting on the validation set through an early stopping mechanism. The intent tags include purchase intent, browsing intent, suspicious intent, and work intent, which are labeled by scene operators based on complete video playback and business rules.

[0066] When the confidence level of the credential validity is lower than a preset threshold (e.g., 0.6), it indicates weak credential evidence, triggering the behavioral language model to infer intent. This requires judging the intent of the target object. The behavioral language model used in this invention is implemented through a lightweight Transformer encoder or a bidirectional LSTM. First, the behavioral primitive sequence is converted into a vector sequence and fused with the context feature vector to obtain an embedding sequence. Then, the fused embedding sequence is fed into a lightweight Transformer encoder or a bidirectional LSTM for encoding to obtain a context vector. Intent classification can then be performed based on the context vector. A predefined intent set includes purchase intent, browsing intent, suspicious intent, and work intent. The context vector is processed by a fully connected layer classifier, which outputs a probability distribution containing the probability value for each type of intent. The maximum probability value is selected from the probability distribution and compared with a preset trigger threshold. If the maximum probability value is greater than the trigger threshold, the corresponding intent type can be determined. Finally, the target object's intent can assist in decision-making. For example, when inferred to be a browsing intent, it can explain the target object's prolonged lingering without payment, thereby suppressing false alarms.

[0067] Preferably, step S6 includes the following specific steps: A multi-layered evidence graph network, comprising a summary layer, a logic layer, and a data layer, is constructed and stored based on a graph database. In the summary layer, the overall confidence level and status labels are combined to describe the complete interaction process of the target object using natural language. In the logic layer, several behavioral primitive nodes are created according to the time sequence of the behavioral primitive sequence and marked with corresponding timestamps. For each behavioral primitive, the interaction confidence, trajectory compliance confidence, and credential validity confidence are associated. Temporal sequence and causal trigger relationship edges are established between behavioral primitive nodes. Causal violation results and context abnormal results are marked on the corresponding behavioral primitive nodes. At the same time, the target object intent is attached to the entire behavioral primitive sequence. In the data layer, video evidence nodes and credential evidence nodes are created. The video evidence nodes include spatiotemporal indexes of multiple video streams, key video segments of behavior, key frames of interaction, video segments of trajectory, and credential presentation frame nodes. The credential evidence nodes include credential OCR text results, semantic similarity results, and logical verification record nodes. The nodes within the logic layer and between the logic layer and the data layer are connected by relational edges. The relational edges between the nodes within the logic layer are temporal or causal relationships, and the relational edges between the nodes in the logic layer and the data layer are proof relationships.

[0068] The summary layer outlines the overall picture and core judgment results of the target object's behavioral events, and is used to quickly browse the event summary. The logic layer carries the behavioral primitives, confidence levels and reasoning conclusions, and is used to construct the behavioral sequence and causal logic chain. The data layer stores evidence materials such as original videos and keyframes, and provides traceable empirical evidence for the upper-level judgment.

[0069] The multi-layered evidence graph network supports on-demand retrieval based on a graph traversal query language: when a quick overview of an event is needed, only summary-level nodes are returned; when performing in-depth audits, starting from the logic-level nodes, the network traverses along the "proof" relationship edges to the video evidence nodes and credential evidence nodes in the data layer to extract the complete evidence chain. The evidence graph network can dynamically generate evidence packages of varying levels of detail according to application requirements, supporting layer-by-layer tracing from the top-level summary to the bottom-level original evidence. The nodes of the multi-level evidence graph network are connected by relational edges such as "contains", "triggers", and "proved by", forming a traversable logical network that supports on-demand querying. It can dynamically generate evidence packages of different levels of detail (such as only the summary layer or the summary layer + logical layer) according to application requirements (quick browsing, in-depth auditing). Finally, the graph traversal query language enables tracing from the top-level summary to the bottom-level evidence.

[0070] See Figure 4 The system shown is based on a behavioral semantic understanding method that combines spatiotemporal causal reasoning and confidence fusion, and includes: At least two edge computing nodes are deployed in the target scene. Each edge computing node integrates an image sensor and a computing unit to perform video stream analysis in real time, extract interaction evidence and trajectory evidence of the target object and items, and generate a structured evidence data package with spatiotemporal stamps. The central server, deployed in the cloud, communicates with the edge computing nodes and is configured as follows: Receive and associate structured evidence data packets from different edge nodes, calculate interaction confidence based on interaction evidence, and calculate trajectory compliance confidence based on trajectory evidence; Run a dynamic fusion decision service, adjust weights based on real-time context, and fuse the interaction confidence and trajectory compliance confidence from edge computing nodes with the credential validity confidence from cloud-recognized data; Run the causal reasoning service, load the scenario-based causal rule library, and reason about the behavioral primitive sequences obtained by parsing the video stream reported from the edge computing node; Run the intent inference service. When the confidence level of the credentials is insufficient, load the behavior language model to infer the intent from the behavior primitive sequence obtained by parsing the video stream reported from the edge computing node, and output the intent category probability distribution as an auxiliary decision-making basis. Run the evidence graph management service, which builds and stores a multi-layered evidence graph network based on a graph database, including a summary layer, a logic layer, and a data layer, and supports layer-by-layer traceability query from the summary to the original evidence.

[0071] Edge computing nodes run lightweight target detection and tracking models for accurate detection and tracking of target objects. Meanwhile, the cloud-based central server deploys dynamic fusion decision services, causal reasoning services, and evidence graph management services in a microservice format. Furthermore, a scenario configuration management interface is provided to inject behavioral primitive definition libraries, causal rule libraries, and confidence fusion benchmark weight strategies for different application scenarios into the cloud-based central server. The edge computing nodes and the cloud-based central server cluster communicate via an event-driven message middleware to ensure the temporal consistency and low-latency transmission of evidence data. The system also includes a scenario configuration management interface for injecting behavioral primitive definition libraries, causal rule libraries, and weight benchmark strategies for different application scenarios into the cloud server.

[0072] Reference Figure 5 As shown, the effectiveness of the present invention will be discussed below through an embodiment: Example: Complete application in a retail loss prevention scenario The system of this invention was deployed in a supermarket to automatically identify the integrity of customers' "take-pay" behavior.

[0073] System Deployment and Data Flow: Edge devices (deployed near shelves and checkout counters): Smart cameras equipped with GPUs, running YOLOv7 object detection, DeepSORT tracking, and MediaPipe hand keypoint detection models.

[0074] Cloud-based central server: Deploys core modules such as fusion decision-making, causal reasoning, and evidence graph management.

[0075] Communication: Edge devices upload structured evidence features (not the original video) to the cloud in real time via an encrypted link.

[0076] Workflow example: Customers enter the store: 14:20:00: The edge camera detected that Zhang San was in front of the "beverage shelf" and his hand interacted with "product A". =0.92. The system generates the GRASP (product A) behavioral primitive.

[0077] 14:20:30: Tracking shows that the customer's movement path highly matches the predicted "to the checkout" path. =0.88.

[0078] 14:21:00: At the checkout area, the OCR system recognized the "Payment Successful" page on Zhang San's phone screen. =0.95.

[0079] Decision fusion: Current pedestrian flow is normal, with weights set to the default values ​​{0.3, 0.3, 0.4}.

[0080] Calculated = 0.92^0.3 * 0.88^0.3 * 0.95^0.4 ≈ 0.92.

[0081] Causal reasoning: The sequence of behavioral primitives [APPROACH (beverage shelf), GRASP (product A), MOVE_TO (cashier), PRESENT (phone)] conforms to the causal rules and has no anomalies.

[0082] Output: Status label is "Strongly Confirmed Completed", evidence image automatically generated. No risk warning.

[0083] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion, characterized in that, Includes the following steps: Step S1: Simultaneously acquire video streams of the target scene from at least two monitoring viewpoints with spatiotemporal correlation, and extract interaction evidence, trajectory evidence, and credential evidence of the target object from the video streams; Step S2: Calculate the confidence level based on interaction evidence, trajectory evidence, and credential evidence to obtain the interaction confidence level, trajectory compliance confidence level, and credential validity confidence level. Step S3: Dynamically weight and fuse the interaction confidence, trajectory compliance confidence, and credential validity confidence through the constructed joint decision-making model. The dynamic weights are adaptively adjusted according to the prior knowledge of the target scenario and the real-time context, and the overall confidence is output and the state label is determined. Step S4: Parse the video stream into discrete behavioral primitive sequences, and use a causal inference engine to perform causal violation detection and context anomaly detection on the behavioral primitive sequences; Step S5: Infer the target object's intent based on the sequence of behavioral primitives using a behavioral language model. The behavioral language model is used to assist decision-making when the confidence level of the credentials is insufficient. Step S6: Construct a multi-layered evidence graph network containing a summary layer, a logic layer, and a data layer to structurally represent and trace the interactive behavior of the target object.

2. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion as described in claim 1, characterized in that, The specific steps of step S1 are as follows: Simultaneously acquire video streams from at least two monitoring viewpoints that are spatially and temporally related; Extract video frames from the video stream that show the entire process of the target object's hand touching the object as evidence of interaction; Extract video frames from the video stream that contain the entire process of the target object moving from the interaction point to the preset operation confirmation area as trajectory evidence; In the operation confirmation area, the text is recognized by an OCR model to obtain credential evidence.

3. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion as described in claim 1, characterized in that, The steps for calculating the interaction confidence are as follows: Calculate the IoU sequence between the target object's hand bounding box and the item bounding box over N consecutive frames, and calculate the sustained contact score using exponential decay weighting. The calculation formula is as follows: in To maintain continuous exposure to scores, As the attenuation factor, Let t be the IoU sequence of the t-th frame; The statistical IoU exceeds the threshold Number of consecutive frames The duration fraction is mapped using an S-shaped function, and the calculation formula is as follows: in Let k be the duration fraction, and k be the slope parameter of the sigmoid function. The reference frame number for determining valid interactions; Based on the item's state before and after the interaction, if a state change occurs, the state change score is calculated. Otherwise, it is 0; The interaction confidence score is obtained by fusing the sustained contact score, duration score, and state change score. The calculation formula is as follows: ,in For interactive confidence, It is an adjustable weight, and .

4. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion according to claim 1, characterized in that, The steps for calculating the confidence level of the trajectory compliance are as follows: Based on the scene topology map of the target object, the theoretically optimal path from the interaction point to the operation confirmation area is generated using the A* algorithm. ; Using a dynamic time warping algorithm, the actual movement path point sequence is transformed. with the optimal path Align and calculate the minimum cumulative distance. ; Path similarity is calculated based on the minimum cumulative distance, and the formula is as follows: in For path similarity, For scale parameters, The length of the optimal path; When it is determined that the target object has entered the geofence of the operation confirmation area within the time window, the arrival score is output. ,otherwise ; The trajectory compliance confidence score is obtained by fusing path similarity and arrival score, and the calculation formula is as follows: in For the confidence level of trajectory compliance, These are the path similarity weight and the arrival score weight, respectively.

5. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion according to claim 1, characterized in that, The steps for calculating the confidence level of the validity of the voucher are as follows: An OCR model is used to obtain the recognition confidence of the text and each character, and the overall text confidence is calculated using the following formula: in For the overall credibility of the text, Let i be the recognition confidence score of the i-th character output by the OCR model. To calculate the average value, To calculate the standard deviation; Using a pre-trained sentence embedding model, the cosine similarity between the text recognized by the OCR model and the expected voucher template is calculated. ; Based on the text recognition of the OCR model, the voucher time and voucher amount are verified to obtain a logical reward score. ; The confidence level of credential validity is calculated based on overall text credibility, cosine similarity, and logical reward score. The calculation formula is as follows: ,in Assuming confidence level for the validity of the document, These are text credibility weight and semantic similarity weight, respectively.

6. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion according to claim 1, characterized in that, The specific steps of step S3 are as follows: A joint decision-making model with dynamic weight adjustment is constructed. The inputs to the joint decision-making model are the interaction confidence, trajectory compliance confidence, and credential validity confidence. The output is the overall confidence of the target object completing the preset operation. The joint decision function of the joint decision-making model adopts the weighted geometric mean form, and the calculation formula is as follows: in For the overall confidence level, Dynamic weights are dynamically adjusted based on prior knowledge of the target scenario and the real-time context. The overall confidence level is compared with a preset threshold range, and a status label representing the completion status of the operation is output.

7. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion according to claim 1, characterized in that, The specific steps of step S4 include: The continuous video stream is parsed to obtain semantic abstract representations of discrete behaviors of several target objects in the target scene, and output as behavior primitives. A sequence of behavior primitives is formed based on the behavior primitives. The behavioral primitive sequence is fed into the causal reasoning engine to detect patterns in the behavioral primitive sequence that violate common sense causal relationships. At the same time, the abnormality of the behavioral primitive is judged by combining contextual information, and the detection results are obtained.

8. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion according to claim 1, characterized in that, In step S5, when the confidence level of the credential validity is lower than a preset threshold, the behavioral language model is triggered to infer intent. The inference process includes: The sequence of behavioral primitives is converted into a vector sequence through an embedding layer, and then concatenated with the context feature vector or fused with attention to obtain the embedding sequence. The embedded sequence is encoded using a lightweight Transformer encoder or a bidirectional LSTM to obtain a context vector representing the overall semantics of the behavioral sequence. The context vector is input into a fully connected layer classifier, which outputs a probability distribution for a predefined set of intents, including purchase intent, browsing intent, suspicious intent, and work intent. The maximum probability value is determined from the probability distribution. When the maximum probability value is greater than the preset trigger threshold, the intent category corresponding to the maximum probability value is determined as the target object intent.

9. The behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion according to claim 1, characterized in that, The specific steps of step S6 include: In the summary layer, the overall confidence level and status labels are combined to describe the complete interaction process of the target object using natural language. In the logic layer, several behavioral primitive nodes are created according to the time sequence of the behavioral primitive sequence and marked with corresponding timestamps. For each behavioral primitive, the interaction confidence, trajectory compliance confidence, and credential validity confidence are associated. Temporal sequence and causal trigger relationship edges are established between behavioral primitive nodes. Causal violation results and context abnormal results are marked on the corresponding behavioral primitive nodes. At the same time, the target object intent is attached to the entire behavioral primitive sequence. In the data layer, video evidence nodes and credential evidence nodes are created. The video evidence nodes include spatiotemporal indexes of multiple video streams, key video segments of behavior, key frames of interaction, video segments of trajectory, and credential presentation frame nodes. The credential evidence nodes include credential OCR text results, semantic similarity results, and logical verification record nodes. The nodes within the logic layer and between the logic layer and the data layer are connected by relational edges. The relational edges between the nodes within the logic layer are temporal or causal relationships, and the relational edges between the nodes in the logic layer and the data layer are proof relationships.

10. A system applying the behavioral semantic understanding method based on spatiotemporal causal reasoning and confidence fusion as described in any one of claims 1-9, characterized in that, include: At least two edge computing nodes are deployed in the target scene. Each edge computing node integrates an image sensor and a computing unit to perform video stream analysis in real time, extract interaction evidence and trajectory evidence of the target object and items, and generate a structured evidence data package with spatiotemporal stamps. The central server, deployed in the cloud, communicates with the edge computing nodes and is configured as follows: Receive and associate structured evidence data packets from different edge nodes, calculate interaction confidence based on interaction evidence, and calculate trajectory compliance confidence based on trajectory evidence; Run a dynamic fusion decision service, adjust weights based on real-time context, and fuse the interaction confidence and trajectory compliance confidence from edge computing nodes with the credential validity confidence from cloud-recognized data; Run the causal reasoning service, load the scenario-based causal rule library, and reason about the behavioral primitive sequences obtained by parsing the video stream reported from the edge computing node; Run the intent inference service. When the confidence level of the credentials is insufficient, load the behavioral language model to infer the intent from the sequence of behavioral primitives. Run the evidence graph management service to build and store a multi-layered evidence graph network containing a summary layer, a logic layer, and a data layer based on a graph database.