Examination room multi-source data fusion abnormal behavior intelligent analysis method and system
By collecting multi-source data in the examination room, performing multi-view fusion and hidden Markov model analysis, combining wavelet decomposition and spatial weight coefficients to identify abnormal behaviors in the examination room, the problem of insufficient recognition ability of complex behaviors and coordinated cheating behaviors in the existing technology is solved, and more efficient and intelligent abnormal behavior monitoring is achieved.
Patent Information
- Application Number
- CN202510494732.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The prior art has multiple shortcomings in the identification of abnormal behavior in the examination room, including the dependence of a single data source, insufficient ability to recognize complex behaviors, limitations in behavior feature extraction and acoustic feature analysis, and failure to consider spatial relationships and dynamic changes between candidates, resulting in weak ability to identify coordinated cheating behaviors.
The intelligent analysis method of multi-source data fusion in the examination room is adopted, and by collecting video data and audio data, multi-view bone point fusion and sound source positioning are performed to extract behavioral feature sequences and acoustic features. Then, based on these features, the hidden Markov model is constructed, and the distribution difference between answering and thinking states is calculated using kernel density estimation and dynamic adaptive weights, the state behavior consistency is evaluated, and abnormal behavior recognition is performed based on the probability distribution entropy value. At the same time, through wavelet decomposition and spatial weight coefficient, the correlation relationship between seat numbers is established to identify the collaborative cheating behavior of multiple people.
It improves the comprehensiveness and accuracy of examination room monitoring, enhances the intelligence and dynamic adaptability of abnormal behavior recognition, can promptly detect and warn of cheating behaviors, and ensures the fairness and impartiality of examination room.
Smart Images

Figure CN120220246A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and particularly to an intelligent analysis method and system for abnormal behaviors through multi-source data fusion in an examination room. Background Art
[0002] In recent years, the use of multi-source data fusion technology for examination room monitoring and abnormal behavior analysis has gradually become a research hotspot. Through the collection and analysis of video and audio data, the behaviors of examinees can be comprehensively monitored, thereby improving the security and fairness of the examination room. However, there are some defects and deficiencies in the existing technology for identifying abnormal behaviors in the examination room. First of all, existing monitoring systems often rely only on a single data source, lacking multi-dimensional information fusion, resulting in insufficient ability to identify complex behaviors. Secondly, there are limitations in behavior feature extraction and acoustic feature analysis in the existing technology, and it is unable to effectively capture subtle changes in abnormal behaviors, affecting the accuracy of identification. Finally, most of the existing abnormal behavior identification methods are static analysis, failing to consider the spatial relationship and dynamic changes among examinees, resulting in weak ability to identify collaborative cheating behaviors. Summary of the Invention
[0003] Embodiments of the present invention provide an intelligent analysis method and system for abnormal behaviors through multi-source data fusion in an examination room, which can solve the problems in the existing technology.
[0004] In the first aspect of the embodiments of the present invention, an intelligent analysis method for abnormal behaviors through multi-source data fusion in an examination room is provided, including: Collecting video data and audio data in the examination room, performing multi-view skeleton point fusion and sound source localization on the detection area corresponding to each seat number, and extracting behavior feature sequences and acoustic features; Constructing a hidden Markov model based on the behavior feature sequences and acoustic features, calculating the distribution difference degree of the answering and thinking states using kernel density estimation to obtain a dynamic adaptive weight, evaluating the state behavior consistency through the matching degree between the state smoothed posterior probability and the personalized normal behavior model, and identifying abnormal behaviors within a sliding time window in combination with the probability distribution entropy value; Calculating a spatial weight coefficient based on the physical distance between the seat numbers determined to have abnormal behaviors and the surrounding seats, performing wavelet decomposition on the behavior feature sequences and acoustic features of the relevant seat numbers to obtain multi-scale coefficients, establishing an association relationship between the seat numbers according to the multi-scale coefficients, recursively calculating the behavior-acoustic association degree between the seat numbers at different time scales, and identifying associated abnormal behaviors and multi-person collaborative cheating behaviors in combination with the spatial weight coefficient; Recording the detected cheating behaviors and pushing warning information in real time.
[0005] In an alternative embodiment, Extracting behavioral feature sequences and acoustic features includes: Collecting video data using multiple cameras, calculating the confidence of human body bone nodes to obtain node visibility weights and selecting the optimal viewing angle, performing temporal prediction and structural constraint on the human body bone nodes under the optimal viewing angle, and performing weighted fusion on the constrained bone nodes and the corresponding bone nodes under other viewing angles to obtain the fusion position; Constructing a spatio-temporal graph structure based on the fusion position and extracting action features, performing temporal segmentation on the action features and mapping them to preset examination scene action primitives to obtain action semantic features; calculating the three-dimensional angle of the head and the fixation center point based on the fusion position, constructing an attention distribution map and performing temporal tracking to obtain an attention transfer sequence; inputting the action semantic features and the attention transfer sequence into a long short-term memory network to obtain an intention state probability distribution; Calculating the similarity between the action semantic features and the attention transfer sequence to obtain a modality consistency score, calculating a feature fusion weight based on the modality consistency score, and performing weighted fusion on the action semantic features, attention transfer sequence and intention state probability distribution and applying a temporal smoothing constraint to obtain a behavioral feature sequence; For the collected audio data, performing sound source localization based on cross-correlation time delay estimation and sound source enhancement algorithm, associating the localization result with the seat number; performing short-time energy calculation and double-threshold speech segment detection on the associated audio data, and extracting volume features and duration features as acoustic features.
[0006] In an alternative embodiment, Constructing a hidden Markov model based on the behavioral feature sequence and acoustic features, calculating the distribution difference degree of the answering and thinking states using kernel density estimation to obtain a dynamic adaptive weight, evaluating the state behavior consistency through the matching degree between the state smoothed posterior probability and the personalized normal behavior model, and performing abnormal behavior recognition within a sliding time window in combination with the probability distribution entropy value, including: Performing normalization processing on the behavioral feature sequence and acoustic features to obtain a normalized feature sequence; constructing a hidden Markov model based on the normalized feature sequence, calculating the state distribution difference degree using kernel density estimation and KL divergence, and obtaining a dynamic adaptive weight through Sigmoid mapping to construct a time-varying state transition probability matrix; in the hidden Markov model, calculating the state smoothed posterior probability using the product of the forward variable and the backward variable, the forward variable is obtained by recursively multiplying the state emission probability and the time-varying state transition probability, and the backward variable is obtained by recursively multiplying the state emission probability at the next moment and the time-varying state transition probability; Construct a personalized normal behavior model for the examinee based on the normalized feature sequence collected within the first answering cycle after the start of the exam, calculate the matching degree between the state smoothed posterior probability and the state parameters of the personalized normal behavior model to obtain a state-behavior consistency score; Calculate the probability distribution entropy value of the state smoothed posterior probability within the sliding windows of multiple time scales, and perform adaptive correction on the anomaly determination in combination with the exam time process and the surrounding seat behavior baseline. When the anomaly scores of multiple consecutive time scales exceed the adaptive anomaly threshold, it is determined as abnormal behavior.
[0007] In an alternative implementation, Construct a hidden Markov model based on the normalized feature sequence, calculate the state distribution difference degree using kernel density estimation and KL divergence, and obtain a dynamic adaptive weight through Sigmoid mapping to construct a time-varying state transition probability matrix, including: Construct a hidden Markov model including the answering state and the thinking state based on the normalized feature sequence, use a state indicator function to mark the states of the examinee's answering behavior sequence, construct an answering state sample set from the feature samples marked as the answering state, and construct a thinking state sample set from the feature samples marked as the thinking state; Use the kernel density estimation method to calculate the feature distribution density functions of the answering state sample set and the thinking state sample set respectively, calculate the KL divergence based on the feature distribution density functions to obtain the feature distribution difference degree, and perform normalization processing on the feature distribution difference degree to obtain the normalized difference degree; Calculate the dynamic adaptive weight based on the normalized difference degree, and the dynamic adaptive weight maps the normalized difference degree through a Sigmoid mapping function; Use the dynamic adaptive weight to construct a time-varying state transition probability matrix, where the state transition probability is obtained by weighting the basic transition probability and the conditional transition probability based on the current observation by the dynamic adaptive weight, and the time-varying state transition probability matrix is used for subsequent state sequence estimation.
[0008] In an alternative implementation, Calculating the probability distribution entropy value of the state smoothed posterior probability within the sliding windows of multiple time scales, and performing adaptive correction on the anomaly determination in combination with the exam time process and the surrounding seat behavior baseline. When the anomaly scores of multiple consecutive time scales exceed the adaptive anomaly threshold, it is determined as abnormal behavior, including: Calculate the probability distribution entropy value of the state smoothed posterior probability within the sliding windows of three time scales: short-term, medium-term, and long-term. The short-term window is used to capture sudden anomalies, the medium-term window is used to detect persistent anomalies, and the long-term window is used to evaluate the overall anomaly trend; Calculate the current exam time process based on the exam start time, divide the exam process into three stages: start, middle, and end, and set different abnormal scoring weights for the probability distribution entropy values in different stages; Obtain the smoothed posterior probability of the status of other seats within a preset range around the determined seat number, calculate the mean value of the probability distribution entropy of the local area as the behavior baseline, and use the behavior baseline to correct the status-behavior consistency score; Calculate the fusion abnormal score based on the probability distribution entropy values at different time scales, the exam stage weights, and the corrected status-behavior consistency score; calculate the adaptive abnormal threshold based on the fusion abnormal score within the first answering cycle. When the fusion abnormal scores of three consecutive time scales all exceed the corresponding adaptive abnormal threshold, it is determined as abnormal behavior.
[0009] In an alternative implementation, Calculate the spatial weight coefficient based on the physical distance between the seat numbers determined to have abnormal behavior and the surrounding seats. Perform wavelet decomposition on the behavior feature sequence and acoustic features of the relevant seat numbers to obtain multi-scale coefficients. Establish the correlation relationship between seat numbers based on the multi-scale coefficients. Identify the correlated abnormal behavior by recursively calculating the behavior-acoustic correlation degree between seat numbers at different time scales. The identification of multi-person collaborative cheating behavior includes: Calculate the spatial weight coefficient based on the physical distance between the seat numbers determined to have abnormal behavior and the surrounding seats. The spatial weight coefficient is obtained by mapping the physical distance between seats through a Gaussian kernel function and normalizing it; Perform wavelet decomposition on the behavior feature sequence and acoustic feature sequence of the relevant seat numbers respectively to obtain multi-scale coefficients. The multi-scale coefficients include the approximation coefficients and detail coefficients of the behavior feature sequence and the approximation coefficients and detail coefficients of the acoustic feature sequence; calculate the cross-correlation coefficient between seat numbers at different time scales based on the multi-scale coefficients. The cross-correlation coefficient is obtained by dividing the covariance of the multi-scale coefficients of the behavior feature sequence and the multi-scale coefficients of the acoustic feature sequence by the product of the standard deviations; perform weighted fusion on the cross-correlation coefficients to obtain the fusion correlation degree, and use the fusion correlation degree as the initial correlation degree; Construct an initial correlation relationship between seat numbers based on the initial correlation degree, use the product of the spatial weight coefficient and the initial correlation degree as the weighted correlation degree between seat numbers, and determine the correlation degree threshold based on the fluctuation range of the weighted correlation degree within the first answering cycle after the exam starts. When the weighted correlation degree exceeds the correlation degree threshold and the duration exceeds the preset time window, it is determined that there is a preliminary collaborative cheating behavior between the relevant seat numbers.
[0010] In an alternative implementation, The further identification of the preliminary collaborative cheating behavior includes: Construct a time-varying graph structure with seat numbers as nodes, take the behavior feature sequence and acoustic features corresponding to the seat numbers as node features, and obtain edge weights based on the adaptive weighting of the physical distance, behavior similarity, and acoustic correlation between seat numbers; Establish a state vector containing behavior state parameters and acoustic state parameters for each seat number, and construct a state transition equation based on the state vector and the edge weights. The state transition equation includes an individual behavior evolution term and an interaction influence term. The individual behavior evolution term is used to calculate the state transition probability distribution, and the interaction influence term obtains the state influence by calculating the state vector difference degree and performing a non-linear mapping with the edge weights. Obtain the group synchronization index based on the state influence; Calculate the correlation degree of the behavior switching sequence between seat numbers as the time series collaboration feature, calculate the consistency of the state transition probability distribution as the state collaboration feature, calculate the neighborhood density with similar behaviors as the group aggregation feature, and calculate the behavior trigger relationship as the causal collaboration feature; Perform multi-level fusion on the time series collaboration feature, state collaboration feature, group aggregation feature, and causal collaboration feature to obtain the collaborative behavior feature, and determine the collaborative cheating behavior based on the collaborative behavior feature and the group synchronization index.
[0011] In the second aspect of the embodiments of the present invention, Provide an intelligent analysis system for abnormal behaviors in multi-source data fusion in an examination room, including: The first unit is used to collect video data and audio data in the examination room, perform multi-view skeleton point fusion and sound source localization on the detection area corresponding to each seat number, and extract the behavior feature sequence and acoustic features; The second unit is used to construct a hidden Markov model based on the behavior feature sequence and acoustic features, calculate the distribution difference degree of the answering and thinking states using kernel density estimation to obtain the dynamic adaptive weight, evaluate the state behavior consistency through the matching degree between the state smoothed posterior probability and the personalized normal behavior model, and identify abnormal behaviors within a sliding time window in combination with the probability distribution entropy value; The third unit is used to calculate the spatial weight coefficient based on the physical distance between the seat numbers determined to have abnormal behaviors and the surrounding seats, perform wavelet decomposition on the behavior feature sequences and acoustic features of the relevant seat numbers to obtain multi-scale coefficients, establish the association relationship between seat numbers according to the multi-scale coefficients, recursively calculate the behavior-acoustic correlation degree between seat numbers at different time scales, and identify the associated abnormal behaviors and multi-person collaborative cheating behaviors in combination with the spatial weight coefficient; The fourth unit is used to record the detected cheating behaviors and push warning information in real time.
[0012] In the third aspect of the embodiments of the present invention, Provided is an electronic device, comprising: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.
[0013] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0014] Advantages of the present application: Improve the comprehensiveness and accuracy of examination room monitoring. Through multi-source data fusion technology, it is possible to comprehensively collect the behavioral and acoustic characteristics in the examination room, providing a richer data basis for the identification of abnormal behaviors.
[0015] Enhance the intelligence and dynamic adaptability of abnormal behavior recognition. By using the hidden Markov model and kernel density estimation method, it is possible to analyze the answering and thinking states of candidates in real time, dynamically adjust the recognition weights, thereby improving the accuracy and sensitivity of recognition.
[0016] Facilitate the timely discovery and warning of cheating behaviors. By calculating the spatial weight coefficient between seats and the behavior-acoustic correlation degree, it is possible to effectively identify multi-person collaborative cheating behaviors and push warning information in real time, ensuring the fairness and impartiality of the examination room. Description of the Drawings
[0017] Figure 1 is a schematic flowchart of an intelligent analysis method for abnormal behaviors of multi-source data fusion in an examination room according to an embodiment of the present invention; Figure 2 is a multi-time scale sliding window entropy value fusion detection graph; Figure 3 is a graph of the change of group synchronization index and multi-dimensional collaborative feature analysis. Detailed Embodiments
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] The technical solution of the present invention will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0020] Figure 1 It is a schematic flowchart of an intelligent analysis method for abnormal behaviors in multi-source data fusion in an examination room according to an embodiment of the present invention. As Figure 1 shown, the method includes: Collect video data and audio data in the examination room, perform multi-viewpoint skeleton point fusion and sound source localization on the detection area corresponding to each seat number, and extract behavior feature sequences and acoustic features; Based on the behavior feature sequences and acoustic features, construct a hidden Markov model, calculate the distribution difference degree of the answering and thinking states using kernel density estimation to obtain a dynamic adaptive weight, evaluate the state behavior consistency through the matching degree between the smoothed posterior probability of the state and the personalized normal behavior model, and identify abnormal behaviors within a sliding time window in combination with the probability distribution entropy value; Based on the physical distance between the seat numbers determined to have abnormal behaviors and the surrounding seats, calculate the spatial weight coefficient, perform wavelet decomposition on the behavior feature sequences and acoustic features of the relevant seat numbers to obtain multi-scale coefficients, establish the correlation relationship between the seat numbers according to the multi-scale coefficients, recursively calculate the behavior-acoustic correlation degree between the seat numbers at different time scales, and identify the associated abnormal behaviors and multi-person collaborative cheating behaviors in combination with the spatial weight coefficient; Record the detected cheating behaviors and push warning information in real time.
[0021] In an alternative embodiment, extracting the behavior feature sequences and acoustic features includes: Use multiple cameras to collect video data, calculate the confidence of human skeleton nodes to obtain node visibility weights and select the optimal view, perform temporal prediction and structural constraint on the human skeleton nodes under the optimal view, and perform weighted fusion on the constrained skeleton nodes and the corresponding skeleton nodes under other views to obtain the fusion position; Based on the fusion position, construct a spatio-temporal graph structure and extract action features, perform temporal segmentation on the action features and map them to the preset action primitives of the examination scene to obtain action semantic features; calculate the three-dimensional angle of the head and the fixation center point based on the fusion position, construct an attention distribution map and perform temporal tracking to obtain an attention migration sequence; input the action semantic features and the attention migration sequence into a long short-term memory network to obtain the intention state probability distribution; Calculate the similarity between the action semantic features and the attention transfer sequence to obtain a modality consistency score, calculate the feature fusion weight based on the modality consistency score, and perform weighted fusion on the action semantic features, attention transfer sequence, and intention state probability distribution, and apply a temporal smoothing constraint to obtain a behavior feature sequence; For the collected audio data, perform sound source localization based on cross-correlation time-delay estimation and sound source enhancement algorithms, and associate the localization result with the seat number; perform short-time energy calculation and double-threshold speech segment detection on the associated audio data, and extract the volume feature and duration feature as acoustic features.
[0022] Exemplarily, video data is collected by multiple cameras, and each camera is arranged at different positions in the examination room, such as the four corners and the central top of the classroom, to ensure comprehensive coverage of the candidates' behaviors. Apply a human body skeleton detection algorithm to the video data from each perspective to obtain human body skeleton nodes, including 25 key points such as the head, neck, shoulders, elbows, and wrists. To determine the optimal perspective, calculate the confidence of each skeleton node. The confidence calculation considers the node detection score and its stability, and the value range is 0-1. For example, the confidence of a candidate's right wrist node is 0.92 in perspective A, 0.78 in perspective B, and 0.45 in perspective C, then perspective A is selected as the optimal perspective for this node. Perform temporal prediction on the skeleton nodes in the optimal perspective, and use the Kalman filter method to predict the possible positions of the nodes in the next frame. If the deviation between the predicted position and the actual detected position exceeds the threshold (such as 15 pixels), then start structural constraint correction. The structural constraint is based on the principle of constant human body skeleton length, and restricts the change in the distance between adjacent skeleton nodes not to exceed 5%. For example, the length of the upper arm should remain relatively stable in the sequence. If there is an abnormal elongation, it is corrected according to the historical average value. Perform weighted fusion on the constrained skeleton nodes and the corresponding nodes in other perspectives, and the fusion weight is proportional to the node visibility weight. For example, the visibility weights of a key point in three perspectives are 0.9, 0.7, and 0.4 respectively, then the corresponding weight ratio is 45%, 35%, and 20% when calculating the fusion position.
[0023] Construct a spatio-temporal graph structure based on the fusion position. The graph nodes are skeleton key points, and the edges include spatial connections (skeleton connections) and temporal connections (connections between the same key points in adjacent frames). Apply a graph convolutional network to this graph structure to extract action features, and the feature dimension is 128. Apply a sliding window (window size is 30 frames, step size is 15 frames) to the action feature sequence for temporal segmentation, and map each segment to a preset examination scene action primitive through a classifier, such as 12 basic actions like "writing", "raising the head", "turning the head", "raising the hand", etc. The mapping result forms action semantic features, with a dimension of 12, and each dimension represents the confidence of the corresponding action primitive.
[0024] Calculate the three-dimensional angles of the head based on the fused position, including the pitch angle (nodding up and down), yaw angle (turning left and right), and roll angle (tilting the head). Determine the head orientation by identifying facial feature points, and then combine with the field of view model to predict the fixation center point. The fixation center point is represented as (x, y, z) in the examination room coordinate system. For example, (150, 200, 50) cm represents the possible desktop position that the examinee may be looking at.
[0025] When constructing the attention distribution map, with the fixation center point as the center and a radius of the visual cone range (about 30-degree view angle), a Gaussian distribution is constructed with the intensity decaying with the distance from the center. Perform temporal tracking on the attention distribution map, record the movement trajectory of the center point within 5 seconds, and generate an attention transfer sequence. For example, the saccade trajectory from the upper left corner to the upper right corner and then to the lower left corner of the test paper can be represented as a sequence of consecutive coordinate points. Input the action semantic features and the attention transfer sequence into a long short-term memory network (the network contains 64 hidden units), and output the probability distribution of the intention states. The intention states include 8 kinds of examination behavior intentions such as "concentrating on answering questions", "thinking", "checking references", "communicating with others", etc., and the probability distribution represents the possibility of each intention.
[0026] Calculate the similarity between the action semantic features and the attention transfer sequence to obtain the modality consistency score. The similarity calculation uses the cosine similarity method. After mapping the two features to the same dimensional space, calculate the cosine value of the vector angle, and the value range is from -1 to 1. For example, the action of "lowering the head to write" is highly consistent with the attention distribution of looking at the test paper area, and the similarity can reach above 0.85; while the action of "raising the head" is inconsistent with the attention distribution of looking at the test paper, and the similarity may be as low as 0.3.
[0027] Calculate the feature fusion weights based on the modality consistency score. The higher the consistency score, the greater the corresponding feature weight. For example, when the consistency score is 0.9, the action feature weight can be set to 0.5, the attention feature weight to 0.3, and the intention state weight to 0.2; when the consistency score is 0.4, it can be adjusted to 0.3, 0.3, 0.4. Perform weighted fusion on the action semantic features, the attention transfer sequence, and the intention state probability distribution to obtain the behavior feature vector. Apply a temporal smoothing constraint to the behavior feature sequence, using the exponential weighted moving average method. The current frame feature value calculation takes into account the influence of historical frames to avoid drastic fluctuations in features. The smoothing parameter is set to 0.8, that is, the current frame retains 20% of the original features and fuses 80% of the historical features.
[0028] In terms of acoustic feature extraction, classroom audio data is collected through multiple microphone arrays. The microphone arrays can be arranged in the front, rear, and both sides of the classroom, and each array contains 4 - 8 microphones. Based on the cross - correlation time - delay estimation algorithm, the sound source position is calculated. The cross - correlation function is calculated for the signals received by two microphones, the time difference corresponding to the cross - correlation peak is found, and the sound source direction is calculated according to the sound propagation speed (about 340 m / s) and the microphone positions. The sound source enhancement algorithm is used to further improve the positioning accuracy. A beamformer is constructed to enhance the signal of the sound source in a specific direction and suppress interference from other directions. The positioning result is associated with the seat number. For example, when the sound source position coordinates are located at (250, 180) cm, the seat distribution table is searched to determine the seat number corresponding to row 3 and column 5. The associated audio data is pre - processed, including denoising, reverberation reduction, etc. Short - time energy calculation is used. The audio signal is framed (frame length 25 ms, frame shift 10 ms), the average energy of each frame is calculated, and the volume feature is obtained after normalization, with a value range of 0 - 1. For example, the volume feature value for a normal answer is about 0.4 - 0.7, and for a quiet conversation is about 0.1 - 0.3. The double - threshold speech segment detection algorithm is used to identify valid speech segments. A high threshold (such as an energy value of 0.4) and a low threshold (such as an energy value of 0.2) are set. When the energy value exceeds the high threshold, it is marked as the start of speech, and when it is below the low threshold for more than 100 ms, it is marked as the end of speech. The duration feature of each speech segment is calculated. For example, a continuous answer explanation may last for 20 - 60 seconds, and a short response may only last for 2 - 5 seconds.
[0029] The present invention improves the accuracy and robustness of human pose recognition through multi - perspective skeleton point fusion and adaptive confidence weighting, and solves the problems of occlusion and perspective change. Through a multi - modal feature fusion mechanism combining action semantic features, attention transfer sequences, and intention states, a fine - grained understanding of the examinee's behavior is achieved. Through modal consistency evaluation and temporal smoothing constraints, the rationality of feature fusion and the stability of behavior recognition are ensured. At the same time, the collaborative positioning of multiple microphone arrays and double - threshold voice detection provide reliable acoustic features, which overall improves the accuracy and real - time performance of abnormal behavior detection in the examination room.
[0030] In an alternative embodiment, a hidden Markov model is constructed based on the behavior feature sequence and acoustic features. The distribution difference degree of the answering and thinking states is obtained by using kernel density estimation to calculate the dynamic adaptive weight. The state behavior consistency is evaluated by the matching degree between the smoothed posterior probability of the state and the personalized normal behavior model. Abnormal behavior recognition is performed within a sliding time window in combination with the probability distribution entropy value, including: Normalize the behavioral feature sequence and acoustic features to obtain a normalized feature sequence; construct a hidden Markov model based on the normalized feature sequence, calculate the state distribution difference degree using kernel density estimation and KL divergence, and obtain a time-varying state transition probability matrix by constructing a dynamic adaptive weight through Sigmoid mapping; in the hidden Markov model, calculate the state smoothed posterior probability using the product of the forward variable and the backward variable, where the forward variable is obtained recursively by multiplying the state emission probability by the time-varying state transition probability, and the backward variable is obtained recursively by multiplying the state emission probability at the next moment by the time-varying state transition probability; Construct a personalized normal behavior model of the examinee based on the normalized feature sequence collected within the first answering cycle after the start of the exam, calculate the matching degree between the state smoothed posterior probability and the state parameters of the personalized normal behavior model to obtain a state behavior consistency score; Calculate the probability distribution entropy value of the state smoothed posterior probability within the sliding windows of multiple time scales, adaptively correct the anomaly determination in combination with the exam time process and the surrounding seat behavior baseline, and determine it as abnormal behavior when the anomaly scores of multiple consecutive time scales exceed the adaptive anomaly threshold.
[0031] Exemplarily, normalize the behavioral feature sequence and acoustic features: collect the behavioral feature data and acoustic feature data within a period of time, and calculate the minimum and maximum values of each feature. Then, using these minimum and maximum values, convert each feature value into a value between 0 and 1, thereby obtaining a normalized feature sequence.
[0032] After obtaining the normalized feature sequence, construct a hidden Markov model based on these sequences. The hidden Markov model consists of states, observations, and transition probabilities. Define the state set and observation set of the model, initialize the states using the normalized feature sequence, calculate the observation probability distribution of each state using the kernel density estimation method, and then obtain the emission probability of the state. At the same time, use KL divergence to evaluate the distribution difference degree between different states. By calculating the state distribution difference degree, obtain dynamic adaptive weights. These weights reflect the importance of different states in the model. Specifically, the Sigmoid mapping method can be used to convert the state distribution difference degree into weight values to ensure that the weights vary between 0 and 1. According to these dynamic adaptive weights, construct a time-varying state transition probability matrix to reflect the transition relationship between states.
[0033] In a Hidden Markov Model, the product of the forward variable and the backward variable is used to calculate the smoothed posterior probability of a state. The forward variable represents the probability that the system is in a certain state given an observed sequence. The forward variable is recursively obtained by multiplying the state emission probability with the time-varying state transition probability. The backward variable represents the probability that the system is in a certain state given future observed sequences. The backward variable is recursively obtained by multiplying the state emission probability at the next time step with the time-varying state transition probability. Multiplying the forward variable and the backward variable gives the smoothed posterior probability of a state, which represents the probability that the system is in a certain state given the entire observed sequence.
[0034] To establish a personalized normal behavior model for examinees, it is necessary to collect behavioral characteristic data of multiple examinees in a similar environment to form a benchmark representing normal behavior. By analyzing these data, state parameters are extracted and the matching degree with the smoothed posterior probability of the state is calculated. The calculation of the matching degree can obtain a state behavior consistency score by comparing the similarity between the smoothed posterior probability of the state and the state parameters of the personalized normal behavior model.
[0035] Within a sliding window of multiple time scales, the probability distribution entropy value of the smoothed posterior probability of the state is calculated. The entropy value reflects the uncertainty of the probability distribution. The higher the entropy value, the greater the uncertainty of the state. Combining the exam time process and the surrounding seat behavior baseline, adaptive correction of the anomaly determination is performed. Specifically, the threshold for anomaly determination can be dynamically adjusted according to the progress of the exam and the behavior patterns of the surrounding examinees. When the anomaly scores of multiple consecutive time scales exceed the adaptive anomaly threshold, it can be determined as abnormal behavior. This process involves monitoring and analyzing the smoothed posterior probability of the state to ensure accurate identification of abnormal behavior at different time scales.
[0036] Existing rule-based anomaly detection methods use a fixed threshold for judgment and do not consider individual differences among examinees. The present invention introduces kernel density estimation and KL divergence to calculate the state distribution difference degree, constructs a dynamic adaptive weight through Sigmoid mapping to realize the dynamic adjustment of the state transition probability, and improves the adaptability of state modeling; constructs a personalized model based on the normal answering behavior sequence at the initial stage of the exam to provide a personalized benchmark for anomaly detection; uses a multi-time scale sliding window to analyze the probability distribution entropy value of the state, and performs adaptive correction in combination with the exam process and the surrounding behavior baseline, enhancing the reliability of anomaly detection.
[0037] In an alternative embodiment, constructing a Hidden Markov Model based on a normalized feature sequence, using kernel density estimation and KL divergence to calculate the state distribution difference degree, and obtaining a dynamic adaptive weight through Sigmoid mapping to construct a time-varying state transition probability matrix includes: Construct a hidden Markov model containing the answering state and the thinking state based on the normalized feature sequence, use a state indicator function to mark the states of the candidate's answering behavior sequence, construct an answering state sample set from the feature samples marked as the answering state, and construct a thinking state sample set from the feature samples marked as the thinking state; Use the kernel density estimation method to calculate the feature distribution density functions of the answering state sample set and the thinking state sample set respectively, calculate the KL divergence based on the feature distribution density functions to obtain the feature distribution difference degree, and perform normalization processing on the feature distribution difference degree to obtain the normalized difference degree; Calculate the dynamic adaptive weight based on the normalized difference degree, and the dynamic adaptive weight maps the normalized difference degree through a Sigmoid mapping function; Use the dynamic adaptive weight to construct a time-varying state transition probability matrix, where the state transition probability is obtained by weighting the basic transition probability and the conditional transition probability based on the current observation by the dynamic adaptive weight, and the time-varying state transition probability matrix is used for subsequent state sequence estimation.
[0038] Exemplarily, construct a hidden Markov model containing the answering state and the thinking state based on the normalized feature sequence. First, define a state indicator function, which judges the candidate's state based on the combination of behavioral features and acoustic features: when a continuous writing action is detected and the head is looking at the test paper area, it is judged as the answering state; when looking up, stopping writing and there are thinking expression features, it is judged as the thinking state. Traverse and mark the feature sequence through this state indicator function to obtain a sample sequence with state labels. Then, classify the feature samples marked as the answering state into the answering state sample set, and classify the feature samples marked as the thinking state into the thinking state sample set. These two sample sets are used for subsequent distribution feature calculations.
[0039] Use the kernel density estimation method to calculate the feature distribution density functions of the answering state sample set and the thinking state sample set respectively. For each feature dimension, select the Gaussian kernel function as the kernel function and perform kernel density estimation on the sample points. The kernel bandwidth parameter is determined by the cross-validation method and takes a value of 0.15 in the example. By applying the kernel function to each sample point in the sample set and accumulating, the feature distribution density function is obtained. Similarly, apply the same method to the thinking state sample set to obtain the feature distribution density function of the thinking state.
[0040] Calculate the KL divergence based on the feature distribution density function to obtain the feature distribution difference degree. Using the discretization method, divide the feature space into a finite number of grid points, calculate the density function value at each grid point, and then calculate the KL divergence. Taking the data of a certain candidate as an example, the feature dimension is 5, and the calculated KL divergence value is 3.27. Normalize the feature distribution difference degree. By analyzing historical data, determine that the theoretical maximum value of the KL divergence is 10 and the minimum value is 0, and map the KL divergence value to the interval from 0 to 1 to obtain the normalized difference degree.
[0041] Calculate the dynamic adaptive weight based on the normalized difference degree. Use the Sigmoid mapping function to map the normalized difference degree and introduce a parameter to adjust the function shape. The parameter is determined through experimental analysis, so that when the normalized difference degree is 0.5, the obtained dynamic adaptive weight is about 0.5. For the normalized difference degree of 0.327 in the example, the calculated dynamic adaptive weight is 0.281.
[0042] Construct a time-varying state transition probability matrix using the dynamic adaptive weight. The state transition probability is calculated by weighting two probabilities with the dynamic adaptive weight: the basic transition probability is obtained by statistical analysis of historical data, representing the transition probability without considering the current observation; the conditional transition probability based on the current observation is calculated by a classifier, representing the transition probability under the current observation value. By combining these two probabilities through weighting, the final time-varying state transition probability matrix is obtained, which is used for subsequent state sequence estimation to achieve accurate identification and prediction of the candidate's answering behavior state.
[0043] The present invention uses the kernel density estimation method to calculate the state distribution characteristics, overcoming the limitations of parametric models in describing complex behavior distributions. By dynamically evaluating the state differences using the KL divergence and performing normalization processing, it realizes the adaptive tracking of the changes in the candidate's personalized behavior patterns.
[0044] In an optional implementation manner, calculate the probability distribution entropy value of the state smoothed posterior probability within the sliding windows of multiple time scales, and perform adaptive correction on the anomaly determination in combination with the exam time process and the surrounding seat behavior baseline. When the anomaly scores of consecutive multiple time scales all exceed the adaptive anomaly threshold, the determined abnormal behaviors include: Calculate the probability distribution entropy value of the state smoothed posterior probability within the sliding windows of three time scales: short-term, medium-term, and long-term. The short-term window is used to capture sudden anomalies, the medium-term window is used to detect persistent anomalies, and the long-term window is used to evaluate the overall anomaly trend; Calculate the current exam time process based on the start time of the exam, divide the exam process into three stages: start, middle, and end, and set different anomaly scoring weights for the probability distribution entropy values in different stages; Obtain the smoothed posterior probability of the status of other seats within a preset range around the determined seat number, calculate the mean value of the probability distribution entropy of the local area as the behavior baseline, and use the behavior baseline to correct the status-behavior consistency score; Calculate the fused anomaly score based on the probability distribution entropy values at different time scales, the exam stage weights, and the corrected status-behavior consistency score; Calculate the adaptive anomaly threshold based on the fused anomaly score within the first answering cycle. When the fused anomaly scores for three consecutive time scales exceed the corresponding adaptive anomaly thresholds, it is determined as abnormal behavior.
[0045] Exemplarily, sliding windows for three time scales of short term, medium term, and long term are defined. The time range of the short-term window is set to the first ten minutes after the start of the exam, the medium-term window is set to the first hour after the start of the exam, and the long-term window is set to the duration of the entire exam process. The probability distribution entropy value of the smoothed posterior probability of the status within each window is obtained by analyzing the status data within each time window. Specifically, collect the status data within each time window and smooth it to eliminate the influence of noise and outliers. Within the short-term window, focus on capturing sudden anomalies. By real-time monitoring the candidate's behavior data, such as the answering speed, frequent raising of hands, etc., calculate the smoothed posterior probability of the status within this window, and calculate the probability distribution entropy value of this window based on it. If this entropy value is significantly higher than the average level of historical data, it can be determined as a potential sudden anomaly. Within the medium-term window, detect continuous abnormal behaviors. By analyzing the candidate's behavior patterns within this window, such as the change in the answering error rate, the increase in the pause time, etc., calculate the smoothed posterior probability of the status, and obtain the probability distribution entropy value of this window. If this entropy value continuously exceeds the preset reference value, it can be determined as a continuous anomaly. The long-term window is used to evaluate the overall anomaly trend. By summarizing the behavior data of the entire exam process, calculate the smoothed posterior probability of the status within the long-term window and its probability distribution entropy value. If this entropy value shows an obvious upward trend, it can be considered that there is a risk of overall abnormal behavior.
[0046] Calculate the current exam time progress based on the start time of the exam, and divide the exam process into three stages: start, middle, and end. For example, the start stage is the first fifteen minutes after the start of the exam, the end stage is the last fifteen minutes of the exam, and the middle stage is the other time. Set different anomaly scoring weights for the probability distribution entropy values in different stages. For example, in the start stage, candidates have poor adaptability and may show higher fluctuations, so the weight for this stage is set to a higher value; while in the end stage, candidates have entered a stable state, and the weight is set to a lower value.
[0047] Obtain the smoothed posterior probability of the status of other seats within a preset range around the determined seat number. By analyzing the behavior data of the surrounding seats, calculate the mean value of the probability distribution entropy of the local area as the behavior baseline. This baseline is used to correct the status-behavior consistency score. Specifically, if the behavior of a certain candidate is significantly different from that of the surrounding candidates, then its status-behavior consistency score will be adjusted to reflect the abnormal behavior of this candidate.
[0048] Calculate the fused anomaly score based on the probability distribution entropy values at different time scales, the exam stage weights, and the corrected status-behavior consistency score. The fused anomaly score comprehensively considers the performance in the short-term, medium-term, and long-term windows and can more comprehensively reflect the candidate's behavior status. Based on the fused anomaly score in the first answering cycle, calculate the adaptive anomaly threshold. If the fused anomaly scores for three consecutive time scales exceed the corresponding adaptive anomaly thresholds, it is determined as abnormal behavior.
[0049] Such as Figure 2As shown in the multi-time-scale sliding window entropy value fusion detection graph, the upper part: Entropy value changes and anomaly detection on multiple time scales: The upper part shows the entropy value changes of the smoothed posterior probabilities of states within three windows of short term (10 min), medium term (30 min), and long term (60 min). It can be observed that between the 25th minute and the 40th minute, the entropy values of all three scales increase significantly and exceed the anomaly threshold, but exhibit different response characteristics: Short-term window (solid line): Responds to anomalies the fastest and has the largest fluctuations, with the entropy value rising to the highest within the anomaly region; Medium-term window (dashed line): The response speed and fluctuation amplitude are between those of the short-term and long-term windows; Long-term window (dash-dotted line): The response is relatively gentle, and the entropy value change curve is smoother; The marked anomaly threshold line in the figure clearly demarcates the boundary between normal and abnormal behaviors. All three curves exceed this threshold within the anomaly region, indicating that the system has detected abnormal behaviors at different time scales. Lower part: Multi-scale fusion anomaly score and detection result: The lower part shows the anomaly score after fusing the entropy values of the three time scales and the final detection result. The system successfully detected the start of the abnormal behavior around the 25th minute and confirmed the end of the anomaly around the 43rd minute. The anomaly score curve is significantly lower than the adaptive anomaly threshold within the anomaly region, indicating the effective recognition of the abnormal state by the system. The black dot markers represent the detected abnormal event points, and the vertical dashed lines indicate the time points of the start and end of the anomaly. As can be seen from the figure, the fusion method effectively reduces false alarms while maintaining a fast response. Compared with a single time scale, the multi-scale fusion method has achieved higher detection rates (92.8% and 95.3% respectively) in detecting sudden and persistent anomalies, while controlling the false alarm rate at 2.4% and the detection delay within 3 minutes. By fusing the entropy value characteristics of the three time scales, the system can simultaneously meet the fast response to sudden cheating behaviors and the reliable detection of persistent anomalies, significantly reducing the false alarm rate while ensuring detection sensitivity, providing a more reliable abnormal behavior detection mechanism for invigilation in examination rooms.
[0050] In an alternative embodiment, a spatial weight coefficient is calculated based on the physical distance between the seat numbers determined to be behaviorally abnormal and the surrounding seats. The behavioral feature sequences and acoustic features of the relevant seat numbers are wavelet decomposed to obtain multi-scale coefficients. An association relationship between the seat numbers is established based on the multi-scale coefficients. By recursively calculating the behavioral-acoustic association degrees between seat numbers at different time scales and combining the spatial weight coefficient, associated abnormal behaviors are identified. Identifying multi-person collaborative cheating behaviors includes: Calculating a spatial weight coefficient based on the physical distance between the seat numbers determined to be behaviorally abnormal and the surrounding seats, where the spatial weight coefficient is obtained by mapping and normalizing the physical distance between seats through a Gaussian kernel function; Perform wavelet decomposition on the behavioral feature sequence and the acoustic feature sequence of the relevant seat numbers respectively to obtain multi-scale coefficients, where the multi-scale coefficients include the approximation coefficients and detail coefficients of the behavioral feature sequence and the approximation coefficients and detail coefficients of the acoustic feature sequence; calculate the cross-correlation coefficients between seat numbers at different time scales according to the multi-scale coefficients, and the cross-correlation coefficients are obtained by dividing the covariance of the multi-scale coefficients of the behavioral feature sequence by the product of the standard deviations of the multi-scale coefficients of the acoustic feature sequence; perform weighted fusion on the cross-correlation coefficients to obtain a fusion correlation degree, and use the fusion correlation degree as the initial correlation degree. Construct an initial association relationship between seat numbers based on the initial correlation degree, take the product of the spatial weight coefficient and the initial correlation degree as the weighted correlation degree between seat numbers, and determine the correlation degree threshold based on the fluctuation range of the weighted correlation degree within the first answering cycle after the start of the exam. When the weighted correlation degree exceeds the correlation degree threshold and the duration exceeds the preset time window, it is determined that there is a preliminary collaborative cheating behavior between the relevant seat numbers.
[0051] Exemplarily, obtain the physical position coordinates of all seat numbers according to the seat layout information in the examination room. When the system detects that a certain seat number exhibits abnormal behavior, calculate the physical distance between this seat and the surrounding seats. Taking a certain examination room as an example, it is detected that seat number 5 exhibits abnormal behavior, and calculate the physical distances between seat number 5 and the surrounding seats: d(5, 1) = 3.2 meters, d(5, 2) = 2.8 meters, d(5, 6) = 1.5 meters, d(5, 9) = 2.1 meters.
[0052] Map the physical distance to a spatial weight coefficient through a Gaussian kernel function: for the physical distance d(i, j) between seat i and seat j, the spatial weight coefficient w(i, j) is calculated through the following steps: first calculate the Gaussian kernel mapping value. For the above example, select the bandwidth parameter σ = 2.0, and obtain the mapping results: w'(5, 1) = 0.60, w'(5, 2) = 0.67, w'(5, 6) = 0.89, w'(5, 9) = 0.78. Then perform normalization processing to obtain the final spatial weight coefficient: w(5, 1) = 0.20, w(5, 2) = 0.23, w(5, 6) = 0.30, w(5, 9) = 0.27.
[0053] Preprocess the behavioral features and acoustic features collected for each relevant seat. The behavioral features mainly include the student's head movement frequency, line-of-sight deviation angle, page-turning frequency, etc., and the acoustic features include sound energy, spectral features, etc. For example, in a certain exam, sample the behavioral feature sequences X(5), X(6) and acoustic feature sequences Y(5), Y(6) of seat number 5 and seat number 6, with a sampling rate of 2 Hz and an observation window of 120 seconds, to obtain feature sequences with a length of 240.
[0054] Perform wavelet decomposition on the obtained feature sequences. Use Daubechies wavelet (db4) to perform 3-layer decomposition on the behavior feature sequence and the acoustic feature sequence respectively to obtain multi-scale coefficients. Taking seat No. 5 as an example, the behavior feature sequence is decomposed to obtain the approximation coefficient A3(X5) and the detail coefficients D1(X5), D2(X5), D3(X5); the acoustic feature sequence is decomposed to obtain the approximation coefficient A3(Y5) and the detail coefficients D1(Y5), D2(Y5), D3(Y5).
[0055] Calculate the cross-correlation coefficients between seat numbers at different time scales based on the multi-scale coefficients. For seat 5 and seat 6, calculate the cross-correlation coefficients of the behavior features and the acoustic features at each scale respectively: First, calculate the covariance between the coefficients, and then divide by the product of the corresponding standard deviations. For example, at the third-layer decomposition scale, calculate the cross-correlation coefficient ρA3(5, 6) = 0.76 between the behavior feature approximation coefficient A3(X5) and the acoustic feature approximation coefficient A3(Y6); similarly, obtain the cross-correlation coefficients between the detail coefficients: ρD3(5, 6) = 0.68, ρD2(5, 6) = 0.45, ρD1(5, 6) = 0.32.
[0056] Perform weighted fusion on the cross-correlation coefficients. The cross-correlation coefficients at different decomposition levels have different importance, and different weights are assigned: wA3 = 0.4, wD3 = 0.3, wD2 = 0.2, wD1 = 0.1. Calculate the fusion correlation degree R(5, 6) = 0.4×0.76 + 0.3×0.68 + 0.2×0.45 + 0.1×0.32 = 0.63. Take this fusion correlation degree as the initial correlation degree between seat 5 and seat 6.
[0057] Construct the initial association relationship between seat numbers. For all relevant seat pairs, calculate the initial correlation degree and establish an association matrix. For example, the initial correlation degrees between seat 5 and its surrounding seats 1, 2, 6, 9 are: R(5, 1) = 0.25, R(5, 2) = 0.31, R(5, 6) = 0.63, R(5, 9) = 0.29.
[0058] Multiply the spatial weight coefficients by the initial correlation degrees to obtain the weighted correlation degrees. Calculate the weighted correlation degrees for the aforementioned example: R'(5, 1) = 0.20×0.25 = 0.05, R'(5, 2) = 0.23×0.31 = 0.07, R'(5, 6) = 0.30×0.63 = 0.19, R'(5, 9) = 0.27×0.29 = 0.08.
[0059] Determine the correlation threshold based on the weighted correlation fluctuation within the first answering period (usually the first 30 minutes) after the start of the exam. Analyzing historical data shows that, under normal circumstances, the standard deviation of the weighted correlation is approximately 0.04, and 3 times the standard deviation is taken as the upper limit of the fluctuation range. The calculated correlation threshold τ = μ + 3σ, where μ is the mean of the weighted correlation within the first answering period. For example, for seats 5 and 6, the mean of the weighted correlation within the first answering period is 0.07, and the standard deviation is 0.04, then the set correlation threshold τ(5, 6) = 0.07 + 3×0.04 = 0.19.
[0060] Continuously monitor the weighted correlation of each pair of seats: When it is detected that the weighted correlation exceeds the threshold and the duration exceeds the preset time window (for example, continuously for 3 minutes), the system determines that there is a preliminary collaborative cheating behavior between the relevant seats. For example, starting from the 45th minute of a certain exam, it is detected that the weighted correlation between seats 5 and 6 continuously remains above 0.22, exceeding the preset threshold of 0.19, and lasts for 4 minutes and 30 seconds, exceeding the preset time window of 3 minutes. The system determines that there is a preliminary collaborative cheating behavior between seats 5 and 6.
[0061] In an alternative implementation, further identifying the preliminary collaborative cheating behavior includes: Construct a time-varying graph structure with seat numbers as nodes, use the behavior feature sequence and acoustic features corresponding to the seat numbers as node features, and obtain the edge weights based on the adaptive weighting of the physical distance, behavior similarity, and acoustic correlation between seat numbers; Establish a state vector for each seat number that includes behavior state parameters and acoustic state parameters, and construct a state transition equation based on the state vector and the edge weights. The state transition equation includes an individual behavior evolution term and an interaction influence term. The individual behavior evolution term is used to calculate the state transition probability distribution, and the interaction influence term obtains the state influence by calculating the state vector difference degree and performing a non-linear mapping with the edge weights. Based on the state influence, obtain the group synchronization index; Calculate the correlation of the behavior switching sequence between seat numbers as the time-series collaboration feature, calculate the consistency of the state transition probability distribution as the state collaboration feature, calculate the neighborhood density of similar behaviors as the group aggregation feature, and calculate the behavior trigger relationship as the causal collaboration feature; Fuse the time-series collaboration feature, state collaboration feature, group aggregation feature, and causal collaboration feature at multiple levels to obtain the collaborative behavior feature, and determine the collaborative cheating behavior based on the collaborative behavior feature and the group synchronization index.
[0062] Exemplarily, a time-varying graph structure with seat numbers as nodes is constructed: Each seat number in the examination room is used as a node in the graph, and each node has two types of features: a behavioral feature sequence and an acoustic feature. The edge weights of the graph are determined by an adaptive weighting mechanism, taking into account three factors: the physical distance between seat numbers, behavioral similarity, and acoustic correlation. The physical distance calculates the Euclidean distance based on the seat arrangement. The behavioral similarity is calculated by comparing the behavioral sequences of two seat numbers, specifically implemented as the proportion of the frequency of the same behavior occurring within the same time window. For example, within 10 minutes, seat A and seat B have behavioral sequences [lower head, lower head, raise head, turn sideways, lower head] and [lower head, raise head, raise head, turn sideways, lower head] respectively, then the similarity is 3 / 5 = 0.6. The acoustic correlation is obtained by calculating the cross-correlation coefficient of the sound signals at two seat positions, with a value range of [-1, 1], where 1 represents complete positive correlation, -1 represents complete negative correlation, and 0 represents no correlation. The final edge weight is the weighted sum of the three: W = α × (1 / physical distance) + β × behavioral similarity + γ × acoustic correlation, where α, β, and γ are adaptive weights that are dynamically adjusted according to the examination stage. For example, at the initial stage of the exam, α = 0.5, β = 0.3, γ = 0.2; in the middle stage of the exam, α = 0.3, β = 0.4, γ = 0.3; at the later stage of the exam, α = 0.2, β = 0.5, γ = 0.3.
[0063] A state vector is established for each seat number and a state transition equation is constructed. The state vector includes behavioral state parameters and acoustic state parameters. The behavioral state parameters include the current behavior type, duration, frequency, etc., and the acoustic state parameters include the volume size, voice activity detection results, etc. The state transition equation consists of an individual behavior evolution term and an interaction influence term. The individual behavior evolution term analyzes the historical behavior sequence to calculate the state transition probability distribution. For example, by statistically calculating the probability of a specific student from "lowering the head" to "raising the head", from "raising the head" to "looking around", etc., a state transition matrix is formed. Taking a certain seat number as an example, according to historical data, the probability of the student at this seat number from "lowering the head to write" to "raising the head to look at the blackboard" is 0.7, the probability of turning sideways is 0.2, and the probability of "lowering the head to flip through the book" is 0.1. The interaction influence term obtains the state influence by calculating the difference degree between the state vectors of two seat numbers and performing a non-linear mapping with the edge weight. The difference degree is calculated using the cosine distance, and the difference degree between two state vectors S1 and S2 is 1 - cos(S1, S2). The non-linear mapping uses the sigmoid function to convert the difference degree and the edge weight into an influence value. Based on this, a group synchronization index is calculated, that is, the average similarity of the state vectors of all seat numbers within a specific time window. Taking a 30-minute window as an example, if the calculated group synchronization index is 0.75, it indicates a relatively high behavioral synchronization.
[0064] Extract multi-dimensional collaborative features: The temporal collaborative feature is obtained by calculating the correlation degree of the behavior switching sequences between seat numbers. For two seat numbers, their behavior switching sequences are extracted respectively (such as "lower the head → raise the head" is regarded as one switch), and then the correlation coefficient of the sequences is calculated. The state collaborative feature is obtained by calculating the consistency of the state transition probability distribution, and the reciprocal of the KL divergence is used as the consistency measure. The group aggregation feature is obtained by calculating the neighborhood density of similar behaviors. Taking seat number i as the center, count the number of neighbors whose behavior similarity exceeds a threshold (such as 0.6), and then divide it by the total number of neighbors to get the density value. The causal collaborative feature is obtained by calculating the behavior triggering relationship, that is, count the probability that seat number j's behavior follows the change within a specific time window (such as 5 seconds) after the behavior change of seat number i. Taking a specific examination room as an example, the correlation coefficient of the behavior switching sequences between seat (2, 3) and seat (2, 4) is 0.82, the reciprocal of the KL divergence of the state transition probability distribution is 0.75, the neighborhood density of seat (2, 3) is 0.65, and the probability that the behavior change of seat (2, 3) triggers the behavior change of seat (2, 4) is 0.58.
[0065] Perform multi-level fusion on the multi-dimensional collaborative features to obtain collaborative behavior features: First, perform feature normalization, and then adopt a hierarchical fusion strategy: In the first layer, fuse the temporal collaborative feature and the state collaborative feature to obtain the behavior collaboration degree; in the second layer, fuse the group aggregation feature and the causal collaborative feature to obtain the spatial association degree; in the third layer, fuse the behavior collaboration degree and the spatial association degree to obtain the final collaborative behavior feature. For example, the behavior collaboration degree calculated for a pair of seats is 0.78, the spatial association degree is 0.62, and the final collaborative behavior feature is 0.71.
[0066] Determine collaborative cheating behaviors based on the collaborative behavior features and the group synchronization index: Set the threshold of the collaborative behavior feature to 0.65 and the threshold of the group synchronization index to 0.7. When the collaborative behavior feature of a pair of seats exceeds 0.65 and the group synchronization index exceeds 0.7, it is determined that there is a collaborative cheating behavior.
[0067] Existing collaborative cheating detections have problems such as insufficient detection accuracy, poor adaptability, insufficient analysis depth, and high false alarm rate, which are mainly manifested as follows: Single features are difficult to accurately identify collaborative behaviors, fixed rules cannot cope with diverse cheating methods, there is a lack of modeling of group behavior dynamics, and spatio-temporal correlation information is not fully utilized.
[0068] Such as Figure 3As shown in the figure of the change of group synchronization index and the analysis of multi-dimensional collaborative features, the upper part shows the "change of group synchronization index". The horizontal axis represents time (in minutes), and the vertical axis represents the synchronization index value (0 - 1.0). There are two curves in the figure: the solid line represents the collaborative cheating group, and the dashed line represents the non-collaborative group. The synchronization index of the collaborative cheating group gradually increases with time, starts to increase significantly after 30 minutes, and reaches a stable high value (about 0.8 - 0.9) at around 60 minutes; while the synchronization index of the non-collaborative group always remains at a relatively low level (about 0.1 - 0.15). The "synchronization threshold" and the "collaborative behavior area" are also marked in the figure. When the synchronization index exceeds the threshold and enters the collaborative behavior area, it indicates that collaborative cheating behavior has been detected. The lower part is a "radar chart of multi-dimensional collaborative features", which shows the collaborative features in five dimensions: temporal collaboration, causal collaboration, state collaboration, group aggregation, and behavior aggregation. The solid-line polygon in the figure represents the feature distribution of the collaborative cheating group, and the dashed-line polygon represents the feature distribution of the non-collaborative group. It can be seen from the radar chart that the feature values of the collaborative cheating group in all five dimensions are significantly higher than those of the non-collaborative group, especially the differences in the dimensions of temporal collaboration and state collaboration are the most significant, which indicates that collaborative cheating behavior has obvious characteristics in terms of time synchronization and state consistency. This figure intuitively shows the time evolution characteristics and multi-dimensional collaborative features of collaborative cheating behavior, providing an important basis for visual analysis of collaborative cheating detection.
[0069] The multi-scale feature fusion framework proposed in this application: combines physical distance, behavioral features, and acoustic features for comprehensive analysis, realizes feature extraction at multiple time scales through wavelet decomposition, and uses an adaptive weight mechanism for feature fusion. This multi-dimensional and multi-scale fusion method significantly improves the feature expression ability, enabling the system to capture more subtle collaborative behavior features. This application innovatively introduces group behavior dynamics modeling: constructs a time-varying graph structure to describe the dynamic relationship among candidates, designs a state transition equation including individual evolution and interaction effects, and extracts multi-dimensional collaborative features to characterize group behavior. At the same time, a multi-level determination mechanism is established, from preliminary anomaly detection to in-depth collaborative analysis, combines temporal features and spatial features, and introduces a group synchronization index to achieve precise identification of collaborative cheating behavior.
[0070] In the second aspect of the embodiments of the present invention, provides an intelligent analysis system for abnormal behaviors of multi-source data fusion in an examination room, including: The first unit is used to collect video data and audio data in the examination room, perform multi-view skeletal point fusion and sound source localization on the detection area corresponding to each seat number, and extract behavioral feature sequences and acoustic features; A second unit, configured to build a hidden Markov model based on the behavior feature sequence and acoustic features, calculate the distribution difference degree of the answering and thinking states by using kernel density estimation to obtain a dynamic adaptive weight, evaluate the state behavior consistency through the matching degree between the state smoothed posterior probability and the personalized normal behavior model, and identify abnormal behaviors within a sliding time window in combination with the probability distribution entropy value; A third unit, configured to calculate a spatial weight coefficient based on the physical distance between the seat numbers determined to have abnormal behaviors and the surrounding seats, perform wavelet decomposition on the behavior feature sequences and acoustic features of the relevant seat numbers to obtain multi-scale coefficients, establish an association relationship between the seat numbers according to the multi-scale coefficients, recursively calculate the behavior-acoustic association degree between the seat numbers at different time scales, and identify associated abnormal behaviors and multi-person collaborative cheating behaviors in combination with the spatial weight coefficient; A fourth unit, configured to record the detected cheating behaviors and push warning information in real time.
[0071] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0072] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0073] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent analysis method for abnormal behavior based on multi-source data fusion in the examination room, characterized by: include: Collect video and audio data in the examination room, perform multi-view skeleton point fusion and sound source localization on the detection area corresponding to each seat number, and extract behavioral feature sequences and acoustic features; A hidden Markov model is constructed based on the behavioral feature sequence and acoustic features, and the distribution difference between the answering state and the thinking state is calculated by using kernel density estimation to obtain a dynamic adaptive weight. The consistency of the state behavior is evaluated by the matching degree between the state smoothed posterior probability and the personalized normal behavior model, and abnormal behavior is identified within a sliding time window in combination with the probability distribution entropy value; The spatial weight coefficient is calculated based on the physical distance between the seat number determined to be abnormal and the surrounding seats, and the behavioral feature sequence and acoustic feature of the relevant seat number are subjected to wavelet decomposition to obtain multi-scale coefficients. The correlation relationship between the seat numbers is established based on the multi-scale coefficients, and the behavioral-acoustic correlation between the seat numbers at different time scales is recursively calculated. The spatial weight coefficient is combined to identify abnormal behaviors with correlation, and to identify multi-person collaborative cheating behaviors; The detected cheating behavior will be recorded and warning information will be pushed in real time.
2. The abnormal behavior intelligent analysis method based on multi-source data fusion in the examination room according to claim 1 is characterized in that: Extracting behavioral feature sequences and acoustic features includes: Multiple cameras are used to collect video data, the confidence of human skeleton nodes is calculated to obtain node visibility weights and the optimal viewing angle is selected, the human skeleton nodes under the optimal viewing angle are subjected to time series prediction and structural constraints, and the constrained skeleton nodes are weightedly fused with the corresponding skeleton nodes under other viewing angles to obtain the fusion position; Based on the fusion position, a spatiotemporal graph structure is constructed and action features are extracted, the action features are segmented in time sequence and mapped to preset test scene action primitives to obtain action semantic features; based on the fusion position, a three-dimensional head angle and a gaze center point are calculated, an attention distribution graph is constructed and time sequence tracking is performed to obtain an attention migration sequence; the action semantic features and the attention migration sequence are input into a long short-term memory network to obtain an intention state probability distribution; Calculating the similarity between the action semantic feature and the attention transfer sequence to obtain a modal consistency score, calculating the feature fusion weight based on the modal consistency score, weightedly fusing the action semantic feature, the attention transfer sequence and the intention state probability distribution, and applying a temporal smoothing constraint to obtain a behavior feature sequence; For the collected audio data, sound source localization is performed based on cross-correlation time delay estimation and sound source enhancement algorithm, and the localization result is associated with the seat number. Short-time energy calculation and double-threshold speech segment detection are performed on the associated audio data, and volume features and duration features are extracted as acoustic features.
3. The abnormal behavior intelligent analysis method based on multi-source data fusion in the examination room according to claim 1 is characterized in that: Based on the behavioral feature sequence and acoustic features, a hidden Markov model is constructed. The distribution difference between the answering state and the thinking state is calculated using kernel density estimation to obtain a dynamic adaptive weight. The consistency of the state behavior is evaluated by the matching degree between the state smoothed posterior probability and the personalized normal behavior model. The abnormal behavior recognition is performed in the sliding time window in combination with the probability distribution entropy value, including: The behavioral feature sequence and the acoustic feature are normalized to obtain a normalized feature sequence; a hidden Markov model is constructed based on the normalized feature sequence, the state distribution difference is calculated using kernel density estimation and KL divergence, and a dynamic adaptive weight is obtained through Sigmoid mapping to construct a time-varying state transition probability matrix; in the hidden Markov model, the state smoothing posterior probability is calculated using the product of the forward variable and the backward variable, the forward variable is recursively obtained by the product of the state emission probability and the time-varying state transition probability, and the backward variable is recursively obtained by the product of the next moment state emission probability and the time-varying state transition probability; Based on the normalized feature sequence collected in the first answering cycle after the start of the exam, a personalized normal behavior model of the examinee is constructed, and the matching degree of the state smoothed posterior probability and the state parameter of the personalized normal behavior model is calculated to obtain a state behavior consistency score; The probability distribution entropy value of the smoothed posterior probability of the state is calculated in a sliding window of multiple time scales, and the abnormal judgment is adaptively corrected in combination with the test time process and the surrounding seat behavior baseline. When the abnormal scores of multiple consecutive time scales exceed the adaptive abnormal threshold, it is determined to be abnormal behavior.
4. The abnormal behavior intelligent analysis method based on multi-source data fusion in the examination room according to claim 3 is characterized in that: Based on the normalized feature sequence, a hidden Markov model is constructed. The kernel density estimation and KL divergence are used to calculate the state distribution difference. The dynamic adaptive weight is obtained through Sigmoid mapping to construct the time-varying state transition probability matrix, including: Based on the normalized feature sequence, a hidden Markov model including answering state and thinking state is constructed, and the state indicator function is used to mark the state of the examinee's answering behavior sequence, and the feature samples marked as answering state are used to construct an answering state sample set, and the feature samples marked as thinking state are used to construct a thinking state sample set; The feature distribution density functions of the answering state sample set and the thinking state sample set are calculated respectively by using a kernel density estimation method, the KL divergence is calculated based on the feature distribution density function to obtain the feature distribution difference, and the feature distribution difference is normalized to obtain the normalized difference; Calculating a dynamic adaptive weight based on the normalized difference, wherein the dynamic adaptive weight maps the normalized difference through a Sigmoid mapping function; The dynamic adaptive weights are used to construct a time-varying state transition probability matrix, wherein the state transition probability is obtained by weighted calculation of the basic transition probability and the conditional transition probability based on the current observation by the dynamic adaptive weights, and the time-varying state transition probability matrix is used for subsequent state sequence estimation.
5. The abnormal behavior intelligent analysis method based on multi-source data fusion in the examination room according to claim 3 is characterized in that: The probability distribution entropy value of the state smoothed posterior probability is calculated in the sliding window of multiple time scales, and the abnormal judgment is adaptively corrected in combination with the test time process and the surrounding seat behavior baseline. When the abnormal scores of multiple consecutive time scales exceed the adaptive abnormal threshold, it is determined as abnormal behavior including: The probability distribution entropy values of the smoothed posterior probability of the state are calculated in sliding windows of three time scales: short-term, medium-term and long-term, wherein the short-term window is used to capture sudden anomalies, the medium-term window is used to detect continuous anomalies, and the long-term window is used to evaluate the overall anomaly trend; The current exam time process is calculated based on the exam start time, and the exam process is divided into three stages: the beginning, the middle, and the end. Different abnormal scoring weights are set for the probability distribution entropy values of different stages. Obtain the state smoothed posterior probability of other seats within a preset range around the determined seat number, calculate the mean of the probability distribution entropy value of the local area as the behavior baseline, and use the behavior baseline to correct the state-behavior consistency score; The fused anomaly score is calculated based on the probability distribution entropy value of different time scales, the weight of the test stage and the corrected state-behavior consistency score; the adaptive anomaly threshold is calculated based on the fused anomaly score in the first answering cycle. When the fused anomaly scores of three consecutive time scales exceed the corresponding adaptive anomaly threshold, it is judged as abnormal behavior.
6. The abnormal behavior intelligent analysis method of examination room multi-source data fusion according to claim 1 is characterized in that: The spatial weight coefficient is calculated based on the physical distance between the seat number determined as abnormal behavior and the surrounding seats, and the behavior feature sequence and acoustic feature of the relevant seat number are subjected to wavelet decomposition to obtain multi-scale coefficients. The correlation relationship between the seat numbers is established according to the multi-scale coefficients, and the behavior-acoustic correlation between the seat numbers at different time scales is recursively calculated. The spatial weight coefficient is combined to identify abnormal behaviors with correlation, and the identification of multi-person collaborative cheating behavior includes: Calculating a spatial weight coefficient based on the physical distance between the seat number determined to be abnormal behavior and the surrounding seats, wherein the spatial weight coefficient is obtained by mapping the physical distance between seats through a Gaussian kernel function and normalizing it; Performing wavelet decomposition on the behavioral feature sequence and the acoustic feature sequence of the relevant seat numbers to obtain multi-scale coefficients, wherein the multi-scale coefficients include the approximate coefficient and detail coefficient of the behavioral feature sequence and the approximate coefficient and detail coefficient of the acoustic feature sequence; calculating the mutual correlation coefficients between the seat numbers at different time scales according to the multi-scale coefficients, wherein the mutual correlation coefficients are obtained by dividing the covariance of the multi-scale coefficients of the behavioral feature sequence and the multi-scale coefficients of the acoustic feature sequence by the product of the standard deviation; performing weighted fusion on the mutual correlation coefficients to obtain a fusion correlation degree, and using the fusion correlation degree as the initial correlation degree; An initial correlation relationship between seat numbers is constructed based on the initial correlation, and the product of the spatial weight coefficient and the initial correlation is used as the weighted correlation between seat numbers. A correlation threshold is determined based on the fluctuation range of the weighted correlation in the first answering cycle after the start of the exam. When the weighted correlation exceeds the correlation threshold and the duration exceeds a preset time window, it is determined that preliminary collaborative cheating behavior exists between the relevant seat numbers.
7. The abnormal behavior intelligent analysis method based on multi-source data fusion in the examination room according to claim 6 is characterized in that: Further identification of the preliminary coordinated cheating behavior includes: A time-varying graph structure with seat numbers as nodes is constructed, and the behavioral feature sequence and acoustic features corresponding to the seat numbers are used as node features. The edge weights are obtained by adaptive weighting based on the physical distance, behavioral similarity and acoustic correlation between seat numbers. A state vector including behavioral state parameters and acoustic state parameters is established for each seat number, and a state transfer equation is constructed based on the state vector and the edge weight, wherein the state transfer equation includes an individual behavior evolution term and an interaction influence term, wherein the individual behavior evolution term is used to calculate the state transfer probability distribution, and the interaction influence term is obtained by calculating the state vector difference and performing nonlinear mapping with the edge weight, and a group synchronization index is obtained based on the state influence; The correlation of the behavior switching sequence between seat numbers is calculated as the temporal coordination feature, the consistency of the state transition probability distribution is calculated as the state coordination feature, the density of neighborhoods with similar behaviors is calculated as the group aggregation feature, and the behavior triggering relationship is calculated as the causal coordination feature; The time series collaborative features, state collaborative features, group aggregation features and causal collaborative features are integrated at multiple levels to obtain collaborative behavior features, and collaborative cheating behavior is determined based on the collaborative behavior features and the group synchronization index.
8. An abnormal behavior intelligent analysis system based on multi-source data fusion in an examination room, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to collect video data and audio data in the examination room, perform multi-view skeleton point fusion and sound source localization on the detection area corresponding to each seat number, and extract behavioral feature sequences and acoustic features; The second unit is used to construct a hidden Markov model based on the behavioral feature sequence and acoustic features, calculate the distribution difference between the answering state and the thinking state by kernel density estimation to obtain a dynamic adaptive weight, evaluate the consistency of the state behavior by matching the state smoothed posterior probability with the personalized normal behavior model, and identify abnormal behavior within a sliding time window in combination with the probability distribution entropy value; The third unit is used to calculate the spatial weight coefficient based on the physical distance between the seat number determined to be abnormal behavior and the surrounding seats, perform wavelet decomposition on the behavioral feature sequence and acoustic features of the relevant seat number to obtain a multi-scale coefficient, establish a correlation relationship between the seat numbers according to the multi-scale coefficient, recursively calculate the behavioral-acoustic correlation between the seat numbers at different time scales, and identify the abnormal behavior with correlation in combination with the spatial weight coefficient, so as to identify the collaborative cheating behavior of multiple people; The fourth unit is used to record the detected cheating behavior and push warning information in real time.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Abnormal behavior detection system and method based on heterogeneous data fusion analysis of big data
CN118378210A
Examinee examination room abnormal behavior analysis method and system for online examination
CN119167281A
Abnormal behavior detection method and system based on cross-modal fusion
CN119169524A
Online examination behavior detection method based on edge calculation and multi-modal fusion
CN119694007A
Artificial intelligence-based behavior monitoring method, program, and device
WO2024106604A1
Cited By
On-off state detection method and system for emergency cut-off valve
CN120667575A
Laying hen feeding scheme adjusting method and system based on moulting image recognition
CN120853826A
Online performance testing method for breather valve for oil and gas storage and transportation
CN121141153A