Smart home emotion interaction method and system based on monitoring camera

By deploying surveillance cameras in living spaces, analyzing expressions and postures in video frames, using dual-stream spatiotemporal attention network weighted features to enhance, assess behavioral patterns and mood changes, setting early warning thresholds, identifying and tracking abnormal emotions or behavioral events, and optimizing monitoring focus, the problems of degradation of recognition capabilities and insufficient optimization of monitoring focus in complex environments in the existing technology are solved, and high-precision abnormal state recognition and efficient utilization of monitoring resources are achieved.

CN120236228APending Publication Date: 2025-07-01SHENZHEN LIGUAN DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510301638.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing technology has reduced its recognition capabilities in complex environments, making it difficult to dynamically optimize the monitoring focus, resulting in insufficient monitoring of key areas and waste of computing resources.

Method used

By deploying surveillance cameras in the living space, capturing video frames and generating preliminary biometric data, analyzing expressions and poses, using dual-stream spatiotemporal attention network weighted feature enhancement, evaluating behavioral patterns and mood changes, setting early warning thresholds, identifying and tracking abnormal emotions or behavioral events, and optimizing surveillance focus.

Benefits of technology

It realizes dynamic assessment of the short-term and long-term emotional states of residents, improves the accuracy of abnormal state recognition, accurately locates the abnormality area, dynamically optimizes the monitoring focus, and reduces the invalid consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236228A_ABST
    Figure CN120236228A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video recognition, in particular to a smart home emotion interaction method and system based on a monitoring camera, and the method comprises the following steps: deploying the monitoring camera in a living space, capturing continuous video frames in the living space, and generating preliminary biological feature data; and based on the preliminary biological characteristic data, analyzing expressions and postures of the residents, and generating expression and posture analysis results. According to the invention, facial expressions and posture changes of residents are captured through video streams, a multi-dimensional emotion recognition system is established, and dynamic evaluation of short-term and long-term emotional states is realized. Expression and posture sequences are processed independently, emotion changes can be recognized, and misjudgment caused by singleness of facial features is avoided. The dynamic early warning mechanism improves the recognition precision of the abnormal state through the comprehensive evaluation of the behavior pattern and the emotion trend, so that the detection of the abnormal emotion not only depends on the single-frame image information, but also carries out the judgment in combination with the time dependence characteristic of the behavior sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video recognition, and particularly to a smart home emotion interaction method and system based on a surveillance camera. Background Art

[0002] The technical field of video recognition mainly studies how to extract, analyze, and understand visual information from video data, and its core includes key technologies such as object detection, action recognition, behavior analysis, facial expression recognition, and scene understanding.

[0003] In the prior art, abnormal behavior detection mainly relies on static object detection, making it difficult to effectively track changes in behavior patterns, resulting in a decline in recognition ability in complex environments. The perspective adjustment of the monitoring system usually based on fixed strategies fails to dynamically optimize the monitoring focus in combination with the distribution characteristics of abnormal behaviors, leading to insufficient monitoring of key areas, while irrelevant areas occupy a large amount of computing resources, reducing computing efficiency and real-time response capabilities. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose a smart home emotion interaction method and system based on a surveillance camera.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions. A smart home emotion interaction method based on a surveillance camera includes the following steps:

[0006] Deploy a surveillance camera inside the living space to capture continuous video frames in the living space and generate preliminary biometric data; based on the preliminary biometric data, analyze the expressions and postures of the occupants to generate an expression and posture analysis result;

[0007] Apply a two-stream spatio-temporal attention network to the expression and posture analysis result. The two-stream spatio-temporal attention network processes the facial expression and posture change sequences respectively, highlights key expressions and actions through weighting, and generates a weighted feature enhancement result; use the weighted feature enhancement result to evaluate behavior patterns and emotional changes, and generate an emotion and behavior evaluation result;

[0008] Collect the emotion and behavior evaluation results, set a warning threshold, determine whether there are abnormal emotions or behaviors, and generate a dynamic warning decision result; according to the dynamic warning decision result, identify and track abnormal emotion or behavior events, and generate an abnormal event tracking record;

[0009] Based on the abnormal event tracking record, real-time locate the occurrence location and time of the abnormal behavior, optimize the monitoring focus, and generate an abnormal location optimization result.

[0010] Preferably, the step of obtaining the preliminary biometric data is:

[0011] Install a surveillance camera in the living space, capture continuous video frames, analyze the video quality and clarity, filter the video frames, and form a preliminary video frame dataset;

[0012] Based on the preliminary video frame dataset, obtain the localization and pose angles of facial feature points to obtain preliminary biometric data.

[0013] Preferably, the steps for obtaining the expression and pose analysis result are as follows:

[0014] Based on the preliminary biometric data, identify the facial features and body postures of the occupants through a deep learning model to obtain feature extraction data of the face and pose;

[0015] According to the feature extraction data of the face and pose, perform emotion and behavior pattern analysis to obtain the expression and pose analysis result.

[0016] Preferably, the steps for obtaining the weighted feature enhancement result are as follows:

[0017] Based on the expression and pose analysis result, deploy a two-stream spatio-temporal attention network to process the sequences of facial expressions and pose changes. Through the hierarchical feature extraction and dynamic analysis of the network structure, extract the temporal dependence characteristics of the expression sequence and the pose sequence to obtain feature sequence data;

[0018] According to the feature sequence data, calculate the feature enhancement values of the expressions and actions. The calculation formula is:

[0019]

[0020] where c k is the confidence of feature k, s k is the stability score of feature k, r k is the response intensity of the feature to the dynamic baseline, Z is the normalization constant, m is the total number of features in the analysis, and E is the feature enhancement value;

[0021] Based on the feature enhancement value, perform the calibration of the actions to generate the weighted feature enhancement result.

[0022] Preferably, the steps for obtaining the emotion and behavior evaluation result are as follows:

[0023] Based on the weighted feature enhancement result, extract the sequences of expression changes and pose changes. Divide multiple behavior segments through time windows, perform pose angle calculation, expression state analysis, and motion trajectory decomposition on each behavior segment to generate behavior pattern data;

[0024] Analyze the trends of facial expressions and postures based on the behavioral pattern data, compare the current behavioral pattern with the past ones, analyze the deviation and change amplitude of behaviors, and generate behavioral pattern evaluation data;

[0025] Based on the behavioral pattern evaluation data, detect the abnormal points of emotional changes through the stability analysis of continuous behavioral sequences, classify them into emotional categories, and generate emotional behavior evaluation results.

[0026] Preferably, the steps for obtaining the dynamic warning decision result are as follows:

[0027] Based on the emotional behavior evaluation result, set the benchmark indicators for different types of emotions and behaviors, and generate abnormal judgment input data by analyzing the emotional change rate, behavior deviation degree, and frequency of continuous anomalies;

[0028] According to the abnormal judgment input data, calculate the abnormal emotion score, and the expression is:

[0029]

[0030] where G is the abnormal emotion score, B j is the current emotional behavior value of the j-th behavioral feature point, M j is the mean value of the j-th behavioral feature point, T j is the time variation factor of the j-th behavioral feature point, and r is the total number of emotional behavior feature points;

[0031] Based on the abnormal emotion score, set the warning threshold. If the abnormal emotion score exceeds the warning threshold, mark the abnormal state and classify and record the abnormal category to generate the dynamic warning decision result.

[0032] Preferably, the steps for obtaining the abnormal event tracking record are as follows:

[0033] Based on the dynamic warning decision result, extract the time series information of abnormal emotion events and abnormal behavior events, analyze the occurrence time, duration, occurrence times, and time intervals of each abnormal event to obtain the abnormal event tracking input data;

[0034] According to the abnormal event tracking input data, calculate the severity score of the abnormal event, and the expression is:

[0035]

[0036] Wherein, J is the severity score of the abnormal event, A is the emotional amplitude of the current abnormal emotional event, CB is the average normal amplitude of the current emotion category, X is the current spatial position coordinate of the abnormal behavior event, Y is the spatial position coordinate of the abnormal behavior event, Z is the time point of the abnormal behavior event, W is the time point of the abnormal behavior event, and U is the change rate of the abnormal event;

[0037] Based on the severity score of the abnormal event, determine whether the abnormal emotional event or the abnormal behavior event meets the preset threshold. If the severity score of the abnormal event exceeds the preset threshold, record the occurrence time, occurrence location and the category to which it belongs of the abnormal event, and generate an abnormal event tracking record.

[0038] Preferably, the steps for obtaining the abnormal positioning optimization result are as follows:

[0039] Based on the abnormal event tracking record, analyze the geographical location and time data where the abnormal behavior occurs, extract the coordinate information and time stamp of the abnormal event, and generate the input data for abnormal behavior positioning;

[0040] According to the input data for abnormal behavior positioning, calculate the abnormal behavior aggregation degree, and the expression is:

[0041]

[0042] Wherein, K is the abnormal behavior aggregation degree, CM is the number of abnormal behavior events within the current time window, N is the average number of abnormal behavior events, R is the number of repeated occurrences of the abnormal behavior in the same area, X is the current spatial position coordinate of the abnormal behavior event, Y is the spatial position coordinate of the abnormal behavior event, T is the time point of the current abnormal behavior, and S is the average time point of the abnormal behavior;

[0043] Based on the abnormal behavior aggregation degree, adjust the viewing range and focus of the monitoring camera to generate the abnormal positioning optimization result.

[0044] The present invention provides a home emotional interaction system, including:

[0045] A monitoring data acquisition module, which installs a monitoring camera in the living space, continuously captures video frames, and extracts the basic biometric data of the occupant from the video frames, including facial features and body postures, to obtain a basic biometric set;

[0046] An expression and posture analysis module, based on the basic biometric set, analyzes facial expressions and body postures, combines continuous expression and posture data streams, emphasizes key actions and expressions through parameter adjustment, obtains weighted features, and generates an expression and posture comprehensive analysis result;

[0047] The emotional behavior assessment module uses the comprehensive analysis results of facial expressions and postures to assess the behavior patterns and emotional changes of the occupant. By comparing parameters and analyzing behavior patterns, it determines the emotional state and predicted changes in behavior patterns, and obtains the emotional behavior analysis results;

[0048] The dynamic early warning determination module sets thresholds based on the emotional behavior analysis results, continuously monitors the behavior and emotions of the occupant, and when abnormal emotions or behaviors are detected, it evaluates and verifies according to preset standards to generate dynamic early warning decision results;

[0049] The abnormal event tracking module locates and tracks abnormal emotion or behavior events according to the dynamic early warning decision results, records the time and location of the events, adjusts the monitoring focus to optimize the monitoring effect, and obtains abnormal event records.

[0050] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0051] In the present invention, by capturing the facial expressions and posture changes of the occupant through the video stream, a multi-dimensional emotion recognition system is established to achieve dynamic assessment of short-term and long-term emotional states. Independently processing the expression and posture sequences can identify emotional changes and avoid misjudgment caused by the simplification of facial features. The dynamic early warning mechanism improves the recognition accuracy of abnormal states through the comprehensive assessment of behavior patterns and emotional trends, enabling the detection of emotional abnormalities not only to rely on single-frame image information but also to combine the time-dependent characteristics of the behavior sequence for judgment. The tracking of abnormal emotions and behaviors combines spatial and temporal information to accurately locate the abnormal occurrence area, and adjusts the monitoring focus based on the aggregation degree of multiple detections, enabling the camera to dynamically optimize the viewing angle and reducing the ineffective consumption of computing resources. The personalized emotion pattern analysis adaptively adjusts the individual emotional fluctuations in combination with historical data, enabling the smart home interaction to adjust the response strategy according to individual habits and improving the intelligent level of human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] Please refer to Figure 1 , the present invention provides a technical solution, a smart home emotional interaction method based on a monitoring camera, including the following steps:

[0055] Deploy surveillance cameras inside the living space to capture continuous video frames of the living space and generate preliminary biometric data; based on the preliminary biometric data, analyze the expressions and postures of the occupants to generate expression and posture analysis results.

[0056] Apply a two-stream spatio-temporal attention network to the expression and posture analysis results. The two-stream spatio-temporal attention network processes the sequences of facial expressions and posture changes separately, highlights key expressions and actions through weighting, and generates weighted feature enhancement results; use the weighted feature enhancement results to evaluate behavior patterns and emotional changes and generate emotion and behavior evaluation results.

[0057] Collect emotion and behavior evaluation results, set a warning threshold, determine whether there are abnormal emotions or behaviors, and generate dynamic warning decision results; based on the dynamic warning decision results, identify and track abnormal emotion or behavior events and generate abnormal event tracking records.

[0058] Based on the abnormal event tracking records, real-time locate the occurrence location and time of abnormal behaviors, optimize the monitoring focus, and generate abnormal location optimization results.

[0059] The steps for obtaining preliminary biometric data are as follows:

[0060] Install surveillance cameras in the living space, capture continuous video frames, analyze the video quality and clarity, screen the video frames, and form a preliminary video frame dataset;

[0061] Based on the preliminary video frame dataset, obtain the positioning and posture angles of facial feature points to obtain preliminary biometric data.

[0062] Specifically, based on the existing surveillance photography devices and continuous frame recording information in the room, first place the cameras in the living area and start the capture function, then extract image frames at a fixed interval frequency and detect the pixel resolution of each frame. Frames with a resolution lower than 1280×720 are marked as having substandard picture quality and are excluded. This resolution standard is obtained by statistically analyzing the resolution distribution of 200 typical sample images of the indoor environment and combining with actual observations. Immediately afterwards, calculate the grayscale mean value for the brightness of the frames. If the grayscale mean value is lower than 50 or higher than 220, it is determined that the lighting is distorted. This interval range is obtained based on the cross interval of the grayscale means of 100 night images and 100 day images and is appropriately widened. For cases where there are large areas of noise or picture defects in the frames, calculate the pixel difference between adjacent frames to determine the degree of abnormality. If the difference exceeds 1000 pixels and occurs continuously three times, it is regarded as abnormal. This numerical threshold is determined based on the statistical distribution of noise in the sampled frames. After the above screening is completed, all qualified frames are re-integrated to form a preliminary video frame dataset.

[0063] Based on the previously obtained preliminary video frame dataset, select the frames containing the portrait area and perform frame-by-frame analysis using the trained facial recognition and key point regression model. This model was trained after collecting 10,000 sample images marked with face bounding boxes and the positions of feature points. During training, a batch size of 16 was used and the initial learning rate was set to 0.001. During a total of 50 rounds of iteration, the classification error of the presence of a face was calculated using the cross-entropy loss function, and the regression deviation of the key point coordinates was evaluated using the mean square error. When the training is completed, a confidence threshold of 0.75 is set in the application phase. This value is the average of the predicted confidence distribution of 500 validation samples, approximately 0.55, plus a fixed value of 0.20. Detection results below this threshold are discarded. For the remaining face regions, the coordinate positions of several feature points such as the left and right eye corners and the tip of the nose are further calculated, and then the facial pose angle is obtained based on the vector angles between the feature points. For example, the angle between the vectors of the left and right eyes and the tip of the nose is detected to identify whether there is a lateral deviation. If the angle exceeds 45°, it is marked as the deflected state in the statistical table. This angle upper limit is given after summarizing the face data of multiple different orientations. Finally, the positioning and pose angle information of the face feature points in all frames are obtained to get the preliminary biometric data.

[0064] The steps for obtaining the expression and pose analysis results are as follows:

[0065] Based on the preliminary biometric data, identify the facial features and body poses of the occupant through a deep learning model to obtain the feature extraction data of the face and pose;

[0066] According to the feature extraction data of the face and pose, perform emotion and behavior pattern analysis to obtain the expression and pose analysis results.

[0067] Specifically, based on the preliminary biometric data obtained previously, training samples containing the face region and the positions of the main limb joints are selected and the corresponding coordinates are marked in them. All samples are divided into an 80% training set and a 20% validation set, and training is performed in a cyclic iteration manner with a fixed step size. At each iteration, the classification error of the face and body joints is identified by calculating the cross-entropy loss, and the accuracy of the key point regression is measured by the mean square error. Then, these error values are weighted and summed and backpropagated to update the model parameters. When the recognition accuracy in the validation set exceeds 85% continuously for five times, the training ends. This value is an empirical benchmark established after statistically analyzing the recognition results of the initial 50 test samples. Then, when using the trained deep learning model for inference, each frame of the image in the preliminary biometric data is input sequentially. If the score of any key point is lower than 0.65, the frame is directly excluded and not included in the subsequent calculation. This threshold is obtained by increasing the average score of 0.55 for the face and limb key points in 20 high-noise scenarios by 0.10. For frames with scores higher than this threshold, key point localization is performed, and the facial contour vector and limb skeleton coordinates are extracted. These vectors are combined to form a feature sequence of a fixed length, and the feature extraction data of the face and posture are obtained.

[0068] According to the facial and posture feature extraction data obtained previously, various expressions and common behavior postures are marked in the training samples respectively, and a corresponding multi-classification table is established. Each key point vector is compared with the feature items in this multi-classification table. If the difference between the feature value and any category is within 0.15, it is temporarily determined to belong to that category. This difference limit is obtained by statistically analyzing the feature vector distributions of 200 expression and action instances, selecting the average deviation of 0.10 and adding 0.05. Subsequently, all feature records that meet the determination conditions are traversed and calculated. If there are multiple similar features with a proportion exceeding 60% within the same time window, they are merged and classified into the corresponding category. This ratio is set based on the proportion of repeated postures shown by most people in 50 posture recognition experiments. Finally, the above classification information is integrated and associated with the emotion marking field and the action marking field to obtain the expression and posture analysis result.

[0069] The steps for obtaining the weighted feature enhancement result are as follows:

[0070] Based on the expression and posture analysis result, a two-stream spatio-temporal attention network is deployed to process the sequences of facial expressions and posture changes. Through the hierarchical feature extraction and dynamic analysis of the network structure, the time-dependent characteristics of the expression sequence and the posture sequence are extracted to obtain the feature sequence data;

[0071] According to the feature sequence data, the feature enhancement values of the expression and the action are calculated, and the calculation formula is:

[0072]

[0073] where, ck is the confidence of feature k, s k is the stability score of feature k, r k is the response intensity of the feature to the dynamic baseline, Z is the normalization constant, m is the total number of features in the analysis, and E is the feature enhancement value;

[0074] Based on the feature enhancement value, the action is calibrated to generate a weighted feature enhancement result.

[0075] Specifically, based on the previously obtained facial expression and gesture analysis results, the time series data required to process the facial expression and gesture change sequence is selected and the input of the two-stream spatio-temporal attention network is prepared. Two parallel input paths are set inside the network to receive the facial expression sequence and the body gesture sequence respectively. Each path reads the feature vectors frame by frame in chronological order, and combines the hierarchical extraction structure in the network to parse the key points contained in each frame layer by layer. During the parsing process, the position coordinate analysis and adjacent frame correlation analysis are performed on each frame feature record to judge the continuous change of the feature over time. When the difference between adjacent frames is greater than the pre-set threshold, this segment of the sequence is marked as a significant fluctuation segment. The threshold is determined by statistically averaging the change amount between adjacent frames after collecting 100 daily activity sequences and then adding a safety redundancy. Subsequently, the marked fluctuation segment is compared and calculated with its adjacent segments before and after. If the change amount of the fluctuation segment is more than 20% higher than the average value of the adjacent segments before and after, it is merged into this segment and included in the subsequent analysis. The 20% ratio is obtained by extracting the maximum fluctuation range of the interaction actions in ten indoor activity scenarios and referring to its average value. After comprehensively evaluating the temporal correlation of each segment through the above hierarchical dynamic analysis, it is recorded on a time axis to complete the temporal feature extraction of the two parallel paths. Finally, the time-dependent characteristics of the facial expression sequence and the gesture sequence are synchronously mapped and uniformly output to obtain the feature sequence data.

[0076] The benefit of the formula is that by comprehensively considering these three types of parameters: confidence, stability score, and dynamic baseline response intensity, the facial expression and action features in different time periods are weighted and superimposed, so that the reliability and fluctuation range of the features are numerically taken into account, and the importance of different key points can be finely reflected when analyzing the multi-dimensional changes of facial expressions or gestures;

[0077] c kThe acquisition steps are as follows: divide the previously obtained feature sequence data frame by frame, count the recognition accuracy of each key point in the same time slice, calculate the ratio of the correct matching number to the total number of frames of the key point in 50 consecutive frame recognition records, and obtain the average matching rate of the key point in this time slice. After statistically analyzing samples in 100 different scenarios and removing outliers from the scenario distribution, the average value of the relatively concentrated part is defined as the confidence level of the key point in the current frame. For example, in one scenario test, if 40 frames determine that the key point is successfully matched and 10 frames are mismatched, then c k Take 40 / 50 = 0.80;

[0078] s k The acquisition steps are as follows: calculate the variance of the coordinate changes between adjacent frames before and after for the key point in consecutive frames within the same time slice. If the variance value is small, it indicates that the key point is relatively stable. First, sort the variance sequences statistically obtained frame by frame for 30 different action scenarios, and select the median as the reference variance for this action scenario. Then, use the ratio of the variance of the specific key point to the reference variance as the basic value of the stability score, and store the stability score after normalizing it to the range of 0 to 1. For example, in a limb stretching action, the variance of this key point is 12.5 and the reference variance is 25. After normalization, s k Take 12.5 / 25 = 0.50;

[0079] r k The acquisition steps are as follows: after confirming that the key point matching is reliable and the stability is at a certain level, subtract the current position of the key point from the previously defined dynamic baseline position. The dynamic baseline position is calculated from the average coordinates of multiple repeated action sequences. Take the absolute value of the coordinate deviation value of this key point and denote it as r k to express the response intensity of the action. For example, when the dynamic baseline coordinates of a certain body joint point are (100, 200) and the current coordinates are (110, 215), the deviation values in the horizontal and vertical directions are 10 and 15 respectively. Through the Euclidean distance, r k = 18.03;

[0080] The acquisition steps of Z are as follows: the normalization factor when normalizing the sum after calculating the confidence levels of all feature points. First, calculate and accumulate the of each key point to obtain the temporary value T sum Compare the range of T sum with the maximum and minimum values of 100 groups of pre-statistical feature data. If T sum is between the two, then set Z to T sum If T sum is lower than the minimum value, then take Z as the minimum value. If T sum is higher than the maximum value, then take Z as the maximum value. For example, in a certain round The sum of the calculation results is 37.2, and this value is within the historical statistical range of 30.0 to 45.0. Therefore, Z = 37.2 at this time;

[0081] m represents the total number of features in the analysis. This value can be determined according to the number of key points to be tracked in different scenarios. When obtaining facial and body postures, usually 20 to 40 key points are tracked simultaneously. When actually obtaining, all key points are de-duplicated and the total quantity is recorded to obtain m. For example, in an ordinary living scenario, 25 valid key points may be recorded, and at this time m = 25;

[0082] Calculation process:

[0083] In the first step, first calculate part, perform corresponding operations on each key point from k = 1 to k = m and then accumulate. In the second step, substitute the response intensity |r k | into ln(1 + |r k |) to calculate the logarithm value, multiply it by the result of the first step and then sum. In the third step, divide the final sum result by Z to obtain the value of E;

[0084] The following is an example. Let m = 3, and the parameters of the three key points are as follows. One is c1 = 0.90, s1

[0085] = 0.30, r1 = 5.0. The second is c2 = 0.70, s2 = 0.60, r2 = 2.0. The third is c3 = 0.85, s3 = 0.45, r3 = 4.0. After statistics, Z = 2.50, then:

[0086]

[0087] Perform product summation on the above calculation results:

[0088] (0.52 × 1.79) + (0.45 × 1.10) + (0.50 × 1.61) = 0.93 + 0.50 + 0.81 = 2.24

[0089] Finally, divide by Z = 2.50:

[0090]

[0091] This result indicates that the feature enhancement value of the expression and action at this time is 0.896. When this value is close to 1, it means that the confidence and stability of the feature points are relatively high and the response intensity is relatively large. When this value is significantly lower than 0.50, it indicates that the confidence or stability of the feature points is insufficient or the response intensity is relatively low. This value can be used to judge the weight distribution of the action or expression feature in the scenario in subsequent action analysis.

[0092] Based on the feature enhancement values obtained previously, in the action calibration step, first conduct a horizontal comparison of the enhancement values corresponding to each key point. Consider the key points with enhancement values higher than 0.70 as action dominant points and assign an identification code. This value is obtained by adding a certain redundancy after the mean of the enhancement value distribution is 0.60 from 30 different action sequences. The remaining key points with enhancement values lower than 0.70 are classified as auxiliary points or background points. When the enhancement values of some key points are between 0.50 and 0.70, perform further analysis to determine whether they are between the dominant and auxiliary roles. If their average enhancement value remains around 0.60 in ten consecutive frames, directly determine them as auxiliary points and record them. Otherwise, label them as transition points to be evaluated and continuously update the enhancement values in subsequent multi-frame data. Configure corresponding action identification information for all key points in the above way. Finally, integrate the action identification information and associate it with the previously defined scene type and time period to generate a weighted feature enhancement result.

[0093] The steps for obtaining the emotional behavior assessment result are as follows:

[0094] Based on the weighted feature enhancement result, extract the expression change sequence and posture change sequence. Divide multiple behavior segments through a time window, calculate the posture angles, analyze the expression states, and decompose the motion trajectories for each behavior segment to generate behavior pattern data;

[0095] According to the behavior pattern data, analyze the trends of expression and posture changes, compare the current behavior pattern with past behavior patterns, analyze the behavior deviation degree and change amplitude, and generate behavior pattern assessment data;

[0096] Based on the behavior pattern assessment data, detect abnormal points of emotional changes through the stability analysis of continuous behavior sequences, classify them into emotional categories, and generate the emotional behavior assessment result.

[0097] Specifically, based on the obtained weighted feature enhancement results, all temporal information is divided into several time windows to obtain continuous facial expression change sequences and posture change sequences. When dividing the time windows, according to the duration statistics of the majority of actions in the past 50 behavioral observations, it is determined that a single window is set to 30 seconds. This duration is obtained by extracting the median from the action persistence distribution in different scenarios and then adding an additional 5-second buffer. Subsequently, the key facial expression points and posture joint points of all frames are sequentially extracted from each time window. The three-dimensional vector angle operation is performed on the extracted posture joint points, and the change curve within the range of 0 degrees to 180 degrees is marked. If there are outliers exceeding this range, they are recorded in the corresponding list for subsequent comparison. For the facial expression change sequence, the feature point vectors corresponding to the facial micro-expressions are smoothed in the time series, and the amplitude of the change between adjacent frames is statistically calculated. When the amplitude exceeds the preset threshold three times in a row, the segment is marked as a high-fluctuation section. This threshold is formed by taking the average value after variance measurement of 30 facial expression data in ordinary scenarios and then adding a safety factor of 0.20. Then, the motion trajectories of each segment are decomposed frame by frame. During the decomposition process, the amplitude and direction of the displacement vector are statistically calculated. Sections with a displacement amplitude greater than 1.3 times the average displacement amplitude are labeled separately for observation. This average amplitude value is obtained after analyzing the daily activity data taken within a week. Finally, the posture angle data, facial expression change distribution, and motion trajectory information of each segment are comprehensively recorded in a behavior pattern list to generate behavior pattern data.

[0098] According to the behavior pattern data obtained previously, first, the facial expression and posture information in the same scenario are sorted on the time axis, and then the main statistical items of the current behavior pattern and the historical behavior pattern are compared longitudinally. When the angle difference exceeds 15 degrees or the facial expression feature vector difference exceeds 25%, a deviation symbol is marked in the comparison item. The upper limit of 15 degrees is the safe range obtained by statistically calculating the joint angles of past regular actions and excluding extreme values. The threshold of 25% is obtained by distinguishing common facial expressions from rare facial expressions through mean deviation analysis of the facial muscle group activity data. Subsequently, calculations are performed on all marked comparison items. If the proportion of the number of deviation items within the same time period exceeds 30%, it is determined that the change amplitude is relatively large. This 30% ratio is obtained through 10 behavioral control tests and repeatedly confirmed in various indoor activity scenarios. Then, according to the number and specific distribution of the deviation items, they are summarized into a behavior deviation degree value. If the deviation degree is between 0.20 and 0.50, it is regarded as a moderate deviation. If it is higher than 0.50, it is regarded as a high deviation. These ranges are set after collecting 100 daily activity data and statistically analyzing their deviation degree distributions. The final deviation degree value and the action fluctuation amplitude are summarized to obtain a change amplitude index. The two values are combined to form a set of comprehensive evaluation outputs to generate behavior pattern evaluation data.

[0099] Based on the behavior pattern evaluation data obtained previously, first perform one-dimensional time series superposition on the records of behavior deviation and change amplitude in each time period, calculate the difference sequence of adjacent time periods and observe its continuity over the entire time span. When the differences in more than two adjacent time periods are all higher than the pre-set 0.15, mark them as potential abnormal points in the sequence. The value of 0.15 is obtained by adding 0.05 to the statistical mean of the common emotional fluctuation levels in the multi-scenario test samples within a week. For the difference intervals higher than 0.20, they are classified as obvious abnormal points and are preferentially included in the subsequent classification process. When performing classification, compare with existing multiple emotion classification reference tables, and combine the emotion behavior data to mark the abnormal points into categories such as joy, anger, sorrow, and happiness. If there are mixed emotion behaviors, encode them as compound emotion states. After completing the stability analysis of all continuous behavior sequences, conduct a final check between the classified emotion categories and the deviation distribution. If a certain emotion category appears repeatedly within a single time period, perform the corresponding label integration. If it remains in this category for three consecutive time windows, mark it as a high-frequency label in the evaluation output. Finally, summarize all the classified emotion abnormal points to generate the emotion behavior evaluation result.

[0100] The steps to obtain the dynamic early warning decision result are as follows:

[0101] Based on the emotion behavior evaluation result, set the benchmark indicators for different types of emotions and behaviors, and generate abnormal determination input data by analyzing the emotion change rate, behavior deviation degree, and frequency of consecutive anomalies;

[0102] According to the abnormal determination input data, calculate the abnormal emotion score, and the expression is:

[0103]

[0104] Among them, G is the abnormal emotion score, B j is the current emotion behavior value of the jth behavior feature point, M j is the mean value of the jth behavior feature point, T j is the time variation factor of the jth behavior feature point, and r is the total number of emotion behavior feature points;

[0105] Based on the abnormal emotion score, set the early warning threshold. If the abnormal emotion score exceeds the early warning threshold, mark the abnormal state, classify and record the abnormal category, and generate the dynamic early warning decision result.

[0106] Specifically, based on the previously obtained emotional behavior assessment results, list the possible emotional categories and behavioral patterns first, and extract their variation rules in the historical data one by one, including the average expression amplitude and limb movement range under each type of emotion. Then record the information of these emotions and behavioral patterns in a reference list to form the benchmark indicators for different types of emotions and behaviors. When setting the indicators, a numerical range of 0 to 1 is selected for the expression amplitude, and the average change rate of different muscle groups is marked as the benchmark value of 0.50. This value is obtained by statistically analyzing the expression activities of samples in 30 ordinary scenarios and observing their distribution. For behavioral patterns, measure the metrics of limb movement in common activities, such as the upper and lower limits of joint offset angles and stride sizes. After comparing 100 daily activities, record the most common offset as the benchmark value. By this method, a benchmark indicator table is constructed. Subsequently, conduct a segmented analysis of the emotional change rate. For example, count the number of fluctuations in expressions and postures within each 10-second segment. If the number of fluctuations is greater than 3 times, mark it as a fast-fluctuation segment. This threshold of 3 times is obtained by comparing the frequency differences between stable emotions and fluctuating emotions in multiple scenarios. Then, perform a weighted process in combination with the degree of behavioral deviation and the frequency of continuous anomalies. If the same type of abnormal event repeats in two consecutive time periods and its deviation value exceeds the previously set 0.40, mark it as a large-deviation segment. This threshold of 0.40 is finally confirmed by observing 20 samples with intense emotional fluctuations. Through the above methods, gradually accumulate various deviation and fluctuation information and merge them to form the abnormal determination input data.

[0107] The advantage of the formula is that it considers the degree of deviation of the behavioral feature points from the mean value, and introduces a time variation factor in the denominator, so that the fluctuations over time can be appropriately amplified or reduced, thus more comprehensively incorporating the temporal changes when evaluating abnormal emotions;

[0108] B j The acquisition steps of are as follows. First, define a quantitative interval for the jth behavioral feature point. For example, decompose the expression features into sub-indicators such as eye muscle activity and mouth corner deformation, and describe them using a numerical range of 0 to 1. Then measure the average amplitude of each expression or movement in the activity scenario for several seconds, integrate at least 50 repeated observations, and record the most common activity amplitude as the current emotional behavior value of this feature point. In an example, it is monitored that the mouth corner deformation range fluctuates between 0.00 and 0.80, and most of the time it is concentrated between 0.30 and 0.50. Therefore, it is set that at this moment B j = 0.45;

[0109] M jThe acquisition steps are as follows: After statistically analyzing each behavioral feature point multiple times before, take its mean value as a reference. If there are more than 100 observation records of this feature point accumulated previously, the corresponding values of each record can be summed up and then the mean value can be calculated. During the process, extreme distorted values are excluded and most of the distribution intervals are retained. In one example, the average value measured for this feature point in the first 100 observations is 0.38, so M j = 0.38;

[0110] T j The acquisition steps are as follows: By comparing the fluctuation frequencies of this feature point at different times and recording them in a time-varying sequence, a time interval of about 30 seconds to 60 seconds can be selected to accumulate the fluctuation values. If the number of fluctuations occurring within this interval is relatively large, it is regarded as a high-variation feature point. According to the previously collected emotional behavior assessment results, sum up the variation curves of this feature point in dozens of observations to obtain a total frequency value, and then divide it by the total observation duration to normalize it to the range of 0 to 1. For example, if the average number of fluctuations of this feature point in a single interval is 4 times and the total observation duration is 300 seconds, then the fluctuation frequency value is approximately 4 / 10 = 0.40, denoted as T j = 0.40;

[0111] r represents the total number of emotional behavior feature points. By listing all possible feature points participating in the calculation during the scene observation stage, a complete list can be obtained, including key points related to facial expressions and joint points related to body postures and movements, etc. In a common residential space monitoring, 30 to 50 feature points will be recorded. If 35 feature points are selected in this monitoring, then r = 35;

[0112] Calculation process:

[0113] Let the current value of the j-th feature point be B j = 0.45, the mean value M j = 0.38, the time-varying factor T j = 0.40, substitute them into the numerator (B j - M j ) 2 = (0.45 - 0.38) 2 = 0.0049, calculate Then the single contribution of this feature point is 0.0049 / 1.6703 ≈ 0.0029. Calculate and sum up all feature points from j = 1 to j = r in this way. If there are 3 feature points in this example and the cumulative values calculated respectively are 0.0029, 0.0015, and 0.0042, and the cumulative sum is 0.0086, then

[0114] The result shows that the abnormal emotion score is approximately 0.0927 at this moment. When this value continuously rises above 0.50, it may indicate obvious emotional or behavioral abnormalities. If it remains between 0.20 and 0.30 for a long time, the subsequent change trend can be continuously tracked. If it is below 0.10, it usually indicates relatively stable or small fluctuations overall.

[0115] Based on the previously obtained abnormal emotion score, after comparing with the previously collected emotion score distribution, a warning threshold is set. If the scores of the same type of scenarios in multiple monitors mostly concentrate below 0.20, the threshold can be tentatively set at 0.25. Once a certain score exceeds this value in subsequent monitors, it is marked as an abnormal state that requires key observation. When it is recorded that the score exceeds 0.25 twice in a row and the change range exceeds 30%, the abnormal category in the classification record is updated. When it exceeds 0.40, it is recorded as a higher-level abnormality. These specific values are set after statistically analyzing the emotion scores of 50 people in different daily scenarios and performing quantile analysis. After all abnormal marks are completed, a dynamic warning decision result can be formed.

[0116] The steps to obtain the abnormal event tracking record are as follows:

[0117] Based on the dynamic warning decision result, extract the time series information of abnormal emotion events and abnormal behavior events, and analyze the occurrence time, duration, occurrence frequency, and time interval of each abnormal event to obtain the input data for abnormal event tracking.

[0118] According to the input data for abnormal event tracking, calculate the severity score of the abnormal event. The expression is:

[0119]

[0120] Among them, J is the severity score of the abnormal event, A is the emotion amplitude of the current abnormal emotion event, CB is the average normal amplitude of the current emotion category, X is the current spatial position coordinate of the abnormal behavior event, Y is the spatial position coordinate of the abnormal behavior event, Z is the time point of the abnormal behavior event, W is the time point of the abnormal behavior event, and U is the change rate of the abnormal event.

[0121] Based on the severity score of the abnormal event, determine whether the abnormal emotion event or abnormal behavior event meets the preset threshold. If the severity score of the abnormal event exceeds the preset threshold, record the occurrence time, occurrence location, and category of the abnormal event to generate an abnormal event tracking record.

[0122] Specifically, based on the previously obtained dynamic early warning decision results, first screen all the marked abnormal emotional or behavioral content and sort out their distribution on the time axis, including the time stamps corresponding to the occurrence moments of each abnormality, the specific durations, and the number of repeated occurrences within a certain observation interval. Then, measure the intervals between adjacent abnormal events. For example, in the current scenario, for each abnormal emotional event, the start second and the end second are recorded. The time interval is obtained by subtracting the end second of the previous event from the start second of the subsequent event. If the time interval is less than 30 seconds, it is marked as an adjacent state. The value of 30 seconds is obtained by adding a part of the redundancy after statistically averaging the intervals of 20 consecutive emotional fluctuations, which is about 25 seconds. Then, similarly, compare the occurrence positions of abnormal behavioral events and record them in the coordinate system. If it is found that they repeatedly fall in the same area for multiple times, continuously monitor this area and conduct coordinate aggregation analysis. By cross-comparing adjacent time periods and adjacent spatial ranges, it can be further confirmed whether there are multiple consecutive abnormalities. If the time intervals of three consecutive abnormalities are all within 30 seconds and the spatial positions are no more than 2 meters apart, these abnormal events are marked as a tightly aggregated type in the tracking input data. The 2-meter distance is the average social safety distance measured from the regular layout of the indoor main activity area and confirmed through multiple scenario tests. Finally, summarize the occurrence moments, durations, occurrence frequencies, and position coordinates of all abnormalities to obtain a tracking input data for abnormal events.

[0123] The benefit of the formula is that by comprehensively considering the difference in emotional amplitude, behavioral coordinates, and time difference, and introducing the dynamic factor of the rate of change in the denominator, it can more accurately reflect the comprehensive degree of abnormal events in terms of intensity and spatio-temporal distribution.

[0124] The steps to obtain A are as follows: For the emotional amplitude of the current abnormal emotional event, perform a quantization process. First, extract the peak value of the expression features or emotional indicators during the event period from the previously identified emotional curve. By statistically calculating the amplitude of its facial muscle activities frame by frame (such as the deformation of the corners of the mouth, the lifting of the corners of the eyes, etc.) and normalizing the results to the interval from 0 to 1, and filtering out sudden noise points, take the peak value during this period as the emotional amplitude. For example, if the peak value of the expression intensity is detected to be about 0.75 during the period from 50 seconds to 60 seconds, then A = 0.75.

[0125] The steps to obtain CB are as follows: For the average normal amplitude of this emotional category, first collect 100 pieces of daily emotional data and classify them into the same category as the current emotion, such as "nervous" or "excited", etc. Extract the peak amplitude from each corresponding record and remove the extreme values in the first 5% and the last 5%, and then take the arithmetic mean of the remaining distribution. For example, through the statistics of 80 valid records of the "nervous" emotion, the average value of this emotional category is finally calculated to be about 0.50, then CB = 0.50.

[0126] The steps for obtaining X and Y are as follows: After calibrating the spatial coordinates of the location where the abnormal behavior event occurs, record the current spatial position coordinates as X, and record the spatial position coordinates in one or more previous similar situations as Y, which is used to measure the difference between the current abnormal behavior and the previous similar behaviors at the spatial level. The coordinates are usually obtained through indoor positioning methods (such as distance measurement or visual recognition) to record the coordinates of the event occurrence. For example, in a living space, the ground is divided into a range of 5 meters for both the X-axis and the Y-axis. If the current position is measured at (2.1, 3.5) in an event, and the historical recorded position is at (2.0, 1.2), then X = 2.1 and Y = 2.0;

[0127] The steps for obtaining Z and W are as follows: After extracting the timestamp of the abnormal behavior event or abnormal emotion event, record the time point of the current event as Z, and record the time of the comparable reference event as W. For example, if the previous event occurred at the 120th second and the current event occurred at the 180th second, then Z = 180 and W = 120;

[0128] The step for obtaining U is defined as the change rate of the abnormal event, which can be calculated based on the recurrence frequency and time difference of the same type of abnormality. If the same type of abnormality appears three times within 300 seconds, it indicates a relatively high rate. First, count the number of repetitions and the interval time, and normalize (number of repetitions / total time) to between 0 and 1 through segmented accumulation, and combine the fluctuation amplitude to form a comprehensive change rate value. For example, if two similar abnormalities occur within two minutes, the rate is calculated as (2 / 120) = 0.0167. If additional quantification of the action or emotion amplitude is considered, appropriate addition or subtraction is performed based on this value. Finally, record the result, for example, U = 0.25;

[0129] Calculation process:

[0130] In the first step, first calculate the numerator (|A - CB|), take the absolute value of the difference between the extracted emotion amplitude A and the corresponding category mean CB, and then substitute information such as the spatial coordinate difference and the time point difference into it, and add it to the previous result. In the second step, obtain 1 + e -U according to the calculated U value. In the third step, divide the previous numerator result by this denominator value to obtain the event severity score J;

[0131] The following is an example: Emotion amplitude A = 0.75, normal amplitude mean CB = 0.50, current position coordinate X = 2.1, past position coordinate Y = 2.0, current time point Z = 180 seconds, past time point W = 120 seconds, change rate U = 0.25, then:

[0132] |A - CB| = |0.75 - 0.50| = 0.25

[0133]

[0134] = 60.0001 ≈ 60.00

[0135] The numerator part is:

[0136] 0.25 + 60.00 = 60.25

[0137] The denominator part is:

[0138] 1 + e -0.25 ≈ 1 + 0.7788 = 1.7788

[0139] Therefore:

[0140]

[0141] This result indicates that the severity score of the abnormal emotion event or abnormal behavior event at this moment is approximately 33.90. When this value exceeds a pre-set threshold, it means that the current space or time period needs to be focused on. If the score is lower than the threshold, it can be temporarily regarded as medium or slightly abnormal;

[0142] Based on the severity scores of the abnormal events obtained previously and combined with the risk preference settings in the scenario, the scores of multiple abnormal events are arranged in order and compared with a reference threshold line. If the score is between 0 and 10, it is mostly mild. When it exceeds 10 but is less than 30, it can be regarded as moderate. When it is higher than 30, it is seriously deviated. These thresholds are obtained by selecting the 30%, 60%, and 90% quantiles after statistics on 50 real-scenario cases and are confirmed after fine-tuning. When it is found that the score of an abnormal event is greater than 30, the timestamp of the event occurrence will be immediately extracted, and the event location will be marked in the coordinate system. Then, the corresponding behavior categories such as "wandering", "impatient", etc. will be matched and sorted. If the same category of events fall into the high-score range multiple times, a high-frequency mark will be added to the result. Finally, after integrating the time and space information, the abnormal event tracking record will be output.

[0143] The steps to obtain the optimized result of abnormal positioning are as follows:

[0144] Based on the abnormal event tracking record, analyze the geographical location and time data of the abnormal behavior occurrence, extract the coordinate information and timestamp of the abnormal event, and generate the input data for abnormal behavior positioning;

[0145] According to the input data for abnormal behavior positioning, calculate the abnormal behavior aggregation degree, and the expression is:

[0146]

[0147] Among them, K is the abnormal behavior aggregation degree, CM is the number of abnormal behavior events within the current time window, N is the average number of abnormal behavior events, R is the number of repeated occurrences of abnormal behavior in the same area, X is the current spatial position coordinate of the abnormal behavior event, Y is the spatial position coordinate of the abnormal behavior event, T is the time point of the current abnormal behavior, and S is the average time point of the abnormal behavior;

[0148] Based on the abnormal behavior aggregation degree, adjust the viewing range and focus of the monitoring camera to generate an optimized abnormal positioning result.

[0149] Specifically, based on the abnormal event tracking records obtained previously, first check each item of the spatial coordinates and occurrence time fields of all abnormal events to confirm whether the coordinate information falls within the indoor monitorable area and classify different abnormal types. During the classification process, emotional abnormalities and behavioral abnormalities are processed separately, and their distribution laws on the time axis are extracted. If some abnormalities occur frequently in a specific time period and the interval is less than 30 seconds, it is recorded as a high-frequency occurrence period. The threshold of 30 seconds is determined by taking the average value of the abnormal event occurrence intervals of multiple residents in 10 residential scenario observations and leaving a certain safety margin. Then, perform a simple spatial position comparison on the coordinate data corresponding to each abnormal record. If abnormalities repeatedly appear within a radius of 2 meters at the same position, it is marked as a potential aggregation point. The 2-meter distance is obtained by field measurement of the average distance between adjacent activity points in common residential layouts and verified in five simulated scenarios. Subsequently, arrange these aggregation points in chronological order and organize the corresponding timestamp information. If the same aggregation point shows abnormalities multiple times within a short period, an additional identifier will be given during summarization for subsequent reference. In this way, a list with coordinates, timestamps, and occurrence frequencies can be formed for subsequent calculation and call. Finally, the above coordinate information and timestamp data together constitute the input data for abnormal behavior positioning.

[0150] The advantage of the formula is that it comprehensively examines the deviation of the number of abnormal behavior events within the current time window from the average level and the correction of this deviation by the number of repeated occurrences in the same area, and incorporates the spatial and temporal coordinate differences into the product operation after quantification, which can reflect the concentration degree of abnormal behavior in numerical form during scenario monitoring;

[0151] The steps to obtain CM are as follows: first determine an observation time window, for example, select every 30 seconds as a unit and count the number of abnormal behavior events during this period, and accumulate all abnormal behavior events in this time period to obtain CM. If 5 abnormal behaviors are recorded in an observation cycle, CM can be set to 5. In order to ensure that the data truly reflects the actual situation of the current period, it is necessary to compare at least 10 or more time windows, and the number of abnormal behaviors in each window is fully counted before parallel analysis with other parameters. If the event value in a single window fluctuates significantly in multiple monitorings, it will affect the credibility of CM. Therefore, it can be combined with other parameters for cross-confirmation in subsequent calculations. In an example scenario, 10 rounds of 30-second observations are conducted on the same indoor monitoring point, and the average value of each round is 4.5. There are 5 abnormalities in one round, so CM=5;

[0152] The step of obtaining N is to summarize the number of abnormal behavior events in the same scene or the same time period in the large number of observation samples collected previously, and calculate their average number after eliminating abnormal extreme values. For example, statistics are performed in the same time period every day for a week, and it is found that there are about 2 abnormal behaviors in this period on average, then N=2. If the scene or time period changes, the mean value needs to be updated. This operation can be achieved by accumulating monitoring data in diversified scenes, and comparing the records of multiple days before and after to obtain a stable distribution. When performing the same type of monitoring again, the mean value can be used as a control value. In an example, if the number of abnormalities in 10 time periods is 1, 2, 3, 2, 2, 2, 3, 1, 1, 2, respectively, and the arithmetic mean value 2 is calculated, then N=2;

[0153] The steps to obtain R are to count the repeated anomalies that occur in the same area or the same location within a certain time range. For example, the room is divided into several coordinate grids and the coordinates of each abnormal event are marked. If multiple similar behaviors occur in a certain area, the number of repetitions is recorded in a list. For example, if three abnormal behaviors occur in the same coordinate range within one day, R = 3;

[0154] The steps of obtaining X and Y are as follows: mark the spatial coordinates of the current abnormal behavior event as X, and mark the coordinates of the reference target position as Y. The target position may be the coordinates of the previous event of the same type or a known reference position. In actual operation, an indoor ranging device can be used to calibrate the coordinates of each event occurrence point. If it is found that the current abnormality is close to the previous abnormality or the reference position, the difference between the two (XY) will be relatively small. In an example, the current event coordinate X = 3.50, the past reference coordinate Y = 2.00, then the difference is 1.50;

[0155] The steps to obtain T and S are as follows: When an event occurs, use the timer in the system to obtain the current time point T, and then select the average time point S of the corresponding type or scenario from the accumulated abnormal event information. This can be obtained by adding up the times when multiple events of the same category occurred in the past and then dividing by the number of occurrences. If a certain type of behavior is usually more common at a fixed time period within a day, the value of S will be relatively concentrated. In one example, if events of this category usually occur between the 100th second and the 150th second, the arithmetic mean can be set to 120. At this time, if the current abnormality occurs at the 180th second, then T = 180 and S = 120;

[0156] The steps to obtain ln(1 + R) are as follows: Add 1 to the aforementioned repeated occurrence times R and then take the natural logarithm. This step converts the influence of the repeated times into a relatively smooth function, avoiding extreme imbalance in the operation results when R is large. If an abnormality occurs 3 times repeatedly in a certain area in a scenario, then R = 3, and ln(1 + 3) = ln(4) ≈ 1.3863;

[0157] Calculation process:

[0158] First step, calculate the numerator First, take the absolute value of |CM - N|, then add ln(1 + R) to the denominator and add 1. Second step, calculate That is, perform the Euclidean distance operation on the differences between the current coordinate and the reference coordinate and between the current time point and the average time point. Finally, multiply this distance by the result calculated in the previous step to obtain K;

[0159] The following are example calculations:

[0160] Let CM = 5, N = 2, R = 2, X = 3.50, Y = 2.00, T = 180, S = 120. First, calculate ln(1 + R) = ln(3) = 1.0986;

[0161] 1 + ln(1 + R) = 1 + 1.0986 = 2.0986

[0162] |CM - N| = |5 - 2| = 3

[0163] The numerator part is:

[0164]

[0165] Then calculate the coordinate and time differences:

[0166] (X - Y) = 3.50 - 2.00 = 1.50, (T - S) = 180 - 120 = 60

[0167]

[0168] Finally:

[0169] K = 1.430 × 60.01 ≈ 85.515

[0170] This result indicates that the aggregation degree of abnormal behaviors is approximately 85.515 in this example. When this value continuously rises above 100, it can be regarded as a more intensive abnormal aggregation phenomenon. If the value is below 10, it means that the abnormal behaviors are relatively dispersed. In each monitoring time window, it can be distinguished according to the reference ranges of different scenarios. For example, in the conventional living scenario, both the range between 10 and 50 and the range between 50 and 100 can be respectively defined as different levels of aggregation degrees.

[0171] Based on the obtained aggregation degree of abnormal behaviors, collect the positions of all cameras that may exhibit abnormalities and their current shooting directions and focal lengths within the same period of time. After sorting out the working parameters of these cameras, compare the areas where the aggregation degree exceeds 50 and match the coverage ranges of these cameras with the coordinate ranges of this area one by one. If it is found that the horizontal viewing angle of a certain camera is not sufficient to fully cover this abnormal area, record the viewing angle adjustment requirements of this camera. For example, in a room that is 8 meters long and 6 meters wide, the default shooting angle of a certain camera only covers a range of 4 meters. If this range partially overlaps with the abnormal aggregation area and there is an uncovered area of 2 meters, then list it as a priority adjustment object. Then, first change the lens direction according to the area with the highest aggregation degree. When rotating the angle, the direction of the line connecting the current aggregation point and the camera center can be referred to and the focal length can be determined. If the area of this area is small, the focal length can be reduced to obtain a local high-definition image. If the area of this area is large, the focal length can be slightly increased. During the whole process, the rotatable range of the camera, the maximum rotation speed of the pan-tilt, etc. will be detected to ensure that all lens adjustments can be completed within 10 seconds. If there are multiple aggregation areas distributed in different directions, the cameras will be grouped and matched in turn. Finally, all adjustment information will be summarized to obtain the optimized result of abnormal positioning.

[0172] The present invention provides a home emotional interaction system, including:

[0173] A monitoring data acquisition module, which installs monitoring cameras in the living space, continuously captures video frames, and extracts the basic biometric data of the occupants from the video frames, including facial features and body postures, to obtain a basic biometric set;

[0174] An expression and posture analysis module, which analyzes facial expressions and body postures based on the basic biometric set, combines continuous expression and posture data streams, emphasizes key actions and expressions through parameter adjustment, obtains weighted processed features, and generates an integrated analysis result of expressions and postures;

[0175] The emotional behavior assessment module uses the comprehensive analysis results of expressions and postures to evaluate the behavior patterns and emotional changes of the occupants. By comparing parameters and analyzing behavior patterns, it determines the emotional state and the estimated changes in behavior patterns, and obtains the emotional behavior analysis results;

[0176] The dynamic early warning determination module sets thresholds based on the emotional behavior analysis results, continuously monitors the behavior and emotions of the occupants, and when abnormal emotions or behaviors are detected, it evaluates and verifies them according to preset criteria to generate dynamic early warning decision results;

[0177] The abnormal event tracking module locates and tracks abnormal emotion or behavior events according to the dynamic early warning decision results, records the time and location of the events, adjusts the monitoring focus to optimize the monitoring effect, and obtains the abnormal event records.

[0178] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as the technical solution content of the present invention is not departed from, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for emotional interaction of smart home based on surveillance camera, characterized in that: The following steps are involved: Deploy surveillance cameras in the living space to capture continuous video frames in the living space and generate preliminary biometric data; Analyzing the expression and posture of the occupant based on the preliminary biometric data to generate an expression and posture analysis result; Applying a dual-stream spatiotemporal attention network to the expression and posture analysis results, the dual-stream spatiotemporal attention network processes facial expressions and posture change sequences respectively, highlights key expressions and actions by weighting, and generates weighted feature enhancement results; using the weighted feature enhancement results, evaluating behavior patterns and emotional changes, and generating emotional behavior evaluation results; Gather the emotional behavior assessment results, set warning thresholds, determine whether there are abnormal emotions or behaviors, and generate dynamic warning decision results; According to the dynamic warning decision results, identify and track abnormal emotions or behavior events, and generate abnormal event tracking records; Based on the abnormal event tracking records, the location and time of the abnormal behavior are located in real time, the monitoring focus is optimized, and the abnormal location optimization result is generated.

2. The method for emotional interaction of smart home based on surveillance camera according to claim 1 is characterized in that: The steps for obtaining the preliminary biometric data are: Install surveillance cameras in the living space, capture continuous video frames, analyze video quality and clarity, filter video frames, and form a preliminary video frame data set; Based on the preliminary video frame data set, the location and posture angle of facial feature points are obtained to obtain preliminary biometric data.

3. The method for emotional interaction of smart home based on surveillance camera according to claim 1 is characterized in that: The steps for obtaining the expression posture analysis result are: Based on the preliminary biometric data, identifying the facial features and body posture of the resident through a deep learning model to obtain facial and posture feature extraction data; Data is extracted based on the facial and posture features, emotion and behavior pattern analysis is performed, and expression and posture analysis results are obtained.

4. The method for emotional interaction of smart home based on surveillance camera according to claim 1, characterized in that: The steps for obtaining the weighted feature enhancement result are: Based on the expression and posture analysis results, a dual-stream spatiotemporal attention network is deployed to process facial expression and posture change sequences, and the time-dependent characteristics of expression sequences and posture sequences are extracted through hierarchical feature extraction and dynamic analysis of the network structure to obtain feature sequence data; According to the feature sequence data, the feature enhancement values ​​of expressions and actions are calculated, and the calculation formula is: Among them, c k is the confidence of feature k, s k is the stability score of feature k, r k is the response strength of the feature to the dynamic baseline, Z is the normalization constant, m is the total number of features in the analysis, and E is the feature enhancement value; Based on the feature enhancement value, the action is calibrated to generate a weighted feature enhancement result.

5. The method for emotional interaction of smart home based on surveillance camera according to claim 1 is characterized in that: The steps for obtaining the emotional behavior assessment result are: Based on the weighted feature enhancement result, the expression change sequence and the posture change sequence are extracted, multiple behavior segments are divided by time window, posture angle calculation, expression state analysis and motion trajectory decomposition are performed on each behavior segment to generate behavior pattern data; Analyze the expression and posture change trends based on the behavior pattern data, compare the current behavior pattern with the past behavior pattern, analyze the behavior deviation and change range, and generate behavior pattern evaluation data; Based on the behavior pattern assessment data, abnormal points of emotion changes are detected through stability analysis of continuous behavior sequences, and are classified into emotion categories to generate emotion behavior assessment results.

6. The method for emotional interaction of smart home based on surveillance camera according to claim 1 is characterized in that: The steps for obtaining the dynamic early warning decision result are: Based on the emotional behavior assessment results, benchmark indicators for different types of emotions and behaviors are set, and abnormality determination input data is generated by analyzing the emotion change rate, the degree of behavioral deviation, and the frequency of continuous abnormalities; According to the abnormal judgment input data, the abnormal emotion score is calculated, and the expression is: Among them, G is the abnormal emotion score, B j is the current emotional behavior value of the jth behavior feature point, M j is the mean value of the jth behavior feature point, T j is the time variation factor of the jth behavior feature point, r is the total number of emotional behavior feature points; Based on the abnormal emotion score, a warning threshold is set. If the abnormal emotion score exceeds the warning threshold, the abnormal state is marked, and the abnormal category is recorded and classified to generate a dynamic warning decision result.

7. The method for emotional interaction of smart home based on surveillance camera according to claim 1, characterized in that: The steps for obtaining the abnormal event tracking record are: Based on the dynamic warning decision results, extract the time series information of abnormal emotional events and abnormal behavioral events, analyze the occurrence time, duration, number of occurrences and time interval of each abnormal event, and obtain abnormal event tracking input data; According to the abnormal event tracking input data, the severity score of the abnormal event is calculated, and the expression is: Wherein, J is the severity score of the abnormal event, A is the emotional amplitude of the current abnormal emotional event, CB is the normal amplitude mean of the current emotional category, X is the current spatial position coordinate of the abnormal behavior event, Y is the spatial position coordinate of the abnormal behavior event, Z is the time point of the abnormal behavior event, W is the time point of the abnormal behavior event, and U is the change rate of the abnormal event; Based on the severity score of the abnormal event, determine whether the abnormal emotional event or abnormal behavioral event meets the preset threshold. If the severity score of the abnormal event exceeds the preset threshold, record the time, location and category of the abnormal event and generate an abnormal event tracking record.

8. The method for emotional interaction of smart home based on surveillance camera according to claim 1 is characterized in that: The steps for obtaining the abnormality positioning optimization result are as follows: Based on the abnormal event tracking record, analyzing the geographical location and time data of the abnormal behavior, extracting the coordinate information and timestamp of the abnormal event, and generating abnormal behavior positioning input data; According to the abnormal behavior location input data, the abnormal behavior aggregation degree is calculated, and the expression is: Where K is the concentration of abnormal behavior, CM is the number of abnormal behavior events in the current time window, N is the average number of abnormal behavior events, R is the number of repeated abnormal behaviors in the same area, X is the current spatial position coordinate of the abnormal behavior event, Y is the spatial position coordinate of the abnormal behavior event, T is the time point of the current abnormal behavior, and S is the average time point of the abnormal behavior; Based on the abnormal behavior concentration, the viewing angle range and focus of the surveillance camera are adjusted to generate an abnormality positioning optimization result.

9. The home emotional interaction system based on the surveillance camera smart home emotional interaction method according to any one of claims 1 to 8, characterized in that: include: The monitoring data acquisition module installs monitoring cameras in the living space, continuously captures video frames, and extracts the basic biometric data of the residents from the video frames, including facial features and body postures, to obtain a basic biometric feature set; The expression and posture analysis module analyzes facial expressions and body postures based on the basic biometric feature set, combines the continuous expression and posture data stream, emphasizes key actions and expressions through parameter adjustment, obtains weighted features, and generates comprehensive expression and posture analysis results; The emotional behavior assessment module uses the comprehensive analysis results of facial expressions and postures to assess the behavior patterns and emotional changes of residents. Through parameter comparison and behavioral pattern analysis, it determines the emotional state and the estimated behavioral pattern changes, and obtains the emotional behavior analysis results; The dynamic early warning judgment module sets thresholds based on the results of emotional behavior analysis, continuously monitors the behavior and emotions of residents, and when abnormal emotions or behaviors are detected, evaluates and verifies them according to preset standards to generate dynamic early warning decision results; The abnormal event tracking module locates and tracks abnormal emotional or behavioral events based on the dynamic warning decision results, records the time and location of the events, adjusts the monitoring focus to optimize the monitoring effect, and obtains abnormal event records.

Citation Information

Cited By

  • Method for detecting moving target in dynamic flight of unmanned aerial vehicle

    CN120510539A

  • Personnel abnormal behavior supervision method and system for important places

    CN121353990A