A police handling process intelligent support method and system based on end-side AI

By using edge AI to extract multimodal features and make risk decisions from audio and video data on law enforcement recorders, the problem of real-time risk warning in law enforcement scenarios in existing technologies has been solved. This enables rapid and secure risk assessment and warning, and improves the safety and adaptability of the law enforcement process.

CN122176600APending Publication Date: 2026-06-09XIAOSHAN DISTRICT BRANCH OF HANGZHOU PUBLIC SECURITY BUREAU +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAOSHAN DISTRICT BRANCH OF HANGZHOU PUBLIC SECURITY BUREAU
Filing Date
2026-03-09
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing law enforcement recorders struggle to assess the emotional fluctuations of individuals involved in sudden and highly confrontational law enforcement scenarios in real time, as well as the potential for violent behavior. This leads to delayed responses or inappropriate actions, posing a significant risk to law enforcement safety. Furthermore, cloud-based AI analysis is limited by network transmission latency and privacy concerns, making it difficult to achieve low-latency, high-reliability real-time risk warnings.

Method used

We use an edge AI-based law enforcement recorder to collect on-site audio and video data and extract multimodal action features. Combined with an interpretable two-layer risk decision-making structure, including a feature triggering mechanism based on hard rules and a fuzzy logic reasoning mechanism, we perform real-time risk scoring and early warning. We construct an abnormal behavior database and optimize historical data by weighted fusion calculation using emotion stability index and cooperation score.

Benefits of technology

It achieves millisecond-level feature analysis and risk assessment, quickly identifies high-risk abnormal behaviors, predicts future violent tendencies, improves the initiative and security of the police response process, reduces network bandwidth consumption and the risk of sensitive data leakage, and is universally applicable to different law enforcement scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176600A_ABST
    Figure CN122176600A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of police intelligence, in particular to a police handling process intelligent support method and system based on end-side AI, which comprises the following steps: collecting real-time audio and video data on the spot based on a law enforcement recorder integrated with an end-side AI engine, and receiving or automatically identifying an event type; performing multi-modal action feature extraction on the audio and video data based on the end-side AI engine, and returning to a back-end system, wherein the multi-modal action features comprise expression features and behavior features; analyzing the received multi-modal action features based on an interpretable double-layer risk decision structure, and outputting a risk score; and sending a multi-level risk warning signal to the law enforcement recorder based on the risk score, so that the application has the real-time risk warning effect of low delay and high reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent policing, and in particular to an intelligent support method and system for the police response process based on edge AI. Background Technology

[0002] Currently, body cameras are standard equipment for frontline police officers and are mainly used to standardize the entire law enforcement process, facilitate post-event review, and protect the legitimate rights and interests of law enforcement officers and law enforcement subjects. Their functions are mainly focused on the real-time display, automatic collection, and storage of audio and video data during the law enforcement process.

[0003] However, in actual police operations, when faced with sudden and highly confrontational law enforcement scenarios, police officers often find it difficult to judge and clearly perceive the emotional fluctuations of the individuals involved and the violent behaviors that may arise, which can easily lead to delayed handling or inappropriate responses, posing a high risk to law enforcement safety.

[0004] In existing technologies, some systems attempt to analyze law enforcement videos using cloud-based AI, but due to limitations in network transmission latency, bandwidth, and privacy and security issues, it is difficult to achieve low-latency, high-reliability real-time risk warnings. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a method and system for intelligent support of the emergency response process based on edge AI.

[0006] On the one hand, this application provides an intelligent support method for the alarm handling process based on edge AI, which adopts the following technical solution: A method for intelligent support of the emergency response process based on edge AI includes the following steps: The law enforcement recorder, which integrates an edge AI engine, collects real-time audio and video data from the scene and receives or automatically identifies the type of event. The system uses an edge AI engine to extract multimodal motion features from audio and video data and transmits them back to the backend system. The multimodal motion features include facial expression features and behavioral features. Based on an interpretable two-layer risk decision-making structure, the received multimodal action features are analyzed and a risk score is output. Based on risk scoring, multi-level risk warning signals are sent to the law enforcement recorder; The interpretable two-layer risk decision structure includes a feature triggering mechanism based on preset hard rules and a fuzzy logic reasoning mechanism based on weighted fusion. The feature triggering mechanism is used to detect in real time whether there are abnormal behaviors in the behavioral features and to determine whether to output a high-risk score based on the recognition results. The fuzzy logic reasoning mechanism is to calculate the risk score by weighted fusion of the emotion stability index and cooperation score generated based on multimodal action features.

[0007] In one embodiment: the database of abnormal behavior is constructed based on historical record analysis, and the construction method includes: Triggering behavioral features are generated based on historical conflict records. By analyzing the proportion of historical conflict records in historical records containing triggering behavioral features, triggering behavioral features with a proportion higher than a first threshold are marked as pseudo-abnormal behavioral features. Here, historical conflict records refer to historical records of conflicts that occurred during law enforcement. Triggering features are characterized by behavioral features that occurred within a preset time period before the conflict occurred in historical conflict records. Filter the first historical record subset containing only a single pseudo-abnormal behavior feature from the historical records. Based on the first historical record subset and the historical conflict records in the first historical record subset, calculate the isolated conflict rate of the single feature. Pseudo-abnormal behavior features with an isolated conflict rate greater than a second threshold are marked as single abnormal behaviors. A second historical data subset containing multiple pseudo-abnormal behavior features that were not marked as anomalous behavior is selected from the historical data. The combined conflict rate of multiple features is calculated. Combined pseudo-abnormal behavior features with a combined conflict rate greater than a third threshold are marked as combined anomalous behavior. A database of abnormal behaviors is constructed based on single and combined abnormal behaviors.

[0008] In one embodiment: the method for obtaining the weights corresponding to the emotional stability index and cooperation score specifically includes: Based on the event type, the causal effect values ​​of the emotional stability index and cooperation score on the conflict event were calculated respectively. The emotional stability weight and cooperation weight are obtained by normalizing the weights according to the causal effect value.

[0009] In one embodiment: by constructing a quantitative model of the emotional fluctuations of the person involved in the case, the facial expression features are analyzed to generate a real-time emotional stability index of the person involved in the case, and an emotional fluctuation curve of the person involved in the case is constructed based on the emotional stability index; the emotional stability index is numerically represented by preset facial expression features.

[0010] In one embodiment: the method for obtaining the facial expression features is as follows: facial data of the person involved in the case is obtained based on audio and video data, and cross-frame tracking is performed through a target tracking algorithm to obtain continuous multi-frame facial video data; based on a facial motion coding system, the facial action unit AU intensity of each frame image is extracted as the facial expression feature from the multi-frame facial video data. The steps to generate a real-time sentiment stability index include: Generate frame-level expression intensity vectors for each time step based on facial expression features; The frame-level expression intensity vector is input into a pre-trained expression classification model to obtain the emotion probability distribution containing multiple emotion category probability values ​​at the corresponding time. Based on the probability sequence of various emotions in the emotional probability distribution within a preset time period, calculate the probability fluctuation index of various emotions. Based on the preset weights of various emotions, the probability fluctuation indicators of the various emotions are weighted and calculated to obtain the emotion stability index.

[0011] In one embodiment: the behavioral feature is the coordinates of human joint points extracted from frame images of audio and video data using a human pose estimation algorithm; The steps for generating a cooperation score specifically include: identifying abnormal action categories and their corresponding confidence levels based on behavioral features using a preset posture classification model; quantifying and deducting points for the identified abnormal actions according to preset deduction rules; and fusing the deduction results from multiple frames to obtain the cooperation score.

[0012] In one embodiment: a mood evolution prediction model trained on historical records is used to perform temporal modeling of multimodal action features, predict changes in multimodal action features within a preset time window in the future, calculate future risk scores based on the predicted multimodal action features, and send multi-level risk warning signals to the law enforcement recorder based on the future risk scores.

[0013] In one embodiment: the multimodal action features also include voice features. By analyzing the voice features of law enforcement officers, verbal stimuli that may trigger emotional agitation in the persons involved in the case are identified. The correlation coefficient between the verbal stimuli and the emotional features of the persons involved in the case is analyzed, and the future risk score is corrected by the correlation coefficient.

[0014] In one embodiment: the method for generating the correlation coefficient includes: An emotion correlation coefficient is constructed based on verbal stimuli and real-time emotion fluctuation curves. Then, based on the probability of abnormal behavior occurring after the same verbal stimulus obtained from historical records, a behavior correlation coefficient is constructed. The correlation coefficient is obtained by adjusting the emotion correlation coefficient based on the behavioral correlation coefficient.

[0015] On the other hand, the intelligent support system for the alarm handling process based on edge AI provided in this application adopts the following technical solution: An intelligent support system for emergency response based on edge AI, comprising: The law enforcement recorder integrates an audio and video acquisition module and an edge AI engine to collect on-site audio and video data in real time, receive or automatically identify event types, extract multimodal action features from the audio and video data through the edge AI engine, and send the multimodal action features to the backend system. The backend system integrates an emotion quantification and prediction module, an abnormal behavior recognition module, and an early warning output module. It is used to form an interpretable two-layer risk decision structure through the emotion quantification and prediction module and the abnormal behavior recognition module, analyze the received multimodal action features to output a risk score, and send multi-level risk warning signals to the law enforcement recorder based on the risk score.

[0016] In summary, this application has the following beneficial effects: 1. By extracting features through the edge AI engine, only lightweight feature data is sent back to the backend system. Compared with existing cloud AI analysis solutions, this completely avoids problems such as network transmission latency and bandwidth limitations, and achieves millisecond-level feature analysis and risk assessment. 2. By using an interpretable two-layer risk decision-making structure, the mechanism of mandatory triggering of abnormal behavior based on hard rules is combined with the fuzzy logic reasoning mechanism based on weighted fusion. This can not only quickly determine sudden high-risk abnormal behavior, but also perform multi-dimensional weighted calculation by combining emotional stability index and cooperation score, so as to achieve refined output of risk score. 3. By combining the emotion evolution prediction model with the temporal modeling of multimodal action features, it is possible to predict the future risk score within a preset time window, realize the advance judgment of the violent behavior tendency of the suspect, and enable the police to take intervention measures in advance, resolve law enforcement conflicts at the bud stage, and greatly improve the initiative in the police response process. 4. By focusing on the behavioral interactions between law enforcement officers and suspects, we can quantitatively analyze the correlation between law enforcement language and the emotional fluctuations of those involved in the case, and use correlation coefficients to correct risk scores, thereby avoiding unnecessary conflicts and promoting the standardization of law enforcement. 4. The abnormal behavior database construction and the weight allocation of emotion and cooperation in this application are all based on historical law enforcement records, and support continuous iterative optimization of the abnormal behavior database and emotion evolution prediction model based on new law enforcement data; at the same time, it can dynamically adjust the weight of emotion stability index and cooperation score according to different event types, meet the differences in law enforcement habits and scenarios in different regions, and greatly improve the universality and practical application value of the system. 5. The original sensitive audio and video data in this application is collected locally by the end-side law enforcement recorder and the feature extraction is completed. Only the extracted multimodal action features and other lightweight data are sent back to the backend system. The original data does not need to be transmitted to the cloud or externally. This reduces the consumption of network bandwidth and effectively avoids the risk of leakage of sensitive data such as the personnel involved in the case and the law enforcement scene during the law enforcement process. It strictly meets the requirements for security management and privacy protection of police data. Attached Figure Description

[0017] Figure 1 This is a logical block diagram of an intelligent support system for the emergency response process based on edge AI in this embodiment; Figure 2 This is a flowchart of an intelligent support method for the emergency response process based on edge AI in this embodiment. Detailed Implementation

[0018] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. In some cases, to avoid obscuring various aspects of this application due to unnecessary description, well-known methods, processes, systems, components, and / or circuits already described at a higher level will not be elaborated upon. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope of protection claimed in this application.

[0019] It should be noted that the descriptions of these embodiments are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0020] In the description of this application, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0021] In the description of this application, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples.

[0022] like Figure 1 As shown, this embodiment provides an intelligent support system for the police response process based on edge AI, including a law enforcement recorder and a backend system, wherein, The law enforcement recorder integrates an audio and video acquisition module and an edge AI engine. The audio and video acquisition module is used to collect on-site audio and video data in real time, providing a basic data source for feature extraction.

[0023] It should be noted that in most cases, before dispatching an officer, the body camera can wirelessly receive the event type of the incident. However, since the initial event type is based on information provided hastily by the officer involved, it may be adjusted and supplemented during the response process. Therefore, the edge AI engine also has the function of automatically recognizing the event type and making real-time adjustments. Specifically, this recognition can be achieved through voice recognition during law enforcement or through manual input by the officer.

[0024] The edge AI engine is mainly used to extract multimodal motion features from audio and video data and send the multimodal motion features to the backend system.

[0025] The backend system is deployed in a local smart cabinet, which integrates an emotion quantification and prediction module, an abnormal behavior recognition module, and an early warning output module.

[0026] A two-tiered, interpretable risk decision-making structure is formed by an emotion quantification and prediction module and an abnormal behavior recognition module. This structure analyzes received multimodal action features to output a risk score. During the process, the emotion quantification and prediction module receives facial expression features from the multimodal action features and uses a trained emotion fluctuation quantification model to calculate an emotion stability index. The abnormal behavior recognition module receives behavioral features from the multimodal action features and compares them with feature templates in an abnormal behavior database to identify abnormal behaviors.

[0027] Finally, the early warning output module sends multi-level risk warning signals to the law enforcement recorder based on the risk score.

[0028] like Figure 2As shown, this application also discloses an intelligent support method for the emergency response process based on edge AI, specifically including: The S100 is a law enforcement recorder with an integrated edge AI engine that collects real-time audio and video data from the scene and receives or automatically identifies the type of event.

[0029] The types of events include theft, disputes, violent conflicts, gang fights, robbery, etc. In addition, the event types also include other information such as the number of participants, the number of injured, the nature of the injuries, and whether it is indoors or outdoors.

[0030] It should be noted that before collecting real-time audio and video data from the scene using a law enforcement recorder, the identity information of the persons involved in the case must be verified. In addition, the entire video collection process is conducted only after the persons involved in the case have given their authorization.

[0031] The S200 uses an edge-side AI engine to extract multimodal motion features from audio and video data and transmits them back to the backend system.

[0032] It should be noted that the multimodal action features include facial expression features and behavioral features.

[0033] Specifically, the method for obtaining the facial features is as follows: facial data of the person involved in the case is obtained based on audio and video data, and cross-frame tracking is performed using target tracking algorithms such as DeepSORT to ensure the continuity of multi-frame facial expression sequences, thereby obtaining continuous multi-frame facial video data.

[0034] Then, based on the facial action coding system, the facial action unit (AU) intensity of each frame of the image is extracted from the multi-frame facial video data using the OpenFace or FaceNet model as expression features. The AU intensity of the action unit is such as AU12 for raising the corners of the mouth and AU1+2 for frowning.

[0035] The behavioral features are the coordinates of 18 / 25 human joints extracted from frame images of audio and video data using human pose estimation algorithms such as OpenPose or AlphaPose.

[0036] In addition, to eliminate interference, the voiceprint characteristics of law enforcement officers can be pre-recorded and cached on the device, thereby enabling rapid screening of law enforcement officers' voices through voiceprint matching.

[0037] S300 analyzes the received multimodal action features based on an interpretable two-layer risk decision structure and outputs a risk score.

[0038] In the above steps, the interpretable two-layer risk decision-making structure includes a feature-triggered mechanism based on preset hard rules and a fuzzy logic reasoning mechanism based on weighted fusion.

[0039] The feature triggering mechanism is used to detect abnormal behavior in behavioral features in real time and determine whether to output a high-risk score based on the recognition results. During the recognition process, the similarity between the real-time detected behavioral features and feature templates in the database of abnormal behaviors is compared to determine the risk level.

[0040] The database of abnormal behavior is built based on historical record analysis, and the specific construction methods include: First, inducement behavior features are generated based on historical conflict records. Then, by analyzing the proportion of historical conflict records in historical records that contain inducement behavior features, inducement behavior features with a proportion higher than a first threshold are marked as pseudo-abnormal behavior features.

[0041] Among them, historical conflict records refer to historical records of conflicts that occurred during law enforcement. These records include anonymized audio and video recordings of law enforcement actions, corresponding multimodal action characteristics, and the outcomes of the actions taken. Trigger characteristics are behavioral features that occurred within a preset timeframe prior to the occurrence of the conflict in the historical conflict records.

[0042] Secondly, a first historical data subset containing only a single pseudo-abnormal behavior feature is selected from the historical records. Based on the first historical data subset and the historical conflict records within it, the isolated conflict rate of a single feature is calculated. Pseudo-abnormal behavior features with an isolated conflict rate greater than a second threshold are marked as single anomalous behaviors. The isolated conflict rate is calculated as the ratio of historical conflict records in the first historical data subset to the total number of historical data records in the first historical data subset. Furthermore, when calculating the isolated conflict rate, the first historical data subset corresponding to the single pseudo-abnormal behavior feature must be larger than the minimum sample size.

[0043] Then, a second historical data subset containing multiple pseudo-abnormal behavior features that were not marked as anomalous behavior is filtered from the historical data. The combined conflict rate of multiple features is calculated, and the combined pseudo-abnormal behavior features with a combined conflict rate greater than a third threshold are marked as combined anomalous behavior.

[0044] Finally, a database of abnormal behaviors is constructed based on single and combined abnormal behaviors.

[0045] Among them, the first threshold, the second threshold and the third threshold can be preset manually based on historical experience.

[0046] The fuzzy logic reasoning mechanism calculates a risk score by weighting and fusing the emotion stability index and cooperation score generated based on multimodal action features.

[0047] The specific methods for obtaining the weights corresponding to the emotional stability index and cooperation score include: First, based on the event type, the causal effect values ​​of the emotional stability index and cooperation score on the conflict event were calculated respectively.

[0048] Specifically, the causal effect value of the emotional stability index and the conflict event: ACE e =P(conflict|E decreases) -P(conflict|E remains unchanged).

[0049] Compatibility score and causal effect value of conflict events: ACE C =P(conflict|C decreases) -P(conflict|C remains unchanged).

[0050] Secondly, the weights are allocated according to the normalization of the causal effect value to obtain the emotional stability weight and the cooperation weight.

[0051] The formula for calculating the weights based on the normalized causal effect value is: w e =ACE e / (ACE e +ACE C ), w C =1-w e .

[0052] In this embodiment, the formula for calculating the risk score is: In the formula, E is the emotional stability index, C is the cooperation score, We is the emotional stability weight, Wc is the cooperation weight, 0≤E≤1, 0≤C≤100.

[0053] Among them, the emotional stability index is generated through facial expression features, and the cooperation score is generated through action features.

[0054] For the calculation of the emotional stability index, a quantitative model of the emotional fluctuations of the persons involved in the case is constructed, facial expression features are analyzed to generate the real-time emotional stability index of the persons involved in the case, and the emotional fluctuation curve of the persons involved in the case is constructed based on the emotional stability index; the emotional stability index is numerically represented by preset facial expression features.

[0055] The steps to generate a real-time sentiment stability index include: First, a frame-level expression intensity vector corresponding to each moment is generated based on the expression features; Secondly, the frame-level expression intensity vector is input into a preset expression classification model to obtain the emotion probability distribution at the corresponding time moment, which includes probability values ​​of multiple emotion categories. Depending on the preset number of emotion categories in the model, multiple probability values ​​can be output for each frame image; for example, when seven emotion categories are preset, the emotion probability distribution P... emo =[p 1 t ,p 2 t ,...,p 7 t ].

[0056] It should be noted that a lightweight temporal convolutional network is preferred for the facial expression classification model. The model can be trained on a modified AffectNet sentiment analysis dataset. For example, the native continuous AU intensity values ​​in the AffectNet dataset are normalized within the [0,1] interval. Samples strongly correlated with emotions in police response scenarios are selected. Simultaneously, the general emotion categories of the dataset are reconstructed into police response-specific emotion categories such as calm, irritability, and mild anger, thus constructing a dedicated labeled dataset between AU intensity and police response scenario emotions. The cross-entropy loss function is used as the model training loss function, combined with mini-batch gradient descent. After training, the model is pruned and optimized for lightweighting to reduce the number of model parameters while maintaining classification accuracy, adapting to the real-time inference requirements of the backend system.

[0057] Then, based on the probability sequences of various emotions in the emotional probability distribution within the preset time period, the probability fluctuation index of each emotion is calculated. In this embodiment, the probability fluctuation index is calculated using the standard deviation (Std) or the mean absolute deviation (MAD). To make outliers more robust, the mean absolute deviation is preferred for calculation. The calculation formula is as follows: Finally, based on the preset weights of various emotions, the probability fluctuation indicators of the various emotions are weighted and calculated to obtain the emotion stability index.

[0058] Specifically, each emotion is pre-defined with a weight vector W. emo =[w 1 emo ,w 2 emo ,...,w n emo The formula for calculating the emotional stability index is: Among them, the larger the calculated emotional stability index E value, the greater the emotional fluctuation and the lower the stability.

[0059] The specific steps for generating a compatibility score include: Based on behavioral characteristics, the spatial position of the key point Jt=[x1t,y1t,...,xkt,ykt] is calculated. Then, using preset violation action templates such as waving, pushing, and escape tendency, the abnormal action category and confidence level are identified based on a preset posture classification model. The confidence level ranges from [0,1], with a higher value indicating a greater likelihood that the person involved in the case will perform the violation action in the current frame.

[0060] It should be noted that the pose classification model can utilize spatial-temporal graph convolutional networks such as ST-GCN, and be trained using publicly available pose datasets modified from the NTU RGB+D dataset, combined with historical law enforcement behavior annotation data. For example, by selecting samples from the dataset related to illegal actions in police response scenarios, extracting the spatial locations of key points in the samples, aligning them with the key point sequence features corresponding to the illegal action templates in police response scenarios, labeling the illegal action categories and corresponding confidence labels to complete the annotation system reconstruction, using the cross-entropy loss function to complete model training, and then pruning and optimizing the model after training to adapt to the real-time inference requirements of the backend system.

[0061] According to the preset deduction rules, the identified abnormal actions are quantified and deducted, and the deduction results of multiple frames are merged to obtain the cooperation score.

[0062] Specifically, based on the previously generated abnormal action categories and confidence levels, a pre-defined deduction vector D=[D1,D2,...,Dj] is used for each abnormal action category j. n The deduction score for a single frame is calculated as follows: In the formula, C t c is the deduction for a single frame at time t. j D represents the confidence level corresponding to the abnormal action category j. j c is the deduction vector corresponding to the abnormal action category j. j ∈[0,1].

[0063] The cooperation score is obtained by fusing the scores from multiple frames, and the specific formula is as follows: S400 sends multi-level risk warning signals to law enforcement recorders based on risk scoring.

[0064] In the above steps, the multi-level risk warning signal includes three warning levels: Level 1 Warning: There is a high probability of violent behavior. It is recommended to take immediate coercive measures and call for support. Level 2 Warning: Significant emotional fluctuations, with a moderate risk of conflict. It is recommended to maintain a safe distance and adjust communication strategies. Level 3 Warning: Slight emotional fluctuations. It is recommended to continue to observe and use calming language.

[0065] The high-risk score in step S200 above corresponds to a Level 1 warning. Different levels of warnings can be differentiated by setting different alert sounds or vibration frequencies on the law enforcement recorder.

[0066] In one embodiment, a pre-defined emotion evolution prediction model is trained using historical records. In this embodiment, the emotion evolution prediction model employs an LSTM (Long Short-Term Memory) network. During training, the multimodal action feature time sequence from the historical records is used as input, and the changes in multimodal action features within the corresponding future pre-defined time window in the historical records are used as the model output.

[0067] Then, the trained emotion evolution prediction model is used to perform temporal modeling of multimodal action features, predict the changes of multimodal action features within a preset time window, calculate the future risk score based on the predicted multimodal action features, and finally send multi-level risk warning signals to the law enforcement recorder based on the future risk score.

[0068] In another embodiment, the multimodal action features also include speech features. In this embodiment, the speech features are obtained as follows: the speech features are extracted by a lightweight acoustic feature extraction model on the edge, such as the MFCC lightweight model. The quantifiable acoustic feature parameters such as volume, speech rate, fundamental frequency, and pitch slope are extracted frame by frame from the law enforcement officer's speech after framing, forming a frame-level acoustic pitch feature vector. Based on the keyword matching library cached on the edge, lightweight semantic analysis is performed by a lightweight NLP algorithm on the edge to generate keyword feature labels to determine whether there are command words or other keywords. Finally, the acoustic pitch feature vector and the keyword feature labels are fused to generate speech features.

[0069] By analyzing the vocal characteristics of law enforcement officers, verbal stimuli that may trigger emotional agitation in those involved in a case are identified. The identification process employs a combination of quantitative parameter threshold determination and keyword tag matching. Specifically, if any acoustic feature parameter exceeds its corresponding threshold, it is determined that a provocative tone is present. Simultaneously, if the generated keyword feature tags contain keywords from a keyword matching database, both the keyword feature tags and the acoustic feature parameters are used as verbal stimuli.

[0070] Then, the correlation coefficient between verbal stimuli and the emotional characteristics of the individuals involved is analyzed, and the future risk score is adjusted based on the correlation coefficient.

[0071] In this embodiment, the method for generating the correlation coefficient includes: An emotion correlation coefficient is constructed based on verbal stimuli and real-time emotion fluctuation curves. In this step, the change in the emotion fluctuation curve before and after the verbal stimuli is determined. Specifically, the difference in the emotion stability index before and after the verbal stimuli can be directly calculated. When the difference is positive, it indicates that the verbal stimuli have had a negative impact, and the larger the difference, the greater the impact.

[0072] Then, based on the probability of abnormal behavior occurring after the same verbal stimulus obtained from historical records, a behavioral correlation coefficient is constructed. When constructing the behavioral correlation coefficient, a baseline probability value can be set. If the probability of abnormal behavior occurring after the verbal stimulus is less than the baseline probability value, the behavioral correlation coefficient is set to a value less than 1; otherwise, it is set to a value greater than 1.

[0073] The correlation coefficient is obtained by adjusting the emotion correlation coefficient based on the behavior correlation coefficient. In this embodiment, the product of the emotion correlation coefficient and the behavior correlation coefficient is directly used as the correlation coefficient.

[0074] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for intelligent support of the emergency response process based on edge AI, characterized in that, include: The law enforcement recorder, which integrates an edge AI engine, collects real-time audio and video data from the scene and receives or automatically identifies the type of event. The system uses an edge AI engine to extract multimodal motion features from audio and video data and transmits them back to the backend system. The multimodal motion features include facial expression features and behavioral features. Based on an interpretable two-layer risk decision-making structure, the received multimodal action features are analyzed and a risk score is output. Based on risk scoring, multi-level risk warning signals are sent to the law enforcement recorder; The interpretable two-layer risk decision structure includes a feature triggering mechanism based on preset hard rules and a fuzzy logic reasoning mechanism based on weighted fusion. The feature triggering mechanism is used to detect in real time whether there are abnormal behaviors in the behavioral features and to determine whether to output a high-risk score based on the recognition results. The fuzzy logic reasoning mechanism is to calculate the risk score by weighted fusion of the emotion stability index and cooperation score generated based on multimodal action features.

2. The intelligent support method for emergency response process based on edge AI according to claim 1, characterized in that, The database of abnormal behavior is constructed based on historical record analysis, and the construction method includes: Triggering behavioral features are generated based on historical conflict records. By analyzing the proportion of historical conflict records in historical records containing triggering behavioral features, triggering behavioral features with a proportion higher than a first threshold are marked as pseudo-abnormal behavioral features. Here, historical conflict records refer to historical records of conflicts that occurred during law enforcement. Triggering features are characterized by behavioral features that occurred within a preset time period before the conflict occurred in historical conflict records. Filter the first historical record subset containing only a single pseudo-abnormal behavior feature from the historical records. Based on the first historical record subset and the historical conflict records in the first historical record subset, calculate the isolated conflict rate of the single feature. Pseudo-abnormal behavior features with an isolated conflict rate greater than a second threshold are marked as single abnormal behaviors. A second historical data subset containing multiple pseudo-abnormal behavior features that were not marked as anomalous behavior is selected from the historical data. The combined conflict rate of multiple features is calculated. Combined pseudo-abnormal behavior features with a combined conflict rate greater than a third threshold are marked as combined anomalous behavior. A database of abnormal behaviors is constructed based on single and combined abnormal behaviors.

3. The intelligent support method for the emergency response process based on edge AI according to claim 1, characterized in that, The specific methods for obtaining the weights corresponding to the emotional stability index and cooperation score include: Based on the event type, the causal effect values ​​of the emotional stability index and cooperation score on the conflict event were calculated respectively. The emotional stability weight and cooperation weight are obtained by normalizing the weights according to the causal effect value.

4. The intelligent support method for the emergency response process based on edge AI according to claim 1, characterized in that: By constructing a quantitative model of the emotional fluctuations of the persons involved in the case, the facial expression features are analyzed to generate a real-time emotional stability index of the persons involved in the case, and an emotional fluctuation curve of the persons involved in the case is constructed based on the emotional stability index; the emotional stability index is numerically represented by preset facial expression features.

5. The intelligent support method for the emergency response process based on edge AI according to claim 4, characterized in that, The method for obtaining the facial expression features is as follows: facial data of the person involved in the case is obtained based on audio and video data, and cross-frame tracking is performed through a target tracking algorithm to obtain continuous multi-frame facial video data; based on a facial motion coding system, the facial action unit AU intensity of each frame image is extracted as the facial expression feature from the multi-frame facial video data. The steps to generate a real-time sentiment stability index include: Generate frame-level expression intensity vectors for each time step based on facial expression features; The frame-level expression intensity vector is input into a pre-trained expression classification model to obtain the emotion probability distribution containing multiple emotion category probability values ​​at the corresponding time. Based on the probability sequence of various emotions in the emotional probability distribution within a preset time period, calculate the probability fluctuation index of various emotions. Based on the preset weights of various emotions, the probability fluctuation indicators of the various emotions are weighted and calculated to obtain the emotion stability index.

6. The intelligent support method for emergency response process based on edge AI according to claim 1, characterized in that, The behavioral features are the coordinates of human joint points extracted from frame images of audio and video data using a human pose estimation algorithm. The steps for generating a cooperation score specifically include: based on behavioral characteristics, identifying abnormal action categories and their corresponding confidence levels through a pre-defined posture classification model; According to the preset deduction rules, the identified abnormal actions are quantified and deducted, and the deduction results of multiple frames are merged to obtain the cooperation score.

7. The intelligent support method for emergency response process based on edge AI according to claim 1, characterized in that: By using an emotion evolution prediction model trained on historical records, the multimodal action features are temporally modeled to predict changes in multimodal action features within a preset time window in the future. Based on the predicted multimodal action features, a future risk score is calculated, and a multi-level risk warning signal is sent to the law enforcement recorder based on the future risk score.

8. The intelligent support method for the emergency response process based on edge AI according to claim 7, characterized in that: The multimodal action features also include voice features. By analyzing the voice features of law enforcement officers, verbal stimuli that may trigger emotional agitation in the persons involved in the case are identified. The correlation coefficient between the verbal stimuli and the emotional features of the persons involved in the case is analyzed, and the future risk score is corrected based on the correlation coefficient.

9. The intelligent support method for the alarm handling process based on edge AI according to claim 8, characterized in that: Methods for generating correlation coefficients include: An emotion correlation coefficient is constructed based on verbal stimuli and real-time emotion fluctuation curves. Then, based on the probability of abnormal behavior occurring after the same verbal stimulus obtained from historical records, a behavior correlation coefficient is constructed. The correlation coefficient is obtained by adjusting the emotion correlation coefficient based on the behavioral correlation coefficient.

10. An intelligent support system for emergency response process based on edge AI, employing the intelligent support method for emergency response process based on edge AI as described in any one of claims 1-9, characterized in that, include: The law enforcement recorder integrates an audio and video acquisition module and an edge AI engine to collect on-site audio and video data in real time, receive or automatically identify event types, and extract multimodal action features from the audio and video data through the edge AI engine, and send the multimodal action features to the backend system. The backend system integrates an emotion quantification and prediction module, an abnormal behavior recognition module, and an early warning output module, which are used to form an interpretable two-layer risk decision structure through the emotion quantification and prediction module and the abnormal behavior recognition module, and analyze the received multimodal action features to output a risk score. Based on risk scores, multi-level risk warning signals are sent to the law enforcement recorder.