Police risk multistage early warning system based on large model cooperative calculation

The multi-level early warning system for police risks, which uses large-scale collaborative computing, solves the problem of insufficient multimodal data fusion in traditional police risk early warning systems. It achieves high-precision identification and multi-level response to police risks, thereby improving the efficiency of police risk management.

CN120973845AActive Publication Date: 2025-11-18JIANGSU LIANFENG GOLDEN SHIELD INTELLIGENT TECH CO
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511495395.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Traditional police risk early warning systems lack multimodal data fusion capabilities, have limited early warning accuracy, struggle to cope with complex and ever-changing police scenarios, and lack modeling of the spatiotemporal evolution of risk events, thus failing to achieve multi-level, adaptive early warning responses.

Method used

A multi-level early warning system for police risks based on large-scale model collaborative computing is adopted. This system collects and preprocesses police data in real time, uses a collaborative scheduling center and a large-scale model cluster for feature extraction and risk value calculation, and generates multi-level early warnings by combining a weighted fusion algorithm with an attention mechanism.

Benefits of technology

It significantly improved the accuracy of risk identification, reduced the false alarm and false alarm rates, achieved full-cycle coverage of police risks from the individual level to the trend level, and enhanced the pertinence and foresight of risk response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973845A_ABST
    Figure CN120973845A_ABST
Patent Text Reader

Abstract

The invention provides a police service risk multi-level early warning system based on large model cooperative computing, and the system comprises a data module which collects and preprocesses multi-source police service data in real time, converts the multi-source police service data into text, visual and space-time initial vectors, stores the initial vectors into a vector database, and recognizes an initial risk event through an event detector to generate an early warning task; the calculation module analyzes tasks through the collaborative scheduling center and calls corresponding vectors, and feature vectors and space-time attributes are generated through text analysis, visual analysis and space-time prediction large model cluster processing adjusted by the police service field; and the early warning module is used for calculating a comprehensive risk value based on an attention mechanism by fusing multi-modal characteristics, and realizing individual-level immediate early warning, regional-level situation early warning and trend-level deduction early warning by combining an early warning threshold value and time-space attributes. The system can realize multi-level, self-adaptive and accurate early warning of police risks, and improves the intelligent level of public security risk prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of public safety, in particular to a police risk multi-level early warning system based on large model collaborative computing. BACKGROUND

[0002] Traditional police risk early warning systems mainly rely on single data sources and rule engines for risk discrimination, and have problems of insufficient multi-modal data fusion capability and limited early warning accuracy. Existing methods usually use static thresholds or simple models for analysis, which are difficult to cope with complex and variable police scenarios, resulting in high false positive and false negative rates. In addition, traditional systems lack the ability to model the spatiotemporal evolution of risk events, and cannot achieve multi-level and adaptive early warning response, which restricts the improvement of public safety risk prevention and control efficiency.

[0003] Current technology has obvious shortcomings in real-time collaborative computing of multi-source heterogeneous police data. Text, vision and spatiotemporal data are often processed independently, lacking cross-modal semantic alignment and attention mechanism driven fusion methods, resulting in insufficient extraction of key risk features. At the same time, traditional models are difficult to support cross-scale risk reasoning from individual level to trend level, and cannot meet the urgent needs of modern police work for intelligent and accurate early warning. SUMMARY

[0004] In order to solve the technical problems mentioned in the background art, the purpose of the present application is to provide a police risk multi-level early warning system based on large model collaborative computing.

[0005] In order to achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:

[0006] The police risk multi-level early warning system based on large model collaborative computing comprises:

[0007] A data module: real-time acquisition of police data, pre-processing of police data, conversion of pre-processed police data into initial vectors, and storage of initial vectors in a preset vector database; setting an event detector to monitor the pre-processed police data in real time, and identifying initial risk events based on preset risk rules to generate early warning tasks;

[0008] A computing module: setting a collaborative scheduling hub and a large model cluster; the collaborative scheduling hub generates subtasks by analyzing the early warning tasks, and retrieves and calls the corresponding initial vectors from the vector database according to the subtasks; the large model cluster includes a text analysis model, a visual analysis model and a spatiotemporal prediction model adjusted by police domain data, and the large model cluster processes the initial vectors to generate feature vectors and spatiotemporal attributes;

[0009] The early warning module calculates a comprehensive risk value based on the feature vector through a weighted fusion algorithm based on an attention mechanism, and generates multi-level early warnings including individual-level immediate early warning, regional-level situation early warning and trend-level deduction early warning according to a warning threshold of the comprehensive risk value and the spatio-temporal attribute.

[0010] Further, the initial vector includes a text initial vector, a visual initial vector and a spatio-temporal initial vector.

[0011] Further, the text initial vector adopts a BERT model to acquire context awareness of text data;

[0012] The visual initial vector adopts a visual encoder based on a ResNet-50 architecture to extract multi-layer convolutional features from image data for acquisition;

[0013] The spatio-temporal initial vector extracts time period features and spatial grid encoding from spatio-temporal data and acquires them through a multi-layer perception machine mapping;

[0014] The initial vector is organized according to a preset time window and data source identifier, stored in a vector database, and an index based on density is established to support similarity retrieval.

[0015] Further, the event detector identifies the initial risk event based on a set of configurable preset rules;

[0016] When the event detector identifies an initial risk event, the early warning task is generated, and the early warning task at least includes an event unique identifier, an event type, an event occurrence time, an event associated core data index and a preliminary description and key features of the event.

[0017] Further, the subtasks include text processing requirements, visual processing requirements and spatio-temporal processing requirements;

[0018] The collaborative scheduling hub decomposes specific data processing requirements according to the event type, preliminary description and event associated core data index of the early warning task, and classifies these requirements into three types of subtasks: text processing requirements, visual processing requirements and spatio-temporal processing requirements;

[0019] The collaborative scheduling hub assigns the retrieved initial vector and the corresponding subtask to the corresponding professional model in the large model cluster for processing.

[0020] Further, the feature vector includes a text feature vector, a visual feature vector and a spatio-temporal feature vector;

[0021] The text analysis model is constructed through a Transformer-XL architecture, specifically including an input layer, a multi-head attention layer, a semantic reinforcement layer and an output layer;

[0022] The text initial vector is input into the input layer, and field self-adaption embedding is performed, and the specific operation is as follows,

[0023] The general semantic space is aligned to the police professional field, the police case terminology recognition capability is enhanced, and the embedding vector aligned to the police field is obtained ;

[0024] The multi-head attention layer captures long-distance semantic dependence and police case event time sequence relationship by calculating 12 heads of relative position attention, and attention features are obtained , and the formula is as follows:

[0025]

[0026] Among them, , and are the query matrix, the key matrix and the value matrix of the first head, is the relative position coding matrix of the first head and the first head, is the number of attention heads, is the dimension of the key or query vector;

[0027] The semantic reinforcement layer fuses the prior knowledge of crime type, and risk features are obtained;

[0028] The output layer fuses the attention features and the risk features, and the text feature vector is obtained.

[0029] Further, the visual analysis model is constructed by a multi-scale attention convolutional network, and specifically includes an input layer, a spatial pyramid layer, an attention layer and an output layer;

[0030] The visual initial vector is input into the input layer, and an optimized feature map Z is obtained by recombining the visual initial vector,

[0031]

[0032] Among them, is a convolution kernel, is a visual initial vector, is a rectified linear unit activation function;

[0033] The spatial pyramid layer performs multi-scale feature fusion on the optimized feature map, captures behavior patterns of different granularities, and obtains multi-scale fusion features;

[0034] The attention layer focuses on the abnormal behavior region, and obtains a behavior attention map;

[0035] The output layer performs channel weighting and dimension normalization to generate a standardized visual feature vector.

[0036] Further, the spatio-temporal prediction model is constructed by a spatio-temporal graph convolution network, specifically including an input layer, a propagation layer, an influence layer and an output layer.

[0037] The spatio-temporal initial vector is input into the input layer to construct a dynamic spatio-temporal graph.

[0038] The propagation layer alternately performs graph convolution and time series modeling to simulate the risk diffusion process and capture the event evolution law to obtain a spatio-temporal state matrix. ;

[0039] The influence layer calculates a diffusion radius to predict the spatial influence range of the risk event.

[0040] The output layer performs graph feature pooling and spatio-temporal attribute analysis to generate a spatio-temporal feature vector and a spatio-temporal attribute triple :

[0041]

[0042] wherein, is a core position, longitude and latitude, t is the duration of the risk event, is a graph-level pooling operation.

[0043] The dimensions of the text feature vector, the visual feature vector and the spatio-temporal feature vector are unified by a preset first fully connected layer.

[0044] Further, the calculation steps of the comprehensive risk value are as follows,

[0045] 1) generating a query vector, a key vector and a value vector of the feature vector by a preset attention perception layer:

[0046]

[0047] wherein, , and are the query vector, the key vector and the value vector, respectively, is the feature vector, represent text, vision and spatio-temporal, respectively, , , , , and are trainable parameters.

[0048] 2) Calculate the attention score of the feature vector by scaling the dot product attention :

[0049]

[0050] where, is the transpose of the key vector, is the dimension of the key vector;

[0051] 3) Based on the attention score weight the value vector to generate modal features ;

[0052] 4) Comprehensive risk value calculated by a preset second fully connected layer and activation function:

[0053]

[0054] where, and are trainable parameters.

[0055] Further, the early warning threshold includes individual threshold, regional threshold, trend threshold and range threshold;

[0056] When the comprehensive risk value is greater than or equal to the individual threshold and the diffusion radius is less than or equal to the range threshold, the individual level immediate early warning is started;

[0057] When the comprehensive risk value is greater than or equal to the regional threshold and less than the individual threshold, or the diffusion radius is greater than the range threshold, the regional level situation early warning is started;

[0058] When the comprehensive risk value is greater than or equal to the trend threshold and less than the regional threshold, and the spatio-temporal attribute shows that the risk event has the potential for persistence or diffusion, the trend level deduction early warning is started.

[0059] Compared with the prior art, the advantages of the present application are:

[0060] 1. The present application integrates multi-source heterogeneous data processing and feature extraction through a large model cluster of text, vision and space-time adjusted in the police field, and adopts a weighted fusion algorithm based on attention mechanism, effectively solving the problem of data isolation and insufficient feature extraction in traditional methods, and significantly improving the risk identification accuracy;

[0061] 2、The present application realizes the whole cycle coverage from real-time disposal to macroscopic deduction of police risk by dynamically triggering early warning of different granularity such as individual level, regional level and trend level according to the comprehensive risk value and its space-time attribute, and enhances the pertinence and foresight of risk response;

[0062] 3、The present application realizes the space-time diffusion law and evolution trend of risk events by the structure of space-time graph convolution network, and the system can effectively capture the space-time diffusion law and evolution trend of risk events, support the quantitative prediction of risk influence range, duration and other attributes, and thus improve the perception and response ability of risk dynamics in complex police scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0064] Figure 1 The system workflow schematic diagram of the present application;

[0065] Figure 2 The large model cluster schematic diagram of the present application;

[0066] Figure 3 The comprehensive risk value generation process schematic diagram of the present application. DETAILED DESCRIPTION

[0067] In order to achieve the above purpose, the present application is realized by the following technical solutions, and the present application provides a police risk multi-level early warning system based on large model collaborative computing, please refer to Figure 1 The system comprises:

[0068] Data module: real-time acquisition of police data, pre-processing of police data, conversion of pre-processed police data into initial vector, and storage of initial vector in preset vector database; setting event detector for real-time monitoring of pre-processed police data, and identifying initial risk event based on preset risk rules to generate early warning task;

[0069] The specific operation of real-time acquisition of police data is as follows:

[0070] The system is connected and real-time accesses to multiple police data sources through the distributed data acquisition agent deployed in the police private network, and the data sources include but are not limited to:

[0071] Police call handling system: real-time streaming 110 or 122 alarm records, police situation description text, alarm person information, preliminary disposal feedback;

[0072] Video monitoring platform: real-time video stream, snapshot images, and structured event snapshots of key areas, such as intelligent analysis and preliminary screening results of fights, crowd gathering, etc.

[0073] Mobile police terminal: real-time text and image reports, positioning information, and on-site audio recordings after text conversion from on-site law enforcement personnel;

[0074] Population, vehicle, and case information database: structured data for on-demand queries or subscription updates, such as personnel files, information on vehicles involved in cases, information on fugitives, and summaries of historical case files;

[0075] Internet of Things sensing devices: such as toll gate vehicle data, WiFi probe data (anonymous MAC addresses), and important infrastructure sensor alert information;

[0076] Internet public opinion monitoring system: integrates network text information related to local public security, public safety, and sensitive events after preliminary screening.

[0077] The collection process emphasizes low latency and high concurrency to ensure that information can be timely processed by the system;

[0078] The collection interface of police data complies with pre-set security protocols and authorization mechanisms.

[0079] Preprocessing includes data desensitization, format unification, and noise removal;

[0080] The specific operations of data desensitization are as follows:

[0081] Text data: For fields in structured text data that explicitly contain personal identity information and sensitive personal information, use pre-defined rules for matching and replacement, such as keeping the first three and last four digits of a mobile phone number and filling in the middle with ****;

[0082] For entities that need to be associated and anonymized, use irreversible encryption hash functions to generate unique and anonymous identifiers, ensuring consistent anonymous representation of the same entity in different records; use data masking techniques, such as keeping only the surname Zhang* or using generalized words;

[0083] For unstructured text data, use named entity recognition (NER) models to identify sensitive information in the text. The identified entities are standardized, replaced with characters, partially masked, or completely deleted according to their sensitivity level;

[0084] Image or video data: Apply deep learning-based detection models to locate sensitive areas and perform pixelization, blurring, or covering mask processing to ensure that the processed image or video frame retains the overall context information of the scene while being unable to identify individual identities. Detection models include face detection and license plate detection, etc.

[0085] Desensitization principle: follow the "minimum necessary principle" and "data minimization principle", only desensitize necessary sensitive information, retain the core value of data for risk analysis, all desensitization rules must comply with the mandatory specifications of the national and public security system on police data security;

[0086] The format is uniformly converted to a standardized technology stack to achieve consistency of multi-source data, the timestamp is unified to ISO 8601 format (including time zone), the geographic coordinate is normalized to WGS84 coordinate system, the text encoding is forcibly converted to UTF-8 and the basic segmentation, the image and video are unified resolution and color space, the numerical data is standardized or normalized;

[0087] Noise removal cleans up noise through a pre-set rule base combined with statistical methods, missing values are selected to be ignored, marked or filled according to field importance, ensuring data reliability, statistical methods include outlier detection and data smoothing, etc.

[0088] Real-time statistics pre-process the indicators of each step, set threshold to trigger alarm.

[0089] The initial vector includes a text initial vector, a visual initial vector and a space-time initial vector;

[0090] The specific operation of converting police data into an initial vector is as follows:

[0091] The text initial vector adopts a BERT model adjusted by police corpus (such as case description, legal documents) to extract context-aware embedding vectors for text data. Let the text input be The text initial vector is represented as:

[0092]

[0093] Wherein, is the dimension of the text initial vector;

[0094] The visual initial vector adopts a visual encoder based on ResNet-50 architecture to extract multi-layer convolutional features for image data. Let the image input be The visual initial vector is represented as:

[0095]

[0096] Wherein, is the dimension of the visual initial vector;

[0097] The space-time initial vector extracts time period features and space grid encoding from space-time data, and maps it to a space-time initial vector through a multi-layer perception (MLP):

[0098]

[0099] where, is the dimension of the spatiotemporal initial vector;

[0100] All initial vectors are organized by time window and data source identification and stored in a vector database (such as Milvus or Weaviate) and a density-based index (HNSW) is established to support subsequent similarity retrieval.

[0101] Event detector and early warning task generation:

[0102] The event detector performs preliminary identification of risk events based on a set of configurable preset rules, which are usually relatively simple, clear, and low-computational-expense logical judgments, such as:

[0103] Keyword or phrase matching: specific keywords such as "armed", "fighting", and "gathering" appear in the alarm text; specific sensitive place and negative sentiment word combinations appear in network public opinion;

[0104] Pattern matching: multiple similar minor alarms occur in the same area within a certain time; specific personnel appear in the control area;

[0105] Threshold alarm: the real-time crowd density in a specific area exceeds the safety threshold; the alarm state of a key infrastructure sensor lasts for more than a certain time;

[0106] Simple model inference: combine preset rules and lightweight statistical models to calculate the probability of risk event occurrence, which can be represented as a rule matching item and a prior probability function:

[0107] P( ∣D)∝P(D∣ )×P( )

[0108] where P( ∣D) is the posterior probability, based on the probability of the event occurrence, P(D∣ ) is the likelihood function, representing the conditional probability of observing the current data D given the event occurs, and P( ) is the prior probability, representing the initial probability of the event occurrence before observing the data;

[0109] When the probability of the risk event exceeds the preset risk threshold, it is determined to be an initial risk event. The risk threshold is set based on the event risk tolerance, historical data and experience, the cost balance of false positives and false negatives, and industry standards and compliance requirements;

[0110] When the event detector identifies an initial risk event, a structured early warning task is generated, which at least includes an event unique identifier, an event type, an event occurrence time, an event associated core data index, and a preliminary description and key features of the event.

[0111] The computing module sets a collaborative scheduling hub and a large model cluster; the collaborative scheduling hub generates a subtask by analyzing the early warning task, and retrieves and calls a corresponding initial vector from the vector database according to the subtask; the large model cluster includes a text analysis model, a visual analysis model, and a space-time prediction model adjusted for police domain data, and the large model cluster processes the initial vector to generate a feature vector and a space-time attribute;

[0112] The subtask includes a text processing requirement, a visual processing requirement, and a space-time processing requirement.

[0113] The feature vector includes a text feature vector, a visual feature vector, and a space-time feature vector.

[0114] The collaborative scheduling hub first performs in-depth analysis on the task, decomposes specific data processing requirements according to the event type, preliminary description, and event associated core data index of the early warning task, and classifies these requirements into three types of subtasks: text processing requirements, visual processing requirements, and space-time processing requirements.

[0115] Subsequently, for each subtask, the scheduling hub generates a structured data retrieval instruction that clearly specifies the data type, time range, and space range of the required initial vector and other key retrieval conditions. This instruction is sent to the vector database for similarity retrieval and calling through HNSW to obtain the most relevant text initial vector, visual initial vector, and space-time initial vector for the current early warning task.

[0116] Finally, the scheduling hub precisely allocates the retrieved initial vector set and corresponding processing requirements to the corresponding professional models in the large model cluster for processing. Please refer to Figure 2 ;

[0117] The text analysis model is built through a Transformer-XL architecture, which specifically includes an input layer, a multi-head attention layer, a semantic reinforcement layer, and an output layer.

[0118] The text initial vector is input into the input layer for domain adaptive embedding, which is specifically operated as follows:

[0119] Align the general semantic space to the police professional domain to enhance the ability to identify police terminology and obtain an embedding vector aligned with the police domain. The formula is as follows:

[0120]

[0121] wherein, and are the police terminology mapping matrix and bias term, respectively, are model trainable parameters, is the text initial vector, and LayerNorm is the layer normalization operation;

[0122] The multi-head attention layer captures long-distance semantic dependencies and event time series relationships by calculating 12 heads of relative position attention, and obtains attention features, as follows:

[0123]

[0124] wherein, , and are the query matrix, key matrix and value matrix of the i-th head, respectively, are model trainable parameters, is the relative position encoding matrix of the i-th head and the j-th head, which is generated according to the police event time difference, is the number of attention heads, which is a preset hyperparameter, is the dimension of the key or query vector of the embedding vector ; The key or query vector of the embedding vector is generated by linear transformation on the embedding vector ;

[0125] The key or query vector of the embedding vector is generated by linear transformation on the embedding vector ;

[0126] The semantic reinforcement layer fuses the prior knowledge of crime types, enhances the high-risk event recognition capability, reduces the false positive rate, and obtains risk features :

[0127]

[0128] wherein, is the crime type vector, which is obtained by encoding the police knowledge base, is the feature fusion matrix, which is a model trainable parameter, is the Gaussian error linear unit activation function;

[0129] The output layer fuses the attention features and the risk features to generate a text feature vector :

[0130]

[0131] wherein, is the dimension transformation matrix, which is a model trainable parameter, is the hyperbolic tangent activation function, which is used to limit the range of feature values.

[0132] The visual analysis model is constructed using a multi-scale attention convolutional network, which specifically includes an input layer, a spatial pyramid layer, an attention layer, and an output layer.

[0133] The initial visual vector is input into the input layer and then reorganized to enhance key spatial features, suppress background interference, and obtain an optimized feature map Z.

[0134]

[0135] in, for The convolution kernel is a set of trainable parameters for the model. As the initial visual vector, To modify the activation function of the linear unit;

[0136] The spatial pyramid layer performs multi-scale feature fusion on the optimized feature map, capturing behavioral patterns at different granularities, such as individual actions and group behaviors, to obtain multi-scale fused features. :

[0137]

[0138] For the average pooling of the s-th downsampling, This is a cross-scale feature splicing operation. This is the preset downsampling ratio;

[0139] The attention layer focuses on areas of abnormal behavior (such as the center of a fight) to obtain a behavioral attention map. :

[0140]

[0141] in, The mask is for human pose estimation, obtained through a pre-trained model; ⊙ represents element-wise multiplication. As the activation function, generate 0-1 attention weights. for Convolution kernel;

[0142] The output layer performs channel weighting and dimension normalization to generate standardized visual feature vectors. :

[0143]

[0144] in, Weighted operations for the channel dimension, Here, is the dimensionality transformation matrix, and are the trainable parameters of the model. This is a channel-level normalization operation.

[0145] The spatio-temporal prediction model is constructed by a spatio-temporal graph convolution network, and specifically includes an input layer, a propagation layer, an influence layer and an output layer;

[0146] The spatio-temporal initial vector is input to the input layer to construct a dynamic spatio-temporal graph :

[0147]

[0148] wherein, is a node set, and a risk-aware entity at a minimum granularity is taken as a node, a single public security camera, a coverage radius of 80 meters, is an edge set;

[0149] The propagation layer alternately performs graph convolution and time series modeling, simulates a risk diffusion process and captures an event evolution rule, and obtains a spatio-temporal state matrix ;

[0150] The influence layer calculates a diffusion radius , predicts a spatial influence range of a risk event, and the calculation formula is as follows:

[0151]

[0152] wherein, is a core area node index set, is a spatial grid node set derived from an event-associated core data index through spatial mapping, and MLP is a multi-layer perception operation, extracts maximum risk features of the core area, is a node of the spatio-temporal graph is a feature vector of the node of the spatio-temporal graph is a kth row of the spatio-temporal state matrix

[0153] The output layer performs graph feature pooling and spatio-temporal attribute analysis to generate a spatio-temporal feature vector and a spatio-temporal attribute triple :

[0154]

[0155] wherein, is a core position, longitude and latitude, is obtained by weighted average of node coordinates of the spatio-temporal graph, t is a duration of the risk event, and is a future prediction value, is a graph-level pooling operation;

[0156] The dimensions of the text feature vector, the visual feature vector and the spatio-temporal feature vector are unified through a preset first full connection layer.

[0157] The early warning module calculates a comprehensive risk value based on the feature vector through a weighting fusion algorithm based on an attention mechanism, and generates multi-level early warnings including individual-level instant early warning, regional-level situation early warning and trend-level deduction early warning according to a warning threshold of the comprehensive risk value and the spatio-temporal attribute;

[0158] The comprehensive risk value is calculated as follows, please refer to Figure 3 :

[0159] 1) The feature vector is passed through a preset attention perception layer to generate query vectors, key vectors and value vectors of text feature vectors, visual feature vectors and spatio-temporal feature vectors respectively:

[0160]

[0161] Among them, , and are query vectors, key vectors and value vectors respectively, is the feature vector, represent text, vision and space respectively, , , , , and are trainable parameters;

[0162] 2) The attention score of the feature vector is calculated by scaling dot product attention :

[0163]

[0164] Among them, is the transpose of the key vector, used to calculate the dot product, is the dimension of the key vector;

[0165] 3) The value vector is weighted based on the attention score to generate weighted modal features ;

[0166] 4) The comprehensive risk value is calculated by the modal features aggregated by all weighted features through a preset second fully connected layer and activation function:

[0167]

[0168] Among them, is the activation function, which compresses the output to the interval.

[0169] After obtaining the comprehensive risk value, a spatio-temporal attribute triple and a preset early warning threshold and spatio-temporal attribute, the early warning threshold is divided into an individual threshold, a regional threshold, a trend threshold and a range threshold, and the following multi-level early warning is generated;

[0170] Individual-level immediate early warning: when the comprehensive risk value is greater than or equal to the individual threshold and the event space range is less than or equal to the range threshold, the individual-level immediate early warning is started;

[0171] In this embodiment, the early warning content includes a target identity and a real-time position, and is directly pushed to front-line law enforcement personnel through a police terminal, a mobile application and the like;

[0172] Regional-level situation early warning: when the comprehensive risk value is greater than or equal to the regional threshold and less than the individual threshold, or the event space range is greater than the range threshold, the regional-level situation early warning is started;

[0173] In this embodiment, the regional-level situation early warning not only includes the geographical range of the risk region, but also integrates real-time situation information such as crowd density, traffic conditions and police deployment, and is pushed to a command center or a district police station for resource scheduling and regional control;

[0174] Trend-level deduction early warning: when the comprehensive risk value is greater than or equal to the trend threshold and less than the regional threshold and the spatio-temporal attribute shows that the risk event has a potential for continuation or diffusion, the trend-level deduction early warning is started;

[0175] In this embodiment, the trend-level deduction early warning does not focus on immediate disposal, but predicts the risk evolution path based on historical patterns and current characteristics, generates a periodic report, and assists the strategic decision-making department in formulating a long-term prevention and control strategy;

[0176] The setting of the early warning threshold is based on historical data verification, industry standards and compliance requirements, expert experience and scene adaptation, and system target orientation.

[0177] In summary, the present application uses a special large model cluster adjusted by police field data to process corresponding data respectively, and innovatively uses a weighted fusion algorithm based on an attention mechanism to calculate a comprehensive risk value, effectively solving the problem of isolated multi-source heterogeneous data and insufficient feature extraction in traditional methods, thereby significantly improving the accuracy of risk identification and reducing the false positive and false negative rates.

[0178] Finally: the above only describes the preferred embodiments of the present application and is not used to limit the present application, and any modifications, equivalent replacements, improvements and the like made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A multi-level early warning system for police risks based on large-scale model collaborative computation, characterized in that, include: Data module: collects police data in real time, preprocesses the police data, converts the preprocessed police data into initial vectors, and stores the initial vectors in a preset vector database; An event detector is set up to monitor the pre-processed police data in real time and identify initial risk events based on preset risk rules to generate early warning tasks; Computation module: Sets up a collaborative scheduling hub and a large model cluster; The collaborative scheduling center generates sub-tasks by parsing the early warning task, and retrieves and calls the corresponding initial vector from the vector database according to the sub-task; the large model cluster includes a text analysis model, a visual analysis model and a spatiotemporal prediction model adjusted by police domain data, and the large model cluster processes the initial vector to generate feature vectors and spatiotemporal attributes; The early warning module calculates a comprehensive risk value based on feature vectors using a weighted fusion algorithm based on an attention mechanism. Based on the early warning threshold of the comprehensive risk value and the spatiotemporal attributes, it generates a multi-level early warning system, including individual-level real-time early warning, regional-level situational early warning, and trend-level deductive early warning.

2. The system according to claim 1, characterized in that, The initial vectors include text initial vectors, visual initial vectors, and spatiotemporal initial vectors.

3. The system according to claim 2, characterized in that, The initial text vector is obtained by using the BERT model to perform context-aware acquisition of the text data. The visual initial vector is obtained by extracting multi-layer convolutional features from the image data using a visual encoder based on the ResNet-50 architecture. The spatiotemporal initial vector is obtained by extracting time period features and spatial grid encoding from spatiotemporal data, and then mapping it through a multilayer perceptron. The initial vectors are organized according to a preset time window and data source identifier, stored in the vector database, and a density-based index is established to support similarity retrieval.

4. The system according to claim 3, characterized in that, The event detector identifies the initial risk event based on a set of configurable preset rules; When the event detector identifies an initial risk event, it generates the early warning task. The early warning task includes at least a unique event identifier, event type, event occurrence time, core data index associated with the event, and a preliminary description and key features of the event.

5. The system according to claim 4, characterized in that, The sub-tasks include text processing requirements, visual processing requirements, and spatiotemporal processing requirements. The collaborative scheduling center decomposes specific data processing requirements based on the event type, preliminary description, and core data index of the event association of the early warning task, and classifies these requirements into three sub-tasks: text processing requirements, visual processing requirements, and spatiotemporal processing requirements. The collaborative scheduling center will assign the retrieved initial vector and corresponding subtask to the appropriate professional models in the large model cluster for processing.

6. The system according to claim 5, characterized in that, The feature vectors include text feature vectors, visual feature vectors, and spatiotemporal feature vectors; The text analysis model is built using the Transformer-XL architecture, which includes an input layer, a multi-head attention layer, a semantic enhancement layer, and an output layer. The initial text vector is input into the input layer for domain-adaptive embedding, as detailed below. Aligning the general semantic space to the policing domain enhances the ability to recognize police terminology and obtains embedding vectors aligned with the policing domain. ; The multi-head attention layer calculates the relative positional attention of 12 heads to capture long-distance semantic dependencies and temporal relationships of alarm events, thereby obtaining attention features. The formula is as follows: in, , and The first The query matrix, key matrix, and value matrix of the header. For the first Head and First Head relative position encoding matrix, For the number of attention heads, The dimension of the key or query vector; The semantic enhancement layer integrates prior knowledge of crime types to obtain risk characteristics; The output layer integrates attention features and risk features, as well as text feature vectors.

7. The system according to claim 6, characterized in that, The visual analysis model is constructed using a multi-scale attention convolutional network, specifically including an input layer, a spatial pyramid layer, an attention layer, and an output layer. The initial visual vector is input into the input layer, and the optimized feature map Z is obtained by reconstructing the initial visual vector. in, for Convolution kernel, As the initial visual vector, To modify the activation function of the linear unit; The spatial pyramid layer performs multi-scale feature fusion on the optimized feature map, capturing behavioral patterns of different granularities to obtain multi-scale fused features. The attention layer focuses on regions of abnormal behavior to obtain a behavioral attention map; The output layer performs channel weighting and dimension normalization to generate standardized visual feature vectors.

8. The system according to claim 7, characterized in that, The spatiotemporal prediction model is constructed using a spatiotemporal graph convolutional network, specifically including an input layer, a propagation layer, an influence layer, and an output layer. The spatiotemporal initial vector is input into the input layer to construct a dynamic spatiotemporal graph; The propagation layer alternately performs graph convolution and temporal modeling to simulate the risk diffusion process and capture the evolution patterns of events, thereby obtaining a spatiotemporal state matrix. ; The influence layer is used to calculate the diffusion radius. Predict the spatial impact range of risk events; The output layer performs graph feature pooling and spatiotemporal attribute parsing to generate spatiotemporal feature vectors. and spacetime attribute triples : in, The core location, latitude and longitude, and t represent the duration of the risk event. This is a graph-level pooling operation; The dimensions of the text feature vector, visual feature vector, and spatiotemporal feature vector are unified by a preset first fully connected layer.

9. The system according to claim 8, characterized in that, The comprehensive risk value The calculation steps are as follows: 1) Generate the query vector, key vector, and value vector of the feature vector through a preset attention perception layer: in, , and These are the query vector, key vector, and value vector, respectively. For feature vectors, Representing text, visual, and time and space respectively. , , , , and These are trainable parameters; 2) Calculate the attention score of the feature vector by scaling the dot product attention. : in, This is the transpose of the key vector. The dimension of the key vector; 3) Based on the attention score For the value vector Weighting is performed to generate modal features. ; 4) Overall Risk Value By using a pre-defined second fully connected layer and The activation function is calculated as follows: in, and These are trainable parameters.

10. The system according to claim 9, characterized in that, The warning thresholds include individual thresholds, regional thresholds, trend thresholds, and range thresholds; The individual-level real-time early warning is activated when the overall risk value is greater than or equal to the individual threshold and the diffusion radius is less than or equal to the range threshold. When the overall risk value is greater than or equal to the regional threshold and less than the individual threshold, or when the diffusion radius is greater than the range threshold, a regional-level situation warning is initiated. When the overall risk value is greater than or equal to the trend threshold and less than the regional threshold, and the spatiotemporal attributes indicate that the risk event has the potential to continue or spread, a trend-level inference warning is initiated.

Citation Information

Patent Citations

  • Internet information auditing system based on artificial intelligence

    CN118035928A

  • Multi-dimensional spatio-temporal data processing method and system for police security

    CN118071100A

  • Internet public opinion intelligent intervention system and method based on big data

    CN120410261A

  • Multi-modal public opinion risk early warning system and method based on dynamic mapping knowledge domain and federal reinforcement learning

    CN120611971A

  • Methods and systems for anomaly and pattern detection of unstructured big data

    US20230186120A1