Alarm processing method and device
By grouping and filtering police report texts from the police information system and classifying them using a large language model, combined with knowledge graphs and hybrid retrieval, the problems of massive data processing and dynamic tag adaptation in the police information system were solved, thereby improving the efficiency of police report processing and retrieval capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG ESHORE TECH
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-26
AI Technical Summary
Police integrated systems face the challenge of processing massive amounts of data. Traditional manual classification is inefficient and cannot meet the needs of dynamically changing classification labels. Furthermore, traditional retrieval functions cannot respond to the automated mining of natural language descriptions and potential relationships, resulting in low efficiency in the analysis of related cases.
By grouping and filtering police incident texts in the police database to identify high-risk and important incident texts, using a large language model for classification and constructing a knowledge graph, and combining keyword retrieval and semantic retrieval for hybrid retrieval, natural language description query and association analysis can be achieved.
It significantly improves the efficiency of police incident handling and classification accuracy, can quickly adapt to dynamic tag requirements, and improves retrieval efficiency and case linkage analysis capabilities.
Smart Images

Figure CN122086949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to methods and devices for handling police incidents. Background Technology
[0002] In related technologies, police integrated systems have long faced the challenge of processing massive amounts of data. Simultaneously, with the development of information technology, the demand for police incident collection and analysis is increasing. On the one hand, the system adds a large amount of new police incident data daily (such as case descriptions and alarm records), and historical data has accumulated to tens of millions of records. Traditional manual classification methods rely on police experience, are inefficient, and struggle to cope with dynamically changing classification tag requirements (such as adding subcategories like "telecom fraud - online order fraud"). On the other hand, traditional search functions only support queries based on structured fields such as time, location, and type, failing to respond to natural language descriptions such as "finding historical police incidents related to 'nighttime car window smashing and theft'," and lack automated methods for mining potential relationships between police incident data (such as similar modus operandi, suspect association, etc.), resulting in low efficiency in analyzing related cases. Summary of the Invention
[0003] To address or partially address the problems existing in related technologies, this application provides a method and apparatus for handling police incidents, which can overcome the technical deficiencies of related technologies in terms of classification efficiency, association analysis, and semantic retrieval.
[0004] The first aspect of this application provides a method for handling police incidents, including: After grouping multiple police incident texts in the police database, high-risk and important police incident texts are selected from each group. After classifying multiple important police incident texts using a large language model, the relationships between multiple important police incident texts in each category are mined, and a corresponding knowledge graph is constructed for at least two important police incident texts in each category that have the relationships. When a query request described in natural language is received, a hybrid search is performed in the police database to obtain the target police information text that matches the query request; the hybrid search includes keyword search and semantic search. The knowledge graph is used to query associated police texts related to the target police text in order to perform case linkage analysis on the target police text and the associated police texts.
[0005] A second aspect of this application provides an alarm processing device, comprising: The filtering module is used to group multiple police incident texts in the police database and then filter out high-risk and important police incident texts from each group. The construction module is used to classify multiple important police incident texts using a large language model, mine the relationships between multiple important police incident texts in each category, and construct a corresponding knowledge graph for at least two important police incident texts in each category that have the relationships. The query module is used to perform a hybrid search in the police database when a query request described in natural language is received, to obtain the target police information text that matches the query request; the hybrid search includes keyword search and semantic search. The analysis module is used to query related police texts associated with the target police text in the knowledge graph, so as to perform case parallel analysis on the target police text and the related police texts.
[0006] A third aspect of this application provides an electronic device, comprising: Processor; and A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.
[0007] A fourth aspect of this application provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described above.
[0008] The technical solution provided in this application may include the following beneficial results: The solution provided in this application groups multiple police incident texts in the police database and then filters out high-risk important police incident texts from each group. After classifying the multiple important police incident texts using a large language model, it mines the relationships between multiple important police incident texts in each category and constructs a corresponding knowledge graph for at least two important police incident texts in each category that have a relationship. When a query request described in natural language is received, a hybrid search is performed in the police database to obtain the target police incident text that matches the query request. The hybrid search includes keyword search and semantic search. The solution also queries the knowledge graph for related police incident texts associated with the target police incident text, enabling case concatenation analysis of the target and related police incident texts. This application dynamically divides massive amounts of police report texts into multiple groups to quickly and accurately filter out important police report texts from small batches after grouping, thereby significantly improving throughput and efficiency while maintaining processing accuracy. Furthermore, large language models perform excellently in natural language understanding tasks; therefore, using large language models to classify important police report texts can improve classification efficiency and accuracy. Further, by integrating a hybrid retrieval architecture combining keyword retrieval and semantic retrieval, it not only balances keyword accuracy and semantic generalization capabilities to meet the needs of natural language description and improve police report retrieval efficiency, but also allows for direct matching and classification based on the semantics of newly added dynamic tags (such as subcategories like "telecom fraud - fake order type") in the police database without requiring manual rule re-engineering or classification system adjustments, quickly adapting to new tag requirements. Finally, by mining the relationships between different important police report texts to construct a corresponding knowledge graph, subsequent queries of related police reports can be conducted through this knowledge graph, thereby improving the efficiency of case analysis.
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0010] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of exemplary embodiments thereof in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments thereof.
[0011] Figure 1 This is a schematic flowchart illustrating the alarm handling method in the embodiments of this application; Figure 2 This is another schematic flowchart illustrating the alarm handling method shown in the embodiments of this application; Figure 3 This is a flowchart illustrating the hybrid retrieval process in an embodiment of this application; Figure 4This is a schematic diagram of the alarm processing device shown in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application. Detailed Implementation
[0012] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.
[0013] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0014] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0015] In related technologies, police integrated systems have long faced the challenge of processing massive amounts of data. Simultaneously, with the development of information technology, the demand for police incident collection and analysis is increasing. On the one hand, the system adds a large amount of new police incident data daily (such as case descriptions and alarm records), and historical data has accumulated to tens of millions of records. Traditional manual classification methods rely on police experience, are inefficient, and struggle to cope with dynamically changing classification tag requirements (such as adding subcategories like "telecom fraud - online order fraud"). On the other hand, traditional search functions only support queries based on structured fields such as time, location, and type, failing to respond to natural language descriptions such as "finding historical police incidents related to 'nighttime car window smashing and theft'," and lack automated methods for mining potential relationships between police incident data (such as similar modus operandi, suspect association, etc.), resulting in low efficiency in analyzing related cases.
[0016] To address the aforementioned issues, this application provides a method for handling police reports. By dynamically dividing massive amounts of police report texts into multiple groups, important police report texts can be quickly and accurately filtered out from the smaller batches of grouped texts. This significantly improves throughput and efficiency while maintaining processing accuracy. Furthermore, large language models perform exceptionally well in natural language understanding tasks; therefore, using large language models to classify important police report texts can improve classification efficiency and accuracy. Further, by integrating a hybrid retrieval architecture combining keyword retrieval and semantic retrieval, not only can keyword accuracy and semantic generalization capabilities be balanced to meet the needs of natural language description and improve police report retrieval efficiency, but also, when faced with newly added dynamic tags (such as adding subcategories like "telecom fraud - fake order type"), the semantics of the tags can be directly matched and classified in the police database without requiring manual re-organization of rules or adjustment of the classification system, quickly adapting to new tag requirements. Furthermore, by mining the relationships between different important police report texts, a corresponding knowledge graph can be constructed. This allows for the retrieval of related police reports through the knowledge graph when querying related reports in the future, thereby improving the efficiency of case analysis.
[0017] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.
[0018] Figure 1 This is a flowchart illustrating the alarm handling method shown in the embodiments of this application.
[0019] See Figure 1 The alarm handling method of this application may include: S110: After grouping multiple police incident texts in the police database, high-risk and important police incident texts are selected from each group.
[0020] In this embodiment, the system can be applied to a police incident intelligent analysis system (hereinafter referred to as the "analysis system"). The analysis system can read massive amounts of recorded police incident texts from the database of the police information system (hereinafter referred to as the "police information database"), and then process the massive amounts of police incident texts using a two-stage processing architecture of "coarse screening followed by fine analysis". In specific implementation, the analysis system can identify the semantic attributes of each police incident text. The semantic attributes can include at least one of the event subject, time, and location. Then, the massive amounts of police incident texts can be grouped according to the semantic attributes. Each group contains a specified number of police incident texts (e.g., 20), but the size can be dynamically adjusted according to the complexity of the texts. For example, if a group contains complex police incident texts, the group can be reduced to 10 texts. Next, the analysis system can filter out ordinary police incident texts and important police incident texts from the specified number of police incident texts in each group. Ordinary police incident texts and important police incident texts are based on the corresponding types and specific descriptions provided by the public security department. Generally speaking, ordinary police incident texts have lower risks and do not require rapid early warning, while important police incident texts have higher risks and require rapid early warning.
[0021] S120 classifies multiple important police reports using a large language model, mines the relationships between multiple important police reports in each category, and constructs a corresponding knowledge graph for at least two important police reports in each category that have a relationship.
[0022] The analysis system is equipped with a Large Language Model (LLM), a deep learning model trained on massive amounts of text data, which possesses natural language understanding and generation capabilities. Therefore, this embodiment can utilize the LLM to complete text classification tasks. To avoid the problems of wasted computing resources (such as processing ordinary police reports one by one) and insufficient real-time response (delayed classification of important police reports) when the LLM is directly applied to police report processing, this embodiment can utilize the LLM to prioritize high-precision classification of multiple important police report texts, and then perform batch classification of multiple ordinary police report texts, thereby avoiding the "resource waste-response delay" paradox that exists when the LLM directly processes police reports.
[0023] The analysis system can automatically mine case entity relationships and construct an iteratively updatable knowledge graph using idle computing power. A knowledge graph is a semantic network structure used to describe entities and their relationships, graphically displaying the correlation between knowledge, and is widely used in search engine optimization, intelligent question answering, recommendation systems, and other fields. In this embodiment, for ordinary police report texts, the analysis system only needs to classify them using a large language model, without constructing a corresponding knowledge graph, thereby reducing the repetitive analysis of ordinary police report texts through batch pre-screening. For important police report texts, after classifying them using a large language model, the analysis system can mine the correlation between multiple important police report texts of the same type, so as to construct a corresponding knowledge graph for at least two important police report texts of the same type that have correlation. The definition of the correlation can closely fit the actual needs of police report analysis, covering three core dimensions: case participation, spatiotemporal correlation, and institutional handling, forming a three-dimensional relationship network, supporting the dynamic adaptation of subsequent addition of relationship types (such as "association of crime tools", "relationship of gang members", etc.).
[0024] S130: When a query request described in natural language is received, a mixed search is performed in the police database to obtain the target police information text that matches the query request; the mixed search includes keyword search and semantic search.
[0025] Users can input queries into the analysis system using natural language descriptions. For example, a query could be: "Find historical police reports related to 'nighttime car window smashing and theft'"; "Case studies of fraud against the elderly around schools"; "Find theft cases involving teenagers around universities in the past week", and so on.
[0026] The analysis system is equipped with a hybrid retrieval system that supports natural language understanding. This system allows for mixed searches within the police database to find target incident texts that match the query request. Specifically, the hybrid retrieval system provides dual search channels (keyword search channel + semantic search channel). The keyword search channel performs keyword searches within the police database, while the semantic search channel performs semantic searches. Therefore, the hybrid retrieval system balances keyword accuracy with semantic generalization capabilities, meeting the needs of natural language description and improving incident retrieval efficiency.
[0027] Furthermore, when faced with newly added dynamic tags (such as new subcategories like "telecom fraud - order brushing"), the hybrid retrieval system can directly match and classify the tags based on their semantics in the police database, without the need for manual re-organization of rules or adjustment of the classification system, thus quickly adapting to new tag requirements.
[0028] S140, query related police texts associated with the target police text in the knowledge graph to perform case concatenation analysis on the target police text and related police texts.
[0029] To uncover deep connections within target police report texts, the analysis system can trigger graph enhancement. The core purpose of graph enhancement is to improve the depth and efficiency of police report correlation analysis. By dynamically adjusting the query scope and priority of the knowledge graph, it achieves a precise balance between business needs and computing resources, aiming to improve the accuracy and response speed of case parallel analysis, especially suitable for scenarios requiring rapid correlation of complex criminal networks. In specific implementation, the analysis system can call the knowledge graph of the target police report text and then expand it based on the key attributes involved in the target police report text. For example, it can dynamically control the query hop count based on urgency coefficient (case type weight), density coefficient (entity association density), and complexity coefficient (number of entities), so that based on the query hop count, it can query related police report texts associated with the target police report text in the called knowledge graph.
[0030] The analysis system can perform case parallel analysis on target and related police report texts. In its implementation, the system generates a structured report containing basic police report information, mixed retrieval scores, and knowledge graph relationships from the target report text. The system also supports visual display of the relationship between the police report list and the knowledge graph, helping users quickly locate key information and significantly improving the efficiency of case parallel analysis.
[0031] As this example demonstrates, the solution provided in this application dynamically divides massive amounts of police report texts into multiple groups, enabling rapid and accurate filtering of important police report texts from small batches of grouped texts. This significantly improves throughput and efficiency while maintaining processing accuracy. Furthermore, large language models perform exceptionally well in natural language understanding tasks; therefore, using large language models to classify important police report texts can improve classification efficiency and accuracy. Moreover, by integrating a hybrid retrieval architecture combining keyword retrieval and semantic retrieval, it not only balances keyword accuracy and semantic generalization capabilities to meet the needs of natural language description and improve police report retrieval efficiency, but also allows for direct matching and classification based on the semantics of newly added dynamic tags (such as subcategories like "telecom fraud - fake order type") in the police database without requiring manual rule re-engineering or classification system adjustments, quickly adapting to new tag requirements. Finally, by mining the relationships between different important police report texts to construct a corresponding knowledge graph, related police reports can be retrieved through this knowledge graph when querying related police reports in the future, thereby improving the efficiency of case analysis.
[0032] Figure 2 This is another flowchart illustrating the alarm handling method shown in this application.
[0033] See Figure 2 The alarm handling method of this application may include: S210: After grouping multiple police incident texts in the police database, high-risk and important police incident texts are selected from each group.
[0034] This step can be found in the description in S110, and will not be repeated here.
[0035] In one implementation, grouping multiple police report texts in the police database may include: Determine the block size and identify the semantic attributes of each police incident text in the police database; wherein the semantic attributes include at least one of the event subject, time and location; divide the multiple police incident texts in the police database into multiple blocks according to the semantic attributes and block size; each block includes a specified number of police incident texts.
[0036] This application provides a packaging-splitting hierarchical processing mechanism, which adopts a two-level processing architecture of "coarse screening followed by fine segmentation": first, a large number of police report texts are packaged and processed in batches using dynamic block segmentation technology to quickly filter out potentially important police report texts; then, these important police report texts are split and handed over to a large language model for accurate classification; finally, the entire process is continuously optimized through closed-loop feedback. This mechanism can improve the processing speed of ordinary police report texts while also achieving real-time response within 500ms for important police report texts, thereby avoiding the "resource waste-response delay" paradox that exists when the large language model directly processes police reports. At the same time, through real-time weight adjustment (forgetting factor) of the rule engine, it ensures that the analysis system can adapt to the identification needs of new crime methods.
[0037] In the coarse screening stage, this application provides a dynamic segmentation strategy, a text structured packaging strategy, and a lightweight rule engine filtering strategy. The dynamic segmentation strategy is used to reduce the computational load, the text structured packaging strategy is used to enable the large language model to clearly distinguish different alarm texts, and the lightweight rule engine filtering strategy is used to quickly identify potentially important alarm texts.
[0038] In the dynamic block segmentation strategy, the analysis system can count the number of words for each alert text separately in order to determine the average number of words per alert (AvgTokenPerCase) and the total number of words in the text (TotalTokens). The block size (BlockSize) is then calculated using the following formula 1: Formula 1 Wherein, BlockSize represents the block size; TotalTokens represents the total number of words in the text; and AvgTokenPerCase represents the average number of words per alert.
[0039] It should be noted that if the alarm text is highly complex, the block size can be reduced as needed (e.g., the block size for complex alarms can be reduced to 10).
[0040] The analysis system identifies the semantic attributes of each police report text. The semantic attributes can include at least one of the event subject, time and location. Then, the massive amount of police report texts can be batch packaged according to the semantic attributes and the block size to obtain multiple blocks. Each block includes a specified number of police report texts (such as 20 or 10).
[0041] In the text structured packaging strategy, in order to ensure that the large language model can distinguish different alarm texts, the analysis system can add a unique identifier (such as [CASE_number]) and an end character ( ) to each alarm text in each block to complete the text structured packaging. For example: [CASE_1] Event description... [CASE_2] Event description... ...
[0042] In one implementation, selecting high-risk, important alert texts from each group may include: Extract key features from each alarm text within each segment; key features include at least one of keyword trigger count, event type probability, and contextual relevance; assign corresponding feature weights to each key feature; calculate the risk score of each alarm text within each segment based on each key feature and its corresponding feature weight; and select important alarm texts from each segment whose risk scores are greater than a preset risk threshold.
[0043] In the lightweight rule engine filtering strategy, the analysis system extracts the key features of each alarm text in each block. The key features may include at least one of the following: keyword trigger count, event type probability, and contextual relevance. Keyword triggers may be, for example, "explosion", "casualties", "emergency", etc. Event type probability may be obtained based on historical data statistics. Contextual relevance may be, for example, "multiple people" or "multiple locations".
[0044] The analysis system can assign corresponding feature weights to each key feature. Then, using key features such as keyword trigger frequency, event type probability, and contextual relevance, along with their corresponding feature weights, a risk score (Score) is calculated for each alert text. The specific calculation formula is shown in Equation 2 below: Formula 2 Wherein, Score represents the risk score; Indicates feature weights; The eigenvalue represents the key feature; n represents the number of features involved in the calculation, i.e., the total number of feature dimensions used to assess the importance of the incident.
[0045] The analysis system can determine whether the risk score of each alert text is greater than a preset risk threshold (e.g., 0.7). If the risk score of an alert text is greater than the preset risk threshold, the analysis system can mark the alert text as a potentially important alert text. It then iterates through each alert text within each block to filter out important alert texts from each block.
[0046] S220 classifies multiple important police reports using a large language model, then mines the relationships between multiple important police reports in each category, and constructs a corresponding knowledge graph for at least two important police reports in each category that have a relationship.
[0047] This step can be found in the description in S120, and will not be repeated here.
[0048] In one implementation, the large language model includes multiple models; classifying multiple important police report texts using the large language model may include: Obtain prompt word templates; input multiple important police report texts and prompt word templates into multiple large language models respectively, so that the multiple large language models can classify the multiple important police report texts, and obtain the structured results of each important police report text output by the multiple large language models under the guidance of prompt word templates; the structured results include classification labels; for the same important police report text, select the classification label with the most occurrences from the classification labels output by the multiple large language models as the target classification label of the important police report text.
[0049] This application's embodiments design a sliding window segmentation algorithm (50 words / window) and a multi-model voting mechanism for important police report texts, enabling real-time response within 500ms. A specially designed Prompt template significantly improves the classification accuracy of the large language model. Specifically, in the fine segmentation stage, this application's embodiments provide a dynamic segmentation strategy, a model input optimization strategy, and a multi-model collaborative classification strategy. The dynamic segmentation strategy is used to split individual police report texts from the packaged blocks; the model input optimization strategy is used to enhance the context of the split individual police report texts; and the multi-model collaborative classification strategy is used to improve the stability (robustness) of the classification results through a multi-model voting mechanism.
[0050] In the dynamic splitting strategy, the analysis system splits multiple police reports within each block into single independent police reports according to semantic boundaries (such as punctuation marks, changes in the subject of the event, and other syntactic structures). Then, it filters out the important police reports with tags so that the sliding window algorithm can be used to segment the important police reports (e.g., segmenting once every 50 words) to ensure that the length of the segmented text is suitable for the input requirements of the large language model.
[0051] In the model input optimization strategy, the analysis system can guide the large language model to output structured results (such as JSON format) through Prompt Engineering. Prompt Engineering is a technique that improves the output quality of generative AI models by optimizing the input prompt. It aims to accurately guide the model to generate content that meets expectations. Its core lies in designing effective instructions that enable the model to accurately understand user intent and reduce erroneous responses. In specific implementation, the analysis system can add a Prompt template before each important alert text, for example: "Please determine if the following alert text belongs to one of the following categories: [Category List]. Alert Content: [Text]". The category list can include, but is not limited to: theft, fraud, robbery, missing persons, fighting, and accidents.
[0052] In the multi-model collaborative classification strategy, the large language model in this embodiment adopts a hybrid model architecture that combines a base model and an enhancement model. The base model (such as a lightweight BERT (Bidirectional Encoder Representations from Transformers)) is used for fast classification, and the enhancement model (such as a fine-tuned qwen3.5 (a large autoregressive pre-trained language model) is used for secondary validation. The final classification result is determined through a multi-model voting mechanism, and the specific determination formula is shown in Equation 3 below: Formula 3 Where FinalLabel represents the final determined target category label; Model1~Model n This indicates multiple models participating in collaborative classification, including both basic models (such as the lightweight BERT) and enhanced models (such as the finely tuned qwen3.5). Each model will independently output its own classification label for the same police report text. Vote indicates a multi-model voting mechanism, which counts the classification labels output by n models and selects the classification label that appears most frequently as the FinalLabel.
[0053] In one example, taking a critical incident text A as an example, the analysis system adds a Prompt template before the critical incident text A: "Please determine if the following incident text belongs to one of the following categories: [Theft, Fraud, Robbery, ...]. Incident content: [Text A]". The analysis system inputs this into n models to obtain structured results from the n models for the critical incident text A. Each structured result contains a classification label, thus yielding n classification labels. The analysis system can select the classification label that appears most frequently from the n classification labels as the target classification label for the critical incident text A. For example, if 3 out of 5 models (i.e., n=5) output the classification label "Theft", and the other 2 models output "Robbery" and "Fraud" respectively, then the analysis system can determine "Theft" as the target classification label for the critical incident text A.
[0054] In one embodiment, the method may further include: Record the classification confidence and processing time of each important police report text, and adjust the block size based on the classification confidence and processing time; and determine the sample set to be updated, so as to update the feature weights using the sample set to be updated; wherein, the adjusted block size is used to balance classification accuracy and classification efficiency, and the updated feature weights are used to balance historical rule experience and new sample patterns.
[0055] The embodiments of this application can enable the analysis system to have continuous self-optimization capabilities by constructing a closed-loop control system of "monitoring-analysis-decision". Its core lies in feeding back key indicators (such as classification confidence, response latency, rule triggering rate, etc.) in the processing process to the strategy engine in real time, dynamically adjusting system parameters through preset optimization algorithms, and dynamically adjusting block strategy and filtering strategy through feedback mechanism.
[0056] To optimize the block segmentation strategy, the analysis system can record the classification confidence level and processing time of each important alarm text in real time, thereby constructing a feedback database. Based on the feedback data (i.e., classification confidence level and processing time), the block size (BlockSize) can be adjusted. The specific adjustment formula is shown in Equation 4 below: Formula 4 Wherein, NewBlockSize represents the adjusted block size; BlockSize represents the original block size (i.e., the block size in Equation 1 above); Accuracy represents the classification accuracy of the current block segmentation strategy, which is determined based on classification confidence and processing time; AvgAccuracy represents the historical average classification accuracy.
[0057] When Accuracy > AvgAccuracy, it means that the current block-splitting strategy is effective, and the block size can be increased to improve batch processing efficiency; when Accuracy ≤ AvgAccuracy, it means that the current block-splitting strategy is ineffective, and the block size can be reduced to ensure classification accuracy.
[0058] To optimize the screening strategy, the analysis system can dynamically update the screening rules using online learning algorithms (such as random forests), supporting incremental learning where new data is gradually integrated into the model without requiring full retraining. In its implementation, the analysis system can extract alert texts that meet the update trigger conditions from the feedback database, serving as the update sample set for the rule engine. This ensures the relevance and effectiveness of the samples. The update trigger conditions may include, but are not limited to: the percentage of alert texts with a classification confidence score < a preset classification threshold (e.g., 0.6) within a block ≥ a first percentage threshold (e.g., 15%); classification accuracy below AvgAccuracy for three consecutive batches exceeding a second percentage threshold (e.g., 10%); and the addition of new alert texts that do not match the current rule (e.g., new fraud methods). The analysis system can use the update sample set to adjust feature weights. The update is performed using the formula shown in Equation 5 below: Formula 5 Among them, W new This represents the updated feature weights; As a forgetting factor, it can balance historical experience and new knowledge; W old This represents the feature weights before the update (i.e., the feature weights in Equation 2 above). (old value); This represents the change in feature weights calculated based on the new label data.
[0059] As can be seen, the adjusted block size NewBlockSize can be used to balance classification accuracy and classification efficiency, and the updated feature weights W new It can be used to balance historical rules and experience with the patterns of new samples.
[0060] In one implementation, mining the relationships between multiple important police report texts of various types may include: For multiple important police reports with the same target classification label, identify the entities of each important police report; the entities include at least one of the following: case type, time, location, person, institution name, and importance level; for multiple important police reports with the same target classification label, based on the entities, mine the relationships between the multiple important police reports; the relationships include at least one of the following: case participation relationship, spatiotemporal relationship, and institutional handling relationship.
[0061] In one implementation, the knowledge graph includes a dedicated knowledge graph; constructing a corresponding knowledge graph for at least two important police report texts that are related across different categories may include: For multiple important police reports with the same target classification label, construct a corresponding exclusive knowledge graph for at least two important police reports that have any correlation relationship.
[0062] In one embodiment, the knowledge graph further includes a global knowledge graph, and the method may further include: For multiple important police reports with different target classification labels, construct a corresponding global knowledge graph for at least two important police reports that have any correlation.
[0063] This application provides an automated knowledge graph construction scheme for important police report texts. The core process includes: data preprocessing, relationship modeling, dynamic update mechanism, and graph storage, thereby realizing the automated mining and real-time updating of police report entity relationships and providing structured knowledge support for case linkage analysis.
[0064] In the data preprocessing stage, the analysis system can input massive amounts of police report text into a large language model for entity annotation. Entities can include at least one of the following: case type, time, location, person, organization name, and importance level. The analysis system can use WordPiece (a sub-word segmentation algorithm) to segment each police report text, generating a standardized segmented word sequence Tokenized_text that begins with [CLS] and ends with [SEP]: Tokenized_text = WordPiece([CLS] + police report text + [SEP]). Then, the analysis system transforms the segmented word sequence Tokenized_text into a 768-dimensional word vector matrix X: X = [x1, x2, ..., x...]. t ], (d=768 dimensions), where x1~x t This represents the word vector corresponding to each token after the police report text has been segmented and vectorized. This completes the text vectorization process to adapt to the input format of subsequent models.
[0065] The embodiments of this application can predefine different types of relationships. These relationships can closely match the actual needs of police situation analysis, covering three core dimensions: case participation, spatiotemporal correlation, and institutional handling, forming a three-dimensional relationship network, and supporting the dynamic adaptation of subsequent new relationship types (such as "relationship of crime tools" and "relationship of gang members").
[0066] The relationships can include at least one of the following: case involvement relationship, spatiotemporal relationship, and institutional handling relationship. The case involvement relationship can be: complainant -> involved party -> suspect, used for personnel association analysis; the spatiotemporal relationship can be: case A -> same area -> case B, used for crime hotspot analysis; the institutional handling relationship can be: police station -> receiving the report -> incident, used for police force efficiency analysis.
[0067] The embodiments of this application can initialize the Span-BERT model (joint extraction model) so that entities and relationships can be identified simultaneously through the Span-BERT model in practical applications. The loss function of the Span-BERT model is shown in Equation 6 below: Formula 6 Where L represents the overall loss function, which consists of two parts: entity recognition loss and relation classification loss; N represents the number of samples (i.e., the total number of entity pairs or relation instances in the training data); and i represents the sample index (traversing all samples from 1 to N). Represents the true label (0 or 1) of the i-th sample, used for binary classification tasks of entity recognition (such as determining whether it is an entity boundary or a specific entity type); This represents the probability that the model predicts the i-th sample to be a positive example (i.e., belonging to the target entity or entity relationship); The weight hyperparameter represents the balance between entity recognition loss and relation classification loss, controlling the proportion of their contributions to the total loss; The cross-entropy loss represents the relationship classification.
[0068] In the relationship modeling stage, this embodiment of the application can input multiple important police reports with the same target classification label after vectorization into the above-mentioned Span-BERT model, and simultaneously complete entity boundary recognition and entity pair matching through the Span-BERT model.
[0069] To improve the accuracy of identifying non-explicit relationships, embodiments of this application can enhance relationship prediction using GAT (Graph Attention Network) to aggregate neighbor information of entities. The specific relationship prediction formula is shown in Equation 7 below: Formula 7 Where r represents the type of association; e1 and e2 are the head entity and tail entity, respectively; W r Represents the relation classification weight matrix; h represents the embedding vector of concatenated entities e1 and e2. e1 h e2 This represents the embedding vectors of entities e1 and e2; the Softmax function is used to convert the output into a probability distribution. This represents the probability distribution of the existence of a relationship r between entities e1 and e2.
[0070] Subsequently, this embodiment of the application can proceed to the knowledge fusion and graph construction stage (covering relationship strength calculation and dynamic update mechanism). This stage achieves structured expression and real-time updating of case-related relationships by dynamically fusing multi-source police data and automatically constructing a knowledge graph. In specific implementation, the analysis system can calculate relationship strength scores for the extracted valid entity-relation pairs in order to assign weights to the edges in the knowledge graph. The specific calculation formula is shown in Equation 8 below: Formula 8 in, Represents entity e i and entity e j The relationship strength score is the weight of the edge connecting two entities in the graph. It is a weighting coefficient used to adjust the importance of co-occurrence frequency in the calculation of relationship strength; It is a weighting coefficient used to adjust the importance of semantic similarity in relation strength calculation, satisfying... ; Represents entity e i and entity e j The co-occurrence frequency reflects how frequently two entities appear together in the data; Represents entity e i With entity e j The semantic similarity is calculated using BERT embedding cosine distance, which measures the semantic similarity between two entities.
[0071] In the dynamic update mechanism, adding new alert text triggers a local graph update. The conflict resolution strategy is shown in Equation 9 below: Formula 9 TrustScore represents the trust score, which is used to determine whether newly added alert texts are trustworthy. New alert texts are only integrated when TrustScore > a preset trust threshold (e.g., 0.5) to avoid erroneous data from entering the knowledge graph. The weights (0.1-0.9) represent the source credibility of newly added police alert texts; SourceCred represents the source credibility of newly added police alert texts; ModelConf represents the model confidence score, which refers to the reliability of the model's prediction or processing results for newly added police alert texts.
[0072] The analysis system can construct a preliminary knowledge graph based on the fused entities (nodes), relationships (edges), and relationship strength (edge weights) according to graph structure rules. Then, it defines basic graph attributes, such as nodes carrying entity type and key attribute values (e.g., location nodes include latitude and longitude, name nodes include identity information), and edges carrying relationship strength scores and associated timestamps. The analysis system can employ a sliding window batch processing strategy to advance the construction, processing 100 police report texts per batch (i.e., window size = 100 police reports / batch), with a window overlap rate set to 15% to avoid missing boundary entities.
[0073] During the graph storage phase, Neo4j is naturally well-suited for efficient querying and association analysis of graph structures, so the analysis system can choose Neo4j as the graph database to store the constructed knowledge graph. Then, the analysis system can use Neo4j's BATCH transaction commit mode to complete the data import, reducing network interaction overhead and improving the storage efficiency of large amounts of data. Next, the analysis system performs global optimization on the imported knowledge graph, such as checking the integrity of node-edge associations, correcting abnormal weights, and establishing entity attribute indexes (e.g., indexes by case type, location, and time), improving subsequent query speed.
[0074] Furthermore, this application provides an entity disambiguation strategy to eliminate duplicate nodes and avoid entity redundancy and confusion in the knowledge graph. In specific implementation, for entities with the same name or similar meanings, the analysis system can combine textual context features (such as entity type) and external knowledge base features (such as standard entity codes provided by the public security department) and use a weighted voting mechanism to determine the final entity: disambiguation score = 0.6 × context similarity + 0.3 × knowledge base matching degree + 0.1 × frequency statistics.
[0075] As can be seen, the embodiments of this application can achieve full lifecycle maintenance of the knowledge graph. Specifically, when new police report text is accessed, the analysis system can trigger only a partial update of the knowledge graph without rebuilding it, thus reducing resource consumption. During the update phase, after the new police report text undergoes preprocessing, entity relationship extraction, knowledge fusion, and trust verification, the analysis system adds the valid entity-relation pairs of the new police report text to the corresponding nodes / edges of the knowledge graph and updates the relationship strength scores of the relevant edges. During the iterative optimization phase, the analysis system can periodically collect graph usage data (query frequency, association analysis accuracy) and reverse-optimize the joint extraction model parameters and entity disambiguation rules to continuously improve the quality of the knowledge graph. The analysis system can dynamically add new types of association relationships (such as gang member relationships, upstream and downstream relationship relationships in cases) through reserved interfaces to adapt to business changes based on new needs in public security operations.
[0076] It should be noted that the knowledge graphs constructed in the embodiments of this application cover both dedicated knowledge graphs of the same category and global knowledge graphs across categories. The dedicated knowledge graphs can solve the problems of business adaptability and analysis efficiency of knowledge graphs, while the global knowledge graphs can solve the problem of cross-category in-depth mining of knowledge graphs.
[0077] S230, when a query request described in natural language is received, the query request is converted into a structured query statement through a large language model, so as to query the target alarm text that matches from the police database using the structured query statement; the target alarm text includes multiple texts.
[0078] This application provides a hybrid retrieval strategy based on the semantic understanding capabilities of a large language model. By integrating structured field retrieval and semantic vector matching, a dual-channel hybrid retrieval system is constructed, which can overcome the limitations of a single retrieval method. Among them, BM25 (Best Match 25) excels at precisely matching keywords, while vector retrieval can capture semantic similarity.
[0079] In its implementation, the analysis system can input natural language-described query requests into a pre-trained SQL (Structured Query Language) generation module or a large language model, so that the SQL generation module or the large language model can automatically convert the query requests into structured query statements. For example, assuming the user inputs a query request of "cases of fraud against elderly people around schools," the SQL generation module or the large language model can convert it into a structured query statement containing the SQL conditional expression "location LIKE '%school%'" and "case_type='fraud'".
[0080] The analysis system can perform structured queries in the police database to obtain a preliminary screening result set, which contains multiple target police information texts that match the query request.
[0081] As can be seen, structured query statements, as retrieval constraints, can quickly narrow down the retrieval scope through field constraints, improving efficiency by more than 60% compared to full retrieval.
[0082] S240, calculate the keyword matching degree between the structured query statement and each target police report text, and calculate the semantic similarity between the structured query statement and each target police report text.
[0083] After obtaining the initial screening result set, the hybrid retrieval process begins, such as... Figure 3 As shown, the hybrid search process covers both keyword search and semantic search channels.
[0084] In the keyword retrieval channel, the analysis system can use the BM25 algorithm to calculate the keyword matching degree BM25(q,d) between the structured query statement and each target alarm text. The specific calculation formula is shown in Equation 10 below: Formula 10 Wherein, BM25(q,d) represents the keyword matching degree; n represents the number of terms in the query, that is, the total number of keywords contained in the query; k1=1.25, which is a parameter in the BM25 algorithm used to control the degree of influence of the frequency of terms in the document on the retrieval score; b=0.75, which is a parameter in the BM25 algorithm used to control the degree of influence of document length on the relevance score, and will normalize the document length to a certain extent, balance the document length factor, and make the score more reasonably reflect the relevance between the document and the query; Representative terms Frequency of occurrence in document d (i.e., word frequency); Representative query terms Inverse document frequency, used to measure the term Distinctiveness The specific calculation formula is shown in Equation 11 below: Formula 11 Where N is the total number of documents; This represents the number of documents containing the term.
[0085] In the semantic retrieval channel, the analysis system can map the initial screening result set and structured query statements to a 768-dimensional semantic space through the ERNIE3.0 model to generate semantic vectors. Then, the cosine similarity algorithm is used to calculate the vector similarity in order to obtain the semantic similarity SemanticSim(q,d) between the structured query statements and each target police text.
[0086] S250. Based on keyword matching degree and semantic similarity, calculate the mixed retrieval score of each target alarm text, and sort the multiple target alarm texts based on the mixed retrieval score to obtain a list of target alarm texts.
[0087] like Figure 3 As shown, the analysis system can perform weighted fusion calculation of keyword matching degree BM25(q,d) and semantic similarity SemanticSim(q,d) to obtain the hybrid retrieval score HybridScore(q,d) for each target police message text. The specific calculation formula is shown in Equation 12 below: Formula 12 HybridScore(q,d) represents the hybrid retrieval score, which is the final score that combines the keyword matching degree BM25(q,d) and the semantic similarity SemanticSim(q,d), and is used to measure the relevance of the query to the document. It is a weighting coefficient used to adjust the contribution ratio of keyword matching degree BM25(q,d) in the mixed search score, and its value ranges from 0 to 1.
[0088] After obtaining the hybrid retrieval score HybridScore(q,d) for each target alert text, the analysis system can sort the multiple target alert texts in descending order of their hybrid retrieval scores HybridScore(q,d) to obtain a list of target alert texts.
[0089] S260, based on the ranking of each target alarm text in the target alarm text list, determine the urgency coefficient and entity authority of each target alarm text respectively.
[0090] This application provides a knowledge graph enhancement strategy, the core purpose of which is to improve the depth and efficiency of police incident correlation analysis. By dynamically adjusting the query scope and priority of the knowledge graph, a balance between precise matching of business needs and computing resources can be achieved. This aims to improve the accuracy and response speed of police incident parallel analysis, and is especially suitable for scenarios that require rapid correlation of complex criminal networks.
[0091] In its implementation, the analysis system can determine the urgency coefficient and entity authority of each target alert text based on its ranking in the target alert text list. The urgency coefficient is mapped through business rules; for example, target alert texts ranked higher (e.g., Top 1-Top 5) are considered high-priority queries, and their urgency coefficient is closer to 1; target alert texts ranked lower are considered regular queries, and their urgency coefficient is closer to 0. The entity authority is derived from the confidence score of graph nodes; for example, core entities in high-ranking target alert texts (such as fugitive suspects or key accounts involved in the case) are assigned higher authority scores. Thus, through dynamic incremental adjustment, both in-depth analysis of high-ranking target alert texts and control of resource consumption for low-ranking target alert texts are ensured.
[0092] S270, obtain the preset hop count threshold, and calculate the maximum hop count for each target alarm text based on the preset hop count threshold, urgency coefficient and entity authority.
[0093] To avoid frequent changes in the jump count benchmark due to ranking fluctuations, this application embodiment can set a preset jump count threshold base_hop (e.g., the default is 2).
[0094] The analysis system can calculate the maximum number of hops for each target alert text based on a preset hop count threshold (base_hop), an urgency coefficient, and entity authority. The specific calculation formula is shown in Equation 13 below: Maximum number of hops = base_hop + (urgency coefficient × 0.5) + (entity authority × 0.3) (Equation 13) High-priority queries (such as telecom fraud warnings) can enable deep expansion (maximum number of hops ≥ 3) to allow penetration of multi-layered relationship networks; regular queries are limited to 1-2 hops to avoid irrelevant entities interfering with the results.
[0095] In one example, suppose the first target alert text A in the target alert text list is a warning of gang telecom fraud (high urgency, high core entity). Then target alert text A belongs to high priority query. The parameter values can be: base_hop = 2 (fixed default value), urgency coefficient = 1 (ranked first, highest priority), entity authority = 1 (core involved account, high confidence). Then the maximum number of hops for target alert text A = 2 + (1×0.5) + (1×0.3) = 2.8 (rounded to 3). Therefore, the correlation query for target alert text A can penetrate multiple layers of relationship network and deeply explore the correlation clues of fraud gangs.
[0096] In another example, suppose target police text B, ranked 20th in the target police text list, is a small-amount theft by a single person (regular priority, ordinary entity). Then target police text B belongs to the regular query, and the parameter values can be: base_hop = 2 (fixed default value), urgency coefficient = 0.2 (lower ranking, low priority), entity authority = 0.3 (ordinary crime scene, low confidence). Then the maximum number of hops for target police text B = 2 + (0.2×0.5) + (0.3×0.3) = 2.19 (rounded to 2). Therefore, the association query for target police text B can only explore the basic clues that are directly related, avoiding irrelevant interference.
[0097] S280: For each target alarm text, starting from the target alarm text, perform an association traversal in the knowledge graph of the target alarm text according to the maximum number of hops of the target alarm text to obtain the associated alarm texts related to the target alarm text.
[0098] The maximum number of hops can be used as a precise execution threshold for knowledge graph association queries, defining the depth boundary of the graph query for each target police report text. This allows the analysis system to strictly follow this value to perform association mining, ultimately achieving a refined and resource-optimized graph query effect where "the higher the police report priority, the deeper the mining; the lower the police report priority, the shallower the mining."
[0099] In the specific implementation, for each target alert text, the analysis system will take the target alert text as the starting node and strictly follow the calculated maximum number of hops to perform a directional and fixed-depth association traversal in the knowledge graph, without skipping hops or exceeding the boundary, so as to obtain the associated alert texts related to the target alert text. Each associated alert text can include at least one of the associated alert texts.
[0100] Taking the target police report text A as an example: if the maximum number of hops for the target police report text A is rounded up to 3, then the analysis system starts from the node of the target police report text A and mines 1 hop (such as directly related suspects / locations / tools) → 2 hops (such as other police reports / gang members related to the suspects) → 3 hops (such as the accounts involved in the case / upstream and downstream cases related to the gang members). After mining 3 hops, it stops immediately and returns all the mining related nodes / relationships, thereby obtaining at least one related police report text associated with the target police report text A.
[0101] Taking the target alert text B as an example: if the maximum number of hops for the target alert text B is rounded up to 2, then the analysis system will only mine 1 hop (such as directly related entities) → 2 hops (such as direct association of entities), and terminate the query after 2 hops without performing deeper traversal, thereby obtaining at least one related alert text associated with the target alert text B.
[0102] As can be seen, by using a dynamic maximum number of hops, we can avoid a "one-size-fits-all" approach to query depth (such as digging 3 hops for all target alert texts). This allows us to dig deeper for clues in high-priority target alert texts while saving computing power for low-priority target alert texts, thus achieving a precise match between query depth and alert value.
[0103] S290 involves performing case parallel analysis on multiple target police report texts and multiple related police report texts.
[0104] The analysis system can perform case parallel analysis on multiple target police report texts and multiple related police report texts. In its specific implementation, the analysis system can generate a structured report containing basic information about each target police report text, mixed search scores, and knowledge graph associations. The analysis system supports the visualization of the relationship between the police report list and the knowledge graph to help users quickly locate key information, thereby greatly improving the efficiency of police report case parallel analysis.
[0105] As can be seen from this example, the solution provided in this application aims to address the technical deficiencies in related technologies regarding classification efficiency, association analysis, and semantic retrieval. (1) Dynamic hierarchical processing mechanism: The embodiments of this application achieve a dynamic balance between batch classification of ordinary police reports and accurate identification of important police reports by dynamic segmentation, multi-level screening and closed-loop optimization. By reducing the repetitive analysis of ordinary police reports through batch pre-screening, it can not only optimize the single processing mode in related technologies into a batch processing mode and solve the problem of low efficiency of the single processing mode in related technologies, but also avoid the limitations of "fixed batch processing" and "single model classification".
[0106] (2) Intelligent knowledge graph construction: This application embodiment utilizes idle computing power to automatically extract entity relationships from police reports through the Span-BERT model, supports real-time association analysis, and automatically mines case entity relationships to construct an iteratively updated knowledge graph, thereby greatly improving the efficiency of police report case serial analysis.
[0107] (3) Hybrid retrieval system: The embodiments of this application provide a hybrid retrieval system that supports natural language understanding (covering keyword detection channel and semantic retrieval channel), which can take into account both keyword accuracy and semantic generalization ability, thereby improving the efficiency of police information retrieval.
[0108] (4) Dynamic label adaptation: This application provides a label management system with dynamic adaptation capabilities, which enables zero-cost migration without retraining the model when the classification system changes, greatly shortening the response time for classification system changes.
[0109] In summary, the solution provided in this application is an intelligent emergency response solution with faster response speed, broader analysis dimensions, and lower operation and maintenance costs.
[0110] To facilitate a clear understanding of the solutions provided in this application, examples of the linkage process and technical implementation, and a scenario for analyzing the joint cases of electric vehicle theft gangs, are provided in this application for illustration: Example 1: The linkage process and technical implementation (including data flow and collaboration mechanisms, dynamic interaction triggering conditions) is based on a linkage mechanism of "knowledge enhancement + reasoning verification". It uses a large language model for initial understanding and classification, then uses a knowledge graph for deep association and verification, and finally uses dynamic interaction triggering conditions to balance efficiency and accuracy. (1) Knowledge enhancement stage: When the large language model processes the police text, it first extracts entities (such as case type, location, time, etc.) and relationships (such as case participation relationship, spatiotemporal relationship, etc.) through the Span-BERT model, and pushes the structured data to the Neo4j knowledge graph in real time. The graph can enhance relationship prediction through graph attention network (GAT) to form a dynamically updated crime network.
[0111] (2) Reasoning and verification stage: After the large language model completes the initial classification of the police report text, it compares the police report classification results (such as theft / fraud) with the historical case patterns in the knowledge graph. If the new case is found to match the crime characteristics (such as the same method, target group) of a gang in the knowledge graph with a degree exceeding the threshold (such as cosine similarity > 0.7), the graph jump number expansion is triggered to deeply explore potential associations.
[0112] (3) Dynamic interaction triggering conditions: For ordinary police texts, only the large language model completes the classification, and the graph does not intervene to ensure processing efficiency; For important police texts: When the classification confidence is <85% or the rule engine score (i.e., risk score) is > the preset risk threshold (e.g., 0.7), the graph is automatically called to perform entity disambiguation and relationship completion (see Equation 9 above for conflict resolution strategy).
[0113] Example 2: Analysis of Cases Involving Electric Bike Theft Gangs (Scenario Description: Multiple electric bike thefts occur in a certain area. The large language model initially classifies them as "theft." The knowledge graph automatically links the cases through spatiotemporal relationships (such as within a 500-meter radius of the crime scene and nighttime hours) to generate a gang crime network diagram): (I) Police Report Text Access and Packaging - Splitting and Hierarchical Processing Suppose the analysis system receives 100 consecutively reported police reports of electric vehicle theft, which include information such as the location, time, and description of the person reporting the crime.
[0114] (1) Perform batch packing and initial screening of police reports: The analysis system first counts the number of words in each police report text and determines the block size (e.g., 15 reports / packet) using Equation 1 above. Subsequently, the analysis system extracts keywords such as "electric vehicle", "theft", and "nighttime" and calculates the risk score for each police report text using Equation 2 above. Assuming that only 22 police report texts have a risk score greater than the preset risk threshold (e.g., 0.7), the analysis system marks these 22 police report texts as potentially important police report texts. The final output includes a packaged block of police report texts containing structured identifiers, as well as the 22 selected important police report texts.
[0115] (2) Segmentation of Important Police Report Texts and Classification by Large Language Model: The analysis system segments 22 important police report texts into individual texts based on semantic boundaries (such as punctuation marks and changes in the subject of the event), and uses a 50-word sliding window to segment each important police report text. Next, the analysis system adds a structured Prompt template to each important police report text to guide the large language model to output structured results. Initial classification is performed using a lightweight BERT, followed by secondary verification using a fine-tuned qwen3.5, and combined with a multi-model voting mechanism. All 22 important police report texts were accurately labeled as "theft," with classification confidence exceeding 90%. Finally, the classification results for the 22 important police report texts are output.
[0116] (II) Dynamic Construction of Knowledge Graphs Based on the categorized police incident data, the analysis system enters the knowledge graph construction phase: (1) Police Report Entity and Relationship Extraction: The Span-BERT model jointly extracts entities (including case type, time, location, people, organization name, and importance level) and relationships (including case participation relationships, spatiotemporal relationships, and organization handling relationships) from 22 important police report texts. Graph Attention Network (GAT) can be used to further aggregate neighbor information of entities, enhancing relationship prediction capabilities and outputting the probability distribution of relationships between entity pairs. Finally, an entity list (e.g., "XX Road XX Community", "2-4 AM", "Li (reporter)") and a relationship list (e.g., "Case A → Same Area → Case B") are obtained.
[0117] (2) Knowledge Fusion and Graph Storage: The analysis system calculates the relationship strength score for each important police report text using Equation 8 above. Simultaneously, entity disambiguation is achieved through a weighted voting mechanism based on contextual similarity, knowledge base matching degree, and frequency statistics, merging duplicate entities. Finally, the structured entity-relationship pairs are stored in the Neo4j graph database, forming a knowledge graph of electric vehicle theft cases containing 22 case nodes, 18 location nodes, and 3 core relationships.
[0118] (III) Hybrid Semantic Retrieval and Graph Enhancement Analysis To uncover deeper connections between cases, the analysis system initiated hybrid semantic retrieval and graph augmentation analysis: (1) Case Association Retrieval and Graph Expansion: Assuming the user inputs a query request: "Find the associations of electric vehicle theft cases in XX District in the past 30 days", the corresponding SQL query statement is automatically generated through the large language model, and the initial screening result set is obtained by querying the police information database. Subsequently, the analysis system performs a hybrid retrieval: the BM25 algorithm is used to calculate the keyword matching degree, and the ERNIE 3.0 model is used to calculate the semantic similarity. Finally, a list of target police information texts sorted based on the hybrid retrieval score is generated. Since this query is of high priority, the maximum number of hops dynamically calculated by the analysis system is ≥3 (e.g., rounded to 3 hops). The analysis system penetrates the 3-layer relationship network and mines the potential association of "case → location → suspect", and finds that cases 1-5 all point to the same suspect's location.
[0119] (2) Gang network generation and output: The analysis system integrates the relationships between cases, locations, and personnel to generate a visualized gang crime network diagram and marks the authority and relationship strength of core entities. Finally, a structured case series analysis report is output, which includes crime patterns, core related points and suspect clues, as well as a visualized gang crime network diagram, directly supporting practical judgment.
[0120] Corresponding to the aforementioned application function implementation method embodiments, this application also provides an alarm processing device, electronic device, and corresponding embodiments.
[0121] Figure 4 This is a schematic diagram of the alarm processing device shown in the embodiments of this application.
[0122] See Figure 4 The alarm handling device provided in this application may include: The filtering module 410 is used to group multiple police incident texts in the police database and then filter out high-risk important police incident texts from each group. The module 420 is used to classify multiple important police incident texts using a large language model, mine the relationships between multiple important police incident texts in each category, and construct a corresponding knowledge graph for at least two important police incident texts in each category that have a relationship. The query module 430 is used to perform a hybrid search in the police database when a query request described in natural language is received, and to obtain the target police information text that matches the query request; the hybrid search includes keyword search and semantic search. Analysis module 440 is used to query related police texts associated with the target police text in the knowledge graph, so as to perform case parallel analysis on the target police text and related police texts.
[0123] In one embodiment, when performing the step of grouping multiple police incident texts in the police database, the filtering module 410 may include: The chunk size determination submodule is used to determine the chunk size, and The semantic attribute recognition submodule is used to identify the semantic attributes of each police incident text in the police database; wherein the semantic attributes include at least one of the event subject, time and location; The partitioning submodule is used to divide multiple police incident texts in the police database into multiple blocks according to semantic attributes and block size; each block includes a specified number of police incident texts.
[0124] In one embodiment, when performing the step of filtering out high-risk important alarm texts from each group, the filtering module 410 may include: The feature extraction submodule is used to extract key features from each of the police alarm texts in each block; the key features include at least one of keyword trigger count, event type probability, and contextual relevance. The weight allocation submodule is used to assign corresponding feature weights to each key feature. The risk score calculation submodule is used to calculate the risk score of each alarm text in each block according to each key feature and its corresponding feature weight. The Important Alarm Text Filtering Submodule is used to filter out important alarm texts with risk scores greater than preset risk thresholds from each block.
[0125] In one embodiment, the large language model includes multiple models; when performing the step of classifying multiple important police report texts using the large language model, the construction module 420 may include: The prompt word template retrieval submodule is used to retrieve prompt word templates; The classification submodule is used to input multiple important police report texts and prompt word templates into multiple large language models, so that the multiple large language models can classify the multiple important police report texts respectively, and obtain the structured results of each important police report text output by the multiple large language models under the guidance of prompt word templates; the structured results include classification labels; The voting submodule is used to select the most frequently occurring classification label from the classification labels output by multiple large language models for the same important police report text as the target classification label for the important police report text.
[0126] In one embodiment, when performing the step of mining the correlation between multiple important police report texts of different types, the construction module 420 may include: The entity recognition submodule is used to identify entities in each of the multiple important police report texts that have the same target classification label; the entities include at least one of the following: case type, time, location, person, organization name, and importance level; The relationship mining submodule is used to mine the relationships between multiple important police reports with the same target classification label based on entities. The relationships include at least one of the following: case participation relationship, spatiotemporal relationship, and institutional handling relationship.
[0127] In one embodiment, the knowledge graph includes a dedicated knowledge graph; when the construction module 420 performs the step of constructing a corresponding knowledge graph for at least two important police report texts that have a relationship in various categories, it may include: The dedicated knowledge graph construction submodule is used to construct a corresponding dedicated knowledge graph for at least two important police incident texts that have any relationship with each other, given multiple important police incident texts that have the same target classification label. The knowledge graph also includes a global knowledge graph, and the building module 420 may also include: The global knowledge graph construction submodule is used to construct a corresponding global knowledge graph for at least two important police reports that have any correlation with each other, given multiple important police reports with different target classification labels.
[0128] In one embodiment, when the query module 430 performs the step of performing a mixed search in the police database to obtain the target police information text that matches the query request upon receiving a query request described in natural language, it may include: The conversion submodule is used to convert a query request described in natural language into a structured query statement through a large language model when a query request is received. The structured query statement is then used to retrieve matching target alarm texts from the police database. The target alarm texts include multiple texts. The dual-channel construction submodule is used to calculate the keyword matching degree between the structured query statement and each target alarm text, and to calculate the semantic similarity between the structured query statement and each target alarm text. The hybrid retrieval submodule is used to calculate the hybrid retrieval score of each target police text based on the keyword matching degree and the semantic similarity, and to sort the multiple target police texts based on the hybrid retrieval score to obtain a list of target police texts.
[0129] In one embodiment, when performing the step of querying related alarm texts associated with the target alarm text in the knowledge graph, the query module 430 may include: The case attribute determination submodule is used to determine the urgency coefficient and entity authority of each target police text based on its ranking in the target police text list. The maximum hop count calculation submodule is used to obtain a preset hop count threshold, and to calculate the maximum hop count of each target alarm text based on the preset hop count threshold, urgency coefficient and entity authority. The association traversal submodule is used to perform association traversal in the knowledge graph of each target alarm text, starting from the target alarm text and according to the maximum number of hops of the target alarm text, to obtain the associated alarm texts related to the target alarm text.
[0130] In one embodiment, the device may further include a feedback optimization module, which may include: The block segmentation strategy optimization submodule is used to record the classification confidence and processing time of each of the aforementioned important alarm texts, so as to adjust the block size based on the classification confidence and processing time; and, The screening strategy optimization submodule is used to determine the sample set to be updated so that the feature weights can be updated using the sample set to be updated. The adjusted block size is used to balance classification accuracy and classification efficiency, while the updated feature weights are used to balance historical rule experience and new sample patterns.
[0131] As this example demonstrates, the solution provided in this application dynamically divides massive amounts of police report texts into multiple groups, enabling rapid and accurate filtering of important police report texts from small batches of grouped texts. This significantly improves throughput and efficiency while maintaining processing accuracy. Furthermore, large language models perform exceptionally well in natural language understanding tasks; therefore, using large language models to classify important police report texts can improve classification efficiency and accuracy. Moreover, by integrating a hybrid retrieval architecture combining keyword retrieval and semantic retrieval, it not only balances keyword accuracy and semantic generalization capabilities to meet the needs of natural language description and improve police report retrieval efficiency, but also allows for direct matching and classification based on the semantics of newly added dynamic tags (such as subcategories like "telecom fraud - fake order type") in the police database without requiring manual rule re-engineering or classification system adjustments, quickly adapting to new tag requirements. Finally, by mining the relationships between different important police report texts to construct a corresponding knowledge graph, related police reports can be retrieved through this knowledge graph when querying related police reports in the future, thereby improving the efficiency of case analysis.
[0132] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated further here.
[0133] Figure 5 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application.
[0134] See Figure 5 The electronic device 500 includes a memory 510 and a processor 520.
[0135] The processor 520 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. Memory 510 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by the processor 520 or other modules of the computer. Permanent storage devices may be read-write storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices may be removable storage devices (e.g., floppy disks, optical drives). System memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory may store some or all of the instructions and data required by the processor during operation. Furthermore, memory 510 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, memory 510 may include a removable storage device that is readable and / or writable, such as a laser disc (CD), a read-only digital multifunction optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-high density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves or transient electronic signals transmitted wirelessly or via wired connections.
[0136] The memory 510 stores executable code, which, when processed by the processor 520, can cause the processor 520 to execute part or all of the methods described above.
[0137] Furthermore, the method according to this application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing some or all of the steps in the method described above.
[0138] Alternatively, this application may be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium) storing executable code (or computer program or computer instruction code) that, when executed by a processor of an electronic device (or server, etc.), causes the processor to perform part or all of the steps of the methods described above according to this application.
[0139] This application also provides a computer program product, which includes computer instructions that, when executed by a processor, implement the method described above.
[0140] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for handling police incidents, characterized in that, include: After grouping multiple police incident texts in the police database, high-risk and important police incident texts are selected from each group. After classifying multiple important police incident texts using a large language model, the relationships between multiple important police incident texts in each category are mined, and a corresponding knowledge graph is constructed for at least two important police incident texts in each category that have the relationships. When a query request described in natural language is received, a hybrid search is performed in the police database to obtain the target police information text that matches the query request; the hybrid search includes keyword search and semantic search. The knowledge graph is used to query associated police texts related to the target police text in order to perform case linkage analysis on the target police text and the associated police texts.
2. The method according to claim 1, characterized in that, The grouping of multiple police report texts in the police database includes: The block size is determined, and the semantic attributes of each police incident text in the police database are identified respectively; wherein, the semantic attributes include at least one of the event subject, time and location; According to the semantic attributes and the block size, the multiple police incident texts in the police database are divided into multiple blocks; each block includes a specified number of police incident texts.
3. The method according to claim 2, characterized in that, The process of selecting high-risk, important alert texts from each group includes: Extract key features from each of the police report texts within each of the respective blocks; the key features include at least one of keyword trigger count, event type probability, and contextual relevance. Assign corresponding feature weights to each of the key features; Based on the key features and their corresponding feature weights, calculate the risk score of each alarm text within each block; Important alert texts with risk scores greater than a preset risk threshold are selected from each of the aforementioned blocks.
4. The method according to claim 3, characterized in that, The large language model includes multiple models; the classification of multiple important police report texts using the large language model includes: Get the prompt word template; Multiple important police report texts and the prompt word templates are respectively input into multiple large language models, so that the multiple large language models classify the multiple important police report texts respectively, and obtain the structured results of each important police report text output by the multiple large language models under the guidance of the prompt word templates; the structured results include classification labels; For the same important police report text, the classification label that appears most frequently is selected from the classification labels output by multiple large language models as the target classification label of the important police report text.
5. The method according to claim 4, characterized in that, The process of mining the relationships between multiple important police report texts of various types includes: For multiple important police report texts that have the same target classification label, the entities of each important police report text are identified respectively; the entities include at least one of case type, time, location, person, organization name and importance level; For multiple important police report texts that have the same target classification label, the relationships between the multiple important police report texts are mined based on the entity; the relationships include at least one of case participation relationship, spatiotemporal relationship and institutional handling relationship.
6. The method according to claim 5, characterized in that, The knowledge graph includes a dedicated knowledge graph; the construction of a corresponding knowledge graph for at least two important police report texts that have the aforementioned relationship in each category includes: For multiple important police reports with the same target classification label, construct a corresponding exclusive knowledge graph for at least two of the important police reports that have any of the aforementioned relationships; The knowledge graph also includes a global knowledge graph, and the method further includes: For multiple important police incident texts with different target classification labels, a corresponding global knowledge graph is constructed for at least two of the important police incident texts that have any of the aforementioned association relationships.
7. The method according to claim 1, characterized in that, When a query request described in natural language is received, a mixed search is performed in the police database to obtain the target police information text that matches the query request, including: When a query request described in natural language is received, the query request is converted into a structured query statement through the large language model, and the structured query statement is used to query the police database for matching target alarm texts; the target alarm texts include multiple texts. Calculate the keyword matching degree between the structured query statement and each of the target alarm texts, and calculate the semantic similarity between the structured query statement and each of the target alarm texts; Based on the keyword matching degree and the semantic similarity, a mixed retrieval score is calculated for each of the target police texts. Based on the mixed retrieval scores, the multiple target police texts are sorted to obtain a list of target police texts.
8. The method according to claim 7, characterized in that, The step of querying the knowledge graph for associated police texts related to the target police text includes: Based on the ranking of each target alert text in the target alert text list, the urgency coefficient and entity authority of each target alert text are determined respectively; A preset hop count threshold is obtained, and the maximum hop count for each target alarm text is calculated based on the preset hop count threshold, the urgency coefficient, and the entity authority. For each target alert text, starting from the target alert text, perform an association traversal in the knowledge graph of the target alert text according to the maximum number of hops of the target alert text to obtain the associated alert texts related to the target alert text.
9. The method according to claim 4, characterized in that, The method further includes: Record the classification confidence level and processing time of each of the aforementioned important police alert texts, and adjust the block size based on the classification confidence level and the processing time; and, Determine the sample set to be updated, and use the sample set to be updated to update the feature weights; The adjusted block size is used to balance classification accuracy and classification efficiency, and the updated feature weights are used to balance historical rule experience and new sample patterns.
10. An alarm processing device, characterized in that, include: The filtering module is used to group multiple police incident texts in the police database and then filter out high-risk and important police incident texts from each group. The construction module is used to classify multiple important police incident texts using a large language model, mine the relationships between multiple important police incident texts in each category, and construct a corresponding knowledge graph for at least two important police incident texts in each category that have the relationships. The query module is used to perform a hybrid search in the police database when a query request described in natural language is received, to obtain the target police information text that matches the query request; the hybrid search includes keyword search and semantic search. The analysis module is used to query related police texts associated with the target police text in the knowledge graph, so as to perform case parallel analysis on the target police text and the related police texts.