A cross-domain data intelligent fusion analysis method for public safety scenarios

CN122734839APending Publication Date: 2026-09-11THE FIRST RES INST OF MIN OF PUBLIC SECURITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610828623.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

实体关联性弱的问题:现有技术主要依赖身份证号等强标识信息进行确定性关联

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734839A_ABST
    Figure CN122734839A_ABST
Patent Text Reader

Abstract

This invention relates to the field of public safety technology and discloses a cross-domain data intelligent fusion analysis method for public safety scenarios. The method includes: collecting multi-source heterogeneous data and preprocessing it; modeling visual events and business events in the multi-source heterogeneous data into computable behavioral units to obtain cross-domain entity association confidence; performing large-scale intelligent agent analysis, adopting an agent-driven hypothesis-verification-reflection analysis paradigm to complete the arrangement from natural language to directed acyclic graphs, obtaining large-scale intelligent agent analysis results; reconstructing cross-domain behavioral chain stories based on the large-scale intelligent agent analysis results, performing multi-dimensional verification on the large-scale intelligent agent analysis results, and obtaining a comprehensive causal confidence score; organizing the evidence chain of the overall analysis results, and outputting judgment results supporting investigative decisions. Using this invention, the intelligence level, depth, and interpretability of cross-domain fusion analysis in public security can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of public safety technology, and specifically to a cross-domain data intelligent fusion and analysis method for public safety scenarios. Background Technology

[0002] Currently, in the fields of public security and criminal investigation, existing information systems are mainly based on structured data query and relationship graph analysis. While these systems have improved investigation efficiency to some extent, they still have the following technical shortcomings: The problem of weak entity association: Existing technologies mainly rely on strong identifying information such as ID card numbers for deterministic association. However, for weak association scenarios such as unverified individuals, vehicles using cloned or fake license plates, it is difficult to establish effective data connections.

[0003] The analytical logic is rigid: the current analysis process relies heavily on police officers manually setting query conditions, such as time range, spatial location, and specific attributes like clothing color. Even though some low-code platforms support building analysis workflows by dragging and dropping nodes, police officers still need to assemble techniques and methods based on their own experience, resulting in poor flexibility and an inability to answer open-ended analysis questions that are not pre-set.

[0004] Lacking in temporal and causal relationships: Existing graph analysis or retrieval techniques are essentially queries targeting static snapshots, and the output results are graph sets, not readable judgment conclusions. They cannot model dynamic crime processes or perform anomaly analysis. Insufficient interpretability: The output of existing systems is in the form of data lists or graphs, lacking chain-of-evidence output, making it difficult to use directly for investigative decision-making.

[0005] To address this, we have invented a cross-domain data intelligent fusion and analysis method for public safety scenarios, which solves the above-mentioned technical problems. Summary of the Invention

[0006] This invention provides a cross-domain data intelligent fusion analysis method for public safety scenarios, which can propose investigative hypotheses, autonomously retrieve data from cross-databases, perform multi-step reasoning, reflective verification, causal judgment, and output traceable analysis conclusions, significantly improving the intelligence level, depth, and interpretability of cross-domain fusion analysis in public security.

[0007] Therefore, the present invention provides the following technical solution: A cross-domain data intelligent fusion and analysis method for public safety scenarios, the method comprising: Step 1: Collect multi-source heterogeneous data and preprocess the multi-source heterogeneous data; Step 2: The visual events and business events in the preprocessed multi-source heterogeneous data are uniformly modeled into computable behavioral units to obtain the cross-domain entity association confidence. Step 3: Perform large-scale agent analysis, adopting an agent-driven hypothesis-to-verification-to-reflection analysis paradigm, complete the arrangement of natural language into directed acyclic graph, and obtain the large-scale agent analysis results. Step 4: Based on the analysis results of the large model agent, reconstruct the cross-domain behavior chain story, perform multi-dimensional verification on the analysis results of the large model agent, and obtain a comprehensive causal confidence score. Step 5: Organize the evidence chain of the overall analysis results and output the judgment results that support the investigation decision.

[0008] Optionally, in step 1, the preprocessing of the multi-source heterogeneous data includes: protocol parsing, format standardization, and quality cleaning.

[0009] Optionally, step 2 includes: Step 21: Extract features from any trajectory within the time window and concatenate them into a comprehensive feature vector; Step 22: The dynamic time warping algorithm is used to calculate the trajectory similarity based on the comprehensive feature vector; Step 23: Based on the trajectory similarity, construct the edge weights between the targets in the video structured data and the targets covered in the business data, thereby constructing a cross-domain association network; Step 24: Perform graph mining on the constructed cross-domain association network based on the Infomap algorithm to divide the network into companion gang communities; Step 25: Calculate the cross-domain entity association confidence based on the trajectory similarity and the community mining results obtained based on the Infomap algorithm.

[0010] Optionally, in step 21, the visual trajectory and the public security business trajectory are mapped to a unified geographic grid, i.e., a time window space. The extracted features include trajectory geometry, movement speed features, and activity pattern features. The trajectory geometry includes: trajectory length, mean turning angle, straightness, and number of covered grids. The movement speed features include: average speed, speed variance, and percentage of stationary time. The activity pattern features include: distribution of active periods, activity frequency, and periodicity indicators.

[0011] Optionally, in step 25, the formula for calculating the cross-domain entity association confidence Conf(u,v) is: Conf(u,v)=α×Sim_DTW+β×I (same community) Wherein, Sim_DTW is the trajectory similarity, with a value of 0-1; u and v represent two personnel entities to be compared; I(same community) is an indicator function; it is 1 if u and v are in the same community, and 0 otherwise. α and β are weighting coefficients, and α+β=1. Optionally, step 3 includes: Step 31: The user raises an open-ended assessment question; Step 32: For the open-ended judgment problem, the large language model generates multiple hypothetical paths based on investigative knowledge and data structure; Step 33: For each hypothetical planning path, convert the hypothesis into a DAG analysis graph. Each node in the graph represents a tool call, the edges represent data dependencies, and the directed edges represent the data flow direction. Step 34: Based on the DAG analysis graph, the tool starts executing to obtain the analysis results of the large model intelligent agent. According to the edge type, it is identified whether to execute serially, in parallel or conditionally, and calls the required tools to obtain data from various business systems. The tools called include video parsing, target retrieval, target deployment and correlation analysis.

[0012] Optionally, step 4 includes: Step 41: Reconstruct the behavior chain by merging the input cross-domain event data along the time axis to generate a continuous behavior chain sequence. Step 42: Using the propensity score matching causal inference method, a control group is constructed to eliminate confounding factors, estimate the true causal effect, and obtain the statistical verification results. Step 43 proposes a causal confidence scoring model, and conducts a comprehensive causal confidence score based on statistical test results to identify the division of labor and crime patterns within the gang.

[0013] Optionally, in step 5, all intermediate results, original data references, statistical indicators, and causal confidence scores in the analysis process are forcibly bound as a structured chain of evidence as the overall analysis result, and an assessment report is output. The chain of evidence adopts a structured JSON format.

[0014] A cross-domain data intelligent fusion and analysis device for public safety scenarios, the device comprising: The data acquisition and preprocessing unit acquires multi-source heterogeneous data and preprocesses the multi-source heterogeneous data; The cross-domain entity association unit models visual events and business events in the preprocessed multi-source heterogeneous data into a computable behavioral unit to obtain the cross-domain entity association confidence. The large model agent analysis unit performs large model agent analysis, adopts an agent-driven hypothesis-to-verification-to-reflection analysis paradigm, completes the arrangement of natural language into directed acyclic graph, and obtains the large model agent analysis results. The causal inference and behavioral chain reconstruction unit reconstructs cross-domain behavioral chain stories based on the analysis results of the large model intelligent agent, performs multi-dimensional verification of the analysis results of the large model intelligent agent, and obtains a comprehensive score of causal confidence. The evidence chain organization and analysis results output unit organizes the evidence chain of the overall analysis results and outputs analysis results that support investigation decisions.

[0015] A computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to perform the steps of the cross-domain data intelligent fusion analysis method for public safety scenarios.

[0016] This invention provides a cross-domain data intelligent fusion analysis method for public safety scenarios. It employs a cross-domain entity association method based on spatiotemporal trajectory fitting and correlation network reasoning, solving the problem of weak entity association. It utilizes an agent-driven "Hypothesis-Verify-Reflect Loop (HVR Loop)" analysis paradigm and automatic orchestration from natural language to directed acyclic graphs (NL2DAGs), breaking through the rigidity of existing analysis logic. It establishes a temporal and causal relationship using cross-domain causal inference and behavioral chain reconstruction techniques. It overcomes the insufficient interpretability of existing technologies by employing an evidence chain-based interpretable output mechanism. This invention breaks through the capability boundaries of traditional big data platforms, graph analysis, and statistical models, enabling the system to proactively propose investigative hypotheses, autonomously retrieve data across databases, perform multi-step reasoning, reflective verification, causal judgment, and output traceable analytical conclusions, much like a seasoned image investigation expert. This significantly improves the intelligence, depth, and interpretability of cross-domain fusion analysis in public security. Compared with existing technologies, this invention has the following technical advantages: To improve the efficiency of analysis, traditional image-based investigation requires multi-agency collaboration and relies on manual sorting of clues and cross-referencing of data, typically taking 3 to 5 days to complete a full analysis. This invention can directly output a complete analysis conclusion with supporting evidence, causal logic explanations, and investigative suggestions within minutes, significantly shortening the investigation cycle and improving emergency response capabilities.

[0017] To enhance the association capabilities, existing technologies are limited by deterministic associations based on strong identifiers such as ID numbers, making it difficult to handle weak association scenarios such as unregistered individuals and vehicles with cloned license plates. This invention, however, adopts a cross-domain entity association method based on spatiotemporal trajectory fitting and association network reasoning, which supports entity association in weak association scenarios such as unregistered individuals and vehicles with cloned license plates. To enhance the depth of analysis, existing technologies remain at the stage of simple spatiotemporal co-occurrence analysis. This invention transcends spatiotemporal co-occurrence analysis to the level of causal inference, and can identify deep relationships such as gang division of labor and crime patterns.

[0018] To improve interpretability, this invention can output a chain of evidence report, with all conclusions supported by original data and statistical indicators, ensuring that the conclusions are traceable, verifiable, and credible.

[0019] Lowering the barrier to entry, this invention supports natural language input analysis, eliminating the need for police officers to master complex tactical assembly skills, making it easy for grassroots police officers to quickly get started and apply. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a cross-domain data intelligent fusion and analysis method for public safety scenarios in a specific embodiment of the present invention; Figure 2 This is a schematic diagram of a cross-domain data intelligent fusion analysis device for public safety scenarios, according to a specific embodiment of the present invention. Detailed Implementation

[0022] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] like Figure 1 The diagram shown is a flowchart of a cross-domain data intelligent fusion and analysis method for public safety scenarios according to an embodiment of the present invention, including the following steps: Step 1: Collect multi-source heterogeneous data and preprocess the multi-source heterogeneous data.

[0025] This system unifies the access, protocol parsing, format standardization, and quality cleaning of multi-source heterogeneous data to obtain cleaned trajectory data and business data, breaking down the isolation barriers between video data and business data at various levels. It supports video protocols such as GB / T 28181 and GA / T1400, as well as API interfaces for various business systems. Step 2: Perform cross-domain entity association, and model visual events and business events in the preprocessed multi-source heterogeneous data into computable behavioral units.

[0026] Cross-domain entity association addresses the heterogeneous alignment issue between visual-side IDs (face identifier `face_id`, license plate identifier `plate_id`) and business-side IDs (ID card number, mobile phone number), achieving weak association matching between cross-domain entities and supporting subsequent intelligent analysis. Cross-domain entities refer to the targets that need to be associated, such as people and vehicles. Specifically, they include: Step 21: Perform feature extraction and map the visual side trajectory (camera ID + timestamp + geographic coordinate sequence) and the public security business side trajectory (spatiotemporal sequence of base station / accommodation / vehicle ETC / transaction / entry and exit) to a unified geographic grid, i.e. time window space, such as a 100m×100m grid with a 1-hour window.

[0027] The geographic grid size of 100m×100m was chosen based on statistics of the average width of urban roads and the coverage of cameras. Experiments have verified that this granularity strikes a balance between computational efficiency and accuracy. The time window of 1 hour was chosen based on statistics of human activity patterns, which can cover most short-term stay behaviors.

[0028] For any trajectory T={(x) in the time window space i ,y i ,t i ,a i )} (i=1 to n), where: (x i , y i ) represents geographic coordinates; t i For timestamps; a i For attribute information (such as dwell time, activity type, etc.), extract three types of features and concatenate them into a comprehensive feature vector: F_T=[f_shape,f_speed,f_temporal] ∈R^d Wherein: f_shape describes the trajectory geometry, including: trajectory length, mean turning angle, straightness, and number of covered grids; f_speed describes the movement speed characteristics, including: average speed, speed variance, and percentage of stationary time; f_temporal describes the activity pattern characteristics, including: distribution of active periods, activity frequency, and periodicity indicators.

[0029] Step 22: Calculate the DTW trajectory similarity.

[0030] Dynamic Time Warping (DTW) algorithm is used to calculate the similarity of trajectories of different lengths. Its core idea is to find the optimal alignment path between two trajectories. Let trajectory A have m points and trajectory B have n points. Construct an m×n distance matrix D, where D[i][j] = d(A[i], B[j]) is the Euclidean distance between two points. The DTW distance is calculated recursively using dynamic programming. DTW[i][j]=d(A[i],B[j])+min{ DTW[i-1][j], / / Insertion DTW[i][j-1], / / Delete DTW[i-1][j-1] / / Match } Where i represents the index of the i-th time point in trajectory A (i=1,2, ...,n), and j represents the index of the j-th time point in trajectory B (j=1,2, ...,m). For example, trajectory A (Li) has 3 points, i=1,2,3; trajectory B (Wang) has 2 points, j=1,2. In the DTW algorithm, i and j are used to traverse all point pairs of the two trajectories to find the optimal alignment path.

[0031] Boundary condition: DTW[0][0]=0, the final DTW distance is DTW[m][n]. Based on the DTW distance, calculate the trajectory similarity Sim_DTW: Sim_DTW=exp(-DTW / σ) Wherein, σ is an empirical parameter that controls the rate of similarity decay, and the recommended value is σ=0.5 (obtained through cross-validation optimization). Step 23: Perform cross-domain network association.

[0032] Based on trajectory similarity, edge weights are constructed between targets in the video structured data and targets covered in the business data to build an association network. The network construction rules are as follows: each cross-domain entity is treated as a node; when Sim_DTW≥0.5, an edge is established between the two entities; the edge weight is w(u,v)=Sim_DTW(u,v), where u represents the first person entity to be compared (e.g., "suspect A"), and v represents the second person entity to be compared (e.g., "suspect B" or "victim").

[0033] Step 24: Perform Infomap community discovery.

[0034] Infomap is an information-theory-based graph community partitioning algorithm that discovers communities by minimizing the amount of information describing random walk paths. This paper uses the Infomap algorithm to perform graph mining on a constructed cross-domain interconnected network, dividing the network into companion groups communities. The Infomap algorithm, based on information theory, discovers community structure by minimizing the amount of information describing random walk paths. Its core idea is that if a community structure exists in the network, then describing the movement of random walks within communities using compressed encoding is more efficient than using global encoding. The steps of the Infomap algorithm include: 1a. Treat the network as an information flow network and define the transition probabilities between nodes; 2a. Assign an initial community tag to each node; 3a. Iterative optimization: For each node, try moving it to an adjacent community and calculate the change in the length of the information description; 4a. The algorithm converges when the information description length no longer decreases. When setting parameters, the random walk step count is set to 10,000 steps; the community merging threshold is: modularity increment > 0.01; and the maximum number of iterations is 100.

[0035] Step 25: Perform association determination.

[0036] Calculate entity association confidence based on DTW trajectory similarity and community mining results: Conf(u,v)=α×Sim_DTW+β×I (same community) Where Sim_DTW is the trajectory similarity (0-1) from step 22; I(same community) is an indicator function; it is 1 if u and v are in the same community, and 0 otherwise. α and β are weighting coefficients, α+β=1 (recommended α=0.7, β=0.3).

[0037] The weighting of α=0.7 and β=0.3 was chosen because historical data validated that this weighting achieves the optimal balance between precision and recall. The threshold of 0.75 was selected based on Receiver Operating Characteristic (ROC) curve analysis, which corresponds to approximately 90% precision. Entity association confidence is used to determine whether two individuals are the same person or whether a connection exists. In this invention, Conf(u,v) ≥ 0.75 is defined as an entity association, supporting file merging.

[0038] Step 3: Perform large-scale model agent analysis, adopting an agent-driven "hypothesis-verification-reflection" analysis paradigm to support automatic planning and execution of open-ended natural language judgment questions, and complete the automatic arrangement of natural language into directed acyclic graph NL2DAG.

[0039] This implementation adopts an agent-driven "hypothesis-validation-reflection" analysis paradigm, supporting automated planning and execution for open-ended natural language processing (NLP) questions. Investigators input NLP questions, and based on the investigative knowledge base and data schema, a large language model (LLM) automatically generates multiple hypotheses to be validated. These hypotheses are then transformed into a directed acyclic graph (DAG) for analysis. Pre-packaged tools for spatiotemporal querying, trajectory collision analysis, and statistical testing are then used to integrate the analysis results. The implemented analysis flow is as follows: After users raise open-ended analysis questions, the LLM (Local Level Management) generates multiple hypotheses based on investigative knowledge and data structures. The verification planner plans paths for each hypothesis, automatically arranging them into a Directed Acyclic Graph (DAG). Each node in the DAG represents a tool call, and edges represent data dependencies. Based on the DAG, the tool execution phase begins, automatically calling the necessary tools to retrieve data from various business systems. Based on edge type, it automatically identifies sequential, parallel, and conditional execution, shortening analysis time. Causal inference and behavioral chain reconstruction are performed on the execution results data to uncover logical relationships, identify spatiotemporal contradictions, find counterexamples, and determine if path replanning is necessary. Finally, the analysis results are summarized and output. Specifically, this includes: Step 31: The user raises an open-ended assessment question. Step 32, Hypothesis Generation: LLM generates multiple hypotheses based on reconnaissance knowledge and data structures. Hypothesis generation follows these rules: hypotheses must be verifiable (with clear data sources and verification methods); hypotheses must be independent or complementary; the number of hypotheses should be controlled to 3-5 to avoid excessive divergence.

[0040] Step 33, Verification Planning: The verification planner automatically arranges each hypothetical planned path into a DAG (Directed Acyclic Graph) analysis diagram. The DAG construction rules are as follows: Nodes: Each node represents a tool call (such as spatiotemporal query, trajectory collision, statistical test, etc.); Edges: Edges represent data dependencies, and directed edges indicate the direction of data flow; Execution strategy: Automatically identify whether to execute sequentially, in parallel, or conditionally based on edge type.

[0041] Step 34, Tool Execution: Based on the Directed Acyclic Graph (DAG), the tool execution phase begins, automatically invoking the necessary tools to retrieve data from various business systems. These tools include video analysis, target retrieval, target deployment, and correlation analysis tools. The tool registration mechanism is as follows: Tools need to be pre-registered with the tool library, including input and output schemas; The execution engine automatically schedules tool calls based on DAG dependencies; Supports retrying on failure and exception handling.

[0042] Step 4 moves from "spatiotemporal accompaniment" to "crime causation" analysis. Based on Agent analysis results, the cross-domain behavioral chain story is reconstructed, the logical relationships in the data are mined, and the analysis results are verified from multiple dimensions, including evidence sufficiency, spatiotemporal consistency, survivor bias, and counterexamples. The verification rules are shown in Table 1. This process identifies gang division of labor and crime patterns. Gang identification is based on a community detection algorithm; gang characteristics include spatiotemporal overlap, frequent communication, and financial connections. Crime pattern identification is based on matching known pattern libraries and anomaly pattern detection. A comprehensive confidence score is calculated based on spatiotemporal overlap, behavioral pattern characteristics, and gang associations to output strong causation and identify the gang.

[0043] Table 1 Reflection and Verification Rules Table

[0044] Traditional analytical methods can only answer whether two people were traveling together, but cannot determine whether the behavior was premeditated or accidental. Practical investigative work requires causal judgment, not simple correlation statistics. This invention proposes a cross-domain causal inference method for public safety scenarios, adapting the Propensity Score Matching (PSM) causal inference algorithm to multi-source heterogeneous trajectory data, and developing behavioral chain temporal reconstruction technology to connect fragmented events into an interpretable crime narrative. Trajectory data refers to spatiotemporal sequence data, including visual and operational trajectories. Specifically, it performs causal inference and behavioral chain reconstruction on execution result data, mining the logical relationships in the data, identifying spatiotemporal contradictions, providing counterexamples, determining whether route replanning is necessary, etc., and finally summarizing and outputting the results. Specifically, this includes: Step 41, Behavioral Chain Reconstruction: Based on the input cross-domain event data, merge the data along the timeline to generate a continuous sequence of behavioral chains. A behavioral chain is a narrative sequence reconstructed from the merged cross-domain events along the timeline. A behavioral chain sequence refers to an ordered sequence of the suspect's actions during the crime, arranged chronologically, in the form of [Behavior 1, Behavior 2, ..., Behavior n], where each behavior includes {timestamp, behavior type, location, associated object}. Behavioral chains are used to reconstruct the crime process (e.g., the complete behavioral chain of reconnaissance, execution, and escape), identify abnormal patterns, and construct evidence chains. The process is as follows: 1b. Event Extraction: Extract the set E={e1,e2,...,e...} of all events of the target entity from the cross-domain entity association graph. n Each event contains: Event types: Visual events / Business events / Financial events / Communication events; Timestamp τ (ISO8601 format); Spatial coordinates or geographic grid ID; Original data reference ID (event_id, record_id, transaction_id).

[0045] 2b. Time sequence sorting: Arrange the event sequence in ascending order by timestamp τ.

[0046] 3b. Anomaly Node Detection: Based on a historical behavior pattern database, identify time points that deviate from normal behavior (such as activities during late nights, deviations from daily activity radius, etc.). Anomaly detection rules include: Time Anomaly: The event occurred outside the 95th percentile historically; Spatial anomaly: The location of the activity exceeds twice the standard deviation of the historical activity radius; Frequency anomaly: The activity frequency exceeds three standard deviations from the historical mean.

[0047] Step 42, Causal Inference: The simultaneous occurrence of "the suspect appearing at the scene" and "abnormal financial activity of the companion" could be a causal relationship or a spurious correlation caused by confounding factors (such as a surge in pedestrian traffic during holidays or normal business operations in the area). Traditional platforms cannot distinguish between these. This invention proposes a propensity score matching (PSM) causal inference method. By constructing a control group, confounding factors are eliminated, and the true causal effect is estimated. Causal inference is performed using behavioral sequences to determine the causal relationship between the behavior and the case. The specific algorithm flow is as follows: 1c. Sample division: treatment group: T=1 (people present at the scene, n1 people), control group: T=0 (people not present at the scene, n0 people).

[0048] 2c. Covariate selection: Select characteristic variables that may affect the outcome, X = {age, criminal record type, activity area, historical activity pattern}.

[0049] 3c. Propensity score calculation: Use logistic regression to estimate the treatment probability for each sample. p(X)=P(T=1|X)=1 / (1+exp(-β0-β1X1-...-β k X k )); In the formula, β0 is the intercept term, β1...β k The regression coefficients are obtained through maximum likelihood estimation (MLE).

[0050] 4c. Sample matching: For each treatment group sample, find the sample in the control group with the closest propensity score (nearest neighbor matching): Matching distance: d(i,j)=|p(X) i )-p(X j )|; 5c. Estimation of the average treatment effect (ATE): ATE=(1 / n1)×Σ i=1 n ¹[Y i(Processing Group)-Y i (Matched control group) Where Y is the outcome variable (such as the degree of financial irregularities).

[0051] 6c. Statistical Test: Use a paired t-test to output the p-value and confidence interval. t = ATE / SE(ATE); p-value = P(|t|>observed value|null hypothesis:ATE=0); Where SE(ATE) is the standard error of ATE; Based on the above calculations, the PSM causal inference results are output as follows: Number of samples in the treatment group: n1; Control group sample size: n0 (after matching); Average treatment effect: ATE; Standard error: SE(ATE); t-statistic: t; p-value: p.

[0052] Step 43, Comprehensive Causal Confidence Score. A single causal inference may be biased; therefore, it is necessary to comprehensively assess the reliability of the causal relationship by considering both statistical significance and effect size. A comprehensive causal confidence score is then calculated based on the statistical test results to identify group division of labor and crime patterns. This invention proposes a causal confidence scoring model that uses the statistical test results output by the PSM (Public Sector Method) for comprehensive scoring: Conf_causal=λ×(1-p / 0.05)×I(p<0.05)+(1-λ)×min(ATE / 0.3, 1); in: Conf_causal: Causal confidence level, with a value range of [0,1]. p: p-value of the PSM test; ATE: Mean treatment effect size; λ: Weighting coefficient, which controls the relative importance of significance and effect size. λ=0.5 is recommended (each accounting for 50%). I(p<0.05): An indicator function that takes the value 1 when p<0.05, and 0 otherwise; min(ATE / 0.3,1): Effect size normalization, 0.3 is the threshold for moderate effect size (refer to Cohen's d criteria). In the formula, (1-p / 0.05)×I (p<0.05): statistical significance score, the smaller the p, the higher the score, and the score is 0 when p≥0.05; min(ATE / 0.3,1): Effect size score. The larger the ATE, the higher the score. When ATE≥0.3, the score is 1. The two weighted sums are used to ensure that Conf_causal is always within the range of [0,1].

[0053] Parameter selection criteria: p<0.05 corresponds to a 95% confidence level and is a commonly used significance standard in statistics; An ATE > 0.3 indicates that the treatment effect exceeds 0.3 times the standard deviation of the outcome variable, which is considered a moderate or greater effect size (refer to Cohen's d criteria). λ=0.5 indicates that significance and effect size are equally important, and can be adjusted according to practical needs (e.g., if more emphasis is placed on significance, λ=0.6 can be set).

[0054] Step 5, Evidence Chain Organization and Analysis Output: This step organizes the evidence chain of the overall analysis results, using behavioral sequences as the core content of the evidence chain. It outputs analysis results to support investigative decisions and achieves human-machine collaborative feedback optimization. All intermediate results, original data references, statistical indicators, confidence scores, etc., from the analysis process are forcibly bound into a structured evidence chain as part of the overall analysis results, outputting an analysis report to ensure the interpretability and traceability of the analysis results. The overall analysis results include: the cleaned trajectory data and business data output in Step 1; the similarity score Sim_DTW and entity association confidence score output in Step 2; and the behavioral sequences, behavioral pattern matching degrees, and comprehensive scores output in Step 4.

[0055] The chain of evidence uses a structured JSON format and includes the following fields: "case_id": "Case number"; "analysis_id": "analysis task ID"; "conclusion": "Analysis and conclusion"; "confidence_level": "Confidence level (high / medium / low)"; The "evidence_chain" chain of evidence specifically includes: "evidence_id": "evidence number", "evidence_type": "evidence type (video / business data / statistical results)", "source": "data source system", "record_id": "original record ID", "timestamp": "evidence timestamp", "description": "evidence description", "relevance_score": "relevance score to the conclusion", "chain_position": "position in the evidence chain"; The "analysis_method" section specifies the analysis method, including: "method_name": "analysis method used", "parameters": "method parameters", and "statistical_significance": "statistical significance index". "confounding_control" includes: "controlled_variables": "controlled confounding variables", "control_method": "control method"; "uncertainty_range": "range of uncertainty"; "human_review": "human review comments". The system supports multiple output formats: Structured JSON: for data exchange between systems; PDF reports: for printing and archiving.

[0056] The system records investigators' feedback on the analysis results (acceptance / partial acceptance / non-acceptance), which is used to optimize subsequent analysis models. Feedback data is used to adjust hypothesis generation strategies, optimize tool invocation priorities, and improve the causal confidence scoring model.

[0057] Accordingly, embodiments of the present invention also provide a cross-domain data intelligent fusion analysis device for public safety scenarios, such as... Figure 2 The image shown is a structural schematic of the device. This cross-domain data intelligent fusion analysis device for public safety scenarios includes the following modules: The data acquisition and preprocessing unit 201 acquires multi-source heterogeneous data and preprocesses the multi-source heterogeneous data; Cross-domain entity association unit 202 models visual events and business events in the preprocessed multi-source heterogeneous data into a computable behavioral unit to obtain cross-domain entity association confidence. The large model agent analysis unit 203 performs large model agent analysis, adopts an agent-driven hypothesis-to-verification-to-reflection analysis paradigm, completes the arrangement of natural language into a directed acyclic graph, and obtains the large model agent analysis results. The Causal Inference and Behavioral Chain Reconstruction Unit 204 reconstructs cross-domain behavioral chain stories based on the analysis results of the large model intelligent agent, performs multi-dimensional verification of the analysis results of the large model intelligent agent, and obtains a comprehensive score of causal confidence. The evidence chain organization and analysis result output unit 205 organizes the evidence chain of the overall analysis results and outputs analysis results that support investigation decisions.

[0058] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0059] The present invention also provides a storage medium, which is a computer-readable storage medium storing a computer program thereon, the computer program being executable when it runs. Figure 1 The method shown may include some or all of the steps. The storage medium may include read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc. The storage medium may also include non-volatile memory or non-transitory memory, etc.

[0060] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data provider to another website, computer, server, or data provider via wired or wireless means.

[0061] The embodiments of the present invention have been described in detail above. Specific implementation methods have been used to illustrate the present invention. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and systems of the present invention, and are merely some, not all, embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention, and the content of this specification should not be construed as a limitation of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A cross-domain data intelligent fusion and analysis method for public safety scenarios, characterized in that, The method includes: Step 1: Collect multi-source heterogeneous data and preprocess the multi-source heterogeneous data; Step 2: The visual events and business events in the preprocessed multi-source heterogeneous data are uniformly modeled into computable behavioral units to obtain the cross-domain entity association confidence. Step 3: Perform large-scale agent analysis, adopting an agent-driven hypothesis-to-verification-to-reflection analysis paradigm, complete the arrangement of natural language into directed acyclic graph, and obtain the large-scale agent analysis results. Step 4: Based on the analysis results of the large model agent, reconstruct the cross-domain behavior chain story, perform multi-dimensional verification on the analysis results of the large model agent, and obtain a comprehensive causal confidence score. Step 5: Organize the evidence chain of the overall analysis results and output the judgment results that support the investigation decision.

2. The cross-domain data intelligent fusion analysis method for public safety scenarios according to claim 1, characterized in that, In step 1, the preprocessing of the multi-source heterogeneous data includes: protocol parsing, format standardization, and quality cleaning.

3. The cross-domain data intelligent fusion analysis method for public safety scenarios according to claim 1, characterized in that, Step 2 includes: Step 21: Extract features from any trajectory within the time window and concatenate them into a comprehensive feature vector; Step 22: The dynamic time warping algorithm is used to calculate the trajectory similarity based on the comprehensive feature vector; Step 23: Based on the trajectory similarity, construct the edge weights between the targets in the video structured data and the targets covered in the business data, thereby constructing a cross-domain association network; Step 24: Perform graph mining on the constructed cross-domain association network based on the Infomap algorithm to divide the network into companion gang communities; Step 25: Calculate the cross-domain entity association confidence based on the trajectory similarity and the community mining results obtained based on the Infomap algorithm.

4. The cross-domain data intelligent fusion analysis method for public safety scenarios according to claim 3, characterized in that, In step 21, the visual trajectory and the public security business trajectory are mapped to a unified geographic grid, i.e., a time window space. The extracted features include trajectory geometry, movement speed features, and activity pattern features. The trajectory geometry includes: trajectory length, mean turning angle, straightness, and number of covered grids. The movement speed features include: average speed, speed variance, and percentage of stationary time. The activity pattern features include: distribution of active periods, activity frequency, and periodicity indicators.

5. The cross-domain data intelligent fusion analysis method for public safety scenarios according to claim 3, characterized in that, In step 25, the formula for calculating the cross-domain entity association confidence Conf(u,v) is as follows: Conf(u,v)=α×Sim_DTW+β×I (same community) Wherein, Sim_DTW is the trajectory similarity, with a value of 0-1; u and v represent two personnel entities to be compared; I(same community) is an indicator function; it is 1 if u and v are in the same community, and 0 otherwise. α and β are weighting coefficients, and α+β=1.

6. The cross-domain data intelligent fusion analysis method for public safety scenarios according to claim 1, characterized in that, Step 3 includes: Step 31: The user raises an open-ended assessment question; Step 32: For the open-ended judgment problem, the large language model generates multiple hypothetical paths based on investigative knowledge and data structure; Step 33: For each hypothetical planning path, convert the hypothesis into a DAG analysis graph. Each node in the graph represents a tool call, the edges represent data dependencies, and the directed edges represent the data flow direction. Step 34: Based on the DAG analysis graph, the tool starts executing to obtain the analysis results of the large model intelligent agent. According to the edge type, it is identified whether to execute serially, in parallel or conditionally, and calls the required tools to obtain data from various business systems. The tools called include video parsing, target retrieval, target deployment and correlation analysis.

7. The cross-domain data intelligent fusion analysis method for public safety scenarios according to claim 1, characterized in that, Step 4 includes: Step 41: Reconstruct the behavior chain by merging the input cross-domain event data along the time axis to generate a continuous behavior chain sequence. Step 42: Using the propensity score matching causal inference method, a control group is constructed to eliminate confounding factors, estimate the true causal effect, and obtain the statistical verification results. Step 43 proposes a causal confidence scoring model, and conducts a comprehensive causal confidence score based on statistical test results to identify the division of labor and crime patterns within the gang.

8. The cross-domain data intelligent fusion analysis method for public safety scenarios according to claim 1, characterized in that, In step 5, all intermediate results, original data references, statistical indicators, and causal confidence scores from the analysis process are forcibly bound into a structured chain of evidence as the overall analysis result, and an assessment report is output. The chain of evidence adopts a structured JSON format.

9. A cross-domain data intelligent fusion analysis device for public safety scenarios, characterized in that, The device includes: The data acquisition and preprocessing unit acquires multi-source heterogeneous data and preprocesses the multi-source heterogeneous data; The cross-domain entity association unit models visual events and business events in the preprocessed multi-source heterogeneous data into a computable behavioral unit to obtain the cross-domain entity association confidence. The large model agent analysis unit performs large model agent analysis, adopts an agent-driven hypothesis-to-verification-to-reflection analysis paradigm, completes the arrangement of natural language into directed acyclic graph, and obtains the large model agent analysis results. The causal inference and behavior chain reconstruction unit reconstructs cross-domain behavior chain stories based on the analysis results of the large model intelligent agent, performs multi-dimensional verification of the analysis results of the large model intelligent agent, and obtains a comprehensive score of causal confidence. The evidence chain organization and analysis results output unit organizes the evidence chain of the overall analysis results and outputs analysis results that support investigative decisions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the cross-domain data intelligent fusion analysis method for public safety scenarios as described in any one of claims 1 to 8.