Law enforcement recording instrument video data management method and system

By constructing a behavioral pattern domain on the law enforcement recorder terminal and performing causal inference and risk management in the cloud, the shortcomings of video data management in the existing system are solved, enabling refined management and efficient retrieval of the law enforcement process, and improving the security and retrieval efficiency of law enforcement data.

CN122137981APending Publication Date: 2026-06-02SHENZHEN QUANDATONG TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN QUANDATONG TECH CO LTD
Filing Date
2026-01-23
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing law enforcement recorder video data management systems lack a collaborative mechanism between terminals and the cloud, making it difficult to identify key nodes and status changes during law enforcement, resulting in low retrieval efficiency, high management costs, and a lack of differentiated storage and protection strategies.

Method used

By constructing a law enforcement behavior pattern domain on the terminal side, identifying changes in behavior status and segmenting them into event units, performing causal inference and risk management on the cloud side, constructing a dispute risk topography map, and achieving efficient retrieval by combining a spatiotemporal inverted index.

Benefits of technology

It enables refined management of the law enforcement process, improves the efficiency of locating and managing key video clips, reduces storage and management costs, and enhances the security and retrieval accuracy of evidence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122137981A_ABST
    Figure CN122137981A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for managing video data from law enforcement recorders, including: collecting video and audio data, generating law enforcement session identifiers, and associating them to form a dataset; jointly analyzing multi-source data on the edge to construct a behavior pattern domain and continuously monitor stability changes; identifying collapse points as boundaries to segment video data and generate a sequence of law enforcement event units; integrating video, audio, context, and semantic data on the cloud to generate a causal consistent description; constructing a risk topography map based on the semantic description to determine differentiated video management strategies; constructing a spatiotemporal inverted index structure to locate event units and output video data results. This invention achieves accurate segmentation, risk identification, and efficient retrieval management of law enforcement recorder video data through edge-cloud collaborative video structuring and causal inference mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video data management technology, and in particular to a method and system for managing video data from law enforcement recorders. Background Technology

[0002] With the continuous improvement of informatization in law enforcement, body cameras have been widely used in on-site law enforcement evidence collection to record video and audio of the entire process. Existing body cameras typically collect, store, and provide basic playback of video and audio data, uploading it to a backend management platform via wired or wireless means for centralized storage and management. This forms a basic application model where terminal devices are responsible for data collection, and the cloud platform handles centralized storage and playback. This technology, to a certain extent, improves the transparency and reliability of law enforcement, providing fundamental data support for post-event review, complaint handling, and judicial evidence collection.

[0003] However, in the existing management model, which primarily relies on terminal data collection and centralized cloud storage, the interaction between terminals and the cloud is mainly based on data uploading and passive storage, lacking collaborative analysis capabilities oriented towards the law enforcement process. The organization of law enforcement recorder video data often uses fixed time periods or single files as management units, making it difficult to effectively identify and characterize the internal stages of the law enforcement process. Video data is typically stored in continuous chronological order, failing to distinguish between different law enforcement behavior states or key law enforcement nodes. When law enforcement behavior abruptly changes from one state to another, existing technologies, whether on the terminal or cloud side, struggle to promptly identify and generate structured markers, resulting in a lack of semantic information about the law enforcement process and hindering refined management.

[0004] Current technologies primarily rely on cloud platforms for post-event analysis of law enforcement recorder video data processing and retrieval, lacking a collaborative mechanism between the terminal and the cloud, and risk-driven management capabilities. Video data retrieval methods are typically based on single criteria such as enforcement time, enforcement personnel, or case number, resulting in coarse-grained retrieval. When locating specific enforcement stages or high-risk segments within a large volume of video data, repeated manual screening is often required, leading to low retrieval efficiency and high management costs. Furthermore, video data of different risk levels generally employs a uniform storage and protection strategy, lacking differentiated management methods based on end-to-cloud collaboration. This increases cloud storage and computing resource consumption and hinders the focused protection of critical evidence.

[0005] Therefore, how to provide methods and systems for managing video data from law enforcement recorders is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a method and system for managing video data from law enforcement recorders. This invention achieves refined management of video data from the entire law enforcement process by structurally processing and intelligently managing the video and audio data collected by law enforcement recorders. It comprehensively utilizes techniques such as law enforcement behavior pattern analysis, causal inference integration, and spatiotemporal inverted indexing. This invention focuses on changes in the state of law enforcement behavior, constructing a law enforcement behavior pattern domain and identifying collapse points in law enforcement behavior patterns. It segments continuous law enforcement videos into law enforcement event units with clear semantic boundaries, generates causally consistent semantic descriptions of law enforcement through four-stream causal inference integration, and further constructs a dispute risk topology map to determine differentiated management strategies. Finally, it achieves efficient retrieval through a video fast retrieval structure based on spatiotemporal inverted indexing. This invention accurately reflects changes in the law enforcement stage, proactively identifies potential dispute risks, improves the location and management efficiency of key video segments, and possesses advantages such as high management precision, high retrieval efficiency, and strong evidence protection capabilities.

[0007] A law enforcement recorder video data management method according to an embodiment of the present invention includes:

[0008] By collecting continuous video and audio data through law enforcement recorders, law enforcement session identifiers are generated. These identifiers are then associated with the video and audio data to form a law enforcement session dataset.

[0009] On the edge, based on the law enforcement session dataset, video data, audio data and law enforcement context information are jointly analyzed to construct a law enforcement behavior pattern domain that reflects the state of law enforcement behavior, and the stability of the law enforcement behavior pattern domain is continuously monitored.

[0010] When a state change is detected in the law enforcement behavior pattern domain, the law enforcement behavior pattern collapse point is determined, and the video data is segmented with the law enforcement behavior pattern collapse point as the boundary to generate law enforcement event units arranged in chronological order.

[0011] A four-stream causal inference integrator is built on the cloud side. For each law enforcement event unit, the video behavior stream, audio speech stream, law enforcement context stream, and semantic inference stream corresponding to the law enforcement event unit are integrated to generate a causally consistent law enforcement semantic description result.

[0012] Based on the semantic description results of law enforcement corresponding to each law enforcement event unit, a dispute risk topography map is constructed to determine the differentiated management strategy for video data corresponding to each law enforcement event unit.

[0013] Based on the time information, geographical location information, semantic features of law enforcement behavior, and intensity of dispute risk of law enforcement event units, index entries are established, and a video fast retrieval structure based on spatiotemporal inverted index is constructed to locate target law enforcement event units and output video data according to the corresponding differentiated management strategies.

[0014] Optionally, the video data and audio data include video frame sequence data continuously collected by the law enforcement recorder during the law enforcement process and synchronously collected voice signal data. The video frame sequence data includes the image, the timestamp, and the corresponding frame number information. The voice signal data includes audio sampling data, audio timestamp, and volume intensity information.

[0015] Optionally, generating a law enforcement session identifier means generating a unique identifier code based on the law enforcement start time, the law enforcement recorder device number, the law enforcement officer's identity identifier, and the law enforcement task number when the law enforcement recorder starts recording law enforcement video, and using the unique identifier code as a unified session identifier for the video and audio data of the law enforcement process.

[0016] Optionally, the construction of a law enforcement behavior pattern domain reflecting the state of law enforcement behavior, and the continuous monitoring of the stability of the law enforcement behavior pattern domain, includes:

[0017] On the device side, based on the law enforcement conversation dataset, video frame sequences, audio signal sequences, time information, geographic location information, law enforcement personnel identity information and law enforcement task information are aligned with a unified time reference to generate aligned multi-source conversation data;

[0018] On the aligned multi-source conversation data, video behavior features, audio intonation features, and law enforcement context state features are extracted by time step, standardized, and a time index is established.

[0019] A law enforcement behavior pattern domain reflecting the state of law enforcement behavior is constructed. The law enforcement behavior pattern domain consists of a video behavior subdomain, an audio speech subdomain, a law enforcement context state subdomain, a device posture and scene structure subdomain, and a cross-channel consistency subdomain. Each subdomain is combined into a pattern domain entry corresponding to a time step according to a unified coding scheme and written into the session data structure.

[0020] Continuous stability monitoring of law enforcement behavior pattern domains is conducted to establish mutation indicator sequences, sliding consistency sequences, phase persistence counting sequences, and context rule conflict sequences.

[0021] Within the synchronization window, the monitoring results are comprehensively judged based on the multi-source triggering conditions. When the mutation indicator sequence, the sliding consistency sequence, and the context rule conflict sequence meet the preset triggering combination rules, the behavior pattern collapse pre-mark is output, and the event unit boundary candidate list is generated.

[0022] Optionally, the generation of law enforcement event units arranged in chronological order includes:

[0023] Obtain the event unit boundary candidate list and the corresponding behavior pattern collapse pre-label, and obtain the mutation indicator sequence, lens motion anomaly indicator sequence, cross-channel consistency deviation sequence, stage persistence count sequence and context rule conflict sequence corresponding to the event unit boundary candidate list;

[0024] The law enforcement session dataset is discretized into time steps arranged in chronological order according to a unified time base, the event unit boundary candidate list is mapped to the boundary candidate time step set, and the index relationship of the boundary candidate time steps is established.

[0025] For each boundary candidate time step, read the values ​​of the mutation indicator sequence, the lens motion anomaly indicator sequence, the cross-channel consistency deviation sequence, the stage persistence count sequence, and the context rule conflict sequence corresponding to the boundary candidate time step to form a boundary candidate judgment element set, and generate the boundary candidate confidence level according to the preset multi-element comprehensive judgment rule.

[0026] When the boundary candidate time step meets the necessary triggering conditions and the confidence level of the boundary candidate reaches the lower limit of the preset confidence threshold range, and the camera motion anomaly indication sequence does not meet the individual triggering conditions, the boundary candidate time step is determined as the law enforcement behavior pattern collapse point, forming a law enforcement behavior pattern collapse point sequence arranged in chronological order.

[0027] Using the collapse point sequence of law enforcement behavior patterns as the segmentation boundary, the video and audio data associated with law enforcement session identifiers are segmented to generate law enforcement event units arranged in chronological order.

[0028] Optionally, generating a causally consistent law enforcement semantic description result includes:

[0029] On the cloud side, a four-stream causal inference integrator is constructed based on law enforcement event units arranged in chronological order. The four-stream causal inference integrator consists of a four-stream alignment unit, a causal adjudication unit, and a semantic generation unit, and establishes a local time index for each law enforcement event unit.

[0030] The four-stream alignment unit receives video behavior stream, audio speech stream, law enforcement context stream, and semantic inference stream respectively, performs unified time reference alignment and sequence correction on the four streams, performs missing segment completion, duplicate segment removal, abnormal timestamp correction and cross-channel delay correction, generates alignment credibility identifier, alignment version identifier and alignment time reference identifier based on device posture and camera motion estimation data, and outputs four-stream alignment entries.

[0031] The causal adjudication unit receives four-flow alignment entries, parses the allowed and prohibited relationships of the enforcement stage, generates a causal constraint set, executes a counterfactual shielding mechanism to shield target stage behavior entries that are not reachable from the current stage, executes context conservation verification to eliminate behavior entries that are inconsistent with the enforcement context flow, and outputs a causal consistency candidate behavior entry set.

[0032] The semantic generation unit receives four-stream aligned entries and a set of causal consistent candidate behavior entries, performs cross-channel divergence resolution, stage boundary inheritance and convergence, merges the two and generates causal consistent law enforcement semantic description entries according to time steps, and merges them into a semantic tag set and a future behavior window according to local time index within the law enforcement event unit.

[0033] The semantic tag set, future behavior window, and corresponding law enforcement event unit's time range, geographical location information, and law enforcement session identifier are associated and stored as a causally consistent law enforcement semantic description result.

[0034] Optionally, the construction of the dispute risk topographic map and the determination of differentiated management strategies for video data corresponding to each law enforcement incident unit include:

[0035] Obtain semantic tag sets and future behavior windows, and combine them with the time range and sequence of law enforcement event units, as well as geographical location information and law enforcement session identifiers, to form time dimensions, semantic behavior dimensions and basic association information for evaluation;

[0036] Based on mutation indicator sequences, camera motion anomaly indicator sequences, cross-channel consistency deviation sequences, phase persistence counting sequences, and context rule conflict sequences, combined with semantic tag sets and future behavior windows, risk feature sets are generated for each law enforcement event unit and standardized and indexed over time.

[0037] Construct a topographic map of dispute risks. Establish a fixed-step time grid on the law enforcement session timeline and a semantic grid on the stage attributes and key behaviors of the semantic tag set. Combine the two types of grids to obtain basic grid units. Aggregate risk feature sets within the corresponding units to generate risk intensity values. Perform continuity and connectivity checks to remove isolated units. Perform regional merging to obtain four types of regions: peaks, slopes, valleys, and troughs, forming a set of topographic map unit entries.

[0038] Based on the unit entry set of the dispute risk topographic map, a differentiated management strategy is generated. When the unit entry is of the peak type, multiple copies are redundant, strong encryption and extended retention period are set. When the unit entry is of the slope type, the protection window is expanded and the redundancy level is increased. When the unit entry is of the valley type, the summary storage and compression strategy is set. When the unit entry is of the trough type, the segmented merging storage and retention period are increased.

[0039] The differentiated management strategy is bound to the video data of the corresponding law enforcement event unit, and the effective parameters of redundancy level, encryption strength and retention period are recorded to generate management strategy index entries.

[0040] Optionally, the step of constructing a video fast retrieval structure based on a spatiotemporal inverted index, locating target law enforcement event units, and outputting video data according to corresponding differentiated management strategies includes:

[0041] Read the management strategy index entries, extract the time range, geographical location information, stage attribute tags or key behavior tags in the semantic tag set, area type and risk intensity value of the dispute risk topographic map, and differentiated management strategy parameters corresponding to each law enforcement event unit, and generate the basic index data;

[0042] The time range is divided into time keys, the geographic location information is rasterized into spatial keys, the stage attribute labels or key behavior labels are encoded into semantic keys, the area type and risk intensity value are classified into risk keys, and the composite keys are combined in a preset order to form composite keys. The composite keys are then mapped to law enforcement event unit identifiers.

[0043] A fast video retrieval structure based on spatiotemporal inverted index is constructed using composite keys as the organizational unit. An inverted list is maintained for each composite key. The inverted list is stored in memory and on disk in a hierarchical manner. The inverted list with the region type of peak and the risk intensity value at the highest level resides in memory. The inverted list with high access frequency is cached with dual copies. The index version number and effective status are recorded for each inverted list.

[0044] Upon receiving a video retrieval request, the request conditions are parsed to generate a query key combination. The results are matched step by step according to the time key, spatial key, semantic key and risk key. The results in the inverted list are deduplicated and filtered by intersection to locate the target law enforcement event unit. Results belonging to the area type "peak" are returned first.

[0045] Based on the location results, a retrieval list is generated. Encryption control, redundant copy reading, and retention period verification are performed according to the corresponding differentiated management strategy parameters. Video data is output and retrieval logs and index version numbers are recorded.

[0046] The law enforcement recorder video data management system according to an embodiment of the present invention includes the following modules:

[0047] The law enforcement session acquisition module is used to collect video and audio data, generate law enforcement session identifiers, and form a law enforcement session dataset.

[0048] The law enforcement behavior pattern construction module is used to construct a law enforcement behavior pattern domain based on the law enforcement session dataset and continuously monitor the stability of the law enforcement behavior pattern domain.

[0049] The law enforcement event unit segmentation module is used to determine the collapse point of the law enforcement behavior pattern when a state change occurs in the law enforcement behavior pattern domain, and to generate law enforcement event units.

[0050] The four-stream causal inference integration module is used to integrate the video behavior stream, audio tone stream, law enforcement context stream, and semantic inference stream corresponding to the law enforcement event unit to generate a causally consistent law enforcement semantic description result.

[0051] The Dispute Risk Terrain and Strategy Generation Module is used to construct dispute risk terrain maps based on law enforcement semantic description results and determine differentiated management strategies.

[0052] The spatiotemporal inverted index retrieval module is used to construct a fast video retrieval structure based on the spatiotemporal inverted index and output video data according to a differentiated management strategy.

[0053] The beneficial effects of this invention are:

[0054] This invention employs an edge-cloud collaborative processing approach, with the edge handling law enforcement behavior perception and event segmentation, while the cloud handles semantic inference, risk management, and video retrieval. By introducing an edge-cloud collaborative method for managing law enforcement recorder video data, this invention achieves structured organization and dynamic management of video data throughout the entire law enforcement process. By constructing law enforcement behavior patterns and monitoring stability on the terminal side of video and audio data, and generating law enforcement event units when sudden changes in law enforcement behavior are detected, continuous law enforcement videos can be accurately segmented according to the stages of law enforcement behavior. This avoids the coarse-grained management methods of traditional approaches based on fixed time periods or single files, improving the accuracy and consistency of law enforcement process stage segmentation.

[0055] This invention constructs a four-stream causal inference integrator on the cloud side, unifying and integrating the video behavior stream, audio speech stream, law enforcement context stream, and semantic inference stream corresponding to law enforcement event units. This generates causally consistent semantic descriptions of law enforcement and further constructs a dispute risk topology map, enabling the proactive identification of potential dispute risks during law enforcement. By mapping law enforcement event units to different regions of the dispute risk topology map and formulating differentiated management strategies accordingly, this invention can prioritize the protection of high-risk law enforcement segments and reasonably compress and manage low-risk segments, effectively reducing storage and management costs while improving evidence security.

[0056] This invention achieves efficient location and retrieval of law enforcement event units by constructing a video fast retrieval structure based on a spatiotemporal inverted index. Combined with an edge-cloud collaborative architecture, the cloud side can significantly improve retrieval efficiency while maintaining retrieval accuracy, reducing manual screening and repetitive operations, and rapidly obtaining key law enforcement segments in application scenarios such as complaint handling, law enforcement review, and judicial evidence collection. In summary, this invention enhances the granularity of law enforcement video data management, risk control capabilities, and retrieval efficiency, while also strengthening the standardization and reliability of law enforcement evidence collection, demonstrating significant practical application value. Attached Figure Description

[0057] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0058] Figure 1 This is a flowchart of the law enforcement recorder video data management method proposed in this invention;

[0059] Figure 2 This is a schematic diagram of the structure of the law enforcement recorder video data management system proposed in this invention. Detailed Implementation

[0060] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0061] refer to Figure 1 Methods for managing video data from law enforcement recorders include:

[0062] By collecting continuous video and audio data through law enforcement recorders, law enforcement session identifiers are generated. These identifiers are then associated with the video and audio data to form a law enforcement session dataset.

[0063] On the edge, based on the law enforcement session dataset, video data, audio data and law enforcement context information are jointly analyzed to construct a law enforcement behavior pattern domain that reflects the state of law enforcement behavior, and the stability of the law enforcement behavior pattern domain is continuously monitored.

[0064] When a state change is detected in the law enforcement behavior pattern domain, the law enforcement behavior pattern collapse point is determined, and the video data is segmented with the law enforcement behavior pattern collapse point as the boundary to generate law enforcement event units arranged in chronological order.

[0065] A four-stream causal inference integrator is built on the cloud side. For each law enforcement event unit, the video behavior stream, audio speech stream, law enforcement context stream, and semantic inference stream corresponding to the law enforcement event unit are integrated to generate a causally consistent law enforcement semantic description result.

[0066] Based on the semantic description results of law enforcement corresponding to each law enforcement event unit, a dispute risk topography map is constructed to determine the differentiated management strategy for video data corresponding to each law enforcement event unit.

[0067] Based on the time information, geographical location information, semantic features of law enforcement behavior, and intensity of dispute risk of law enforcement event units, index entries are established, and a video fast retrieval structure based on spatiotemporal inverted index is constructed to locate target law enforcement event units and output video data according to the corresponding differentiated management strategies.

[0068] In this embodiment, the video data and audio data include video frame sequence data continuously collected by the law enforcement recorder during the law enforcement process and audio signal data collected synchronously. The video frame sequence data includes the image, the timestamp, and the corresponding frame number information. The audio signal data includes audio sampling data, audio timestamp, and volume intensity information.

[0069] In this embodiment, generating an enforcement session identifier means that when the law enforcement recorder starts recording enforcement video, a unique identifier is generated based on the enforcement start time, the law enforcement recorder device number, the law enforcement officer's identity identifier, and the enforcement task number, and the unique identifier is used as a unified session identifier for the video data and audio data of the enforcement process.

[0070] In this embodiment, the construction of a law enforcement behavior pattern domain reflecting the state of law enforcement behavior, and the continuous monitoring of the stability of the law enforcement behavior pattern domain, includes:

[0071] On the device side, based on the law enforcement conversation dataset, video frame sequences, audio signal sequences, time information, geographic location information, law enforcement personnel identity information and law enforcement task information are aligned with a unified time reference to generate aligned multi-source conversation data;

[0072] On the aligned multi-source conversation data, video behavior features, audio intonation features, and law enforcement context state features are extracted by time step, standardized, and a time index is established.

[0073] A law enforcement behavior pattern domain reflecting the state of law enforcement behavior is constructed. The law enforcement behavior pattern domain consists of a video behavior subdomain, an audio speech subdomain, a law enforcement context state subdomain, a device posture and scene structure subdomain, and a cross-channel consistency subdomain. Each subdomain is combined into a pattern domain entry corresponding to a time step according to a unified coding scheme and written into the session data structure.

[0074] Continuous stability monitoring is performed on the law enforcement behavior pattern domain. A mutation indicator sequence, a sliding consistency sequence, a stage persistence count sequence, and a context rule conflict sequence are established. The mutation indicator sequence is used to record the markers when the difference between adjacent time steps exceeds the rule threshold. The sliding consistency sequence is used to record cross-channel consistency deviations within a fixed-length sliding window. The stage persistence count sequence is used to record the time step counts of the same state. The context rule conflict sequence is used to record state markers that are inconsistent with the law enforcement process rule base.

[0075] Within the synchronization window, the monitoring results are comprehensively judged based on multi-source triggering conditions. When the mutation indicator sequence, the sliding consistency sequence, and the context rule conflict sequence meet the preset triggering combination rules, the behavior pattern collapse pre-mark is output, and an event unit boundary candidate list is generated. The preset triggering combination rules are as follows:

[0076] When mutation indicator sequences appear consecutively within the same synchronization window and the duration exceeds the minimum stability threshold, a significant trend of behavioral state change is identified.

[0077] Within the synchronization window, if the sliding consistency sequence changes from a stable state to an unstable state, and the unstable state continues to cover the main time period of the synchronization window, the stability of the current behavior pattern is considered to be disrupted.

[0078] Within the synchronization window, if a context rule conflict sequence is triggered and the behavior state indicated by the conflict sequence is inconsistent with the current law enforcement context rule, a context constraint conflict is identified.

[0079] When significant behavioral state change trends, behavioral pattern instability, and context constraint conflicts occur simultaneously within the same synchronization window, the time position corresponding to the synchronization window is marked as a behavioral pattern collapse pre-marker, and the corresponding time position is added to the event unit boundary candidate list.

[0080] The law enforcement behavior pattern domain refers to a comprehensive behavioral state space constructed during a law enforcement session to depict the overall state of law enforcement behavior over time. Based on video and audio data collected by law enforcement recorders and corresponding law enforcement context information, the law enforcement behavior pattern domain jointly represents multi-source information such as law enforcement actions, language intonation, environmental changes, and task attributes. It maps the law enforcement process onto a continuous time axis as a set of continuously monitorable behavioral states. The law enforcement behavior pattern domain can reflect the stability and changing trends of law enforcement behavior at different stages, transforming law enforcement behavior from an unstructured process presented only in raw video into an analyzable object with state boundaries and evolutionary characteristics. This provides a foundation for identifying sudden changes in law enforcement behavior states, segmenting law enforcement event units, and subsequent semantic analysis and risk management.

[0081] In this embodiment, the generation of law enforcement event units arranged in chronological order includes:

[0082] Obtain the event unit boundary candidate list and the corresponding behavior pattern collapse pre-label, and obtain the mutation indicator sequence, lens motion anomaly indicator sequence, cross-channel consistency deviation sequence, stage persistence count sequence and context rule conflict sequence corresponding to the event unit boundary candidate list;

[0083] The law enforcement session dataset is discretized into time steps arranged in chronological order based on a unified time base. The event unit boundary candidate list is mapped to a set of boundary candidate time steps, and an index relationship is established for the boundary candidate time steps. Specifically, the mapping of the event unit boundary candidate list to the boundary candidate time step set is as follows:

[0084] Using the timestamps of video data in the law enforcement session dataset as a unified time reference, the video and audio data are synchronized and aligned, and the law enforcement session dataset is discretized in time according to a preset time granularity to form a time step sequence arranged in chronological order.

[0085] For each candidate boundary time point in the event unit boundary candidate list, locate the time step that is closest to and no later than the corresponding timestamp in the time step sequence based on the corresponding timestamp, and mark the time step as a boundary candidate time step;

[0086] When multiple candidate boundary time points are mapped to the same time step, the time steps are merged into a single candidate boundary time step, and the source information of the corresponding multiple candidate boundaries is recorded.

[0087] All the boundary candidate time steps obtained by mapping are sorted in chronological order to form a set of boundary candidate time steps;

[0088] For each boundary candidate time step, the values ​​of the mutation indicator sequence, the lens motion anomaly indicator sequence, the cross-channel consistency deviation sequence, the stage persistence count sequence, and the context rule conflict sequence corresponding to the boundary candidate time step are read to form a boundary candidate judgment element set. The boundary candidate confidence level is generated according to the preset multi-element comprehensive judgment rule. The multi-element comprehensive judgment rule includes taking the mutation indicator sequence and the context rule conflict sequence as the necessary triggering condition, taking the cross-channel consistency deviation sequence and the stage persistence count sequence as the enhancement condition, and taking the lens motion anomaly indicator sequence as the suppression condition.

[0089] When a candidate boundary time step meets the necessary triggering conditions and the confidence level of the candidate boundary reaches the lower limit of the preset confidence threshold range, and the abnormal camera movement indication sequence does not meet the individual triggering conditions, the candidate boundary time step is determined as the law enforcement behavior pattern collapse point, forming a law enforcement behavior pattern collapse point sequence arranged in chronological order, wherein:

[0090] The confidence level of a boundary candidate is obtained by statistically analyzing the triggering status, duration, and coverage ratio of mutation indication sequences, sliding consistency sequences, and context rule conflict sequences at the boundary candidate time step within the synchronization window. The statistical results corresponding to each sequence are aggregated according to a pre-set weighting relationship to form a comprehensive confidence result that reflects the overall triggering strength of the boundary candidate time step. The comprehensive confidence result is then verified for consistency by combining the continuity of confidence changes between the boundary candidate time step and adjacent time steps, and determined as the boundary candidate confidence level.

[0091] Using the collapse point sequence of law enforcement behavior patterns as the segmentation boundary, the video and audio data associated with law enforcement session identifiers are segmented to generate law enforcement event units arranged in chronological order. Specifically, the segmentation of the video and audio data associated with law enforcement session identifiers is as follows:

[0092] According to the time sequence of each collapse point in the law enforcement behavior pattern collapse point sequence, the video data and audio data associated with the same law enforcement session identifier are time-aligned, and the time interval between adjacent collapse points is determined.

[0093] Extract consecutive video frame data and corresponding audio sampling data within each time interval into an independent data segment, while maintaining the time synchronization relationship between the video frame data and the audio sampling data;

[0094] Each data segment is assigned a corresponding start time, end time, and law enforcement session identifier. The data segments are treated as law enforcement event units and arranged in chronological order to form a law enforcement event unit sequence.

[0095] In this embodiment, generating a causally consistent law enforcement semantic description result includes:

[0096] On the cloud side, a four-stream causal inference integrator is constructed based on law enforcement event units arranged in chronological order. The four-stream causal inference integrator consists of a four-stream alignment unit, a causal adjudication unit, and a semantic generation unit, and establishes a local time index for each law enforcement event unit.

[0097] The four-stream alignment unit receives video behavior stream, audio speech stream, law enforcement context stream, and semantic inference stream respectively, performs unified time reference alignment and sequence correction on the four streams, performs missing segment completion, duplicate segment removal, abnormal timestamp correction and cross-channel delay correction, generates alignment credibility identifier, alignment version identifier and alignment time reference identifier based on device posture and camera motion estimation data, and outputs four-stream alignment entries.

[0098] The causal adjudication unit receives the four-flow alignment entries, parses the permitted and prohibited relationships of the enforcement stages, generates a causal constraint set, executes a counterfactual masking mechanism to mask target stage behavior entries that are not reachable from the current stage, executes a context conservation check to remove behavior entries inconsistent with the enforcement context flow, and outputs a causal consistency candidate behavior entry set. Specifically, the counterfactual masking mechanism for masking target stage behavior entries that are not reachable from the current stage involves:

[0099] Based on the set of causal constraints, read the stage identifier corresponding to the current law enforcement stage, and determine the set of allowed subsequent stages associated with the stage identifier in the set of causal constraints.

[0100] For each target stage behavior entry marked in the four-stream alignment entry, compare one by one whether the stage identifier corresponding to the target stage behavior entry is included in the set of allowed successor stages;

[0101] When the stage identifier of the target stage behavior entry is not included in the set of allowed successor stages, the target stage behavior entry is marked as an unreachable behavior entry and is masked from the candidate set, and only behavior entries whose stage identifiers satisfy the constraints of the set of allowed successor stages are retained.

[0102] The semantic generation unit receives four-stream aligned entries and a set of causal consistent candidate behavior entries, performs cross-channel divergence resolution, stage boundary inheritance, and convergence, merges the two, and generates causal consistent law enforcement semantic description entries according to time steps. Within the law enforcement event unit, these are merged according to local time indices to form a semantic tag set and a future behavior window. Specifically, the cross-channel divergence resolution, stage boundary inheritance, and convergence are performed as follows:

[0103] Based on the video behavior stream, audio speech stream, law enforcement context stream and semantic inference stream within the same time step, the behavior description entries corresponding to each channel are aligned and compared to identify the divergent entries that differ in behavior orientation, stage attributes or contextual constraints.

[0104] When a cross-channel deviation entry is detected, the deviation entries are filtered and merged according to a preset channel priority order and temporal continuity rules. Behavioral description entries that are consistent with the majority of channels and have higher temporal continuity are retained. The preset channel priority order is: law enforcement context stream takes precedence over video behavior stream, video behavior stream takes precedence over audio speech stream, and audio speech stream takes precedence over semantic inference stream. The temporal continuity rules are as follows:

[0105] The behavior description entries within adjacent time steps are compared. When the same behavior description entry remains consistent or only undergoes limited changes in consecutive time steps, the behavior description entry is deemed to satisfy the temporal continuity.

[0106] When a behavior description item appears only in a single time step and is not supported by similar description items in the preceding and following time steps, the behavior description item is deemed not to satisfy temporal continuity.

[0107] When multiple divergent entries exist simultaneously, the behavioral description entries that appear more frequently and last longer within consecutive time steps are preferred as entries that satisfy the time continuity rule.

[0108] At the start time step of the law enforcement event unit, the stage attribute tag at the end of the previous law enforcement event unit is inherited and used continuously in the current event unit until a new stage switching condition is detected.

[0109] When a phase switching condition is detected within the current law enforcement event unit, the inheritance process of the phase attribute label is terminated, and phase boundary convergence is completed at the corresponding time step to generate a new phase attribute label.

[0110] Based on the behavioral description entries after completing the divergence resolution and stage boundary inheritance and convergence, causally consistent law enforcement semantic description entries are generated according to time steps.

[0111] The semantic tag set, future behavior window, and corresponding law enforcement event unit's time range, geographical location information, and law enforcement session identifier are associated and stored as a causally consistent law enforcement semantic description result.

[0112] In this embodiment, the construction of the dispute risk topographic map and the determination of differentiated management strategies for video data corresponding to each law enforcement incident unit include:

[0113] Obtain semantic tag sets and future behavior windows, and combine them with the time range and sequence of law enforcement event units, as well as geographical location information and law enforcement session identifiers, to form time dimensions, semantic behavior dimensions and basic association information for evaluation;

[0114] Based on mutation indicator sequences, camera motion anomaly indicator sequences, cross-channel consistency deviation sequences, phase persistence counting sequences, and context rule conflict sequences, combined with semantic tag sets and future behavior windows, risk feature sets are generated for each law enforcement event unit and standardized and indexed over time.

[0115] A topographic map of disputed risks is constructed by establishing a fixed-step time grid on the law enforcement session timeline and a semantic grid based on the stage attributes and key behaviors of the semantic tag set. These two types of grids are combined to obtain basic grid units. Risk intensity values ​​are generated by aggregating risk feature sets within corresponding units. Continuity and connectivity checks are performed to eliminate isolated units. Region merging is then performed to obtain four types of regions: peaks, slopes, valleys, and troughs, forming a set of topographic map unit entries. The four types of regions—peaks, slopes, valleys, and troughs—refer to:

[0116] Peak regions refer to areas in dispute risk topographic maps where the corresponding risk intensity value is at the highest level among adjacent basic grid cells, and exhibits risk clustering characteristics in both the time axis and semantic grid directions.

[0117] Slope areas refer to regions in disputed risk topographic maps where the corresponding risk intensity value shows a unidirectional or multidirectional increasing or decreasing trend relative to adjacent basic grid units, located between peak and valley areas, and where the risk intensity changes continuously.

[0118] Valley regions refer to areas in dispute risk topographic maps where the corresponding risk intensity value is at the lowest level among adjacent basic grid cells, and exhibits low-risk clustering characteristics in both the time axis and semantic grid directions.

[0119] The groove area refers to a low-risk band-shaped area in the dispute risk topographic map where the corresponding risk intensity value is lower than the adjacent areas on both sides, but it extends continuously in the time axis direction, forming a distribution along the time axis.

[0120] Based on the unit entry set of the disputed risk topographic map, a differentiated management strategy is generated. When a unit entry belongs to the peak type, multiple copies are used for redundancy, strong encryption, and extended retention period. When a unit entry belongs to the slope type, the protection window is expanded and the redundancy level is increased. When a unit entry belongs to the valley type, summary storage and compression strategies are implemented. When a unit entry belongs to the trough type, segmented merging and extended retention period are implemented. Specifically, the differentiated management strategy generated based on the unit entry set of the disputed risk topographic map is as follows:

[0121] Read the region type identifier of each topographic map unit entry, and limit the region type identifier to peak type, slope type, valley type, or trough type;

[0122] Based on the region type identifier, a corresponding set of management policy parameters is selected from a pre-established management policy mapping relationship. This set of management policy parameters includes redundancy parameters, encryption parameters, storage format parameters, and retention period parameters. The pre-established management policy mapping relationship is as follows:

[0123] When the region type is identified as peak type, the corresponding video data entry is associated with multiple copy redundancy parameters, strong encryption parameters, and extended retention period parameters;

[0124] When the area type is identified as slope type, the corresponding video data entry is associated with the protection window extension parameter and the redundancy level enhancement parameter;

[0125] When the region type is identified as valley type, the corresponding video data entry is associated with summary storage parameters and compressed storage parameters;

[0126] When the region type is identified as a groove type, the corresponding video data entries are associated with the segmentation and merging saving parameters and the saving period adjustment parameters.

[0127] The selected set of management strategy parameters is associated with the corresponding video data entries to form a set of differentiated management strategies that correspond one-to-one with the topographic map unit entries.

[0128] The differentiated management strategy is bound to the video data of the corresponding law enforcement event unit, and the effective parameters of redundancy level, encryption strength and retention period are recorded to generate management strategy index entries.

[0129] In this embodiment, the construction of a video fast retrieval structure based on a spatiotemporal inverted index, locating the target law enforcement event unit, and outputting video data according to the corresponding differentiated management strategy includes:

[0130] Read the management strategy index entries, extract the time range, geographical location information, stage attribute tags or key behavior tags in the semantic tag set, area type and risk intensity value of the dispute risk topographic map, and differentiated management strategy parameters corresponding to each law enforcement event unit, and generate the basic index data;

[0131] The time range is divided into time keys, the geographic location information is rasterized into spatial keys, the stage attribute tags or key behavior tags are encoded into semantic keys, and the area type and risk intensity value are classified into risk keys. These are combined in a preset order to form a composite key, and the composite key is mapped to the law enforcement event unit identifier. The preset order is as follows: first, the time key is used as the first-level index key; second, the spatial key is used as the second-level index key; third, the semantic key is used as the third-level index key; and finally, the risk key is used as the fourth-level index key. The composite key is formed by combining the time key, spatial key, semantic key, and risk key in that order.

[0132] A fast video retrieval structure based on spatiotemporal inverted index is constructed using composite keys as the organizational unit. An inverted list is maintained for each composite key. The inverted list is stored in memory and on disk in a hierarchical manner. The inverted list with the region type of peak and the risk intensity value at the highest level resides in memory. The inverted list with high access frequency is cached with dual copies. The index version number and effective status are recorded for each inverted list.

[0133] Upon receiving a video retrieval request, the system parses the request conditions to generate a query key combination. This combination is then matched step-by-step by time key, spatial key, semantic key, and risk key. The inverted list results are deduplicated and subjected to intersection / union filtering to locate the target law enforcement event unit. Results belonging to the "peak" region type are returned first. Specifically, the deduplication and intersection / union filtering of the inverted list results involves:

[0134] Based on the time key, spatial key, semantic key, and risk key in the query key combination, the corresponding inverted index is used to read the list of matching law enforcement event unit identifiers.

[0135] Perform deduplication on the read list of multiple law enforcement event unit identifiers to eliminate duplicate identifiers caused by multi-key hits, and obtain a deduplicated candidate identifier set;

[0136] According to the matching order of the query keys, the intersection operation is performed on the deduplicated candidate identifier set in turn to remove law enforcement event unit identifiers that do not simultaneously meet the currently parsed query conditions;

[0137] When there are optional matching items in the query conditions, perform a union operation on the corresponding law enforcement event unit identifier set, and include the union result into the candidate identifier set;

[0138] The set of law enforcement event unit identifiers after deduplication and intersection filtering is used as the target law enforcement event unit location result;

[0139] Based on the location results, a retrieval list is generated. Encryption control, redundant copy reading, and retention period verification are performed according to the corresponding differentiated management strategy parameters. Video data is output and retrieval logs and index version numbers are recorded.

[0140] refer to Figure 2 The law enforcement recorder video data management system includes the following modules:

[0141] The law enforcement session acquisition module is used to collect video and audio data, generate law enforcement session identifiers, and form a law enforcement session dataset.

[0142] The law enforcement behavior pattern construction module is used to construct a law enforcement behavior pattern domain based on the law enforcement session dataset and continuously monitor the stability of the law enforcement behavior pattern domain.

[0143] The law enforcement event unit segmentation module is used to determine the collapse point of the law enforcement behavior pattern when a state change occurs in the law enforcement behavior pattern domain, and to generate law enforcement event units.

[0144] The four-stream causal inference integration module is used to integrate the video behavior stream, audio tone stream, law enforcement context stream, and semantic inference stream corresponding to the law enforcement event unit to generate a causally consistent law enforcement semantic description result.

[0145] The Dispute Risk Terrain and Strategy Generation Module is used to construct dispute risk terrain maps based on law enforcement semantic description results and determine differentiated management strategies.

[0146] The spatiotemporal inverted index retrieval module is used to construct a fast video retrieval structure based on the spatiotemporal inverted index and output video data according to a differentiated management strategy.

[0147] Example 1:

[0148] To verify the feasibility of this invention in practice, it was applied to security patrols in a large commercial park. The park covers approximately 200,000 square meters and includes multiple office buildings, underground parking, and public activity areas. Security personnel within the park are required to conduct round-the-clock patrols to prevent emergencies and record the patrol process. During patrols, security personnel wear body-worn recording devices with video and audio capture capabilities to record the on-site situation, personnel interactions, and the handling of any abnormal events.

[0149] In traditional management methods, inspection videos generated by personal recording devices are stored as complete files, with each video typically lasting 20 to 40 minutes. When customer complaints, internal disputes, or the need to review the inspection process occur, managers need to replay the entire video and manually search for key segments. This is especially problematic when personnel disputes, equipment malfunctions, or sudden environmental changes occur during the inspection process, resulting in low efficiency in locating key segments and high costs for video management and use.

[0150] In this embodiment, the video data management method described in this invention is introduced. At the start of an inspection, the personal recording device collects continuous video and audio data at the device's edge and generates a corresponding inspection session identifier. This session identifier is then locally associated with the collected video and audio data to form an inspection session dataset. This process is completed locally on the device, without the need to immediately upload the complete original video.

[0151] During the inspection process, the personal recording device at the edge performs joint analysis on video data, audio data, and contextual information such as inspection time, inspection area, and task type based on the inspection session dataset to construct a behavior pattern domain reflecting the inspection behavior status. The device continuously monitors the stability of the behavior pattern domain. When the inspection status changes from normal inspection to abnormal handling or personnel communication, the behavior pattern domain will show a significant change. At this time, the device at the edge identifies the corresponding behavior pattern collapse point and segments the local video data using this collapse point as the boundary, generating multiple inspection event units arranged in chronological order. Each event unit corresponds to a relatively independent inspection behavior stage.

[0152] After the inspection is completed, the personal recording device uploads the generated inspection event unit and its associated time and location information to the cloud management platform. The cloud platform constructs a four-stream causal inference integrator to integrate the video behavior stream, audio context stream, inspection context stream, and semantic inference stream corresponding to the inspection event unit, generating a causally consistent inspection semantic description result. Based on the inspection semantic description result, the cloud constructs a risk assessment result graph to determine the video management strategy for different inspection event units. Event units involving abnormal handling or personnel conflicts are marked as high-risk and adopt enhanced storage and access control strategies; ordinary inspection event units adopt conventional storage strategies.

[0153] During internal reviews or customer dispute resolution, managers can initiate video query requests through the management platform. The system builds index entries based on the time information, spatial location information, behavioral semantic features, and risk intensity of the inspection event unit, and uses a video fast retrieval structure based on spatiotemporal inverted index to locate the target inspection event unit, thereby directly retrieving relevant video clips and avoiding repeated manual screening of complete inspection videos.

[0154] Table 1. Comparison of the Implementation Effects of Video Data Management for Commercial Park Inspections

[0155]

[0156] As shown in Table 1, when the number of inspection videos and the average duration of each video are basically the same, the method of this invention does not change the original inspection operation. The difference in the number and duration of videos generated by the two management methods within the statistical period is small, indicating that this invention optimizes the video data management process without increasing the inspection burden or affecting the continuity of video acquisition, and has good business adaptability and feasibility.

[0157] Regarding video management effectiveness, traditional management methods store and use each inspection video as a single management unit. However, with the method of this invention, each video can be divided into approximately 2.8 event units on average, clearly defining different states and stages in the inspection process. This event-unit-based structured management improves the efficiency of locating key segments, reducing the average location time of key segments from 11.6 minutes to 3.9 minutes, effectively lowering the cost of manual screening and improving the actual utilization efficiency of video data.

[0158] In terms of risk identification and resource utilization, the proportion of high-risk events identified by the method of this invention is close to that of traditional methods, falling within a reasonable fluctuation range, indicating that risk assessment is more stable and reliable. Through differentiated management strategies, the overall storage growth rate of video data is reduced by approximately 9%. While ensuring the integrity and retrieval of high-risk video data, the storage resource consumption of low-risk video data is reduced, demonstrating the comprehensive effectiveness of this invention in refined management and resource optimization.

[0159] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for managing video data from law enforcement recorders, characterized in that, include: By collecting continuous video and audio data through law enforcement recorders, law enforcement session identifiers are generated. These identifiers are then associated with the video and audio data to form a law enforcement session dataset. On the edge, based on the law enforcement session dataset, video data, audio data and law enforcement context information are jointly analyzed to construct a law enforcement behavior pattern domain that reflects the state of law enforcement behavior, and the stability of the law enforcement behavior pattern domain is continuously monitored. When a state change is detected in the law enforcement behavior pattern domain, the law enforcement behavior pattern collapse point is determined, and the video data is segmented with the law enforcement behavior pattern collapse point as the boundary to generate law enforcement event units arranged in chronological order. A four-stream causal inference integrator is built on the cloud side. For each law enforcement event unit, the video behavior stream, audio speech stream, law enforcement context stream, and semantic inference stream corresponding to the law enforcement event unit are integrated to generate a causally consistent law enforcement semantic description result. Based on the semantic description results of law enforcement corresponding to each law enforcement event unit, a dispute risk topography map is constructed to determine the differentiated management strategy for video data corresponding to each law enforcement event unit. Based on the time information, geographical location information, semantic features of law enforcement behavior, and intensity of dispute risk of law enforcement event units, index entries are established, and a video fast retrieval structure based on spatiotemporal inverted index is constructed to locate target law enforcement event units and output video data according to the corresponding differentiated management strategies.

2. The law enforcement recorder video data management method according to claim 1, characterized in that, The video and audio data include video frame sequence data continuously collected by the law enforcement recorder during the law enforcement process and audio signal data collected synchronously. The video frame sequence data includes the image, the timestamp, and the corresponding frame number information. The audio signal data includes audio sampling data, audio timestamp, and volume intensity information.

3. The law enforcement recorder video data management method according to claim 1, characterized in that, The generation of law enforcement session identifiers refers to generating a unique identifier code based on the law enforcement start time, law enforcement recorder device number, law enforcement officer identification, and law enforcement task number when the law enforcement recorder begins recording law enforcement video. This unique identifier code serves as a unified session identifier for the video and audio data of the law enforcement process.

4. The law enforcement recorder video data management method according to claim 1, characterized in that, The construction of a law enforcement behavior pattern domain reflecting the state of law enforcement behavior, and the continuous monitoring of the stability of the law enforcement behavior pattern domain, includes: On the device side, based on the law enforcement conversation dataset, video frame sequences, audio signal sequences, time information, geographic location information, law enforcement personnel identity information and law enforcement task information are aligned with a unified time reference to generate aligned multi-source conversation data; On the aligned multi-source conversation data, video behavior features, audio intonation features, and law enforcement context state features are extracted by time step, standardized, and a time index is established. A law enforcement behavior pattern domain reflecting the state of law enforcement behavior is constructed. The law enforcement behavior pattern domain consists of a video behavior subdomain, an audio speech subdomain, a law enforcement context state subdomain, a device posture and scene structure subdomain, and a cross-channel consistency subdomain. Each subdomain is combined into a pattern domain entry corresponding to a time step according to a unified coding scheme and written into the session data structure. Continuous stability monitoring of law enforcement behavior pattern domains is conducted to establish mutation indicator sequences, sliding consistency sequences, phase persistence counting sequences, and context rule conflict sequences. Within the synchronization window, the monitoring results are comprehensively judged based on the multi-source triggering conditions. When the mutation indicator sequence, the sliding consistency sequence, and the context rule conflict sequence meet the preset triggering combination rules, the behavior pattern collapse pre-mark is output, and the event unit boundary candidate list is generated.

5. The law enforcement recorder video data management method according to claim 1, characterized in that, The generation of law enforcement event units arranged in chronological order includes: Obtain the event unit boundary candidate list and the corresponding behavior pattern collapse pre-label, and obtain the mutation indicator sequence, lens motion anomaly indicator sequence, cross-channel consistency deviation sequence, stage persistence count sequence and context rule conflict sequence corresponding to the event unit boundary candidate list; The law enforcement session dataset is discretized into time steps arranged in chronological order according to a unified time base, the event unit boundary candidate list is mapped to the boundary candidate time step set, and the index relationship of the boundary candidate time steps is established. For each boundary candidate time step, read the values ​​of the mutation indicator sequence, the lens motion anomaly indicator sequence, the cross-channel consistency deviation sequence, the stage persistence count sequence, and the context rule conflict sequence corresponding to the boundary candidate time step to form a boundary candidate judgment element set, and generate the boundary candidate confidence level according to the preset multi-element comprehensive judgment rule. When the boundary candidate time step meets the necessary triggering conditions and the confidence level of the boundary candidate reaches the lower limit of the preset confidence threshold range, and the camera motion anomaly indication sequence does not meet the individual triggering conditions, the boundary candidate time step is determined as the law enforcement behavior pattern collapse point, forming a law enforcement behavior pattern collapse point sequence arranged in chronological order. Using the collapse point sequence of law enforcement behavior patterns as the segmentation boundary, the video and audio data associated with law enforcement session identifiers are segmented to generate law enforcement event units arranged in chronological order.

6. The law enforcement recorder video data management method according to claim 1, characterized in that, The generated causally consistent law enforcement semantic description results include: On the cloud side, a four-stream causal inference integrator is constructed based on law enforcement event units arranged in chronological order. The four-stream causal inference integrator consists of a four-stream alignment unit, a causal adjudication unit, and a semantic generation unit, and establishes a local time index for each law enforcement event unit. The four-stream alignment unit receives video behavior stream, audio speech stream, law enforcement context stream, and semantic inference stream respectively, performs unified time reference alignment and sequence correction on the four streams, performs missing segment completion, duplicate segment removal, abnormal timestamp correction and cross-channel delay correction, generates alignment credibility identifier, alignment version identifier and alignment time reference identifier based on device posture and camera motion estimation data, and outputs four-stream alignment entries. The causal adjudication unit receives four-flow alignment entries, parses the allowed and prohibited relationships of the enforcement stage, generates a causal constraint set, executes a counterfactual shielding mechanism to shield target stage behavior entries that are not reachable from the current stage, executes context conservation verification to eliminate behavior entries that are inconsistent with the enforcement context flow, and outputs a causal consistency candidate behavior entry set. The semantic generation unit receives four-stream aligned entries and a set of causal consistent candidate behavior entries, performs cross-channel divergence resolution, stage boundary inheritance and convergence, merges the two and generates causal consistent law enforcement semantic description entries according to time steps, and merges them into a semantic tag set and a future behavior window according to local time index within the law enforcement event unit. The semantic tag set, future behavior window, and corresponding law enforcement event unit's time range, geographical location information, and law enforcement session identifier are associated and stored as a causally consistent law enforcement semantic description result.

7. The law enforcement recorder video data management method according to claim 1, characterized in that, The construction of the dispute risk topography map and the determination of differentiated management strategies for video data corresponding to each law enforcement incident unit include: Obtain semantic tag sets and future behavior windows, and combine them with the time range and sequence of law enforcement event units, as well as geographical location information and law enforcement session identifiers, to form time dimensions, semantic behavior dimensions and basic association information for evaluation; Based on mutation indicator sequences, camera motion anomaly indicator sequences, cross-channel consistency deviation sequences, phase persistence counting sequences, and context rule conflict sequences, combined with semantic tag sets and future behavior windows, risk feature sets are generated for each law enforcement event unit and standardized and indexed over time. Construct a topographic map of dispute risks. Establish a fixed-step time grid on the law enforcement session timeline and a semantic grid on the stage attributes and key behaviors of the semantic tag set. Combine the two types of grids to obtain basic grid units. Aggregate risk feature sets within the corresponding units to generate risk intensity values. Perform continuity and connectivity checks to remove isolated units. Perform regional merging to obtain four types of regions: peaks, slopes, valleys, and troughs, forming a set of topographic map unit entries. Based on the unit entry set of the dispute risk topographic map, a differentiated management strategy is generated. When the unit entry is of the peak type, multiple copies are redundant, strong encryption and extended retention period are set. When the unit entry is of the slope type, the protection window is expanded and the redundancy level is increased. When the unit entry is of the valley type, the summary storage and compression strategy is set. When the unit entry is of the trough type, the segmented merging storage and retention period are increased. The differentiated management strategy is bound to the video data of the corresponding law enforcement event unit, and the effective parameters of redundancy level, encryption strength and retention period are recorded to generate management strategy index entries.

8. The law enforcement recorder video data management method according to claim 1, characterized in that, The construction of a video fast retrieval structure based on a spatiotemporal inverted index, locating target law enforcement event units and outputting video data according to corresponding differentiated management strategies, includes: Read the management strategy index entries, extract the time range, geographical location information, stage attribute tags or key behavior tags in the semantic tag set, area type and risk intensity value of the dispute risk topographic map, and differentiated management strategy parameters corresponding to each law enforcement event unit, and generate the basic index data; The time range is divided into time keys, the geographic location information is rasterized into spatial keys, the stage attribute labels or key behavior labels are encoded into semantic keys, the area type and risk intensity value are classified into risk keys, and the composite keys are combined in a preset order to form composite keys. The composite keys are then mapped to law enforcement event unit identifiers. A fast video retrieval structure based on spatiotemporal inverted index is constructed using composite keys as the organizational unit. An inverted list is maintained for each composite key. The inverted list is stored in memory and on disk in a hierarchical manner. The inverted list with the region type of peak and the risk intensity value at the highest level resides in memory. The inverted list with high access frequency is cached with dual copies. The index version number and effective status are recorded for each inverted list. Upon receiving a video retrieval request, the request conditions are parsed to generate a query key combination. The results are matched step by step according to the time key, spatial key, semantic key and risk key. The results in the inverted list are deduplicated and filtered by intersection to locate the target law enforcement event unit. Results belonging to the area type "peak" are returned first. Based on the location results, a retrieval list is generated. Encryption control, redundant copy reading, and retention period verification are performed according to the corresponding differentiated management strategy parameters. Video data is output and retrieval logs and index version numbers are recorded.

9. A law enforcement recorder video data management system, comprising the law enforcement recorder video data management method according to any one of claims 1 to 8, characterized in that, Includes the following modules: The law enforcement session acquisition module is used to collect video and audio data, generate law enforcement session identifiers, and form a law enforcement session dataset. The law enforcement behavior pattern construction module is used to construct a law enforcement behavior pattern domain based on the law enforcement session dataset and continuously monitor the stability of the law enforcement behavior pattern domain. The law enforcement event unit segmentation module is used to determine the collapse point of the law enforcement behavior pattern when a state change occurs in the law enforcement behavior pattern domain, and to generate law enforcement event units. The four-stream causal inference integration module is used to integrate the video behavior stream, audio tone stream, law enforcement context stream, and semantic inference stream corresponding to the law enforcement event unit to generate a causally consistent law enforcement semantic description result. The Dispute Risk Terrain and Strategy Generation Module is used to construct dispute risk terrain maps based on law enforcement semantic description results and determine differentiated management strategies. The spatiotemporal inverted index retrieval module is used to construct a fast video retrieval structure based on the spatiotemporal inverted index and output video data according to a differentiated management strategy.