Intelligent abnormal situation early warning system and method for traffic police law enforcement audio and video
Patent Information
- Application Number
- CN202610669279.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-18
AI Technical Summary
然而,现有技术中对交警执法音视频的管理与分析,通常仍以人工回看、人工抽检或基于单一数据源的识别方式为主,难以对现场视频、现场音频、设备事件记录和执法处置记录进行统一时间对齐、关联分析和综合判断
本发明通过将视频、音频、设备日志和执法处置记录进行统一校时、融合分析和语境建模,能够从多维度识别执法现场异常演化过程,不再局限于单一画面或单一语音特征判断。其不仅能够提高异常识别的准确性和实时性,还能通过渐进异常累积、滞后响应和联动收敛等机制,更早发现由轻微对抗逐步升级的风险事件。与此同时,本发明还能输出对应时间区间的证据片段、关键标注和证据摘要,增强预警结果的可解释性与可复核性,并通过误触发剔除、重复预警合并和可信度调节,降低误报率,提升系统在复杂执法环境中的适用性与实战价值。
Smart Images

Figure CN122598062A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent video analysis and law enforcement supervision and early warning technology, and more specifically, to an intelligent early warning system and method for abnormal situations in traffic police law enforcement audio and video. Background Technology
[0002] With the widespread application of body cameras, vehicle-mounted data collection devices, and business terminals in traffic police enforcement activities, enforcement scenes typically generate various types of data, including video, audio, equipment operation logs, and enforcement handling records, providing a foundation for record-keeping, process supervision, and post-event review. However, current technologies for managing and analyzing traffic police enforcement audio and video typically rely on manual review, random checks, or identification based on a single data source. This makes it difficult to perform unified time alignment, correlation analysis, and comprehensive judgment of on-site video, audio, equipment event records, and enforcement handling records. For abnormal situations that occur during enforcement, such as escalating disputes, gatherings of people, abnormal vehicle movement, camera obstruction, equipment malfunctions, and inconsistencies between recorded content and the actual situation, current technologies often suffer from problems such as untimely detection, insufficient identification accuracy, numerous false alarms and omissions, and difficulty in effectively reconstructing abnormal processes. These shortcomings make it difficult to meet the application needs of traffic police enforcement scenarios for real-time early warning of abnormal situations and linked evidence output.
[0003] To address the above problems, this invention proposes a solution. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent early warning system and method for abnormal situations of traffic police law enforcement audio and video, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: In a preferred embodiment, it includes: Retrieve on-site video streams, on-site audio streams, environmental video streams or supplementary audio streams, equipment operation logs and law enforcement handling records corresponding to the same law enforcement process, establish a unified timeline based on time correction anchor points, and generate a standardized time sequence fragment sequence; Feature extraction is performed on on-site video data, on-site audio data, equipment event data, and law enforcement handling record data in standardized time-series segments, and time alignment and fusion are performed according to a unified time grid to generate law enforcement context representation records; The law enforcement context representation record is read sequentially along the time sequence of the time segment. Anomaly triggering conditions are established based on the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state, and candidate events for abnormal situations are constructed. For the candidate events for abnormal situations, the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state between adjacent time units are compared. Voice adversarial change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change are extracted. Change driving quantity is constructed based on each change quantity. Combined with the gradual anomaly accumulation, lag response, linkage convergence, on-site deviation, and credibility adjustment results, the anomaly evolution intensity is constructed, and then the anomaly warning event is determined. Establish a relationship for generating warning levels for abnormal situation warning events, construct evidence fragments, key content marking sequences, and evidence summaries around the time interval corresponding to the abnormal situation warning events, correct the warning result records, and generate the final warning result records.
[0006] In a preferred embodiment, the on-site video stream, on-site audio stream, environmental video stream or supplementary audio stream, equipment operation log, and law enforcement handling record corresponding to the same law enforcement process are collected to construct the original data set corresponding to a single law enforcement process; the original timestamps from each data source and event records that can appear in more than two data sources are extracted as time correction anchors to establish a unified timeline and the timestamps of each data source are corrected; the on-site video stream, on-site audio stream, equipment operation log, and law enforcement handling record are subjected to integrity verification, and low-confidence markers are added to the corresponding time intervals on the unified timeline.
[0007] In a preferred embodiment, a sliding time window is set along a unified timeline, and the original data set is segmented into time-series segments and boundary corrections are performed by combining the start recording event, stop recording event, record generation time, and voice activity change time points to generate a time-series segment sequence. According to the start and end times of each time-series segment on the unified timeline, the on-site video subsequence, on-site audio subsequence, environmental video subsequence, equipment operation log subsequence, and law enforcement handling record subsequence are multi-source aligned. Furthermore, sections with completely black screens, lens-obstructed sections, and sections with long-term silence and no event changes in the equipment operation log are identified, and invalid section markers are added at the corresponding time positions.
[0008] In a preferred embodiment, video frame extraction and image content recognition are performed on the on-site video data in standardized time-series segments to construct video feature records containing personnel status, body movement status, and scene status; audio segmentation, speech content recognition, speaker separation, semantic tagging, and sound event recognition are performed on the on-site audio data to construct audio feature records containing speech text content, speaker status, semantic type, and sound event status; event sequence parsing is performed on the equipment operation logs to construct equipment event feature records corresponding to equipment event status; text segmentation and semantic classification are performed on the law enforcement handling records to construct handling record feature records corresponding to law enforcement action status; and the video feature records, audio feature records, equipment event feature records, and handling record feature records are grouped into continuous time units in a unified time grid, and jointly organized to generate law enforcement context representation records containing speech text content, speaker status, personnel status, scene status, sound event status, equipment event status, and law enforcement action status, and the law enforcement context representation records in adjacent time units are concatenated in chronological order to form a time-series segment-level law enforcement context representation sequence.
[0009] In a preferred embodiment, the law enforcement context representation record is read sequentially along the time sequence of the time segment. The voice text content, speaker status, personnel status, scene status, sound event status, device event status, and law enforcement action status are checked item by item to establish abnormal triggering conditions for visual, voice, sound events, devices, and law enforcement actions. Specifically, based on the correspondence between changes in the number of personnel, changes in relative distance, tense scene conditions, adversarial expressions, high-frequency alternating voices from multiple speakers, continuous shouting, changes in camera status, and the law enforcement action status and the scene status, abnormal triggering time units related to continuous changes are identified. Using the time unit that satisfies the abnormal triggering condition as the starting time unit, the process expands forward and backward according to the continuous occurrence of similar or related abnormal triggering conditions, constructing candidate events for abnormal situations with start time, end time, and abnormal triggering type.
[0010] In a preferred embodiment, candidate events of abnormal situations are sequentially analyzed across modalities and states within continuous time units. The changes in speech text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state between adjacent time units are compared. This constructs speech adversarial change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change. Instead of directly judging a single explicit abnormal phenomenon such as pushing, falling, collision, or device interruption, multiple weak changes that gradually appear in the normal law enforcement phase, such as increased adversarialness, frequent speaker switching, personnel gathering and closing, close-range confrontation in the scene, and intermittent fluctuations in device state, are incorporated into the same continuous analysis framework. Furthermore, based on these multiple changes, change-driving quantities are constructed, and combined with the progressive abnormality accumulation quantity of the previous time unit and the law enforcement action maintenance quantity of the current time unit, a recursive calculation relationship for the progressive abnormality accumulation quantity is established to characterize the continuous superposition process of multiple weak abnormal changes within continuous time units.
[0011] In a preferred embodiment, during the calculation of the progressive anomaly accumulation, the composition of the change-driving quantity is adaptively adjusted based on the change in visibility perturbation. When video visibility is stable, spatial compression change and scene tightening change are emphasized. When video visibility decreases, the participation of speech adversarial change, speaker alternation change, and sound perturbation change is increased, thereby establishing a multi-factor dynamic analysis structure for visibility fluctuation scenarios. On this basis, the time units of the first significant increase of various change quantities and their sequential adjacency relationships are sequentially compared to construct a hysteresis response quantity to characterize the forward response chain formed sequentially between speech adversarial change, spatial compression change, and visibility perturbation change. The process of different types of changes further transitioning from sequential responses to clustering in the same time interval is characterized as a linkage convergence quantity.
[0012] In a preferred embodiment, the state of law enforcement actions is compared with the changes in voice confrontation, spatial compression, scene tightening, sound disturbance, and visibility disturbance within the same time unit to construct a field deviation. This deviation represents the state deviation relationship where, while the law enforcement action appears to be still in the stage of verifying documents, asking questions, informing about matters, or collecting evidence on-site, the internal changes on-site have been continuously tightening. Then, combined with low-confidence markers, invalid segment markers, and device event states, the confidence of continuous changes, gradual anomaly accumulation, delayed response, field deviation, and linkage convergence is adjusted to form corresponding confidence correction results. Subsequently, according to the judgment hierarchy of gradual anomaly accumulation, sequential response, local convergence, and surface stable deviation, a layer-by-layer opening relationship of anomaly evolution intensity is constructed. Based on the continuous distribution of anomaly evolution intensity within the candidate event time range, event-level aggregation is performed to distinguish between anomaly evolution intervals and general law enforcement fluctuations. Finally, combined with duration thresholds and peak intensity thresholds, anomaly warning events are determined.
[0013] In a preferred embodiment, the abnormal trigger type and abnormal scoring result of the abnormal situation warning event are read, a warning level generation relationship is established, and the start time, end time, abnormal trigger type, abnormal scoring result, and warning level are written into the warning result record. An evidence extraction time interval is established around the start and end times of the abnormal situation warning event. On-site video data, on-site audio data, supplementary environmental data, equipment event data, and law enforcement handling record data are located and extracted from a standardized time-series segment sequence to construct evidence segments. Personnel status, body movement status, scene status, voice and text content, semantic type, sound event status, equipment event status, and law enforcement action status are marked to the corresponding time positions in the evidence segments to construct a key content marking sequence. The key content marking sequence is read in chronological order to organize and generate an evidence summary corresponding to the abnormal situation warning event. Multiple warning result records that are temporally adjacent are subjected to overlap checks, continuation checks, and key content continuity checks. Merging rules for the same abnormal process are established, and warning result records are merged, false triggers are eliminated, and abnormal scores are corrected to generate a final warning result record containing the abnormal trigger type, abnormal scoring result, warning level, evidence segment index, evidence summary, and correction status.
[0014] In a preferred embodiment, it includes: a multi-source law enforcement data standardization module, a law enforcement context representation module, an anomaly evolution early warning module, an early warning result processing module, and signal connections between the modules; The multi-source law enforcement data standardization module is used to retrieve on-site video streams, on-site audio streams, environmental video streams or supplementary audio streams, equipment operation logs and law enforcement handling records corresponding to the same law enforcement process, establish a unified timeline based on time correction anchor points, and generate a standardized time sequence fragment sequence. The law enforcement context representation module is used to extract features from on-site video data, on-site audio data, equipment event data, and law enforcement handling record data in a standardized time-series segment sequence, and to perform time alignment and fusion according to a unified time grid to generate law enforcement context representation records. The abnormal evolution early warning module is used to read the law enforcement context representation record in chronological order along the time segment. It establishes abnormal triggering conditions based on the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state, and constructs candidate events for abnormal situations. For the candidate events for abnormal situations, it compares the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state between adjacent time units, and extracts the voice confrontation change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change. Based on each change, it constructs change driving quantity, and combines it with the gradual abnormal accumulation quantity, lag response quantity, linkage convergence quantity, on-site deviation quantity, and credibility adjustment result to construct the abnormal evolution intensity, and then determines the abnormal situation early warning event. The early warning result processing module is used to establish an early warning level generation relationship for abnormal situation early warning events, construct evidence fragments, key content marking sequences and evidence summaries around the time interval corresponding to the abnormal situation early warning events, correct the early warning result records, and generate the final early warning result records.
[0015] The technical effects and advantages of the intelligent early warning system and method for abnormal situations in traffic police law enforcement audio and video of the present invention are as follows: This invention unifies the time synchronization, integrates analysis, and performs contextual modeling of video, audio, device logs, and law enforcement records. This enables multi-dimensional identification of abnormal evolution processes at law enforcement scenes, moving beyond judgments based on single visual or auditory features. It not only improves the accuracy and real-time performance of anomaly identification but also detects escalating risks earlier through mechanisms such as gradual anomaly accumulation, delayed response, and coordinated convergence. Simultaneously, this invention outputs evidence fragments, key annotations, and evidence summaries for corresponding time intervals, enhancing the interpretability and verifiability of warning results. Furthermore, it reduces false alarm rates through false trigger removal, duplicate warning merging, and credibility adjustment, thereby improving the system's applicability and practical value in complex law enforcement environments. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the overall structure of the intelligent early warning system and method for abnormal situations in traffic police law enforcement audio and video according to the present invention.
[0017] Figure 2 This is a schematic diagram illustrating the multi-source feature extraction and law enforcement context representation of the intelligent early warning system and method for abnormal situations in traffic police law enforcement audio and video of the present invention.
[0018] Figure 3 This is a graph showing the abnormal evolution intensity and warning threshold of the intelligent early warning system and method for abnormal situations in traffic police law enforcement audio and video.
[0019] Figure 4 This is a time series curve of the lag response and linkage convergence of the intelligent early warning system and method for abnormal situations in traffic police law enforcement audio and video of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] In this embodiment, the present invention discloses an intelligent early warning method for abnormal situations in traffic police enforcement audio and video, such as... Figure 1 As shown, it includes: Step 1: Retrieve the on-site video stream, on-site audio stream, environmental video stream or supplementary audio stream, equipment operation logs and law enforcement handling records corresponding to the same law enforcement process, establish a unified timeline based on time correction anchor points, and generate a standardized time sequence fragment sequence; Step 2: Extract features from the on-site video data, on-site audio data, equipment event data, and law enforcement handling record data in the standardized time-series segment sequence, and perform time alignment and fusion according to a unified time grid to generate law enforcement context representation records; Step 3: Read the law enforcement context representation record in chronological order along the time sequence of the time segment. Establish anomaly triggering conditions based on the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state, and construct candidate events for anomalies. For the candidate events for anomalies, compare the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state between adjacent time units. Extract voice adversarial change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change. Construct change driving quantities based on each change quantity, and combine them with progressive anomaly accumulation, hysteresis response, linkage convergence, on-site deviation, and credibility adjustment results to construct the anomaly evolution intensity, and then determine the anomaly warning event. Step 4: Establish a warning level generation relationship for abnormal situation warning events, construct evidence fragments, key content marking sequences and evidence summaries around the time interval corresponding to the abnormal situation warning events, correct the warning result records, and generate the final warning result records.
[0022] In step one, the raw audio and video data corresponding to the same law enforcement process is retrieved first. This raw audio and video data includes the on-site video and audio streams generated by the law enforcement recorder, the environmental video stream or supplementary audio stream generated by the vehicle-mounted device, the device operation log generated by the law enforcement equipment, and the law enforcement handling record generated by the business terminal. The on-site video stream contains continuous video frames, frame numbers, and frame timestamps; the on-site audio stream contains continuous audio sampling blocks and sampling timestamps; the environmental video stream or supplementary audio stream is used to supplement on-site content not covered by a single acquisition device; the device operation log records events such as device startup, shutdown, location refresh, storage anomalies, and network status changes, along with their occurrence times; and the law enforcement handling record records the operational items during the law enforcement process and their recording times. The above data is then aggregated according to the law enforcement task identifier, device identifier, and acquisition time range to form the raw data set corresponding to a single law enforcement process.
[0023] After obtaining the original dataset, time base calibration is performed on each data source. Since law enforcement recorders, vehicle-mounted devices, and business terminals operate independently, timestamps for the same moment may differ across different data sources. Therefore, original timestamps are first extracted from the on-site video stream, on-site audio stream, equipment operation logs, and law enforcement handling records. Then, event records with a clear occurrence time are extracted from the equipment operation logs and law enforcement handling records as time calibration anchors. These time calibration anchors are events that can occur in more than two data sources during the same law enforcement process. These events include recording start events, recording stop events, equipment reconnection events, location refresh events, and handling record generation events. For any original timestamp in any data source... The time offset of the data source is calculated based on the difference between the anchor time corresponding to the data source and the corresponding anchor time on the unified time axis. And correct the original timestamp to: ; in, The original timestamp of the third original data record. This is the corrected timestamp. This represents the time offset of data source s. When multiple anchor events exist within the same data source, the differences are statistically processed, and the median is taken as the time offset of the data source to eliminate deviations caused by individual abnormal records.
[0024] After time base calibration, the integrity of the original dataset is verified. Frame sequence numbers, timestamps, and encoding information are read frame by frame from the on-site video stream to determine if there are frame sequence number reversals, consecutive frame breaks, or encoding corruption. Segment by segment, sampling block timestamps and sampled data are read from the on-site audio stream to determine if there are sampling block interruptions, timestamp jumps, or abnormal silence. Event types and times are read from each item in the equipment operation log to determine if there are missing key events, abnormal time jumps, or prolonged periods without event recordings. The creation time, modification time, and content of law enforcement records are read to determine if there are any conflicts where the recording time is earlier than the start time or later than the end time of law enforcement. For data that passes verification, the original time position is retained. For data that fails verification, it is not directly deleted but instead a low-confidence marker is added to the corresponding time interval. The low-confidence marker is used to identify data gaps, equipment malfunctions, or discontinuous recordings within that time interval, and the marker's position directly corresponds to the start and end times on a unified timeline.
[0025] After integrity verification, the on-site video and audio streams are segmented into time-series segments. During segmentation, a sliding time window is set along a unified time axis to obtain initial candidate time-series segments. The length L of the sliding time window is the time range covered by a single candidate time-series segment, and the step size s is the time interval between the start times of adjacent candidate time-series segments; both are preset during processing. After the initial candidate time-series segments are formed, their boundaries are corrected. The data used for boundary correction includes start recording events and stop recording events from the device operation log, record generation times from law enforcement handling records, and time points of voice activity changes in the on-site audio stream. The timing of changes in speech activity is obtained through speech activity detection. Specifically, the following method extracts short-time energy, zero-crossing rate, and spectral entropy from the live audio stream using continuous short-time windows. Short-time energy characterizes the intensity of sound within the current time window, zero-crossing rate characterizes the degree of frequency change, and spectral entropy characterizes the dispersion of the sound spectrum. When these parameters in the continuous short-time window change from a low-change state to a high-change state, it is determined that speech activity has transitioned from silence; conversely, when they change from a high-change state to a low-change state, it is determined that speech activity has transitioned from silence. When both short-time energy and spectral entropy increase synchronously within a continuous time period, it is determined that there is a strong change in speech activity during that period. Using these timing points as candidate boundaries, the start and end positions of the initial candidate time segments are adjusted to obtain a corrected time segment sequence.
[0026] After obtaining the time-series segment sequence, multi-source alignment is performed on each time-series segment. For the k-th time-series segment, based on the start and end times of the segment on a unified time axis, the corresponding time ranges are extracted into on-site video subsequences, on-site audio subsequences, environmental video subsequences, equipment operation log subsequences, and law enforcement handling record subsequences. Since on-site video is sampled frame by frame, on-site audio is sampled block by block, equipment operation logs are recorded by event, and law enforcement handling records are recorded by text, it is necessary to perform unified mapping on different data according to the same time range of the segment. For the on-site video subsequence, the corresponding video frame is read according to the time position within the segment; for the on-site audio subsequence, audio sampling blocks are aggregated according to the time window corresponding to the video frame; for the equipment operation log subsequence, events are marked by falling into the corresponding time window according to their occurrence time; for the law enforcement handling record subsequence, each record is assigned to the corresponding time window according to its recording time.
[0027] After multi-source alignment, the time-series segments undergo purification processing. For the on-site video subsequence, image jitter suppression is performed. Specifically, corner points or local feature points are extracted from adjacent video frames, global motion vectors between adjacent frames are calculated, and the video frame positions are compensated based on these global motion vectors to reduce image shift caused by device shaking. For low-light video frames, brightness enhancement is performed. Specifically, the brightness components of the video frames are extracted, dynamically stretched, and contrast enhancement is performed on local areas to make previously difficult-to-identify dark targets visible. For the on-site audio subsequence, noise reduction is performed. Specifically, stable background noise frequency bands are first identified in the audio spectrum, and then background noise components are suppressed using spectral subtraction or adaptive filtering, while retaining speech and sudden sound components. Invalid segments in the segments are identified. Invalid segments include completely black segments, segments where the camera is obstructed, and segments that are silent for extended periods and show no event changes in the device operation log. Identified invalid segments are not deleted; only their original time positions are retained and marked as invalid segments.
[0028] In step two, the on-site video data in the standardized time-series segments is first subjected to video frame extraction processing. The video frame extraction processing is performed according to the duration of the time-series segments on the same time axis. For each predetermined sampling time within the duration, the corresponding video frame is read; when multiple video frames exist near the sampling time, the video frame with the timestamp closest to that sampling time is selected as the current sampling frame. The predetermined sampling time is generated by a fixed time interval, which is a pre-set frame sampling interval used to control the video frame extraction density.
[0029] The extracted video frame sequences are processed for image content recognition. This processing includes personnel target recognition, vehicle target recognition, vehicle recognition, body movement recognition, and scene state recognition. Personnel target recognition is achieved by performing target detection on the video frames. Target detection takes the video frame image as input and outputs the location region and target category of each target in the image; target categories include at least law enforcement personnel, parties involved in the case, and bystanders. Vehicle target recognition is achieved by detecting the vehicle outline, body area, and license plate area in the video frames. The detection results include the vehicle's location region and vehicle category in the video frame. Vehicle recognition is used to identify objects related to law enforcement activities, such as warning cones, police vehicles, law enforcement documents, and hand gesture aids; the recognition method also employs video target detection. Body movement recognition is achieved by extracting key points of the human body from consecutive video frames and calculating the displacement changes of these key points. Key points refer to the positions of the head, shoulders, elbows, wrists, hips, knees, and ankles in the image; these positions are extracted using human pose estimation methods. The displacement amplitude, direction, and duration of the key points over consecutive time constitute the criteria for determining body movement. Scene state recognition is used to determine whether the scene corresponding to the current video frame is in a normal parking state, a close-range confrontation state, a crowd gathering state, a fast-moving state, or a state where the camera is obstructed. The scene state is determined by comprehensively considering the results of human target recognition, vehicle target recognition, vehicle recognition, and body movement recognition at the same time.
[0030] During the image content recognition process, a video feature record is generated for each sampling moment. This video feature record includes the video frame timestamp at that sampling moment, the category and location of personnel, vehicles, and transportation equipment, the location of human key points, the category of body movements, and the category of scene state. All of the above information is directly derived from the video frame recognition results at the corresponding sampling moment. Specifically, the categories of personnel, vehicles, and transportation equipment are given by the target detection results; the location of human key points is given by the human pose estimation results; the category of body movements is calculated from the changes in human key points across consecutive video frames; and the category of scene state is determined by the joint determination of all video recognition results at the current moment.
[0031] Audio segmentation is performed on the live audio data in the standardized time-series segments. Audio segmentation is conducted using a continuous short-time-window approach; that is, the current audio segment is extracted along the time axis of the time-series segment by a preset audio analysis window length, and then moved forward according to a preset sliding interval to form a sequentially arranged audio segment sequence. The audio analysis window length represents the duration covered by a single audio segment, and the sliding interval represents the time interval between the start times of two adjacent audio segments. For each audio segment, basic acoustic features are extracted, including short-time energy, zero-crossing rate, fundamental frequency, formants, spectral centroid, and Mel-frequency cepstral coefficients. Short-time energy is obtained by calculating the sum of squares of the amplitudes of the sampling points within the current audio segment, and is used to characterize the intensity of the sound within the current audio segment; zero-crossing rate is obtained by counting the number of times the waveform crosses zero within the current audio segment, and is used to characterize the frequency of sound changes; fundamental frequency is obtained by performing periodic detection on the audio segment, and is used to characterize the pitch of the main sound; formants are obtained by analyzing the spectral envelope of the audio segment, and are used to characterize the position of the spectral peaks formed by the vocal organs; spectral centroid is obtained by calculating the center of the spectral energy distribution, and is used to characterize the brightness of the sound; Mel frequency cepstral coefficients are obtained by performing Mel filtering and cepstral transform on the audio segment spectrum, and are used to characterize the speech spectral structure.
[0032] After extracting basic acoustic features, speech content recognition and sound event recognition processing are performed on the on-site audio data. Speech content recognition processing converts the speech components in the audio into text content. Specifically, speech activity detection is first performed on the audio segments to filter out audio segments containing continuous speech components; then, speech recognition is performed on the filtered audio segments to convert them into corresponding text. During speech activity detection, the start and end positions of speech segments are identified using short-time energy, zero-crossing rate, and spectral distribution changes; during speech recognition, the correspondence between the acoustic features in the audio segments and speech units is used to output the recognized text. For the recognized text, speaker separation processing and semantic tagging processing are further performed. Speaker separation processing divides continuous speech segments into speech sequences belonging to different speakers by comparing the timbre features, fundamental frequency features, and formant distributions between different speech segments; semantic tagging processing marks the imperative statements, interrogative statements, persuasive statements, warning statements, argumentative statements, and requests for help statements in the recognized text, based on keywords, sentence structure, and semantic strength in the text. Sound event recognition processing is used to identify non-speech sound events, including continuous horn sounds, sudden braking sounds, collision sounds, falling sounds, violent knocking sounds, continuous shouting sounds, and prolonged noise interference sounds. Sound event recognition is obtained by jointly analyzing the time-domain and frequency-domain features of an audio segment. The analysis result is the sound event category corresponding to the current audio segment and its occurrence time.
[0033] During the on-site audio data processing, an audio feature record is generated for each audio analysis moment. This audio feature record includes the start and end times of the audio segment, basic acoustic features, speech recognition text, speaker markers, semantic markers, and sound event categories. Specifically, the start and end times of the audio segment are derived from the audio segmentation process; the basic acoustic features are derived from the acoustic calculation results of the current audio segment; the speech recognition text is derived from the speech content recognition results; the speaker markers are derived from the speaker separation results; the semantic markers are derived from the semantic analysis results of the recognized text; and the sound event categories are derived from the sound event recognition results.
[0034] Event sequence parsing is performed on device event data in standardized time-series segments. The device event data originates from the device operation log, and each record in the log contains an event type and an event occurrence time. During parsing, event records in the device operation log are read chronologically, and event types are converted into unified event markers. These unified event markers include recording start markers, recording stop markers, storage error markers, network interruption markers, network recovery markers, positioning refresh markers, device restart markers, and lens status change markers. For each device event marker, its corresponding occurrence time is retained and projected onto the corresponding time position within its time-series segment, forming a device event feature record. This device event feature record includes the event occurrence time, event type, and event duration. When the log contains both the event start time and the event end time, the event duration is determined by the time length between the start and end times. When the log contains only a single event moment, the event duration is recorded as an instantaneous event.
[0035] The data from standardized time-series law enforcement incident records is parsed and processed. This data originates from incident record text generated in the business terminal. For each incident record, the record generation time, modification time, and text content are read. Text segmentation and semantic classification are performed on the text content to extract the content representing law enforcement actions. These actions include at least stopping vehicles, checking documents, informing the recipient, questioning, directing vehicle relocation, collecting evidence on-site, issuing verbal warnings, and transferring the case. Text segmentation divides continuous text into word sequences with independent semantics, and semantic classification maps these word sequences to corresponding law enforcement action categories. For each incident record, a incident record feature record is generated, including the record generation time, modification time, law enforcement action category, and key semantic fragments in the text. Key semantic fragments are text segments that directly represent the content of law enforcement actions, such as text representing the object being checked, the action being taken, or the on-site state.
[0036] After completing video feature extraction, audio feature extraction, device event feature extraction, and law enforcement handling record feature extraction, time alignment and fusion processing is performed on various features within the same time segment. The time alignment and fusion processing is based on a unified time grid within the time segment. The unified time grid is a series of continuous time units formed by dividing the start and end times of the time segment according to preset time intervals, with each time unit corresponding to a fixed-length time range within the time segment. For video feature records, they are assigned according to the sampling time falling into the corresponding time unit; for audio feature records, they are assigned according to the overlap between the start and end times of the audio segment and the time unit; for device event feature records, they are assigned according to the event occurrence time falling into the corresponding time unit; and for handling record feature records, they are assigned according to the record generation time or record modification time falling into the corresponding time unit.
[0037] Multiple features within each time unit are jointly organized to generate a law enforcement context representation. This representation is a unified description of the law enforcement scene state, voice communication state, sound event state, equipment operation state, and law enforcement action state within the current time unit. Specifically, the generation process involves: first, reading the video feature set of the current time unit to determine the personnel distribution state, vehicle distribution state, body language state, and scene state; then, reading the audio feature set of the current time unit to determine the speech text content, speaker changes, semantic type, and sound event status; next, reading the equipment event feature set of the current time unit to determine if there are any recording interruptions, network anomalies, location changes, or camera state changes; and finally, reading the handling record feature set of the current time unit to determine the recorded law enforcement actions. These results are then combined within the same time unit to form the law enforcement context representation record for that time unit. This record includes the start and end times of the time unit, personnel status, vehicle status, body language state, scene state, speech text content, speaker status, semantic type, sound event status, equipment event status, and law enforcement action status.
[0038] When there is continuity in the law enforcement context representation records across multiple adjacent time units, the adjacent time units are sequentially concatenated to form a temporal segment-level law enforcement context representation. Continuity refers to the absence of significant changes in the main personnel targets, the consistency of the scene state, the continuous occurrence of voice communication, or the continuous existence of the same law enforcement action within adjacent time units. During the temporal concatenation process, the law enforcement context representation records in adjacent time units are arranged chronologically, and the start and end times of each time unit are retained, thereby forming a continuous law enforcement context representation sequence corresponding to that temporal segment.
[0039] like Figure 2As shown, in step three, when sequentially parsing the law enforcement context representation records, the records are first read sequentially along the time sequence of each time segment. Each law enforcement context representation record includes the start and end times of the time unit, personnel status, vehicle status, body language status, scene status, voice text content, speaker status, semantic type, sound event status, device event status, and law enforcement action status. Each of these items is checked to identify whether predefined abnormal triggering conditions appear. These abnormal triggering conditions refer to judgment conditions that characterize possible abnormal situations during law enforcement. These abnormal triggering conditions are categorized according to the source of the abnormality into visual abnormality triggering conditions, voice abnormality triggering conditions, sound event abnormality triggering conditions, device event abnormality triggering conditions, and law enforcement action abnormality triggering conditions.
[0040] The abnormal scene triggering conditions are determined through personnel status, vehicle status, limb movement status, and scene status. Personnel status characterizes the changes in the number, spatial distribution, and relative distance of law enforcement officers, parties involved, and bystanders within the current time unit; changes in number are obtained by comparing personnel target recognition results within consecutive time units, changes in spatial distribution are obtained by changes in the position of each personnel target in the video frame, and changes in relative distance are obtained by changes in the image distance between the center points of different personnel target position areas. Vehicle status characterizes whether a vehicle remains stationary, suddenly starts, rapidly deviates, or approaches law enforcement officers; vehicle status is determined based on changes in the vehicle target's position, contour displacement, and position relative to personnel targets in consecutive video frames. Limb movement status characterizes whether there are actions such as waving arms, pushing, grabbing, running, falling, rapidly approaching, or rapidly retreating; limb movement status is determined based on the displacement direction, displacement speed, and duration of key human body points in consecutive video frames. Scene state is used to characterize whether the current scene is in a state of close confrontation between people, a state of gathering of people, a state of rapid movement, or a state of camera obstruction. The scene state is jointly determined by the state of people, the state of vehicles, and the state of body movements. When the number of people significantly increases in a short period of time, the relative distance between the parties involved and law enforcement officers decreases sharply in a continuous time unit, the vehicle suddenly shifts in the vicinity of law enforcement officers, or the body movements include pushing, grabbing, running, falling, etc., the corresponding time unit is determined to meet the scene abnormal triggering conditions.
[0041] The speech anomaly triggering conditions are determined based on the speech text content, speaker status, and semantic type. The speech text content is the text obtained through speech recognition in step two; the speaker status is the speaker change result obtained through speaker separation processing in step two; and the semantic type is the sentence category obtained after semantic tagging of the text content in step two. When determining the speech anomaly triggering conditions, the speech text content within the current time unit is first read, and then checked for keywords and sentence patterns related to disputes, non-cooperation, threats, insults, cries for help, or requests for assistance. Simultaneously, the speaker status is checked for situations such as multiple speakers alternating at high speed, a single speaker continuously speaking at high intensity, or a sudden increase in the number of speakers. The semantic type is then checked for consecutive occurrences of argumentative statements, warning statements, requests for assistance, or high-intensity command statements within the current time unit. When the speech text content contains consecutive confrontational expressions, and the speaker status shows multiple speakers alternating at high frequency, or the semantic type continuously shows argumentative statements and requests for assistance, the corresponding time unit is determined to meet the speech anomaly triggering conditions.
[0042] The abnormal triggering conditions for sound events are determined through the sound event status. The sound event status originates from the sound event recognition processing performed on the on-site audio data in step two. The sound event status includes continuous horn blasts, sudden braking sounds, collision sounds, falling sounds, violent knocking sounds, continuous shouting sounds, and prolonged noise interference sounds. When determining the abnormal triggering conditions for sound events, the sound event category and duration within the current time unit are read. If a collision sound, falling sound, violent knocking sound, or continuous shouting sound occurs within the current time unit, the current time unit is directly determined to meet the abnormal triggering conditions for sound events. If a continuous horn blast, sudden braking sound, or prolonged noise interference sound occurs within the current time unit, further determination is made by combining the vehicle status, personnel status, and scene status in adjacent time units. When the above sound events occur simultaneously with sudden vehicle displacement, rapid personnel movement, or personnel gathering, the corresponding time unit is determined to meet the abnormal triggering conditions for sound events.
[0043] The abnormal triggering conditions for device events are determined through device event states. Device event states originate from the device event feature records parsed from the device operation log in step two. These device event states include recording start markers, recording stop markers, storage error markers, network interruption markers, network recovery markers, location refresh markers, device restart markers, and camera status change markers. When determining abnormal triggering conditions for device events, it checks whether a recording stop marker, storage error marker, network interruption marker, device restart marker, or camera status change marker exists within the current time unit. If any of these markers exists, the occurrence time of the marker is read and compared with the video and audio content states within the same time unit. When a recording stop marker, device restart marker, or camera status change marker appears within the time interval of argumentative statements, pushing and shoving actions, collision sounds, or a gathering of people, the corresponding time unit is determined to meet the abnormal triggering conditions for device events. Similarly, when a network interruption marker or storage error marker appears within the time interval of a continuous law enforcement action where the video content suddenly disappears, the corresponding time unit is also determined to meet the abnormal triggering conditions for device events.
[0044] The abnormal triggering conditions for law enforcement actions are determined through the status of the law enforcement actions. The law enforcement action status is derived from the results of parsing the law enforcement handling records in step two. The law enforcement action status includes stopping vehicles, verifying documents, informing others of matters, questioning and explaining, directing vehicle relocation, collecting evidence on-site, issuing verbal warnings, and handing over for further processing. When determining abnormal triggering conditions for law enforcement actions, the law enforcement action status is compared with the personnel status, voice / text content, and equipment event status at the same time and location. If the law enforcement action status indicates that it is in the stage of verifying documents, questioning and explaining, or informing others of matters, and strong argumentative statements, pushing and shoving, rapid approaching actions, or collision sounds occur within the same time unit, then that time unit is determined to meet the abnormal triggering conditions for law enforcement actions. If the law enforcement action status changes continuously and the corresponding time interval lacks video content, audio content, or equipment event records for a long period, then that time interval is recorded as an abnormal time interval where the law enforcement action record and the on-site record are inconsistent, and it is determined to meet the abnormal triggering conditions for law enforcement actions.
[0045] After identifying a time unit that meets any abnormal trigger condition, an abnormal situation candidate event is generated for that time unit. An abnormal situation candidate event refers to an event segment consisting of one or more adjacent time units that initially exhibits abnormal characteristics. When generating an abnormal situation candidate event, the first time unit that meets the abnormal trigger condition is used as the starting time unit, and then the process expands to check the adjacent time units before and after it. If there are still similar abnormal trigger conditions in adjacent time units, or if there are other abnormal trigger conditions that are directly related to that type of abnormal trigger condition, then the adjacent time units are merged into the same abnormal situation candidate event. The direct correlation means that different abnormal trigger conditions occur consecutively in time and are mutually interpretable in content; for example, a pushing action and a continuous shouting sound occur simultaneously, or a disputed statement is immediately followed by a device restart marker and a screen interruption. After expansion processing, an abnormal situation candidate event with a start time, end time, and abnormal trigger type is obtained.
[0046] It should be noted that during traffic police enforcement, some abnormal situations do not appear directly as obvious abnormal phenomena such as pushing, falling, collisions, or equipment malfunctions. Instead, while the enforcement actions are still related to checking documents, asking questions, or informing, changes gradually emerge, such as increased confrontational content in the voice and text messages, frequent changes in the speaker's state, a tendency for people to gather together, a tendency for close confrontation in the scene, and intermittent fluctuations in the status of equipment events. When these changes occur individually, they are usually insufficient to be identified as abnormal situations on their own; however, when these changes occur continuously and are interconnected within a continuous time unit, it can easily cause a lag in the identification of abnormal situations or lead to confusion between normal enforcement fluctuations and abnormal situations.
[0047] Therefore, in this embodiment, continuous change quantity extraction processing is first performed on candidate events of abnormal situations. Continuous change quantity refers to the amplitude of state change exhibited by candidate events of abnormal situations between adjacent time units. Continuous change quantity is obtained from the law enforcement context representation record formed in step two. For each time unit, the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state are read respectively, and the current time unit is compared with the previous time unit. Voice text content comparison is used to determine whether adversarial expressions have increased; the adversarial expressions refer to argumentative statements, expressions of refusal to cooperate, threatening expressions, and high-intensity warning expressions. Speaker state comparison is used to determine whether the frequency of alternating voices has increased and whether the number of speakers has increased. Personnel state comparison is used to determine whether the number of personnel targets has increased, whether the relative distance between personnel has shortened, and whether the personnel distribution range has narrowed towards the vicinity of law enforcement personnel. Scene state comparison is used to determine whether the general communication state has changed to a state of close confrontation between personnel or a state of personnel gathering. Sound event state comparison is used to determine whether continuous shouting, short-term violent knocking sounds, or continuous environmental noise have changed from none to present or from weak to strong. Equipment event status comparison is used to determine whether there has been an increase in changes in lens status, short-term image shift, partial obstruction, audio distortion, or instantaneous field of view deviation. Law enforcement action status comparison is used to determine whether law enforcement actions are still at the stage of verifying documents, asking questions, informing, or collecting evidence on site.
[0048] Based on the above comparison results, the following variables were generated: speech adversarial variation, speaker alternation variation, spatial compression variation, scene tightening variation, sound disturbance variation, visibility disturbance variation, and law enforcement action maintenance variation. Speech adversarial variation represents the increase in adversarial expression within the current time unit compared to the previous time unit, derived from comparisons of speech text content and semantic types in adjacent time units. Speaker alternation variation represents the degree of change resulting from variations in the number of speakers and alternating frequency of speech alternation within the current time unit, derived from comparisons of speaker states in adjacent time units. Spatial compression variation represents the degree of spatial convergence resulting from changes in the number of personnel targets, relative distances between personnel, and personnel distribution range, derived from comparisons of personnel states in adjacent time units. Scene tightening variation represents the shift in scene state from general communication to close-range confrontation between personnel. The change in the state of personnel gathering is determined by comparing the scene states in adjacent time units; the change in sound disturbance indicates the degree of enhancement of continuous shouting, violent knocking, and environmental noise, and is determined by comparing the sound event states in adjacent time units; the change in visibility disturbance indicates the degree of enhancement of changes in camera state, short-term image shift, partial obstruction, and sound pickup distortion, and is determined by comparing the equipment event states in adjacent time units; the maintenance of law enforcement actions indicates the degree to which the law enforcement action state remains at the level of verifying documents, asking questions, informing, or collecting evidence within consecutive time units, and is determined by comparing the law enforcement action states in adjacent time units. When the law enforcement action state corresponding to the current time unit has not escalated, the maintenance of law enforcement actions takes a higher value; when the law enforcement action state corresponding to the current time unit has changed to a verbal warning or a more intense enforcement action, the maintenance of law enforcement actions decreases.
[0049] After the continuous change extraction process is completed, the progressive anomaly accumulation is calculated for candidate events of abnormal situations. The progressive anomaly accumulation is used to characterize whether multiple weak anomalies continuously superimpose and form an evolutionary trend within consecutive time units. For the nth time unit, the change driving force of the current time unit is calculated based on the changes in speech adversarial behavior, speaker alternation, spatial compression, scene tightening, sound disturbance, and visibility disturbance. The change driving quantity Used to characterize the overall activity level of anomalous changes within the current time unit.
[0050] To clarify the formation process of the change-driving quantities, in a preferred embodiment, the voice adversarial change quantity, speaker alternation change quantity, spatial compression change quantity, scene tightening change quantity, sound disturbance change quantity, and visibility disturbance change quantity are first normalized to ensure that each change quantity falls within a uniform numerical range. The normalization process is performed by linearly mapping the minimum and maximum values of the corresponding change quantities within the continuous time units covered by the current abnormal situation candidate events. After normalization, the change-driving quantity for the nth time unit is defined as: Dn = wa·Aa,n + ws·Sa,n + wc·Ca,n + wj·Ja,n + wd·Da,n + wv·Va,n;
[0051] Wherein, Dn represents the change driving quantity in the nth time unit, Aa,n represents the normalized result of the speech adversarial change quantity in the nth time unit, Sa,n represents the normalized result of the speaker alternation change quantity in the nth time unit, Ca,n represents the normalized result of the spatial compression change quantity in the nth time unit, Ja,n represents the normalized result of the scene tightening change quantity in the nth time unit, Da,n represents the normalized result of the sound perturbation change quantity in the nth time unit, Va,n represents the normalized result of the visibility perturbation change quantity in the nth time unit, and wa, ws, wc, wj, wd, and wv represent the basic participation coefficients of the corresponding change quantities; the basic participation coefficients are pre-set non-negative numbers, and the sum of wa, ws, wc, wj, wd, and wv is 1.
[0052] Furthermore, when adaptively adjusting the composition of the change-driving quantity based on the change in visibility perturbation, when the change in visibility perturbation in the nth time unit is lower than the preset visibility stabilization threshold, the basic participation coefficients corresponding to the spatial compression change and the scene tightening change are increased, while the basic participation coefficients corresponding to the speech adversarial change, speaker alternation change, and sound perturbation change are decreased accordingly. When the change in visibility perturbation in the nth time unit is higher than the preset visibility decrease threshold, the basic participation coefficients corresponding to the speech adversarial change, speaker alternation change, and sound perturbation change are increased, while the basic participation coefficients corresponding to the spatial compression change and the scene tightening change are decreased accordingly. When the change in visibility perturbation is between the preset visibility stabilization threshold and the preset visibility decrease threshold, the basic participation coefficients remain unchanged. The change-driving quantity obtained after adaptive adjustment is still denoted as Dn.
[0053] In one example, wa, ws, wc, wj, wd, and wv can be set to 0.18, 0.12, 0.22, 0.20, 0.13, and 0.15, respectively. When video visibility is stable, wc and wj are increased by 0.05, and the corresponding values are proportionally deducted from wa, ws, and wd. When video visibility decreases, wa, ws, and wd are increased by 0.05, and the corresponding values are proportionally deducted from wc and wj. The above values are merely examples, and those skilled in the art can make equivalent substitutions based on different law enforcement environments, but this does not affect the implementation method of constructing change-driving quantities based on each change.
[0054] Then, based on the asymptotic accumulation of the previous time unit... The driving force of changes in the current time unit and the amount of law enforcement actions maintained in the current time unit. Calculate the cumulative asymptotic anomaly in the current time unit. The formula for its calculation is: ; in, This represents the cumulative asymptotic anomaly at the nth time unit. This represents the cumulative amount of asymptotic anomalies in the previous time unit. This represents the driving force of change in the nth time unit. This represents the duration of law enforcement actions in the nth time unit. This represents the historical retention coefficient. The historical retention coefficient controls the degree to which the cumulative results of the previous time unit are retained in the current time unit; a larger historical retention coefficient indicates that previously formed abnormal cumulative trends are retained more in the current time unit; a smaller historical retention coefficient indicates that the current time unit emphasizes newly emerging state changes. Because... and The products are multiplied and then used in the calculation. Therefore, the gradual accumulation of anomalies will only increase significantly when there are abnormal changes in the current time unit and the law enforcement action status remains in the normal law enforcement stage. If the law enforcement action status has been upgraded, the anomaly accumulation trend in that time unit will no longer continue to increase in a gradual evolutionary manner.
[0055] In the calculation of asymptotic cumulative anomalous quantities, the driving force of change is... Instead of using a fixed composition method, it adaptively adjusts based on the visibility state of the current time unit. When the change in visibility perturbation within the current time unit is lower than a preset visibility stability threshold, spatial compression change and scene tightening change are used as the main judgment criteria. When the change in visibility perturbation within the current time unit reaches or exceeds the preset visibility stability threshold, the influence of video-related changes on the change-driving factors is reduced, while the influence of speech adversarial changes, speaker alternation changes, and sound perturbation changes on the change-driving factors is increased. After this processing, when the camera shifts, partially obscures, or the sound pickup is distorted, it no longer relies on the already unstable video state for judgment, but continues to depict the abnormal evolution process based on the speech and sound changes that can still be stably obtained on-site.
[0056] After calculating the cumulative anomalous amount, the hysteresis response is calculated for candidate anomalous events. The hysteresis response is used to characterize the response relationship formed when different types of changes do not occur simultaneously, but rather sequentially within a continuous time unit. Specifically, within the continuous time unit covered by the candidate anomalous event, the time unit in which the speech adversarial change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, and visibility disturbance change first show a significant increase is determined. Then, the sequential adjacency relationship between different changes is checked. If the speech adversarial change increases significantly first, and then the spatial compression change or scene tightening change increases within a preset response time interval, it is determined that the speech adversarial change forms a forward response to the spatial change. If the spatial compression change or scene tightening change increases significantly first, and then the visibility disturbance change increases within a preset response time interval, it is determined that the spatial compression change forms a forward response to the visibility disturbance. If the above two types of forward responses occur consecutively, it is determined that there is a progressive linkage evolution within the candidate anomalous event.
[0057] To quantify this type of sequential response relationship, the lag response quantity at the nth time unit is defined. This represents the forward response intensity formed within the current time unit. The hysteresis response increases when the changes in speech adversarial behavior, spatial compression, and visibility perturbation within the current time unit satisfy the forward response condition; conversely, the hysteresis response decreases if these changes are temporally disjointed or separated by a time interval exceeding a preset response time interval. The preset response time interval is a continuous time range predetermined based on the time unit length used for time segmentation, used to limit the maximum interval between different changes that can be considered part of the same evolutionary process.
[0058] In a preferred embodiment, to quantify the sequential adjacency relationships between different changes, a significant increase threshold is first set for each of the following: speech adversarial change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, and visibility disturbance change. When the increment of any change in the current time unit relative to the previous time unit reaches the corresponding significant increase threshold, the change is considered to have significantly increased for the first time in the current time unit. Then, the sequential adjacency relationships between different changes are checked within a preset response time interval. If, after the speech adversarial change has significantly increased for the first time, a spatial compression change or a scene tightening change has significantly increased for the first time within the preset response time interval, a first-type forward response marker is generated. If, after the spatial compression change or a scene tightening change has significantly increased for the first time, a visibility disturbance change has significantly increased for the first time within the preset response time interval, a second-type forward response marker is generated. If the first-type and second-type forward response markers appear consecutively in consecutive time units, the corresponding abnormal situation candidate event is considered to have a progressively linked evolution.
[0059] To characterize the above sequential response relationship, the lag response of the nth time unit is defined as: Ln=p1·R1,n+p2·R2,n+p3·R3,n;
[0060] Where Ln represents the lag response in the nth time unit, R1,n represents the response result of the forward response formed by the change in speech adversarial power towards the change in spatial compression or the change in scene tightening in the nth time unit, R2,n represents the response result of the forward response formed by the change in spatial compression or the change in scene tightening towards the change in visibility perturbation in the nth time unit, R3,n represents the response result of the above two types of forward responses being consecutively true in the nth time unit, p1, p2, and p3 represent the participation coefficients of the corresponding response results, and the sum of p1, p2, and p3 is 1; when the corresponding response result is true, the value is 1, and when it is false, the value is 0, or a continuous value inversely proportional to the response time interval. The shorter the response time interval, the higher the value of the corresponding response result; the longer the response time interval, the lower the value of the corresponding response result.
[0061] In a preferred embodiment, the convergence factor is quantified by assessing the concentration of multiple types of changes within the same time interval. The convergence factor for the nth time unit is defined as: Cn = Qn / M;
[0062] Where Cn represents the convergence amount of the nth time unit, Qn represents the number of change categories that occur synchronously and reach a significant increase in the judgment threshold within the nth time unit and its adjacent time units, and M represents the total number of change categories participating in the convergence judgment. In this embodiment, M is 5, corresponding to speech adversarial change, spatial compression change, scene tightening change, sound disturbance change, and visibility disturbance change. When Qn increases, the convergence amount of the convergence increases; when various changes occur scattered and there is a lack of aggregation between different time intervals, the convergence amount of the convergence decreases.
[0063] In a preferred embodiment, the on-site deviation is quantified by the degree of deviation between the state of law enforcement action and the internal changes on-site. The on-site deviation for the nth time unit is defined as: Pn=G(n)·(u1·Aa,n+u2·Ca,n+u3·Ja,n+u4·Da,n+u5·Va,n); Where Pn represents the on-site deviation in the nth time unit, and G(n) represents the state maintenance function for the law enforcement action state corresponding to the nth time unit, which remains in the stage of verifying documents, inquiring about explanations, informing matters, or collecting on-site evidence. When the law enforcement action state in the nth time unit belongs to the stage of verifying documents, inquiring about explanations, informing matters, or collecting on-site evidence, G(n) is 1; when the law enforcement action state in the nth time unit has changed to a verbal warning or a higher intensity handling state, G(n) takes a decay value less than 1 or is 0. u1, u2, u3, u4, and u5 represent the participation coefficients of voice confrontation change, spatial compression change, scene tightening change, sound disturbance change, and visibility disturbance change in the on-site deviation, respectively, and the sum of u1, u2, u3, u4, and u5 is 1. After this processing, when the law enforcement action state is still in a stable stage on the surface, but the internal changes on the scene continue to tighten, the on-site deviation increases; when the law enforcement action state has been upgraded synchronously, the on-site deviation decreases.
[0064] After the hysteresis response is calculated, the on-site deviation is calculated for candidate events of abnormal situations. The on-site deviation is used to characterize the degree of inconsistency between the state of law enforcement action and the state of change on-site. Specifically, the state of law enforcement action within the current time unit is read, along with the voice text content, speaker status, personnel status, scene status, sound event status, and device event status within the same time unit. If the state of law enforcement action still corresponds to verifying documents, inquiring about explanations, informing about matters, or collecting evidence on-site, and at least two of the following changes in the same time unit—voice confrontation change, spatial compression change, scene tightening change, sound disturbance change, or visibility disturbance change—continue to increase, then on-site deviation is identified for that time unit. The on-site deviation increases with the duration of the aforementioned changes; when the state of law enforcement action has changed to a verbal warning or a higher-intensity handling state, the on-site deviation decreases. The on-site deviation reflects the degree to which the law enforcement action appears to be in a stable phase, while the internal changes on-site have continued to tighten.
[0065] After calculating the continuous change, gradual anomaly accumulation, hysteresis response, and on-site deviation, the linkage convergence value is calculated for candidate anomaly events. The linkage convergence value characterizes whether multiple types of changes gradually shift from scattered occurrences to concentrated occurrences within the same time interval. Specifically, each continuous time unit covered by the candidate anomaly event is examined. When voice adversarial changes, spatial compression changes, scene tightening changes, sound disturbance changes, and visibility disturbance changes initially appear scattered but then gradually converge within adjacent time units, the candidate anomaly event is considered to have linkage convergence. The linkage convergence value increases with the number of change categories occurring synchronously within the same time interval; it decreases when all types of changes remain scattered and do not approach each other. The linkage convergence value differs from the hysteresis response value. The hysteresis response value describes the forward-driving relationship formed sequentially between different changes, while the linkage convergence value describes the degree of temporal concentration of different changes after gradual evolution.
[0066] like Figure 4As shown, after the above calculations are completed, the credibility adjustment process is performed on the candidate events of abnormal situations. The data used for credibility adjustment comes from the low-credibility markers and invalid segment markers retained in step one, as well as the device event states formed in step two. If there are long-term low-credibility markers or invalid segment markers within the continuous time unit covered by the current candidate event of abnormal situations, the credibility of the video and audio content corresponding to that time interval is reduced; if the current candidate event of abnormal situations has changes in camera state or partial occlusion, but the device event state and the voice text content and sound event state remain continuous, the credibility of the voice-related changes is maintained; if the device event state shows a short-term screen deflection while the environmental supplementary data remains continuous, the credibility of the corresponding changes in the environmental supplementary data is maintained. After credibility adjustment, the continuous change amount, the gradual anomaly accumulation amount, the hysteresis response amount, the on-site deviation amount, and the linkage convergence amount all correspond to the credibility correction results.
[0067] To ensure the repeatability of the credibility adjustment results, in a preferred embodiment, a credibility adjustment coefficient Kn is first generated based on low-credibility markers, invalid segment markers, and device event states for the nth time unit; the credibility adjustment coefficient Kn ranges from 0 to 1. If a low-credibility marker exists within the time range corresponding to the nth time unit, a first attenuation is applied to Kn; if an invalid segment marker exists within the time range corresponding to the nth time unit, a second attenuation is applied to Kn; if there are changes in camera state or partial occlusion within the time range corresponding to the nth time unit, but the device event state, voice text content, and sound event state remain continuous, a maintenance correction is applied to the credibility adjustment coefficient corresponding to the voice-related changes; if there is a short-term screen deflection in the device event state display while the environmental supplementary data remains continuous, a compensation correction is applied to the corresponding changes in the environmental supplementary data.
[0068] Furthermore, the credibility adjustment coefficient for the nth time unit is defined as: Kn = kb·ki·ke; Where Kn represents the reliability adjustment coefficient for the nth time unit, kb represents the first reliability factor obtained based on the low reliability marker, ki represents the second reliability factor obtained based on the invalid segment marker, and ke represents the third reliability factor obtained based on the continuity of equipment event status and supplementary environmental data. If there is no low reliability marker, kb is 1; if there is a low reliability marker, kb is a decay value less than 1. If there is no invalid segment marker, ki is 1; if there is an invalid segment marker, ki is a decay value less than 1. If the equipment event status shows that the relevant data still has continuous support, ke is 1 or a compensation value greater than the basic decay value; if the equipment event status shows that the relevant data cannot form continuous support, ke is 1 or a decay value less than 1.
[0069] Based on this, credibility corrections are applied to continuous changes, gradual anomaly accumulation, lag response, on-site deviation, and linkage convergence, yielding corresponding credibility correction results. Specifically, the credibility correction results for speech adversarial changes, speaker alternation changes, spatial compression changes, scene tightening changes, sound disturbance changes, and visibility disturbance changes are obtained by multiplying their respective original results by the credibility adjustment coefficient Kn; the credibility correction results for gradual anomaly accumulation, lag response, on-site deviation, and linkage convergence are obtained by multiplying their respective original results by the average of the credibility adjustment coefficients of multiple time units within their coverage time range.
[0070] In one example, if both a low-confidence marker and an invalid segment marker exist simultaneously in the nth time unit, then kb and ki are set to 0.7 and 0.6, respectively; if the device event status displays continuous on-site audio data and continuous environmental supplementary data, then ke is set to 0.9 or 1. The above values are merely examples, and those skilled in the art can make equivalent settings based on different device stability and law enforcement environments.
[0071] After the credibility adjustment process is completed, anomaly evolution intensity is generated for candidate events of abnormal situations. The anomaly evolution intensity is not a direct, fixed-weighted sum of the results, but rather generated layer by layer according to the judgment order of continuous increase, sequential response, local convergence, and surface-level stable deviation. Specifically, it first checks whether the cumulative asymptotic amount continuously increases within a continuous time unit; when the cumulative asymptotic amount continuously increases, it checks whether the lag response amount reaches the preset response condition; when the lag response amount reaches the preset response condition, it checks whether the linkage convergence amount reaches the preset convergence condition; when the linkage convergence amount reaches the preset convergence condition, it checks whether the on-site deviation amount reaches the preset deviation condition. Only when the previous layer's judgment is valid does the next layer's judgment enter the calculation.
[0072] To characterize the above-mentioned layer-by-layer generation relationship, the anomalous evolution intensity of the nth time unit is defined. for: ; in, This represents the intensity of the anomalous evolution in the nth time unit. This represents the cumulative asymptotic anomaly at the nth time unit. This represents the lag response in the nth time unit. This represents the response enabling function generated based on the hysteresis response. This represents the convergence quantity of the linkage in the nth time unit. This represents the field deviation in the nth time unit. The response activation function controls the degree to which the anomaly evolution intensity is activated when the hysteresis response does not meet the preset response conditions. When the hysteresis response does not meet the preset response conditions, the response activation function outputs a lower value; when the hysteresis response meets the preset response conditions, the response activation function outputs an increased value. Due to the use of a product-based layer-by-layer activation relationship, when only scattered weak anomaly changes occur in a certain time unit, or when sequential response, linkage convergence, and field deviation are not formed, the anomaly evolution intensity will not be increased individually; however, when multiple changes simultaneously form a continuous increase, sequential response, local convergence, and field deviation in consecutive time units, the anomaly evolution intensity will increase rapidly.
[0073] In a preferred embodiment, the preset response condition, preset convergence condition, and preset deviation condition are implemented using a threshold comparison method. If the hysteresis response of the nth time unit is not lower than the preset response threshold, the response activation function is set to 1; otherwise, the response activation function is set to 0. If the linkage convergence of the nth time unit is not lower than the preset convergence threshold, local convergence is considered to be successful; otherwise, local convergence is considered to be unsuccessful. If the on-site deviation of the nth time unit is not lower than the preset deviation threshold, surface stationary deviation is considered to be successful; otherwise, surface stationary deviation is considered to be unsuccessful. The preset response threshold, preset convergence threshold, and preset deviation threshold can be determined based on historical law enforcement sample statistics, labeled sample training results, or results set by manual experience.
[0074] Furthermore, the anomaly scoring result is generated based on the distribution of anomaly evolution intensity within the time range corresponding to the anomaly warning event. The anomaly scoring result for the anomaly warning event is defined as: S = r1·Emax + r2·Eavg + r3·Tdur;
[0075] Where S represents the anomaly score result, Emax represents the peak anomaly evolution intensity within the time range of the anomaly warning event, Eavg represents the average anomaly evolution intensity within the time range of the anomaly warning event, Tdur represents the duration of the anomaly warning event after normalization, and r1, r2 and r3 represent the participation coefficients of the peak anomaly evolution intensity, average anomaly evolution intensity and duration in the anomaly score result, respectively, and the sum of r1, r2 and r3 is 1.
[0076] like Figure 3As shown, after the abnormal evolution intensity is generated, event-level aggregation processing is performed on candidate events for abnormal situations. Event-level aggregation processing refers to continuously checking the abnormal evolution intensity of each time unit along the start time to the end time of the candidate events for abnormal situations. If the abnormal evolution intensity of multiple adjacent time units all reach the preset evolution threshold, then the continuous time interval is identified as an abnormal evolution interval; if the abnormal evolution intensity only increases instantaneously in a single time unit and then immediately decreases, then the time unit is identified as corresponding to a general law enforcement fluctuation. For abnormal evolution intervals, their start time, end time, peak time unit, and peak abnormal evolution intensity are recorded; for general law enforcement fluctuations, only their corresponding time positions are retained and they are not used as abnormal situation warning events.
[0077] After event-level aggregation processing is completed, abnormal situation warning events are determined for the abnormal evolution interval. If the duration of the abnormal evolution interval reaches a preset duration threshold, and the peak abnormal evolution intensity within the abnormal evolution interval reaches a preset intensity threshold, then the corresponding abnormal evolution interval is determined as an abnormal situation warning event; if the duration of the abnormal evolution interval is less than the preset duration threshold, or the peak abnormal evolution intensity does not reach the preset intensity threshold, then it is recorded as an event that does not meet the warning conditions. For the determined abnormal situation warning events, their start time, end time, gradual abnormal accumulation trajectory, hysteresis response trajectory, linkage convergence trajectory, on-site deviation trajectory, and peak abnormal evolution intensity are recorded; for events that do not meet the warning conditions, their start time, end time, and peak abnormal evolution intensity are recorded.
[0078] In a preferred embodiment, the significantly increased judgment threshold, preset response time interval, preset response threshold, preset convergence threshold, preset deviation threshold, duration threshold, peak intensity threshold, and each scoring threshold in the warning level generation relationship are set based on the statistical distribution results of historical law enforcement samples, or based on the differentiation effect of labeled abnormal samples and normal samples; when the law enforcement environment, equipment model, or collection conditions change, the above thresholds are recalibrated.
[0079] In step four, when triggering an abnormal situation warning event, the start time, end time, abnormal trigger type, time correlation verification result, content correlation verification result, evidence consistency verification result, and abnormal score result of the abnormal situation warning event are first read. The abnormal trigger type is one or more of the following identified in step three: video abnormal trigger conditions, voice abnormal trigger conditions, sound event abnormal trigger conditions, equipment event abnormal trigger conditions, and law enforcement action abnormal trigger conditions. The time correlation verification result is the verification result given in step three regarding the concentration of various abnormal trigger contents over time. The content correlation verification result is the verification result given in step three regarding the semantic mutual support of various abnormal trigger contents. The evidence consistency verification result is the verification result given in step three regarding whether there are conflicts between on-site video data, on-site audio data, equipment event data, and law enforcement handling record data. The abnormal score result is the score result obtained after comprehensive calculation of the abnormal situation warning event in step three. After reading the above content, a warning level is generated according to the abnormal score result and the abnormal trigger type. The warning level is used to characterize the urgency of the abnormal situation warning event.
[0080] To make the relationship between warning levels clearer, in a preferred embodiment, warning levels are generated according to the combination of anomaly scoring results and anomaly trigger types. Specifically, the anomaly scoring results are first divided into multiple scoring intervals, and then the warning levels are determined by considering whether the anomaly trigger types include multiple simultaneous triggering conditions among visual anomaly triggering conditions, voice anomaly triggering conditions, sound event anomaly triggering conditions, device event anomaly triggering conditions, and law enforcement action anomaly triggering conditions.
[0081] In one implementation, a first warning level is generated when the abnormal score result is lower than a first scoring threshold; a second warning level is generated when the abnormal score result is not lower than the first scoring threshold and is lower than a second scoring threshold; a third warning level is generated when the abnormal score result is not lower than the second scoring threshold and is lower than a third scoring threshold; and a fourth warning level is generated when the abnormal score result is not lower than the third scoring threshold. If the abnormal trigger type includes at least two of the following: visual abnormal trigger conditions, voice abnormal trigger conditions, and device event abnormal trigger conditions; or if it includes at least two of the following: voice abnormal trigger conditions, sound event abnormal trigger conditions, and law enforcement action abnormal trigger conditions, then the warning level is increased by one level. If the abnormal trigger type includes only a single abnormal trigger condition, and there is a conflict in the evidence consistency verification results, then the original warning level is maintained or decreased by one level.
[0082] In one example, the first, second, and third scoring thresholds can be set to 0.30, 0.55, and 0.80, respectively; the warning levels can correspond to general concern, key concern, immediate action, and priority action. The above classification is merely an example, and those skilled in the art can make equivalent substitutions based on the needs of law enforcement supervision.
[0083] The warning level is determined as follows: when the anomaly score is in the high range and the anomaly trigger type includes two or more of the following: visual anomaly trigger condition, voice anomaly trigger condition, and sound event anomaly trigger condition, the anomaly warning event is designated as a high-level warning; when the anomaly score is in the middle range and the anomaly trigger type is mainly concentrated in a single source or dual source, the anomaly warning event is designated as a medium-level warning; when the anomaly score just reaches the preset anomaly threshold and the anomaly trigger type is relatively few, the anomaly warning event is designated as a low-level warning. After the warning level is determined, the start time, end time, anomaly trigger type, anomaly score result, and warning level of the anomaly warning event are written into the warning result record.
[0084] After the early warning triggering process is completed, evidence fragment generation is performed on the abnormal situation early warning event. An evidence fragment refers to a data segment extracted from the time interval corresponding to the abnormal situation early warning event that can characterize the occurrence process of the abnormal situation. When generating evidence fragments, firstly, based on the start and end times of the abnormal situation early warning event, the on-site video data, on-site audio data, supplementary environmental data, equipment event data, and law enforcement handling record data corresponding to the time interval are located from the standardized time-series fragment sequence obtained in step one. Then, using the start time of the abnormal situation early warning event as the central starting point, a preset backtracking time is performed, and using the end time of the abnormal situation early warning event as the central ending point, a preset extension time is performed to determine the evidence extraction time interval. The preset backtracking time is used to retain the on-site changes before the abnormal situation early warning event occurred, and the preset extension time is used to retain the state changes after the abnormal situation early warning event occurred. Based on the aforementioned evidence extraction time interval, corresponding video evidence segments are extracted from the on-site video data, corresponding audio evidence segments are extracted from the on-site audio data, corresponding supplementary evidence segments are extracted from the supplementary environmental data, equipment event records within the corresponding time interval are extracted from the equipment event data, and law enforcement handling record content within the corresponding time interval is extracted from the law enforcement handling record data.
[0085] After extracting the evidence data, key content annotation is performed on the evidence segments. Key content annotation refers to marking the abnormal triggering content already identified within the abnormal situation warning event to the corresponding time position in the evidence segment. Specifically, for video evidence segments, the personnel status, vehicle status, limb movement status, and scene status formed in steps two and three are read, and pushing actions, grabbing actions, running actions, falling actions, sudden vehicle displacement status, personnel gathering status, and lens obstruction status are marked to the corresponding video time positions; for audio evidence segments, the voice text content, semantic type, and sound event status are read, and argumentative statements, requests for help statements, warning statements, continuous shouting sounds, collision sounds, falling sounds, and violent knocking sounds are marked to the corresponding audio time positions; for equipment event records, recording stop markers, storage abnormality markers, network interruption markers, equipment restart markers, and lens status change markers are read and marked to the corresponding event occurrence time; for law enforcement handling records, the law enforcement action status and its generation time are read, and law enforcement actions such as verifying documents, questioning explanations, informing matters, on-site evidence collection, verbal warnings, and transfer for handling are marked to the corresponding time positions.
[0086] After the key content is annotated, the evidence fragments are processed to generate an evidence summary. The evidence summary is a continuous textual description of the key facts in an abnormal situation warning event. When generating the evidence summary, the key content markers corresponding to the abnormal situation warning event are first read in chronological order, and then organized into a continuous event description according to the chronological order. The continuous event description includes at least the on-site state before the abnormal situation warning event begins, the time and location of the first occurrence of the abnormal trigger content, the continuous change process of the abnormal state, the changes in equipment events, and the on-site state at the end of the abnormal situation warning event. For example, in an abnormal situation warning event, if a disputed statement appears first, followed by pushing and shoving, then continuous shouting and a change in camera status markers, the evidence summary will write these contents sequentially according to the chronological order, retaining the occurrence time of each content. All content used in the evidence summary is directly derived from the key content markers in the evidence fragments, without introducing any additional information outside the evidence fragments.
[0087] After the evidence fragment generation process, the early warning result records undergo early warning result correction processing. This correction process is used to eliminate false triggers in abnormal situation early warning events, merge cases where the same abnormal process is repeatedly identified, and correct the abnormality scoring results. During early warning result correction processing, multiple time-adjacent early warning result records are first read to check for overlapping or consecutive time intervals. If the time intervals of two early warning result records overlap and the abnormal trigger types are identical, or if the end time of the previous early warning result record and the start time of the subsequent early warning result record are only separated by a preset short interval, and their key content markers are consecutive, then they are considered to belong to the same abnormal process. Multiple early warning result records belonging to the same abnormal process are then merged. The merged early warning result record uses the earliest start time as the start time and the latest end time as the end time, and all abnormal trigger types are combined and integrated, with all key content markers rearranged in chronological order.
[0088] After completing the merging of duplicate warnings, a false trigger removal process is performed on the warning result records. A false trigger refers to a situation where, although the anomaly score reaches a preset anomaly threshold, a subsequent check reveals that the anomaly warning event lacks stable evidence support. When removing false triggers, the evidence fragments and summaries corresponding to the current warning result record are read. The system checks whether there are at least two mutually supporting anomalies among the conditions for triggering visual, audio, sound events, device events, and law enforcement actions. Simultaneously, it checks whether there are significant conflicts in the evidence consistency verification results. If the current warning result record is triggered by only a single source, and the evidence consistency verification results show significant conflicts, or if the key content markers cannot continuously appear in adjacent time positions, then the warning result record is identified as a false trigger record and removed from the valid warning result records. If the current warning result record is supported by multiple sources, and the evidence consistency verification results do not show significant conflicts, then it is retained as a valid warning result record.
[0089] Anomaly score correction processing is performed on the retained valid early warning results. This process adjusts the anomaly score based on the results of merging duplicate early warnings, removing false triggers, and the completeness of the evidence fragments. The completeness of the evidence fragments refers to the completeness of the on-site video data, on-site audio data, equipment event data, and law enforcement handling records within the current evidence fragment. This completeness is determined by checking for low-confidence markers, invalid segment markers, and data interruption segments within the current evidence fragment. When merging duplicate early warnings results in a longer-term continuous anomaly process, and this continuous anomaly process contains multiple high-intensity anomaly triggers, the anomaly score is increased. When removing false triggers reveals that some anomalies lack effective support, or when there are long-term data interruption segments in the evidence fragment, the anomaly score is decreased. After anomaly score correction, the corrected anomaly score is re-compared with the preset anomaly threshold and early warning level division range to obtain the corrected early warning level.
[0090] After the early warning results are corrected, a final early warning result record is generated. This record includes the start time, end time, trigger type, score, warning level, evidence fragment index, evidence summary, and correction status of the abnormal situation warning event. The evidence fragment index is a corresponding identifier used to locate the storage location of video evidence fragments, audio evidence fragments, supplementary environmental evidence fragments, equipment event records, and law enforcement handling records in the original dataset. The correction status indicates whether the final early warning result record has undergone duplicate warning merging, false trigger removal, and abnormal score correction processing. The generated final early warning result records are arranged in chronological order of the start time of the abnormal situation warning events, forming an abnormal situation warning result sequence corresponding to a single law enforcement process.
[0091] This invention also proposes an intelligent early warning system for abnormal situations in traffic police law enforcement audio and video, including: a multi-source law enforcement data standardization module, a law enforcement context representation module, an abnormal evolution early warning module, an early warning result processing module, and signal connections between the modules; The multi-source law enforcement data standardization module is used to retrieve on-site video streams, on-site audio streams, environmental video streams or supplementary audio streams, equipment operation logs and law enforcement handling records corresponding to the same law enforcement process, establish a unified timeline based on time correction anchor points, and generate a standardized time sequence fragment sequence. The law enforcement context representation module is used to extract features from on-site video data, on-site audio data, equipment event data, and law enforcement handling record data in a standardized time-series segment sequence, and to perform time alignment and fusion according to a unified time grid to generate law enforcement context representation records. The abnormal evolution early warning module is used to read the law enforcement context representation record in chronological order along the time segment. It establishes abnormal triggering conditions based on the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state, and constructs candidate events for abnormal situations. For the candidate events for abnormal situations, it compares the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state between adjacent time units, and extracts the voice confrontation change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change. Based on each change, it constructs change driving quantity, and combines it with the gradual abnormal accumulation quantity, lag response quantity, linkage convergence quantity, on-site deviation quantity, and credibility adjustment result to construct the abnormal evolution intensity, and then determines the abnormal situation early warning event. The early warning result processing module is used to establish an early warning level generation relationship for abnormal situation early warning events, construct evidence fragments, key content marking sequences and evidence summaries around the time interval corresponding to the abnormal situation early warning events, correct the early warning result records, and generate the final early warning result records.
[0092] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0093] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0094] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0095] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0096] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0097] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An intelligent abnormal situation early warning method for traffic police law enforcement audio and video, characterized in that, include: Retrieve on-site video streams, on-site audio streams, environmental video streams or supplementary audio streams, equipment operation logs and law enforcement handling records corresponding to the same law enforcement process, establish a unified timeline based on time correction anchor points, and generate a standardized time sequence fragment sequence; Feature extraction is performed on on-site video data, on-site audio data, equipment event data, and law enforcement handling record data in standardized time-series segments, and time alignment and fusion are performed according to a unified time grid to generate law enforcement context representation records; The law enforcement context representation record is read sequentially along the time sequence of the time segment. Anomaly triggering conditions are established based on the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state, and candidate events for abnormal situations are constructed. For the candidate events for abnormal situations, the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state between adjacent time units are compared. Voice adversarial change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change are extracted. Change driving quantity is constructed based on each change quantity. Combined with the gradual anomaly accumulation, lag response, linkage convergence, on-site deviation, and credibility adjustment results, the anomaly evolution intensity is constructed, and then the anomaly warning event is determined. Establish a relationship for generating warning levels for abnormal situation warning events, construct evidence fragments, key content marking sequences, and evidence summaries around the time interval corresponding to the abnormal situation warning events, correct the warning result records, and generate the final warning result records.
2. The method of claim 1, wherein the method further comprises: determining whether the audio and video of the traffic police law enforcement is abnormal; and if the audio and video of the traffic police law enforcement is abnormal, sending an abnormality warning to the traffic police law enforcement terminal. Collect on-site video streams, on-site audio streams, environmental video streams or supplementary audio streams, equipment operation logs, and law enforcement handling records corresponding to the same law enforcement process to construct the original data set corresponding to a single law enforcement process; extract the original timestamps from each data source and event records that can appear in more than two data sources as time correction anchors to establish a unified timeline and correct the timestamps of each data source; perform integrity verification on the on-site video streams, on-site audio streams, equipment operation logs, and law enforcement handling records, and add low-confidence markers to the corresponding time intervals on the unified timeline. 3.The method of claim 2, wherein the method further comprises: determining whether the audio-video of the traffic police law enforcement is abnormal based on the audio-video of the traffic police law enforcement. A sliding time window is set along a unified timeline, and the original data set is segmented into time-series segments and boundary corrections are performed by combining the start recording event, stop recording event, record generation time, and voice activity change time points. A time-series segment sequence is generated. According to the start and end times of each time-series segment on the unified timeline, the on-site video subsequence, on-site audio subsequence, environmental video subsequence, equipment operation log subsequence, and law enforcement handling record subsequence are aligned from multiple sources. Furthermore, sections with completely black screens, lens-obstructed sections, and sections with long periods of silence and no event changes in the equipment operation log are identified, and invalid section markers are added at the corresponding time positions.
4. The intelligent early warning method for abnormal situations in traffic police enforcement audio and video as described in claim 3, characterized in that: Video frame extraction and image content recognition are performed on the on-site video data in the standardized time sequence segment to construct video feature records containing personnel status, body movement status, and scene status; audio segmentation, speech content recognition, speaker separation, semantic tagging, and sound event recognition are performed on the on-site audio data to construct audio feature records containing speech text content, speaker status, semantic type, and sound event status. Perform event sequence parsing on the device operation logs to construct device event feature records corresponding to device event states; The law enforcement handling records are segmented into words and classified semantically to construct handling record feature records corresponding to the law enforcement action status. The video feature records, audio feature records, equipment event feature records and handling record feature records are grouped into continuous time units in a unified time grid. The law enforcement context representation records containing speech text content, speaker status, personnel status, scene status, sound event status, equipment event status and law enforcement action status are jointly generated. The law enforcement context representation records in adjacent time units are concatenated in chronological order to form a temporal segment-level law enforcement context representation sequence.
5. The intelligent early warning method for abnormal situations in traffic police enforcement audio and video as described in claim 4, characterized in that: The law enforcement context representation records are read sequentially along the time sequence segments. Each item—voice text content, speaker status, personnel status, scene status, sound event status, equipment event status, and law enforcement action status—is examined to establish abnormal trigger conditions for visual, voice, sound events, equipment events, and law enforcement actions. Specifically, based on the correspondence between changes in the number of personnel, changes in relative distance, tense scene conditions, adversarial expressions, high-frequency alternating voices from multiple speakers, continuous shouting, changes in camera status, and the state of law enforcement actions and the scene, abnormal trigger time units related to continuous changes are identified and extracted. Using the time unit that satisfies the abnormal trigger condition as the starting time unit, the process expands forward and backward according to the continuous occurrence of similar or related abnormal trigger conditions, constructing candidate events for abnormal situations with start time, end time, and abnormal trigger type.
6. The intelligent early warning method for abnormal situations in traffic police enforcement audio and video as described in claim 5, characterized in that: For candidate events of abnormal situations, sequential analysis is performed across modalities and states within continuous time units. The changes in speech text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state are compared between adjacent time units. This constructs speech adversarial change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change. Notably, instead of directly judging a single explicit abnormal phenomenon such as pushing, falling, collision, or device interruption, multiple weak changes that gradually emerge in the normal law enforcement phase, such as increased adversarialness, frequent speaker switching, personnel gathering and closing, close-range confrontation in the scene, and intermittent fluctuations in device state, are incorporated into the same continuous analysis framework. Furthermore, based on these multiple changes, change-driving quantities are constructed, and combined with the progressive anomaly accumulation quantity of the previous time unit and the law enforcement action maintenance quantity of the current time unit, a recursive calculation relationship for the progressive anomaly accumulation quantity is established to characterize the continuous superposition process of multiple weak anomaly changes within continuous time units.
7. The intelligent early warning method for abnormal situations in traffic police enforcement audio and video as described in claim 6, characterized in that: In the process of calculating the cumulative anomalous amount, the composition of the driving factors of change is adaptively adjusted according to the change in visibility disturbance. When video visibility is stable, spatial compression change and scene tightening change are highlighted. When video visibility decreases, the participation of speech adversarial change, speaker alternation change and sound disturbance change is increased, thereby establishing a multi-factor dynamic analysis structure for visibility fluctuation scenarios. On this basis, the time units of the first significant increase of various change amounts and their sequential adjacency relationships are sequentially compared to construct hysteresis response quantities to characterize the forward response chain formed in sequence between speech adversarial change, spatial compression change and visibility disturbance change. The process of different types of change further transitioning from sequential response to clustering in the same time interval is characterized as a linkage convergence quantity.
8. The intelligent early warning method for abnormal situations in traffic police enforcement audio and video as described in claim 7, characterized in that: By comparing the state of law enforcement actions with the changes in voice confrontation, spatial compression, scene tightening, sound disturbance, and visibility disturbance within the same time unit, a field deviation is constructed to characterize the state deviation relationship where the internal changes of the scene have been continuously tightening while the law enforcement actions are still superficially in the stage of verifying documents, asking questions, informing matters, or collecting evidence on site. Then, by combining low-confidence markers, invalid segment markers, and equipment event status, the confidence of continuous changes, gradual anomaly accumulation, delayed response, field deviation, and linkage convergence is adjusted to form corresponding confidence correction results. Subsequently, according to the judgment hierarchy of gradual anomaly accumulation, sequential response, local convergence, and surface stable deviation, a hierarchical opening relationship of anomaly evolution intensity is constructed. Based on the continuous distribution of anomaly evolution intensity within the candidate event time range, event-level aggregation is performed to distinguish between anomaly evolution intervals and general law enforcement fluctuations. Finally, by combining duration thresholds and peak intensity thresholds, anomaly warning events are determined.
9. The intelligent early warning method for abnormal situations in traffic police enforcement audio and video as described in claim 8, characterized in that: Read the abnormal trigger type and abnormal score result of the abnormal situation warning event, establish the warning level generation relationship, and write the start time, end time, abnormal trigger type, abnormal score result and warning level into the warning result record; An evidence collection time interval is established around the start and end times of the abnormal situation warning event. On-site video data, on-site audio data, supplementary environmental data, equipment event data, and law enforcement handling records are located and extracted from a standardized time-series fragment sequence to construct evidence fragments. Personnel status, body language status, scene status, voice and text content, semantic type, sound event status, equipment event status, and law enforcement action status are marked to their corresponding time positions within the evidence fragments to construct a key content marking sequence. The key content marking sequence is read in chronological order to generate an evidence summary for the corresponding abnormal situation warning event. The system performs overlap checks, continuity checks, and key content continuity checks on multiple adjacent early warning result records, establishes merging rules for the same abnormal process, merges early warning result records, removes false triggers, and corrects abnormal scores, generating a final early warning result record that includes abnormal trigger type, abnormal score result, early warning level, evidence fragment index, evidence summary, and correction status.
10. An intelligent early warning system for abnormal situations in traffic police enforcement audio and video, used to implement the intelligent early warning method for abnormal situations in traffic police enforcement audio and video as described in any one of claims 1-9, characterized in that, include: The module includes a multi-source law enforcement data standardization module, a law enforcement context representation module, an anomaly evolution early warning module, an early warning result processing module, and signal connections between the modules. The multi-source law enforcement data standardization module is used to retrieve on-site video streams, on-site audio streams, environmental video streams or supplementary audio streams, equipment operation logs and law enforcement handling records corresponding to the same law enforcement process, establish a unified timeline based on time correction anchor points, and generate a standardized time sequence fragment sequence. The law enforcement context representation module is used to extract features from on-site video data, on-site audio data, equipment event data, and law enforcement handling record data in a standardized time-series segment sequence, and to perform time alignment and fusion according to a unified time grid to generate law enforcement context representation records. The abnormal evolution early warning module is used to read the law enforcement context representation record in chronological order along the time segment. It establishes abnormal triggering conditions based on the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state, and constructs candidate events for abnormal situations. For the candidate events for abnormal situations, it compares the voice text content, speaker state, personnel state, scene state, sound event state, device event state, and law enforcement action state between adjacent time units, and extracts the voice confrontation change, speaker alternation change, spatial compression change, scene tightening change, sound disturbance change, visibility disturbance change, and law enforcement action maintenance change. Based on each change, it constructs change driving quantity, and combines it with the gradual abnormal accumulation quantity, lag response quantity, linkage convergence quantity, on-site deviation quantity, and credibility adjustment result to construct the abnormal evolution intensity, and then determines the abnormal situation early warning event. The early warning result processing module is used to establish an early warning level generation relationship for abnormal situation early warning events, construct evidence fragments, key content marking sequences and evidence summaries around the time interval corresponding to the abnormal situation early warning events, correct the early warning result records, and generate the final early warning result records.