A video stream real-time frame extraction system and method for public security interrogation
Patent Information
- Application Number
- CN202611265570.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]当被审讯人员实施举手、指认、拿取或展示物品等持续性动作时,动作幅度可能在发展过程中多次变化,并可能出现短暂停顿、再次增强或逐步恢复,仅依据局部峰值或单帧状态容易在动作尚未完整结束时提前确定抽帧位置
[0064]由于本发明不是依据单帧动作幅度或动作置信度立即抽取视频帧,而是基于连续视频帧中动作特征的变化方向、变化幅度及持续时间识别启动、发展、峰值、保持和恢复状态,并在形成峰值状态后结合第一恢复状态帧、恢复连续帧数及动作幅度下降比例确定有效动作完成节点,因此能够减少将动作尚未充分展开、短暂停顿或恢复过程中的视频帧误判为动作完成帧;又由于在动作启动后对动作相关历史视频帧设置缓存保护,并在受保护视频帧占用数量达到压缩触发阈值时转入阶段代表帧保护方式,按照实际形成的动作阶段保留基线代表帧、启动代表帧、发展代表帧、峰值候选帧、保持代表帧和恢复代表帧,因此能够在限定容量环形帧缓存条件下降低关键历史帧被循环覆盖的概率;进一步地,由于在有效动作完成后从受保护历史帧中回溯形成候选帧集合,并结合主体完整度、动作关键部位可见度、相关物品可见度、动作状态代表性、图像清晰度、遮挡程度及前后帧一致性评价证据信息完整度,因此能够避免仅以峰值状态帧作为最终抽帧结果,进而提高目标证据帧对动作主体、关键部位及相关物品信息的完整表达能力,同时通过关联保存动作区间和原视频时间戳,实现目标证据帧与连续原始审讯视频之间的可追溯复核。
Smart Images

Figure CN122802729A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of video image processing and computer vision technology, specifically to a real-time frame extraction system and method for video streams used in police interrogations. Background Technology
[0002] Police interrogation processes typically require real-time recording of continuous video footage, and the extraction of key video frames reflecting the interrogated person's actions from the video stream for subsequent review, verification, and evidence location. Existing video frame extraction methods usually select video frames based on fixed time intervals, image differences, action amplitude, or single-frame recognition confidence, rarely considering the continuous state evolution relationship of the same action from stable, initiation, development, peak, maintenance to recovery.
[0003] When the person being interrogated performs continuous actions such as raising their hand, pointing, taking or displaying items, the amplitude of the action may change multiple times during the process, and there may be short pauses, re-intensification, or gradual recovery. It is easy to determine the frame drop position in advance before the action has completely ended based solely on local peaks or single-frame states.
[0004] Meanwhile, real-time video processing equipment is usually only configured with a limited frame buffer. As video frames are continuously written, early action frames, frames near the peak, and video frames that can fully display the subject, key parts of the action, and related items may be overwritten in a loop before the action is determined to be completed. Simply extending the buffer will increase the storage resource consumption when multiple interrogation videos are processed in parallel.
[0005] Furthermore, frames with the largest movements may suffer from motion blur, subject occlusion, or incomplete display of related objects, and therefore do not necessarily possess high completeness of evidentiary information. Thus, identifying the evolution trend of action states online under limited frame buffering conditions, dynamically determining effective action completion nodes, and retaining key historical video frames for retrospective selection of target evidence frames with high completeness of evidentiary information after action completion have become the technical challenges that real-time interrogation video frame extraction needs to address.
[0006] In view of this, the present invention provides a real-time frame extraction system and method for video streams used in public security interrogations. Summary of the Invention
[0007] The purpose of this invention is to provide a real-time frame extraction system and method for video streams used in public security interrogations. This addresses the technical problem that in the existing real-time frame extraction process of public security interrogation videos, it is difficult to accurately identify the state evolution and actual completion time of continuous actions, and key historical frames are easily overwritten under limited frame buffer conditions, making it difficult to trace back and obtain video frames with high integrity of evidence information.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] In a first aspect, the present invention provides a real-time frame extraction method for video streams used in police interrogations, comprising the following steps:
[0010] Acquire real-time video stream of the interrogation scene, convert the real-time video stream into video frames with original video timestamps, and write them into a circular frame buffer of limited capacity in chronological order;
[0011] The action features of the same interrogator are extracted from consecutive video frames to form an action state observation sequence. Based on the direction, magnitude and duration of the change of the action features, the initiation state, development state, peak state, maintenance state and recovery state of the action are identified online.
[0012] Once the startup state is detected, a cache protection flag is set for the video frames related to the current action, and the protected video frames are updated as the action state evolves.
[0013] After the current action has formed the initiation state, development state and peak state, when the peak state is followed by the maintenance state or directly enters the recovery state, the effective action completion node is determined according to the degree of recovery and whether the same action is enhanced again.
[0014] When the number of protected video frames corresponding to the current action reaches the preset compression trigger threshold, the representative frame of the action stage is retained according to the action stage.
[0015] Using the effective action completion node as the backtracking starting point, the candidate frame corresponding to the current action is determined from the circular frame buffer. The target evidence frame is determined based on the completeness of the evidence information of the candidate frame, and the target evidence frame, action interval, and original video timestamp are associated and saved.
[0016] Furthermore, the action features of the same interrogator are extracted from consecutive video frames, including:
[0017] Perform subject detection and subject tracking on video frames, and configure the same subject identifier for the same interrogation subject in consecutive video frames;
[0018] Extract the key point location, key point displacement direction, key point displacement amount, movement range, movement speed, and torso posture of the interrogating subject;
[0019] Detect items related to the current action, determine the location of the items, the relative position of the items to the key parts of the action, and the visibility status of the items;
[0020] When the amplitude of the motion is continuously lower than the baseline update threshold, a stable baseline is established or updated based on the corresponding video frame, and the update of the stable baseline is stopped after the start state is detected.
[0021] Furthermore, the online recognition system identifies the initiation, development, peak, maintenance, and recovery states of actions, including:
[0022] When the amplitude of the action continuously reaches the number of activation confirmation frames and exceeds the activation threshold, the first video frame that exceeds the activation threshold is determined as the action activation frame.
[0023] When the motion characteristics continue to increase relative to the stable baseline but do not reach the peak determination criteria, it is determined to be in a developmental state;
[0024] When the amplitude of the motion reaches the local maximum value within the current motion range and no video frame with a larger amplitude of motion appears within the subsequent peak confirmation frame number, the video frame corresponding to the local maximum value is determined as the peak state frame.
[0025] When the change in the action feature relative to the peak state is continuously within the peak stability threshold range, it is determined to be a sustained state;
[0026] After the peak state or the hold state, when the motion amplitude continues to decrease and the key point moves towards the stable baseline position and meets the recovery determination condition, the video frame that first meets the recovery determination condition is determined as the first recovery state frame and enters the recovery state.
[0027] Further, determining the effective action completion node includes:
[0028] After the current action has reached the start state, development state, and peak state, when it enters the hold state or recovery state after the peak state, the number of consecutive recovery frames is accumulated starting from the first recovery state frame;
[0029] When the number of consecutive recovery frames reaches the number of recovery confirmation frames, and the decrease ratio of the current action amplitude relative to the peak action amplitude reaches the recovery threshold, the current video frame is determined as a valid action completion node.
[0030] If the motion amplitude is detected to increase again and the motion enhancement condition is met before the recovery confirmation frame count is reached, the accumulated recovery consecutive frame count is cleared, and the increased motion is returned to the current motion range.
[0031] When the recovery determination condition is met again, the first recovery state frame is re-determined, and the number of consecutive recovery frames is re-accumulated.
[0032] Furthermore, set cache protection flags for the current action-related video frames, including:
[0033] The video frames in the circular frame buffer are divided into normal frames, replaceable protection frames, and locked protection frames according to the buffer protection status.
[0034] The ordinary frames are allowed to be overwritten preferentially in the cyclic writing order. The replaceable protection frames are allowed to be replaced according to the integrity of the evidence information and the repetition of the action features, provided that the minimum number of representative frames for the corresponding action stage is maintained. The locked protection frames are prohibited from being overwritten in the ordinary cyclic writing method before the corresponding locking condition is released.
[0035] When the startup state is detected, a preset number of video frames before the action startup frame, the action startup frame, and the action-related video frames written after the action starts are set as protected video frames.
[0036] Furthermore, when a new video frame is written to the circular frame buffer:
[0037] If a normal frame is stored at the cache location pointed to by the write pointer, then the normal frame is overwritten with the new video frame;
[0038] If the cache location stores a replaceable protection frame, the completeness of the evidence information and the repetition of the action features of the new video frame are compared with those of the video frames already retained in the same action phase. Replacement is performed when the completeness of the evidence information of the new video frame is high. If the repetition of the action features reaches the repetition threshold and the completeness of the evidence information is not improved, no continuous protection mark is set for the new video frame.
[0039] If the cache location stores a lock protection frame, then overwriting the cache location is prohibited, and the write pointer is deferred to the next cache location.
[0040] Furthermore, when the number of protected video frames corresponding to the current action reaches the preset compression trigger threshold, the protection mode is switched from continuous protection mode to stage representative frame protection mode.
[0041] The compression trigger threshold is less than or equal to the maximum protection capacity that the current action is allowed to occupy in the circular frame buffer;
[0042] According to the action state, the protected video frames are divided into the pre-action baseline, start, development, peak, hold and recovery stages, and the baseline representative frame, start representative frame, development representative frame, peak candidate frame, hold representative frame and recovery representative frame are retained respectively.
[0043] When there are multiple video frames in the same action phase, the video frames with higher evidence information completeness are retained first, and the continuous protection of the corresponding video frames is lifted when the action feature repetition reaches the repetition threshold and the evidence information completeness is not improved. Among them, at least a preset minimum number of representative frames are retained for each action phase that has occurred.
[0044] The peak candidate frame consists of the peak state frame and video frames within a preset adjacent range that meet the minimum visibility requirement, and is updated accordingly when the peak state is updated.
[0045] Further, determining the completeness of the evidence information of the candidate frames includes:
[0046] Read peak candidate frames and hold-state frames from the protected video frames corresponding to the current action to form a candidate frame set;
[0047] When the current action has entered the stage representative frame protection mode, a candidate frame set is formed from the peak candidate frame, the hold representative frame, and the protected representative frame that is adjacent to the peak candidate frame in time;
[0048] Candidate frames that do not meet the minimum visibility requirements for the main body, the minimum visibility requirements for key parts of the action, or the minimum visibility requirements for related items involved in the current action are removed.
[0049] The remaining candidate frames are normalized and weighted based on the integrity of the main body, the visibility of key parts of the action, the visibility of related items, the representativeness of the action state, the image clarity, the degree of occlusion, and the consistency between consecutive frames to obtain the completeness of the evidence information. The candidate frame with the highest completeness of the evidence information is then determined as the target evidence frame.
[0050] If the current action does not involve related items, the visibility of related items will not be used as a valid evaluation item, and the weights of the remaining valid evaluation items will be renormalized.
[0051] Secondly, the present invention provides a real-time frame extraction system for video streams used in police interrogation, for executing the real-time frame extraction method for video streams described in the first aspect, including:
[0052] The video stream access module is used to receive real-time video streams from the interrogation site and generate video frames with the original video timestamps.
[0053] The finite frame buffer module is used to cyclically write to a circular frame buffer of limited capacity in chronological order, and to protect, replace, and retain the stage representative frame for the current action-related video frame.
[0054] The action feature extraction module is used to extract action features of the same interrogating subject;
[0055] The state evolution recognition module is used to identify the start-up state, development state, peak state, maintenance state, and recovery state based on the direction, amplitude, and duration of the change in action characteristics.
[0056] The completion node determination module is used to accumulate the number of consecutive recovery frames starting from the first recovery state frame, and determine the valid action completion node based on the recovery duration and whether the same action is enhanced again;
[0057] The candidate frame backtracking and completeness evaluation module is used to determine candidate frames by taking the effective action completion node as the backtracking starting point, and to determine the target evidence frame based on the completeness of the evidence information of the candidate frames.
[0058] An evidence frame storage device is used to associate and save the target evidence frame, action range, and original video timestamp.
[0059] Furthermore, it also includes interrogation camera equipment, original video storage equipment, and interrogation review terminals;
[0060] The interrogation camera device is used to send the same real-time video stream to the original video storage device and the video stream access module respectively;
[0061] The original video storage device is used to continuously store the original video stream and its corresponding video file identifier, video channel identifier, and original video timestamp.
[0062] The interrogation review terminal is used to locate the corresponding original video file from the original video storage device based on the video channel identifier associated with the target evidence frame and the original video timestamp, and to obtain the original video segment containing the action interval.
[0063] Compared with the prior art, the beneficial effects of the present invention are:
[0064] Because this invention does not immediately extract video frames based on the motion amplitude or motion confidence of a single frame, but rather identifies the initiation, development, peak, maintenance, and recovery states based on the direction, amplitude, and duration of motion feature changes in continuous video frames, and determines the valid motion completion node after the peak state is formed by combining the first recovery state frame, the number of consecutive recovery frames, and the percentage decrease in motion amplitude, it can reduce the misclassification of video frames where the motion has not fully unfolded, are briefly paused, or are in the recovery process as motion completion frames. Furthermore, because it sets up cache protection for historical video frames related to the motion after the motion is initiated, and switches to a stage representative frame protection mode when the number of protected video frames reaches the compression trigger threshold, it retains the baseline representative frame, initiation representative frame, and development representative frame according to the actual formed motion stage. The system uses a circular frame buffer with a limited capacity to reduce the probability of key historical frames being overwritten. Furthermore, by backtracking from protected historical frames to form a candidate frame set after an effective action is completed, and combining the integrity of the subject, the visibility of key parts of the action, the visibility of related items, the representativeness of the action state, the image clarity, the degree of occlusion, and the consistency between previous and subsequent frames to evaluate the completeness of the evidence information, it can avoid using only the peak state frame as the final frame extraction result. This improves the ability of the target evidence frame to fully express the information of the subject, key parts, and related items of the action. At the same time, by associating and saving the action interval and the original video timestamp, it enables traceable verification between the target evidence frame and the continuous original interrogation video. Attached Figure Description
[0065] Figure 1 This is a schematic diagram of the actual deployment structure and data flow of the real-time frame extraction system for video streams according to the present invention.
[0066] Figure 2 This is a schematic diagram of the overall process of the real-time frame extraction method for video streams used in public security interrogation according to the present invention.
[0067] Figure 3 This is a schematic diagram illustrating the evolution of the interrogation subject's actions from a stable state, an initiation state, a development state, a peak state, a maintenance state, to a recovery state, as well as the temporal relationship between the action initiation frame, the peak state frame, the peak candidate frame, the first recovery state frame, and the effective action completion node.
[0068] Figure 4 This is a schematic diagram illustrating the cyclic writing, cache protection state, and cache overwrite priority relationship of the ring frame cache with limited capacity in this invention.
[0069] Figure 5 This diagram illustrates the process of overwriting, replacing, and determining the write pointer continuation for ordinary frames, replaceable protection frames, and locked protection frames when writing new video frames to the circular frame buffer.
[0070] Figure 6 This is a schematic diagram illustrating the process of switching from continuous protection mode to stage representative frame protection mode when the number of protected video frames corresponding to the current action of the present invention reaches a preset compression trigger threshold, and selecting, updating and retaining stage representative frames according to the actual action stage formed by the current action.
[0071] Figure 7 This is a schematic diagram illustrating the process of determining the target evidence frame after candidate frames are screened based on minimum visibility requirements, the completeness of evidence information is evaluated, and they are sorted according to the present invention.
[0072] In the diagram: 110, interrogation camera equipment; 120, raw video storage device; 200, video stream access module; 300, real-time frame extraction processing device; 310, limited frame buffer module; 320, action feature extraction module; 330, state evolution recognition module; 340, completion node determination module; 350, candidate frame backtracking and integrity evaluation module; 400, evidence frame storage device; 500, interrogation review terminal. Detailed Implementation
[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0074] Example 1
[0075] like Figure 1As shown, this embodiment provides a real-time frame extraction system for video streams used in public security interrogations, including an interrogation camera device 110, an original video storage device 120, a video stream access module 200, a real-time frame extraction processing device 300, an evidence frame storage device 400, and an interrogation review terminal 500.
[0076] The interrogation camera device 110 is used to collect real-time video streams from the interrogation site and sends the same real-time video stream to the original video storage device 120 and the video stream access module 200 respectively. The original video storage device 120 is used to continuously store the original video stream without frame extraction processing and record the video file identifier, video channel identifier, and original video timestamp, so that the target evidence frames extracted later can be relocated to the original video stream according to the corresponding original video timestamp.
[0077] The video stream access module 200 receives the real-time video stream output by the interrogation camera device 110, decodes the real-time video stream, generates continuous video frames according to the video acquisition sequence, and configures a frame identifier, original video timestamp, and video channel identifier for each video frame. Video frame sequences are established separately for different video channels, and video frames from different video channels are not merged into the same action state observation sequence.
[0078] The real-time frame extraction processing device 300 includes a limited frame buffer module 310, an action feature extraction module 320, a state evolution recognition module 330, a completion node determination module 340, and a candidate frame backtracking and integrity evaluation module 350.
[0079] The limited frame buffer module 310 employs a circular frame buffer with a limited capacity. The circular frame buffer includes multiple buffer locations arranged in a circular write order, and is configured with a write pointer to indicate the next buffer location to be written. Each buffer record includes at least a frame identifier, original video timestamp, video image, interrogation subject identifier, subject detection area, motion features, motion status, image quality information, and buffer protection flag.
[0080] like Figure 4 As shown, based on the relationship between the cache record and the current action, as well as the cache protection status, the video frames in the circular frame cache are divided into ordinary frames, replaceable protection frames, and locked protection frames.
[0081] Normal frames are video frames that do not have continuous protection requirements set, and newly written video frames are allowed to be overwritten preferentially according to predetermined overwrite rules.
[0082] Replaceable protection frames are video frames that belong to the current action range and have been marked with protection. However, while maintaining the minimum number of representative frames for the corresponding action phase, replacement is allowed based on the completeness of evidence information and the repetition of action features.
[0083] A lockable protection frame is a video frame that must not be overwritten by a normal cyclic write operation before the current action is completed. A lockable protection frame includes at least the currently valid peak candidate frame and representative frames that must be retained after being determined according to the minimum retention quantity for each action phase. For baseline frames used to represent a stable state before the action starts, they are used as lockable protection frames if they are the only baseline representative frame for the current action.
[0084] Normal frames, replaceable guard frames, and locked guard frames represent the cached state of a record during the current action processing phase, rather than corresponding to a fixed cache location. The cached state of the same record can change as the action state changes, as the representative frame is updated, and as the action is completed.
[0085] The minimum retention quantity for each stage is a pre-configured minimum retention quantity of at least one frame for each identified action stage. This ensures that at least one video frame representing the action stage is retained during buffer cycle overwriting and stage representative frame compression. The minimum retention quantities for different action stages can be the same or different, and are pre-configured based on the circular frame buffer capacity and the importance of each action stage.
[0086] The motion feature extraction module 320 sequentially performs subject detection, subject tracking, human keypoint detection, related object detection, and image quality detection on the video frames. Subject tracking is used to assign the same subject identifier to the same interrogation subject in consecutive video frames.
[0087] Human body key points include at least the head, shoulders, elbows, wrists, and torso key points. Motion features include at least the location of key points, the direction of key point displacement, the amount of key point displacement, the amplitude of the movement, the speed of movement, the torso posture, the location of related objects, the relative position of related objects to the key body parts involved in the movement, and the confidence level of the movement recognition. Image quality information includes at least image sharpness and the degree of occlusion.
[0088] The motion feature extraction module 320 establishes a stable baseline based on continuous video frames when the interrogator is in a stable state. The stable baseline includes at least the reference positions of key human body points and the reference posture of the torso; when the current action involves related objects, the stable baseline also includes the reference position of the related objects.
[0089] When the range of motion of the interrogator continuously falls below the baseline update threshold, a stable baseline is established or updated based on the corresponding video frames. Once an action is detected as initiated, the update of the stable baseline stops until the current action is completed or determined to have terminated abnormally, thus preventing ongoing actions from being incorrectly incorporated into the stable baseline.
[0090] The state evolution recognition module 330 arranges the action features according to the same subject identifier and the original video timestamp to form a continuously updated action state observation sequence. Based on the direction, magnitude and duration of the change of the action features, the action state of the interrogating subject is identified as a stable state, an initiation state, a development state, a peak state, a maintained state or a recovery state.
[0091] In a stable state, the range of motion of the interrogator remains within a preset fluctuation range corresponding to the stable baseline. When the range of motion continuously reaches the number of activation confirmation frames and exceeds the activation threshold, the first video frame that exceeds the activation threshold is determined as the action activation frame, and the current action is transferred from the stable state to the activation state.
[0092] After the action is initiated, if at least one of the following is continuously increasing relative to the stable baseline: the magnitude of the action, the displacement of key points, the displacement of related objects, or the confidence level of the action recognition, and the peak determination condition has not yet been met, the corresponding video frame is determined as a development state frame.
[0093] When the motion amplitude of a video frame reaches a local maximum within the current motion range, and no video frame with a larger motion amplitude appears within the subsequent peak confirmation frames, the video frame corresponding to this local maximum value is determined as the peak state frame. If a video frame with a larger motion amplitude appears during the subsequent motion development, the current peak marker of the original peak state frame is canceled, and a new peak state frame is determined.
[0094] After the peak state frame is determined, a peak candidate range is determined within a preset adjacent range before and after the peak state frame, using the peak state frame as the time center. Video frames that meet the minimum visibility requirement within this peak candidate range are then determined as peak candidate frames. The peak state frame is used to determine the location of the local maximum motion amplitude of the current action; the peak candidate frames are used to form a key evaluation range near the peak. Therefore, the peak candidate frames can include the peak state frame and multiple video frames located before and after the peak state frame.
[0095] Once the peak state is formed, when the change in motion characteristics of consecutive video frames relative to the peak state remains within the peak stability threshold range, the corresponding video frames are designated as hold-up frames. Therefore, hold-up states are established based on the already determined peak state and are not considered as alternative states when no peak state exists.
[0096] After the peak state, the current action can either enter a sustained state and then a recovery state, or it can directly enter the recovery state without forming a clear sustained state. When the motion amplitude shows a continuous decreasing trend, the key points of the human body continuously move towards the position corresponding to the stable baseline, and the above changes meet the preset recovery judgment conditions, the video frame that first meets the recovery judgment conditions is determined as the first recovery state frame, and the current action is identified as a recovery state starting from the first recovery state frame. The decrease in motion amplitude in a single video frame is not used as the sole criterion for determining the recovery state.
[0097] The peak state frame, peak candidate frame, first recovery state frame, and effective action completion node each perform different functions: the peak state frame is used to locate the local maximum position of the action amplitude; the peak candidate frame is used to establish the candidate evaluation range near the peak; the first recovery state frame is used to determine the cumulative starting point of the number of consecutive recovery frames; and the effective action completion node is used to confirm that the current action has met the preset completion conditions and trigger the historical cache backtracking.
[0098] like Figure 3 As shown, in order to uniformly represent the evolution process of action states with different durations and absolute action amplitudes, the time position in the current action interval is converted into normalized time, and the action amplitude of the interrogator relative to the stable baseline is converted into normalized action amplitude. The values of normalized time and normalized action amplitude are both in the range of 0 to 1. Figure 3 The horizontal axis represents normalized time, and the vertical axis represents normalized motion amplitude.
[0099] Figure 3 The motion evolution process is shown sequentially as follows: stable state, initiation state, development state, peak state, maintenance state, and recovery state. The initiation frame corresponds to the starting position of the transition from the stable state to the initiation state; the peak state frame corresponds to the position of the local maximum motion amplitude determined after peak confirmation; the peak candidate frame is located within a preset proximity range near the peak state frame; the first recovery state frame corresponds to the video frame that first meets the recovery determination conditions; and the effective motion completion node is located after the first recovery state frame, corresponding to video frames that continuously meet the motion completion conditions during the recovery process.
[0100] Figure 3 The state boundaries and the corresponding x-coordinates of each keyframe or node are used to illustrate the temporal relationships, but do not imply that the corresponding state or node must occur at a fixed normalized time value. The actual start time, end time, and duration of each state are dynamically determined based on the real-time action state observation sequence. For actions that directly enter the recovery state after the peak state, the duration of the hold state can be shortened to the point where it does not form an independent hold state, but this does not affect the determination of the first recovery state frame and the effective action completion node.
[0101] After the state evolution recognition module 330 recognizes the action start state, the limited frame buffer module 310 sets a preset number of video frames before the action start frame, the action start frame, and the action-related video frames continuously written after the action starts as protected video frames, and continuously updates the protected video frames according to the change of the action state.
[0102] The node determination module 340 performs a valid action completion judgment after the current action has formed an initiation state, a development state, or a peak state. When a holding state is formed after the peak state, recovery confirmation is performed after the holding state ends and the recovery state is entered; when no obvious holding state is formed after the peak state and the recovery judgment condition is met directly, recovery confirmation is performed directly.
[0103] When the recovery determination condition is met for the first time, the node determination module 340 determines the corresponding video frame as the first recovery state frame and starts accumulating the number of consecutive frames to be recovered from the first recovery state frame.
[0104] When the number of consecutive recovered frames reaches the preset number of recovery confirmation frames, and the decrease ratio of the current motion amplitude to the peak motion amplitude reaches the recovery threshold, the current video frame that meets the above conditions will be determined as a valid motion completion node.
[0105] The first recovery state frame only indicates that the current action has started the continuous recovery process, and does not mean that the current action has been completed. The valid action completion node must be located after the first recovery state frame, and is determined by the number of consecutive recovery frames and the decrease ratio of the action amplitude relative to the peak action amplitude.
[0106] Before the number of consecutive recovery frames reaches the number of recovery confirmation frames, if the motion amplitude increases again and meets the conditions for motion enhancement, the accumulated number of consecutive recovery frames is cleared, the increased motion is returned to the current motion range, and motion state evolution recognition continues. When the recovery determination conditions are met again subsequently, the first recovery state frame is redefined, and the number of consecutive recovery frames is re-accumulated starting from the redefined first recovery state frame.
[0107] Therefore, a brief pause or temporary drop during the action can be distinguished from the actual completion of the action.
[0108] When the subject of the interrogation is briefly occluded, causing some key human features to be temporarily undetectable, the recovery status or completion of the action is not determined solely by a decrease in action recognition confidence. If the number of frames of occlusion does not exceed the preset occlusion tolerance, the current action state before occlusion is maintained, and the historical video frames corresponding to the current action are preserved. After the subject of the interrogation reappears, the current action state is updated continuously based on the subject identifier, key feature positions, and the continuity of action features before and after occlusion in the reappearing video frames.
[0109] After determining the valid action completion node, the candidate frame backtracking and integrity evaluation module 350 reads the protected video frame corresponding to the current action from the limited frame cache module 310 based on the action start frame, the valid action completion node, the action instance identifier, and the cache protection flag.
[0110] For each current action, an action instance identifier is established, and the same action instance identifier is associated with the same action from the start until the effective action is completed. When a new independent action occurs after the action is completed, a new action instance identifier is established for the new action, so that multiple actions occurring consecutively by the same subject can be cached, completed, and backtracked separately.
[0111] Candidate frames consist of peak candidate frames and hold-state frames actually formed by the current action. When the current action has not formed a clear hold-state, it is not required that the candidate frame set must contain hold-state frames; instead, the candidate frame set is formed by peak candidate frames. There is no state determination path that "has a hold-state but has not formed a peak state."
[0112] The candidate frame backtracking and integrity evaluation module 350 first performs a minimum visibility requirement judgment on the candidate frames. The minimum visibility requirements include the interrogation subject meeting the minimum visibility requirement, the key parts performing the current action meeting the minimum visibility requirement for key parts of the action, and the relevant items meeting the minimum visibility requirement when the current action involves relevant items.
[0113] Video frames that do not meet any of the applicable minimum visibility requirements are removed from the candidate frame set.
[0114] For candidate frames that meet the minimum visibility requirements, the following criteria are determined: subject integrity, visibility of key parts of the action, visibility of related objects, representativeness of the action state, image clarity, degree of occlusion, and consistency between consecutive frames.
[0115] Subject integrity indicates whether the body area of the interrogating subject related to the current action is completely located in the valid frame; Visibility of key parts of the action indicates the visibility of the hand, wrist, upper limb, or other key parts used to complete the current action; Visibility of related objects indicates the completeness of the detection of objects involved in the current action and the degree to which the spatial relationship between the objects and the key parts of the action can be identified; Representativeness of action state indicates the degree to which the candidate frame expresses the main behavioral features of the current action; Image sharpness reflects motion blur and imaging quality; Occlusion degree reflects the degree to which the interrogating subject or related objects are occluded by other targets; Consistency between consecutive frames is used to exclude instantaneous detection errors or single-frame anomalies.
[0116] The integrity of the subject, the visibility of key parts of the action, the visibility of related objects, the representativeness of the action state, the image clarity, and the consistency between consecutive frames are converted into normalized values that increase with the degree of information integrity. The degree of occlusion is converted into an evaluation value that increases with the degree of effective visibility. These values are then combined according to preset weights to obtain the completeness of the evidence information of the candidate frames.
[0117] The weights corresponding to each evaluation item are pre-configured based on the degree of influence of the corresponding evaluation item on the completeness of the action evidence information, and are then normalized. If the current action does not involve related items, the visibility of related items is not used as a valid evaluation item for that candidate frame, and the weights of the remaining valid evaluation items are re-normalized to avoid the absence of related items affecting the evaluation results of the completeness of the evidence information.
[0118] The candidate frame backtracking and completeness evaluation module 350 arranges candidate frames from high to low according to the completeness of evidence information, and determines the candidate frame with the highest completeness of evidence information as the target evidence frame. When at least two candidate frames have the same completeness of evidence information, the video frame with the smaller time distance from the peak state frame is selected first; if the time distance is still the same, the video frame with the earlier original video timestamp is selected.
[0119] Therefore, the peak state frame is not directly equivalent to the peak candidate frame, nor is the peak candidate frame directly equivalent to the target evidence frame. The peak state frame solves the problem of determining the location of the action peak, the peak candidate frame solves the problem of which historical frames near the peak enter the subsequent evaluation range, and the target evidence frame is the video frame finally determined from the candidate frames based on the completeness of the evidence information.
[0120] The evidence frame storage device 400 stores the target evidence frame and associates it with the subject identifier, action instance identifier, action start time, peak time, time corresponding to the first recovery state frame, effective action completion time, video channel identifier, original video timestamp, action range, and evidence information completeness.
[0121] The interrogation review terminal 500 locates the corresponding original video file from the original video storage device 120 based on the video channel identifier associated with the target evidence frame and the original video timestamp, and obtains the original video segment containing the action interval, so as to realize the corresponding review between the target evidence frame and the continuous original interrogation video.
[0122] Example 2
[0123] like Figure 2 As shown in the figure, this embodiment provides a real-time frame extraction method for video streams used in public security interrogations, including the following steps.
[0124] S101, acquire the real-time video stream and establish a limited frame buffer. The video stream access module 200 acquires the real-time video stream output by the interrogation camera device 110, decodes the real-time video stream into video frames with the original video timestamp, and writes them into the limited frame buffer module 310 according to the video frame acquisition time order.
[0125] In this embodiment, the real-time video stream is captured at a frame rate of 25 frames per second, and the circular frame buffer has a capacity of 300 frames, which corresponds to storing approximately 12 seconds of video frames. A write pointer is set in the circular frame buffer to indicate the target buffer location for the next video frame.
[0126] The frame rate and ring frame buffer capacity described above are used to illustrate the specific processing procedure of this embodiment. In other embodiments, the ring frame buffer capacity can be configured based on the real-time video frame rate, the buffer space that the processing device can allocate, the number of video channels processed in parallel, and the allowable video processing latency.
[0127] S102, extract action features and establish action state observation sequence. The action feature extraction module 320 performs subject detection and subject tracking on continuous video frames, configures the same subject identifier for the same interrogation subject, and extracts the key point position, key point displacement direction, key point displacement amount, action amplitude, movement speed and torso posture of the interrogation subject.
[0128] When there are objects related to the current action in the video frame, the position of the objects, the relative position between the objects and the key parts of the action, and the visibility status of the objects are extracted simultaneously.
[0129] The motion feature extraction module 320 also extracts the integrity of the subject, image clarity, and degree of occlusion, and writes the corresponding information, along with the frame identifier and the original video timestamp, into the corresponding cache record.
[0130] The state evolution recognition module 330 establishes a sliding observation interval using consecutive video frames of the same interrogation subject. In this embodiment, 15 consecutive frames are used as the sliding observation interval. Each time a new video frame is received, the new video frame is added to the sliding observation interval, and the earliest video frame is removed to form a continuously updated action state observation sequence.
[0131] When the motion amplitude is consistently below the baseline update threshold, a stable baseline is established or updated based on the corresponding consecutive video frames; the update of the stable baseline is stopped after the motion is detected.
[0132] S103, online recognition of the evolution trend of action state, such as Figure 3As shown, in a stable state, the normalized motion amplitude of the interrogator remains near the stable baseline. When the motion amplitude exceeds the activation threshold for 5 consecutive frames, the first video frame that exceeds the activation threshold is identified as the action activation frame, an action instance identifier is established for the current action, and the current action is transitioned from the stable state to the activation state.
[0133] When an action is initiated, the 25 frames before the action initiation frame, the action initiation frame, and the action-related video frames continuously received after the action is initiated are set as protected video frames, so that when determining the valid action completion node later, the historical information before and after the action is initiated can still be obtained from the limited historical cache.
[0134] After the action is initiated, the development status is determined based on the continuous changing trends of the action amplitude, key point displacement, and related object displacement in consecutive video frames.
[0135] When the amplitude of the motion reaches the local maximum value in the current motion range, and no video frame with a larger amplitude of motion appears in the subsequent 5 frames, the video frame corresponding to the local maximum value is determined as the peak state frame.
[0136] Centered on the peak state frame, the five frames before and after the peak state frame are taken as the peak candidate range in this embodiment, and the video frames that meet the minimum visibility requirements within this range are determined as peak candidate frames.
[0137] If a larger movement occurs in a subsequent video frame, the new local maximum value will be re-confirmed for peak value. After the new local maximum value meets the peak value determination condition, the peak state frame will be re-determined, and the peak candidate range and the current valid peak candidate frame will be updated accordingly.
[0138] Once the peak state is established, if the change ratio of the motion amplitude of consecutive video frames relative to the peak motion amplitude remains within the peak stability threshold range, the corresponding video frames are identified as hold-up frames. Hold-up frames are only identified after the peak state has been determined.
[0139] After the peak state, if consecutive video frames remain within the peak stability threshold range, a hold state is formed first; if a hold state that meets the conditions for continuity is not formed and the amplitude of the motion has continued to decrease, the peak state can be directly transitioned to the recovery state.
[0140] When the amplitude of the movement continues to decrease, the key points of the main body continue to move towards the position corresponding to the stable baseline, and the preset recovery judgment condition is met for the first time, the corresponding video frame is determined as the first recovery state frame, and the recovery state is identified from the video frame.
[0141] Figure 3Normalized time (0-1) is used as the x-axis, and normalized motion amplitude (0-1) is used as the y-axis. The actual durations of the steady state, initiation state, development state, peak state, hold state, and recovery state, as well as the boundary positions between each state, are dynamically determined based on the motion characteristics in consecutive video frames, not based on... Figure 3 The fixed horizontal axis value shown is used as the state switching threshold.
[0142] Figure 3 The action start frame is used to mark the start position of the action interval; the peak state frame is used to mark the position of the local maximum action amplitude; the peak candidate frame is used to form a candidate evaluation range near the peak; the first recovery state frame is used to mark the cumulative starting point of the number of consecutive recovery frames; and the effective action completion node is used to mark the position where the recovery process has reached the preset action completion condition. There is a temporal relationship between the above frames or nodes, but a fixed time interval is not required.
[0143] S104, Dynamically determine the valid action completion node. The completion node determination module 340 performs the valid action completion judgment after the current action has reached the start state, development state, and peak state.
[0144] When a hold state is formed after a peak state, the system enters the recovery state after the hold state; when no hold state that meets the continuity condition is formed after a peak state, the system can directly enter the recovery state from the peak state. In both cases, the video frame that first meets the recovery determination condition is taken as the first recovery state frame.
[0145] The number of consecutive recovery frames is accumulated starting from the first recovery status frame. In this embodiment, the number of recovery confirmation frames is set to 8 frames.
[0146] When the number of consecutive frames recovered reaches 8, and the decrease ratio of the current motion amplitude to the current motion peak amplitude reaches the preset recovery threshold, the current video frame is determined as a valid motion completion node.
[0147] The first recovery status frame is only used to determine the starting position of the recovery confirmation process, while the effective action completion node is used to determine that the recovery process has continued to meet the action completion requirements. Therefore, the two correspond to different video frame positions.
[0148] If a renewed increase in motion amplitude is detected before the 8-frame recovery confirmation is completed, and the conditions for further motion enhancement are met, then the accumulated number of consecutive recovered frames is cleared, and the enhanced video frames are added back to the current motion instance.
[0149] The enhanced action continues to identify the development state, peak state, maintenance state, and recovery state. When the recovery determination condition is met again, the first recovery state frame is re-determined, and a new number of consecutive recovery frames is accumulated starting from the re-determined first recovery state frame. Only after the recovery confirmation condition is met again is a valid action completion node determined.
[0150] Therefore, the brief pauses that occur when the interrogator raises their hand, displays an item, or makes an identification will not be mistakenly classified as two separate actions.
[0151] When a subject is briefly occluded, if the number of frames the occlusion lasts does not reach the allowable number of frames, the action instance identifier and action state before the occlusion are maintained, and the historical video frames corresponding to that action are preserved. After the subject is detected again, state recognition continues based on the subject identifier and the continuity of action features before and after the occlusion.
[0152] S105 maintains a circular frame buffer according to write, protect, and overwrite rules, such as Figure 4 As shown, the limited frame buffer module 310 writes video frames cyclically in a fixed direction and classifies the stored video frames into ordinary frames, replaceable protection frames, and locked protection frames according to the buffer protection status.
[0153] Normal frames are allowed to be overwritten preferentially in the order of circular writes; replaceable protection frames are only allowed to be replaced if they can still maintain the minimum number of representative frames during the corresponding action phase and meet the integrity replacement condition or the duplicate frame elimination condition; locked protection frames are prohibited from being overwritten until their locking condition is released.
[0154] like Figure 5 As shown, when a new video frame arrives, the target buffer location is first determined by the write pointer.
[0155] If the target cache location stores a normal frame, then the normal frame is overwritten, and a new video frame is written to the target cache location.
[0156] If the target cache location stores replaceable protection frames, then compare the completeness of evidence information and the repetition of action features between the new video frame and the video frames already preserved in the same action phase.
[0157] When a new video frame belongs to the same action phase and its evidence information completeness is higher than that of the current replaceable protection frame, the new video frame replaces the corresponding replaceable protection frame.
[0158] When the repetition of action features between a new video frame and a previously retained video frame in the same action phase reaches the repetition threshold, and the completeness of the evidence information in the new video frame is not improved, the current protection frame will not be replaced by the new video frame; after the new video frame completes the current real-time processing, no continuous protection flag will be set.
[0159] If the target cache location stores a lock protection frame, the lock protection frame will not be overwritten, and the write pointer will continue to move to the next cache location to re-execute the cache location status judgment.
[0160] When the write pointer needs to determine the overwrite object from multiple overwriteable buffer locations, the overwrite object is first determined from ordinary frames that are not protected and do not belong to the current action interval; if no such ordinary frames exist, video frames with low evidence information integrity and whose replacement will not cause the corresponding action stage to fall below the minimum retention number are selected from the replaceable protection frames; if multiple replaceable protection frames still exist that meet the conditions, video frames with high repetition of action features with the representative frames already retained in the same action stage are given priority as the overwrite object.
[0161] The current valid peak candidate frame has a higher cache protection priority than ordinary action-related video frames. When the peak state is updated, the peak candidate range and the current valid peak candidate frame are re-determined based on the updated peak state frame; the current unique valid peak candidate frame is not deleted according to the ordinary cyclic overwrite rule before a new valid peak candidate frame is formed or the current action ends.
[0162] Before an action is performed, at least one baseline representative frame must be retained as the stable baseline. Each action phase that has occurred must maintain at least the corresponding minimum number of representative frames to prevent the complete loss of identified action phases due to limited buffer cycle overwriting.
[0163] S106, after the compression trigger condition is met, a representative frame is retained according to the action stage, such as... Figure 6 As shown, the limited frame buffer module 310 continuously counts the number of protected video frames occupied corresponding to the current action instance and compares the number of protected video frames occupied with a preset compression trigger threshold.
[0164] The compression trigger threshold is less than or equal to the maximum protected capacity allowed by the circular frame buffer for the current action. By setting the compression trigger threshold to not exceed the maximum protected capacity, stage representative frame compression can be initiated before the protected video frames occupied by the current action exhaust the available protected capacity, thereby reserving buffer space for continuous writing of subsequent real-time video frames.
[0165] When the number of protected video frames corresponding to the current action has not reached the compression trigger threshold, the protection and replacement rules of S105 continue to maintain the video frames related to the action.
[0166] When the number of protected video frames reaches the compression trigger threshold, the limited frame buffer module 310 switches from continuous protection mode to stage representative frame protection mode, and no longer implements continuous protection for all action-related video frames generated after the current action.
[0167] After entering the stage-representing frame protection mode, the already protected video frames are divided into six groups according to the current action status: pre-action baseline, start-up, development, peak, hold, and recovery.
[0168] For the pre-action baseline group, video frames with high evidence information completeness are selected from video frames near the stable baseline as baseline representative frames.
[0169] For the startup phase group, video frames that can indicate the action has started to deviate from the stable state and have a high degree of completeness of evidence information are selected from the video frames that have been identified as startup frames as startup representative frames.
[0170] For the development stage group, video frames with high action representativeness and completeness of evidence information were selected from video frames whose action features continued to improve as development representative frames.
[0171] For the peak stage group, peak candidate frames are formed by the current peak state frame and video frames within the preset adjacent range that meet the minimum visibility requirements, and peak candidate frames with higher evidence information integrity are retained first; when the peak state frame is updated, the peak candidate frames are updated synchronously.
[0172] For the hold phase group, video frames with high evidence completeness are selected from those already identified as hold states as representative hold frames. If the current action does not result in a hold state, a hold phase group is not established, and it is not required to reserve representative frames for the hold phase.
[0173] For the recovery phase group, video frames with high integrity of evidence information are selected from the recovery state video frames starting from the first recovery state frame as representative recovery frames.
[0174] When there are multiple video frames to be protected in the same stage, video frames with higher evidence information completeness are retained first; for video frames whose action features are repeated to the same threshold as the retained representative frames and whose evidence information completeness has not improved, continuous protection markers are no longer set.
[0175] When a new video frame with higher completeness of evidence appears in the subsequent same action phase, the replaceable representative frame with lower completeness in that phase is replaced with the new video frame.
[0176] For each action phase that has occurred, at least the minimum number of video frames corresponding to that phase must be retained; in this embodiment, at least one representative frame must be retained for each action phase that has occurred. "Each action phase that has occurred" refers to the action phase that has actually been identified and formed in the current action. For phases such as the hold state that have not actually formed, it is not required to configure representative frames.
[0177] The currently valid peak candidate frames and representative frames that cannot be reduced further after reaching the minimum retention number for each stage are set as locked protection frames. The total number of representative frames retained across all action stages must not exceed the maximum protection capacity allowed for the current action. When the total number of representative frames approaches the maximum protection capacity, video frames with low evidence information completeness and high repetition of action features with other representative frames in the same stage are prioritized for elimination, provided that the minimum retention number for each stage is not lower than the minimum retention number for each stage.
[0178] Figure 6 This indicates that after the number of protected video frames reaches the compression trigger threshold, the protection mode switches from continuous protection to stage representative frame protection, and the representative frames are selected, updated, and protected according to each actual action stage. The candidate frame backtracking and target evidence frame determination after normal completion are executed according to S107 to S109 respectively, and stage representative frame protection itself is not equated with the final target evidence frame output.
[0179] S107, candidate frames are determined by backtracking from the limited historical cache. After determining the effective action completion node, the candidate frame backtracking and integrity evaluation module 350 takes the effective action completion node as the backtracking starting point, reads the protected video frames backward along the original video timestamp according to the current action instance identifier, until the protected baseline representative frame before the action start frame, thereby determining the historical frame range corresponding to the current action.
[0180] When the current action has not entered the stage of the representative frame protection mode, the peak state frame is taken as the center and the five frames before and after it are read to form the peak candidate range. The video frames that meet the minimum visibility requirements are determined as peak candidate frames. When the current action actually forms a hold state, the hold state frame and the peak candidate frame are combined to form a candidate frame set. When the hold state is not formed, the peak candidate frame forms a candidate frame set.
[0181] The current action has entered Figure 6 When the stage shown represents the frame protection mode, a candidate frame set is formed from the peak candidate frames still retained in the limited frame buffer, the representative frames that actually exist in the current action, and the protected representative frames that are adjacent to the peak candidate frames in time.
[0182] During the phase of frame protection, video frames that have been released according to the protection rules and covered by subsequent video frames are not rewritten into the current limited frame buffer, thus enabling the candidate frame backtracking process to be executed under the condition of limited buffer capacity.
[0183] S108, Evaluate the completeness of the evidence information in the candidate frames, such as Figure 7As shown, the candidate frame backtracking and integrity evaluation module 350 first performs a minimum visibility requirement judgment on the candidate frames. For limb actions that do not involve objects, at least the minimum visibility requirement of the subject and the minimum visibility requirement of the key parts of the action are judged. For actions such as picking up, handing over, displaying, pointing, or other actions involving objects, the minimum visibility requirement of the relevant objects is also judged. Candidate frames that do not meet any applicable minimum visibility requirement are removed from the candidate frame set. For the remaining candidate frames, the integrity of the subject, the visibility of the key parts of the action, the visibility of the relevant objects, the representativeness of the action state, the image clarity, the degree of occlusion, and the consistency between consecutive frames are determined, and each evaluation quantity is normalized.
[0184] The higher the integrity of the main body, the higher the visibility of key parts of the action, the higher the visibility of related objects, the higher the representativeness of the action state, the higher the image clarity, and the higher the consistency between consecutive frames, the higher the corresponding evaluation value; the larger the occlusion area, the lower the corresponding effective evaluation value.
[0185] For actions involving related items, the completeness of the evidence information is determined based on preset weights corresponding to the integrity of the subject, the visibility of key parts of the action, the visibility of related items, the representativeness of the action state, the image clarity, the degree of occlusion, and the consistency between consecutive frames. For actions not involving related items, the visibility of related items is removed from the current evaluation items, and the completeness of the evidence information is determined after the weights of the remaining evaluation items are renormalized.
[0186] The preset weights are used to represent the relative influence of each evaluation item on the integrity of the candidate frame evidence information, and the sum of the weights of each valid evaluation item participating in the current candidate frame evaluation is normalized to 1.
[0187] Based on the above evaluation results, candidate frames are arranged from high to low according to the completeness of evidence information.
[0188] For example, when the interrogator picks up and displays an item, the item is raised to its maximum height in a certain peak state frame, but the video frame has obvious motion blur. Although the movement amplitude of another peak candidate frame or the hold state frame formed by the current action is slightly lower, it can completely display the interrogator, the hand performing the action, and the displayed item. Moreover, the image is clear and the degree of obstruction is low. Therefore, the evidence information completeness of the peak candidate frame or the hold state frame can be higher than that of the peak state frame.
[0189] Therefore, the peak state frame is only used to determine the peak position of the action and is not directly used as the final target evidence frame.
[0190] S109, the target evidence frame is determined and saved. The candidate frame backtracking and integrity evaluation module 350 determines the candidate frame with the highest integrity of evidence information as the target evidence frame.
[0191] When two or more candidate frames have the same level of evidence information completeness, the candidate frame with the smaller time distance from the current peak state frame is selected first; if the time distance is still the same, the candidate frame with the earlier original video timestamp is selected.
[0192] Action start frame is used to determine the start position of action interval, peak state frame is used to determine the local maximum position of action amplitude, peak candidate frame is used to determine the key evaluation range near the peak, first recovery state frame is used to start recovery continuous frame counting, effective action completion node is used to determine the end position of normal action interval, and target evidence frame is used to represent the video frame finally selected for saving after the evidence information integrity evaluation. The above types of frames or nodes do not replace each other.
[0193] The identified target evidence frame is associated with the subject identifier, action instance identifier, action start time, peak time, first recovery state frame time, effective action completion time, action interval, video channel identifier, original video timestamp, and evidence information integrity, and then sent to the evidence frame storage device 400.
[0194] After the evidence frame storage device 400 completes the saving of the target evidence frame and the corresponding positioning information, the limited frame buffer module 310 removes the buffer protection mark of the video frame corresponding to the current action instance.
[0195] The cached records that have been deprotected are re-enacted in the cyclic overlay of subsequent video frames; the target evidence frames that have been transferred to the evidence frame storage device 400 and their original video positioning information are not affected by the subsequent overlay operation of the limited frame cache.
[0196] The interrogation review terminal 500 locates the corresponding original video file from the original video storage device 120 based on the video channel identifier associated with the target evidence frame and the original video timestamp, and determines the original video segment containing the current action based on the action start time and the effective action completion time, thereby realizing the corresponding review between the target evidence frame, the action interval and the continuous original interrogation video.
[0197] For abnormally terminated actions, the evidence frame storage device 400 associates and saves the abnormal termination marker, the time corresponding to the last video frame with a definite action state before the abnormal termination, and the review candidate frame information; the interrogation review terminal 500 locates the corresponding original video segment based on the above time information, and the reviewer judges the actual state of the abnormally terminated action by combining the continuous original video, and does not record the abnormal termination position as a valid action completion node.
[0198] Once an action is successfully completed and the corresponding cache protection is released, the state evolution recognition module 330 allows the updating of the stable baseline again and continues to process subsequent video frames. When the same interrogation subject meets the action initiation conditions again, a new action instance identifier is generated for the new action, and new action state evolution recognition, cache protection, completion node determination, historical frame backtracking, and target evidence frame selection are performed according to S103 to S109.
[0199] Through the above processing, under the condition of maintaining only a limited capacity of historical frame buffer in real-time video processing, the action range can be determined according to the continuous state evolution relationship of the interrogation subject's actions from initiation, development, peak, maintenance to recovery, and after the action is completed, the target evidence frame with high evidence information completeness can be selected by backtracking from the limited historical video frames.
[0200] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A real-time frame extraction method for video streams used in police interrogations, characterized in that, Includes the following steps: Acquire real-time video stream of the interrogation scene, convert the real-time video stream into video frames with original video timestamps, and write them into a circular frame buffer of limited capacity in chronological order; The action features of the same interrogator are extracted from consecutive video frames to form an action state observation sequence. Based on the direction, magnitude and duration of the change of the action features, the initiation state, development state, peak state, maintenance state and recovery state of the action are identified online. Once the startup state is detected, a cache protection flag is set for the video frames related to the current action, and the protected video frames are updated as the action state evolves. After the current action has formed the initiation state, development state and peak state, when the peak state is followed by the maintenance state or directly enters the recovery state, the effective action completion node is determined according to the degree of recovery and whether the same action is enhanced again. When the number of protected video frames corresponding to the current action reaches the preset compression trigger threshold, the representative frame of the action stage is retained according to the action stage. Using the effective action completion node as the backtracking starting point, the candidate frame corresponding to the current action is determined from the circular frame buffer. The target evidence frame is determined based on the completeness of the evidence information of the candidate frame, and the target evidence frame, action interval, and original video timestamp are associated and saved.
2. The real-time frame extraction method for video streams used in police interrogation according to claim 1, characterized in that, Extracting action features of the same interrogator from consecutive video frames, including: Perform subject detection and subject tracking on video frames, and configure the same subject identifier for the same interrogation subject in consecutive video frames; Extract the key point location, key point displacement direction, key point displacement amount, movement amplitude, movement speed, and torso posture of the interrogating subject; Detect items related to the current action, determine the location of the items, the relative position of the items to the key parts of the action, and the visibility status of the items; When the amplitude of the motion is continuously lower than the baseline update threshold, a stable baseline is established or updated based on the corresponding video frame, and the update of the stable baseline is stopped after the start state is detected.
3. The real-time frame extraction method for video streams used in police interrogation according to claim 2, characterized in that, The online system identifies the initiation, development, peak, maintenance, and recovery states of an action, including: When the amplitude of the action continuously reaches the number of activation confirmation frames and exceeds the activation threshold, the first video frame that exceeds the activation threshold is determined as the action activation frame. When the motion characteristics continue to increase relative to the stable baseline but do not reach the peak determination criteria, it is determined to be in a developmental state; When the amplitude of the motion reaches the local maximum value within the current motion range and no video frame with a larger amplitude of motion appears within the subsequent peak confirmation frame number, the video frame corresponding to the local maximum value is determined as the peak state frame. When the change in the action feature relative to the peak state is continuously within the peak stability threshold range, it is determined to be a sustained state; After the peak state or the hold state, when the motion amplitude continues to decrease and the key point moves towards the stable baseline position and meets the recovery determination condition, the video frame that first meets the recovery determination condition is determined as the first recovery state frame and enters the recovery state.
4. The real-time frame extraction method for video streams used in police interrogation according to claim 3, characterized in that, Determining the effective action completion node includes: After the current action has reached the start state, development state, and peak state, when it enters the hold state or recovery state after the peak state, the number of consecutive recovery frames is accumulated starting from the first recovery state frame; When the number of consecutive recovery frames reaches the number of recovery confirmation frames, and the decrease ratio of the current action amplitude relative to the peak action amplitude reaches the recovery threshold, the current video frame is determined as a valid action completion node. If the motion amplitude is detected to increase again and the motion enhancement condition is met before the recovery confirmation frame count is reached, the accumulated recovery consecutive frame count is cleared, and the increased motion is returned to the current motion range. When the recovery determination condition is met again, the first recovery state frame is re-determined, and the number of consecutive recovery frames is re-accumulated.
5. A real-time frame extraction method for video streams used in police interrogations according to claim 1, characterized in that, Set buffer protection flags for video frames related to the current action, including: The video frames in the circular frame buffer are divided into normal frames, replaceable protection frames, and locked protection frames according to the buffer protection status. The ordinary frames are allowed to be overwritten preferentially in the cyclic writing order. The replaceable protection frames are allowed to be replaced according to the integrity of the evidence information and the repetition of the action features, provided that the minimum number of representative frames for the corresponding action stage is maintained. The locked protection frames are prohibited from being overwritten in the ordinary cyclic writing method before the corresponding locking condition is released. When the startup state is detected, a preset number of video frames before the action startup frame, the action startup frame, and the action-related video frames written after the action starts are set as protected video frames.
6. A real-time frame extraction method for video streams used in police interrogation according to claim 5, characterized in that, When a new video frame is written to the circular frame buffer: If a normal frame is stored at the cache location pointed to by the write pointer, then the normal frame is overwritten with the new video frame; If the cache location stores a replaceable protection frame, the completeness of the evidence information and the repetition of the action features of the new video frame are compared with those of the video frames already retained in the same action phase. Replacement is performed when the completeness of the evidence information of the new video frame is high. If the repetition of the action features reaches the repetition threshold and the completeness of the evidence information is not improved, no continuous protection mark is set for the new video frame. If the cache location stores a lock protection frame, then overwriting the cache location is prohibited, and the write pointer is deferred to the next cache location.
7. A real-time frame extraction method for video streams used in police interrogations according to claim 5, characterized in that, When the number of protected video frames corresponding to the current action reaches the preset compression trigger threshold, the protection mode is switched from continuous protection mode to stage representative frame protection mode. The compression trigger threshold is less than or equal to the maximum protection capacity that the current action is allowed to occupy in the circular frame buffer; According to the action state, the protected video frames are divided into the pre-action baseline, start, development, peak, hold and recovery stages, and the baseline representative frame, start representative frame, development representative frame, peak candidate frame, hold representative frame and recovery representative frame are retained respectively. When there are multiple video frames in the same action phase, the video frames with higher evidence information completeness are retained first, and the continuous protection of the corresponding video frames is lifted when the action feature repetition reaches the repetition threshold and the evidence information completeness is not improved. Among them, at least a preset minimum number of representative frames are retained for each action phase that has occurred. The peak candidate frame consists of the peak state frame and video frames within a preset adjacent range that meet the minimum visibility requirement, and is updated accordingly when the peak state is updated.
8. A real-time frame extraction method for video streams used in police interrogations according to claim 7, characterized in that, Determining the completeness of the evidence information of the candidate frames includes: Read peak candidate frames and hold-state frames from the protected video frames corresponding to the current action to form a candidate frame set; When the current action has entered the stage representative frame protection mode, a candidate frame set is formed from the peak candidate frame, the hold representative frame, and the protected representative frame that is adjacent to the peak candidate frame in time; Candidate frames that do not meet the minimum visibility requirements for the main body, the minimum visibility requirements for key parts of the action, or the minimum visibility requirements for related items involved in the current action are removed. The remaining candidate frames are normalized and weighted based on the integrity of the main body, the visibility of key parts of the action, the visibility of related items, the representativeness of the action state, the image clarity, the degree of occlusion, and the consistency between consecutive frames to obtain the completeness of the evidence information. The candidate frame with the highest completeness of the evidence information is then determined as the target evidence frame. If the current action does not involve related items, the visibility of related items will not be used as a valid evaluation item, and the weights of the remaining valid evaluation items will be renormalized.
9. A real-time frame extraction system for video streams used in police interrogation, for executing the real-time frame extraction method for video streams as described in any one of claims 1-8, characterized in that, include: The video stream access module is used to receive real-time video streams from the interrogation site and generate video frames with the original video timestamps. The finite frame buffer module is used to cyclically write to a circular frame buffer of limited capacity in chronological order, and to protect, replace, and retain the stage representative frame for the current action-related video frame. The action feature extraction module is used to extract action features of the same interrogating subject; The state evolution recognition module is used to identify the start-up state, development state, peak state, maintenance state, and recovery state based on the direction, amplitude, and duration of the change in action characteristics. The completion node determination module is used to accumulate the number of consecutive recovery frames starting from the first recovery state frame, and determine the valid action completion node based on the recovery duration and whether the same action is enhanced again; The candidate frame backtracking and completeness evaluation module is used to determine candidate frames by taking the effective action completion node as the backtracking starting point, and to determine the target evidence frame based on the completeness of the evidence information of the candidate frames. An evidence frame storage device is used to associate and save the target evidence frame, action range, and original video timestamp.
10. A real-time frame extraction system for video streams for public security interrogation according to claim 9, characterized in that, It also includes interrogation camera equipment, original video storage equipment, and interrogation review terminals; The interrogation camera device is used to send the same real-time video stream to the original video storage device and the video stream access module respectively; The original video storage device is used to continuously store the original video stream and its corresponding video file identifier, video channel identifier, and original video timestamp. The interrogation review terminal is used to locate the corresponding original video file from the original video storage device based on the video channel identifier associated with the target evidence frame and the original video timestamp, and to obtain the original video segment containing the action interval.