Computer vision-based integrated vehicle and gun close-range burst action warning method

By unifying the processing of multi-view video footage and improving the PoseC3D model, the problem of unifying multi-view footage in the scene of guarding and escorting vehicles has been solved, enabling accurate identification of firearms and door areas and early warning of continuous movements, thereby improving the security of guarding and escorting.

CN122493518APending Publication Date: 2026-07-31ZHEJIANG ANBANG SECURITY TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG ANBANG SECURITY TECH SERVICE CO LTD
Filing Date
2026-04-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies lack unified mapping processing for multi-view images in video surveillance of the handover side, door side, and guard side of guarded vehicles. They are difficult to identify dedicated areas such as the firearm control zone, door buffer zone, and handover allowance zone, resulting in insufficient ability to identify fine-grained risky actions such as approaching, reaching out, deflecting, and disengaging.

Method used

By acquiring multi-view video footage, generating close-up monitoring sequence, separating human and vehicle structural regions, constructing object pairing and cross-frame thermal models, and using an improved PoseC3D model for motion analysis, we can identify continuous actions related to seizure, containment, and loss of control.

Benefits of technology

It enables accurate identification and timely warning of sudden actions of guard vehicles at close range, improving the accuracy of identification and the practicality of warning, and is suitable for integrated guard vehicle and gun security protection scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493518A_ABST
    Figure CN122493518A_ABST
Patent Text Reader

Abstract

This invention discloses a computer vision-based method for early warning of sudden close-range actions by a vehicle-mounted gun integrated with a guard unit, comprising the following steps: Step 1: Acquiring continuous video footage, calculating planar mapping relationships, and generating a sequence of close-range monitoring footage; Step 2: Separating the human body region and vehicle structure region frame by frame, and determining the vehicle-mounted gun action region; Step 3: Numbering and pairing the vehicle-mounted gun action regions, and generating relationship connection results in the current frame; Step 4: Performing cross-frame matching to generate joint point heatmaps and limb heatmaps, and forming a sequence of action analysis windows; Step 5: Obtaining segment labeling results through an improved PoseC3D model; Step 6: Forming continuous early warning segments based on the segment labeling results, determining the early warning type of each continuous early warning segment, and generating early warning information. This invention achieves early warning of sudden close-range actions by integrating the vehicle-mounted gun-human relationship and posture timing through an improved PoseC3D model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video image analysis technology, and in particular to a computer vision-based method for early warning of sudden close-range actions of a vehicle-mounted gun integrated with a guard unit. Background Technology

[0002] During escort operations, the handover side, door side, and guard side of the escort vehicle are typically subject to frequent close-range personnel activity, significant vehicle structural obstruction, and sudden changes in movement. Especially when doors are open, handover procedures are being conducted, escort personnel are adjusting their positions, or external personnel are approaching, there is a high risk of unauthorized close proximity, intrusion, obstruction, and weapons losing stable control. To enhance security at escort sites, existing technologies are beginning to employ video surveillance, target detection, human recognition, and behavior analysis to automatically monitor activities around escort vehicles and attempt to trigger alarms through image recognition results.

[0003] However, most existing technologies focus on identifying human targets, vehicle targets, or simple action categories in single-channel video. They typically prioritize monitoring the appearance of personnel, their approach, area intrusion, or general abnormal behavior, lacking unified mapping processing of multi-view images from the handover side, door side, and guard side of the guarded vehicle. This makes it difficult to stably analyze the spatial relationships between guards, external approaching personnel, firearms, holsters, and vehicle door structures within the same scene plane. Furthermore, existing technologies generally lack dedicated area constraint processing tailored to guarded operations scenarios, such as firearm control zones, door buffer zones, and handover allowance zones, resulting in insufficient ability to identify fine-grained risky actions such as approaching, reaching out, deflecting, and disengaging from holsters.

[0004] Therefore, how to provide a computer vision-based early warning method for close-range sudden actions of integrated guard and escort vehicles with firearms is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a computer vision-based early warning method for close-range sudden actions of integrated guard and escort vehicle-mounted guns. This invention achieves effective identification and early warning of continuous action segments related to seizure, encirclement, and loss of control by unifying multi-view close-range monitoring images and combining vehicle-mounted gun action area separation, object pairing, cross-frame thermal body construction, and improved PoseC3D model processing. This improves the identification accuracy, temporal continuity, and early warning practicality of close-range sudden actions of integrated guard and escort vehicle-mounted guns.

[0006] The computer vision-based early warning method for close-range sudden actions of a vehicle-mounted gun integrated with a guard unit, according to an embodiment of the present invention, includes the following steps: Step 1: Collect continuous video footage from the handover side, door side, and warning side of the guarded vehicle, calculate the planar mapping relationship, and generate a sequence of close-range monitoring footage; Step 2: Separate the human body area and vehicle structure area frame by frame in the close-up monitoring image sequence, and determine the vehicle gun action area; Step 3: Number and pair objects in the vehicle-gun action area, and generate the relationship connection results in the current frame; Step 4: Read the object numbers and relationship connection results in adjacent frames in chronological order, perform cross-frame matching to generate joint heatmaps and limb heatmaps, and form a motion analysis window sequence; Step 5: Input the motion analysis window sequence into the improved PoseC3D model, and obtain the fragment labeling results through the thermal body segmentation module, branch extraction module, temporal splicing module, and fragment labeling module; Step 6: Based on the segment marking results, form continuous warning segments, determine the warning type of each continuous warning segment, and generate warning information.

[0007] Optionally, step one specifically includes: Collect continuous video footage from the handover side, door side, and guard side of the guarded vehicles, and write time stamps to each frame of each continuous video footage. Extract the pixel positions of the door hinge edge, door lock handle, junction window boundary, and pedal outer edge from each continuous video frame; Calculate the planar mapping relationship between consecutive video frames based on the corresponding pixel positions of the same fixed reference in different video frames; Transform each continuous video frame onto a unified scene plane according to the planar mapping relationship; By cropping the outer area of ​​the vehicle door, the outer area of ​​the handover window, and the area where the escort is stationed within a unified scene plane, a sequence of close-up monitoring images is obtained.

[0008] Optionally, step two specifically involves: In the close-up monitoring image sequence, the human body area and the vehicle structure area are separated frame by frame, and the human body area is further distinguished from the main body area of ​​the escort personnel and the area of ​​external approaching personnel. The firearms area was extracted from the upper limbs of the escort personnel's main body area, and the holster area was extracted from the waist side of the escort personnel's main body area. A local sub-region is cropped at the end of the forearm near the personnel area, and the external hand region is extracted; A gun-holding baseline is formed by connecting the center point of the guard's torso with the point of the wrist holding the gun, and the gun-holding baseline is widened at a fixed width on both sides along the main axis of the firearm to generate a firearm control belt; A door buffer zone is formed by the door hinge edge and the outer edge of the pedal in the vehicle structure area. The width is expanded outward according to the boundary of the handover window to generate the handover allowance zone. The main area for escort personnel, the area for external approaching personnel, the area for firearms, the area for holsters, the area for external hands, the firearm control strap, the buffer zone of the vehicle door, and the handover allowance zone are designated as the vehicle-gun movement area.

[0009] Optionally, step three specifically includes: Assign object numbers to the escort body area, external approaching personnel area, firearm area, holster area, external hand area, firearm control strip, vehicle door buffer strip, and handover allowance strip in each frame of the vehicle-gun action area. The first type of pairing consists of the firearms area and the main body area of ​​the escort; the second type of pairing consists of the firearms area and the holster area; the third type of pairing consists of the main body area of ​​the escort and the door structure area; and the fourth type of pairing consists of the external approach personnel area and the firearms area. The fifth type of pairing consists of the external hand area and the firearm control strap; the sixth type of pairing consists of the external hand area and the door buffer strip; the seventh type of pairing consists of the external approach personnel area and the main body area of ​​the escort; and the eighth type of pairing consists of the external approach personnel area and the door buffer strip. For each type of pairing, calculate the position of the center points at both ends, the distance between the center points, the direction angle of the object connection line, the overlapping area, the boundary contact length, and the relative position order, and record the calculation results as the relationship connection results in the current frame according to the pairing category.

[0010] Optionally, step four specifically involves: The object numbers and relationship connection results in two adjacent frames are read in chronological order. Cross-frame matching is then performed on the main body area of ​​the escort, the area of ​​approaching external personnel, the firearm area, the holster area, and the external hand area. The cross-frame matching specifically includes: Extract the coordinates of key points on the shoulders, elbows, wrists, hips, knees, and ankles of the escort and the personnel approaching from outside; Using the center point of the escort officer's torso within a unified scene plane as the local reference origin, the coordinates of each key point are translated and normalized. Write the coordinates of each key point in a series of consecutive frames into the key point sequence in chronological order. A joint heatmap is generated within the corresponding frame, centered on the coordinates of each key point. According to the order of limb connection, generate limb heat maps between the shoulder and elbow, between the elbow and wrist, between the hip and knee, and between the knee and ankle; The joint heatmaps and limb heatmaps from several consecutive frames are stacked in chronological order to form a three-dimensional posture heatmap. The sequence of distances between external personnel approaching the firearm, the sequence of entry lengths of external hands into the firearm control strap, the sequence of entry lengths of external hands into the vehicle door buffer strap, the sequence of angles between the main axis of the firearm and the central axis of the escort's torso, and the sequence of distances from the firearm area to the holster area within a series of consecutive frames are bound to the three-dimensional attitude thermosphere according to the time window number to form a motion analysis window sequence.

[0011] Optionally, the improved PoseC3D model is specifically as follows: The motion analysis window sequence is input into the thermal body segmentation module. The 3D attitude key point thermal bodies in the motion analysis window sequence are read sequentially according to the time window. Part segmentation processing is then performed on the 3D attitude key point thermal bodies. The part segmentation processing specifically includes: Extract the joint point heat maps corresponding to the left and right shoulders, left and right elbows, and left and right wrists, as well as the limb heat maps corresponding to the left and right shoulders and left and right elbows, and the left and right elbows and left and right wrists. Combine the joint point heat maps and limb heat maps in the original time sequence to form the upper limb gun-holding heat body. Extract the joint point heat maps corresponding to the left and right shoulders and the left and right hips, as well as the connection area heat map between the midpoint of the left and right shoulders and the midpoint of the left and right hips, and combine the joint point heat map and the connection area heat map in the original time sequence to form a stable thermodynamic body of the torso; Extract the joint point heat maps corresponding to the left and right hips, left and right knees, and left and right ankles, as well as the limb heat maps corresponding to the left and right hips and left and right knees, and the left and right knees and left and right ankles. Combine the joint point heat maps and limb heat maps in the original time sequence to form a lower limb displacement heat body. The upper limb holding gun thermal body, the trunk stabilization thermal body, and the lower limb displacement thermal body are input into the branch extraction module. Three-dimensional convolution extraction processing is performed on the three thermal bodies respectively. The three-dimensional convolution extraction processing is to perform local convolution scanning and inter-layer downsampling processing in the time dimension, the horizontal dimension, and the vertical dimension respectively, and output the gun holding action body features, the trunk support body features, and the standing position movement body features. The features of the gun-holding action, the features of the torso support, and the features of the standing and moving body are input into the timing splicing module according to the same time window number. The features of the gun-holding action, the features of the torso support, and the features of the standing and moving body are read according to the time window number and spliced ​​together in chronological order to form the posture timing features. Read the distance sequence from the external approaching personnel to the firearm, the entry length sequence of the external hand into the firearm control strap, the entry length sequence of the external hand into the vehicle door buffer strap, the angle sequence between the main axis of the firearm and the central axis of the escort's torso, and the distance sequence from the firearm area to the holster area, which are the same as the time window number. Perform length alignment on each relation sequence according to the time frame position until the start frame and end frame of each relation sequence are consistent with the start frame and end frame of the attitude timing feature; The length-aligned relational sequences are sequentially concatenated to the post-pose temporal features in a preset order to form a joint action sequence; The joint action sequence is input into the fragment labeling module, and window-by-window scanning and fragment classification processing are performed on the joint action sequence according to the time window order to obtain the fragment labeling result.

[0012] Optionally, the step of performing window-by-window scanning and segment classification processing on the joint action sequence according to the time window order to obtain segment labeling results is as follows: Within each time window, read the upper limb movement change segment, trunk support change segment, and lower limb displacement change segment in the combined action sequence; Synchronously compare the segments of upper limb movement changes with the segments of angle changes between the main axis of the firearm and the central axis of the escort's torso, and mark the time segments in which the upper limb rotation direction changes continuously and the angle between the firearm and the escort's torso increases continuously and is greater than a preset threshold as the gun-holding deflection segment. The shoulder-to-elbow and elbow-to-wrist transition segments in the upper limb movement change segments are compared synchronously. The time segments in which the shoulder-to-elbow and elbow-to-wrist transitions continuously increase and exceed a preset threshold within the same time window are marked as upper limb abrupt transition segments. Synchronously compare the torso support change segment with the displacement change segment in the stationary movement characteristics, and combine the distance sequence change from the approaching personnel to the firearm to mark the time period when the torso center shift and the lower limbs move backward consecutively as the forced retreat segment. The entry length sequence of external hands entering the gun control area and the entry length sequence of external hands entering the door buffer area are checked window by window, and the time period when the entry length continuously increases is marked as the hand intrusion segment. Synchronously compare the changes in the physical characteristics of the gun-holding action with the changes in the distance sequence from the gun area to the holster area. Mark the time period in which the changes in the gun-holding action occur continuously and the distance between the gun area and the holster area increases continuously as the gun unholstering segment. The segments of sudden upper limb turning, gun deflection, forced retreat, hand intrusion, and gun disarming were used as segment labeling results.

[0013] Optionally, step six specifically includes: Read the upper limb sudden turn segment, gun deflection segment, forced retreat segment, hand intrusion segment, and gun disarming segment from the segment marking results in chronological order, and record the start frame, end frame, and object number of each segment. Segments that overlap in time or are consecutive in sequence are merged to form continuous warning segments, and the warning type of each continuous warning segment is determined, specifically including: When a continuous warning segment includes segments of sudden upper limb turning and gun deflection, the corresponding continuous warning segment is identified as a robbery warning type. When a continuous warning segment contains both a probe intrusion segment and a forced retreat segment, the corresponding continuous warning segment is identified as a containment warning type. When a continuous warning segment includes both a weapon deflection segment and a weapon dislodgement segment, the corresponding continuous warning segment is identified as a disengagement warning type. Extract the distance sequence from external approaching personnel to firearms, the entry length sequence of external hands into the firearm control strap, the entry length sequence of external hands into the vehicle door buffer strip, the distance sequence from the firearm area to the holster area, and the change sequence of the escort personnel's position within the time range corresponding to each continuous warning type, and calculate the change amount of each sequence within the time range. The change amount corresponding to each continuous warning type is compared with a preset threshold. When the change amount reaches the preset threshold, the corresponding warning level is generated. Write the warning type, warning level, start frame, end frame, and object number into the warning record to generate warning information.

[0014] The beneficial effects of this invention are: This invention addresses the challenges of unifying multi-view images in close-range scenarios involving guarded vehicles, accurately representing the relationship between personnel, vehicles, and firearms, and the lag in early warning of sudden actions. It achieves continuous and targeted technical effects. By time-stamping, extracting fixed references, and mapping continuous video footage from the handover side, door side, and guard side of the guarded vehicle, the hinged edges of doors, door handles, handover window boundaries, and pedal edges from different perspectives can be unified into a single scene plane for analysis. This improves the spatial consistency of close-range monitoring images and avoids recognition bias caused by unstable target positional relationships from a single perspective. Furthermore, this invention performs frame-by-frame separation of the human body area, vehicle structure area, guard body area, external approaching personnel area, firearm area, holster area, and external hand area, constructing firearm control zones, door buffer zones, and handover allowance zones. This allows key objects and their constrained areas in the guarded scenario to be described synchronously, providing clear spatial references for risky actions such as external approach, hand intrusion, firearm deflection, and firearm dismounting. By recording object numbers, classification matching, and relationship connection results, this invention transforms the originally scattered personnel actions, firearm status, and vehicle structural relationships into continuously trackable relationship information. Combined with cross-frame matching, joint heatmaps, limb heatmaps, and three-dimensional posture heatmap construction, it achieves continuous representation of the guard's posture changes, the process of external personnel approaching, and the evolution of the vehicle-gun relationship.

[0015] Further utilizing the improved PoseC3D model, branch extraction, temporal splicing, and segment labeling are performed on the upper limb gun-holding thermosphere, torso stability thermosphere, and lower limb displacement thermosphere. This enables accurate differentiation of upper limb abrupt turn segments, gun-holding deflection segments, forced retreat segments, hand intrusion segments, and weapon disarming segments. By merging continuous warning segments and determining the warning type, risks related to robbery, encirclement, and loss of control are promptly converted into directly output warning information. Therefore, this invention not only enhances the ability to identify complex target relationships and continuous movement changes in close-range guarding scenarios but also improves the pertinence, continuity, and practicality of the warning results, making it more suitable for application in integrated vehicle-mounted gun security scenarios. Attached Figure Description

[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of the computer vision-based integrated vehicle-mounted gun close-range sudden action early warning method proposed in this invention. Figure 2 This is a schematic diagram illustrating the action analysis window sequence generation steps of the computer vision-based vehicle-mounted gun close-range sudden action early warning method proposed in this invention. Figure 3 This is a flowchart of the improved PoseC3D model processing procedure for the computer vision-based integrated vehicle-mounted gun close-range sudden action early warning method proposed in this invention. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0018] refer to Figures 1-3 A computer vision-based early warning method for close-range sudden actions of a vehicle-mounted gun integrated with guards includes the following steps: Step 1: Collect continuous video footage from the handover side, door side, and warning side of the guarded vehicle, calculate the planar mapping relationship, and generate a sequence of close-range monitoring footage; Step 2: Separate the human body area and vehicle structure area frame by frame in the close-up monitoring image sequence, and determine the vehicle gun action area; Step 3: Number and pair objects in the vehicle-gun action area, and generate the relationship connection results in the current frame; Step 4: Read the object numbers and relationship connection results in adjacent frames in chronological order, perform cross-frame matching to generate joint heatmaps and limb heatmaps, and form a motion analysis window sequence; Step 5: Input the motion analysis window sequence into the improved PoseC3D model, and obtain the fragment labeling results through the thermal body segmentation module, branch extraction module, temporal splicing module, and fragment labeling module; Step 6: Based on the segment marking results, form continuous warning segments, determine the warning type of each continuous warning segment, and generate warning information.

[0019] In this embodiment, step one specifically includes: Fixed video capture devices are installed on the handover side, door side, and guard side of the guarded vehicle, respectively, so that the shooting direction of each video capture device covers the door hinge edge, door lock handle, handover window, and the area where the pedal is located; after each video capture device is started, the video frames output by each video capture device are continuously read, and the corresponding time stamp is written for each video frame according to the order of capture. In each continuous video frame, edge extraction is first performed on the current video frame. Then, the polygonal boundary corresponding to the door rotation connection position is selected from the edge extraction results to determine the pixel position of the door hinge edge. The outer contour boundary of the door lock handle is extracted in the area near the door hinge edge to determine the pixel position of the door lock handle. The closed boundary of the junction window is extracted in the area with opening features on the side of the vehicle body to determine the pixel position of the junction window boundary. The outer edge of the pedal is extracted in the horizontal load-bearing boundary below the door near the bottom of the vehicle body to determine the pixel position of the pedal outer edge. In each continuous video frame, the positions of the door hinge edge, door lock handle, junction window boundary, and pedal outer edge extracted in consecutive frames are aligned; boundary points whose position jump variables exceed the preset range in consecutive frames are deleted; the retained boundary points are fitted with lines and closed according to fixed reference categories to obtain the set of door hinge edge positions, door lock handle positions, junction window boundary positions, and pedal outer edge positions in the current video frame; Select synchronized video frames with the same time stamp from different continuous video frames, and read the sets of door hinge edge positions, door lock handle positions, junction window boundary positions, and pedal outer edge positions from each synchronized video frame; establish cross-viewpoint corresponding point pairs according to fixed reference categories, wherein the endpoints and turning points of the door hinge edges are taken as the first set of corresponding points, the center point and boundary corner points of the outer contour of the door lock handle are taken as the second set of corresponding points, the corner points of the junction window boundary are taken as the third set of corresponding points, and the endpoints of the pedal outer edge are taken as the fourth set of corresponding points. Using each set of corresponding point pairs, calculate the first plane mapping relationship from the handover side video image to the unified scene plane, the second plane mapping relationship from the door side video image to the unified scene plane, and the third plane mapping relationship from the warning side video image to the unified scene plane; then perform coordinate transformation on each video frame in each continuous video image according to the corresponding plane mapping relationship, so that the same fixed reference falls into the same position area in the unified scene plane in the transformed different video images. In a unified scene plane, the outer area of ​​the car door is defined by taking the transformed door hinge edge as one boundary and the outer edge of the pedal as the other boundary. The outer region of the intersection window is formed by extending a preset distance outward from the transformed intersection window boundary; The common adjacent area between the door hinge edge, the junction window boundary, and the outer edge of the pedal is designated as the guard's standing area. Based on the outer area of ​​the vehicle door, the outer area of ​​the handover window, and the area where the escort personnel are stationed, the video frames that have completed coordinate transformation are captured to obtain a sequence of close-range monitoring images.

[0020] In this embodiment, step two specifically includes: Read each frame of the close-range monitoring image sequence in chronological order. In the current frame, extract continuous closed boundaries and areas with continuous grayscale changes. Then, include the areas that are connected to or adjacent to the door hinge edge, door lock handle, junction window boundary, and pedal outer edge into the vehicle structure candidate area. Regions located outside the vehicle door, outside the handover window, and within the guard's standing area, whose boundary shapes change with the frame, are classified as human body candidate regions. The overlapping parts between the vehicle structure candidate region and the human body candidate region are split according to their adjacency relationship with the fixed reference to obtain the human body region and the vehicle structure region in the current frame. In the human body region of the current frame, first read the center position, outer contour height, outer contour width, and relative position with the guard's standing area of ​​each human body region; The main area that is continuously located within the guard's station area and is adjacent to the outer area of ​​the vehicle door or the outer area of ​​the handover window is defined as the guard's main area. The remaining human body areas located outside the main area of ​​the escort personnel, connected to the outer area of ​​the vehicle door or the outer area of ​​the handover window, or gradually approaching the main area of ​​the escort personnel are defined as the external approach personnel area. When there are multiple external approaching personnel areas in the same frame, the independent boundaries of each external approaching personnel area are preserved. Within the main body area of ​​the escort, first extract the area from the shoulder to the wrist as the upper limb search range, and then extract a narrow closed contour with a length greater than its width within the upper limb search range; For each narrow closed contour, calculate the contour length, contour width, main axis direction, and contour center position; retain the narrow closed contour that is adjacent to the guard's wrist and whose main axis direction is close to the forearm direction, and determine it as the firearm area; The waist-side search range is cut off on both sides of the waist of the escort's main body area, and the local protruding contour that fits the outer contour of the human body is extracted within the waist-side search range. For each local protruding contour, calculate the contour area, the length of the edge, and the center position of the contour. Retain the local protruding contours that have reached the preset length of the edge and whose center is located within the waist area. The retained local protruding contours are determined as the holster area. In each external area close to people, first determine the forearm extension direction according to the connection direction from shoulder to elbow and from elbow to wrist, and then cut out a local sub-region at the end of the forearm along the forearm extension direction. Extract local contours within a local sub-region where the number of undulations at the end boundary is higher than the number of undulations at the forearm boundary, and perform boundary closure processing on the local contours. Calculate the area, perimeter, center position, and distance from the end of the forearm for each closed local contour. The local contour with the smallest distance from the end of the forearm is retained, and the retained local contour is defined as the external hand region. When there are two forearm end positions in the same external approaching person region, the two local sub-regions are extracted respectively, and the corresponding external hand regions are extracted respectively.

[0021] Read the coordinates of the guard's torso center point, the coordinates of the wrist point holding the gun, and the main axis direction of the firearm area in the current frame; connect the guard's torso center point and the wrist point holding the gun to form the gun-holding baseline; Using the main axis of the firearm area as the outward expansion direction and the firearm holding baseline as the center line, equal-distance outward expansion is performed on both sides of the firearm holding baseline; then the two outward expansion boundary lines are closed and connected with the starting end and the ending end of the firearm holding baseline to form a strip-shaped area, and the strip-shaped area is defined as the firearm control zone; Read the position of the door hinge edge and the position of the pedal edge in the unified scene plane, take the door hinge edge as the inner boundary, take the pedal edge as the outer boundary, and take the connecting line between the two ends of the above two boundaries as the front and rear closed boundaries to form a closed area. When the length of the outer edge of the pedal is greater than the length of the door hinge edge, the corresponding segment is cut off according to the vertical projection position of the two ends of the door hinge edge on the outer edge of the pedal, and then the cut-off pedal outer edge segment and the door hinge edge form a closed area; the closed area is defined as the door buffer zone. Read the boundary position of the handover window in the unified scene plane, and use the outer contour of the handover window boundary as a reference to perform fixed-width outward expansion along the side away from the vehicle body to form an extended boundary parallel to the outer contour of the handover window boundary. The area between the boundary of the junction window and the extended boundary is taken as the junction allowable zone. When there is a corner at the boundary of the junction window, the fixed-width outward expansion is performed according to the adjacent boundary segments. Then, the endpoints of each outward expansion boundary segment are connected in sequence to form a continuous closed strip area. The continuous closed strip area is determined as the junction allowable zone.

[0022] In this embodiment, step three specifically includes: Read the current frame in the close-range monitoring image sequence in chronological order, extract the main body area of ​​the escort, the area of ​​external approaching personnel, the firearm area, the holster area, the external hand area, the firearm control strip, the door buffer strip, the handover allowance strip, and the door structure area, and perform independent boundary closure processing on each area to generate the corresponding closed contour of each area.

[0023] For each closed contour in the current frame, calculate the contour area and contour center coordinates, and assign object numbers according to the region category. The main area of ​​the escort is numbered as the first object, the area of ​​external approaching personnel is numbered as the second object, the firearm area is numbered as the third object, the holster area is numbered as the fourth object, the external hand area is numbered as the fifth object, the firearm control strap is numbered as the sixth object, the door buffer strip is numbered as the seventh object, the handover allowance strip is numbered as the eighth object, and the door structure area is numbered as the ninth object. When there are multiple areas in the same category, sub-numbers are added sequentially according to the distance from the center point of each area to the center point of the escort's main area.

[0024] Extract the center coordinates from the third object and the first object in the current frame to form the first type of pairing; Extract the center coordinates from the third and fourth objects to form a second type of pairing; Extract the center coordinates from the first object and the ninth object to form the third type of pairing; Extract the center coordinates from the second and third objects to form a fourth type of pairing; Extract the center coordinates from the fifth and sixth objects to form the fifth type of pairing; Extract the center coordinates from the fifth and seventh objects to form the sixth type of pairing; Extract the center coordinates from the second object and the first object to form the seventh type of pairing; Extract the center coordinates from the second and seventh objects to form the eighth type of pairing.

[0025] For each type of pairing, connect the center coordinates of the objects at both ends of the pairing to generate object lines; Calculate the orientation angle of the object connection line within the unified scene plane based on the starting and ending coordinates of the object connection line; Calculate the center point distance between the paired objects based on the difference in horizontal and vertical coordinates between the starting and ending coordinates. For each pairing, read the closed contours of the two objects at both ends and calculate the overlapping area of ​​the closed contours of the two objects at both ends; When the closed contours of the two objects overlap, extract the closed boundary of the overlapping part and calculate the area value of the overlapping part; When the closed contours of the two objects do not overlap, the overlapping area is recorded as zero. For each pairing, extract the boundary segments in the closed contours of the two objects whose distance from each other is less than the preset boundary distance; Calculate the boundary interval between corresponding points along the boundary segment point by point; The lengths of consecutive boundary segments with a boundary spacing less than the preset boundary distance are summed to obtain the boundary contact length of this type of pairing. When there are no continuous boundary segments, the boundary contact length is recorded as zero; For each pairing, compare the horizontal and vertical positions of the center coordinates of the two objects in the unified scene plane; When the center coordinates of one object are located closer to the vehicle body side of the center coordinates of the other object, that object is recorded as the inner object. When the center coordinates of one object are located on the side of the other object that is far from the vehicle body, that object is recorded as the outer object. When the center coordinates of one object are located near the center coordinates of the other object in the direction of the vehicle's front, that object is recorded as the forward object. When the center coordinates of one object are located near the rear of the vehicle, the object at that end is recorded as the rear object. Based on the determination results of the inner object, outer object, forward object, and backward object, the relative position order of the corresponding pair is generated.

[0026] The object numbers, center point positions at both ends, center point spacing, object connection direction angle, overlap area, boundary contact length, and relative position order corresponding to the first to eighth pairings are recorded as the relationship connection results in the current frame according to the pairing category. In this embodiment, step four specifically includes: Read the object numbers, closed contours, center coordinates, and corresponding relationship connection results of the escort body area, external approaching personnel area, firearm area, holster area, and external hand area in the current frame and the next frame in chronological order; Starting from the center coordinates of each object in the current frame, search for the center coordinates of objects of the same category in the next frame; For each object in the current frame, calculate the center point distance, outline overlap area, and difference in bounding rectangle size between it and other objects of the same category in the next frame; The object with the smallest center point distance, the largest outline overlap area, and the difference in the size of the outer rectangle is less than a preset range is retained as the matching object of the current object in the next frame; Key points were extracted from the shoulder, elbow, wrist, hip, knee, and ankle of the matched escort personnel's main area and the external approach personnel area, respectively. Specifically: First, determine the head and shoulder transition position on the human body's outer contour, and take the local peak points on both sides below the head and shoulder transition position as the left and right shoulder key points; then, search downwards along the connected area between the shoulder key points and the middle section of the human body's outer contour for contour transition points, and take the bending position of the middle section of the upper limb as the left and right elbow key points; continue to search along the direction of the upper limb's end for the forearm termination position, and take the position where the end of the forearm connects with the hand as the left and right wrist key points; determine the bifurcation position where the torso transitions to the legs on both sides of the lower half of the human body's outer contour, and take the upper edge of the bifurcation position as the left and right hip key points; search downwards along the hip key points for the middle section of the leg transition position, and take the corresponding transition position as the left and right knee key points; then search along the direction of the lower leg termination position, and take the end position of the leg before it contacts the ground as the left and right ankle key points; when a key point on one side is occluded, read the corresponding key point position of the same numbered object in the previous frame, and combine it with the unoccluded adjacent key point position in the current frame to perform point filling; Read the key points of the left and right shoulders, the left and right hips, and the center position of the weapon area of ​​the escort officer's main body in the current frame; connect the midpoints of the left and right shoulder key points and the midpoints of the left and right hip key points to obtain the central axis of the escort officer's torso; The midpoint between the midpoint of the left and right shoulder key points and the midpoint of the left and right hip key points is taken as the center point of the escort's torso; then, the center point of the escort's torso is used as the local reference origin, and the coordinates of the key points of the shoulder, elbow, wrist, hip, knee and ankle in the escort's main body area and the external approaching personnel area in the current frame are translated respectively. The translation process specifically involves subtracting the horizontal coordinate of the escort's torso center point from the horizontal coordinate of each key point, and subtracting the vertical coordinate of the escort's torso center point from the vertical coordinate of each key point to obtain normalized key point coordinates. A preset number of video frames are read sequentially in chronological order; key point extraction and translation are performed in each frame; then, the normalized key point coordinates of the left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles of the escort's main body area and the external approaching personnel area are written into the key point sequence table in frame order; the key point sequence table is arranged in chronological order, and the time stamp and object number of each frame are written at the corresponding position of each frame. In each frame corresponding to the keypoint sequence list, the normalized keypoint coordinates of the left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles are used as the center points to generate a joint heatmap in the local plane of the corresponding frame. The method for generating the key point heat map is as follows: taking the key point coordinates as the center, assign values ​​to the surrounding areas in a local area of ​​a preset size, so that the pixel value at the center of the key point is the largest, and fill the pixel values ​​outward from the center in a decreasing manner to obtain a two-dimensional heat map corresponding to a single key point; then store the two-dimensional heat map corresponding to all key points in the same frame according to the key point category. According to the order of limb connection, in each frame, connect the left and right shoulders with the left and right elbows, the left and right elbows with the left and right wrists, the left and right hips with the left and right knees, and the left and right knees with the left and right ankles respectively; with each connecting line segment as the center, fill the pixels on both sides of the connecting line segment with a preset width to generate the corresponding limb heat map. For each pixel on a connecting line segment, assign a decreasing value based on its distance from the center line of the connecting line segment; after filling all connecting line segments, store the limb heatmaps corresponding to each connecting line segment in the same frame according to limb category. The joint point heatmaps and limb heatmaps from several consecutive frames are superimposed frame by frame in chronological order. Specifically, joint point heatmaps of the same key point category in different time frames are stacked in chronological order, and limb heatmaps of the same limb category in different time frames are stacked in chronological order. Then, the stacked results of the heatmaps corresponding to all key point categories and all limb categories are combined to form a three-dimensional posture heatmap corresponding to the current time window. Read the relationship connection results and object displacement record table corresponding to the current time window; extract the distance between the center points at both ends of the pairing results between the external approaching personnel area and the firearm area frame by frame, and write the distance sequence from the external approaching personnel to the firearm in chronological order; For the pairing results of the external hand region and the firearm control strap, the boundary length of the external hand region entering the firearm control strap is extracted frame by frame and written into the entry length sequence of the external hand entering the firearm control strap in chronological order. For the pairing results of the external hand area and the door buffer strip, the boundary length of the external hand area entering the door buffer strip is extracted frame by frame, and written into the entry length sequence of the external hand entering the door buffer strip in chronological order. For the pairing results of the firearm area and the escort's main body area, calculate the angle between the main axis of the firearm and the central axis of the escort's torso frame by frame, and write the angle sequence between the main axis of the firearm and the central axis of the escort's torso in chronological order. For the pairing results of the gun area and the holster area, the distance between the center points at both ends is extracted frame by frame and written into the distance sequence from the gun area to the holster area in chronological order.

[0027] Each time window's corresponding three-dimensional attitude thermogram is assigned a time window number. The sequence of distances from external personnel to the firearm, the sequence of external hands entering the firearm's control strap, the sequence of external hands entering the vehicle door buffer strap, the sequence of angles between the firearm's main axis and the escort's torso axis, and the sequence of distances from the firearm area to the holster area, all with the same time window number, are written into the corresponding sequence recording area. The three-dimensional attitude thermograms are then bound to each sequence recording area according to the time window number to form a motion analysis window sequence.

[0028] In this embodiment, the improved PoseC3D model is specifically as follows: The motion analysis window sequence is input into the thermal body segmentation module. The 3D attitude key point thermal bodies in the motion analysis window sequence are read sequentially according to the time window. Part segmentation processing is then performed on the 3D attitude key point thermal bodies. The part segmentation processing specifically includes: Extract the joint point heat maps corresponding to the left and right shoulders, left and right elbows, and left and right wrists, as well as the limb heat maps corresponding to the left and right shoulders and left and right elbows, and the left and right elbows and left and right wrists. Combine the joint point heat maps and limb heat maps in the original time sequence to form the upper limb gun-holding heat body. Extract the joint point heat maps corresponding to the left and right shoulders and the left and right hips, as well as the connection area heat map between the midpoint of the left and right shoulders and the midpoint of the left and right hips, and combine the joint point heat map and the connection area heat map in the original time sequence to form a stable thermodynamic body of the torso; Extract the joint point heat maps corresponding to the left and right hips, left and right knees, and left and right ankles, as well as the limb heat maps corresponding to the left and right hips and left and right knees, and the left and right knees and left and right ankles. Combine the joint point heat maps and limb heat maps in the original time sequence to form a lower limb displacement heat body. The thermal models of the upper limb holding the gun, the trunk stabilization, and the lower limb displacement are input into the branch extraction module. Three-dimensional convolution extraction processing is then performed on each of the three thermal models. The three-dimensional convolution extraction processing specifically includes: The upper limb holding a gun thermal body is input into the first three-dimensional convolution branch, and local convolution scanning and inter-layer downsampling are performed sequentially in the time dimension, horizontal dimension and vertical dimension to output the gun holding action body features; Input the stable thermodynamic body of the torso into the second three-dimensional convolutional branch, and perform local convolutional scanning and inter-layer downsampling processing in the time dimension, horizontal dimension and vertical dimension in sequence to output the torso support body features; The lower limb displacement thermodynamic body is then input into the third three-dimensional convolution branch, and local convolution scanning and inter-layer downsampling processing are performed sequentially in the time dimension, horizontal dimension and vertical dimension to output the station movement body features; The features of the gun-holding action, the features of the torso support, and the features of the standing and moving body are input into the timing splicing module according to the same time window number. The features of the gun-holding action, the features of the torso support, and the features of the standing and moving body are read according to the time window number and spliced ​​together in chronological order to form the posture timing features. Read the distance sequence from the external approaching personnel to the firearm, the entry length sequence of the external hand into the firearm control strap, the entry length sequence of the external hand into the vehicle door buffer strap, the angle sequence between the main axis of the firearm and the central axis of the escort's torso, and the distance sequence from the firearm area to the holster area, which are the same as the time window number. Perform length alignment on each relation sequence according to the time frame position until the start frame and end frame of each relation sequence are consistent with the start frame and end frame of the attitude timing feature; The length-aligned relational sequences are sequentially concatenated to the post-pose temporal features in a preset order to form a joint action sequence; The combined action sequence is input into the segment labeling module, and the upper limb movement change segment, trunk support change segment, lower limb displacement change segment, and each relation sequence change segment in the combined action sequence are read within each time window. Synchronously compare the segments of upper limb movement changes with the segments of angle changes between the main axis of the firearm and the central axis of the escort's torso, and mark the time segments in which the direction of upper limb rotation changes continuously and the angle between the firearm and the escort's torso as the gun-holding deflection segments. The shoulder-to-elbow and elbow-to-wrist transition segments in the upper limb movement change segments were compared synchronously, and the time periods in which the shoulder-to-elbow and elbow-to-wrist transitions increased continuously within the same time window were marked as upper limb abrupt transition segments. Synchronously compare the torso support change segment with the displacement change segment in the stationary movement characteristics, and combine the distance sequence change from the approaching personnel to the firearm to mark the time period when the torso center shift and the lower limbs move backward consecutively as the forced retreat segment. The entry length sequence of external hands entering the gun control area and the entry length sequence of external hands entering the door buffer area are checked window by window, and the time period when the entry length continuously increases is marked as the hand intrusion segment. Synchronously compare the changes in the physical characteristics of the gun-holding action with the changes in the distance sequence from the gun area to the holster area. Mark the time period in which the changes in the gun-holding action occur continuously and the distance between the gun area and the holster area increases continuously as the gun unholstering segment. The segments of sudden upper limb turning, gun deflection, forced retreat, hand intrusion, and gun disarming were used as segment labeling results.

[0029] The improved PoseC3D model proposed in this step shares similarities with the traditional PoseC3D model in that both use human pose information as the core input. Both first convert keypoint information from consecutive video frames into a 3D pose heatmap with temporal and spatial dimensions. Then, they use 3D convolution to extract features from the heatmap along the temporal, lateral, and longitudinal dimensions, thereby obtaining a spatiotemporal representation of motion changes. Neither model directly performs whole-frame recognition on the original image; instead, they highlight joint positions and limb connections through keypoint and limb heatmaps, allowing the model to more effectively represent pose changes during movement. Furthermore, both the traditional and improved PoseC3D models retain the basic mechanism of motion analysis based on temporal windows. They rely on a heatmap input consisting of several consecutive frames to extract continuous motion features of a person over a period of time, rather than relying solely on the static pose results of a single frame. Furthermore, both methods extract motion body features through 3D convolution and then perform subsequent analysis and recognition of these features. This indicates that the improved PoseC3D model used in this step is still an extension of PoseC3D based on the existing technology. It maintains the basic framework and core ideas of traditional PoseC3D in terms of pose thermal body modeling, 3D convolution spatiotemporal feature extraction, and time window motion representation.

[0030] The difference lies in the fact that the improved PoseC3D model proposed in this step does not follow the traditional PoseC3D approach of inputting the entire three-dimensional posture thermosphere into a single path for unified feature extraction. Instead, in the thermosphere segmentation module, according to the operational requirements of close-range emergency action warning for integrated vehicle and gun units, the original three-dimensional posture key point thermosphere is first split into upper limb gun-holding thermosphere, torso stability thermosphere, and lower limb displacement thermosphere, so that the different action semantics corresponding to different body parts are preserved and processed separately. Subsequently, the improved model sets up three independent three-dimensional convolutional branches in the branch extraction module, corresponding to the action feature extraction of three different directions: gun-holding action, torso support, and stance movement, respectively, instead of using a single convolutional path to extract all posture information. Furthermore, the improved model does not stop at the posture features themselves. In the temporal stitching module, it stitches together the features of the gun-holding action, the torso support, and the standing movement according to time window numbers to form a temporal posture feature. It further introduces the distance sequence from the approaching person to the firearm, the entry length sequence of the external hand into the firearm control strap, the entry length sequence of the external hand into the vehicle door buffer strap, the angle sequence between the main axis of the firearm and the central axis of the escort's torso, and the distance sequence from the firearm area to the holster area, and performs length alignment and joint stitching with the temporal posture feature. Finally, in the segment labeling module, the improved model does not output general action classification results, but directly performs window-by-window scanning and segment classification processing on segments such as gun-holding deflection, upper limb abrupt turning, forced retreat, hand intrusion, and firearm disengagement. Therefore, its output objects are more closely aligned with the risk action recognition needs in this patent scenario.

[0031] The beneficial effect of the improvements lies in the fact that by segmenting the 3D posture key point thermospheres and constructing separate thermospheres for upper limb gun-holding, torso stability, and lower limb displacement, the motion information that is prone to interference under the traditional unified whole-body representation can be processed separately. This allows for clearer and more independent representations of different types of posture changes, such as upper limb gun-holding deflection, torso support changes, and lower limb stance shifts, thereby enhancing the model's ability to identify fine-grained risky actions in guard and escort scenarios. By setting three independent 3D convolutional branches, the improved model can extract features of gun-holding actions, torso support, and stance movement separately, avoiding the mixing and dilution of motion changes from different body parts within the same feature channel, and making the expression of key motion changes more prominent in both the temporal and spatial dimensions. Furthermore, by aligning and stitching the aforementioned three types of body features with sequences of relationships such as gun distance, control strip entry length, vehicle door buffer strip entry length, gun main axis angle, and holster spacing, the model can combine changes in human posture with changes in close-range relationships between the vehicle, gun, and person. This allows the model output to go beyond human movement itself and directly serve the labeling of risky scenarios such as gun deflection, hand intrusion, forced retreat, and gun disarming. In this way, the model retains PoseC3D's advantages in posture thermal spatiotemporal modeling while improving its adaptability to close-range, sudden action scenarios involving integrated guard and escort vehicles and guns, thus enhancing the accuracy, continuity, and practicality of the segment labeling results.

[0032] In this embodiment, step six specifically includes: Read the upper limb sudden turn segment, gun deflection segment, forced retreat segment, hand intrusion segment, and gun disarming segment from the segment marking results in chronological order, and record the start frame, end frame, and object number of each segment. Segments that overlap in time or are consecutive in sequence are merged to form continuous warning segments, and the warning type of each continuous warning segment is determined, specifically including: When a continuous warning segment includes segments of sudden upper limb turning and gun deflection, the corresponding continuous warning segment is identified as a robbery warning type. When a continuous warning segment contains both a probe intrusion segment and a forced retreat segment, the corresponding continuous warning segment is identified as a containment warning type. When a continuous warning segment includes both a weapon deflection segment and a weapon dislodgement segment, the corresponding continuous warning segment is identified as a disengagement warning type. Extract the distance sequence from external approaching personnel to firearms, the entry length sequence of external hands into the firearm control strap, the entry length sequence of external hands into the vehicle door buffer strip, the distance sequence from the firearm area to the holster area, and the change sequence of the escort personnel's position within the time range corresponding to each continuous warning type, and calculate the change amount of each sequence within the time range. The change amount corresponding to each continuous warning type is compared with a preset threshold. When the change amount reaches the preset threshold, the corresponding warning level is generated. Write the warning type, warning level, start frame, end frame, and object number into the warning record to generate warning information.

[0033] Example 1: To verify the feasibility of this invention in practice, it was applied to an area where a core financial district and a high-density commercial area intersect in a provincial capital city. When security vehicles perform cash box handover tasks in front of branch offices, the vehicles typically need to make a brief stop in designated parking areas. The security guard disembarks from the vehicle door and moves to the handover side to complete the delivery. Security personnel must maintain close-range vigilance between the vehicle door and the handover side. In this scenario, the vehicle body, door structure, handover window, pedal area, security guard position, approach path of external personnel, and the positional relationship between firearms and holsters change continuously within a very short time. Especially in areas where branch entrances / exits, non-motorized vehicle parking areas, and pedestrian walkways overlap, ordinary passersby, queuing customers, branch security guards, and the objects being escorted will move densely within a few meters, easily leading to high-risk situations such as rapid close-range approach, hands reaching beyond boundaries, blocking at the vehicle door, forced retreat of security guards, and firearms deviating from stable control positions. In existing methods, most on-site operations still rely on manual monitoring, ordinary video playback, or simple area intrusion judgment. This makes it difficult to unify vehicle structure, personnel movements, weapon status, and continuous temporal changes within a single analytical framework. Consequently, problems arise such as delayed warnings, high false alarm rates, and insufficient identification of pre-robbery, pre-encirclement, and pre-loss-from-control warnings. This embodiment addresses these issues by deploying the computer vision-based integrated vehicle-gun close-range emergency action warning method of this invention in a real-world, close-range security operation environment. It uniformly processes continuous video footage from the vehicle handover side, door side, and guard side, generating warning information that can be directly used for on-site security coordination.

[0034] In practical applications, fixed video acquisition devices are first installed on the handover side, door side, and guard side of the escort vehicle. All three acquisition directions cover the areas including the door hinge edge, door handle, handover window boundary, and outer edge of the pedal. After the system is operational, it continuously acquires video feeds from each channel and writes a unified time stamp to each frame. Then, the pixel positions of the door hinge edge, door handle, handover window boundary, and outer edge of the pedal are extracted from each video feed. Using the corresponding positional relationships of the same fixed reference under different viewpoints, the planar mapping relationship from each video feed to a unified scene plane is calculated. After mapping, the outer area of ​​the door, the outer area of ​​the handover window, and the guard's position area are cropped within the unified scene plane to obtain a sequence of close-up monitoring images. Because the ground projections from different camera directions are unified to the same scene plane, the vehicle structure and personnel activity areas, previously scattered across different images, are pulled into the same spatial coordinate system. Subsequent actions, such as determining if an external person is approaching a weapon or if an external hand enters the door buffer zone, can be performed at a unified scale, avoiding deviations caused by perspective errors and differences in occlusion angles in single-channel images.

[0035] After obtaining the close-range monitoring image sequence, the system separates the human body region and the vehicle structure region in each frame, and then distinguishes the escort officer's main body region and the external approaching personnel region within the human body region. The system extracts the firearm region within the upper limb search range of the escort officer's main body region, the holster region within the waist search range, and extracts a local sub-region from the forearm end of the external approaching personnel and their external hand region. Then, it connects the center point of the escort officer's torso with the wrist point holding the firearm to form a firearm-holding baseline, and extends it outwards along the main axis of the firearm to generate a firearm control zone. A door buffer zone is then formed based on the door hinge edge and the outer edge of the pedal, and a handover allowance zone is generated by extending it outwards according to the boundary of the handover window. In this way, the key objects in the scene are no longer just coarse-grained recognition results of "someone," "vehicle," and "movement," but rather form a vehicle-gun movement region composed of the escort officer's main body region, the external approaching personnel region, the firearm region, the holster region, the external hand region, the firearm control zone, the door buffer zone, and the handover allowance zone. For security operations, such zoning is closer to actual operational rules. For example, it is normal for customers to move within the handover zone, but it is clearly outside the scope of normal operations for external hands to enter the firearms control zone or the vehicle door buffer zone.

[0036] In the relationship modeling phase, the system further numbers each object in the vehicle-gun action area of ​​each frame and establishes eight types of pairing relationships: firearms and escort personnel, firearms and holsters, escort personnel and vehicle door structure, external approach personnel and firearms, external hands and firearm control straps, external hands and vehicle door buffer strips, external approach personnel and escort personnel, and external approach personnel and vehicle door buffer strips. For each pairing, the system calculates the position of the center points at both ends, the distance between the center points, the direction angle of the object connection line, the overlap area, the boundary contact length, and the relative position order to form the relationship connection result in the current frame. Subsequently, the system reads the object numbers and relationship connection results in adjacent frames in chronological order, performs cross-frame matching, and obtains the continuous trajectory of the escort personnel, external approach personnel, firearms, holsters, and external hands; and extracts the coordinates of key points of the shoulder, elbow, wrist, hip, knee, and ankle of the escort personnel and external approach personnel to generate joint heatmaps, limb heatmaps, and three-dimensional posture heatmaps. Simultaneously, the sequence of distances from external personnel to the firearm, the sequence of entry lengths of external hands into the firearm's control strap, the sequence of entry lengths of external hands into the vehicle door buffer strap, the sequence of angles between the firearm's main axis and the escort's torso axis, and the sequence of distances from the firearm area to the holster area are bound to a three-dimensional posture thermosphere according to time window numbers to form a motion analysis window sequence. Then, the system inputs the motion analysis window sequence into the improved PoseC3D model, and through thermosphere segmentation, branch extraction, temporal splicing, and segment labeling, obtains segment labeling results such as upper limb abrupt turn segments, firearm deflection segments, forced retreat segments, hand intrusion segments, and firearm disengagement segments. Finally, based on the segment labeling results, continuous warning segments are generated to determine three warning types: robbery, encirclement, and loss of control, and the warning level, start frame, end frame, and object number are output, directly fed back to the vehicle terminal and monitoring center.

[0037] To verify the practical effectiveness of the method of this invention, 18 operation points were selected in the financial security escort operation route of a certain city, including core urban business outlets, community-type business outlets, and business district-type business outlets. Close-range video of the security vehicles was continuously collected, resulting in 216 sets of effective close-range monitoring video sequences, with a total video duration of 158.4 hours. The samples included normal handover scenarios, ordinary people passing by scenarios, close-range observation scenarios, sudden approach scenarios, hand reaching into scenarios, gathering at vehicle doors scenarios, and firearms being removed from their holsters scenarios. After on-site verification, 462 key events that could be used for risk verification were identified, including 286 normal handover actions, 91 ordinary close-range contact actions without crossing boundaries, 31 pre-robbery warning events, 28 pre-encirclement warning events, and 26 pre-loss warning events. The method of this invention was compared and tested with existing conventional area intrusion early warning schemes and ordinary posture thermal body recognition schemes without improved PoseC3D. The conventional area intrusion warning scheme primarily uses boundary crossings outside the vehicle doors and in fixed areas around the vehicle as the alarm basis; the ordinary posture thermal body recognition scheme uses the overall human body thermal body for recognition, but does not incorporate the joint action sequence processing based on the vehicle-gun relationship sequence as described in this invention. The comparison results are shown in Table 1.

[0038] Table 1. Comparative Test Table of Close-Range Sudden Action Warning for Integrated Guard and Escort Vehicle-Mounted Guns As shown in Table 1, while conventional area intrusion warning schemes have a relatively fast processing speed, their reliance on fixed area boundary judgments results in a significantly high false alarm rate in normal handover, ordinary passing, and close-range contact scenarios, making it difficult to distinguish between normal approach and dangerous approach. Conventional posture thermal body recognition schemes have shown improvement, but due to the lack of a joint action sequence built around the close-range operation scenario of integrated guard and escort vehicles, the coupling degree of human posture, external proximity relationships, and changes in weapon status is insufficient, leading to missed alarms in the three directions of seizure, containment, and loss of control. The method of this invention achieves recognition rates of 93.5%, 92.9%, and 92.3% for the three types of risk precursors, respectively, with a comprehensive false alarm rate reduced to 2.4% and a comprehensive missed alarm rate reduced to 6.7%. The consistency rate between warning information and on-site verification reaches 96.4%, indicating that it has higher relevance and stability for close-range operation scenarios of integrated guard and escort vehicles.

[0039] Further analysis revealed that the method of this invention achieved a recognition accuracy of 95.8% in normal handover scenarios, significantly higher than the comparative scheme. The key reason for this is that this invention does not simply treat all approaching behavior as risk, but rather distinguishes between "permitted approach" and "approaching beyond the boundary" through spatial constraints such as the handover allowance zone, the vehicle door buffer zone, and the firearm control zone. For example, in front of a community business outlet, if a customer lingers outside the handover window and keeps their hands within the handover allowance zone, the system will not generate a risk warning; however, when an outsider suddenly extends their hand into the firearm control zone, or when multiple people rapidly gather towards the vehicle door buffer zone, the system can promptly generate a hand intrusion segment or a related containment segment within a continuous time window. For example, in front of a commercial area, where people move quickly and there are many background obstructions, conventional methods often misjudge ordinary passersby as approaching danger. However, the method of this invention further combines the main body area of ​​the escort, the firearm area, the external hand area, and the corresponding relationship sequence changes. Only when segments such as sudden upper limb turning, gun deflection, hand intrusion, or forced retreat meet the conditions simultaneously will a continuous warning segment be formed, thus significantly reducing false alarms.

[0040] This embodiment demonstrates that the method of the present invention effectively solves the problems of difficulty in unifying multi-view close-range scenes, difficulty in expressing the relationship between vehicles, guns, and personnel, difficulty in identifying continuous risky actions, and unstable early warning results in the prior art. By mapping multiple images onto a unified scene plane, separating the human body area and vehicle structure area frame by frame, and establishing object numbers and pairing relationships around the escort, external approaching personnel, firearms, holsters, and external hands, and then utilizing cross-frame matching, thermal body construction, and improved PoseC3D model fragment labeling processing, it can more accurately identify continuous early warning fragments related to robbery, encirclement, and loss of control, and further output early warning information with type, level, and object number. This method is particularly suitable for complex scenes such as close handover of escort vehicles, dense personnel in the vehicle door area, and limited warning space, and has high engineering application value and practical promotion significance.

[0041] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A computer vision-based method for early warning of sudden close-range actions of integrated guard and escort vehicle-mounted guns, characterized in that, Includes the following steps: Step 1: Collect continuous video footage from the handover side, door side, and warning side of the guarded vehicle, calculate the planar mapping relationship, and generate a sequence of close-range monitoring footage; Step 2: Separate the human body area and vehicle structure area frame by frame in the close-up monitoring image sequence, and determine the vehicle gun action area; Step 3: Number and pair objects in the vehicle-gun action area, and generate the relationship connection results in the current frame; Step 4: Read the object numbers and relationship connection results in adjacent frames in chronological order, perform cross-frame matching to generate joint heatmaps and limb heatmaps, and form a motion analysis window sequence; Step 5: Input the motion analysis window sequence into the improved PoseC3D model, and obtain the fragment labeling results through the thermal body segmentation module, branch extraction module, temporal splicing module, and fragment labeling module; Step 6: Based on the segment marking results, form continuous warning segments, determine the warning type of each continuous warning segment, and generate warning information.

2. The method for early warning of sudden close-range actions of a vehicle-mounted gun integrated with guard equipment based on computer vision according to claim 1, characterized in that, Step one specifically involves: Collect continuous video footage from the handover side, door side, and guard side of the guarded vehicles, and write time stamps to each frame of each continuous video footage. Extract the pixel positions of the door hinge edge, door lock handle, junction window boundary, and pedal outer edge from each continuous video frame; Calculate the planar mapping relationship between consecutive video frames based on the corresponding pixel positions of the same fixed reference in different video frames; Transform each continuous video frame onto a unified scene plane according to the planar mapping relationship; By cropping the outer area of ​​the vehicle door, the outer area of ​​the handover window, and the area where the escort is stationed within a unified scene plane, a sequence of close-up monitoring images is obtained.

3. The method for early warning of sudden close-range actions of a vehicle-mounted gun integrated with guard service based on computer vision according to claim 1, characterized in that, Step two specifically involves: In the close-up monitoring image sequence, the human body area and the vehicle structure area are separated frame by frame, and the human body area is further distinguished from the main body area of ​​the escort personnel and the area of ​​external approaching personnel. The firearms area was extracted from the upper limbs of the escort personnel's main body area, and the holster area was extracted from the waist side of the escort personnel's main body area. A local sub-region is cropped at the end of the forearm near the personnel area, and the external hand region is extracted; A gun-holding baseline is formed by connecting the center point of the guard's torso with the point of the wrist holding the gun, and the gun-holding baseline is widened at a fixed width on both sides along the main axis of the firearm to generate a firearm control belt; A door buffer zone is formed by the door hinge edge and the outer edge of the pedal in the vehicle structure area. The width is expanded outward according to the boundary of the handover window to generate the handover allowance zone. The main area for escort personnel, the area for external personnel to approach, the area for firearms, the area for holsters, the area for external hands, the firearm control strap, the buffer zone of the vehicle door, and the handover allowance zone are designated as the vehicle-gun movement area.

4. The method for early warning of sudden close-range actions of a vehicle-mounted gun integrated with guard equipment based on computer vision according to claim 1, characterized in that, Step three specifically involves: Assign object numbers to the escort body area, external approaching personnel area, firearm area, holster area, external hand area, firearm control strip, vehicle door buffer strip, and handover allowance strip in each frame of the vehicle-gun action area. The first type of pairing consists of the firearms area and the main body area of ​​the escort; the second type of pairing consists of the firearms area and the holster area; the third type of pairing consists of the main body area of ​​the escort and the door structure area; and the fourth type of pairing consists of the external approach personnel area and the firearms area. The fifth type of pairing consists of the external hand area and the firearm control strap; the sixth type of pairing consists of the external hand area and the door buffer strip; the seventh type of pairing consists of the external approach personnel area and the main body area of ​​the escort; and the eighth type of pairing consists of the external approach personnel area and the door buffer strip. For each type of pairing, calculate the position of the center points at both ends, the distance between the center points, the direction angle of the object connection line, the overlapping area, the boundary contact length, and the relative position order, and record the calculation results as the relationship connection results in the current frame according to the pairing category.

5. The method for early warning of sudden close-range actions of a vehicle-mounted gun integrated with guard equipment based on computer vision according to claim 1, characterized in that, Step four specifically involves: The object numbers and relationship connection results in two adjacent frames are read in chronological order. Cross-frame matching is then performed on the main body area of ​​the escort, the area of ​​approaching external personnel, the firearm area, the holster area, and the external hand area. The cross-frame matching specifically includes: Extract the coordinates of key points on the shoulders, elbows, wrists, hips, knees, and ankles of the escort and the personnel approaching from outside; Using the center point of the escort officer's torso in a unified scene plane as the local reference origin, the coordinates of each key point are translated and normalized. Write the coordinates of each key point in a series of consecutive frames into the key point sequence in chronological order. A joint heatmap is generated within the corresponding frame, centered on the coordinates of each key point. According to the order of limb connection, generate limb heat maps between the shoulder and elbow, between the elbow and wrist, between the hip and knee, and between the knee and ankle; The joint heatmaps and limb heatmaps from several consecutive frames are stacked in chronological order to form a three-dimensional posture heatmap. The sequence of distances between external personnel approaching the firearm, the sequence of entry lengths of external hands into the firearm control strap, the sequence of entry lengths of external hands into the vehicle door buffer strap, the sequence of angles between the main axis of the firearm and the central axis of the escort's torso, and the sequence of distances from the firearm area to the holster area within a series of consecutive frames are bound to the three-dimensional attitude thermosphere according to the time window number to form a motion analysis window sequence.

6. The method for early warning of sudden close-range actions of a vehicle-mounted gun integrated with guard service based on computer vision according to claim 1, characterized in that, The improved PoseC3D model is specifically as follows: The motion analysis window sequence is input into the thermal body segmentation module. The 3D attitude key point thermal bodies in the motion analysis window sequence are read sequentially according to the time window. Part segmentation processing is then performed on the 3D attitude key point thermal bodies. The part segmentation processing specifically includes: Extract the joint point heat maps corresponding to the left and right shoulders, left and right elbows, and left and right wrists, as well as the limb heat maps corresponding to the left and right shoulders and left and right elbows, and the left and right elbows and left and right wrists. Combine the joint point heat maps and limb heat maps in the original time sequence to form the upper limb gun-holding heat body. Extract the joint point heat maps corresponding to the left and right shoulders and the left and right hips, as well as the connection area heat map between the midpoint of the left and right shoulders and the midpoint of the left and right hips, and combine the joint point heat map and the connection area heat map in the original time sequence to form a stable thermodynamic body of the torso; Extract the joint point heat maps corresponding to the left and right hips, left and right knees, and left and right ankles, as well as the limb heat maps corresponding to the left and right hips and left and right knees, and the left and right knees and left and right ankles. Combine the joint point heat maps and limb heat maps in the original time sequence to form a lower limb displacement heat body. The upper limb holding gun thermal body, the trunk stabilization thermal body, and the lower limb displacement thermal body are input into the branch extraction module. Three-dimensional convolution extraction processing is performed on the three thermal bodies respectively. The three-dimensional convolution extraction processing is to perform local convolution scanning and inter-layer downsampling processing in the time dimension, the horizontal dimension, and the vertical dimension respectively, and output the gun holding action body features, the trunk support body features, and the standing position movement body features. The features of the gun-holding action, the features of the torso support, and the features of the standing and moving body are input into the timing splicing module according to the same time window number. The features of the gun-holding action, the features of the torso support, and the features of the standing and moving body are read according to the time window number and spliced ​​together in chronological order to form the posture timing features. Read the distance sequence from the external approaching personnel to the firearm, the entry length sequence of the external hand into the firearm control strap, the entry length sequence of the external hand into the vehicle door buffer strap, the angle sequence between the main axis of the firearm and the central axis of the escort's torso, and the distance sequence from the firearm area to the holster area, which are the same as the time window number. Perform length alignment on each relation sequence according to the time frame position until the start frame and end frame of each relation sequence are consistent with the start frame and end frame of the attitude timing feature; The length-aligned relational sequences are sequentially concatenated to the post-pose temporal features in a preset order to form a joint action sequence; The joint action sequence is input into the fragment labeling module, and window-by-window scanning and fragment classification processing are performed on the joint action sequence according to the time window order to obtain the fragment labeling result.

7. The computer vision-based method for early warning of sudden close-range actions of integrated guard and escort vehicle-mounted guns according to claim 6, characterized in that, The process of performing window-by-window scanning and segment classification on the joint action sequence according to the time window order to obtain segment labeling results is as follows: Within each time window, read the upper limb movement change segment, trunk support change segment, and lower limb displacement change segment in the combined action sequence; Synchronously compare the segments of upper limb movement changes with the segments of angle changes between the main axis of the firearm and the central axis of the escort's torso, and mark the time segments in which the upper limb rotation direction changes continuously and the angle between the firearm and the escort's torso increases continuously and is greater than a preset threshold as the gun-holding deflection segment. The shoulder-to-elbow and elbow-to-wrist transition segments in the upper limb movement change segments are compared synchronously. The time segments in which the shoulder-to-elbow and elbow-to-wrist transitions continuously increase and exceed a preset threshold within the same time window are marked as upper limb abrupt transition segments. Synchronously compare the torso support change segment with the displacement change segment in the stationary movement characteristics, and combine the distance sequence change from the approaching personnel to the firearm to mark the time period when the torso center shift and the lower limbs move backward consecutively as the forced retreat segment. The entry length sequence of external hands entering the gun control area and the entry length sequence of external hands entering the door buffer area are checked window by window, and the time period when the entry length continuously increases is marked as the hand intrusion segment. Synchronously compare the changes in the physical characteristics of the gun-holding action with the changes in the distance sequence from the gun area to the holster area. Mark the time period in which the changes in the gun-holding action occur continuously and the distance between the gun area and the holster area increases continuously as the gun unholstering segment. The segments of sudden upper limb turning, gun deflection, forced retreat, hand intrusion, and gun disarming were used as segment labeling results.

8. The method for early warning of sudden close-range actions of a vehicle-mounted gun integrated with guard equipment based on computer vision according to claim 1, characterized in that, Step six specifically involves: Read the upper limb sudden turn segment, gun deflection segment, forced retreat segment, hand intrusion segment, and gun disarming segment from the segment marking results in chronological order, and record the start frame, end frame, and object number of each segment. Segments that overlap in time or are consecutive in sequence are merged to form continuous warning segments, and the warning type of each continuous warning segment is determined, specifically including: When a continuous warning segment includes segments of sudden upper limb turning and gun deflection, the corresponding continuous warning segment is identified as a robbery warning type. When a continuous warning segment contains both a probe intrusion segment and a forced retreat segment, the corresponding continuous warning segment is identified as a containment warning type. When a continuous warning segment includes both a weapon deflection segment and a weapon dislodgement segment, the corresponding continuous warning segment is identified as a disengagement warning type. Extract the distance sequence from external approaching personnel to firearms, the entry length sequence of external hands into the firearm control strap, the entry length sequence of external hands into the vehicle door buffer strip, the distance sequence from the firearm area to the holster area, and the change sequence of the escort personnel's position within the time range corresponding to each continuous warning type, and calculate the change amount of each sequence within the time range. The change amount corresponding to each continuous warning type is compared with a preset threshold. When the change amount reaches the preset threshold, the corresponding warning level is generated. Write the warning type, warning level, start frame, end frame, and object number into the warning record to generate warning information.