Climbing behavior warning method and device, electronic equipment, and storage medium
By using video image data processing and a lightweight detection model to identify climbing behavior, the problem of low accuracy in climbing behavior early warning in scenic area monitoring systems has been solved, and efficient climbing behavior management has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2021-07-22
- Publication Date
- 2026-05-15
AI Technical Summary
The existing scenic area monitoring system has poor accuracy in detecting uncivilized behaviors such as climbing sculptures, and security personnel are prone to fatigue, resulting in low early warning efficiency.
By acquiring video image data, the system identifies and tracks the behavior of objects entering the target area, uses a lightweight detection model and spatiotemporal relationships to determine climbing behavior, marks climbing video frames, and generates early warning information.
It improved the accuracy of early warning and management efficiency of climbing behavior, reduced the fatigue of security personnel, and enabled the timely detection of uncivilized behavior.
Smart Images

Figure CN115917589B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a climbing behavior early warning method and device, electronic device, and storage medium. Background Technology
[0002] As the number of tourists in scenic areas increases, so too does uncivilized behavior, such as graffiti on cultural relics and climbing sculptures. Taking climbing sculptures as an example, it can damage the sculpture, injure the tourist, and negatively impact other visitors.
[0003] To promptly detect and address such uncivilized behaviors, existing scenic areas typically install video surveillance systems, with security personnel monitoring the monitor screens in real time to ensure timely detection of such behavior.
[0004] However, security personnel are prone to fatigue when monitoring multiple scenes at the same time, and since uncivilized behavior is an occasional phenomenon, the accuracy of early warnings is relatively poor. Summary of the Invention
[0005] This disclosure provides a climbing behavior early warning method and device, electronic device, and storage medium to address the shortcomings of related technologies.
[0006] According to a first aspect of the present disclosure, a climbing behavior early warning method is provided, the method comprising:
[0007] Acquire video image data, the video image data including the detected target and at least one object;
[0008] When it is determined that the object has entered the target area corresponding to the detected target, the behavior information of the object is obtained;
[0009] When the behavioral information is determined to indicate that the object is climbing the detected target, the video frame in which the object is located is marked.
[0010] Optionally, determining that the object enters the target region corresponding to the detected target includes:
[0011] The target region where the detected target is located is obtained from multiple video frames of the video image data, and the object region where the target object is located is obtained; the head of the target object is located within the target region.
[0012] Obtain the spatiotemporal relationship between the object region and the target region; the spatiotemporal relationship refers to the relative spatial position of the object region and the target region at different times;
[0013] When it is determined that the spatiotemporal relationship satisfies the first preset condition, it is determined that the target object enters the target area;
[0014] The first preset condition includes at least one of the following: the object region is within the target region and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the object region successively touches the edge of the target region and two marker lines and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; wherein the two marker lines are set between the edge of the target region and the detected target.
[0015] Optionally, the spatiotemporal relationship includes at least one of the following:
[0016] The following conditions are met: the object region is within the target region; the object region successively touches the edge and two marker lines of the target region; the object region successively touches the two marker lines and the edge of the target region; the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the distance between the bottom edge of the object region and the bottom edge of the target region is less than a set distance threshold; or the object region is outside the target region.
[0017] Optionally, the object region containing the target object is obtained, including:
[0018] Obtain the position of the head of each object and the object region where each object is located within multiple video frames in the video image data;
[0019] Select the object whose head is located within the target area as the target object, and obtain the object area where the target object is located.
[0020] Optionally, obtaining the position of the head of each object within multiple video frames in the video image data includes:
[0021] Obtain preset image features of each video frame within the multiple video frames;
[0022] Based on the preset image features, the identification position of the head in the current video frame is identified, and the predicted position of the head in the next video frame is predicted.
[0023] The identified position and the predicted position are matched, and when the match is successful, the predicted position is updated to the identified position to obtain the position of the same head in two adjacent video frames.
[0024] Optionally, obtaining the object's behavior information includes:
[0025] The location of key parts of the target object's behavior information in multiple video frames of the video image data is obtained; the head of the target object is located within the target area; the behavior information includes human posture.
[0026] According to the preset order of presentation, generate a one-dimensional vector for the key parts of the behavioral information in each video frame;
[0027] The corresponding one-dimensional vectors in each video frame are concatenated to obtain an RGB image; the RGB channels in the RGB image correspond to the xyz axis coordinates of each key part of the behavioral information.
[0028] The behavioral information of the target object is obtained from the RGB image.
[0029] Optionally, determining the behavioral information characterizing the object climbing the detected target includes:
[0030] The location of a specified part of the target object is determined based on the behavioral information; the behavioral information includes human posture.
[0031] When the location of the specified part is within the target area and the distance between it and the bottom edge of the target area exceeds a set distance threshold, the behavioral information is determined to represent the target object climbing the detected target.
[0032] Optionally, after marking the video frame where the object is located, the method further includes:
[0033] Obtain the facial image of the target object;
[0034] When the facial image meets preset requirements, an identification code matching the facial image is obtained; the preset requirements include obtaining key points of the face and the confidence level of the identification result exceeding a set confidence level threshold.
[0035] When it is determined that no object matching the identification code exists in the specified database, an early warning message is generated.
[0036] According to a second aspect of the present disclosure, a climbing behavior warning device is provided, the device comprising:
[0037] The data acquisition module is used to acquire video image data, which includes the detected target and at least one object;
[0038] The information acquisition module is used to acquire the behavior information of the object when it is determined that the object has entered the target area corresponding to the detected target;
[0039] The video tagging module is used to tag the video frame where the object is located when the behavioral information is determined to represent the object climbing the detected target.
[0040] Optionally, the information acquisition module includes:
[0041] The region acquisition submodule is used to acquire the target region where the detected target is located in multiple video frames of the video image data, and to acquire the object region where the target object is located; the head of the target object is located within the target region;
[0042] The relationship acquisition submodule is used to acquire the spatiotemporal relationship between the object region and the target region; the spatiotemporal relationship refers to the relative spatial position of the object region and the target region at different times;
[0043] The region determination submodule is used to determine that the target object enters the target region when the spatiotemporal relationship is determined to meet the first preset condition;
[0044] The first preset condition includes at least one of the following: the object region is within the target region and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the object region successively touches the edge of the target region and two marker lines and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; wherein the two marker lines are set between the edge of the target region and the detected target.
[0045] Optionally, the spatiotemporal relationship includes at least one of the following:
[0046] The following conditions are met: the object region is within the target region; the object region successively touches the edge and two marker lines of the target region; the object region successively touches the two marker lines and the edge of the target region; the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the distance between the bottom edge of the object region and the bottom edge of the target region is less than a set distance threshold; or the object region is outside the target region.
[0047] Optionally, the region acquisition submodule includes:
[0048] The location acquisition unit is used to acquire the position of the head of each object in multiple video frames in the video image data and the object region where each object is located.
[0049] The object selection unit is used to select an object whose head is located within the target area as the target object, and to obtain the object area where the target object is located.
[0050] Optionally, the location acquisition unit includes:
[0051] The feature acquisition subunit is used to acquire preset image features of each video frame within the multiple video frames;
[0052] The position prediction subunit is used to identify the head's position in the current video frame based on the preset image features, and to predict the head's position in the next video frame.
[0053] The location acquisition subunit is used to match the identified location and the predicted location, and when the match is successful, update the predicted location to the identified location to obtain the location of the same head in two adjacent video frames.
[0054] Optionally, the information acquisition module includes:
[0055] The location acquisition submodule is used to acquire the location of key parts of the target object's behavioral information in multiple video frames of the video image data; the head of the target object is located within the target area; the behavioral information includes human posture.
[0056] The vector generation submodule is used to generate one-dimensional vectors from key parts of behavioral information in each video frame according to a preset expression order.
[0057] The image acquisition submodule is used to concatenate the corresponding one-dimensional vectors in each video frame to obtain an RGB image; the RGB channels in the RGB image correspond to the xyz axis coordinates of each key part of the behavioral information.
[0058] The behavior information acquisition submodule is used to acquire the behavior information of the target object based on the RGB image.
[0059] Optionally, the video tagging module includes:
[0060] The location determination submodule is used to determine the location of a specified part of the target object based on the behavioral information; the behavioral information includes human posture.
[0061] The target determination submodule is used to determine the behavioral information representing the target object climbing the detected target when the location of the specified part is within the target area and the distance between it and the bottom edge of the target area exceeds a set distance threshold.
[0062] Optionally, the device further includes:
[0063] The image acquisition module is used to acquire facial images of the target object;
[0064] The identification code acquisition module is used to acquire an identification code that matches the facial image when the facial image meets preset requirements; the preset requirements include acquiring key points of the face and the confidence level of the recognition result exceeding a set confidence level threshold;
[0065] The signal generation module is used to generate a warning message when it is determined that no object matching the identification code exists in the specified database.
[0066] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0067] processor;
[0068] Memory for storing computer programs executable by the processor;
[0069] The processor is configured to execute a computer program in the memory to implement the method described above.
[0070] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when an executable computer program in the storage medium is executed by a processor, enables the implementation of the above-described method.
[0071] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0072] As can be seen from the above embodiments, the solution provided in this disclosure can acquire video image data; the video image data includes a detected target and at least one object; when it is determined that the object enters the target area corresponding to the detected target, the behavior information of the object is acquired; when it is determined that the behavior information indicates that the object is climbing the detected target, the video frame where the object is located is marked. Thus, by marking video frames in the video image data, this embodiment can promptly detect the behavior of an object climbing the detected target, improving management efficiency.
[0073] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0074] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0075] Figure 1 This is a flowchart illustrating a climbing behavior early warning method according to an exemplary embodiment.
[0076] Figure 2 This is a flowchart illustrating the determination of the current behavior of a target object according to an exemplary embodiment.
[0077] Figure 3 This is a flowchart illustrating the tracking of the same head according to an exemplary embodiment.
[0078] Figure 4 This is a flowchart illustrating the current behavior of a target object according to an exemplary embodiment.
[0079] Figure 5 This is a schematic diagram illustrating the effect of acquiring the target object according to an exemplary embodiment.
[0080] Figure 6 This is a flowchart illustrating, according to an exemplary embodiment, a process for determining whether behavioral information characterizes an object climbing a detected target.
[0081] Figure 7This is a schematic diagram illustrating the spatiotemporal relationship between an object region and a target region according to an exemplary embodiment.
[0082] Figure 8 This is a flowchart illustrating another climbing behavior warning method according to an exemplary embodiment.
[0083] Figure 9 This is a block diagram illustrating a climbing behavior warning device according to an exemplary embodiment. Detailed Implementation
[0084] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described below by way of example do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatus consistent with some aspects of this disclosure as detailed in the appended claims.
[0085] To address the aforementioned technical problems, this disclosure provides a climbing behavior early warning method applicable to electronic devices. Figure 1 This is a flowchart illustrating a climbing behavior early warning method according to an exemplary embodiment. See also... Figure 1 A climbing behavior early warning method includes steps 11 to 13.
[0086] In step 11, video image data is acquired, which includes the detected target and at least one object.
[0087] In this embodiment, the electronic device can be connected to a camera and receive video image data output by the camera. That is, when the camera is turned on, it can capture video frames to form a video frame stream, then encode and compress the video frames before sending them to the electronic device. The electronic device can obtain the aforementioned video image data by decoding the received image data.
[0088] Given that the solution provided in this disclosure is to monitor certain target behaviors, such as uncivilized behaviors like climbing and graffiti, the shooting range of the aforementioned cameras is usually directed at the designated target to be detected. The target to be detected may include, but is not limited to, statues in scenic spots, cultural relics in museums, safety railings, etc., or the video image data acquired by the electronic device may include the target to be detected.
[0089] It is understood that video image data may or may not include objects, where objects may be tourists or management personnel. Considering that the solution provided in this disclosure applies to scenarios including objects, subsequent embodiments only consider scenarios where the video image data includes at least one object.
[0090] In step 12, when it is determined that the object has entered the target area corresponding to the detected target, the behavior information of the object is obtained.
[0091] In this embodiment, the electronic device can process the aforementioned video image data to determine whether an object has entered the target area corresponding to the detected target. See [link to relevant documentation]. Figure 2 This includes steps 21 to 23.
[0092] In step 21, the electronic device can acquire the target region where the detected target is located in the multiple video frames of the video image data, and acquire the object region where the target object is located.
[0093] Taking the acquisition of the target region as an example, the electronic device can pre-store a target recognition model, such as a convolutional network (CNN) model. The electronic device can input each video frame of the video image data into the target recognition model, which can identify the detected target in each video frame of the video image data. Then, it generates a minimum bounding rectangle based on the shape of the detected target. The area in the video frame corresponding to this minimum bounding rectangle is the target region. In other words, the target region where the detected target is located in multiple video frames can be obtained through the above recognition process. It is understood that the above minimum bounding rectangle can also be replaced by other preset shapes, such as circles, rhombuses, etc. If the target region can be obtained, the corresponding solution falls within the protection scope of this disclosure.
[0094] Taking object region acquisition as an example, electronic devices can pre-store head detection models, such as convolutional network models. In this example, the head detection model is a lightweight CNN-based model, which is suitable for scenarios with low resource configuration of electronic devices or for upgrading existing monitoring systems. Thus, by setting the aforementioned lightweight detection model in this example, the recognition performance is maintained while reducing the number of parameters of the lightweight detection model, and the detection results can have a high confidence level.
[0095] In this example, the lightweight detection model can utilize model compression and model pruning. Model compression involves compressing the parameters of the already trained model, reducing the number of model parameters and thus minimizing memory usage, thereby improving processing efficiency.
[0096] Model pruning refers to retaining important weights and removing unimportant weights while maintaining the accuracy of a CNN. Generally, the closer a weight's value is to 0, the less important it is. Model pruning can include: 1. Modifying the blob structure or not, directly defining a diagonal mask, and rewriting the original matrix into a sparse matrix storage method; 2. Using a new method to calculate the multiplication of sparse matrices and vectors. In other words, there are two starting points for pruning: one is starting from the blob and modifying it, storing the diagonal mask within the blob structure. The blob-based approach allows the diagonal mask-related operations to be run directly on the CPU or GPU, resulting in higher efficiency. The second approach is starting from the layer and directly defining the diagonal mask; this method is simpler but relatively less efficient.
[0097] It should be noted that when setting the pruning rate, you can either set a global pruning rate or set a pruning rate for each layer individually. In practical applications, the actual value of the pruning rate can be obtained through experimentation.
[0098] It should also be noted that, generally speaking, removing unimportant weights will decrease the model's accuracy. However, removing unimportant weights increases the model's sparsity, which can reduce overfitting, and the model's accuracy can also be improved after fine-tuning.
[0099] When performing pruning, there are two starting points: one is to start from the blob, modify the blob, and store the diagonal mask in the blob structure; the other is to start from the layer and directly define the diagonal mask. Each method has its own advantages. The blob-based approach allows the diagonal mask-related operations to be performed directly on the CPU or GPU, resulting in higher efficiency, but requires a strong understanding of the source code. The layer-based approach is simpler, but relatively less efficient.
[0100] This disclosure allows for optimization of the confidence level in the aforementioned lightweight detection model. For example, firstly, the confidence threshold for the head is gradually reduced from a preset value (e.g., 0.7) until the recall of the head detection results exceeds the recall threshold. Then, combining the tracking results of the head tracking model with the aforementioned detection results, focusing on the recall and precision of the same head, the confidence threshold for the head is further adjusted (fine-tuned) until the recall of the same head exceeds the recall threshold and the precision exceeds the precision threshold; for example, both the recall and precision thresholds are set to exceed 0.98. Thus, by optimizing the confidence level of the head in this example, a good recall and precision can be achieved for the same head during target object tracking, ultimately achieving a balance between recall and precision.
[0101] In this example, the electronic device can input each video frame into the lightweight detection model. This model can detect the heads of objects in each video frame, such as heads from various angles including front, back, side, and top. Based on the one-to-one correspondence between heads and objects, and combined with the shape of the objects, it generates a minimum bounding rectangle and the object region where each object is located. In other words, the electronic device can obtain the position of the head of each object and the object region where each object is located within multiple video frames in the video image data. Then, the electronic device can combine the above target regions to select objects whose heads are located within the target regions as target objects, and at the same time, select the object region corresponding to the minimum bounding rectangle of the target object, thus obtaining the object region where the target object is located.
[0102] Understandably, the head detection model described above can detect the heads of objects in each video frame, but it cannot determine whether the heads in two adjacent video frames belong to the same object. Therefore, the process of obtaining the head position within each video frame by the electronic device can include obtaining the positions of the same object's head in different video frames, see [reference needed]. Figure 3 This includes steps 31 to 33.
[0103] In step 31, for each video frame in the multi-video frame, the electronic device can acquire preset image features of the current video frame, such as color features or histogram of oriented gradients (HARQ) features. The preset image features can be selected according to the specific scene. As long as the preset image features can effectively distinguish the heads of different objects and reduce computational complexity, it falls within the protection scope of this disclosure. It is understood that by reducing computational complexity in this step, the resource requirements of the electronic device under this disclosure can be reduced, which is beneficial to expanding the application scope of this disclosure.
[0104] In step 32, the electronic device can identify the head's position in the current video frame based on the preset image features. Step 32 can be implemented using the lightweight detection model described above, which will not be elaborated further here. In this step, the lightweight detection model can quickly identify the head's position, which is beneficial for achieving real-time detection.
[0105] In step 32, the electronic device can also predict the predicted position of the head in the next video frame of the current video frame. For example, the electronic device can use a fast tracking processing video frame based on a Kalman filter model to preset the position and speed of the head movement. It should be noted that since this example only focuses on the predicted position of the head, how to utilize the speed of movement is not described in detail. It can be processed according to the requirements of the Kalman filter model, and the corresponding solution falls within the protection scope of this disclosure.
[0106] In step 33, the electronic device can match the identified position and the predicted position. This matching can be achieved using the cosine distance of feature vectors. For example, if the cosine value of the feature vectors corresponding to the identified and predicted positions exceeds a cosine threshold (which can be set, such as 0.85 or higher), the identified and predicted positions are considered to have matched. After successful matching, the electronic device can update the predicted position to the identified position, obtaining the position of the same head in the current and next video frames. In this example, by tracking the same head, object loss can be avoided, which helps improve detection accuracy.
[0107] For example, the process of head tracking by electronic devices is as follows:
[0108] Video frame 0: The head detection model detected 3 head detections in Frame 0. There are currently no tracks, so these 3 detections are initialized as tracks.
[0109] Video Frame 1: The head detection model detected 3 more detections; for the tracks in Frame 0, first predict to obtain new tracks; then, match the new tracks with the detections, the matching model can include using the Hungarian model, so as to obtain (track, detection) matching pairs; finally, update the corresponding track with the detection in each matching pair.
[0110] In step 22, the electronic device can acquire the spatiotemporal relationship between the object region and the target region; the spatiotemporal relationship refers to the relative spatial position of the object region and the target region at different times.
[0111] In this embodiment, the electronic device can set two marker lines inside the target area, wherein the first marker line is closer to the edge of the target area than the second marker line, that is, the second marker line is located between the first marker line and the target being detected. The principle is as follows:
[0112] (1) Identify the situation where the object enters or exits the target area directly vertically by setting two horizontal marker lines at the top edge of the target area;
[0113] (2) By setting two vertical marker lines on the left side of the target area, the situation of the object entering and exiting the target area from the left is identified;
[0114] (3) By setting two vertical marker lines on the right side of the target area, the situation of the object entering and exiting the target area from the right and parallel can be identified;
[0115] (4) By setting a horizontal line at the bottom edge of the target area, the distance between the object and the ground can be identified, thereby distinguishing whether the object is passing by the detected target or may climb the detected target.
[0116] In this embodiment, the electronic device can determine the spatiotemporal relationship between the object region and the target region based on two marker lines. This spatiotemporal relationship refers to the relative spatial positions of the object region and the target region at different times. Specifically, the spatiotemporal relationship includes at least one of the following: the object region is within the target region; the object region successively touches the edge of the target region and both marker lines; the object region successively touches both marker lines and the edge of the target region; the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the distance between the bottom edge of the object region and the bottom edge of the target region is less than a set distance threshold; or the object region is outside the target region.
[0117] Taking the object region entering the target region as an example, over time, the object region will move from the outside of the target region to the inside of the target region. That is, the object region will first "touch" the first marker line, and then "touch" the second marker line. Taking the object region leaving the target region as an example, over time, the object region will move from the inside of the target region to the outside of the target region. That is, the object region will first "touch" the second marker line, and then "touch" the first marker line.
[0118] In step 23, when it is determined that the spatiotemporal relationship meets the first preset condition, the electronic device can determine that the current behavior of the target object does not belong to the target behavior.
[0119] In this embodiment, the electronic device can pre-store a first preset condition, which includes at least one of the following: the object area is within the target area and the distance between the bottom edge of the object area and the bottom edge of the target area exceeds a set distance threshold; the object area successively touches the edge of the target area and two marker lines and the distance between the bottom edge of the object area and the bottom edge of the target area exceeds a set distance threshold; wherein the two marker lines are set between the edge of the target area and the detected target. The first preset condition can be set according to the specific scenario. When it can be determined that the target object is passing by the detected target and does not constitute uncivilized behavior, the corresponding solution falls within the protection scope of this disclosure.
[0120] In this embodiment, the electronic device can determine whether the spatiotemporal relationship determined in step 22 satisfies the first preset condition. When the spatiotemporal relationship satisfies the first preset condition, the electronic device can determine that the current behavior of the target object does not belong to the target behavior, that is, the target object is a passing detected target. When the spatiotemporal relationship does not satisfy the first preset condition, that is, satisfies the second preset condition, the electronic device can determine that the current behavior of the target object may belong to the target behavior. At this time, the electronic device can obtain the behavior information of the object entering the target area. It is understood that this behavior information includes at least human posture. See also Figure 4 This includes steps 41 to 44.
[0121] In step 41, for each video frame in the multi-video image data, the electronic device can obtain the location of key parts of the target object's behavior information in each video frame. For example, the electronic device can pre-store a keypoint extraction model, and then input each video frame into the keypoint extraction model, which can then extract the key points of the target object in each video frame. These key points may include the left arm bone points, right arm bone points, left leg bone points, right leg bone points, and torso bone points.
[0122] In step 42, the electronic device can generate a one-dimensional vector from the key parts of the behavioral information in each video frame according to a preset presentation order. The one-dimensional vector can be found in [reference needed]. Figure 5 The vectors below the second and third rows of figures shown are, for example, [63,64,97,103,121,124]. The order of these representations can include at least one of the following: left arm bone points, right arm bone points, left leg bone points, right leg bone points, and torso bone points; left arm bone points, right arm bone points, torso bone points, left leg bone points, and right leg bone points; left arm bone points, torso bone points, left leg bone points, right arm bone points, and right leg bone points. In other words, adjusting the arrangement of key points for the left and right hands, left and right legs, and torso falls within the protection scope of this disclosure.
[0123] In step 43, the electronic device can concatenate the corresponding one-dimensional vectors in each video frame of the video image data to obtain a frame of RGB image; the RGB channels in the RGB image correspond to the xyz axis coordinates of each key part of the behavioral information.
[0124] In step 44, the electronic device can acquire behavioral information of the target object based on the RGB image. In one example, the electronic device can classify behavioral information based on a 3D skeleton point-based behavioral information detection method, including: behavioral information representation based on keypoint coordinates (effect as shown). Figure 5 As shown in the first row of the figure), including spatial descriptors (effect as shown in the figure). Figure 5 The leftmost graphic in the third row (as shown in the image) represents a geometric descriptor (the effect is as follows). Figure 5 (As shown in the middle image of the third row), keyframe descriptor (effect as shown) Figure 5 (As shown in the rightmost figure in the third row of the middle section); after considering the correlation of key points in the subspace to improve the discriminative power and the matching degree of different video sequences based on the dynamic programming model, the behavioral information of the target object can finally be obtained.
[0125] In step 13, when it is determined that the behavioral information represents the object climbing the detected target, the video frame in which the object is located is marked.
[0126] In this embodiment, after determining the behavioral information of the target object, the electronic device can determine whether the behavioral information indicates that the object is climbing the detected target. See [link to relevant documentation]. Figure 6 The process includes steps 61 and 62. In step 61, the electronic device can determine the position of a specified part of the target object based on the behavioral information. Taking the specified part as the object's leg as an example, after determining the target object's movement, the positions of the target object's left and right legs can be determined. See also... Figure 7 In the sculpture located in the center, the right leg of the target object on the left is within the target area; the left and right legs of the target object on the right are also within the target area; both legs of the target object on the left are within the target area. It should be noted that in practical applications, the target area does not need to be displayed. Figure 7 The edges of the target area are represented by dashed lines to facilitate understanding of the scheme disclosed herein. In step 62, when the location of the specified part is within the target area and the distance from the bottom edge of the target area exceeds a set distance threshold, the electronic device can determine that the above behavioral information indicates that the target object is climbing the detected target.
[0127] Understandably, when a target object passes by a detected target, the bottom edge of the target object's area and the bottom edge of the target area should theoretically overlap, meaning the distance is 0. However, considering the target object's walking motion, its legs may lift to a certain height, potentially causing the bottom edge of the target object's area to be slightly higher than the target area's. Therefore, there is a certain distance (e.g., 10-30cm, adjustable) between the bottom edges of the target and the target areas. This distance threshold is set to eliminate the influence of the object passing by the detected target. Alternatively, when a specified part is located within the target area and the distance to the bottom edge of the target area exceeds the set distance threshold, the electronic device can determine that the target object is climbing the detected target.
[0128] In this embodiment, when it is determined that a target object is climbing the detected target, the video frame containing the target object is marked. In some examples, when marking the corresponding video frame, the facial image of the target object can also be extracted and associated with the video frame. This allows managers to see the facial image while reviewing the video frame, thus enabling timely identification of the target object. In this way, by marking video frames in the video image data, this embodiment can promptly detect preset target behaviors (i.e., uncivilized behavior), improving management efficiency.
[0129] In one embodiment, after step 13, the electronic device may also generate a warning signal, see [link to relevant documentation]. Figure 8 This includes steps 81 to 83.
[0130] In step 81, the electronic device can acquire a facial image of the target object. This facial image can be acquired simultaneously during the head recognition process, or after determining that the target object's current behavior is the target behavior. It is understood that not all objects within the target area need to have their behavior determined; therefore, the latter requires fewer facial images than the former, thus reducing the amount of data processing.
[0131] In step 82, when the facial image meets preset requirements, the electronic device can acquire an identification code matching the facial image; the preset requirements include obtaining key points of the face and the confidence level of the recognition result exceeding a set confidence threshold. For example, the electronic device can acquire attribute information of an area image, wherein the attribute information may include, but is not limited to, gender, age, height, skin color, and the location of key facial points. Then, the electronic device can generate an identification code matching the facial image based on the attribute information and store it in a designated database.
[0132] In step 83, when it is determined that no object matching the aforementioned identification code exists in the designated database, it can be determined that the target object is not a manager but a tourist. At this time, the electronic device can generate a warning message: "If a tourist is climbing the sculpture, please stay alert." Of course, the electronic device can also provide the aforementioned warning message to the appropriate personnel, for example, by notifying the manager via telephone or SMS, or by directly alarming the police.
[0133] As can be seen, by identifying the target object in this embodiment, the scenario where managers use target behaviors to maintain the detected target can be excluded, thereby improving the accuracy of the early warning.
[0134] Based on the climbing behavior early warning method provided in the above embodiments, this disclosure also provides a climbing behavior early warning device, see [link to relevant documentation]. Figure 9 The device includes:
[0135] Data acquisition module 91 is used to acquire video image data, the video image data including the detected target and at least one object;
[0136] Information acquisition module 92 is used to acquire the behavior information of the object when it is determined that the object has entered the target area corresponding to the detected target;
[0137] The video tagging module 93 is used to tag the video frame where the object is located when it is determined that the behavior information represents the object climbing the detected target.
[0138] In one embodiment, the information acquisition module includes:
[0139] The region acquisition submodule is used to acquire the target region where the detected target is located in multiple video frames of the video image data, and to acquire the object region where the target object is located; the head of the target object is located within the target region;
[0140] The relationship acquisition submodule is used to acquire the spatiotemporal relationship between the object region and the target region; the spatiotemporal relationship refers to the relative spatial position of the object region and the target region at different times;
[0141] The region determination submodule is used to determine that the target object enters the target region when the spatiotemporal relationship is determined to meet the first preset condition;
[0142] The first preset condition includes at least one of the following: the object region is within the target region and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the object region successively touches the edge of the target region and two marker lines and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; wherein the two marker lines are set between the edge of the target region and the detected target.
[0143] In one embodiment, the spatiotemporal relationship includes at least one of the following:
[0144] The following conditions are met: the object region is within the target region; the object region successively touches the edge and two marker lines of the target region; the object region successively touches the two marker lines and the edge of the target region; the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the distance between the bottom edge of the object region and the bottom edge of the target region is less than a set distance threshold; or the object region is outside the target region.
[0145] In one embodiment, the region acquisition submodule includes:
[0146] The location acquisition unit is used to acquire the position of the head of each object in multiple video frames in the video image data and the object region where each object is located.
[0147] The object selection unit is used to select an object whose head is located within the target area as the target object, and to obtain the object area where the target object is located.
[0148] In one embodiment, the location acquisition unit includes:
[0149] The feature acquisition subunit is used to acquire preset image features of each video frame within the multiple video frames;
[0150] The position prediction subunit is used to identify the head's position in the current video frame based on the preset image features, and to predict the head's position in the next video frame.
[0151] The location acquisition subunit is used to match the identified location and the predicted location, and when the match is successful, update the predicted location to the identified location to obtain the location of the same head in two adjacent video frames.
[0152] In one embodiment, the information acquisition module includes:
[0153] The location acquisition submodule is used to acquire the location of key parts of the target object's behavioral information in multiple video frames of the video image data; the head of the target object is located within the target area; the behavioral information includes human posture.
[0154] The vector generation submodule is used to generate one-dimensional vectors from key parts of behavioral information in each video frame according to a preset expression order.
[0155] The image acquisition submodule is used to concatenate the corresponding one-dimensional vectors in each video frame to obtain an RGB image; the RGB channels in the RGB image correspond to the xyz axis coordinates of each key part of the behavioral information.
[0156] The behavior information acquisition submodule is used to acquire the behavior information of the target object based on the RGB image. The behavior information includes human posture;
[0157] In one embodiment, the video tagging module includes:
[0158] The location determination submodule is used to determine the location of a specified part of the target object based on the behavioral information.
[0159] The target determination submodule is used to determine the behavioral information representing the target object climbing the detected target when the location of the specified part is within the target area and the distance between it and the bottom edge of the target area exceeds a set distance threshold.
[0160] In one embodiment, the device further includes:
[0161] The image acquisition module is used to acquire facial images of the target object;
[0162] The identification code acquisition module is used to acquire an identification code that matches the facial image when the facial image meets preset requirements; the preset requirements include acquiring key points of the face and the confidence level of the recognition result exceeding a set confidence level threshold;
[0163] The signal generation module is used to generate a warning message when it is determined that no object matching the identification code exists in the specified database.
[0164] It should be noted that the device shown in this embodiment is similar to... Figure 1 The content of the method embodiment shown is consistent with that of the above method embodiment, and will not be repeated here.
[0165] In an exemplary embodiment, an electronic device is also provided, comprising:
[0166] processor;
[0167] Memory for storing computer programs executable by the processor;
[0168] The processor is configured to execute a computer program in the memory to achieve, for example, Figure 1 The steps of the method are described.
[0169] In an exemplary embodiment, a computer-readable storage medium, such as a memory including instructions, is also provided, wherein the executable computer program described above can be executed by a processor to achieve, for example... Figure 1 The steps of the method are described above. The readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0170] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0171] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for early warning of climbing behavior, characterized in that, The method includes: Acquire video image data, the video image data including the detected target and at least one object; When it is determined that the object has entered the target area corresponding to the detected target, the behavior information of the object is obtained; When the behavioral information is determined to represent the object climbing the detected target, the video frame in which the object is located is marked; Determining that the object enters the target region corresponding to the detected target includes: The target region where the detected target is located is obtained from multiple video frames of the video image data, and the object region where the target object is located is obtained; the head of the target object is located within the target region. Based on the target area and two marker lines within the target area, the spatiotemporal relationship between the object area and the target area is obtained; the spatiotemporal relationship refers to the relative spatial position of the object area and the target area at different times; the two marker lines are set as follows: the two marker lines are set inside the target area, and the first marker line is closer to the edge of the target area than the second marker line, so that the second marker line is located between the first marker line and the detected target; When it is determined that the spatiotemporal relationship satisfies the first preset condition, it is determined that the target object enters the target area; The first preset condition includes at least one of the following: the object region is within the target region and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the object region successively touches the edge of the target region and two marker lines and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; wherein the two marker lines are located between the edge of the target region and the detected target.
2. The method according to claim 1, characterized in that, The spatiotemporal relationship includes at least one of the following: The following conditions are met: the object region is within the target region; the object region successively touches the edge and two marker lines of the target region; the object region successively touches the two marker lines and the edge of the target region; the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the distance between the bottom edge of the object region and the bottom edge of the target region is less than a set distance threshold; or the object region is outside the target region.
3. The method according to claim 1, characterized in that, Retrieve the object region containing the target object, including: Obtain the position of the head of each object and the object region where each object is located within multiple video frames in the video image data; Select the object whose head is located within the target area as the target object, and obtain the object area where the target object is located.
4. The method according to claim 3, characterized in that, Obtaining the position of the head of each object within multiple video frames in the video image data includes: Obtain preset image features of each video frame within the multiple video frames; Based on the preset image features, the identification position of the head in the current video frame is identified, and the predicted position of the head in the next video frame is predicted. The identified position and the predicted position are matched, and when the match is successful, the predicted position is updated to the identified position to obtain the position of the same head in two adjacent video frames.
5. The method according to claim 1, characterized in that, Obtaining the behavior information of the object includes: The location of key parts of the target object's behavior information in multiple video frames of the video image data is obtained; the head of the target object is located within the target area; the behavior information includes human posture. According to the preset order of presentation, generate a one-dimensional vector for the key parts of the behavioral information in each video frame; The corresponding one-dimensional vectors in each video frame are concatenated to obtain an RGB image; the RGB channels in the RGB image correspond to the xyz axis coordinates of each key part of the behavioral information. The behavioral information of the target object is obtained from the RGB image.
6. The method according to claim 1, characterized in that, Determining the behavioral information characterizing the object climbing the detected target includes: The location of a specified part of the target object is determined based on the behavioral information; the behavioral information includes human posture. When the location of the specified part is within the target area and the distance between it and the bottom edge of the target area exceeds a set distance threshold, the behavioral information is determined to represent the target object climbing the detected target.
7. The method according to claim 1, characterized in that, After marking the video frame where the object is located, the method further includes: Obtain the facial image of the target object; When the facial image meets preset requirements, an identification code matching the facial image is obtained; the preset requirements include obtaining key points of the face and the confidence level of the identification result exceeding a set confidence level threshold. When it is determined that no object matching the identification code exists in the specified database, an early warning message is generated.
8. A climbing behavior early warning device, characterized in that, The device includes: The data acquisition module is used to acquire video image data, which includes the detected target and at least one object; The information acquisition module is used to acquire the behavior information of the object when it is determined that the object has entered the target area corresponding to the detected target; The video tagging module is used to tag the video frame where the object is located when the behavioral information is determined to represent the object climbing the detected target. The information acquisition module includes: The region acquisition submodule is used to acquire the target region where the detected target is located in multiple video frames of the video image data, and to acquire the object region where the target object is located; the head of the target object is located within the target region; The relationship acquisition submodule is used to acquire the spatiotemporal relationship between the object region and the target region based on the target region and two marker lines within the target region. The spatiotemporal relationship refers to the relative spatial position of the object region and the target region at different times. The two marker lines are set as follows: the two marker lines are set inside the target region, and the first marker line is closer to the edge of the target region than the second marker line, so that the second marker line is located between the first marker line and the detected target. The region determination submodule is used to determine that the target object enters the target region when the spatiotemporal relationship is determined to meet the first preset condition; The first preset condition includes at least one of the following: the object region is within the target region and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the object region successively touches the edge of the target region and two marker lines and the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; wherein the two marker lines are located between the edge of the target region and the detected target.
9. The apparatus according to claim 8, characterized in that, The spatiotemporal relationship includes at least one of the following: The following conditions are met: the object region is within the target region; the object region successively touches the edge and two marker lines of the target region; the object region successively touches the two marker lines and the edge of the target region; the distance between the bottom edge of the object region and the bottom edge of the target region exceeds a set distance threshold; the distance between the bottom edge of the object region and the bottom edge of the target region is less than a set distance threshold; or the object region is outside the target region.
10. The apparatus according to claim 8, characterized in that, The region acquisition submodule includes: The location acquisition unit is used to acquire the position of the head of each object in multiple video frames in the video image data and the object region where each object is located. The object selection unit is used to select an object whose head is located within the target area as the target object, and to obtain the object area where the target object is located.
11. The apparatus according to claim 10, characterized in that, The location acquisition unit includes: The feature acquisition subunit is used to acquire preset image features of each video frame within the multiple video frames; The position prediction subunit is used to identify the head's position in the current video frame based on the preset image features, and to predict the head's position in the next video frame. The location acquisition subunit is used to match the identified location and the predicted location, and when the match is successful, update the predicted location to the identified location to obtain the location of the same head in two adjacent video frames.
12. The apparatus according to claim 8, characterized in that, The information acquisition module includes: The location acquisition submodule is used to acquire the location of key parts of the target object's behavioral information in multiple video frames of the video image data; the head of the target object is located within the target area; the behavioral information includes human posture. The vector generation submodule is used to generate one-dimensional vectors from key parts of behavioral information in each video frame according to a preset expression order. The image acquisition submodule is used to concatenate the corresponding one-dimensional vectors in each video frame to obtain an RGB image; the RGB channels in the RGB image correspond to the xyz axis coordinates of each key part of the behavioral information. The behavior information acquisition submodule is used to acquire the behavior information of the target object based on the RGB image.
13. The apparatus according to claim 8, characterized in that, The video tagging module includes: The location determination submodule is used to determine the location of a specified part of the target object based on the behavioral information; the behavioral information includes human posture. The target determination submodule is used to determine the behavioral information representing the target object climbing the detected target when the location of the specified part is within the target area and the distance between it and the bottom edge of the target area exceeds a set distance threshold.
14. The apparatus according to claim 8, characterized in that, The device further includes: The image acquisition module is used to acquire facial images of the target object; The identification code acquisition module is used to acquire an identification code that matches the facial image when the facial image meets preset requirements; the preset requirements include acquiring key points of the face and the confidence level of the recognition result exceeding a set confidence level threshold. The signal generation module is used to generate a warning message when it is determined that no object matching the identification code exists in the specified database.
15. An electronic device, characterized in that, include: processor; Memory for storing computer programs executable by the processor; The processor is configured to execute a computer program in the memory to implement the method as described in any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, When the executable computer program in the storage medium is executed by a processor, it can implement the method as described in any one of claims 1 to 7.