Surveillance video identification method, medium and system

By extracting and comparing the identification characteristics of personnel between multiple camera devices, the problem of being unable to track personnel between multiple monitoring devices in the prior art is solved, and the effect of simplifying operations and improving monitoring efficiency is achieved.

CN120070513AInactive Publication Date: 2025-05-30YUNNAN SHENGSHANG MECHANICAL & ELECTRICAL ENG CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510135360.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing explosion-proof recording and monitoring devices cannot track personnel between multiple monitoring equipment, resulting in the need of staff to view each equipment one by one in the later stage, which is cumbersome to operate.

Method used

By obtaining video files of multiple cameras, the identification characteristics of the personnel are extracted, including facial movements, eye movements, sounds, height, body shape, body proportions, postures, gaits, etc., and the characteristics are compared to determine whether they are the same person, thereby tracking the personnel.

Benefits of technology

It realizes tracking personnel between multiple monitoring devices, without the need to view each device one by one, simple operation and improves monitoring efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070513A_ABST
    Figure CN120070513A_ABST
Patent Text Reader

Abstract

The invention provides a surveillance video identification method, medium and system, and belongs to the technical field of surveillance, the surveillance video identification method, medium and system comprises the following steps: obtaining a first video file in a first camera; first recognition features of the person in the first video file are extracted, the first recognition features comprise face movement, eyeball movement, sound, height, body shape, body proportion, posture, gait and the like, and the first recognition features can record the identity of the person; acquiring a second video file of a second camera; extracting a second identification feature of the personnel in the second video file, wherein the second identification feature cannot judge the identity of the personnel in the video; and comparing the first identification feature with the second identification feature, and judging whether the personnel in the second video file is the personnel in the first video file or not according to the first identification feature, so that the personnel in the shooting range can be tracked when the plurality of cameras are connected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of monitoring, and more specifically, relates to a method, medium and system for identifying surveillance videos. Background Art

[0002] With the development of technology and social progress, surveillance systems have become an indispensable part of modern society. The installation of surveillance can prevent and reduce criminal activities, protect the safety of people and property, enhance safety awareness and sense of responsibility, provide evidence and investigation support, improve work efficiency and supervision management, providing convenience for people's lives and playing an essential role in daily life.

[0003] Chinese Utility Model Patent with Publication No. CN2503480Y (Application No. CN01270335.4) discloses an explosion-proof video surveillance device. The lid of this explosion-proof video surveillance device is fixedly installed on the housing and forms a sealed box body; a camera and a TV adjacent-channel modulator are fixedly installed inside the housing; the camera is connected to the TV adjacent-channel modulator through a wire; at least two holes are provided on the housing, with glass installed on the wall of one of the holes, and a cable passes through the other hole; one end of the cable is connected to the TV adjacent-channel modulator.

[0004] The above-mentioned explosion-proof video surveillance device cannot track personnel within the monitoring ranges of multiple surveillance devices, and later requires staff to check each device one by one, which is cumbersome to operate. Summary of the Invention

[0005] In view of this, the present invention provides a method, medium and system for identifying surveillance videos, which can track personnel within the monitoring ranges of multiple surveillance devices and do not require later checking of each device one by one, with simple operation.

[0006] The present invention is implemented as follows:

[0007] In a first aspect of the present invention, there is provided a method for identifying surveillance videos, including:

[0008] S10: Obtain a first video file in a first camera;

[0009] S20: Extract first identification features of a person in the first video file, where the first identification features include facial movements, eye movements, voice, height, body shape, body proportion, body posture, gait, etc., and the first identification features can identify the identity of the person in the video;

[0010] S30: Obtain a second video file of a second camera;

[0011] S40: Extract second identification features of a person in the second video file, where the second identification features cannot determine the identity of the person in the video;

[0012] S50: Compare the first identification feature with the second identification feature, and determine whether the person in the second video file is the person in the first video file according to the first identification feature.

[0013] The technical effects of a surveillance video recognition method provided by the present invention are as follows: By obtaining the first video file in the first camera, the first camera is connected to the server, which facilitates the server to obtain the video file in the first camera; By extracting the first identification feature of the person in the first video file, it is convenient to identify the person information in the first video file, so as to track the person in the first video file; By obtaining the second video file of the second camera, the server is connected to the second camera, which facilitates the server to obtain the video file in the second camera; By extracting the second identification feature of the person in the second video file, it is convenient to obtain the feature information of the person in the second video file; By comparing the first identification feature with the second identification feature, and determining whether the person in the second video file is the person in the first video file according to the first identification feature, it is convenient to compare the second identification feature in the second video file with the first identification feature in the first video file to determine whether the person in the second video file is the person in the first video file, so as to track the person in the first video file.

[0014] The first camera and the second camera are cameras with human body recognition functions.

[0015] On the basis of the above technical solution, a surveillance video recognition method of the present invention can also be improved as follows:

[0016] Among them, the specific steps of extracting the first identification feature of the person in the first video file and the second identification feature of the person in the second video file include:

[0017] S21: Extract the person information in the first video file and the second video file, and the person information includes facial movements, eye movements, voice, height, body shape, body proportion, body posture, gait, etc.;

[0018] S22: Clean the person information to remove the noise, invalid pictures and abnormal pictures in the images of the first video file and the second video file;

[0019] S23: Align the person information;

[0020] S24: Extract the features in the person information to convert them into more representative features.

[0021] Further, the specific steps of obtaining the first video file in the first camera and obtaining the second video file of the second camera include

[0022] S11: The server establishes a communication connection with the camera;

[0023] S12: Obtain and download the video files stored in the camera;

[0024] S13: Segment the video file to obtain multiple video segment files;

[0025] S14: Calculate the degree of change of the image frames within each video segment file;

[0026] S15: Perform video compression processing on the video segment files with the degree of change of the image frames less than the threshold;

[0027] S16: Combine multiple video segment files into one video file.

[0028] Further, the step of calculating the degree of change of the image frames within each video segment file specifically includes:

[0029] S141: Extract the image frame sequence of the video segment file;

[0030] S142: Perform uniform processing on the brightness of the image frame sequence;

[0031] S143: Divide each image frame into multiple pixel blocks;

[0032] S144: Combine multiple pixel blocks into color block areas according to colors;

[0033] S145: Establish a three-dimensional coordinate system, where the X-axis is the horizontal axis of the image frame sequence, the Y-axis is the vertical axis of the image frame, and the Z-axis is the time axis, and mark the color block areas of each image frame of the image frame sequence in the three-dimensional coordinate system;

[0034] S146: Orthogonally project the pixel blocks on each image frame onto the first image frame to form a projection area;

[0035] S147: Take the ratio of the area of the projection area to the pixel block as the change rate of the pixel block;

[0036] S148: Take the sum of the change rates of all pixel blocks as the degree of change of the image frame.

[0037] Further, the step of performing video compression processing on the video segment files with the degree of change of the image frames less than the threshold specifically includes:

[0038] S151: Delete some image frames of the video segment files with the degree of change of the image frames less than the threshold;

[0039] S152: Save the remaining image frames as new video segment files.

[0040] Further, the step of deleting partial image frames of the video segment file with an image frame change degree less than the threshold is specifically as follows: deleting the image frames with odd numbers in the video segment file with an image frame change degree less than the threshold and other image frames except the first image frame and the last image frame.

[0041] Further, the step of deleting partial image frames of the video segment file with an image frame change degree less than the threshold is specifically as follows: deleting other image frames of the video segment file with an image frame change degree less than the threshold except the first image frame.

[0042] Further, in the step of merging multiple pixel blocks into a color block area according to colors, the color tolerance for merging pixel blocks is 20-40.

[0043] A second aspect of the present invention provides a computer-readable storage medium, on which computer program instructions are stored; when the computer program instructions are executed by a processor, the above-mentioned monitoring video recording acquisition method is implemented.

[0044] A third aspect of the present invention provides a monitoring video recognition system, including a server and multiple camera devices. The camera devices include a first camera and a second camera. The server is communicatively connected to the camera devices, and the server contains the code of the above-mentioned computer-readable storage medium.

[0045] Compared with the prior art, the beneficial effects of the monitoring video recognition method, medium and system provided by the present invention are as follows: By obtaining the first video file in the first camera, the first camera is connected to the server, which is convenient for the server to obtain the video file in the first camera; By extracting the first recognition feature of the person in the first video file, it is convenient to identify the person information in the first video file, so as to track the person in the first video file; By obtaining the second video file of the second camera, the server is connected to the second camera, which is convenient for the server to obtain the video file in the second camera; By extracting the second recognition feature of the person in the second video file, it is convenient to obtain the feature information of the person in the second video file; By comparing the first recognition feature with the second recognition feature, and judging whether the person in the second video file is the person in the first video file according to the first recognition feature, it is convenient to compare the second recognition feature in the second video file with the first recognition feature in the first video file to judge whether the person in the second video file is the person in the first video file, so as to track the person in the first video file. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Figure 1 It is a flowchart of a method for identifying surveillance videos;

[0048] Figure 2 It is a flowchart of extracting video file features of a method for identifying surveillance videos;

[0049] Figure 3 It is a flowchart of obtaining video files in a camera of a method for identifying surveillance videos;

[0050] Figure 4 It is a flowchart of calculating the change degree of image frames of a method for identifying surveillance videos;

[0051] Figure 5 It is a schematic diagram of a system for identifying surveillance videos;

[0052] In the accompanying drawings, the list of components represented by each reference numeral is as follows: Specific Embodiments

[0053] To make the purpose, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention.

[0054] As Figure 1 shown, it is the first embodiment of a method for identifying surveillance videos provided by the first aspect of the present invention. In this embodiment, it includes:

[0055] S10: Obtain the first video file in the first camera;

[0056] S20: Extract the first identification features of the person in the first video file. The first identification features include facial movements, eye movements, voice, height, body shape, body proportion, posture, gait, etc. The first identification features can identify the identity of the person in the video;

[0057] S30: Obtain the second video file of the second camera;

[0058] S40: Extract the second identification features of the person in the second video file. The second identification features cannot determine the identity of the person in the video;

[0059] S50: Compare the first identification feature with the second identification feature, and determine whether the person in the second video file is the same as the person in the first video file according to the first identification feature.

[0060] By obtaining the first video file in the first camera, the first camera is connected to the server, which facilitates the server to obtain the video file in the first camera; by extracting the first identification feature of the person in the first video file, it is convenient to identify the person information in the first video file, so as to track the person in the first video file; by obtaining the second video file of the second camera, the server is connected to the second camera, which facilitates the server to obtain the video file in the second camera; by extracting the second identification feature of the person in the second video file, it is convenient to obtain the feature information of the person in the second video file; by comparing the first identification feature with the second identification feature, and determining whether the person in the second video file is the same as the person in the first video file according to the first identification feature, it is convenient to compare the second identification feature in the second video file with the first identification feature in the first video file to determine whether the person in the second video file is the same as the person in the first video file, so as to track the person in the first video file.

[0061] Such as Figure 2 shown, in the above technical solution, the specific steps of extracting the first identification feature of the person in the first video file and the second identification feature of the person in the second video file include:

[0062] S21: Extract the person information in the first video file and the second video file, including facial movements, eye movements, voice, height, body shape, body proportion, body posture, gait, etc.;

[0063] S22: Clean the person information to remove noise, invalid pictures and abnormal pictures in the images of the first video file and the second video file;

[0064] S23: Align the person information;

[0065] S24: Extract the features in the person information to convert them into more representative features.

[0066] Such as Figure 3 shown, further, in the above technical solution, the specific steps of obtaining the first video file in the first camera and the second video file of the second camera include

[0067] S11: The server establishes a communication connection with the camera;

[0068] S12: Obtain the video file stored in the camera and download it;

[0069] S13: Segment the video file to obtain multiple video segment files;

[0070] S14: Calculate the change degree of the image frames within each video segment file;

[0071] S15: Perform video compression processing on the video segment files with the change degree of the image frames less than the threshold;

[0072] S16: Combine the multiple video segment files into one video file.

[0073] As Figure 4 shown, further, in the above technical solution, the step of calculating the change degree of the image frames within each video segment file specifically includes:

[0074] S141: Extract the image frame sequence of the video segment file;

[0075] S142: Perform uniform processing on the brightness of the image frame sequence;

[0076] S143: Divide each image frame into multiple pixel blocks;

[0077] S144: Merge the multiple pixel blocks into color block areas according to colors;

[0078] S145: Establish a three-dimensional coordinate system, where the X-axis is the horizontal axis of the image frame sequence, the Y-axis is the vertical axis of the image frame, and the Z-axis is the time axis, and mark the color block areas of each image frame of the image frame sequence in the three-dimensional coordinate system;

[0079] S146: Orthogonally project the pixel blocks on each image frame onto the first image frame to form a projection area;

[0080] S147: Take the area ratio of the projection area to the pixel blocks as the change rate of the pixel blocks;

[0081] S148: Take the sum of the change rates of all pixel blocks as the change degree of the image frame.

[0082] Among them, the key frame is one or several frames of images that reflect the main information content in a group of shots, and can concisely express the shot content. When browsing or retrieving video materials, the content can be quickly located through the non-linear browsing of the key frames. Therefore, applying the key frame technology can well browse and search for video materials.

[0083] The existing video key frame extraction methods can be mainly divided into the following three types:

[0084] The first one is the specified position method. This method does not consider the specific content of the video and the changing trend of the video, but uses a relatively fixed position as the key frame. For example, after determining the start and end points of the shot, directly take the first frame, the last frame, the middle frame, or the frame closest to the average value of all frames as the key frame. Although this method is simple to operate and fast to calculate, and can obtain key frames in real time, it cannot ensure that there is at least one key frame for all important segments in the video, nor can it ensure the representativeness of the key frame to the shot content.

[0085] The second one is the method of analyzing significant content changes within the shot. This method processes the video sequence sequentially and only focuses on the degree of significant changes in the video on the time axis. The first key frame usually takes the first frame of the shot. Traverse all frames in order. When the change reaches a certain degree (that is, reaches the threshold), the frame that reaches the threshold is taken as the next key frame. For example, in the method published in the article Detection and Representation of Scenes in Videos[J]IEEE Transaction on Multimedia, vol.7, no.6, 1097 - 1105, 2005., start searching backward from the previous reference frame until a frame whose distance to the reference frame is greater than the threshold is found, and then take the previous frame of this frame as the new key frame. Then start searching backward from the key frame until a frame whose distance to this new key frame is greater than the threshold is found, and take the previous frame of this frame as the next reference frame. The key frames obtained in this way represent all the frames between the previous reference frame and the next reference frame. However, the key frames extracted by this method have a great relationship with the starting position and the threshold setting. If the cumulative change method is used, even a long video with very small changes will generate more key frames. Therefore, the representativeness of the key frames to the important segments of the video may not be sufficient. Moreover, due to the cumulative change, the result of key frame extraction is also related to the direction of processing the video, resulting in different results when processing the video from back to front and from front to back.

[0086] The third method is to divide the frames of the video shot into several categories through clustering analysis, select the points closest to the cluster center to represent the points of the cluster, and finally form a set of key frames for the video sequence. However, for the current main clustering methods, such as applying methods like fuzzy C - means clustering, the similarity between clusters is relatively low, and it cannot effectively make the similarity within the cluster large enough, which cannot ensure that the extracted key frames have good representativeness to the shot content.

[0087] The ViBe algorithm uses neighboring pixels to create a background model and detects the foreground by comparing the background model and the current input pixel value. Generally speaking, it is a background subtraction method that first models the background and then detects the foreground.

[0088] The specific idea is to store a sample set for each pixel. The sampled values in the sample set are the pixel values of the pixel itself in the past and the pixel values of its neighboring points. Then, each new pixel value is compared with the sample set to determine whether it belongs to the background points.

[0089] The ViBe algorithm can be divided into the following three steps:

[0090] 1. Initialization of the background model

[0091] Initialize the background model for each pixel in a single-frame image. Assume that the pixel values of each pixel and its neighboring pixels have a similar distribution in the spatial domain. Based on this assumption, each pixel model can be represented by the pixels in its neighborhood. To ensure that the background model conforms to statistical laws, the neighborhood range should be large enough. For a pixel point, randomly select the pixel values of its neighboring points as its model sample values.

[0092] 2. Update the background model

[0093] For each pixel in the input frame, match it with the background samples at the corresponding position. If the current pixel matches at least one sample successfully, increment the counter. If the value of the counter is greater than a predefined threshold (usually a small fraction of the number of matches), mark the pixel as a background pixel and update a random sample in the background sample set. If the value of the counter does not exceed the threshold, mark the pixel as a foreground pixel.

[0094] 3. Foreground detection process

[0095] For each pixel, generate a two-dimensional image as a foreground mask according to whether it is marked as background or foreground. Morphological operations or other processing can be further performed on the foreground mask as needed to remove noise or fill holes. The extracted foreground region can be obtained by performing a bitwise AND operation at the pixel level with the original input image.

[0096] Furthermore, in the above technical solution, the steps of performing video compression processing on the video segment files with the degree of change of image frames less than the threshold specifically include:

[0097] S151: Delete some image frames of the video segment files with the degree of change of image frames less than the threshold;

[0098] S152: Save the remaining image frames as new video segment files.

[0099] Furthermore, in the above technical solution, the step of deleting some image frames of the video segment files with the degree of change of image frames less than the threshold is specifically: Delete the image frames with odd numbers in the video segment files with the degree of change of image frames less than the threshold and other image frames except the first image frame and the last image frame.

[0100] Further, in the above technical solution, the step of deleting partial image frames of a video segment file with an image frame change degree less than a threshold is specifically: deleting other image frames of the video segment file with an image frame change degree less than the threshold except the first image frame.

[0101] Further, in the above technical solution, in the step of merging multiple pixel blocks into a color block area according to colors, the color tolerance for merging pixel blocks is 20-40.

[0102] As Figure 5 shown, it is the first embodiment of a surveillance video recognition system provided by the third aspect of the present invention. In this embodiment, it includes a server and multiple camera devices. The camera devices include a first camera and a second camera. The server is communicatively connected to the camera devices, and the server contains the code of the above computer-readable storage medium.

[0103] Specifically, the principle of the present invention is: by obtaining the first video file in the first camera, the first camera is connected to the server, which facilitates the server to obtain the video file in the first camera; by extracting the first recognition feature of the person in the first video file, it is convenient to identify the person information in the first video file, so as to track the person in the first video file; by obtaining the second video file of the second camera, the server is connected to the second camera, which facilitates the server to obtain the video file in the second camera; by extracting the second recognition feature of the person in the second video file, it is convenient to obtain the feature information of the person in the second video file; by comparing the first recognition feature with the second recognition feature, it is convenient to judge whether the person in the second video file is the person in the first video file according to the first recognition feature in the first video file, so as to compare the second recognition feature in the second video file with the first recognition feature in the first video file to judge whether the person in the second video file is the person in the first video file, thereby tracking the person in the first video file.

Claims

1. A surveillance video recognition method, characterized in that: include: S10: Obtain a first video file in a first camera; S20: extracting a first identification feature of the person in the first video file, wherein the first identification feature includes facial movements, eye movements, voice, height, body shape, body proportions, posture, gait, etc. The first identification feature may be used to identify the person in the video file; S30: Obtain a second video file of the second camera; S40: extracting a second identification feature of the person in the second video file, where the second identification feature cannot determine the identity of the person in the video; S50: Compare the first identification feature with the second identification feature, and determine whether the person in the second video file is the person in the first video file according to the first identification feature.

2. A surveillance video recognition method according to claim 1, characterized in that: The specific steps of extracting the first identification feature of the person in the first video file and extracting the second identification feature of the person in the second video file include: S21: extracting character information from the first video file and the second video file, the character information including facial movements, eye movements, voice, height, body shape, body proportions, posture, gait, etc.; S22: Cleaning the character information to remove noise, invalid images and abnormal images in the first video file and the second video file; S23: aligning the character information; S24: Extract features from the character information and convert them into more representative features.

3. The surveillance video recognition according to claim 2, characterized in that: The specific steps of obtaining the first video file in the first camera and obtaining the second video file in the second camera include: S11: The server establishes a communication connection with the camera; S12: Obtain the video file stored in the camera and download it; S13: Segment the video file to obtain multiple video segment files; S14: Calculate the degree of change of the image frames in each video segment file; S15: Performing video compression processing on the video segment files whose image frame change degree is less than the threshold; S16: Combining multiple video segment files into one video file.

4. A surveillance video recognition method according to claim 3, characterized in that: The step of calculating the degree of change of the image frames in each video segment file specifically includes: S141: extracting the image frame sequence of the video segment file; S142: performing uniform processing on the brightness of the image frame sequence; S143: Divide each image frame into a plurality of pixel blocks; S144: merging multiple pixel blocks into a color block area according to color; S145: establishing a three-dimensional coordinate system, wherein the X axis is the horizontal axis of the image frame sequence, the Y axis is the vertical axis of the image frame, and the Z axis is the time axis, and marking the color block area of ​​each image frame in the image frame sequence in the three-dimensional coordinate system; S146: positively projecting the pixel blocks on each image frame onto the first image frame to form a projection area; S147: taking the area ratio of the projection area to the pixel block as the change rate of the pixel block; S148: The sum of the change rates of all pixel blocks is taken as the degree of change of the image frame.

5. A surveillance video recognition method according to claim 4, characterized in that: The step of performing video compression processing on the video segment files whose image frame change degree is less than the threshold specifically includes: S151: deleting some image frames of the video segment file whose image frame change degree is less than a threshold; S152: Save the remaining image frames as a new video segment file.

6. A surveillance video recognition method according to claim 5, characterized in that: The step of deleting some image frames of the video segment file whose image frame change degree is less than the threshold is specifically: deleting the odd-numbered image frames of the video segment file whose image frame change degree is less than the threshold and other image frames except the first image frame and the last image frame.

7. A surveillance video recognition method according to claim 6, characterized in that: The step of deleting some image frames of the video segment file whose image frame change degree is less than the threshold is specifically: deleting other image frames of the video segment file whose image frame change degree is less than the threshold except the first image frame.

8. A surveillance video recognition method according to claim 7, characterized in that: In the step of merging the plurality of pixel blocks into a color block area according to color, the color tolerance of the merged pixel blocks is 20-40.

9. A computer-readable storage medium having computer program instructions stored thereon; when the computer program instructions are executed by a processor, the monitoring video collection method according to any one of claims 1 to 8 is implemented.

10. A surveillance video recognition system, comprising a server and a plurality of camera devices, wherein the camera devices comprise a first camera and a second camera, the server and the camera devices are communicatively connected, and the server contains the code of the computer-readable storage medium as claimed in claim 9.

Citation Information

Patent Citations

  • Explosion camera control monitor

    CN2503480Y