Region location update method, security system, and computer-readable storage medium
By updating the position of the target area in the security system, the position misalignment of the target area caused by camera posture changes is solved, and more accurate object recognition and false alarms are achieved.
Patent Information
- Application Number
- CN202210770654.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-29
- Filing Date
- 2022-06-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-30
AI Technical Summary
When the camera posture changes in the existing security system, the position of the target area is easily misaligned, resulting in the problem of false alarms in object recognition.
A method for updating the area position is provided, by obtaining the initial coordinate data of the target area in the video image, tracking the position change of the target area, determining whether the camera posture has changed, and updating the initial coordinate data after the change is made, so as to synchronously update the position of the target area.
Ensure that the target area remains unchanged after the camera moves and/or rotates, avoid misidentification and false alarms, and improve identification efficiency and user experience.
Smart Images

Figure CN115063750B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This disclosure claims the priority of the Chinese patent application with the application number 202210474787.X, titled "Region Location Update Method, Security System and Computer - Readable Storage Medium", filed on April 29, 2022. The full text of this patent application is incorporated into the solution of this disclosure by reference. Technical Field
[0003] This disclosure relates to the field of data processing technologies, and in particular, to a region location update method, a security system, and a computer - readable storage medium. Background Art
[0004] With the rapid development of security technologies, security systems are deployed in many key areas. Cameras in the security system can monitor the security area all - day long by collecting and recording videos. Moreover, existing security systems also allow users to plan out Region A as a restricted area in the video image through a web - based interface and conduct key monitoring on the restricted area. Summary of the Invention
[0005] This disclosure provides a region location update method, a security system, and a computer - readable storage medium to solve the deficiencies of related technologies.
[0006] According to the first aspect of the embodiments of this disclosure, a region location update method is provided. The method includes:
[0007] Obtain the initial coordinate data of a target region in a video image;
[0008] Track the position of the target region in each video image according to the initial coordinate data and each video image to obtain an identification result;
[0009] When the identification result contains target coordinate data, determine whether the attitude of the camera has changed;
[0010] After the attitude of the camera has changed, update the initial coordinate data according to the target coordinate data to update the position of the target region in the video image.
[0011] Optionally, obtaining the initial coordinate data of a target region in a video image includes:
[0012] In response to detecting an operation representing the drawing of a target region, obtain the coordinate data of each trigger position;
[0013] Connect each trigger position in sequence to obtain the target region;
[0014] When the shape of the target area is rectangular, the coordinate data of each trigger position is used as the initial coordinate data of the target area; when the shape of the target area is other than rectangular, the minimum bounding rectangle of the other shape is obtained, and the coordinate data of each vertex of the minimum bounding rectangle is used as the initial coordinate data of the target area.
[0015] Optionally, tracking the position of the target area in each video image according to the initial coordinate data to obtain an identification result, including:
[0016] Based on the initial coordinate data, an image of the target area corresponding to the initial coordinate data in the target video image is obtained to obtain a reference image;
[0017] Based on the initial coordinate data, a first tracking image is obtained, where the first tracking image refers to an image containing the target area in each video image after the target video image;
[0018] The reference image and the first tracking image are input into a preset region tracking model to obtain an identification result, where the identification result includes the probability values of each video image containing at least one candidate region and its coordinate data.
[0019] Optionally, the region tracking model includes a Siamese network module, a region candidate network module, and an identification result module;
[0020] The Siamese network module includes an upper branch network and a lower branch network; the network structures and parameters of the upper branch network and the lower branch network are the same; the upper branch network outputs a feature image of a first size, and the lower branch network outputs a feature image of a second size;
[0021] The region candidate network module includes a classification branch network and a regression branch network; the classification branch network is used to distinguish the target and the background according to the feature image of the first size and the feature image of the second size; the regression branch network is used to adjust the position of the candidate region;
[0022] The identification result module includes a class output unit and a coordinate data output unit; the class output unit is connected to the classification branch network and is used to output the probability value of each candidate region; the coordinate data output unit is connected to the regression branch network and is used to output the coordinate data of each candidate region.
[0023] Optionally, the method further includes a step of judging whether the identification result contains target coordinate data, specifically including:
[0024] Obtaining the maximum value of the probability values of the at least one candidate region;
[0025] When the maximum value exceeds a preset probability threshold, determine the candidate region corresponding to the maximum value as the target region tracked in each video image, and obtain the target coordinate data of the target region.
[0026] Optionally, the method further includes:
[0027] When the maximum value is less than the preset probability threshold, determine that the target region is not tracked in each video image.
[0028] Optionally, determining that the target region is not tracked in each video image includes:
[0029] Determine whether the target region in the first video image is located at the vertex of the first video image; the first video image refers to the video image before the video image in which the target region is not tracked.
[0030] When the target region is located at the vertex of the first video image, obtain at least one target pixel point of the target region located within the first video image.
[0031] Obtain a first distance between the at least one target pixel point and the boundary of the first video image.
[0032] When the first distance is less than a preset distance threshold, determine that the type of the target region not tracked in each video image is that the target region has shifted outside the video image.
[0033] Optionally, determining that the target region is not tracked in each video image includes:
[0034] Determine whether the target region in the first video image is located at the boundary of the first video image; the first video image refers to the video image before the video image in which the target region is not tracked.
[0035] When there is a vertex of the target region located at the boundary of the first video image, obtain a second distance between the vertex of the target region far from the boundary and the boundary.
[0036] When the second distance is less than a preset distance threshold, determine that the type of the target region not tracked in each video image is that the target region has shifted outside the video image.
[0037] Optionally, when the target region is not tracked and the target region is within the first video image due to an abnormal tracking model, the method further includes:
[0038] When the target area is not tracked in the respective video images, decrease the tracking matching threshold by a preset step size, and perform the step of obtaining an identification result by tracking the position of the target area in the respective video images according to the initial coordinate data and the respective video images, until it is determined that the target area is tracked in the respective video images or the tracking matching threshold is equal to the first probability threshold, where the first probability threshold refers to the minimum value of the tracking matching threshold.
[0039] Optionally, when the target area is not tracked and the target area is within the first video image due to an abnormal tracking model, the method further includes:
[0040] Generate a plurality of second tracking images with the vertices of the first tracking image corresponding to the respective video images as the centers and the length and width of the first tracking image as the benchmarks, and perform the step of inputting the reference image and the first tracking image into a preset area tracking model.
[0041] Optionally, the method further includes:
[0042] Obtain the distance between preset points of the target area in two adjacent video images;
[0043] When the distance between the preset points is less than the central distance threshold, update it as the coordinate data of the newly identified target area;
[0044] When the distance between the preset points exceeds the central distance threshold, keep the target area of the video image in which the target area is not tracked as the previous frame of the video image or adopt a constructed area; the constructed area refers to the weighted value of the coordinate data of the target area in multiple video images before the video image in which the target area is not tracked.
[0045] Optionally, determining whether the posture of the camera has changed includes:
[0046] Obtain the amount of change in the angle of the camera;
[0047] When the amount of change in the angle satisfies a preset condition, determine that the posture of the camera has changed.
[0048] Optionally, determining whether the posture of the camera has changed includes:
[0049] Obtain the distances between the respective pixel points within the target area in two adjacent video images;
[0050] When the distance of at least one pixel point exceeds the pixel point distance threshold, determine that the posture of the camera has changed.
[0051] Optionally, determining whether the posture of the camera has changed includes:
[0052] Obtain the distance between the preset points in the target area of two adjacent video frames;
[0053] When the distance between the preset points exceeds the central threshold, it is determined that the posture of the camera has changed.
[0054] Optionally, updating the initial coordinate data according to the target coordinate data includes:
[0055] When the shape of the target area is a rectangle, updating the initial coordinate data to the target coordinate data; or,
[0056] When the shape of the target area is other than a rectangle, obtain the relative position data between the preset target area and the minimum circumscribed rectangle; calculate the target recovery data of the target area according to the target coordinate data and the relative position data; update the initial coordinate data to the target recovery data.
[0057] According to the second aspect of the embodiments of the present disclosure, a security system is provided, and the system includes a region configuration module, a region tracking module, an update judgment module, and a coordinate feedback module;
[0058] The region configuration module is configured to obtain the initial coordinate data of the target area in the video image and send the initial coordinate data to the region tracking module;
[0059] The region tracking module is configured to track the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result, and send the target coordinate data to the update judgment module when the identification result includes the target coordinate data;
[0060] The update judgment module is configured to judge whether the posture of the camera has changed, and send the target coordinate data to the coordinate feedback module after the posture of the camera has changed;
[0061] The coordinate feedback module is configured to feedback the target coordinate data to the region configuration module, so that the region configuration module updates the initial coordinate data according to the target coordinate data to update the position of the target area in the video image.
[0062] Optionally, the region configuration module includes:
[0063] A coordinate data acquisition unit, configured to acquire the coordinate data of each trigger position in response to detecting an operation representing drawing a target area;
[0064] A target area acquisition unit, configured to connect each trigger position in sequence to obtain a target area;
[0065] An initial coordinate acquisition unit, configured to use the coordinate data of each trigger position as the initial coordinate data of the target area when the shape of the target area is a rectangle; when the shape of the target area is other than a rectangle, obtain the minimum circumscribed rectangle of the other shape, and use the coordinate data of each vertex of the minimum circumscribed rectangle as the initial coordinate data of the target area.
[0066] Optionally, the area tracking module is configured to track the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result, including:
[0067] Obtain an image of the target area corresponding to the initial coordinate data in the target video image based on the initial coordinate data to obtain a reference image;
[0068] Obtain images of the target area included in each video image after the target video image based on the initial coordinate data to obtain first tracking images corresponding to each video image;
[0069] Input the reference image and the first tracking images into a preset area tracking model to obtain an identification result output by the area tracking model, where the identification result includes probability values and coordinate data of at least one candidate area included in each video image.
[0070] Optionally, the area tracking module is configured to send the target coordinate data to the update judgment module when the identification result includes target coordinate data, including:
[0071] Obtain the maximum value of the probability values of the at least one candidate area;
[0072] When the maximum value exceeds a preset probability threshold, determine that the candidate area corresponding to the maximum value is the target area tracked in each video image, and obtain the target coordinate data of the target area;
[0073] Send the target coordinate data of the target area to the update judgment module.
[0074] Optionally, the area tracking module is further configured to:
[0075] When the maximum value is less than the preset probability threshold, determine that the target area is not tracked in each video image.
[0076] Optionally, the area tracking module is configured to determine that the target area is not tracked in each video image, including:
[0077] Determine whether the target area in the first video image is located at the boundary of the first video image; the first video image refers to the video image before the target area is not tracked.
[0078] When there is a vertex of the target area located at the boundary of the first video image, obtain the second distance between the vertex far from the boundary in the target area and the boundary.
[0079] When the second distance is less than the preset distance threshold, determine that the type of the target area not being tracked in each video image is that the target area has shifted outside the video image.
[0080] Optionally, when the target area is not tracked and the tracking model is abnormal and the target area is within the first video image, after the area tracking module is used to determine that the target area is not tracked in each video image, the area tracking module is further used for:
[0081] When the target area is not tracked in each video image, reduce the tracking matching threshold according to a preset step size, and perform the step of obtaining the recognition result according to the initial coordinate data and tracking the position of the target area in each video image in each video image, until it is determined that the target area is tracked in each video image or the tracking matching threshold is equal to the first probability threshold, and the first probability threshold refers to the minimum value of the tracking matching threshold.
[0082] Optionally, when the target area is not tracked and the tracking model is abnormal and the target area is within the first video image, after the area tracking module is used to determine that the target area is not tracked in each video image, the area tracking module is further used for:
[0083] Generate a plurality of second tracking images with the vertices of the first tracking image corresponding to each video image as the centers and the length and width of the first tracking image as the benchmarks, and perform the step of inputting the reference image and the first tracking image into a preset area tracking model.
[0084] Optionally, the area tracking module is further used for:
[0085] Obtain the distance between preset points of the target area in two adjacent video images.
[0086] When the distance between the preset points is less than the center distance threshold, update it to the coordinate data of the newly recognized target area.
[0087] When the distance of the preset point exceeds the central distance threshold, the video image of the untracked target area is maintained as the target area of the previous frame of video image or a constructed area is adopted; the constructed area refers to the weighted value of the coordinate data of the target area in multiple frames of video images before the video image of the untracked target area.
[0088] Optionally, the update judgment module is used to judge whether the posture of the camera changes, including:
[0089] Obtain the amount of change in the angle of the camera;
[0090] When the amount of change in the angle meets the preset condition, it is determined that the posture of the camera has changed.
[0091] Optionally, the update judgment module is used to judge whether the posture of the camera changes, including:
[0092] Obtain the distances between the pixel points in the target area in two adjacent frames of video images;
[0093] When the distance of at least one pixel point exceeds the pixel point distance threshold, it is determined that the posture of the camera has changed.
[0094] Optionally, the update judgment module is used to judge whether the posture of the camera changes, including:
[0095] Obtain the distance of the preset point in the target area in two adjacent frames of video images;
[0096] When the distance of the preset point exceeds the central threshold, it is determined that the posture of the camera has changed.
[0097] Optionally, the area configuration module includes:
[0098] The first configuration module is used to directly update the initial coordinate data according to the target coordinate data when the shape of the target area is a rectangle; or,
[0099] The second configuration module is used to obtain the relative position data of the preset target area and the minimum bounding rectangle when the shape of the target area is other than a rectangle; calculate the target restoration data of the target area according to the target coordinate data and the relative position data; update the initial coordinate data to the target restoration data.
[0100] According to the third aspect of the embodiments of the present disclosure, a security system is provided, including at least one camera, at least one configuration terminal and a server; the camera is used to collect images and send them to the server; the configuration terminal is used to obtain the initial coordinate data of the target area and send it to the server; the server includes:
[0101] A processor;
[0102] A memory for storing a computer program executable by the processor;
[0103] Wherein, the processor is configured to execute the computer program in the memory to implement the method described in the first aspect.
[0104] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, which can implement the method described in the first aspect when the executable computer program in the storage medium is executed by a processor.
[0105] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0106] As can be seen from the above embodiments, in the solution provided by the embodiments of the present disclosure, the initial coordinate data of the target area in the video image can be obtained; then, the position of the target area in each video image is tracked according to the initial coordinate data and each video image to obtain an identification result; after that, when the identification result contains target coordinate data, it is determined whether the posture of the camera has changed; finally, after the posture of the camera has changed, the initial coordinate data is updated according to the target coordinate data to update the position of the target area in the video image. In this way, the position of the target area in the video image remains unchanged when the posture of the camera does not change, and after the posture of the camera changes, the coordinate data of the target area is updated to the target coordinate data, that is, the position of the target area is synchronously updated after the camera moves and / or rotates, so that the target area will not be misaligned as the camera moves and / or rotates, thereby avoiding the problems of misidentification and false alarm in the subsequent process of identifying the object in the target area, which is beneficial to improving the identification efficiency and further improving the user experience.
[0107] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0109] Figure 1 is a flowchart of a method for updating the position of an area shown according to an exemplary embodiment.
[0110] Figure 2 is a schematic diagram of a target area being a polygon shown according to an exemplary embodiment.
[0111] Figure 3Another schematic diagram with a circular target area shown according to an exemplary embodiment.
[0112] Figure 4 Effect schematic diagram of configuring the target area shown according to an exemplary embodiment.
[0113] Figure 5 Flowchart of obtaining an identification result shown according to an exemplary embodiment.
[0114] Figure 6 Structural schematic diagram of an area tracking model shown according to an exemplary embodiment.
[0115] Figure 7 Flowchart of obtaining target coordinate data shown according to an exemplary embodiment.
[0116] Figure 8 Flowchart of a tracking stability mechanism shown according to an exemplary embodiment.
[0117] Figure 9 Flowchart of another tracking stability mechanism shown according to an exemplary embodiment.
[0118] Figure 10 Schematic diagram of a target area located at the edge of the current video image shown according to an exemplary embodiment.
[0119] Figure 11 Flowchart of obtaining that the target area deviates outside the video image shown according to an exemplary embodiment.
[0120] Figure 12 Flowchart of another obtaining that the target area deviates outside the video image shown according to an exemplary embodiment.
[0121] Figure 13 Flowchart of obtaining the target coordinate data of the target area shown according to an exemplary embodiment.
[0122] Figure 14 Workflow diagram of a security system shown according to an exemplary embodiment.
[0123] Figure 15 Workflow diagram of another security system shown according to an exemplary embodiment.
[0124] Figure 16 Workflow diagram of yet another security system shown according to an exemplary embodiment.
[0125] Figure 17 Effect schematic diagram of obtaining the target area shown according to an exemplary embodiment.
[0126] Figure 18 It is a block diagram of a security system shown according to an exemplary embodiment.
[0127] Figure 19 It is a block diagram of a server shown according to an exemplary embodiment. Detailed implementation manners
[0128] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The exemplary embodiments described below do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims. It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0129] In practical applications, when setting target areas such as restricted areas in a video image, since the objects in the target area will move out of the above target area, at this time, the user needs to adjust the orientation of the camera to monitor the security of different areas. When moving and / or rotating the orientation of the camera, the above target area will change synchronously, that is, the coverage range of the target area changes from area A to area B. At this time, the camera can identify the objects in area B and give an alarm. However, area B is not the expected target area A to be monitored, thus causing false alarms and reducing the user experience.
[0130] To solve the above technical problems, the embodiments of the present disclosure provide a method for updating the area position, which can be applied to a security system. In one example, the security system includes at least one camera and at least one configuration terminal. In another example, the security system includes at least one camera, a server, and at least one configuration terminal. Among them, the configuration terminal can be used as a web configuration terminal to perform corresponding configurations on the video image, such as setting a target area (such as a restricted area) to prevent objects from entering the target area. The server can communicate with any one of the cameras in the security system, and the communication methods include wired methods or wireless methods. Taking the wireless method as an example, it includes but is not limited to Bluetooth, WiFi, Zigbee, etc. The server can obtain the video images (pictures or video images) collected by the camera through the above communication method and distribute them to each configuration terminal for display. Of course, when the processing resources of the camera are sufficient, it can also replace the configuration terminal to distribute the collected images to each configuration terminal for display. That is to say, in the present disclosure, both the camera and the server can execute a method for updating the area position, which can be set according to the specific scenario. In the subsequent embodiments, the scheme of each embodiment is described by taking the camera only collecting images and uploading the images to the server and the server executing a method for updating the area position as an example.
[0131] Figure 1 It is a flowchart of a method for updating regional location shown according to an exemplary embodiment. Refer to Figure 1 , a method for updating regional location, including steps 11 to 14.
[0132] In step 11, initial coordinate data of a target region in a video image is obtained.
[0133] In this embodiment, the configuration terminal can display the video image collected by the camera, and the user can select at least one background in the video image as the target region. The target region is a region corresponding to a part of the video image displayed by the configuration terminal, and is used to determine the recognition range to determine whether an object enters or leaves the target region. Taking the target region as a restricted area as an example, the restricted area is used to determine the area where the object is prohibited from entering. When it is detected that an object enters this area, the security system can give an alarm.
[0134] In this embodiment, the shape of the above-mentioned target region can be a rectangle or other shapes other than a rectangle.
[0135] In one embodiment, when the target region is a rectangle (for example, the user selects a rectangular component), the configuration terminal can obtain the initial coordinate data of the target region, including: when it is detected an operation representing drawing the target region, the configuration terminal can obtain the coordinate data of the trigger position in the current video image. The coordinate data of the above-mentioned trigger position can include the trigger positions of multiple single triggers detected within a preset duration, and can be applicable to the scenario of discrete touch operations. For example, the coordinate data of 4 points, namely point A, point B, point C, and point D, are used as the coordinate data of the trigger position. The coordinate data of the above-mentioned trigger position can include the coordinate data of each position collected at a set period between the position where the user first presses and the position where the user finally bounces up detected by the configuration terminal, and can be applicable to the scenario of continuous touch operations. For example, when pressing down at point A, passing through point B and point C, and finally bouncing up at point D, the configuration terminal collects the coordinate data of points A, B, C, and D at a set period as the coordinate data of the trigger position. When it is detected an operation representing saving the target region, the configuration terminal can obtain the coordinate data of all trigger positions to obtain the initial coordinate data.
[0136] In one embodiment, when the target area is a shape other than a rectangle (for example, the user selects other shape components or does not select components), the configuration terminal can obtain the initial coordinate data of the target area, including: when detecting an operation representing the drawing of the target area, the configuration terminal can obtain the coordinate data of each trigger position in the current video image, and connect the trigger positions in sequence to form a closed candidate area. When the shape of the candidate area is a rectangle, the configuration terminal can use the vertex coordinate data of the candidate area as the initial coordinate data of the target area. When the shape of the candidate area is a shape other than a rectangle, obtain the minimum bounding rectangle of the other shape, and use the coordinate data of each vertex of the minimum bounding rectangle as the initial coordinate data of the target area.
[0137] It should be noted that after using the coordinate data of the minimum bounding rectangle as the initial coordinate data of the target area, it is equivalent to replacing the image in the area where the minimum bounding rectangle is located with the image in the target area as the data to be processed during tracking. To ensure that the target area can be accurately restored after tracking, in this embodiment, during or after determining the initial coordinate data of the target area, the relative coordinate data between the target area and the minimum bounding rectangle can also be obtained, so as to restore the target area when the position of the minimum bounding rectangle is updated, achieving the effect of updating the position of the target area.
[0138] See Figure 2 , the target area 202 in the current video image 201, and the minimum bounding rectangle 203 of the target area 202. Assume that the upper left corner coordinate of the minimum bounding rectangle 203 is (0, 0), and the lower right corner coordinate is (1, 1). At this time, the configuration terminal can calculate the relative coordinate data of each point in the target area with respect to the minimum bounding rectangle. For example, the relative coordinate of the bottommost point is (0.65, 1). See Figure 3 , the target area 302 in the current video image 301. When the target area 302 is a circle / ellipse or a manually drawn continuous irregular shape, after the configuration terminal finds the minimum bounding rectangle 303, it can calculate the relative coordinate data of all discrete points with respect to its minimum bounding rectangle 303. For example, the discrete point set is [(0.5, 0), (0.45, 0.05), …… (0.55, 0.05)].
[0139] In addition, during the process of obtaining the initial coordinate data, when detecting an operation representing the clearing of the target area, the configuration terminal can delete the coordinate data of all trigger positions or the coordinate data of the most recent trigger position. In this way, in this embodiment, manual setting of the target area is allowed, which can ensure the accuracy of the target area and enhance the fun and practicality of human-computer interaction.
[0140] In practical applications, function components are usually set in the configuration interface. When an operation of selecting a function component in the video image is detected, the configuration terminal in the security system can display the configuration interface corresponding to the above function component. The configuration interface may include a brush component and a save component. Refer to Figure 4 , and the configuration interface may include a brush component 21 and a save component 23.
[0141] When an operation of selecting the brush component is detected, the configuration terminal can obtain the trigger position of the brush component in the video image and use this trigger position as a vertex of the target area. At this time, the coordinate data of this trigger position can be obtained.
[0142] The user can repeatedly operate (i.e., click multiple times) within the video image using the brush component, and the configuration terminal can detect multiple trigger positions and the coordinate data of each trigger position. In practical applications, after detecting more than 3 trigger positions, the configuration terminal can connect these trigger positions in sequence to form a closed candidate area and display it within the video image for the user to view.
[0143] In an embodiment, when the configuration terminal detects an operation of selecting the save component, it displays a preset prompt message in the video image to prompt that the configuration of the target area is completed. In an example, the configuration terminal can display a preset prompt message of "Area configuration completed" in the upper left corner of the video image and use an animation effect of fading out within 3 seconds to remind the user, so that the user can determine that the target area is successfully matched and the user experience is improved. Then, the configuration terminal can record the coordinate data of each vertex of the target area in clockwise order based on the center of the target area, and the coordinate data of all vertices of the target area can form the initial position data. Finally, the configuration terminal can upload the above initial position data to the server. In this way, the server can obtain the initial coordinate data of the target area in the video image.
[0144] In another embodiment, the server may also obtain the target coordinate data of the target area after the camera moves and / or rotates, and update it to the above initial coordinate data. The solution for obtaining the target coordinate data will be described in subsequent embodiments and will not be elaborated here. The camera can move and / or rotate in three dimensions. Based on the pivot point of the camera, it can rotate upward, downward, leftward, rightward, clockwise along the optical axis, and counterclockwise along the optical axis. Of course, the camera can also move and / or rotate in seven dimensions. In addition to moving and / or rotating in three dimensions based on the pivot point of the camera, it can also move and / or rotate in four dimensions with respect to the fixed end of the object to which the camera is attached (such as a column), such as moving in the front-back dimension of the fixed end (causing the camera to move along the X-axis), left-right dimension (causing the camera to move along the Y-axis), up-down dimension (causing the camera to move along the Z-axis), and rotating counterclockwise or clockwise around the Z-axis of the column (causing the camera to rotate left or right), etc. It can be understood that no matter in which dimension the camera moves and / or rotates, the target coordinate data is obtained based on the images it captures, which does not affect the implementation of the solution of the present disclosure.
[0145] In step 12, the position of the target area in each video image is tracked according to the initial coordinate data and each video image to obtain an identification result.
[0146] In this embodiment, the configuration terminal can display the video images captured by at least one camera. For example, each camera uploads the captured video images to the server. The server can obtain the configuration information and push the above video images to the configuration terminal specified in the above configuration information. Or rather, the server can determine each video image displayed by the configuration terminal.
[0147] Then, the server can track the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result. See Figure 5 , including steps 31 to 33.
[0148] In step 31, the server can obtain the image of the target area corresponding to the initial coordinate data in the target video image based on the initial coordinate data to obtain a reference image. The target video image refers to the first frame of video image obtained after obtaining the initial coordinate data. For example, after the configuration terminal uploads the updated initial coordinate data to the server, the server will pull the stream after receiving the above initial coordinate data. The first frame of video image pulled at this time is the target video image. After obtaining the target video image, the server can find each vertex corresponding to the initial coordinate data in the target video image, and then connect each vertex (in a clockwise or counterclockwise manner) in sequence to obtain a closed area. The image within this closed area is the reference image.
[0149] In step 32, the server can obtain a first tracking image based on the initial coordinate data, where the first tracking image is an image containing the target area in each video image after the target video image. Assuming the number of the target video image is 1, then the numbers of each frame of video image are n, where n = 2, 3, 4, ……, that is, n is an integer greater than or equal to 2. The server can determine the area corresponding to the initial coordinate data in the video image according to the solution in step 31. It can be understood that the area corresponding to the initial coordinate data at this time may or may not be the same as the position of the target area in the target video image. Therefore, the solution of the present disclosure needs to predict the position of the target area in each frame of video image after the target video image. In this step, the server can generate a larger area including the area corresponding to the above initial coordinate data. For example, when the area corresponding to the initial coordinate data is a rectangle, its length and width can be doubled respectively to obtain a larger rectangle with an area 4 times that of the previous area. Then, the server can use the image within the above larger area as the first tracking image. By repeating the above steps, the server can obtain the first tracking image corresponding to each frame of video image after the target video image.
[0150] In step 33, the server can input the reference image and the first tracking image into a preset region tracking model to obtain an identification result, where the identification result includes the probability values and coordinate data of at least one candidate region in each video image.
[0151] In this step, a preset region tracking model can be stored in the server. The region tracking model has been pre-trained and can track the target area. In this example, the region tracking model includes a Siamese network module, a region candidate network module, and an identification result module. Among them,
[0152] The Siamese network module includes an upper branch network and a lower branch network; the network structures and parameters of the upper branch network and the lower branch network are the same, and the network structures of the upper branch network and the lower branch network do not include an output layer. Therefore, the difference between the upper branch network and the lower branch network is that the upper branch network outputs a feature image of the first size, and the lower branch network outputs a feature image of the second size; the region candidate network module includes a classification branch network and a regression branch network; the classification branch network is respectively connected to the upper branch network and the lower branch network, and is used to distinguish the target and the background according to the feature image of the first size and the feature image of the second size; the regression branch network is respectively connected to the upper branch network and the lower branch network, and is used to adjust the position of the candidate region. The recognition result module includes a category output unit and a coordinate data output unit; the category output unit is connected to the classification branch network and is used to output the probability value of each candidate region; the coordinate data output unit is connected to the regression branch network and is used to output the coordinate data of each candidate region. In one example, the above region tracking model may adopt a region tracking model based on deep features, including but not limited to SiamFC, siamRPN, DaSiamRPN, siamRPN++, etc. In this example, the above region tracking model may adopt the siamRPN++ algorithm.
[0153] See Figure 6 , the left part of the region tracking model is a Siamese network structure 41, and the network structures and parameters of the upper branch network and the lower branch network are exactly the same. Moreover, the input data of the upper branch network is the reference image, and the object to be tracked is determined based on the reference image. Or rather, the feature data of the reference image is obtained as the reference feature data. The input data of the lower branch network is the first tracking image, or the video image to be detected. Obviously, the area of the first tracking image is larger than that of the reference image, that is, the search area of the first tracking image is larger than that of the reference image, so as to ensure that the offset target area is still within the search area. The two branches of the Siamese network structure 41 respectively obtain the feature data of the reference image and the first tracking image, and obtain the similarity of the two feature vectors. The greater the similarity, the more likely it is that the test image and the reference image are of the same category.
[0154] Continue to see Figure 6, the middle part of the region tracking model is the region candidate network 42, which is composed of two branches. The upper branch is a classification branch used to distinguish the target from the background (such as the content within the target region in subsequent embodiments). The feature data of the reference image and the first tracking image after passing through the Siamese network is then passed through a convolutional layer to become 2k * 256 channels, where k is the number of anchor boxes, and 2k means being divided into two categories. The lower branch is a regression branch for fine-tuning the candidate region, which is the bounding box regression branch. Since there are four quantities [x, y, w, h], the number of channels is 4k * 256. Among them, x, y, w, and h respectively refer to the horizontal coordinate offset, vertical coordinate offset, width offset, and height offset of the target region. In practical applications, the lower branch can also output coordinate data based on the above coordinate offset data, which is not limited herein.
[0155] Continue to refer to Figure 6 , the right part of the region tracking model is the tracked target region.
[0156] It should be noted that the concept of the region tracking model in this disclosure is to process the video image sequence collected by the camera, calculate the position of the object (such as a parking space, road, etc. within the restricted area) in the target region (such as the restricted area) of the target video image in each frame of the video image; then, according to the feature values related to the object, associate the same object in the video image sequence to obtain the motion parameters of the object in each frame of the video image and the corresponding relationship between adjacent frames of the object, so as to obtain the motion trajectory of the object. Or rather, the concept of the region tracking model in this disclosure is to find the object existing in the reference image in the first tracking image, and the region where the found object is located is the target region in the first tracking image. In other words, this disclosure takes the immovable object in the physical world corresponding to the target region as the tracking target and combines the principle that the imaging of the above tracking target in the camera is basically unchanged to find the tracking target in each video image and determine the target region corresponding to the tracking target, that is, the target region is found in the first tracking image. It should be noted that when there are some movable objects and immovable objects in the target region, considering that when the proportion of the immovable objects in the target region (area) exceeds the preset proportion threshold (such as 60%), the subsequent preset probability threshold can be adjusted according to the proportion of the movable objects. For example, the larger the corresponding proportion of the movable objects, the smaller the preset probability threshold, so as to select the matching target coordinate data.
[0157] In this embodiment, the server may call the above-mentioned preset region tracking model, input the reference image and the first tracking image into the above-mentioned region tracking model, that is, input the reference image into the upper branch of the Siamese network and input the first tracking image into the lower branch of the Siamese network. Then, the region tracking model may process the above-mentioned reference image and the first tracking image and output an identification result. It can be understood that the above-mentioned identification result includes the probability values and their coordinate data (i.e., the coordinate data of each region) of at least one candidate region in each video image. In this way, the server can obtain the above-mentioned identification result.
[0158] In step 13, when the target coordinate data is included in the identification result, it is judged whether the posture of the camera has changed.
[0159] In this embodiment, after obtaining the identification result, the server may judge whether the above-mentioned identification result includes target coordinate data. Refer to Figure 7 , which includes step 51 and step 52.
[0160] In step 51, the server may obtain the maximum value of the probability values of at least one candidate region in the above-mentioned identification result. For example, the server can directly sort the above-mentioned at least one candidate region to obtain the maximum value. A preset probability threshold can be stored in the server, and the range of the preset probability threshold is 0.6 to 1.0. Then, the server may compare the maximum value with the above-mentioned preset probability threshold to obtain the size relationship between the maximum value and the preset probability threshold.
[0161] In step 52, when the maximum value exceeds the preset probability threshold, the server may determine that the candidate region corresponding to the maximum value is the target region tracked in each video image and obtain the target coordinate data of the target region. When the maximum value is less than the above-mentioned preset probability threshold, the server may determine that the target region has not been tracked in each video image. In this step, by selecting the maximum value and the preset probability threshold to determine whether the target region is tracked, the accuracy of the result can be improved.
[0162] Considering that the region tracking model may not be able to ensure accurate tracking of the target region in all cases. For example, when the video image appears severely blurred or has a color distortion and other abnormal conditions, the region tracking model may fail. Considering that there are two cases where the target region is not tracked in each video image: the first case is that the region tracking model is normal but the target region has shifted outside the video image range; the second case is that the target region is within the video image range but the region tracking model is abnormal. To address the above problems, the present disclosure embodiment also provides a tracking stability mechanism to ensure normal tracking of the target region in the case of an abnormal region tracking model. Refer to Figure 8 and Figure 9 , the server tracks the target region and determines whether the target region is tracked.
[0163] After determining that the target area has been tracked, it is determined whether the above target area is mis-tracked. If it is determined that it is not a mis-track, the server can determine that the above target area is accurate. At this time, the server can determine whether the target area is located at the edge of the current video image. If the target area is not at the edge of the current video image, the target coordinate data of the latest target area is used. If the target area is at the edge of the current video image, cropping and compensation are performed on the target area, that is, the coordinate data of the part of the target area located within the current video image is obtained. Refer to Figure 10 , the server can crop and compensate the target area according to the boundary of the current video image, and determine the coordinate data of the polygon ABCDE as the target coordinate data. If it is determined that it is a mis-track, the server can determine that the target area of the current video image (that is, the video image where the target area has not been tracked) remains the target area of the previous frame of video image, or use a constructed area. The constructed area refers to the weighted value of the coordinate data of the target area in multiple frames of video images before the video image where the target area has not been tracked.
[0164] After determining that the target area has been tracked, it is determined whether the above target area goes out of bounds. Going out of bounds includes going out of bounds from the vertex of the first video image or from the boundary of the first video image. If it is determined that the target area goes out of bounds, it is determined that there is no target area in the video image. If it is determined that the target area does not go out of bounds, it is determined whether a re-search is required. If a re-search is not required, the target area of the previous frame of video image is maintained. If a re-search is required, the tracking matching threshold is lowered or the first tracking image is updated to re-search.
[0165] In an embodiment, in the first case, the server can determine whether the target area goes out of bounds (from the vertex of the first video image). Refer to Figure 11 , which includes steps 71 to step 74.
[0166] In step 71, the server can determine whether the target area in the first video image is located at the vertex of the first video image; the first video image refers to the previous frame of video image before the target area is not tracked. For example, the server can obtain the vertices of the target area in the first video image. It is understandable that when a part of the target area offsets out of the first video image, the target area can be located at the upper left corner / upper right corner / lower left corner / lower right corner of the first video image. At this time, at least one vertex of the target area coincides with the vertex of the first video image. Therefore, the server can determine whether the target area is located at a corner of the video image by judging whether the coordinate data of the upper left vertex of the target area is [0,0], whether the abscissa x of the lower left vertex is 0, whether the ordinate y of the upper right vertex is 0, and whether the abscissa of the lower right vertex is the maximum value of the abscissa value and the ordinate is the maximum value of the ordinate value.
[0167] In step 72, when the target area is located at the vertex of the first video image, the server can obtain at least one target pixel point in the target area that is located within the first video image. That is to say, the server can obtain the target pixel points in that part of the target area within the first video image.
[0168] In step 73, the server can obtain the first distance between the at least one target pixel point and the boundary of the first video image. The first distance from the target pixel point to the boundary of the first video image can be converted into the distance from a point to a line in mathematics. For details, reference can be made to the related technology and will not be elaborated here.
[0169] In step 74, when the first distance is less than the preset distance threshold, the server can determine that the type of the target area not tracked in each video image is that the target area has offset out of the video image (offset out of the first video image), and assign the coordinate data of the target area to a null value. Among them, the range of the above preset distance threshold can be 5 to 20 pixels. In one example, the value of the above preset distance threshold is 10 pixels. Since the coordinate data of the target area is forced to be assigned a null value, it can be determined that the target area has moved out of the boundary when a null value is read subsequently.
[0170] In another embodiment, in the first case, the server can determine whether the target area goes out of bounds (from the boundary of the first video image), see Figure 12 , including steps 81 to 83.
[0171] In step 81, the server can determine whether each vertex of the target area in the first video image is located on the boundary of the first video image; the first video image refers to the previous frame of video image before the target area is not tracked. For example, the server can obtain whether each vertex of the target area in the first video image is located on the boundary of the first video image (i.e., the boundary). It is understandable that to determine whether a certain vertex of the target area is located on the boundary of the first video image, the server can judge whether the abscissa of the upper left vertex of the target area is 0, and whether the abscissa x of the lower left vertex is 0 to determine whether the target area is located on the left boundary of the first video image. For another example, the server can judge whether the ordinate of the upper left vertex of the target area is 0, and whether the ordinate of the upper right vertex is 0, to determine whether the target area is located on the upper boundary of the first video image. For another example, the server can judge whether the abscissa of the upper right vertex of the target area is the maximum value of the abscissa value, and whether the abscissa of the lower right vertex is the maximum value of the abscissa value, to determine whether the target area is located on the right boundary of the first video image. For another example, the server can judge whether the ordinate of the lower left vertex of the target area is the maximum value of the ordinate value, and whether the ordinate of the lower right vertex is the maximum value of the ordinate value, to determine whether the target area is located on the lower boundary of the first video image.
[0172] In step 82, when the target area is located on the boundary of the first video image, the server can obtain the second distance between the vertex of the target area far from the boundary and the boundary. When the server determines that the target area is located on a certain boundary of the first video image, it can obtain the second distance between the vertex far from the boundary and this boundary. The calculation of the second distance can refer to the calculation of the distance from a point to a side in the data, which will not be elaborated here.
[0173] In step 83, when the second distance is less than the preset distance threshold, the server can determine that the type of the target area not tracked in each video image is that the target area has shifted outside the video image (shifted out of the first video image), and assign the coordinate data of the target area to a null value. This preset distance threshold can refer to Figure 7 the content of the embodiments shown.
[0174] In this embodiment, by determining that the target area has shifted outside the video image, it can be determined that the area tracking model can work properly, ensuring the accuracy of the detection result.
[0175] For the case where it has not shifted outside the video image, that is, the second case, the server can judge whether it is necessary to search again.
[0176] In one embodiment, the reason for not matching the target area may be that the tracking matching threshold is relatively large, so no matching target area is searched. At this time, the server can reduce the tracking matching threshold. Suppose the range of the tracking matching threshold is 0.3 to 0.9, and the current tracking matching threshold is 0.6 when determining the unmatched target area. Then the server can reduce the tracking matching threshold according to a preset step size (such as 0.1); then importantly execute step 12, that is, the step of obtaining the recognition result by tracking the position of the target area in each video image according to the initial coordinate data and each video image, and then determine whether there is a target area in the video image; if there is no target area, continue to reduce the tracking matching threshold, and repeat this way until it is determined that the target area is tracked in each video image or the tracking matching threshold is equal to the first probability threshold. The first probability threshold refers to the minimum value of the tracking matching threshold, that is, the minimum reference value for the recognition result of the above-mentioned region tracking model to output a credible or effective candidate area.
[0177] In another embodiment, considering that the area of the first video image may be relatively small, which is 4 times that of the reference image, there is a certain probability that the target area cannot be searched. In this embodiment, the search range can be updated. For example, the server can generate multiple second tracking images with the vertices of each first tracking image corresponding to each video image as the center and the length and width of the first tracking image as the reference, such as 2 to 4, and execute the step of inputting the reference image and the first tracking image into a preset region tracking model. In this way, the number of second tracking images in this embodiment is much larger than that of the first tracking images, which can increase the search range of the target area, thereby increasing the probability of searching for the target area. It should be noted that when solving the above two situations where the target area is not tracked according to the above multiple solutions, if the target area of the current video image still cannot be tracked, the server can use the target area of the previous frame of the video image for the target area of the current video image, thereby avoiding the problem of false tracking and being beneficial to improving the accuracy of the tracking result.
[0178] It can be understood that in the embodiments of the present disclosure, by judging whether the target area deviates out of the video image or whether the area recognition model is abnormal, it is ensured that the area position update method provided by the present disclosure can work reliably and ensure the accuracy of tracking the target area.
[0179] In one embodiment, after determining that the target area is tracked in each video image, the server can judge whether there is false tracking, see Figure 13 , including step 91 to step 93.
[0180] In step 91, the server can obtain the distance between the preset points in the target area of two adjacent video images. The preset points can be set according to the target area, such as the vertices, center points, centroids, etc. of the target area, which are not limited here. For example, when the target area is a regular figure, such as a rectangle, the preset point can be the center point. When the target area is an irregular image, the preset point can be one of the vertices of the target area. The distance between two preset points can be converted into the Euclidean distance between two points in mathematics. For the specific calculation method of the Euclidean distance, reference can be made to the relevant technology and will not be elaborated here.
[0181] In step 92, when the distance between the preset points is less than the center distance threshold, the server can update the coordinate data of the newly recognized target area. The range of the center distance threshold is 1 to 10 pixels. In one example, the above center distance threshold is 5 pixels. When the distance between the preset points is less than the center distance threshold, the server can determine that the target area in the video image is not mis-tracked. At this time, the server can adopt the latest target coordinate data for the target area of the video image.
[0182] It should be noted that the above center distance threshold is related to the acquisition frequency of the camera. When the acquisition frequency of the camera is larger, the center distance threshold is smaller. For example, when the acquisition frequency of the camera is 25 Hz, the center distance threshold can be set to 10 pixels. When the acquisition frequency of the camera is 50 Hz, the center distance threshold can be set to 5 pixels. Technicians can set it according to the specific scenario, which is not limited here.
[0183] In step 93, when the distance between the preset points exceeds the center distance threshold, the server can keep the target area of the previous frame of the video image for the video image where the target area is not tracked or adopt a constructed area; the constructed area refers to the weighted value of the coordinate data of the target area in multiple frames of video images before the video image where the target area is not tracked. For example, the server can record the offsets of the target area in the x and y directions at least 5 times. The offsets of the historical restricted area in the x direction for five times are [1, 2, -1, 0, 1], and the historical offsets in the y direction are [1, 1, -1, 1, 0]; then, by taking the average value, the offset of the current video image relative to the target area in the previous video image is predicted. The offset of the target area of the current video image is [1, 0] (the result after rounding [0.6, 0.4]). Combining with the coordinate data of the reference image, the target coordinate data of the target area in the current video image can be obtained.
[0184] In step 14, after the posture of the camera changes, update the initial coordinate data according to the target coordinate data to update the position of the target area in the video image.
[0185] In this embodiment, after determining the target area and target coordinate data in each video image, the server can determine whether the posture of the camera has changed. For example, the server can obtain the amount of change in the angle of the camera. The server and the camera can communicate to obtain the amount of change in the movement and / or rotation angle of the camera, and determine the amount of change in the angle based on the amount of change in the movement and / or rotation angle. For example, when the amount of change in the movement and / or rotation angle is 0, it is determined that the camera is in a stationary state, and when the amount of change in the movement and / or rotation angle is a certain value not equal to 0, it is determined that the posture of the camera has changed. Among them, the above preset condition means that the camera moves from a stationary state to a moving state and then to a stationary state and remains stationary for a certain period of time (such as 30 to 100 seconds), or the amount of change in the angle exceeds a preset angle threshold (such as 5 degrees).
[0186] For another example, the server can obtain the distances of each pixel point in the target area in two adjacent video images. Then, the server can compare the distances of each pixel point with a preset pixel point distance threshold. If the distances of at least one pixel point exceed the pixel point distance threshold, the server can determine that the posture of the camera has changed; if the distances of all pixel points are less than the pixel point distance threshold, the server can determine that the posture of the camera has not changed.
[0187] For another example, the server can obtain the distance of a preset point in the target area in two adjacent video images. When the distance of the preset point exceeds the center threshold, it is determined that the posture of the camera has changed; when the distance of the preset point is less than the center threshold, the server can determine that the posture of the camera has not changed.
[0188] In this embodiment, after determining that the posture of the camera has changed, the server can update the initial coordinate data according to the target coordinate data. For example, when the shape of the target area is a rectangle, the server can update the initial coordinate data to the target coordinate data. For another example, when the shape of the target area is other than a rectangle, the server can obtain the relative position data of the preset target area and the minimum circumscribed rectangle. Among them, the acquisition method of the above-mentioned relative position data of the preset target area and the minimum circumscribed rectangle can be referred to in steps 11 and Figure 2 and Figure 3 the content shown in, which will not be elaborated here. Then, the server can calculate the target restoration data of the target area according to the target coordinate data and the relative position data; update the initial coordinate data to the above target restoration data.
[0189] That is to say, the server can obtain a new initial coordinate data by updating the initial coordinate data, and re-execute steps 11 to 14 to update the position of the target area in the video image.
[0190] So far, in this embodiment, the target area in the video image remains unchanged in position when the camera does not move and / or rotate, and after the camera moves and / or rotates, the coordinate data of the target area is updated to the target coordinate data, that is, the position of the target area is synchronously updated after the camera rotates, so that the target area will not be misaligned as the camera moves and / or rotates, thereby avoiding the problems of mis-identification and false alarm during the subsequent process of identifying the object in the target area, which is beneficial to improving the identification efficiency and further enhancing the user experience.
[0191] Next, a method for updating the area position provided by the embodiments of the present disclosure will be described in combination with the scene of restricted area intrusion recognition, where the restricted area is the above-mentioned target area. Refer to Figures 14 to 16 , a security system provided by the embodiments of the present disclosure may include an area configuration module, an area tracking module, an update judgment module, and a coordinate feedback module. Among them,
[0192] The area configuration module may include displaying a video image, manually configuring a restricted area, automatically receiving the restricted area configuration sent by the coordinate feedback module, and sending the restricted area coordinates to the area tracking module.
[0193] Area configuration module
[0194] The web page in the area configuration module can display the video image of the camera to be configured. In practical applications, refer to Figure 4 , the web page may include three interactive operation buttons: a brush component 21, an eraser component 22, and a save component 23. The user can click on the brush component 21 to draw the restricted area point by point, and the eraser component 22 can be used to erase the drawn vertices during the drawing process. After the drawing is completed, the save component 23 can be clicked to obtain the coordinate data of all vertices of the restricted area, that is, the initial coordinate data in the above embodiment. Specifically, refer to Figure 1 the content of step 11 shown in the example. The area configuration module can send the restricted area coordinates to the area tracking module. In this way, an operation of manually configuring the restricted area is completed.
[0195] In addition, the area configuration module can wait in real time to receive the latest restricted area coordinates sent by the coordinate feedback module, that is, the target coordinate data. When the target coordinate data is received, the area update module can update the initial coordinate data of the restricted area to the above-mentioned target coordinate data, and send the updated initial coordinate data to the matching tracking module. At the same time, the restricted area is redrawn in the displayed video image according to the updated coordinate data, as shown by the restricted area A1A2A3A4 in Figure 17 .
[0196] Area tracking module
[0197] The working process of the area tracking module can be referred to Figure 8 and Figure 9The content of the example. And the area tracking model uses a tracking method based on depth features to track the restricted area using the restricted area coordinate data sent by the area configuration module and the pulled video stream.
[0198] The specific process of restricted area tracking is as follows:
[0199] First, the area tracking model obtains the restricted area coordinates from the area configuration module and obtains the latest video image, i.e., the above-mentioned target image.
[0200] Then, the area tracking model uses the content within the restricted area box in the latest video image as a template, i.e., the above-mentioned reference image, and extracts the features of the template; and sends the content of the area that is twice the size of the template in each frame of the video image after the latest video image (i.e., the above-mentioned first tracking image) and the template into the Siamese network. Then, the Siamese network sends the extracted features into the classification branch and regression branch of the Region Proposal Network (RPN) respectively; the classification branch outputs the probability of each region belonging to the background and the target (i.e., the content within the restricted area), and the regression branch outputs the predicted [x, y, w, h] offset values of each region (i.e., the above-mentioned target coordinate data).
[0201] Finally, the area tracking module can take the region with the maximum probability value as the tracked restricted area; if the probability values of all regions are less than the preset probability threshold, it can be determined that there is no restricted area in this frame of video image.
[0202] Tracking stability mechanism
[0203] The area tracking model cannot guarantee accurate tracking of the restricted area in all cases. For example, in cases of severely blurred or pixelated images, the area tracking model will fail. The present disclosure also provides a tracking stability mechanism to ensure normal tracking of the restricted area in case of abnormal conditions of the area tracking model, which can improve tracking stability.
[0204] First, it is determined whether the area tracking model has tracked the restricted area. If the restricted area has not been tracked, there are two cases at this time: the area tracking model is normal but the restricted area has shifted outside the video image; the restricted area has not shifted out of the boundary of the video image but the area tracking model is abnormal. The specific principle of the tracking stability mechanism includes:
[0205] It is determined whether the restricted area in the previous frame is located in the upper left corner / upper right corner / lower left corner / lower right corner of the video image. By determining whether the upper left corner of the restricted area is [0, 0]; whether the abscissa x of the lower left corner is 0; and whether the y coordinate of the upper right corner is 0, it is determined whether the restricted area is located in the upper left corner of the video image. The judgment of other corner points is similar. If the restricted area is located at a corner point, it is determined whether the pixel points of the restricted area that are not on the image boundary are less than 10 pixel values away from the boundary of the video image. If less, it is determined that the restricted area has shifted outside the video image.
[0206] Determine whether the restricted area in the previous frame is located at the boundary of the video image by judging whether the abscissa x of the upper left corner of the restricted area is 0; whether the abscissa x of the lower left corner is 0 to determine whether the restricted area is located at the left boundary of the video image. The same applies to other boundaries. If the restricted area is located at the boundary of the video image, judge whether the distances from the other two vertices of the restricted area that are not at the video image boundary to the boundary of the video image are less than 10 pixel values. If less, it is judged that the restricted area has shifted outside the video image.
[0207] For the case where it has shifted outside the video image, directly judge that there is no restricted area in the current video image and assign the restricted area information in the current video image to be empty.
[0208] For the case where it has not shifted outside the video image, it is necessary to judge whether re-search is needed. The embodiments of the present disclosure provide two re-search methods:
[0209] 1. Reduce the tracking matching threshold. Since it has been judged that the restricted area is in the video image but not found, it may be caused by a relatively high tracking matching threshold setting. At this time, the tracking matching threshold in the region tracking model can be reduced and re-searched. Assume that the current tracking matching threshold is 0.6 and the minimum value of the tracking matching threshold is 0.3, and the preset step size is 0.1. If it is necessary to reduce the tracking matching threshold, reduce the tracking matching threshold by 0.1 and re-judge in the region tracking model. If the restricted area is not tracked, reduce the tracking matching threshold again and re-judge until the minimum value of the tracking matching threshold is reached or the restricted area is tracked.
[0210] 2. Change the search area. Since the search area of the region tracking model is an area that is twice the size of the area where the current template is located, there is a probability that the restricted area cannot be searched. At this time, the search area can be changed and re-searched, including:
[0211] Taking the four vertices of the current search area (i.e., the above-mentioned first tracking image) as the centers and using the length and width of the current search area, 4 new search areas can be constructed to obtain the above-mentioned second tracking image). Then, the above four second tracking images are sent into the region tracking model one by one and the restricted area is re-tracked.
[0212] If the restricted area of the current video image still cannot be searched by the above two methods, at this time, the restricted area of the current video image remains the restricted area of the previous frame of the video image, that is, the coordinate data of the restricted area of the current video image uses the coordinate data of the previous frame of the video image, so as to avoid the problem of inaccurate tracking results caused by inaccurate region tracking models and is beneficial to improving the accuracy of tracking results.
[0213] Then, judge whether there is mis-tracking in the tracking result, including:
[0214] Determine whether the distance between the preset points of the restricted area in the current video image and the restricted area in the previous video image is greater than the preset point threshold (such as 5 pixels). If it is less, it can be determined that there is no mis-tracking. Then, determine whether the target area is located at the edge of the current video image. When the target area is located at the edge of the current video image, crop and compensate the restricted area (such as filling the edge to form a closed area for the part within the current video image), match and restore the coordinates of the restricted area, that is, the latest coordinate data can be adopted; if it is greater, it is determined as mis-tracking. The processing methods for mis-tracking can include:
[0215] 1. Keep the restricted area of the previous video image;
[0216] 2. Take a constructed area. For example, obtain the offsets of the restricted area in the x and y directions for five historical frames, and predict the offset of the restricted area in the current video image relative to the restricted area in the previous video image by taking the average. For example, if the offset in the x direction is [1, 2, -1, 0, 1] and the offset in the y direction is [1, 1, -1, 1, 0], then the offset of the restricted area in the current video image is [0.6, 0.4], and after rounding, it is [1, 0].
[0217] Update the judgment module
[0218] The area tracking module will obtain the coordinate data of the restricted area in each frame of the video image. Since the area tracking model needs to be processed in real time, the position of the target box of the restricted area will be predicted for each frame of the video image, but the coordinate data of not every frame of the video image needs to be passed back to the area configuration module. Then, the methods for the area tracking module to judge whether to pass back the coordinate data include:
[0219] 1. Judge whether to pass back the coordinates of the restricted area by judging the movement and / or rotation state of the camera. The update judgment module obtains the movement and / or rotation angle of the camera in real time. After judging that the movement and / or rotation state changes to a stationary state and after a set duration, it can be determined that the camera has undergone a movement and / or rotation once, and at this time, the coordinate data of the restricted area needs to be updated. In addition, to improve the accuracy of this method, a timed update of the coordinates can also be set in this method, such as forcibly updating the coordinates of the restricted area every hour.
[0220] 2. Judge whether to pass back the coordinates of the restricted area by judging the relative position change amount of the coordinates of the restricted area box in the image. If the camera moves and / or rotates, there will be a deviation between the target coordinate data and the initial coordinate data of the restricted area. Therefore, by comparing the deviation of the coordinates of the restricted area, it can be judged whether the camera moves and / or rotates, including: comparing each pixel point in the two restricted areas point by point, and when the position of a pixel point exceeds the pixel point distance threshold (such as 5 pixels), it is determined that the posture of the camera has changed. Or calculate the distance between the preset points of the two restricted areas. If the preset point distance exceeds the center threshold, it can be determined that the posture of the camera has changed.
[0221] Coordinate transmission module
[0222] When the update judgment module determines that the restricted area coordinates need to be transmitted (that is, the target coordinate data of the restricted area is transmitted), the coordinate transmission module will obtain the target coordinate data of the restricted area in the current video image and send the target coordinate data to the area configuration module. The area configuration module receives
[0223] After the target coordinate data, it can update the value of the initial coordinate data to the value of the above target coordinate data, and update the restricted area in the display interface and send it to the area tracking module. Thus, an automatic update of the restricted area coordinates is completed.
[0224] Based on the area position update method provided in the embodiments of the present disclosure, the embodiments of the present disclosure further provide a security system. See Figure 18 , the system includes: an area configuration module 131, an area tracking module 132, an update judgment module 133, and a coordinate transmission module 134;
[0225] The area configuration module 131 is used to obtain the initial coordinate data of the target area in the video image and send the initial coordinate data to the area tracking module;
[0226] The area tracking module 132 is used to track the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result, and send the target coordinate data to the update judgment module when the identification result contains the target coordinate data;
[0227] The update judgment module 133 is used to judge whether the posture of the camera has changed, and send the target coordinate data to the coordinate transmission module 134 after the posture of the camera has changed;
[0228] The coordinate transmission module 134 is used to transmit the target coordinate data back to the area configuration module 131, so that the area configuration module 131 updates the initial coordinate data according to the target coordinate data to update the position of the target area in the video image.
[0229] In one embodiment, the area tracking module is used to track the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result, including:
[0230] Based on the initial coordinate data, obtain the image of the target area corresponding to the initial coordinate data in the target video image to obtain a reference image;
[0231] After obtaining the target video image based on the initial coordinate data, obtain the images containing the target area in each subsequent video image of the target video image, and obtain the first tracking image corresponding to each video image;
[0232] Input the reference image and the first tracking image into a preset region tracking model, and obtain the recognition result output by the region tracking model. The recognition result includes the probability values and coordinate data of at least one candidate region in each video image.
[0233] In one embodiment, when the recognition result contains target coordinate data, the region tracking module is used to send the target coordinate data to the update judgment module, including:
[0234] Obtain the maximum value of the probability values of the at least one candidate region;
[0235] When the maximum value exceeds a preset probability threshold, determine the candidate region corresponding to the maximum value as the target region tracked in each video image, and obtain the target coordinate data of the target region;
[0236] Send the target coordinate data of the target region to the update judgment module.
[0237] In one embodiment, the region tracking module is further used for:
[0238] When the maximum value is less than the preset probability threshold, determine that the target region is not tracked in each video image.
[0239] In one embodiment, when the region tracking module determines that the target region is not tracked in each video image, it includes:
[0240] Determine whether the target region in the first video image is located at the vertex of the first video image; the first video image refers to the previous video image before the video image in which the target region is not tracked;
[0241] When the target region is located at the vertex of the first video image, obtain at least one target pixel point in the target region that is located within the first video image;
[0242] Obtain the first distance between the at least one target pixel point and the boundary of the first video image;
[0243] When the first distance is less than a preset distance threshold, determine that the situation where the target region is not tracked in each video image is of the type that the target region has shifted outside the video image, and assign the coordinate data of the target region as a null value.
[0244] In one embodiment, the area tracking module is used to determine that the target area is not tracked in the video images, including:
[0245] Determine whether each vertex of the target area in the first video image is located on the boundary of the first video image; the first video image refers to the previous frame of video image before the video image where the target area is not tracked.
[0246] When the target area is located on the boundary of the first video image, obtain the second distance between the vertex of the target area far from the boundary and the boundary.
[0247] When the second distance is less than the preset distance threshold, determine that the situation where the target area is not tracked in the video images is of the type that the target area has shifted outside the video image, and assign the coordinate data of the target area as a null value.
[0248] In one embodiment, when the situation that the target area is not tracked is that the tracking model is abnormal and the restricted area is within the first video image, after the area tracking module is used to determine that the target area is not tracked in the video images, the area tracking module is further used for:
[0249] When the target area is not tracked in the video images, reduce the tracking matching threshold according to a preset step size, and perform the step of obtaining the recognition result according to the initial coordinate data and tracking the position of the target area in each video image in each video image until it is determined that the target area is tracked in each video image or the tracking matching threshold is equal to the first probability threshold.
[0250] In one embodiment, when the situation that the target area is not tracked is that the tracking model is abnormal and the restricted area is within the first video image, after the area tracking module is used to determine that the target area is not tracked in the video images, the area tracking module is further used for:
[0251] Generate a plurality of second tracking images with the vertices of the first tracking image corresponding to each video image as the centers and the length and width of the first tracking image as the benchmarks, and perform the step of inputting the reference image and the first tracking image into a preset area tracking model.
[0252] In one embodiment, the area tracking module is further used for:
[0253] Obtain the distance between preset points of the target area in two adjacent frames of video images.
[0254] When the distance between the preset points is less than the central distance threshold, update it as the coordinate data of the newly recognized target area.
[0255] When the distance of the preset point exceeds the central distance threshold, the video image of the untracked target area is kept as the target area of the previous frame of video image or a constructed area is adopted; the constructed area refers to the weighted value of the coordinate data of the target area in multiple frames of video images before the video image of the untracked target area.
[0256] In one embodiment, the update judgment module is used to judge whether the posture of the camera changes, including:
[0257] Obtain the amount of change in the angle of the camera;
[0258] When the amount of change in the angle meets the preset condition, it is determined that the posture of the camera changes.
[0259] In one embodiment, the update judgment module is used to judge whether the posture of the camera changes, including:
[0260] Obtain the distances of each pixel point in the target area in two adjacent frames of video images;
[0261] When the distance of at least one pixel point exceeds the pixel point distance threshold, it is determined that the posture of the camera changes.
[0262] In one embodiment, the update judgment module is used to judge whether the posture of the camera changes, including:
[0263] Obtain the distance of the preset point in the target area in two adjacent frames of video images;
[0264] When the distance of the preset point exceeds the central threshold, it is determined that the posture of the camera changes.
[0265] It should be noted that the device shown in this embodiment matches the content of the method embodiment, and the content of the above method embodiment can be referred to and will not be elaborated here.
[0266] In an exemplary embodiment, a security system is further provided, including at least one camera, at least one configuration terminal and a server. The camera is used to collect images and send them to the server; the configuration terminal is used to obtain the initial coordinate data of the target area and send it to the server; see Figure 19 , the server includes:
[0267] A processor 141; a memory 142 for storing computer programs executable by the processor;
[0268] Wherein, the processor is configured to execute the computer program in the memory to implement as Figures 1 to 17 the method described above.
[0269] In an exemplary embodiment, a computer-readable storage medium is further provided, such as a memory including an executable computer program, and the executable computer program can be executed by a processor to implement the method of the embodiment as shown in Figures 1 to 12 the embodiment. Among them, the readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0270] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0271] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for updating regional location, characterized in that The method includes: Obtaining initial coordinate data of a target area in a video image; Tracking the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result; When the identification result contains target coordinate data, determining whether the posture of the camera has changed; After the posture of the camera changes, updating the initial coordinate data according to the target coordinate data to update the position of the target area in the video image; The method further includes: when the identification result does not contain target coordinate data, determining that the target area is not tracked in each video image, including: Determining whether the target area in the first video image is located at the boundary of the first video image; the first video image refers to the previous video image before the video image in which the target area is not tracked; When the target area is located at the boundary of the first video image, obtaining a second distance between the vertex of the target area far from the boundary and the boundary; When the second distance is less than a preset distance threshold, determining that the fact that the target area is not tracked in each video image is a type that the target area has shifted outside the video image; 2. The method according to claim 1, wherein Obtaining initial coordinate data of a target area in a video image includes: In response to detecting an operation representing drawing a target area, obtaining coordinate data of each trigger position; Sequentially connecting each trigger position to obtain a target area; When the shape of the target area is a rectangle, using the coordinate data of each trigger position as the initial coordinate data of the target area; when the shape of the target area is other than a rectangle, obtaining the minimum bounding rectangle of the other shape and using the coordinate data of each vertex of the minimum bounding rectangle as the initial coordinate data of the target area.
3. The method according to claim 1 or 2, characterized in that, Tracking the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result includes: Obtaining an image of the target area corresponding to the initial coordinate data in a target video image based on the initial coordinate data to obtain a reference image; Obtaining a first tracking image based on the initial coordinate data, where the first tracking image refers to an image containing the target area in each video image after the target video image; Inputting the reference image and the first tracking image into a preset region tracking model to obtain an identification result, where the identification result includes probability values and coordinate data of at least one candidate region in each video image; 4. The method according to claim 3, characterized in that, The region tracking model includes a Siamese network module, a region candidate network module, and an identification result module; The Siamese network module includes an upper branch network and a lower branch network; the network structures and parameters of the upper branch network and the lower branch network are the same; the upper branch network outputs a feature image of a first size, and the lower branch network outputs a feature image of a second size; The region candidate network module includes a classification branch network and a regression branch network; the classification branch network is used to distinguish the target and the background according to the feature image of the first size and the feature image of the second size; the regression branch network is used to adjust the position of the candidate region; The recognition result module includes a category output unit and a coordinate data output unit; the category output unit is connected to the classification branch network and is used to output the probability values of each candidate region; the coordinate data output unit is connected to the regression branch network and is used to output the coordinate data of each candidate region.
5. The method according to claim 3, wherein The method further includes a step of judging whether the recognition result contains target coordinate data, specifically including: Obtaining the maximum value of the probability values of the at least one candidate region; When the maximum value exceeds a preset probability threshold, determining the candidate region corresponding to the maximum value as the target region tracked in each video image, and obtaining the target coordinate data of the target region.
6. The method according to claim 5, characterized in that, The method further includes: When the maximum value is less than the preset probability threshold, determining that the target region is not tracked in each video image or a part of the target region is tracked.
7. The method according to claim 6, characterized in that When the target region is not tracked and the tracking model is abnormal and the target region is within the first video image, the method further includes: When the target region is not tracked in each video image, reducing the tracking matching threshold according to a preset step length, and performing the step of obtaining the recognition result according to the initial coordinate data and tracking the position of the target region in each video image in each video image until it is determined that the target region is tracked in each video image or the tracking matching threshold is equal to the first probability threshold, where the first probability threshold refers to the minimum value of the tracking matching threshold.
8. The method according to claim 6, characterized in that, When the target region is not tracked and the tracking model is abnormal and the target region is within the first video image, the method further includes: Generating a plurality of second tracking images with the vertices of the first tracking image corresponding to each video image as the centers and the length and width of the first tracking image as the benchmarks, and performing the step of inputting the reference image and the first tracking image into a preset region tracking model.
9. The method according to any one of claims 6 to 8, characterized in that The method further includes: Obtaining the distance between preset points of the target region in two adjacent video images; When the distance between the preset points is less than the center distance threshold, updating it as the coordinate data of the newly recognized target region; When the distance between the preset points exceeds the center distance threshold, keeping the target region of the video image in which the target region is not tracked as the target region of the previous video image or adopting a constructed region; the constructed region refers to the weighted value of the coordinate data of the target region in multiple video images before the video image in which the target region is not tracked.
10. The method according to claim 1, wherein Judging whether the posture of the camera changes, including: Obtaining the angle change amount of the camera; When the angle change amount meets a preset condition, determining that the posture of the camera changes.
11. The method according to claim 1, characterized in that, Judging whether the posture of the camera changes, including: Obtaining the distances between each pixel point in the target region in two adjacent video images; When there is at least one pixel point whose distance exceeds the pixel point distance threshold, determining that the posture of the camera changes.
12. The method according to claim 1, wherein Judging whether the posture of the camera changes, including: Obtaining the distance between preset points of the target region in two adjacent video images; When the distance between the preset points exceeds the center threshold, determining that the posture of the camera changes.
13. The method according to claim 1, characterized in that, Updating the initial coordinate data according to the target coordinate data includes: When the shape of the target area is a rectangle, updating the initial coordinate data to the target coordinate data; or, When the shape of the target area is other than a rectangle, obtaining relative position data of a preset target area and its minimum bounding rectangle; calculating target restoration data of the target area according to the target coordinate data and the relative position data; updating the initial coordinate data to the target restoration data.
14. A security system, characterized in that, The system includes an area configuration module, an area tracking module, an update judgment module, and a coordinate feedback module; The area configuration module is configured to obtain initial coordinate data of a target area in a video image and send the initial coordinate data to the area tracking module; The area tracking module is configured to track the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result, and send the target coordinate data to the update judgment module when the identification result includes the target coordinate data; The update judgment module is configured to judge whether the posture of the camera has changed, and send the target coordinate data to the coordinate feedback module after the posture of the camera has changed; The coordinate feedback module is configured to feedback the target coordinate data to the area configuration module, so that the area configuration module updates the initial coordinate data according to the target coordinate data to update the position of the target area in the video image; When the identification result does not include the target coordinate data, the area tracking module is configured to determine that the target area is not tracked in each video image, including: Determining whether the target area in the first video image is located at the boundary of the first video image; the first video image refers to the previous frame of video image before the video image where the target area is not tracked; When the target area is located at the boundary of the first video image, obtaining a second distance between a vertex of the target area far from the boundary and the boundary; When the second distance is less than a preset distance threshold, determining that the situation that the target area is not tracked in each video image is a type that the target area has shifted outside the video image.
15. The system according to claim 14, wherein, The area configuration module includes: A coordinate data acquisition unit, configured to respond to detecting an operation of representing drawing the target area and acquire coordinate data of each trigger position; A target area acquisition unit, configured to sequentially connect each trigger position to obtain a target area; An initial coordinate acquisition unit, configured to use the coordinate data of each trigger position as the initial coordinate data of the target area when the shape of the target area is a rectangle; when the shape of the target area is other than a rectangle, obtaining the minimum bounding rectangle of the other shape and using the coordinate data of each vertex of the minimum bounding rectangle as the initial coordinate data of the target area.
16. The system according to claim 14 or 15, characterized in that, The area tracking module is configured to track the position of the target area in each video image according to the initial coordinate data and each video image to obtain an identification result, including: Based on the initial coordinate data, obtain the image of the target area corresponding to the initial coordinate data in the target video image to obtain a reference image; Based on the initial coordinate data, obtain a first tracking image, where the first tracking image refers to the images containing the target area in each video image after the target video image; Input the reference image and the first tracking image into a preset region tracking model to obtain an identification result, where the identification result includes the probability values and coordinate data of at least one candidate region in each video image.
17. The system according to claim 16, wherein The region tracking model includes a Siamese network module, a region candidate network module, and an identification result module; The Siamese network module includes an upper branch network and a lower branch network; the network structures and parameters of the upper branch network and the lower branch network are the same; the upper branch network outputs a feature image of a first size, and the lower branch network outputs a feature image of a second size; The region candidate network module includes a classification branch network and a regression branch network; the classification branch network is used to distinguish the target and the background according to the feature image of the first size and the feature image of the second size; the regression branch network is used to adjust the position of the candidate region; The identification result module includes a class output unit and a coordinate data output unit; the class output unit is connected to the classification branch network and is used to output the probability values of each candidate region; the coordinate data output unit is connected to the regression branch network and is used to output the coordinate data of each candidate region.
18. The system according to claim 16, wherein The region tracking module is used to send the target coordinate data to the update judgment module when the identification result contains the target coordinate data, and includes: Obtain the maximum value of the probability values of the at least one candidate region; When the maximum value exceeds a preset probability threshold, determine the candidate region corresponding to the maximum value as the target region tracked in each video image, and obtain the target coordinate data of the target region; Send the target coordinate data of the target region to the update judgment module.
19. The system according to claim 18, wherein, The region tracking module is further used for: When the maximum value is less than the preset probability threshold, determine that the target region is not tracked in each video image.
20. The system according to claim 19, wherein When the situation that the target region is not tracked is due to an abnormal tracking model and the target region does not go out of bounds, after the region tracking module is used to determine that the target region is not tracked in each video image, the region tracking module is further used for: When the target region is not tracked in each video image, reduce the tracking matching threshold according to a preset step size, and execute the step of obtaining the identification result according to the initial coordinate data and tracking the position of the target region in each video image until it is determined that the target region is tracked in each video image or the tracking matching threshold is equal to the first probability threshold, where the first probability threshold refers to the minimum value of the tracking matching threshold.
21. The system according to claim 19, wherein, When the situation that the target region is not tracked is due to an abnormal tracking model and the target region does not go out of bounds, after the region tracking module is used to determine that the target region is not tracked in each video image, the region tracking module is further used for: Generate a plurality of second tracking images centered on each vertex of the first tracking image corresponding to each video image, with the length and width of the first tracking image as the benchmark, and perform the step of inputting the benchmark image and the first tracking image into a preset region tracking model.
22. The system according to any one of claims 19 to 21, characterized in that, The region tracking module is further configured to: Obtain the distance between preset points of the target region in two adjacent video images; When the distance between the preset points is less than the center distance threshold, update it to the coordinate data of the newly recognized target region; When the distance between the preset points exceeds the center distance threshold, keep the target region of the previous frame of the video image without tracking the target region or adopt a constructed region; the constructed region refers to the weighted value of the coordinate data of the target region in multiple frames of video images before the video image without tracking the target region.
23. The system according to claim 14, wherein The update judgment module is used to judge whether the posture of the camera has changed, including: Obtain the angle change amount of the camera; When the angle change amount meets the preset conditions, determine that the posture of the camera has changed.
24. The system according to claim 14, wherein The update judgment module is used to judge whether the posture of the camera has changed, including: Obtain the distance between each pixel point in the target region of two adjacent video images; When the distance of at least one pixel point exceeds the pixel point distance threshold, determine that the posture of the camera has changed.
25. The system according to claim 14, characterized in that, The update judgment module is used to judge whether the posture of the camera has changed, including: Obtain the distance between preset points of the target region in two adjacent video images; When the distance between the preset points exceeds the center threshold, determine that the posture of the camera has changed.
26. The system according to claim 14, wherein The region configuration module includes: The first configuration module is used to directly update the initial coordinate data according to the target coordinate data when the shape of the target region is a rectangle; or, The second configuration module is used to obtain the relative position data between the preset target region and the minimum circumscribed rectangle when the shape of the target region is other than a rectangle; calculate the target recovery data of the target region according to the target coordinate data and the relative position data; update the initial coordinate data to the target recovery data.
27. A security system, characterized in that, Include at least one camera, at least one configuration terminal and a server; The camera is used to collect images and send them to the server; the configuration terminal is used to obtain the initial coordinate data of the target region and send it to the server; The server includes: A processor; A memory for storing computer programs executable by the processor; Wherein, the processor is configured to execute the computer program in the memory to implement the method according to any one of claims 1 to 13.
28. A computer-readable storage medium, characterized in that, When the executable computer program in the storage medium is executed by the processor, it can implement the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Visual multi-target tracking method and device based on deep learning
CN111161311A
Region adjustment method and device, camera and storage medium
CN113923420A