A target tracking and positioning system based on video analysis
By designing a target tracking and positioning system based on video analysis in the video surveillance system, using video analysis technology to calculate the human grid characteristics and contribution sequence, and adjust the camera angle, the problem of difficulty in adjusting the angle of the camera in the existing technology is solved, and the clarity and accuracy of the monitoring effect are improved.
Patent Information
- Application Number
- CN202510088432.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing video surveillance system handles different movement speeds and different ways of handling characters, making it difficult for the camera to adjust the angle in a timely and accurate manner, resulting in poor monitoring results.
A target tracking and positioning system based on video analysis is designed. Through the data detection module, feature generation module, calculation module and tracking and positioning module, the pre-trained character recognition network model is used to calculate the human grid feature map and contribution sequence, and adjust the camera angle to achieve tracking and positioning of the target.
It realizes that the camera automatically adjusts the angle when the character moves, ensures that the character's screen-occupation ratio is appropriate, fully monitors the character's movement trajectory, and improves the clarity and accuracy of the monitoring effect.
Smart Images

Figure CN119540875B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of video surveillance. More specifically, the present invention relates to an object tracking and positioning system based on video analysis. Background Art
[0002] A reasonable camera shooting angle can obtain better shooting effects and visual experiences. By adjusting the angle of the camera, users can obtain a better visual experience, clearer images, and more accurate information; for a video surveillance system, a suitable camera angle can increase the monitoring range and field of view, improving the monitoring effect and security.
[0003] In related technologies, for example, a patent with the application publication number CN118071826A discloses a target positioning method and system based on video surveillance, including obtaining video stream data, establishing multiple target detection frames, using each target detection frame to extract multi-channel gradient features and multi-channel color features from different target regions of the video image to generate multiple feature sub-maps; splicing the feature sub-maps according to the intersection over union between any two target detection frames, calculating the motion blur kernel of the spliced target feature map to calculate the blur degree, compensating and optimizing the motion blur kernel until the blur degree reaches a preset value, performing a de-blurring operation on the target feature map, and calculating the position coordinates of the target in consecutive frame images in combination with a correlation filter to complete target tracking and positioning.
[0004] In the prior art, the camera generally needs to maintain a certain distance from the person to achieve a suitable person occupancy ratio on the screen. Then, calibrate the position and range of the person, and based on the recognition of several actions of the person, calculate the angle that needs to be adjusted, and then manually adjust the angle of the camera to achieve a better monitoring effect.
[0005] This solution can control the camera to adjust the angle according to the moving direction of the person when someone moves, so that the camera always adjusts the angle following the changes of the person, solving the problem that the camera cannot achieve a timely and accurate monitoring effect due to different action speeds and action methods of the person. Summary of the Invention
[0006] To solve the above one or more technical problems, the present invention proposes an object tracking and positioning system based on video analysis. For this purpose, the present invention provides solutions in the following aspects.
[0007] The present invention provides a target tracking and positioning system based on video analysis, including: a data detection module, configured to input video data captured by a camera in different scenarios into a pre-trained human recognition network model for detection to obtain a set of human region feature images; a feature generation module, configured to calculate the portrait feature data of each image in the set of human region feature images, and divide a grid for the human region feature map based on the portrait feature data to obtain a human grid feature map; a calculation module, configured to calculate the contribution degree of each grid in the human grid feature map to human classification to construct a contribution degree sequence, where the human grid feature map and the contribution degree sequence are in one-to-one correspondence; a tracking and positioning module, configured to calculate the distance between contribution degree sequences of adjacent images in the set of human region feature images, and adjust the camera angle based on the distance to achieve tracking and positioning of the target.
[0008] By controlling the camera to change the angle as the person moves, making the person have a suitable screen occupation ratio, the action trajectory of the person can be completely monitored. At the same time, dividing the grid for the image can more sensitively judge the action information of the person, and enable the camera to quickly adjust to a suitable angle, reducing the angle shooting deviation and making the monitoring effect clearer and more accurate.
[0009] Preferably, the portrait feature data includes the proportion of the area of the human circumscribed matrix in the entire image area, the proportion of the human area within the external matrix, the average density of the human body in the row direction, and the average density of the human body in the column direction, including:
[0010] The proportion of the area of the human circumscribed rectangle in the entire image area satisfies the relational expression:
[0011] ;
[0012] The proportion of the human area within the external matrix satisfies the relational expression:
[0013] ;
[0014] The average density of the human body in the row direction satisfies the relational expression:
[0015] ;
[0016] The average density of the human body in the column direction satisfies the relational expression:
[0017] ;
[0018] Among them, represents the proportion of the area of the human circumscribed rectangle in the entire image area, represents the row of the image, represents the column of the image, represents the proportion of the human area within the circumscribed rectangle, Represents the average density of the human body in the row direction, Represents the average density of the human body in the column direction, Represents the area of the human body region in the circumscribed rectangle, Represents the area of the circumscribed rectangle, Represents the width of the circumscribed rectangle, Represents the height of the circumscribed rectangle.
[0019] Preferably, the grid for dividing the human body region feature map includes the number of row grids and the number of column grids, including:
[0020] The number of row grids satisfies the relational expression:
[0021] ;
[0022] The number of column grids satisfies the relational expression:
[0023] ;
[0024] Wherein, Represents the number of row grids of the image, Represents the number of column grids of the image, Represents the ratio of the area of the human body circumscribed rectangle to the area of the entire image, Represents the proportion of the human body region within the circumscribed rectangle, Represents the average density of the human body in the row direction, Represents the average density of the human body in the column direction.
[0025] In this way, dividing the image into grids can more sensitively judge the action information of the person, and enable the camera to adjust to a suitable angle, reduce the shooting angle deviation, and make the monitoring effect clearer and more accurate.
[0026] Preferably, the contribution degree of the grid human body classification satisfies the relational expression:
[0027] ;
[0028] Wherein, Represents the contribution degree of the th grid to the human body classification, Represents the rd grid, the th row and the th column position pixel value.
[0029] Preferably, the method for adjusting the camera angle based on distance includes: calculating the distance between two adjacent contribution sequences based on the contribution sequence, comparing the distance with a preset threshold, outputting a first signal if the distance is less than or equal to the preset threshold, and outputting a second signal if the distance is greater than the preset threshold; in response to the first signal, since the person in the image has not moved, there is no need to adjust the camera angle; in response to the second signal, since the person in the image has moved, the camera angle needs to be adjusted.
[0030] Preferably, the method for adjusting the camera angle includes: based on the change in the contribution of the same-position grids within the adjacent frame images, extracting the grids with changes in the person area, and screening the grids with increased contribution to obtain the moving direction of the person; calculating the center point coordinates of the overall grids with increased contribution, and controlling the camera to adjust to this center point.
[0031] The present invention has the following technical effects:
[0032] 1. It can control the camera to change the angle as the person moves, so that the person has an appropriate screen occupancy ratio and can completely monitor the movement trajectory of the person.
[0033] 2. Dividing the image into grids can more sensitively judge the movement information of the person, and enable the camera to quickly adjust to an appropriate angle, reducing the shooting angle deviation and making the monitoring effect clearer and more accurate.
[0034] 3. By calculating the average value of the grid divisions corresponding to all human body areas on the image, the overall monitoring screen can be kept relatively stable, preventing the situation where the person appears in different sizes and avoiding the person being blocked. Description of the Drawings
[0035] By reading the following detailed description with reference to the drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where:
[0036] Figure 1 is the system block diagram of a target tracking and positioning system based on video analysis according to an embodiment of the present invention. Detailed Embodiments
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] It should be understood that when terms such as "first", "second", etc. are used in the claims, the description and the drawings of the present invention, they are only used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the description and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0039] The present invention provides a target tracking and positioning system based on video analysis. As Figure 1 shown, a target tracking and positioning system based on video analysis includes modules 101 - 104, which will be specifically described below.
[0040] The data detection module 101 is used to input the video data captured by the camera in different scenarios into a pre-trained human recognition network model for detection to obtain a set of human region feature images.
[0041] Among them, collect the videos captured by the camera in different environments. Different environments include whether the lights are on and whether there are people in the room, etc. Convert the videos into images, give each image a classification label, and the labels are divided into a "person present" label and a "person absent" label to establish a human body classification data set.
[0042] In one embodiment, simulate the application environment of the camera. The application environment of the camera includes the lights being on or off in the room and whether there are people in the room, etc. Shoot videos in different environments, convert the videos into images, extract one image every 10 frames of images, give each extracted image a label, and the labels are divided into a "person present" label and a "person absent" label to establish a human body classification data set. Shooting videos in different environments is to enhance the diversity of the images, so that the trained model has stronger robustness.
[0043] Among them, the human body classification model can use a convolutional neural network. Train the convolutional neural network using the human body classification data set. After training is completed, classify the newly captured images by the camera and use the cam (feature visualization technology) algorithm to obtain the cam grayscale image of the new image.
[0044] In one embodiment, a classification convolutional neural network is trained using a human body classification dataset. The input is an image, and the output is 1 for someone present and 0 for no one present. The loss function of the model uses the cross-entropy loss function. Training stops when the model reaches the maximum number of training times or the model loss is less than a preset loss threshold. The optimal classification model is selected according to the model evaluation metric accuracy. After completing the training of the classification convolutional neural network, the CAM grayscale image of the new image is obtained using feature visualization technology. The larger the pixel value in the grayscale image, the greater the accuracy of the model in determining that there is someone in the image. Connected component extraction is performed on the grayscale image to obtain the human body region in the image, which is the human body region feature map. In this solution, all the obtained human body region feature maps are used as the human body region feature map set.
[0045] In this alternative embodiment, portrait feature data is calculated based on the human body region feature map. Among them, the portrait feature data includes the proportion of the area of the circumscribed matrix of the human body in the area of the entire image, the proportion of the human body region within the external matrix, the average density of the human body in the row direction, and the average density of the human body in the column direction.
[0046] In one embodiment, the circumscribed rectangle of each connected component is determined, and the portrait feature data is calculated. The portrait feature data includes the proportion of the area of the circumscribed rectangle of the human body in the area of the entire image, the proportion of the human body region within each circumscribed rectangle, and the average density of the human body in the row direction and the average density of the human body in the column direction of the circumscribed rectangle. The number of pixel points in the connected component within each circumscribed rectangle is counted to obtain the area size S of each human body region. According to S, the height and width of the circumscribed rectangle, the proportion of the area of the circumscribed rectangle of the human body in the area of the entire image, the proportion of the human body region within the circumscribed rectangle, the average density of the human body in the row direction, and the average density of the human body in the column direction are obtained.
[0047] The proportion of the area of the circumscribed rectangle of the human body in the area of the entire image satisfies the relation:
[0048] ;
[0049] The proportion of the human body region within the external matrix satisfies the relation:
[0050] ;
[0051] The average density of the human body in the row direction satisfies the relation:
[0052] ;
[0053] The average density of the human body in the column direction satisfies the relation:
[0054] ;
[0055] Among them, represents the proportion of the area of the circumscribed rectangle of the human body in the area of the entire image, represents the row of the image, Represents the column of the image, Represents the proportion of the human body area within the bounding rectangle, Represents the average density of the human body in the row, Represents the average density of the human body in the column, Represents the area of the human body area in the bounding rectangle, Represents the area of the bounding rectangle, Represents the width of the bounding rectangle, Represents the height of the bounding rectangle.
[0056] Among them, when dividing the image into grids, if the grid division is too large, the changes in the human body area cannot be reflected in the grid; if the grid division is too small, the computational requirements are relatively high. The purpose of dividing the image into grids is to make the human body features in each grid more prominent. The human body features in the grid specifically refer to being more sensitive to the changes in the human body area within the grid.
[0057] Preferably, when the grid division is the same as the size of the image, that is, one image is divided into one grid, the movement of the human body area in the grid does not change the area of the human body area in the entire grid. Therefore, according to the proportion of the area of the human body bounding rectangle in the entire image area, the proportion of the human body area within the bounding rectangle, and the average density of the human body in the row and column, the image is divided into grids. The smaller the proportion of the area of the human body bounding rectangle in the entire image area, the more grids are divided; the smaller the proportion of the human body area within the bounding rectangle, the more grids are divided; the smaller the average density in the row, the more row grids in the image; the smaller the average density in the column, the more column grids in the image. The proportion of the human body area within the bounding rectangle determines the overall number of grid divisions of the image, and the average density of the human body in the row and column respectively determines the number of grid divisions in the row and column of the image.
[0058] Among them, the grid division of the human body area feature map includes the number of row grids and the number of column grids, including:
[0059] The number of row grids satisfies the relationship:
[0060] ;
[0061] The number of column grids satisfies the relationship:
[0062] ;
[0063] Among them, Represents the number of row grids of the image, Represents the number of column grids of the image, Represents the proportion of the area of the human body bounding rectangle in the entire image area, Represents the proportion of the human body area within the bounding rectangle, Represents the average density of the human body in the row, Represents the average density of the human body in the column.
[0064] Preferably, the ratio of the area of the circumscribed rectangle of the human body to the area of the entire image is 1 / 6, the ratio of the human body area within the circumscribed rectangle is 1 / 2, the average row and column densities of the human body are h / 2 and w / 2 respectively, and the image is divided into a 24×24 grid.
[0065] Preferably, when there are multiple human body areas on an image, each human body area corresponds to a grid division, and the grid division of the entire image is the average value of each grid division.
[0066] The calculation module 103 is used to calculate the contribution degree of each grid in the human body grid feature map to human body classification to construct a contribution degree sequence.
[0067] Among them, the contribution degree of grid human body classification satisfies the relational expression:
[0068] ;
[0069] Among them, represents the contribution degree of the th grid to human body classification, represents the th grid, the th row and the th column position pixel value. The greater the contribution degree, the greater the influence of the grid on human body classification. Calculating the human body classification contribution degree of each grid obtains the human body classification contribution degree sequence of the image.
[0070] The tracking and positioning module 104 is used to calculate the distance between contribution degree sequences and adjust the camera angle based on the distance to achieve tracking and positioning of the target.
[0071] Among them, calculate the distance between the human body classification contribution degree sequences of the front and rear frame images, compare the distance between the human body classification contribution degree sequences of the front and rear frame images with a preset threshold. If it is less than or equal to the preset threshold, output the first signal. If it is greater than the preset threshold, output the second signal; in response to the first signal, the person in the image has not moved and there is no need to adjust the camera angle; in response to the second signal, the person in the image has moved and the camera angle needs to be adjusted.
[0072] Preferably, the threshold is set to S / 6, where S is the area size of the human body area.
[0073] Among them, if the person moves, then judge which grids the human body in moves according to the contribution of each grid in the image to the distance, and move the angle of the camera to move the center of the camera to the new human body center grid.
[0074] Specifically, if the sequence distance of the human body classification contribution degrees of the front and rear frame images is greater than a set threshold value, it indicates that the human body in the image has moved. Determine which grid areas the characters in move based on the contribution of each grid to the distance, extract the grids with an increased contribution degree of the human body classification in the grid to obtain the moving direction of the human body, calculate the center point coordinates of these grids as a whole, adjust the camera angle to this center point. After one adjustment is completed, determine the new grid division result of the image based on the initial frame captured by the camera after the adjustment angle. When there is no longer a human body area in the entire camera, that is, no human body area is detected within the range that can be captured at all camera angles, the camera angle is adjusted to the central position to prevent the situation where the camera cannot detect the human body area after rotating to a certain angle and then keeps shooting at an extreme angle.
[0075] As can be seen from the above description of the modular design of the present invention, the system of the present invention can be flexibly arranged according to the application scenario or requirements without being limited to the architecture shown in the drawings. Further, it should also be understood that any module, unit, component, server, computer or device that executes operations in the examples of the present invention may include or otherwise access a computer-readable medium, such as a storage medium, a computer storage medium or a data storage device (removable) and / or non-removable), such as a magnetic disk, an optical disk or a magnetic tape. The computer storage medium may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information, such as computer-readable instructions, data structures, program modules or other data.
[0076] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0077] The above-described embodiments only represent several implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.
Claims
1. A target tracking and positioning system based on video analysis, characterized in that: include: A data detection module is used to input the video data captured by the camera in different scenes into a pre-trained person recognition network model for detection to obtain a human body region feature image set; A feature generation module, used to calculate the portrait feature data of each image in the human body region feature image set, and divide the human body region feature map into grids based on the portrait feature data to obtain a human body grid feature map; The portrait feature data includes the ratio of the area of the human body circumscribed matrix to the entire image area, the ratio of the human body area in the circumscribed matrix, the average human body density in the rows and the average human body density in the columns; The ratio of the area of the human body's circumscribed rectangle to the entire image area satisfies the relationship: ; The proportion of the human body area in the external matrix satisfies the relationship: ; The average density of pedestrians satisfies the relationship: ; The average density of the human body satisfies the relationship: ; It represents the ratio of the area of the human body's circumscribed rectangle to the entire image area. represents the rows of the image, represents the columns of the image, Indicates the proportion of the human body area within the circumscribed rectangle. represents the average density of pedestrians, represents the average density of human body, Represents the area of the human body in the circumscribed rectangle, represents the area of the circumscribed rectangle, Indicates the width of the bounding rectangle. Indicates the height of the circumscribed rectangle; The grid division of the human body region feature map includes the number of row grids and the number of column grids; the number of row grids satisfies the relationship: ; The number of column grids satisfies the relationship: ; Represents the number of row grids of the image, Represents the number of column grids of the image; A calculation module, used for calculating the contribution of each grid in the human body grid feature map to human body classification to construct a contribution sequence, wherein the human body grid feature map and the contribution sequence correspond one to one; The contribution of each grid to human body classification satisfies the relationship: ; Indicates The contribution of each grid to human body classification. Indicates In the grid Line The pixel value at the column position; A tracking and positioning module, used to calculate the distance of contribution degree sequences between adjacent images in the human body region feature image set; The camera angle is adjusted based on the distance, including: based on the contribution changes of the grids at the same position in the adjacent images, the grids with changed character areas are extracted, and the grids with increased contribution are screened to obtain the direction of character movement; the coordinates of the center point of the grids with increased contribution are calculated as a whole, and the camera is controlled to adjust to the center point; so as to achieve tracking and positioning of the target.
2. The target tracking and positioning system based on video analysis according to claim 1, characterized in that: The adjusting the camera angle based on the distance comprises: Based on the contribution sequence, a distance between two adjacent contribution sequences is calculated, and the distance is compared with a preset threshold. If the distance is less than or equal to the preset threshold, a first signal is output; if the distance is greater than the preset threshold, a second signal is output; In response to the first signal, the person in the image does not move, and there is no need to adjust the camera angle; In response to the second signal, the person in the image moves and the camera angle needs to be adjusted.
Citation Information
Patent Citations
Target positioning method and system based on video monitoring
CN118071826A
Video monitoring target tracking method and device, computer equipment and storage medium
CN118784984A