Image information transmission method based on video monitoring

By analyzing the probability of motion in the content area and the probability of changes in the bounding box state of the surveillance video image, the sharpness is dynamically adjusted, which solves the problem of blurry key images caused by coarse sharpness adjustment in the existing technology, and realizes efficient video image transmission and identification of specific personnel.

CN120881239BActive Publication Date: 2026-01-23XIAN XINGXUN INTELLIGENT COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511367396.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-23
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing methods for adjusting the clarity of surveillance video images are crude and cannot meet the high-definition requirements of key areas and the low-definition compression requirements of non-key areas, resulting in blurry key images and making it difficult to meet the tracking and identification requirements of specific personnel.

Method used

By acquiring the probability of motion in the content area and the probability of changes in the bounding box state of the surveillance video image, and combining this with the predicted center position, the resolution requirements are dynamically adjusted to achieve resolution adjustment and transmission of the surveillance video image.

Benefits of technology

While ensuring transmission efficiency, the clarity of important content in surveillance video images has been improved, enhancing the ability to identify and track specific individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120881239B_ABST
    Figure CN120881239B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, and discloses an image information transmission method based on video monitoring, which comprises the following steps: obtaining a motion probability of a content area according to the positional relationship between a pixel point and the edge of the content area; obtaining a motion vector and a predicted central position of a surrounding frame according to the positional relationship between the surrounding frame and its matching frame in the latest monitoring video image; obtaining the overall change degree of the pixel point according to the difference between the motion vectors of the surrounding frame and its matching frame; obtaining the state change degree of the pixel point according to the overall change degree of the pixel point in the state evaluation image; obtaining the state change probability of the surrounding frame according to the state change degree of the pixel point; obtaining a predicted surrounding frame according to the state change probability of the surrounding frame; and obtaining the clear demand degree according to the predicted surrounding frame and the motion probability. The next frame of the monitoring video image is processed and transmitted through the clear demand degree, so that the transmission efficiency is ensured, and the definition of important image content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to an image information transmission method based on video monitoring. BACKGROUND

[0002] In a security monitoring scene, a specific area and a specific person need to be video image collected by deploying a monitoring camera and the like, and video information needs to be quickly and safely transmitted to an analysis platform in combination with network communication technology. While ensuring the video data transmission efficiency, the clarity of the person in the video needs to be ensured, thereby providing reliable support for the analysis platform to research and judge the behavior of the specific area and the specific person.

[0003] In the prior art, when transmitting the picture captured by the monitoring camera, the clarity of part of the content with low importance is generally reduced, thereby improving the transmission efficiency of the video data. However, due to the high requirement for appearance recognition of the specific person in the security monitoring scene, the existing method is relatively rough in adjusting the clarity of the monitoring video image, and often cannot balance the high-definition requirement of the key area and the low-definition compression requirement of the non-key area, and it is difficult to meet the tracking and recognition requirement for the specific person, and the key picture is prone to be blurred. SUMMARY

[0004] The present application provides an image information transmission method based on video monitoring to solve the problem of blurred key picture caused by rough clarity adjustment of the existing monitoring video image, and the technical solution adopted is as follows:

[0005] The present application provides an image information transmission method based on video monitoring to solve the problem of blurred key picture caused by rough clarity adjustment of the existing monitoring video image, and the technical solution adopted is as follows:

[0006] A plurality of continuous monitoring video images are acquired;

[0007] A plurality of content areas and a gray difference value of each pixel point are acquired according to the latest monitoring video image; a motion probability of each content area is obtained according to the positional relationship between the pixel point and the edge of the content area and the gray difference value of the pixel point;

[0008] A plurality of bounding boxes and their confidence degrees in each monitoring video image are acquired, and a matching box of each bounding box is acquired; a motion vector and a predicted center position of each bounding box are obtained according to the positional relationship between the bounding box and its matching box in the latest monitoring video image; an overall change degree of each pixel point in each bounding box is obtained according to the difference of the motion vectors of the bounding box and its matching box; a state evaluation image of each pixel point is acquired in the latest monitoring video image; a state change degree of each pixel point is obtained according to the overall change degree of the pixel point in its state evaluation image; a state change probability of each bounding box is obtained according to the distance between the bounding box and each pixel point and the state change degree of the pixel point;

[0009] In the latest monitoring video image, according to the state change probability of the bounding box, in combination with the predicted center position, a predicted bounding box of each bounding box is obtained; according to the confidence of the bounding box and the predicted bounding box, in combination with the motion probability of the content area, the clear demand degree of each pixel point is obtained; the next frame of the monitoring video image obtained by collection is adjusted and transmitted in terms of the clear demand degree of the pixel point.

[0010] Further, the method for obtaining the motion probability of each content area according to the positional relationship between the pixel point and the edge of the content area and the gray difference value of the pixel point comprises the following specific method:

[0011] A monitoring gray image of each frame of the monitoring video image is obtained; any one pixel point in the monitoring gray image corresponding to the latest monitoring video image is recorded as a target pixel point;

[0012] A content area where the target pixel point is located is obtained, and the nearest edge pixel point to the target pixel point is obtained from all edge pixel points in the content area and is recorded as a corresponding edge point of the target pixel point;

[0013] The linear normalization result of the average value of the gray difference values of all pixel points on the connection path from the target pixel point to the corresponding edge point of the target pixel point is recorded as the motion performance index of the target pixel point;

[0014] The corresponding edge point of the target pixel point is recorded as a target edge point, and the target pixel point is recorded as a motion performance pixel point of the target edge point;

[0015] The product of the weight normalization result of the gray difference value of the target pixel point and the motion performance index of the target pixel point is recorded as the motion contribution index of the target pixel point; and the sum value of the motion contribution indexes of all motion performance pixel points of the target edge point is recorded as the motion performance degree of the target edge point;

[0016] For any one content area, all edge pixel points of the content area are DBSCAN clustered according to the motion performance degree of the edge pixel points, to obtain a plurality of edge clusters;

[0017] An interval formed by the edge pixel points belonging to the same edge cluster on the edge of the content area is recorded as an edge interval; the average value of the motion performance degrees of all edge pixel points in each edge interval is recorded as the motion index of each edge interval; and the edge interval with the largest motion index in the content area is recorded as the motion representation interval of the content area;

[0018] The linear normalization result of the ratio of the number of edge pixel points in the motion representation interval of the content area to the number of all edge pixel points in the content area is recorded as the motion performance proportion of the content area.

[0019] The product of the motion performance proportion of the content area and the motion index of the motion representation interval of the content area is recorded as the motion probability of the content area.

[0020] Further, the specific method for obtaining the motion vector and the predicted center position of each bounding box according to the positional relationship between the bounding box and its matching box in the latest monitoring video image comprises:

[0021] For any one bounding box in the latest monitoring video image, the vector from the center position of the matching box of the bounding box to the center position of the bounding box is recorded as the motion vector of the bounding box.

[0022] The position obtained by moving the center position of the bounding box along the motion vector of the bounding box is recorded as the predicted center position of the bounding box.

[0023] Further, the specific method for obtaining the overall change degree of each pixel point in each bounding box according to the difference of the motion vector of the bounding box and its matching box comprises:

[0024] For any one bounding box in any one frame of monitoring video image, the linear normalization result of the module length of the difference vector between the motion vector of the bounding box and the motion vector of the matching box in the previous frame of monitoring video image is recorded as the state change index of the bounding box.

[0025] Any one frame of monitoring video image is recorded as a target monitoring video image; for any one bounding box in the target monitoring video image and any one pixel point in the bounding box, the state change index of the bounding box is recorded as the change contribution index of the bounding box to the pixel point; the average value of the change contribution indices of all bounding boxes containing the pixel point in the target monitoring video image to the pixel point is recorded as the overall change degree of the pixel point in the target monitoring video image.

[0026] Further, the specific method for obtaining the state evaluation image of each pixel point in the latest monitoring video image comprises:

[0027] For any one pixel point in the latest monitoring video image, the pixel point is corresponded to any one frame of monitoring video image; if the pixel point exists in any one bounding box of the monitoring video image, the monitoring video image is recorded as the state evaluation image of the pixel point.

[0028] Further, the specific method for obtaining the state change degree of each pixel point according to the overall change degree of the pixel point in its state evaluation image comprises:

[0029] For any one pixel point in the latest monitoring video image, the variance of the overall change degree of the pixel point in all state evaluation images is recorded as the state change coefficient of the pixel point;

[0030] The product of the maximum value of the overall change degree of the pixel point in all state evaluation images and the state change coefficient of the pixel point is recorded as the fluctuation change index of the pixel point;

[0031] The mean value of the overall change degree of the pixel point in all state evaluation images is obtained, and the product of the difference between 1 and the state change coefficient of the pixel point and the mean value is recorded as the stable change index of the pixel point;

[0032] The sum of the fluctuation change index and the stable change index of the pixel point is recorded as the state change degree of the pixel point.

[0033] Further, the state change probability of each bounding box is obtained according to the distance between the bounding box and each pixel point and the state change degree of the pixel point, including the specific method that:

[0034] For any one bounding box and any one pixel point in the latest monitoring video image, the linear normalization result of the reciprocal of the Euclidean distance between the center position of the bounding box and the pixel point is obtained, and the product of the linear normalization result and the state change degree of the pixel point is recorded as the change contribution degree of the pixel point to the bounding box; the mean value of the change contribution degree of all pixel points to the bounding box in the latest monitoring video image is recorded as the state change probability of the bounding box.

[0035] Further, the predicted bounding box of each bounding box is obtained according to the state change probability of the bounding box and the predicted center position in the latest monitoring video image, including the specific method that:

[0036] For any one bounding box in the latest monitoring video image, the product of the sum of 1 and the state change probability of the bounding box and the length of the bounding box is obtained as the predicted length of the bounding box;

[0037] The predicted width of the bounding box is obtained;

[0038] The predicted center position of the bounding box is taken as the center, and the predicted length and the predicted width are used to construct the predicted bounding box of the bounding box.

[0039] Further, the clear demand degree of each pixel point is obtained according to the confidence of the bounding box and the predicted bounding box, combined with the motion probability of the content area, including the specific method that:

[0040] For any one pixel point in the latest monitoring video image, obtain the difference between 1 and the motion probability of the content area where the pixel point is located, multiply the confidence of the prediction bounding box where the pixel point is located in the corresponding bounding box in the latest monitoring video image by the difference, and record the product as the clarity promotion degree of the pixel point.

[0041] Add the clarity promotion degree of the pixel point and the motion probability of the content area where the pixel point is located, and record the sum as the clarity demand degree of the pixel point.

[0042] Further, the method for obtaining the gray difference value of each pixel point in the latest monitoring video image comprises:

[0043] In the obtained monitoring video image, perform a gray-scale operation on each frame of the monitoring video image respectively to obtain a plurality of monitoring gray images; obtain a difference image between the monitoring gray image corresponding to the latest monitoring video image and the monitoring gray image corresponding to the previous frame of the monitoring video image, and record the difference image as a current difference image;

[0044] Perform DBSCAN clustering on the monitoring gray image corresponding to the latest monitoring video image to obtain a plurality of gray clusters; in the monitoring gray image corresponding to the latest monitoring video image, a single area formed by the pixel points belonging to the same gray cluster is recorded as a content area;

[0045] Record any one pixel point in the monitoring gray image corresponding to the latest monitoring video image as a target pixel point, and record the gray value of the target pixel point at the corresponding position in the current difference image as the gray difference value of the target pixel point.

[0046] The beneficial effects of the present application are: when transmitting the monitoring video image, in order to ensure the transmission efficiency of the monitoring video image, and at the same time improve the definition of important content in the monitoring video image, the motion probability of each content area is obtained, the possibility of motion of each content area in the monitoring video image is judged, and the definition requirement of each content area is basically judged; in the public security monitoring scene, not only the motion information in the monitoring video needs to be analyzed, but also the identity of the person in the monitoring video needs to be identified, the bounding box of the person returned by the monitoring server is used to predict the position and size of the bounding box in the next frame, and in the process of obtaining the predicted bounding box, the state change degree of each pixel point is obtained according to the overall change degree of the pixel point in the state evaluation image, the state change probability of each bounding box is obtained by combining the distance between the bounding box and each pixel point, and then the size of the predicted bounding box is adjusted. Thus, the definition requirement of each pixel point is obtained according to the motion probability of the content area where each pixel point is located and the predicted bounding box where each pixel point is located, and then the definition of the next frame of monitoring video image collected is adjusted and transmitted, so that the definition of important image content is improved on the premise of ensuring the transmission efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0048] Figure 1 The flow chart of the image information transmission method based on video monitoring provided by an embodiment of the present application. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0050] Please refer to Figure 1 which shows the flow chart of the image information transmission method based on video monitoring provided by an embodiment of the present application, and the method comprises the following steps:

[0051] Step S001, acquiring continuous multiple frames of monitoring video images.

[0052] It should be noted that in the public safety monitoring scene, in order to effectively identify and track the characters, it is necessary to first acquire the monitoring video images of the area to be monitored.

[0053] Specifically, a monitoring camera is arranged near the position to be monitored, the lens of the monitoring camera is aimed at the area to be monitored, video acquisition is performed on the area to be monitored, and continuous multiple frames of monitoring video images are obtained; wherein the acquisition frequency of the monitoring camera is 30Hz, and this embodiment is described taking this as an example.

[0054] Step S002, obtaining a plurality of content regions and a gray difference value of each pixel point according to the latest monitoring video image; obtaining a motion probability of each content region according to the positional relationship between the pixel point and the edge of the content region and the gray difference value of the pixel point.

[0055] It should be noted that when the monitoring camera is used to acquire video of the area to be monitored, since the position of the monitoring camera is fixed, the position of the static object in the video content acquired by the monitoring camera is stable, in order to improve the transmission efficiency, the static object and the dynamic object in the video content can be distinguished, and the motion probability of different regions is measured, so as to judge the basic requirement degree of each region to the clarity.

[0056] Specifically, in the acquired monitoring video images, each frame of monitoring video image is subjected to a gray scale operation to obtain a plurality of monitoring gray scale images; a difference image between the monitoring gray scale image corresponding to the latest monitoring video image and the monitoring gray scale image corresponding to the previous frame of monitoring video image is obtained, which is recorded as a current difference image; wherein the gray scale operation method and the difference image acquisition method are known technologies, and the specific method is not introduced herein;

[0057] It should be noted that the pixel point with a higher gray value in the current difference image is often caused by the edge of the dynamic object moving, so it is necessary to analyze the pixel point with a higher gray value in the current difference image to judge the motion probability of each region.

[0058] Specifically, the DBSCAN clustering is performed on the monitoring gray scale image corresponding to the latest monitoring video image to obtain a plurality of gray scale clusters; in the monitoring gray scale image corresponding to the latest monitoring video image, a single region composed of pixel points belonging to the same gray scale cluster is recorded as a content region; wherein the DBSCAN clustering is a known technology, and the specific method is not introduced herein;

[0059] Any pixel point in the monitoring gray scale image corresponding to the latest monitoring video image is recorded as a target pixel point, and the gray value of the target pixel point at the corresponding position in the current difference image is recorded as the gray difference value of the target pixel point.

[0060] It should be noted that the higher the gray value of the pixel point in the current difference image indicates that the nearby content region has a real displacement, and at the same time, when the content region moves, the pixel point with a higher gray value in the current difference image will adhere to the edge of the content region, so the motion probability of each content region is measured accordingly.

[0061] Specifically, the content region where the target pixel point is located is obtained, and the edge pixel point closest to the target pixel point is obtained from all edge pixel points in the content region, which is recorded as the corresponding edge point of the target pixel point.

[0062] The linear normalization result of the mean value of the gray difference values of all pixel points on the connection path from the target pixel point to the corresponding edge point of the target pixel point is recorded as the motion performance index of the target pixel point; wherein the linear normalization object is the mean value of the gray difference values of all pixel points on the connection path from each pixel point in all monitoring gray images to the corresponding edge point thereof.

[0063] It should be noted that the greater the motion performance index of the target pixel point, the stronger the adhesion of the target pixel point to the corresponding edge point thereof, and if at the same time, the greater the gray difference value of the target pixel point, the greater the probability of real movement of the corresponding edge point of the target pixel point.

[0064] Further, the corresponding edge point of the target pixel point is recorded as the target edge point, and the target pixel point is recorded as the motion performance pixel point of the target edge point.

[0065] The product of the weight normalization result of the gray difference value of the target pixel point and the motion performance index of the target pixel point is recorded as the motion contribution index of the target pixel point; and the sum of the motion contribution indices of all motion performance pixel points of the target edge point is recorded as the motion performance degree of the target edge point; wherein the weight normalization object is the gray difference value of all motion performance pixel points of the target edge point.

[0066] It should be noted that if the object corresponding to the content region moves, a high difference value area will be formed in the current difference image near the edge of the content region, which is manifested as the formation of continuous high motion performance degree edge pixel points on the edge of the content region, so for any content region, the motion performance degrees of all edge pixel points need to be integrated to determine whether the object corresponding to the content region has a real movement.

[0067] Specifically, for any content region, all edge pixel points of the content region are DBSCAN clustered according to the motion performance degree of the edge pixel points to obtain a plurality of edge clusters; wherein the DBSCAN clustering is a known technology, and the specific method is not introduced here.

[0068] An interval formed by the edge pixel points belonging to the same edge cluster on the edge of the content region is recorded as an edge interval; the mean of the motion performance degree of all edge pixel points in each edge interval is recorded as the motion index of each edge interval; the edge interval with the largest motion index of the content region is recorded as the motion representation interval of the content region;

[0069] A linear normalization result of the ratio of the number of edge pixel points of the motion representation interval of the content region to the number of all edge pixel points of the content region is recorded as the motion performance proportion of the content region; wherein, the object of linear normalization is the ratio of the number of edge pixel points of the motion representation interval of each content region in each frame of monitoring gray-scale image to the number of all edge pixel points;

[0070] The product of the motion performance proportion of the content region and the motion index of the motion representation interval of the content region is recorded as the motion probability of the content region.

[0071] It should be noted that the higher the motion probability of the content region, the higher the clarity required for the content region, thereby achieving more effective monitoring of the content region.

[0072] Step S003, obtaining a plurality of bounding boxes and their confidence degrees in each frame of monitoring video image, and a matching box of each bounding box; obtaining a motion vector and a predicted center position of each bounding box according to the positional relationship between the bounding box and its matching box in the latest monitoring video image; obtaining the overall change degree of each pixel point in each bounding box according to the difference of the motion vectors of the bounding box and its matching box; obtaining a state evaluation image of each pixel point in the latest monitoring video image; obtaining the state change degree of each pixel point according to the overall change degree of the pixel point in its state evaluation image; obtaining the state change probability of each bounding box according to the distance between the bounding box and each pixel point and the state change degree of the pixel point.

[0073] It should be noted that in the public security monitoring scene, when using a monitoring camera to collect video of the area to be monitored, not only the motion information in the monitoring video needs to be analyzed, but also the identity of the person in the monitoring video needs to be identified, so as to accurately identify the public security concerned person in the monitoring video, and since the computing power of the monitoring camera itself is relatively weak, it is necessary to combine the person identification data returned by the monitoring server to track the person in the monitoring video image.

[0074] Specifically, the obtained each frame of monitoring video image is transmitted to a monitoring server in real time, the monitoring server uses a neural network to perform person recognition on each frame of monitoring video image to obtain bounding box information of each person in each frame of monitoring video image; wherein, the bounding box information of each person includes a center position, a size and a confidence of the bounding box; wherein, the size of the bounding box includes a length and a width of the bounding box, and the confidence of the bounding box represents a similarity between the person and a public security concerned person, and the confidence has a value range of 0 to 1; wherein, the neural network adopts a YOLOv8n model, and the technology of the neural network for person recognition is a known technology, and the specific method is not introduced here.

[0075] For any one bounding box in any frame of monitoring video image, an intersection over union of the bounding box and each bounding box in a previous frame of monitoring video image is obtained, and a bounding box with the highest intersection over union is recorded as a matching box of the bounding box.

[0076] It should be noted that only the motion probability of the content area is used to adjust the definition of the monitoring camera transmission image, and since the position of the person in the next frame will change, the content area with the actual demand for high definition cannot be completely covered, so it is necessary to predict the position of the bounding box where the person is located in the next frame of monitoring video image.

[0077] It should be further noted that in most cases, the speed of the person moving is uniform, so the possible position of the person in the next frame is basically predicted according to the difference between the center positions of each bounding box and its matching box.

[0078] Specifically, for any one bounding box in the latest monitoring video image, a vector from the center position of the matching box of the bounding box to the center position of the bounding box is recorded as a motion vector of the bounding box.

[0079] A position obtained by moving the center position of the bounding box along the motion vector of the bounding box is recorded as a predicted center position of the bounding box.

[0080] It should be noted that the predicted center position of the bounding box represents the most likely position of the person in the next frame of image when the person continues to move in the current moving state, but the moving state of the person can mutate, so the size of the predicted bounding box can reflect the fault tolerance of the position prediction. When the person approaches an obstacle, an intersection or a doorway, the person can suddenly turn, and when the person needs to complete a certain behavior, the person can also change the speed or direction, such as entering a store, going up or down stairs, crossing a zebra crossing, so the moving state of the person can mutate in some positions. Therefore, on the basis of predicting the center position of the bounding box, the size of the bounding box needs to be enlarged to ensure that the range where the person is located can be completely covered.

[0081] It needs to be further explained that since the position of the monitoring camera is fixed, the motion state of the person in the image is likely to change at some positions, which is reflected in the historical monitoring video images collected by the monitoring camera. When the person reaches around some pixel points, there is often a large change in the motion state. Since the position of the monitoring camera is fixed, the same position pixel points in different frames of monitoring video images correspond to the same position in the real scene, so the same position pixel points in different frames of monitoring video images are analyzed next.

[0082] Specifically, for any one bounding box in any one frame of monitoring video image, the linear normalization result of the modulus of the difference vector between the motion vector of the bounding box and the motion vector of the matching box in the previous frame of monitoring video image is recorded as the state change index of the bounding box; wherein the linear normalization object is the modulus of the difference vector between the motion vector of each bounding box in each frame of monitoring video image and the motion vector of the matching box in the previous frame of monitoring video image.

[0083] Any one frame of monitoring video image is recorded as a target monitoring video image; for any one bounding box in the target monitoring video image and any one pixel point in the bounding box, the state change index of the bounding box is recorded as the change contribution index of the bounding box to the pixel point; the average of the change contribution indexes of all bounding boxes containing the pixel point in the target monitoring video image to the pixel point is recorded as the overall change degree of the pixel point in the target monitoring video image; the target monitoring video image is recorded as the state evaluation image of the corresponding pixel point in the latest monitoring video image, that is, for any one pixel point in the latest monitoring video image, the pixel point is corresponded to any one frame of monitoring video image, if the pixel point exists in any one bounding box of the monitoring video image, the monitoring video image is recorded as the state evaluation image of the pixel point; wherein the pixel point in the latest monitoring video image corresponds to the pixel point in any one frame of monitoring video image according to the same position.

[0084] It needs to be noted that in some cases, even if a certain pixel point belongs to a position that is easy to cause the motion state of the person to change, the overall change degree of the corresponding pixel point in the actual scene may still be high or low, which is mainly related to external environmental factors. For example, in a red light scene, some people will stop because of the red light, causing the motion vector to change suddenly and the overall change degree to increase, while some people will directly pass through the red light when it is green, causing the state to change less and the overall change degree to decrease; at the entrance of a shopping mall or a station, some people will stop and queue, causing the speed to change suddenly, while some people will quickly pass through, etc. Therefore, the state change degree of the pixel point is obtained according to the overall change degree of the pixel point.

[0085] Specifically, for any one pixel point in the latest monitoring video image, the variance of the overall change degree of the pixel point in all state evaluation images is recorded as the state change coefficient of the pixel point;

[0086] The product of the maximum value of the overall change degree of the pixel point in all state evaluation images and the state change coefficient of the pixel point is recorded as the fluctuation change index of the pixel point;

[0087] The mean value of the overall change degree of the pixel point in all state evaluation images is obtained, and the product of the difference between 1 and the state change coefficient of the pixel point and the mean value is recorded as the stable change index of the pixel point;

[0088] The sum of the fluctuation change index and the stable change index of the pixel point is recorded as the state change degree of the pixel point.

[0089] It should be noted that the greater the state change coefficient of the pixel point, the more unstable the motion state change of the person near the pixel point, so the mean level of the overall change degree of all state evaluation images cannot be used to measure the state change degree. In order to improve the effectiveness of the identification and tracking of the public security concerned person, the mean level needs to be improved to ensure the clarity of the video transmission.

[0090] It should be noted that for any one bounding box in the latest monitoring video image, the closer the bounding box is to the pixel point with a high state change degree, the greater the possibility of motion state mutation of the person in the bounding box, so in order to improve the prediction fault tolerance of the predicted center position of the bounding box, the prediction size of the bounding box needs to be improved.

[0091] Specifically, for any one bounding box and any one pixel point in the latest monitoring video image, the linear normalization result of the reciprocal of the Euclidean distance between the center position of the bounding box and the pixel point is obtained, and the product of the linear normalization result and the state change degree of the pixel point is recorded as the change contribution degree of the pixel point to the bounding box. The mean value of the change contribution degree of all pixel points to the bounding box in the latest monitoring video image is recorded as the state change probability of the bounding box; wherein the object of linear normalization is the reciprocal of the Euclidean distance between the center position of the bounding box and all pixel points.

[0092] Step S004, in the latest monitoring video image, the predicted bounding box of each bounding box is obtained according to the state change probability of the bounding box combined with the predicted center position; the clarity demand degree of each pixel point is obtained according to the motion probability of the content area combined with the bounding box confidence and the predicted bounding box; the clarity of the next frame of monitoring video image collected is adjusted and transmitted according to the clarity demand degree of the pixel point.

[0093] It should be noted that the greater the state change probability of the bounding box, the greater the possibility of state mutation of the person in the bounding box, and thus the prediction size of the bounding box needs to be enlarged to improve the prediction tolerance of the predicted center position of the person.

[0094] Specifically, for any one bounding box in the latest monitoring video image, the sum of 1 and the state change probability of the bounding box is multiplied by the length of the bounding box to obtain the predicted length of the bounding box.

[0095] According to the method for obtaining the predicted length of the bounding box, the predicted width of the bounding box is obtained.

[0096] The predicted bounding box of the bounding box is constructed using the predicted length and the predicted width with the predicted center position of the bounding box as the center.

[0097] It should be noted that after obtaining the predicted bounding box of each bounding box in the latest monitoring video image, the clarity of the transmission image of the next frame of monitoring video image needs to be adjusted in combination with the motion probability of the content area.

[0098] Specifically, for any one pixel point in the latest monitoring video image, the difference between 1 and the motion probability of the content area where the pixel point is located is obtained, and the product of the confidence of the bounding box corresponding to the predicted bounding box of the pixel point in the latest monitoring video image and the difference is recorded as the clarity improvement degree of the pixel point. It should be noted that if the pixel point does not exist in any predicted bounding box, the clarity improvement degree of the pixel point is 0.

[0099] The sum of the clarity improvement degree of the pixel point and the motion probability of the content area where the pixel point is located is recorded as the clarity demand degree of the pixel point.

[0100] The clarity demand degrees of all pixel points in the latest monitoring video image are used to generate a clarity mask, and the next frame of monitoring video image collected is subjected to pooling processing to achieve the purpose of clarity reduction, and the pooling processing result is transmitted to achieve the purpose of ensuring transmission efficiency while improving the clarity of the monitoring video image of the public security concerned person. The pooling is a known technology, and the specific method is not described here.

[0101] The above only describes the preferred embodiments of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for transmitting image information based on video surveillance, characterized in that, The method includes the following steps: Acquire multiple consecutive frames of surveillance video images; Based on the latest surveillance video images, obtain several content regions and the grayscale difference value of each pixel; based on the positional relationship between the pixel and the edge of the content region and the grayscale difference value of the pixel, obtain the motion probability of each content region; Obtain several bounding boxes and their confidence scores in each frame of the surveillance video image, as well as the matching box for each bounding box; based on the positional relationship between the bounding box and its matching box in the latest surveillance video image, obtain the motion vector and predicted center position of each bounding box; based on the difference between the motion vectors of the bounding box and its matching box, obtain the overall change degree of each pixel in each bounding box; in the latest surveillance video image, obtain the state evaluation image of each pixel; based on the overall change degree of the pixel in its state evaluation image, obtain the state change degree of each pixel; based on the distance between the bounding box and each pixel and the state change degree of the pixel, obtain the state change probability of each bounding box; In the latest surveillance video image, based on the state change probability of the bounding box and combined with the predicted center position, the predicted bounding box of each bounding box is obtained; based on the confidence of the bounding box and the predicted bounding box, combined with the motion probability of the content area, the sharpness requirement of each pixel is obtained; based on the sharpness requirement of the pixel, the sharpness of the next frame of the surveillance video image is adjusted and transmitted. The specific method for obtaining several content regions and the grayscale difference value of each pixel based on the latest surveillance video image is as follows: In the acquired surveillance video images, each frame of the surveillance video image is converted to grayscale to obtain several surveillance grayscale images; the difference image between the surveillance grayscale image corresponding to the latest surveillance video image and the surveillance grayscale image corresponding to the previous frame of the surveillance video image is obtained and recorded as the current difference image. DBSCAN clustering is performed on the grayscale image corresponding to the latest surveillance video image to obtain several grayscale clusters; in the grayscale image corresponding to the latest surveillance video image, a single region composed of pixels belonging to the same grayscale cluster is denoted as a content region. Record any pixel in the grayscale image corresponding to the latest surveillance video image as the target pixel, and record the grayscale value of the target pixel at the corresponding position in the current difference image as the grayscale difference value of the target pixel. The specific method for obtaining the matching box of each bounding box is as follows: For any bounding box in any frame of the surveillance video image, obtain the intersection-union ratio (IUU) of the bounding box with each bounding box in the previous frame of the surveillance video image, and record the bounding box with the highest IUU as the matching box of the bounding box. The method for obtaining the motion vector and predicted center position of each bounding box based on the positional relationship between the bounding box and its matching box in the latest monitoring video image includes the following specific methods: For any bounding box in the latest surveillance video image, the vector from the center position of the matching box of the bounding box to the center position of the bounding box is denoted as the motion vector of the bounding box; The position obtained by moving the center of the bounding box along the motion vector of the bounding box is denoted as the predicted center position of the bounding box. The method for obtaining the overall change degree of each pixel in each bounding box based on the difference in motion vectors between the bounding box and its matching box includes the following specific methods: For any bounding box in any frame of the surveillance video image, the linear normalized result of the magnitude of the difference vector between the motion vector of the bounding box and the motion vector of the matching box in the previous frame of the surveillance video image is denoted as the state change index of the bounding box. Let any frame of the surveillance video image be the target surveillance video image; for any bounding box in the target surveillance video image and any pixel in the bounding box, the state change index of the bounding box is recorded as the change contribution index of the bounding box to the pixel; the average of the change contribution indices of all bounding boxes containing the pixel in the target surveillance video image is recorded as the overall change degree of the pixel in the target surveillance video image. The specific method for obtaining the state evaluation image of each pixel in the latest monitoring video image is as follows: For any pixel in the latest surveillance video image, map that pixel to any frame of the surveillance video image. If the pixel exists in any bounding box of the surveillance video image, then record that surveillance video image as the state evaluation image of that pixel.

2. The image information transmission method based on video surveillance according to claim 1, characterized in that, The method for obtaining the motion probability of each content region based on the positional relationship between the pixel and the edge of the content region, and the grayscale difference value of the pixel, includes the following specific methods: Obtain the grayscale image of each frame of the surveillance video; Record any pixel in the grayscale image corresponding to the latest surveillance video image as the target pixel. Get the content region where the target pixel is located. Among all the edge pixels in the content region, get the edge pixel that is closest to the target pixel and record it as the corresponding edge point of the target pixel. The linear normalization result of the mean gray level difference of all pixels on the path connecting the target pixel to its corresponding edge point is denoted as the motion performance index of the target pixel. The corresponding edge point of the target pixel is recorded as the target edge point, and the target pixel is recorded as the motion performance pixel of the target edge point; The product of the weighted normalized result of the gray-level difference value of the target pixel and the motion performance index of the target pixel is denoted as the motion contribution index of the target pixel; the sum of the motion contribution indices of all motion performance pixels of the target edge point is denoted as the motion performance degree of the target edge point. For any content region, all edge pixels of the content region are clustered by DBSCAN according to the degree of motion of the edge pixels to obtain several edge clusters; An edge interval is defined as the range of edge pixels belonging to the same edge cluster on the edge of the content region; the average value of the motion performance of all edge pixels in each edge interval is defined as the motion index of each edge interval; the edge interval with the largest motion index of the content region is defined as the motion representation interval of the content region. The linearly normalized result of the ratio of the number of edge pixels in the motion representation interval of the content region to the total number of edge pixels in the content region is denoted as the motion representation ratio of the content region. The product of the percentage of motion performance in a content area and the motion index of the motion representation interval of that content area is denoted as the motion probability of that content area.

3. The image information transmission method based on video surveillance according to claim 1, characterized in that, The method for evaluating the overall change in the image based on the state of each pixel to obtain the state change of each pixel includes the following specific methods: For any pixel in the latest surveillance video image, the variance of the overall change of the image in all states of that pixel is denoted as the state variation coefficient of that pixel. The product of the maximum value of the overall change of the image in all states of the pixel and the state change coefficient of the pixel is denoted as the fluctuation change index of the pixel. The mean of the overall change of the image is obtained by evaluating the pixel in all its states. The product of the difference between 1 and the state change coefficient of the pixel and the mean is recorded as the stability change index of the pixel. The sum of the fluctuation index and the stable index of a pixel is recorded as the degree of change in the state of that pixel.

4. The image information transmission method based on video surveillance according to claim 1, characterized in that, The method for obtaining the state change probability of each bounding box based on the distance between the bounding box and each pixel and the degree of state change of the pixel includes the following: For any bounding box and any pixel in the latest surveillance video image, obtain the linearly normalized result of the reciprocal of the Euclidean distance between the center position of the bounding box and the pixel. Multiply the linearly normalized result by the degree of state change of the pixel and record it as the degree of contribution of the pixel to the change of the bounding box. Record the mean of the degree of contribution of all pixels in the latest surveillance video image to the change of the bounding box as the probability of state change of the bounding box.

5. The image information transmission method based on video surveillance according to claim 1, characterized in that, The method for obtaining the predicted bounding box for each bounding box in the latest surveillance video image by combining the state change probability of the bounding box with the predicted center position includes the following specific methods: For any bounding box in the latest surveillance video image, multiply the sum of 1 and the state change probability of the bounding box by the length of the bounding box to obtain the predicted length of the bounding box. Get the predicted width of the bounding box; Using the predicted center position of the bounding box as the center, the predicted bounding box of the bounding box is constructed using the predicted length and predicted width.

6. The image information transmission method based on video surveillance according to claim 1, characterized in that, The method for obtaining the sharpness requirement of each pixel based on the confidence level of the bounding box, the predicted bounding box, and the motion probability of the content region includes the following specific methods: For any pixel in the latest surveillance video image, obtain the difference between 1 and the motion probability of the content region where the pixel is located. Multiply the confidence of the bounding box corresponding to the pixel in the latest surveillance video image by the difference, and record the degree of sharpness improvement of the pixel. The sharpness requirement of a pixel is the sum of the sharpness enhancement level of the pixel and the motion probability of the content area where the pixel is located.

Citation Information

Patent Citations

  • Intelligent human shape recognition alarm system and method based on monitoring camera

    CN117809379A

  • Video transmission system for optimizing 5G network topology structure

    CN118101943A