Dynamic object monitoring method and device, equipment and storage medium

By using HSV color space and optical flow algorithm in dynamic object detection, combined with depth information and pre-trained models, the problem of RGB color space being sensitive to light changes is solved, and low-cost and high-performance dynamic object monitoring is achieved in environments with large changes in light conditions.

CN120279477APending Publication Date: 2025-07-08GUANGDONG MECHANICAL & ELECTRICAL COLLEGE
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510243035.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In environments where lighting conditions are changing greatly, it is difficult to maintain high-performance and low-cost stable monitoring. In particular, traditional RGB color spaces are sensitive to lighting changes, resulting in a decrease in recognition accuracy and it is difficult to flexibly deploy expensive hardware equipment or complex lighting correction technologies.

Method used

Image processing is performed using HSV color space and optical flow algorithms. By acquiring multi-frame HSV images, calculating optical flow fields, combining depth information and pre-trained classification models, dynamic objects are identified and tracked, and the sensitivity to light changes is reduced.

Benefits of technology

In an environment with large changes in lighting conditions, low-cost, stable and accurate dynamic object monitoring is achieved, which improves recognition accuracy and robustness, reduces system hardware costs and improves real-time and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279477A_ABST
    Figure CN120279477A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic object monitoring method and device, equipment and a storage medium, and relates to the technical field of image recognition, and the method comprises the steps: obtaining a plurality of frames of HSV images of a target scene; for any frame of HSV image, calculating an optical flow field of the HSV image based on an image difference between the HSV image and an adjacent frame of HSV image; and determining a dynamic object of the target scene according to the displacement of each pixel point in the optical flow field. The HSV color space and the optical flow algorithm are combined to carry out image processing and dynamic object detection, the situation that the object recognition precision is reduced due to the fact that a traditional RGB color space is sensitive to illumination change is avoided, expensive hardware equipment or a complex illumination correction technology is adopted to overcome the influence of the illumination change, and the object recognition precision is improved. The dynamic object can be monitored stably and accurately at low cost in an environment with large illumination condition changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to a method, device, equipment, and storage medium for monitoring dynamic objects. Background Art

[0002] Currently, in existing dynamic object detection technologies, most systems use image processing methods based on the RGB (Red Green Blue) color space for object recognition. However, since the RGB color space is relatively sensitive to changes in illumination, especially changes in ambient light intensity and light source direction, it is easy to cause image color deviation or brightness changes. Such illumination changes often lead to unstable color features of objects, thereby affecting the recognition accuracy of objects. Furthermore, it causes the dynamic object detection technology based on the RGB color space to perform poorly in environments with large changes in illumination conditions, making it difficult for the RGB image processing method to work stably in places with large changes in illumination conditions. The usual solution is to require expensive hardware devices or complex illumination correction technologies to overcome the influence of illumination changes, but the cost and complexity of these solutions are relatively high, and it is difficult to achieve flexible deployment in practical applications. This not only increases the working burden of the system but also may lead to unstable performance in different environments and time periods. Especially when the system faces complex illumination conditions or rapidly changing light sources, traditional RGB image processing methods often have difficulty maintaining consistent detection accuracy.

[0003] Therefore, how to monitor dynamic objects in a scene while taking into account both low cost and high performance is an urgent problem to be solved currently. Summary of the Invention

[0004] The main objective of this application is to provide a method, device, equipment, and storage medium for monitoring dynamic objects, aiming to solve the technical problem of how to monitor dynamic objects in a scene while taking into account both low cost and high performance.

[0005] To achieve the above objective, this application proposes a method for monitoring dynamic objects, and the method for monitoring dynamic objects includes:

[0006] Obtain multiple frames of HSV images of the target scene;

[0007] For any frame of HSV image, calculate the optical flow field of the HSV image based on the image difference between the HSV image and the adjacent frame of HSV image;

[0008] Determine the dynamic objects in the target scene according to the displacements of the pixel points in the optical flow field.

[0009] In one embodiment, the image difference includes an image gradient and a pixel displacement. The step of calculating the optical flow field of the HSV image based on the image difference between the HSV image and the adjacent-frame HSV image includes:

[0010] Perform multiple downsamplings on the HSV image to obtain multiple images to be processed at different resolutions;

[0011] For any image to be processed, determine the candidate optical flow field of the image to be processed according to the image gradient and pixel displacement between the image to be processed and the adjacent-frame image to be processed, where the adjacent-frame image to be processed is obtained by downsampling the adjacent-frame HSV image and has the same resolution as the image to be processed;

[0012] Perform data fusion on the candidate optical flow fields obtained after traversing each image to be processed to obtain the optical flow field of the HSV image.

[0013] In one embodiment, the step of determining the candidate optical flow field of the image to be processed according to the image gradient and pixel displacement between the image to be processed and the adjacent-frame image to be processed includes:

[0014] Obtain the depth map of the target scene, and perform feature extraction on the depth map to obtain the depth information of the target scene;

[0015] Perform error compensation on the hue information of the image to be processed based on the depth information to obtain the adjusted hue information;

[0016] Calculate the candidate optical flow field of the image to be processed according to the image gradient and pixel displacement between the image to be processed and the adjacent-frame image to be processed, and the saturation information and the adjusted hue information of the image to be processed.

[0017] In one embodiment, the step of obtaining multiple frames of HSV images of the target scene includes:

[0018] Obtain multiple frames of RGB images of the target scene, and convert the color space of each frame of RGB image into the HSV color space to obtain the original HSV images corresponding to each frame of the target scene;

[0019] Filter the pixel points in each frame of the original HSV image according to a preset hue threshold interval and saturation threshold interval to obtain multiple frames of HSV images of the target scene, where the hue values of the pixel points in any one frame of HSV image are within the hue threshold interval, and the saturation values of the pixel points are within the saturation threshold interval.

[0020] In one embodiment, before the step of screening the pixel points in each frame of the original HSV image according to the preset hue threshold range and saturation threshold range, the following steps are included:

[0021] Detect the illumination direction and illumination intensity of the light source in the target scene, as well as the reflectivity of the surfaces of the objects in the target scene, and calculate the reference illumination intensity of the target scene based on the illumination direction, the illumination intensity, and the reflectivity.

[0022] Obtain the actual illumination intensity corresponding to the HSV image, and adjust the lower limits of the preset original hue threshold range and original saturation threshold range based on the ratio between the reference illumination intensity and the actual illumination intensity to obtain the hue threshold range and the saturation threshold range.

[0023] In one embodiment, the step of determining the dynamic objects in the target scene according to the displacements of the pixel points in the optical flow field includes:

[0024] Determine the displacement amount of each pixel point according to the optical flow field.

[0025] In the case where the displacement amount of the target pixel point group among the pixel points exceeds a preset displacement threshold, it is determined that there are dynamic objects in the target scene, where the target pixel point group is a plurality of adjacent target pixel points among the pixel points, and the displacement amount of each target pixel point exceeds the displacement threshold.

[0026] In one embodiment, after the step of determining that there are dynamic objects in the target scene, the following steps are included:

[0027] Extract the target pixel point group and input the target pixel point group into a pre-trained classification model, where the pre-trained classification model is trained based on training samples, the training samples are composed of sample features and sample labels, the sample features include the hue information, saturation information, and depth information of the sample object, and the sample label is the object type of the sample object.

[0028] Identify the target pixel point group through the pre-trained classification model.

[0029] In addition, to achieve the above object, the present application also proposes a dynamic object monitoring device, and the dynamic object monitoring device includes:

[0030] An acquisition module, configured to acquire multiple frames of HSV images of a target scene;

[0031] A calculation module, configured to calculate the optical flow field of any frame of HSV image based on the image difference between the HSV image and the adjacent frame of HSV image.

[0032] A determination module, configured to determine a dynamic object of the target scene according to displacements of pixel points in the optical flow field.

[0033] In addition, to achieve the above object, the present application further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the dynamic object monitoring method as described above.

[0034] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the dynamic object monitoring method as described above.

[0035] One or more technical solutions provided by the present application have at least the following technical effects:

[0036] The present application first obtains multiple frames of HSV (Hue Saturation Value) images of a target scene, and effectively eliminates the interference of light changes (such as changes in ambient light intensity and light source direction) on object recognition based on the strong adaptability of the hue and saturation channels in the HSV color space to light changes; then for any frame of HSV image, based on the image difference between the HSV image and the adjacent frame of HSV image, the optical flow field of the HSV image is calculated to analyze the movement of pixel points between consecutive frames, so as to accurately identify and track dynamic objects; finally, according to the displacements of pixel points in the optical flow field, the dynamic objects of the target scene are determined, ensuring real-time monitoring and efficient object tracking in a changing lighting environment.

[0037] In summary, by combining the HSV color space and the optical flow algorithm for image processing and dynamic object detection, the present application avoids the problem that the traditional RGB color space is sensitive to light changes, resulting in a decrease in object recognition accuracy, and the problem of using expensive hardware devices or complex light correction technologies to overcome the influence of light changes, and realizes low-cost, stable and accurate monitoring of dynamic objects in an environment with large changes in lighting conditions. Description of the Drawings

[0038] The drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0039] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a schematic flowchart provided for the first embodiment of the dynamic object monitoring method of the present application;

[0041] Figure 2 It is a schematic flowchart provided for the second embodiment of the dynamic object monitoring method of the present application;

[0042] Figure 3 It is a brief schematic flowchart of the dynamic object monitoring method provided for the second embodiment of the present application;

[0043] Figure 4 It is a schematic module structure diagram of the dynamic object monitoring device in the embodiment of the present application;

[0044] Figure 5 It is a schematic device structure diagram of the hardware operating environment involved in the dynamic object monitoring method in the embodiment of the present application.

[0045] The implementation, functional features, and advantages of the object of the present application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Specific Embodiments

[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0047] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0048] The main solution of the embodiment of the present application is: obtaining multiple frames of HSV images of the target scene; for any frame of HSV image, calculating the optical flow field of the HSV image based on the image difference between the HSV image and the adjacent frame of HSV image; and determining the dynamic objects in the target scene according to the displacements of the pixel points in the optical flow field.

[0049] In the existing dynamic object detection technology, most systems use image processing methods based on RGB (Red Green Blue) color space for object recognition. However, since the RGB color space is sensitive to changes in illumination, especially changes in ambient light intensity and light source direction, it is easy to cause image color deviation or brightness changes. Such illumination changes often lead to unstable color features of objects, thereby affecting the recognition accuracy of objects, which in turn leads to poor performance of dynamic object detection technology based on RGB color space in environments with large changes in illumination conditions, making it difficult for RGB image processing methods to work stably in places with large changes in illumination conditions. The usual solution is to require expensive hardware equipment or complex illumination correction technology to overcome the impact of illumination changes, but these solutions are costly and complex, and are difficult to implement flexible deployment in practical applications, which not only increases the workload of the system, but also may lead to unstable performance in different environments and time periods. In particular, when the system faces complex illumination conditions or rapidly changing light sources, traditional RGB image processing methods often find it difficult to maintain consistent detection accuracy. Therefore, how to monitor dynamic objects in the scene while taking into account both low cost and high performance is a problem that needs to be solved urgently.

[0050] The present application provides a solution by combining the HSV color space and the optical flow algorithm for image processing and dynamic object detection, thereby avoiding the problem that the traditional RGB color space is sensitive to lighting changes, thereby reducing the accuracy of object recognition, and using expensive hardware equipment or complex lighting correction technology to overcome the impact of lighting changes. This achieves low-cost, stable and accurate monitoring of dynamic objects in an environment with large changes in lighting conditions.

[0051] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device that can realize the above functions. The following takes an electronic device as an example to illustrate this embodiment and the following embodiments.

[0052] Based on this, the present application embodiment provides a dynamic object monitoring method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the dynamic object monitoring method of the present application.

[0053] In this embodiment, the dynamic object monitoring method includes steps S10 to S30:

[0054] Step S10, acquiring multiple frames of HSV images of the target scene;

[0055] It is understandable that since the RGB color space is relatively sensitive to changes in light, especially changes in ambient light intensity and light source direction, it is prone to cause color deviation or brightness change in the image. This kind of light change often leads to unstable color features of objects, thus affecting the recognition accuracy of objects. Therefore, step S10 is carried out. Due to the strong adaptability of the hue and saturation channels in the HSV color space to light changes, it can avoid the problems of RGB image color deviation or brightness change caused by changes in ambient light intensity and light source direction, which affect the stability of object color features, effectively eliminating the interference of light changes on object recognition and providing a highly stable data basis for object recognition.

[0056] Exemplarily, first, a high-frame-rate camera is used to collect video of the target scene, and the collected video stream images are transmitted to the Raspberry Pi processing unit in real time by the camera. During the image transmission process, first, noise removal and preprocessing are performed on the original image to remove some noise caused by environmental interference. Gaussian filtering or median filtering is used to smooth the image and reduce the impact of noise on subsequent processing. A series of RGB images are obtained in the Raspberry Pi processing unit. Then, using an image processing software or an image processing library in a programming language (such as OpenCV), each frame of RGB image is converted into the HSV color space through a color space conversion function. This conversion process calculates the corresponding hue (H), saturation (S), and value (V) values by normalizing the RGB values and according to specific conversion formulas, thereby obtaining each frame of HSV image corresponding to the target scene. For example:

[0057]

[0058] V = max(R, G, B)

[0059] where R, G, and B are the pixel values of the red, green, and blue channels of the RGB image respectively. Through the above conversion formulas, the hue and saturation information of the image will be able to effectively resist light changes and maintain the recognition ability of objects. These HSV images retain the color information of the original scene while reducing the impact of light changes on subsequent processing.

[0060] In addition, using the Raspberry Pi as an edge computing platform and combining it with a low-cost camera, a low-power and high-performance dynamic object detection system is realized. As a computing platform, the Raspberry Pi not only has the advantage of flexible deployment but also can be quickly installed and configured in different monitoring places, reducing the hardware cost of the system.

[0061] Step S20, for any frame of HSV image, calculate the optical flow field of the HSV image based on the image difference between the HSV image and the adjacent frame of HSV image;

[0062] It should be noted that adjacent frame HSV images refer to two or more consecutive HSV images in a time series; the optical flow field of an HSV image refers to the displacement vector field of each pixel point between consecutive frames in the image represented in the HSV color space, which reflects the motion direction and speed of the pixel points.

[0063] It can be understood that since the existing motion detection algorithms often have poor performance in real-time monitoring of changing lighting environments and are difficult to accurately identify and track dynamic objects, step S20 is performed. By calculating the optical flow field of the HSV image to describe the motion trend of pixel points in the image, the motion information of dynamic objects can be captured, avoiding the problem that object detection relying only on single-frame image information is easily affected by factors such as noise and occlusion, and providing effective motion information for the detection and tracking of dynamic objects.

[0064] Exemplarily, any one frame of HSV image is selected as the current frame and compared with the previous or next adjacent frame of HSV image in the time series. Using an optical flow calculation algorithm (such as the Lucas-Kanade optical flow algorithm), based on the pixel differences between the two frames of images, the optical flow field of the current frame is calculated. Specifically, the algorithm will first select a series of feature points in the current frame and the adjacent frame, and then solve the displacement vectors of these feature points between the two frames through an iterative optimization method, and finally obtain an optical flow field containing the displacement vectors of all pixel points. This optical flow field reflects the motion trend and speed of the objects in the image.

[0065] In a feasible implementation manner, the image differences in step S20 include image gradients and pixel point displacements. The step of calculating the optical flow field of the HSV image based on the image differences between the HSV image and the adjacent frame HSV image may include steps S21 to S23:

[0066] Step S21, perform multiple downsamplings on the HSV image to obtain multiple images to be processed at different resolutions;

[0067] It should be noted that an image to be processed refers to an HSV image with a resolution smaller than the original image after downsampling processing, which is used for subsequent optical flow field calculation.

[0068] It can be understood that directly calculating the optical flow field on a high-definition image will result in excessive computational complexity and poor real-time performance, and it is also difficult to capture the motion of objects at large and small scales simultaneously. Therefore, step S21 is performed. Through the downsampling operation, the problems of large computational complexity and poor real-time performance caused by directly calculating the optical flow field on a high-definition image can be avoided, thereby improving the efficiency of optical flow field calculation. And multiple downsampling operations generate images with different resolutions, which can avoid the inability to capture the motion at different scales and ensure the integrity of multi-scale motion information.

[0069] Exemplarily, downsampling methods such as Gaussian pyramid or mean pyramid are adopted to perform layer-by-layer downsampling on the original HSV image. For example, in each layer of downsampling, the width and height of the image are both reduced to half of the original, thereby obtaining a series of processed images with different resolutions. In specific operations, the downsampling function in an image processing library (such as OpenCV) can be used to perform downsampling on each channel of the HSV image separately to ensure that color information remains consistent during the resolution reduction process. Assuming the original image is I0 and the image of the i-th layer of its pyramid structure is I i , multiple downsamplings are performed through the following formula:

[0070] I i = f(I i-1 ), i = 1, 2, …, N

[0071] where f(·) is the downsampling operation and N is the number of pyramid layers. In this way, a sequence of processed images with different resolutions from high to low is generated, providing a data basis for subsequent multi-scale optical flow field calculation.

[0072] Step S22: For any processed image, determine the candidate optical flow field of the processed image according to the image gradient and pixel point displacement between the processed image and the adjacent-frame processed image, where the adjacent-frame processed image is obtained by downsampling the adjacent-frame HSV image and has the same resolution as the processed image;

[0073] It should be noted that the adjacent-frame processed image refers to the processed image obtained by performing the same downsampling process on the adjacent-frame HSV image; the image gradient refers to the gray-scale difference between adjacent pixel points in the image, reflecting the edge and texture information of the image; the pixel point displacement refers to the position change of the same pixel point between consecutive frames, which can be represented by a displacement vector; the candidate optical flow field refers to the optical flow field calculated based on the processed image, reflecting the motion information of the objects in the image at this resolution.

[0074] It can be understood that since the image gradient and pixel point displacement are key information for calculating the optical flow field, and the motion speed and direction of the objects in the image can be estimated through this information, so step S22 is carried out. By combining the image gradient and pixel point displacement information, the problems of insufficient accuracy and poor robustness caused by calculating the optical flow field relying only on a single information source can be avoided, thereby improving the accuracy and stability of the optical flow field calculation.

[0075] Exemplarily, for any image to be processed, first calculate the image gradient between it and the corresponding image to be processed in the adjacent frame. This can be achieved by applying the Sobel operator or the gradient operator to obtain the gradient components in the horizontal and vertical directions. At the same time, use the optical flow method (such as the Lucas-Kanade optical flow method) to estimate the displacement vector of pixel points between consecutive frames. For example: At each layer of the pyramid, use the optical flow method to calculate the object motion between adjacent frames. Assume that in pyramid layer i, the images of two adjacent frames are I i (t) and I i (t + 1), and the optical flow can be calculated by the following equation:

[0076]

[0077] where, is the image gradient, and Δx is the displacement of the object. By solving this equation, the optical flow field of each layer of the image can be obtained. In specific operations, the optical flow method can be implemented in an image processing library. Input the image to be processed and its adjacent frame image to be processed, and output the displacement vector of each pixel point. Combining the image gradient and the displacement vector, determine the optical flow field of the image to be processed by minimizing the optical flow constraint equation. This process can be performed in parallel on images to be processed at different resolutions to improve the calculation efficiency.

[0078] In a feasible implementation manner, step S22 may include steps S221 to S223:

[0079] Step S221, obtain the depth map of the target scene, and perform feature extraction on the depth map to obtain the depth information of the target scene;

[0080] It should be noted that the depth map of the target scene refers to a grayscale or color image obtained through a depth sensor (such as Kinect, LiDAR, etc.), which represents the distance between each point in the scene and the sensor; the depth information of the target scene refers to the key features extracted from the depth map, which describe the three-dimensional position and structure of objects in the scene, such as depth values, gradients, edges, etc.

[0081] It can be understood that since the depth map provides the three-dimensional position information of objects in the scene, which helps to more accurately understand the scene structure and object motion, so by performing step S221 and extracting the depth information, it is possible to avoid problems such as incorrect target recognition or inaccurate motion estimation that may occur when relying solely on two-dimensional image information for object detection due to the lack of depth information. Thus, it enhances the system's ability to understand the three-dimensional structure of the scene and provides an important reference basis for subsequent hue information error compensation and optical flow field calculation.

[0082] Exemplarily, a depth camera (such as Kinect) is used to scan the target scene to obtain a depth map containing the depth information of the scene. Then, an edge detection algorithm (such as the Canny operator) is used to extract edges from the depth map to highlight the object contours. At the same time, a depth gradient calculation method is applied to analyze the change trend of the depth values of each pixel point in the depth map, so as to extract the depth gradient features. In addition, region growing or segmentation algorithms can also be used to segment the depth map into different object regions, and further extract the depth statistical features of each region (such as average depth, depth variance, etc.). Through these feature extraction means, the depth information of the target scene is finally obtained, providing a three-dimensional structure reference for subsequent processing.

[0083] Step S222, perform error compensation on the hue information of the image to be processed based on the depth information to obtain the adjusted hue information;

[0084] It should be noted that the hue information of the image to be processed refers to the hue (H) component of the image to be processed in the HSV color space, representing the basic attribute of the color.

[0085] It can be understood that since factors such as illumination changes and object materials may affect the accuracy of the hue information, and the depth information can provide additional constraints to help correct these errors, so step S222 is performed. Through error compensation, the problem of incorrect object recognition or optical flow field calculation deviation caused by inaccurate hue information can be avoided, thereby improving the accuracy and stability of the hue information and providing a more reliable data basis for subsequent optical flow field calculations.

[0086] Exemplarily, first convert the image to be processed from the RGB color space to the HSV color space, and extract the hue (H) component as the hue information. Then, using the obtained depth information, analyze the relationship between the depth change on the object surface and the hue change, and establish an error compensation model. For example, regression analysis or machine learning methods can be used to predict and correct the hue value according to the depth features. Specifically, for each pixel point, a correction coefficient is calculated based on its depth value and the depth information of the surrounding pixels, and this coefficient is applied to the original hue value to obtain the adjusted hue information. For example: use the depth information D(x, Y) to compensate for the error of the hue information H(x, y) caused by illumination changes:

[0087]

[0088] where α is the illumination compensation factor, representing the influence of depth change on the hue, is the adjusted hue information. This method effectively reduces the influence of factors such as illumination and material on the hue information and improves the accuracy of the hue information.

[0089] Step S223: Calculate the candidate optical flow field of the image to be processed based on the image gradient and pixel displacement between the image to be processed and the adjacent-frame images to be processed, as well as the saturation information and adjusted hue information of the image to be processed.

[0090] It can be understood that since optical flow field calculation needs to comprehensively consider image gradient, pixel displacement, and color information to accurately estimate the motion of an object, step S223 is performed. By comprehensively considering various information, it is possible to avoid the problem of inaccurate optical flow field caused by insufficient information or noise interference when relying solely on a single information source for optical flow field calculation, thereby improving the accuracy and robustness of optical flow field calculation and providing a more reliable motion estimate for dynamic object detection.

[0091] Exemplarily, after obtaining the adjusted hue information, perform image gradient calculation on the image to be processed and its adjacent-frame images to be processed, and use the Sobel operator or Gaussian gradient operator to extract the gradient components in the horizontal and vertical directions. At the same time, adopt an optical flow estimation algorithm (such as the Lucas-Kanade method) to calculate the preliminary optical flow field according to the displacement of pixel points between adjacent frames. Then, combine the saturation information of the image to be processed and the adjusted hue information to optimize the preliminary optical flow field. Specifically, an energy function can be constructed, taking image gradient, pixel displacement, saturation information, and hue information as constraint conditions, and obtaining a more accurate optical flow field by minimizing the energy function. Finally, through this multi-information fusion method, calculate the candidate optical flow field of the image to be processed, providing a reliable motion estimate for dynamic object detection.

[0092] In this embodiment, by performing depth map acquisition and feature extraction, hue information error compensation, and multi-information fusion optical flow field calculation, the problems of inaccurate target recognition, deviation in motion estimation, and poor robustness caused by the lack of depth information, hue error, and reliance on a single information source in traditional dynamic object monitoring are avoided, achieving the technical effects of improving the accuracy, stability, and real-time performance of dynamic object detection in complex lighting and dynamic scenarios. Specifically, the introduction of the depth map provides three-dimensional structure information of the scene, enhancing the accuracy of object recognition; the error compensation of hue information reduces the influence of factors such as lighting and material on color information, improving the reliability of color information; and the multi-information fusion optical flow field calculation comprehensively utilizes image gradient, pixel displacement, saturation information, and adjusted hue information, improving the accuracy and robustness of motion estimation, thereby realizing more efficient and accurate dynamic object detection.

[0093] Step S23: Perform data fusion on the candidate optical flow fields obtained after traversing each image to be processed to obtain the optical flow field of the HSV image.

[0094] It can be understood that since the optical flow fields at different resolutions contain motion information at different scales, these information can be integrated through data fusion to obtain a more complete and accurate optical flow field. Therefore, step S23 is performed. Through data fusion, a complete optical flow field containing multi-scale motion information is obtained, which can avoid the problems of incomplete optical flow field information and easy loss of small-scale motion information at a single resolution, thereby improving the accuracy and robustness of dynamic object detection.

[0095] Exemplarily, after obtaining the candidate optical flow fields of each image to be processed, a data fusion technique is used to integrate them into the optical flow field of the original HSV image. A specific method is to first perform weighted averaging on the optical flow vectors of each pixel point at different resolutions, and the weights can be set according to the confidence level of the optical flow vectors or the resolution levels. For example, a higher-resolution optical flow field can be assigned a higher weight. Then, an interpolation method (such as bilinear interpolation) is used to restore the fused optical flow field to the resolution of the original image. For example: Assume u i and v i are the horizontal and vertical optical flow components of the image at the i-th layer of the pyramid respectively, then the total optical flow can be fused by weighted averaging:

[0096]

[0097] where w i is the weight of each layer of the pyramid. Finally, through this multi-scale data fusion method, a complete optical flow field that contains both large-scale motion information and retains small-scale details is obtained, providing accurate motion estimation for dynamic object detection.

[0098] In this embodiment, through multi-scale downsampling, the problems that the calculation of the optical flow field at a single resolution may lose small-scale motion information or be interfered by large-scale noise are avoided; through the calculation of the optical flow field based on image gradients and pixel displacements, the problems of large computational complexity and poor real-time performance caused by directly calculating the optical flow field on the original high-resolution image are avoided; through data fusion, the problems of incomplete or inaccurate information that may exist in a single optical flow field are avoided. It realizes obtaining a complete and accurate optical flow field containing multi-scale motion information while ensuring the calculation efficiency, thereby improving the accuracy and robustness of dynamic object detection, enabling the system to stably and efficiently monitor dynamic objects in complex lighting and dynamic scenes.

[0099] Step S30, determine the dynamic objects in the target scene according to the displacements of the pixel points in the optical flow field.

[0100] It should be noted that the displacement of a pixel point refers to the position change of a certain pixel point between consecutive frames in the optical flow field, usually represented by a displacement vector, including the direction and magnitude of the displacement.

[0101] It can be understood that during the movement of a dynamic object, the corresponding pixel points will show significant displacements in the optical flow field. By analyzing these displacements, dynamic objects can be detected. Therefore, performing step S30 can avoid the problem in traditional methods that object detection only relies on static features such as color and shape and is easily affected by factors such as illumination and background, and can accurately and quickly detect dynamic objects in the target scene through the displacement information of pixel points in the optical flow field, improving the accuracy and robustness of detection.

[0102] Exemplarily, analyze the calculated optical flow field and focus on the displacement vectors of pixel points. Set a displacement threshold. When the displacement vector of a certain pixel point in the optical flow field exceeds this threshold, it is considered that there is a dynamic object in the area corresponding to this pixel point. Further, adjacent dynamic pixel points can be clustered together through a clustering algorithm to form candidate regions of dynamic objects. Perform subsequent morphological processing and classification recognition on these candidate regions to finally determine the dynamic objects in the target scene. This method utilizes the displacement information in the optical flow field, effectively distinguishes dynamic objects from the static background, and realizes the accurate detection of dynamic objects.

[0103] In a feasible implementation manner, step S30 may include steps S31 to S32:

[0104] Step S31, determining the displacement amount of each pixel point according to the optical flow field;

[0105] It can be understood that since the optical flow field reflects the movement trend of each pixel point in the image, the displacement amount of each pixel point can be quantitatively determined by analyzing the optical flow field, which is the basis for judging whether an object is moving and the degree of movement. Therefore, performing step S31 can avoid the problems of subjective error and uncertainty that may be caused by simply visually judging the movement of an object by accurately calculating the displacement amount of each pixel point, thus providing objective and accurate data support for subsequent judgment of dynamic objects.

[0106] Exemplarily, utilize the calculated optical flow field to analyze each pixel point in the image. Each vector in the optical flow field represents the movement direction and speed of the corresponding pixel point. By extracting the length of these vectors, the displacement amount of each pixel point can be obtained. In specific operations, all vectors in the optical flow field can be traversed, and the Euclidean norm of the vector (i.e., the length of the vector) can be used to represent the displacement amount. This process can be implemented through programming. For example, the NumPy library in Python can be used to calculate the length of the vector. Finally, a displacement amount value is assigned to each pixel point in the image to form a displacement amount map for subsequent dynamic object detection.

[0107] Step S32: When the displacement amount of the target pixel point group among the respective pixel points exceeds a preset displacement threshold, it is determined that there is a dynamic object in the target scene, where the target pixel point group is a plurality of adjacent target pixel points among the respective pixel points, and the displacement amount of each target pixel point exceeds the displacement threshold.

[0108] It should be noted that the target pixel point group refers to a group of adjacent pixel points in the image. These pixel points not only have obvious displacements themselves, but also their displacement amounts exceed the preset displacement threshold, jointly constituting a region representing the movement of the object.

[0109] It can be understood that since the displacement of a single pixel point may be caused by noise or local micro-movement and is not sufficient to judge the movement of the entire object, while the pixel point group, as a group of adjacent pixel points with displacement amounts exceeding the threshold, can better represent the overall movement of the object. Therefore, in step S32, by setting the displacement threshold and considering the overall movement of the target pixel point group, it is possible to avoid misjudging the dynamic object due to the displacement of a single pixel point or local noise interference, thereby improving the accuracy and robustness of dynamic object detection.

[0110] Exemplarily, first, a reasonable displacement threshold is set, which is determined according to the actual application scenario and requirements and is used to distinguish static and moving pixel points. Then, the displacement map is traversed to find a plurality of adjacent pixel points whose displacement amounts exceed the preset displacement threshold, forming a target pixel point group. Specifically, in implementation, region growing or connected component labeling algorithms can be used to identify these target pixel point groups. Once such a target pixel point group is detected, it can be determined that there is a dynamic object in the target scene. For example, a window can be set in the displacement map, and the window is slid to traverse the image to check whether the pixel points in the window meet the conditions of the target pixel point group. If they meet, this region is recorded as the dynamic object region. In this way, dynamic objects can be effectively detected from the image.

[0111] In this embodiment, by using optical flow field analysis to determine the pixel point displacement amount and combining with the displacement threshold to judge the target pixel point group, the problems of false detection and missed detection caused by misjudgment of a single pixel point displacement or local noise interference in traditional dynamic object detection are avoided, and the technical effect of accurately and stably detecting dynamic objects in a complex dynamic scene is achieved. Specifically, the displacement amount of each pixel point is calculated using the optical flow field, providing objective and accurate motion information; and by setting the displacement threshold and identifying a plurality of adjacent target pixel points to form a target pixel point group, the real moving object and background noise are effectively distinguished, improving the robustness and accuracy of dynamic object detection, thereby achieving efficient and reliable dynamic object detection in a complex environment.

[0112] This embodiment provides a method for monitoring dynamic objects. By combining the HSV color space and the optical flow algorithm for image processing and dynamic object detection, it avoids the problem that the traditional RGB color space is sensitive to light changes, resulting in a decrease in object recognition accuracy, and the problem of using expensive hardware devices or complex light correction techniques to overcome the influence of light changes. It realizes the low-cost, stable and accurate monitoring of dynamic objects in an environment with large light condition changes.

[0113] In a feasible implementation manner, after the step of determining that there are dynamic objects in the target scene in step S32, steps S33 to S34 may further be included:

[0114] Step S33, extract the target pixel point group and input the target pixel point group into a pre-trained classification model, where the pre-trained classification model is trained based on training samples, the training samples are composed of sample features and sample labels, the sample features include the hue information, saturation information and depth information of the sample object, and the sample label is the object type of the sample object;

[0115] It should be noted that the pre-trained classification model refers to a deep learning model for object classification that has been pre-trained on a large amount of training data, such as a convolutional neural network (CNN), etc. This model establishes a mapping relationship from image features to object categories by learning the hue, saturation and depth information in the training data; the training samples refer to the data sets used to train the pre-trained classification model, which contain objects under various different environments and light conditions, as well as the corresponding hue, saturation and depth information, and are used to train the model to achieve accurate recognition of different types of objects.

[0116] It can be understood that since the determined target pixel point group represents a potential dynamic object area, further identification and classification are required. The pre-trained classification model has learned how to distinguish different types of objects based on hue, saturation and depth information. Therefore, by performing step S33, extracting and inputting the target pixel point group into the pre-trained classification model, it can avoid the problems of insufficient feature representation and inaccurate classification that may occur when directly performing object recognition based on the original image or simple features, thereby making full use of the complex feature representation and classification capabilities learned by the model and improving the accuracy and efficiency of object recognition.

[0117] Exemplarily, the region where the target pixel point group is located is extracted from the image to form a series of small image patches. At the same time, a pre-trained classification model is prepared, which is pre-trained using a large number of training sample data, including hue information, saturation information, and depth information, as well as the corresponding sample labels. During the training process, the model learns how to distinguish different types of objects based on this information. The extracted target pixel point group image patches are input into the pre-trained classification model, and the model will perform feature extraction and preliminary classification on them.

[0118] Step S34, identify the target pixel point group through the pre-trained classification model.

[0119] It can be understood that since the pre-trained classification model has been trained with a large amount of training data and already has strong object recognition capabilities, and by inputting the target pixel point group into the model, accurate classification of potential dynamic objects can be achieved. Therefore, step S34 is carried out. Through the recognition of the pre-trained classification model, the problems of poor recognition effect and weak generalization ability caused by the possible need to manually design features or use simple classifiers in traditional methods can be avoided, thus achieving accurate and rapid classification of dynamic objects and providing a reliable basis for subsequent tracking, analysis, and other processes.

[0120] Exemplarily, after the target pixel point group image patches are input into the pre-trained classification model, the model will perform further processing on them. First, the model extracts features from the image patches through structures such as convolutional layers and pooling layers to obtain high-dimensional feature representations, and then inputs these features into the fully connected layer for classification prediction, outputting the probabilities of each image patch belonging to different classes. Finally, according to the probability distribution, the class with the highest probability is selected as the recognition result of the image patch. In this way, the pre-trained classification model can achieve accurate recognition of the target pixel point group, thereby judging the dynamic objects and their types in the image. In addition, when the system detects an abnormal object, it will control alarm devices (such as buzzers or LED lights) to emit alarms through the Raspberry Pi, and at the same time, it will be linked with the monitoring center through the network to notify the management personnel. The management personnel can view the real-time image or video stream through the remote monitoring platform to further determine whether measures need to be taken.

[0121] In this embodiment, by adopting a pre-trained classification model in deep learning and combining hue information, saturation information, and depth information, feature extraction and classification recognition are performed on the target pixel point group, avoiding the problems of insufficient feature representation ability, low classification accuracy, and poor generalization ability in traditional dynamic object detection methods, and achieving the technical effects of efficient and accurate recognition of dynamic objects in complex scenarios. Specifically, through learning on a large-scale training scenario data, the pre-trained classification model has obtained powerful feature extraction and classification capabilities, can fully exploit the multi-dimensional information in the target pixel point group, effectively distinguish different types of dynamic objects, thereby improving the accuracy and robustness of dynamic object detection, and achieving efficient and reliable dynamic object recognition in complex environments.

[0122] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , step S10 may include steps S11 to S12:

[0123] Step S11, obtaining multiple frames of RGB images of the target scene, and converting the color space of each frame of RGB image into the HSV color space to obtain each frame of original HSV image corresponding to the target scene;

[0124] Step S12, screening the pixel points in each frame of the original HSV image according to a preset hue threshold interval and saturation threshold interval to obtain multiple frames of HSV images of the target scene, wherein the hue value of each pixel point in any one frame of HSV image is within the hue threshold interval, and the saturation value of each pixel point is within the saturation threshold interval.

[0125] It should be noted that multiple frames of HSV images of the target scene refer to the images obtained after screening by the hue and saturation thresholds, which only contain pixel points that meet the preset threshold interval, and these pixel points are likely to belong to the target object.

[0126] It can be understood that since different objects usually have specific hue and saturation characteristics. By setting specific hue and saturation threshold intervals, pixel points that meet these characteristics can be screened out, thereby highlighting potential target objects. Therefore, step S12 is performed. Through the screening of hue and saturation, the problem of inaccurate recognition of target objects caused by the interference of background pixels in a complex background is avoided, and the target HSV image only containing pixel points with specific hue and saturation is obtained, simplifying subsequent processing and improving the recognition accuracy and processing efficiency of the target object.

[0127] Exemplarily, first, according to the color characteristics of the target object, a hue threshold range and a saturation threshold range are preset. For example, if the target object is red, the hue threshold range can be set as [0, 10] ∪ [340, 360] (considering the cyclicity of the hue circle), and the saturation threshold range can be set as [100, 255]. Then, each pixel point in the HSV image is traversed to check whether its hue value and saturation value are both within the preset threshold range. If so, the pixel point is retained in the target HSV image; otherwise, the pixel point is set to black or other background colors. In this way, the pixel points that conform to the color characteristics of the target object are screened out, and multiple frames of HSV images of the target scene are obtained. After obtaining the target HSV image, the original HSV image is directly replaced with this HSV image to obtain a new HSV image. Then, based on this new HSV image, the optical flow field calculation step is performed, that is, the image difference between the new HSV image and the adjacent frame HSV image is calculated, and the optical flow field is calculated accordingly.

[0128] Exemplarily, since the filtered target HSV image highlights the target object more, but may lose some background information, in order to restore the background information as much as possible while retaining the information of the target object, after obtaining multiple frames of HSV images of the target scene, they are fused with the original HSV image to update the original HSV image, which can avoid the problem of inaccurate optical flow field calculation caused by incomplete information when calculating the optical flow field only based on the filtered HSV image, thereby obtaining a new HSV image that not only highlights the target object but also retains some background information. Specifically, a weighted fusion method can be adopted, giving a higher weight to the pixel points in the filtered HSV image and a lower weight to the background pixel points in the original HSV image. For example, the weight of the filtered HSV image can be set to 0.7, and the weight of the original HSV image can be set to 0.3. Then, the values of the corresponding pixel points of the two images are weighted and averaged to obtain a new HSV image. Based on this new HSV image, the optical flow field calculation step is performed, that is, the image difference between the new HSV image and the adjacent frame HSV image is calculated, and the optical flow field is calculated accordingly. In this way, both the information of the target object is retained and some background information is restored, improving the accuracy and robustness of the optical flow field calculation.

[0129] In this embodiment, by presetting the hue and saturation threshold ranges for screening, the problem of inaccurate recognition of the target object caused by background interference in a complex background is avoided, further achieving the technical effect of accurately and efficiently detecting and tracking dynamic objects in a complex background, and improving the accuracy and robustness of dynamic object detection.

[0130] In a feasible implementation manner, steps S100 to S200 may further be included before step S12:

[0131] Step S100: Detect the illumination direction and intensity of the light source in the target scene, as well as the reflectivity of the surfaces of various objects in the target scene, and calculate the reference illumination intensity of the target scene based on the illumination direction, the illumination intensity, and the reflectivity.

[0132] It should be noted that the illumination direction refers to the propagation direction of the light rays emitted by the light source in space; the illumination intensity refers to the brightness or energy intensity of the light rays emitted by the light source; the reflectivity of the object surface refers to the ability of the object surface to reflect light rays, which is usually related to the material and surface characteristics of the object; the reference illumination intensity of the target scene refers to the theoretical illumination intensity calculated based on the illumination direction, the illumination intensity, and the object surface reflectivity, and is used as a reference benchmark for the actual illumination intensity.

[0133] It can be understood that since the illumination direction and intensity directly affect the brightness and color performance of the object surface, and the reflectivity of the object surface determines its light reflection ability, performing Step S100 and detecting these parameters can avoid the problem of inaccurate color recognition of the target object caused by ignoring the changes in the actual illumination conditions, can more accurately simulate and calculate the theoretical illumination conditions of the target scene, provide a benchmark for the subsequent adjustment of the hue and saturation thresholds, make the threshold setting more in line with the actual illumination conditions, and improve the accuracy of color screening.

[0134] Exemplarily, first, use a combination device of a light sensor and a direction sensor to scan the target scene to detect the illumination direction and intensity of the light source. For example, measure the intensity of the light rays through the light sensor, while the direction sensor determines the source direction of the light rays. At the same time, use a spectral reflectometer to measure the reflectivity of the surfaces of various objects in the scene to obtain the reflectivity data of different materials and colors. Then input these measurement data into a pre-established illumination model, which takes into account factors such as the illumination direction, the illumination intensity, and the object surface reflectivity, and calculate the reference illumination intensity I of the target scene ref :

[0135] I ref (x,y) = R(x,y) · I light · cos(d light · n(x,y))

[0136] where (x,y) is the pixel position coordinate in the target scene, R(x,y) is the reflectivity of the object surface, I light is the illumination direction of the light source, d light is the illumination intensity of the light source, and n(x,y) is the normal vector of the object surface, which reflects the angular relationship between the surface and the light source direction. This reference illumination intensity represents the illumination effect that the objects in the scene should exhibit under ideal conditions and provides a benchmark for subsequent image processing.

[0137] Step S200: Obtain the actual illumination intensity corresponding to the HSV image, and adjust the lower limits of the preset original hue threshold interval and original saturation threshold interval based on the ratio between the reference illumination intensity and the actual illumination intensity, so as to obtain the hue threshold interval and the saturation threshold interval.

[0138] It should be noted that the actual illumination intensity refers to the illumination intensity reflected in the HSV image under actual shooting conditions, and can usually be estimated or measured through image processing techniques.

[0139] It can be understood that since the actual illumination conditions may differ from the reference illumination conditions, this difference will affect the color performance of objects in the image. Therefore, by performing Step S200 and compensating for this difference by adjusting the threshold interval, it is possible to avoid problems such as incorrect color recognition or missed detection of the target object caused by the mismatch between the actual illumination conditions and the reference illumination conditions, thereby improving the adaptability and accuracy of color screening and enabling the dynamic object detection to maintain stable performance under complex and variable illumination conditions.

[0140] Exemplarily, after obtaining the HSV image, first use an image processing algorithm, such as histogram equalization or a brightness-based statistical method, to estimate the actual illumination intensity in the image. This actual illumination intensity reflects the true illumination situation of the scene under the current shooting conditions. Then compare the calculated reference illumination intensity with the actual illumination intensity to obtain the ratio between the two. Based on this ratio, dynamically adjust the lower limits of the preset original hue threshold interval and original saturation threshold interval. For example, if the actual illumination intensity is lower than the reference illumination intensity, the threshold lower limits of hue and saturation can be appropriately reduced to adapt to color changes in a dark environment; conversely, the threshold lower limits are increased. For example: assume that the illumination intensity of the current image is I current , and the reference illumination I ref has been estimated through a reflection model, then the lower limit H adj of the adjusted hue threshold interval and the lower limit S adj of the saturation threshold interval can be corrected according to the illumination difference:

[0141]

[0142] where H min is the lower limit of the original hue threshold interval [H min , H max , and S min is the lower limit of the original saturation threshold interval [S min , S maxThe lower limit of the interval, and β is the light adjustment factor, which enables the model to still work stably when the light changes greatly. Through this adaptive threshold adjustment mechanism, the accuracy and stability of color screening are ensured under different light conditions, thereby improving the robustness of dynamic object detection.

[0143] In this embodiment, by detecting the light condition, measuring the surface reflectivity of the object, calculating the reference light intensity, estimating the actual light intensity, and the dynamic threshold adjustment mechanism, problems such as inaccurate color recognition of the target object, false detection, or missed detection caused by changes in the light condition are avoided, and accurate and stable detection of dynamic objects in a complex light environment is achieved, improving the adaptability and accuracy of color screening, and further enhancing the robustness and reliability of the dynamic object detection system.

[0144] Exemplarily, to help understand the implementation process of the dynamic object monitoring method obtained by combining the above-mentioned Embodiment 1, please refer to Figure 3 , Figure 3 A brief flow schematic diagram of a dynamic object monitoring method is provided. Specifically:

[0145] First, multiple frames of RGB images of the target scene are obtained, and the color space of each frame of RGB image is converted into the HSV color space to obtain the original HSV images corresponding to the target scene. Then, pixel points of each frame of the original HSV image are screened based on a preset hue (H) threshold interval and saturation (S) threshold interval to obtain the HSV images corresponding to the target scene. Based on any one frame of the HSV image, multiple downsampling operations are performed to obtain n images to be processed, and the optical flow field of each image to be processed is calculated in combination with the depth map of the target scene. Thus, after fusing the candidate optical flow fields of each image to be processed, the optical flow field corresponding to the target scene is obtained. Finally, according to the displacements of the pixel points in the optical flow field, it is monitored whether there are dynamic objects in the target scene.

[0146] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the dynamic object monitoring method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0147] This application also provides a dynamic object monitoring device. Please refer to Figure 4 , the dynamic object monitoring device includes:

[0148] An acquisition module 10, configured to acquire multiple frames of HSV images of the target scene;

[0149] A calculation module 20, configured to calculate the optical flow field of any frame of HSV image based on the image difference between the HSV image and the adjacent frame of HSV image;

[0150] A determination module 30, configured to determine a dynamic object of the target scene according to displacements of pixel points in the optical flow field.

[0151] Optionally, the image difference includes an image gradient and a pixel point displacement, and the calculation module 20 is further configured to:

[0152] Perform multiple downsamplings on the HSV image to obtain multiple to-be-processed images at different resolutions;

[0153] For any to-be-processed image, determine a candidate optical flow field of the to-be-processed image according to the image gradient and pixel point displacement between the to-be-processed image and an adjacent-frame to-be-processed image, where the adjacent-frame to-be-processed image is obtained by downsampling an adjacent-frame HSV image and has the same resolution as the to-be-processed image;

[0154] Perform data fusion on the candidate optical flow fields obtained after traversing each to-be-processed image to obtain the optical flow field of the HSV image.

[0155] Optionally, the calculation module 20 is further configured to:

[0156] Obtain a depth map of the target scene, and perform feature extraction on the depth map to obtain depth information of the target scene;

[0157] Perform error compensation on the hue information of the to-be-processed image based on the depth information to obtain adjusted hue information;

[0158] Calculate a candidate optical flow field of the to-be-processed image according to the image gradient and pixel point displacement between the to-be-processed image and an adjacent-frame to-be-processed image, and the saturation information and adjusted hue information of the to-be-processed image.

[0159] Optionally, the acquisition module 10 is further configured to:

[0160] Obtain multiple frames of RGB images of the target scene, and convert the color space of each frame of RGB image into the HSV color space to obtain each frame of original HSV image corresponding to the target scene;

[0161] Filter pixel points in each frame of original HSV image according to a preset hue threshold interval and saturation threshold interval to obtain multiple frames of HSV images of the target scene, where the hue value of each pixel point in any one frame of HSV image is within the hue threshold interval, and the saturation value of each pixel point is within the saturation threshold interval.

[0162] Optionally, the acquisition module 10 is configured to:

[0163] Detect the illumination direction and illumination intensity of the light source in the target scene, as well as the reflectivity of the surfaces of various objects in the target scene, and calculate the reference illumination intensity of the target scene based on the illumination direction, the illumination intensity, and the reflectivity.

[0164] Obtain the actual illumination intensity corresponding to the HSV image, and adjust the lower limits of the preset original hue threshold interval and the original saturation threshold interval based on the ratio between the reference illumination intensity and the actual illumination intensity to obtain the hue threshold interval and the saturation threshold interval.

[0165] Optionally, the determination module 30 is further configured to:

[0166] Determine the displacement amount of each pixel point according to the optical flow field;

[0167] When the displacement amount of the target pixel point group among the pixel points exceeds a preset displacement threshold, it is determined that there are dynamic objects in the target scene, where the target pixel point group is a plurality of adjacent target pixel points among the pixel points, and the displacement amount of each target pixel point exceeds the displacement threshold.

[0168] Optionally, the determination module 30 is further configured to:

[0169] Extract the target pixel point group and input the target pixel point group into a pre-trained classification model, where the pre-trained classification model is trained based on training samples, the training samples are composed of sample features and sample labels, the sample features include the hue information, saturation information, and depth information of the sample object, and the sample label is the object type of the sample object;

[0170] Identify the target pixel point group through the pre-trained classification model.

[0171] The dynamic object monitoring device provided by the present application adopts the dynamic object monitoring method in the above embodiment, and can solve the technical problem of how to monitor dynamic objects in a scene while taking into account both low cost and high performance. Compared with the prior art, the beneficial effects of the dynamic object monitoring device provided by the present application are the same as those of the dynamic object monitoring method provided by the above embodiment, and other technical features in the dynamic object monitoring device are the same as the features disclosed in the above embodiment method, and will not be described in detail here.

[0172] The present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the dynamic object monitoring method in the first embodiment above.

[0173] Refer to the following Figure 5 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0174] As Figure 5 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.

[0175] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0176] The electronic device provided by the present application adopts the dynamic object monitoring method in the above embodiment, and can solve the technical problem of how to monitor dynamic objects in a scene while taking into account both low cost and high performance. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as those of the dynamic object monitoring method provided in the above embodiment, and other technical features in the electronic device are the same as those disclosed in the method of the previous embodiment, and will not be described in detail here.

[0177] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0178] As described above, only the specific embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0179] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the dynamic object monitoring method in the above embodiment.

[0180] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0181] The above computer-readable storage medium can be included in an electronic device; it can also exist separately without being assembled into the electronic device.

[0182] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device is caused to: obtain multiple frames of HSV images of a target scene; for any frame of HSV image, calculate the optical flow field of the HSV image based on the image difference between the HSV image and an adjacent frame of HSV image; and determine the dynamic objects of the target scene according to the displacements of the pixel points in the optical flow field.

[0183] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0184] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0185] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0186] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned dynamic object monitoring method, and can solve the technical problem of how to monitor dynamic objects in a scene while taking into account both low cost and high performance. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the dynamic object monitoring method provided in the above embodiments, and will not be elaborated here.

[0187] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. A method for monitoring dynamic objects, characterized in that, The dynamic object monitoring method includes: Obtaining multiple frames of HSV images of a target scene; For any frame of HSV image, calculating the optical flow field of the HSV image based on the image difference between the HSV image and the adjacent frame of HSV image; Determining the dynamic objects in the target scene according to the displacements of the pixel points in the optical flow field.

2. The dynamic object monitoring method according to claim 1, wherein The image difference includes image gradient and pixel point displacement. The step of calculating the optical flow field of the HSV image based on the image difference between the HSV image and the adjacent frame of HSV image includes: Performing multiple downsamplings on the HSV image to obtain multiple images to be processed at different resolutions; For any image to be processed, determining the candidate optical flow field of the image to be processed according to the image gradient and pixel point displacement between the image to be processed and the adjacent frame of image to be processed, where the adjacent frame of image to be processed is obtained by downsampling the adjacent frame of HSV image and has the same resolution as the image to be processed; Performing data fusion on the candidate optical flow fields obtained after traversing each image to be processed to obtain the optical flow field of the HSV image.

3. The dynamic object monitoring method according to claim 2, wherein The step of determining the candidate optical flow field of the image to be processed according to the image gradient and pixel point displacement between the image to be processed and the adjacent frame of image to be processed includes: Obtaining the depth map of the target scene and performing feature extraction on the depth map to obtain the depth information of the target scene; Performing error compensation on the hue information of the image to be processed based on the depth information to obtain the adjusted hue information; Calculating the candidate optical flow field of the image to be processed according to the image gradient and pixel point displacement between the image to be processed and the adjacent frame of image to be processed, and the saturation information and the adjusted hue information of the image to be processed.

4. The dynamic object monitoring method according to claim 1, characterized in that, The step of obtaining multiple frames of HSV images of the target scene includes: Obtaining multiple frames of RGB images of the target scene and converting the color space of each frame of RGB image into the HSV color space to obtain the original HSV images corresponding to the target scene; Screening the pixel points in each frame of the original HSV image according to a preset hue threshold interval and saturation threshold interval to obtain multiple frames of HSV images of the target scene, where the hue values of the pixel points in any frame of HSV image are within the hue threshold interval, and the saturation values of the pixel points are within the saturation threshold interval.

5. The dynamic object monitoring method according to claim 4, wherein Before the step of screening the pixel points in each frame of the original HSV image according to a preset hue threshold interval and saturation threshold interval includes: Detecting the illumination direction and illumination intensity of the light source in the target scene, and the reflectivity of the surfaces of the objects in the target scene, and calculating the reference illumination intensity of the target scene based on the illumination direction, the illumination intensity, and the reflectivity; Obtaining the actual illumination intensity corresponding to the HSV image, and adjusting the lower limits of the preset original hue threshold interval and original saturation threshold interval based on the ratio between the reference illumination intensity and the actual illumination intensity to obtain the hue threshold interval and the saturation threshold interval.

6. The dynamic object monitoring method according to claim 1, wherein, The step of determining the dynamic object in the target scene according to the displacements of the pixel points in the optical flow field includes: Determining the displacement amounts of the pixel points according to the optical flow field; When the displacement amounts of the target pixel point group among the pixel points exceed a preset displacement threshold, determining that there is a dynamic object in the target scene, where the target pixel point group is a plurality of adjacent target pixel points among the pixel points, and the displacement amount of each target pixel point exceeds the displacement threshold.

7. The dynamic object monitoring method according to claim 6, characterized in that After the step of determining that there is a dynamic object in the target scene includes: Extracting the target pixel point group and inputting the target pixel point group into a pre-trained classification model, where the pre-trained classification model is trained based on training samples, the training samples are composed of sample features and sample labels, the sample features include the hue information, saturation information, and depth information of the sample object, and the sample label is the object type of the sample object; Identifying the target pixel point group through the pre-trained classification model.

8. A dynamic object monitoring device, characterized in that, The dynamic object monitoring device includes: An acquisition module for acquiring multiple frames of HSV images of a target scene; A calculation module for calculating the optical flow field of any frame of HSV image based on the image difference between the HSV image and the adjacent frame of HSV image; A judgment module for determining the dynamic object in the target scene according to the displacements of the pixel points in the optical flow field.

9. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the dynamic object monitoring method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the dynamic object monitoring method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Endoscope rotation angle determination method and device, endoscope system and storage medium

    CN120370537A

  • Rotor wing flow field determination method, device and equipment and storage medium

    CN120808054A

  • Rotor flow field determination method, device, equipment and storage medium

    CN120808054B