A method and system for occupancy monitoring based on visual images
By analyzing light conditions and media types to generate structured grating trigger commands, projecting near-infrared gratings and performing feature enhancement processing, the problem of personnel positioning deviation in traditional visual surveillance under complex environments is solved, achieving high-precision personnel spatial positioning and dynamic headcount.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional visual surveillance methods are susceptible to image quality issues in environments with drastic changes in lighting or interference from media (such as glass, water mist, or raindrops), leading to blurred outlines or false positives and false negatives, which affects the accuracy and reliability of people counting.
By capturing RGB images of the monitored scene, analyzing the lighting conditions and media types, generating structured grating trigger commands, projecting near-infrared coded gratings, acquiring grating deformation images, combining media optical parameters for feature enhancement processing, outputting clear personnel outline images, establishing a mapping relationship between two-dimensional image coordinates and three-dimensional spatial absolute positions, and statistically analyzing personnel position coordinates.
It achieves high-precision spatial distribution reconstruction of personnel in complex monitoring scenarios, improves the applicability and stability of personnel identification, and provides reliable data support for dynamic personnel statistics and regional behavior analysis.
Smart Images

Figure CN120808272B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, and particularly relates to a method and system for personnel monitoring based on visual images. BACKGROUND
[0002] With the wide application of video monitoring in public security, intelligent building and traffic management, personnel recognition and quantity statistics based on visual images gradually become a research hotspot. Traditional personnel monitoring methods mainly rely on visible light cameras to cooperate with target detection algorithms to count the number of people, and the core is to realize the positioning and counting of personnel in the scene through image feature extraction and classification recognition. In recent years, the introduction of deep learning models has significantly improved the accuracy of human body detection in complex scenes, making crowd density estimation and individual segmentation technology based on RGB images mature. However, in the environment with large changes in light or medium interference (such as glass, water mist, raindrops), the image quality is easily affected, resulting in frequent contour blurring or false detection and missing detection, thereby affecting the reliability of the final statistical results.
[0003] The current mainstream usually adopts multi-frame fusion, adaptive enhancement or light compensation to improve the contrast of images when dealing with high reflection or scattering areas, but these methods are difficult to fundamentally eliminate the brightness unevenness and edge distortion caused by the difference in optical properties of the medium surface. Especially in the case of significant deviation between environmental brightness and medium reflection characteristics, conventional image processing methods cannot effectively restore the true human body contour information, thereby affecting the accuracy of subsequent spatial coordinate mapping and position judgment. SUMMARY
[0004] In view of the above existing problems, the present application is proposed.
[0005] Therefore, the present application provides a personnel monitoring method based on visual images to fundamentally eliminate the brightness unevenness and edge distortion caused by the difference in optical properties of the medium surface.
[0006] To solve the above technical problems, the present application provides the following technical solutions:
[0007] In a first aspect, the present application provides a personnel monitoring method based on visual images, which comprises,
[0008] capturing an original RGB image of a monitoring scene, analyzing the light conditions and medium types of the original RGB image to obtain an environmental brightness value and a medium reflection characteristic; comparing the environmental brightness value and the medium reflection characteristic to generate a light difference, and generating a structured grating trigger instruction when the light difference exceeds a preset difference index;
[0009] Start the near-infrared coded grating projection action according to the structured grating trigger instruction, synchronously collect the grating deformation image of the target area, call the medium optical parameter set to perform feature enhancement processing on the grating deformation image, and output a clear boundary personnel contour image;
[0010] Set a measurement reference point in the monitoring scene, obtain the three-dimensional space absolute position of the measurement reference point, establish a mapping relationship between the image coordinates of the measurement reference point and the three-dimensional space absolute position, extract the human body shape features from the clear boundary personnel contour image, determine the personnel position coordinates in combination with the mapping relationship and the human body shape features, and output a position coordinate list;
[0011] Count the number of position coordinates in the position coordinate list, compare the number of position coordinates with the preset personnel accommodation reference, and start the alarm when the number of position coordinates continuously exceeds the personnel accommodation reference.
[0012] As a preferred scheme of the personnel monitoring method based on visual images, wherein: the ambient brightness value is the brightness value of all pixel points in the original RGB image after extracting the pixel-level brightness distribution state from the color-brightness feature set, and the average processing is used to generate an index reflecting the overall illumination intensity of the current monitoring scene;
[0013] The color-brightness feature set is a data set generated by combining the color features and brightness features of each pixel point in the original RGB image.
[0014] As a preferred scheme of the personnel monitoring method based on visual images, wherein: the medium reflection characteristic refers to generating a medium reflection characteristic describing the optical behavior of the glass and rain and fog medium surface in the monitoring scene by analyzing the color, texture and edge features of the original RGB image.
[0015] As a preferred scheme of the personnel monitoring method based on visual images, wherein: the near-infrared coded grating projection action is started according to the structured grating trigger instruction, and the grating deformation image of the target area is synchronously collected, and the specific steps are as follows,
[0016] The structured grating trigger instruction is used as the starting signal of the near-infrared coded grating projection;
[0017] Read the projection parameters in the structured grating trigger instruction to determine the coding mode;
[0018] Configure the emission assembly of the near-infrared light source, and adjust the projection direction and pattern structure of the near-infrared light source according to the coding mode;
[0019] After adjusting the projection direction and pattern structure, activate the near-infrared light source, emit the near-infrared grating pattern to the target area, and generate a structured light field covering the target area;
[0020] The grating deformation image of the target region is collected by the near-infrared imaging component while the near-infrared grating pattern is projected.
[0021] As a preferred scheme of the personnel monitoring method based on visual images, the method comprises the following steps:
[0022] The glass refractive index and the rain and fog scattering coefficient are extracted from the medium optical parameter set.
[0023] The glass refractive index and the rain and fog scattering coefficient are extracted from the medium optical parameter set.
[0024] The brightness value of each pixel point in the grating deformation image is adjusted according to the glass refractive index and the rain and fog scattering coefficient.
[0025] The Sobel operator edge detection is applied to identify the sharp boundary of the glass reflection area and the fuzzy boundary of the rain and fog scattering area, and a candidate boundary pixel set is generated.
[0026] The brightness value of each pixel point in the grating deformation image is adjusted according to the glass refractive index and the rain and fog scattering coefficient.
[0027] The morphological filtering is performed on the grating deformation image after the edge sharpening processing to remove noise interference, and an enhanced grating deformation image is generated.
[0028] The continuous and closed target boundary contour data are obtained from the enhanced grating deformation image.
[0029] The boundary contour that meets the human body size range is filtered from the target boundary contour data as the boundary clear personnel contour image through the human body size reference standard.
[0030] As a preferred scheme of the personnel monitoring method based on visual images, the mapping relationship between the image coordinates of the measurement reference point and the absolute position in the three-dimensional space is generated by associating the pixel coordinates of the measurement reference point and the absolute position in the three-dimensional space.
[0031] As a preferred scheme of the personnel monitoring method based on visual images, the mapping relationship between the image coordinates of the measurement reference point and the absolute position in the three-dimensional space is generated by associating the pixel coordinates of the measurement reference point and the absolute position in the three-dimensional space.
[0032] The structure points of the human body contour are detected from the boundary clear personnel contour image as the human body shape features.
[0033] Determine the pixel horizontal coordinate and the pixel vertical coordinate of the human center point based on the pixel position of the structural point in the clear boundary personnel contour image, and form the image coordinate;
[0034] Call the mapping relationship table between the two-dimensional image coordinate and the three-dimensional space absolute position, and convert the image coordinate into the longitude, latitude and height of the three-dimensional space absolute position;
[0035] Record the converted three-dimensional space absolute position as the personnel position coordinate;
[0036] Integrate all three-dimensional space absolute positions of the personnel to generate a position coordinate list containing all personnel position coordinates.
[0037] In the second aspect, the present application provides a personnel monitoring system based on visual images, comprising,
[0038] The instruction module is used for capturing the original RGB image of the monitoring scene, analyzing the light conditions and medium types of the original RGB image to obtain the environmental brightness value and the medium reflection characteristics; comparing the environmental brightness value with the medium reflection characteristics to generate light differences, and generating a structured grating trigger instruction when the light differences exceed the preset difference index;
[0039] The processing module is used for starting the near-infrared coded grating projection action according to the structured grating trigger instruction, synchronously collecting the grating deformation image of the target area, calling the medium optical parameter set to perform feature enhancement processing on the grating deformation image, and outputting the clear boundary personnel contour image;
[0040] The mapping module is used for setting a measurement reference point in the monitoring scene, obtaining the three-dimensional space absolute position of the measurement reference point, establishing the mapping relationship between the image coordinate and the three-dimensional space absolute position of the measurement reference point, extracting the human body shape features from the clear boundary personnel contour image, combining the mapping relationship and the human body shape features to determine the personnel position coordinate, and outputting the position coordinate list;
[0041] The alarm module is used for counting the number of position coordinates in the position coordinate list, comparing the number of position coordinates with the preset personnel accommodation reference, and starting the alarm when the number of position coordinates continuously exceeds the personnel accommodation reference.
[0042] In the third aspect, the present application provides a computer device comprising a memory and a processor, and the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the personnel monitoring method based on visual images according to the first aspect of the present application is realized.
[0043] In a fourth aspect, the present application provides a computer readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements any step of the method for crowd monitoring based on visual images according to the first aspect of the present application.
[0044] The present application has the beneficial effects that: by establishing the mapping relationship between the two-dimensional image coordinates and the three-dimensional absolute position, the accurate restoration of the spatial distribution of personnel in a complex monitoring scene is realized. Not only does it effectively overcome the positioning deviation problem brought by the traditional visual monitoring which relies on fixed viewing angle or ideal environment, but also it improves the applicability and stability of the crowd recognition under non-standard installation conditions; based on this spatial mapping mechanism, high-precision human spatial positioning can be realized without additional depth perception hardware support, by using the corresponding relationship between image features and surveying data, thereby providing reliable data support for dynamic people counting and regional behavior analysis, and finally achieving the beneficial effect of improving the overall monitoring judgment ability. BRIEF DESCRIPTION OF DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0046] Fig. 1 The flowchart of the method for crowd monitoring based on visual images.
[0047] Fig. 2 The schematic diagram of the system for crowd monitoring based on visual images.
[0048] Fig. 3 The flowchart of feature enhancement processing.
[0049] Fig. 4 The flowchart of position coordinate determination. DETAILED DESCRIPTION
[0050] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.
[0051] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0052] Second, the "one embodiment" or "an embodiment" referred to herein can include a particular feature, structure, or characteristic. The various embodiments utilized in the description of the specification are not necessarily all mutually exclusive, but a single embodiment can be employed with a variety of alternatives.
[0053] Referring to Figs. 1-4 For one embodiment of the present application, the embodiment provides a visual image-based crowd monitoring method, comprising the following steps:
[0054] S1, capture the original RGB image of the monitoring scene, analyze the light conditions and medium types of the original RGB image to obtain the environmental brightness value and medium reflection characteristics; compare the environmental brightness value with the medium reflection characteristics to generate a light difference, and generate a structured grating trigger instruction when the light difference exceeds a preset difference indicator.
[0055] Further, start the visual sensor, continuously shoot the monitoring area, obtain the original RGB image containing the visible light band, and store the original RGB image to the local buffer;
[0056] Read the original RGB image from the local buffer, retain the color information of the original RGB image, and directly extract the color feature and the brightness feature; specifically: first, extract the stored original RGB image from the local buffer, ensure that the read original RGB image contains complete red, green and blue three-channel pixel information; then, extract the red, green and blue three-channel values of each pixel point of the original RGB image as the color feature, and generate the brightness value of each pixel point; combine the color feature and the brightness feature of all pixel points to form a color-brightness feature set;
[0057] Extract the pixel-level brightness distribution state from the color-brightness feature set and convert it into the environmental brightness value of the current monitoring scene; specifically: first, extract the brightness value of each pixel point in the original RGB image from the color-brightness feature set to form a brightness set containing the brightness values of all pixel points; then, arrange the brightness values in the brightness set in size order to form a pixel-level brightness distribution state, which represents the distribution of the brightness values in the original RGB image; then, based on the pixel-level brightness distribution state, integrate the brightness values of all pixel points in the original RGB image, generate a single brightness indicator through average processing, and use it as the environmental brightness value of the current monitoring scene, reflecting the overall illumination intensity of the monitoring area; finally, generate the environmental brightness value of the current monitoring scene;
[0058] The color features are obtained from the color-luminance feature set and quantitatively processed in combination with the texture features to generate a color histogram. Specifically, first, the red, green and blue three-channel values of each pixel point are extracted from the color-luminance feature set as color features; then, the texture features are obtained by analyzing the gray level change mode of the neighborhood of the pixel point from the original RGB image through the local binary pattern method; specifically, a target pixel point is selected from the original RGB image, and a neighborhood range centered on the target pixel point is determined, for example, a 3x3 pixel window including the target pixel point and the surrounding 8 adjacent pixel points; the luminance value of the target pixel point is taken as the target pixel point luminance threshold, and the size relationship between the luminance values of the 8 adjacent pixel points in the neighborhood and the target pixel point luminance threshold is compared one by one: if the luminance value of a certain adjacent pixel point is greater than or equal to the target pixel point luminance threshold, the position of the adjacent pixel point is marked as a local binary pattern binary mark value 1; if the luminance value of a certain adjacent pixel point is less than the target pixel point luminance threshold, the position of the adjacent pixel point is marked as a local binary pattern binary mark value 0; in a clockwise order, the local binary pattern binary mark values of the 8 adjacent pixel points are combined into an 8-bit binary number to generate a local binary pattern feature value; secondly, the red, green and blue three-channel values of the color features are divided into predefined color intervals, for example, each channel is divided into 256 intervals, the number of pixel points in each interval is counted to form color distribution statistical data; at the same time, the local binary pattern feature value is mapped to discrete texture categories, for example, 8 typical texture modes, the frequency of occurrence of each texture category is counted to form texture distribution statistical data; then, the color distribution statistical data and the texture distribution statistical data are integrated to generate a color histogram containing color and texture information; wherein the color histogram records the pixel point distribution of each color interval in the color-luminance feature set and the texture categories in the original RGB image; the predefined color interval is to uniformly divide the range of the red, green and blue three-channel values (for example, 0 to 255) into multiple subintervals, for example, 256 intervals, each interval represents a specific color intensity range, which is used to count the pixel point distribution.
[0059] The color histogram is matched with glass and rain / fog features in a pre-stored media feature library to output the media type of the monitored scene as glass or rain / fog. Specifically, the frequency values of each interval of color distribution statistics and texture distribution statistics are read from the color histogram. Each interval frequency value is compared with glass feature color histogram samples and rain / fog feature color histogram samples in the pre-stored media feature library. The similarity between color distribution and texture distribution is evaluated using the histogram intersection method to generate glass feature similarity scores and rain / fog feature similarity scores. The glass feature similarity scores are compared with a preset glass recognition threshold: if the glass feature similarity score is higher than the preset glass recognition threshold (based on the glass feature color histogram in the pre-stored media feature library), the score is considered higher. If the average similarity score of the image samples is set, the medium type of the monitoring scene is determined to be glass. The rain and fog feature similarity scores are compared with a preset rain and fog recognition threshold (based on the average similarity score of the rain and fog feature color histogram samples in the pre-stored medium feature library): if the rain and fog feature similarity score is higher than the preset rain and fog recognition threshold, the medium type of the monitoring scene is determined to be rain and fog; if both the glass feature similarity score and the rain and fog feature similarity score are higher than their respective preset recognition thresholds, the medium type with the higher score is selected as the medium type of the monitoring scene. The medium feature library is a custom library built based on specific experimental data and scene analysis, used to store glass feature color histogram samples and rain and fog feature color histogram samples for medium type matching.
[0060] The glass feature similarity score is expressed as follows:
[0061] ;
[0062] The similarity score for rain and fog features is expressed as follows:
[0063] ;
[0064] in, A similarity score is assigned to the glass features, with values ranging from 0 to 1. 1 indicates a perfect match between the color histogram and the glass features, while 0 indicates no similarity. The similarity score for rain and fog features ranges from 0 to 1, where 1 indicates a perfect match between the color histogram and the rain / fog features, and 0 indicates no similarity. This represents the total number of interval frequency values for color distribution statistics and texture distribution statistics. This is a range index, specifically referring to the sequence number of each range in the color distribution statistics and texture distribution statistics. The first color in the histogram Pixel frequency values in each interval For the first glass feature in the pre-stored medium feature library Frequency values in each interval, The first interval frequency value of the rain and fog feature in the pre-stored medium feature library is determined as The minimum value function is determined as
[0065] According to the medium type, a corresponding high-reflection region or scattering region in the original RGB image is selected, pixel values (i.e., red, green, and blue three channels of the original RGB image) in the high-reflection region or scattering region are extracted and composed into a sub-image block;
[0066] In the sub-image block, an average luminance value of all pixels in the region is calculated as a luminance mean value, and the expression is as follows:
[0067]
[0068] wherein, is the average luminance value, is the width (in pixels) of the image block, is the height (in pixels) of the image block, is the total number of pixels in the image block, i.e. × , is a horizontal pixel coordinate (column number) from 1 to , is a vertical pixel coordinate (row number) from 1 to , is a luminance value of a pixel position in the image;
[0069] An edge detection algorithm is applied to the sub-image block to extract an edge pixel set in the sub-image block; specifically, first, a Sobel operator is respectively convolved in a horizontal direction and a vertical direction on each pixel point in the sub-image block to obtain gradient component values of each pixel point in the two directions; then, a gradient amplitude value of each pixel point is determined according to the gradient component values in the two directions;
[0070] Next, non-maximum suppression is performed on the gradient magnitude of all the pixel points, and the pixel points with the local maximum gradient are retained, and other non-key pixel points are removed; specifically: first, according to the gradient direction of each pixel point, the pixel points are classified into one of the four fixed directions, and the four fixed directions are 0°, 45°, 90° and 135°; then, select two pixel points adjacent to the current pixel point as comparison objects in the corresponding direction; for example, when the gradient direction of the current pixel point is 0°, select the left and right two pixel points in the horizontal direction as the comparison objects; when the gradient direction of the current pixel point is 90°, select the upper and lower two pixel points in the vertical direction as the comparison objects; when the gradient direction of the current pixel point is 45° or 135°, select two diagonal adjacent pixel points in the corresponding diagonal direction as the comparison objects; then, compare the gradient magnitude of the current pixel point with the gradient magnitudes of the two adjacent pixel points selected; if the gradient magnitude of the current pixel point is greater than or equal to any one or both of the two adjacent pixel points, the gradient magnitude of the current pixel point is retained; otherwise, the gradient magnitude of the current pixel point is set to zero;
[0071] Subsequently, the remaining pixel points are classified and judged by using a double-threshold method, and a high threshold and a low threshold are set, wherein the pixel points with a gradient magnitude higher than the high threshold are marked as strong edge points, the pixel points with a gradient magnitude between the low threshold and the high threshold are marked as weak edge points, and the pixel points lower than the low threshold are excluded; finally, in the connection stage, only the weak edge points connected with the strong edge points are retained, thereby forming a continuous and complete edge structure; the final output is a set composed of all the edge pixels screened and retained, which is the edge pixel set in the sub-image block; wherein the high threshold and the low threshold are set based on the distribution range of the gradient magnitudes of the pixel points in the sub-image block and the sensitivity requirement of edge detection, for example, the high threshold is set at 60% to 80% of the range of the gradient magnitudes, and the low threshold is set at 20% to 40% of the range of the gradient magnitudes.
[0072] The total number of pixels in the edge pixel set and the average gradient magnitude are counted to generate an edge intensity index;
[0073] The brightness mean value and the edge intensity index are combined into an image feature vector;
[0074] The image feature vector is mapped using a preset medium optical response table, and description data reflecting the current medium surface optical behavior characteristics is output as the medium reflection characteristic; Specifically, first, read the medium optical response table in the pre-stored medium feature library, the medium optical response table is established based on experimental data, and contains the corresponding relationship between the brightness mean value and the edge intensity index and the medium surface optical behavior, for example, the high reflection characteristic of glass and the scattering characteristic of rain and fog; Then, the brightness mean value and the edge intensity index of the image feature vector are taken as input, and are compared with the preset feature interval in the medium optical response table one by one; The feature interval is set based on the distribution range of the brightness mean value and the edge intensity index in the experimental data and the typical optical behavior characteristics of the medium such as glass and rain and fog, for example, the brightness mean value is divided into a sub-interval of 0 to 255, and the edge intensity index is divided into a sub-interval of 0 to 100; For the brightness mean value, the brightness interval to which the brightness mean value belongs is determined, for example, the brightness mean value is in a certain sub-interval within the range of 0 to 255; For the edge intensity index, the intensity interval to which the edge intensity index belongs is determined, for example, the edge intensity index is in a certain sub-interval within the range of 0 to 100; Then, according to the mapping rule of each interval in the medium optical response table, the combination of the brightness mean value and the edge intensity index is mapped to the corresponding optical behavior description, for example, high brightness mean value combined with high edge intensity index is mapped to the high reflection characteristic of glass, and low brightness mean value combined with low edge intensity index is mapped to the scattering characteristic of rain and fog; Finally, the medium reflection characteristic description data containing the optical behavior description is output as the medium reflection characteristic.
[0075] Determine the image position information of the corresponding high reflection area or scattering area according to the medium reflection characteristic;
[0076] Extract the local brightness distribution state corresponding to the ambient brightness value in the high reflection area or scattering area indicated by the image position information;
[0077] Compare the local brightness distribution state with the ambient brightness value to identify the brightness change trend; Specifically, first, obtain the set of all pixel brightness values in the local brightness distribution state in the sub-image block, and arrange them into a brightness value sequence; Then, compare each pixel brightness value in the brightness value sequence with the ambient brightness value one by one to determine the deviation degree of each pixel brightness value relative to the ambient brightness value, for example, the pixel brightness value is higher or lower than the ambient brightness value; Then, count the number of pixels higher than the ambient brightness value and the number of pixels lower than the ambient brightness value in the brightness value sequence to generate deviation distribution statistical data; Then, according to the proportion of high brightness pixels and low brightness pixels in the deviation distribution statistical data, judge the brightness change trend, for example, when the proportion of high brightness pixels is higher than the proportion of low brightness pixels, it is determined that the brightness change trend is enhanced; Conversely, when the proportion of low brightness pixels is higher than the proportion of high brightness pixels in the deviation distribution statistical data, it is determined that the brightness change trend is weakened.
[0078] Based on the medium type, the local brightness distribution state and the ambient brightness value, it is judged whether the image area corresponding to the glass medium or the rain and fog medium exists a significant brightness deviation phenomenon; Specifically, first, the medium type is read to determine that the current image area is a glass medium or a rain and fog medium; Then, the sequence of pixel brightness values in the local brightness distribution state is obtained, and the ambient brightness value is taken as a reference standard; Then, the sequence of pixel brightness values of the local brightness distribution state is compared with the ambient brightness value, and the proportion of pixels deviating from the ambient brightness value in the sequence of brightness values is counted, for example, the proportion of pixels deviating more than ± 20% of the ambient brightness value; Then, for the glass medium, it is checked whether the proportion of high-brightness pixels is significantly higher than the ambient brightness value, for example, the proportion of high-brightness pixels is more than 60%, if it is satisfied, it is determined that the glass medium image area exists a significant brightness deviation phenomenon, which is manifested as a high-reflection characteristic; For the rain and fog medium, it is checked whether the proportion of low-brightness pixels is significantly lower than the ambient brightness value, for example, the proportion of low-brightness pixels is more than 60%, if it is satisfied, it is determined that the rain and fog medium image area exists a significant brightness deviation phenomenon, which is manifested as a scattering characteristic; Finally, the judgment result of the brightness deviation phenomenon is output, which is used to generate a light difference signal;
[0079] According to the judgment result of the brightness deviation phenomenon, the proportion of pixels deviating from the ambient brightness value is counted as the intensity of the brightness deviation phenomenon to generate a light difference;
[0080] The light difference is compared with a preset difference index; wherein the difference index is set based on experimental data and scene analysis, and the value range is 0 to 1;
[0081] When the brightness deviation degree represented by the light difference is greater than or equal to the difference index, a structured grating trigger instruction is generated.
[0082] S2, according to the structured grating trigger instruction, a near-infrared coded grating projection action is started, a grating deformation image of the target area is synchronously collected, a medium optical parameter set is called to perform feature enhancement processing on the grating deformation image, and a clear boundary personnel contour image is output.
[0083] Further, the structured grating trigger instruction is used as a starting signal of the near-infrared coded grating projection;
[0084] The structured light trigger instruction is used to control the near-infrared light source to emit the near-infrared grating pattern according to a preset coding mode, and complete the projection action. Specifically, first, the projection parameters contained in the structured light trigger instruction are read to determine the coding mode, for example, the coding mode is a parallel line pattern or a grid pattern with uniform stripe interval. Then, the emission components of the near-infrared light source are configured, the projection direction and pattern structure of the near-infrared light source are adjusted according to the coding mode, and it is ensured that the near-infrared grating pattern covers the target monitoring area. Then, the near-infrared light source is activated, and the near-infrared grating pattern with a wavelength of, for example, 750 nanometers is emitted to the target area to generate a structured light field covering the target area. Finally, the near-infrared grating pattern is projected to the target area to provide a basis for the subsequent near-infrared imaging component to collect the grating deformation image. The coding mode is set based on the scene characteristics of the target area and the grating deformation detection requirement, for example, the coding mode is a parallel line pattern or a grid pattern with uniform stripe interval.
[0085] The grating deformation image of the target area is collected by the near-infrared imaging component at the same time as the near-infrared grating pattern is projected.
[0086] The medium optical parameter set matched with the current medium type is called from the pre-stored medium feature library.
[0087] The grating deformation image is subjected to contrast enhancement and edge sharpening processing by using the medium optical parameter set to highlight the medium influence area. Specifically, first, the parameters corresponding to the current medium type are extracted from the medium optical parameter set, including the glass refractive index and the rain and mist scattering coefficient, as the reference basis for processing the grating deformation image. Then, for each pixel point of the grating deformation image, the brightness value of the pixel point is adjusted according to the glass refractive index or the rain and mist scattering coefficient in the medium optical parameter set, so that the brightness distribution of the area affected by glass reflection or rain and mist scattering in the grating deformation image is more distinct, thereby enhancing the contrast of the grating deformation image. Then, for the pixel points in the grating deformation image after brightness adjustment, an edge detection method based on the Sobel operator is applied to identify the sharp boundary of the glass reflection area or the blurred boundary of the rain and mist scattering area in the grating deformation image by analyzing the brightness change of the neighborhood of the pixel point. Specifically, first, each pixel point in the grating deformation image is selected as a target pixel point, and a 3x3 pixel window centered on the target pixel point is determined, containing the target pixel point and the surrounding eight adjacent pixel points. Then, the horizontal direction template of the Sobel operator is used to compare the brightness values of each pixel point in the 3x3 pixel window one by one, generating the brightness change feature in the horizontal direction. At the same time, the vertical direction template of the Sobel operator is used to compare the brightness values of each pixel point in the 3x3 pixel window one by one, generating the brightness change feature in the vertical direction. Then, the horizontal direction brightness change feature and the vertical direction brightness change feature are integrated to generate the brightness change intensity value of the target pixel point, representing the boundary feature of the target pixel point in the grating deformation image. After that, the brightness change intensity values of all pixel points in the grating deformation image are sorted, and the pixel points with higher brightness change intensity values are retained as candidate points for the sharp boundary of the glass reflection area or the blurred boundary of the rain and mist scattering area. Finally, a candidate boundary pixel set containing the sharp boundary of the glass reflection area and the blurred boundary of the rain and mist scattering area is generated.
[0088] For each pixel point in the candidate boundary pixel set, the brightness value of the pixel point in the grating deformation image is increased to make the boundary of the glass reflection area clearer and the boundary of the rain and mist scattering area more prominent, thereby completing the edge sharpening processing. Finally, the grating deformation image after contrast enhancement and edge sharpening processing is generated, highlighting the medium influence area such as glass or rain and mist.
[0089] The processed grating deformation image is subjected to morphological filtering to remove noise interference and retain the main structural features of the grating deformation image, to generate an enhanced grating deformation image. Specifically, first, each pixel point in the processed grating deformation image is selected as a target pixel point, and a 3x3 pixel window centered on the target pixel point is determined, containing the target pixel point and its eight adjacent pixel points. Then, a morphological dilation operation is performed, the brightness values of each pixel point in the 3x3 pixel window are compared, the pixel point with the highest brightness value in the window is retained, and the brightness value of the target pixel point is replaced with the highest brightness value, to fill the small cracks or gaps in the grating deformation image caused by noise. Then, a morphological erosion operation is performed, the brightness values of each pixel point in the 3x3 pixel window are compared again, the pixel point with the lowest brightness value in the window is retained, and the brightness value of the target pixel point is replaced with the lowest brightness value, to remove isolated bright spots or small protrusions in the grating deformation image caused by noise. Then, the morphological dilation operation is repeated once to restore the continuity and smoothness of the sharp boundaries of the glass reflection region and the blurred boundaries of the rain and mist scattering region in the grating deformation image, while maintaining the integrity of the main structural features. Finally, the enhanced grating deformation image is generated, retaining the main structural features of the grating deformation image and removing noise interference.
[0090] Continuous and closed target boundary contour data are obtained from the enhanced grating deformation image.
[0091] The target boundary contour data are filtered through a human body size reference standard to select boundary contours that meet the human body size range as a clear boundary personnel contour image. Specifically, first, the geometric features of each closed boundary contour, including the width, height and aspect ratio of the closed boundary contour, are extracted from the target boundary contour data as a size description of the closed boundary contour. Then, the preset human body size range parameters are retrieved from the human body size reference standard, including the typical width range, height range and aspect ratio range of a standing adult in the grating deformation image, based on the camera angle and distance settings in the actual scene, such as the width range, height range and aspect ratio range of a standing adult in pixels. Then, the width, height and aspect ratio of each closed boundary contour are compared one by one with the typical width range, height range and aspect ratio range in the human body size reference standard to determine whether the size description of the closed boundary contour meets the constraint conditions of the typical width range, height range and aspect ratio range. Then, the closed boundary contours that meet the constraint conditions of the typical width range, height range and aspect ratio range are retained as boundary contours that meet the human body size range, and the closed boundary contours that do not meet the constraint conditions are removed. Finally, the closed boundary contours that meet the human body size range are integrated to generate a clear boundary personnel contour image.
[0092] S3, set a measurement reference point in the monitoring scene, obtain a three-dimensional space absolute position of the measurement reference point, establish a mapping relationship between an image coordinate of the measurement reference point and the three-dimensional space absolute position, extract a human body shape feature from the boundary clear personnel contour image, determine a personnel position coordinate in combination with the mapping relationship and the human body shape feature, and output a position coordinate list.
[0093] Further, a plurality of fixed reference positions with stable physical features are selected in the monitoring area as the measurement reference points;
[0094] A high-precision positioning device (such as a Trimble R12i global navigation satellite receiver) is used to field survey each measurement reference point to obtain a three-dimensional space absolute position of the measurement reference point;
[0095] A grating deformation image containing the measurement reference points is collected by a visual sensor, a pixel position of the measurement reference point is detected in the grating deformation image, and the pixel position is recorded as an image coordinate;
[0096] The image coordinate is one-to-one corresponding to the corresponding three-dimensional space absolute position to form a coordinate pairing sample set;
[0097] A mapping relationship table between the two-dimensional image coordinate and the three-dimensional space absolute position is constructed based on the coordinate pairing sample set. Specifically, first, the image coordinate and the three-dimensional space absolute position of each pair of measurement reference points are extracted from the coordinate pairing sample set to form a list containing multiple sets of pairing data, and each set of pairing data includes a pixel horizontal coordinate, a pixel vertical coordinate of the measurement reference point, and a longitude, a latitude and an altitude of the corresponding three-dimensional space absolute position. Then, for the pairing data in the coordinate pairing sample set, a fitting method based on the least square method is applied to determine a conversion relationship between the two-dimensional image coordinate and the three-dimensional space absolute position, and a parameter set describing the corresponding relationship between the pixel horizontal coordinate and the pixel vertical coordinate and the longitude, the latitude and the altitude is generated;
[0098] The fitting of the least square method specifically comprises: selecting pixel horizontal coordinates, pixel vertical coordinates and corresponding longitude, latitude and height of each set of paired data from the paired data list of the coordinate pairing sample set to construct a sample pair containing multiple sets of input and output; assuming that the longitude, latitude and height are in linear relationship with the pixel horizontal coordinates and the pixel vertical coordinates, initializing the coefficient parameters of the linear relationship; adjusting the coefficient parameters of the initialized linear relationship through iteration to make the predicted longitude, latitude and height in the sample pair as close as possible to the actual three-dimensional absolute position, generating an optimal set of coefficient parameters as a set of coordinate conversion parameters describing the corresponding relationship between the pixel horizontal coordinates and the pixel vertical coordinates and the longitude, latitude and height; then, storing the generated set of coordinate conversion parameters as a mapping relationship table between the two-dimensional image coordinates and the three-dimensional absolute position, and the mapping relationship table records the longitude, latitude and height values corresponding to each set of pixel horizontal coordinates and pixel vertical coordinates in tabular form; thereafter, for the pixel coordinates not directly included in the coordinate pairing sample set, an interpolation method is applied to derive the three-dimensional absolute position corresponding to the non-included pixel coordinates according to the set of coordinate conversion parameters in the mapping relationship table, and the three-dimensional absolute position is supplemented into the mapping relationship table;
[0099] The interpolation method specifically comprises: extracting paired data containing pixel horizontal coordinates, pixel vertical coordinates and corresponding longitude, latitude and height from the mapping relationship table to form a set of known coordinate points; for the pixel coordinates not included in the mapping relationship table in the raster deformation image, selecting each non-included pixel horizontal coordinate and pixel vertical coordinate as a target pixel coordinate; from the set of known coordinate points, based on the pixel horizontal coordinates and the pixel vertical coordinates of the target pixel coordinate, finding the four nearest known coordinate points to the target pixel coordinate to form a four-neighborhood coordinate point set with the target pixel coordinate as the center; applying a bilinear interpolation method to determine the longitude, latitude and height of the target pixel coordinate according to the pixel horizontal coordinates, pixel vertical coordinates and corresponding longitude, latitude and height of each known coordinate point in the four-neighborhood coordinate point set; determining the relative position of the target pixel coordinate with respect to the four-neighborhood coordinate point set according to the pixel horizontal coordinates and the pixel vertical coordinates of the four known coordinate points in the four-neighborhood coordinate point set; based on the relative position, combining the longitude, latitude and height of each known coordinate point in the four-neighborhood coordinate point set according to the weight distribution principle of the bilinear interpolation to generate the longitude, latitude and height values of the target pixel coordinate; adding the pixel horizontal coordinates, pixel vertical coordinates and generated longitude, latitude and height values of the target pixel coordinate to the mapping relationship table as new paired data; repeating the above steps until all non-included pixel coordinates in the raster deformation image generate corresponding longitude, latitude and height values; finally, generating a complete mapping relationship table between the two-dimensional image coordinates and the three-dimensional absolute position.
[0100] The key structure points of the human body contour detected from the clear boundary personnel contour image are taken as the human body shape features;
[0101] determining the image coordinates of the human center point based on the pixel position of the key structure point in the clear boundary personnel contour image;
[0102] calling the mapping relationship table between the two-dimensional image coordinates and the three-dimensional space absolute position, and converting the image coordinates of the human center point into the three-dimensional space absolute position;
[0103] record the converted three-dimensional space absolute position as the personnel position coordinates, and generate a position coordinate list containing all personnel position coordinates.
[0104] S4, count the number of position coordinates in the position coordinate list, and compare the number of position coordinates with the preset personnel accommodation benchmark. When the number of position coordinates continuously exceeds the personnel accommodation benchmark, start the alarm.
[0105] Further, extract the total number of position coordinates from the position coordinate list containing all personnel position coordinates as the current statistical population;
[0106] numerically compare the current statistical population with the preset personnel accommodation benchmark; wherein the personnel accommodation benchmark is set based on the physical space capacity of the monitoring area and the safety management requirements, for example, the bus is 80 people;
[0107] When the current statistical population is greater than the personnel accommodation benchmark, start the time accumulation mechanism to record the duration of the overstaffing state;
[0108] continuously obtain a new position coordinate list and update the current statistical population during the operation of the time accumulation mechanism;
[0109] If the current statistical population is still greater than the personnel accommodation benchmark and the overstaffing state duration reaches the preset alarm triggering threshold, an alarm signal is generated and sent to the monitoring terminal to prompt the personnel over-limit state; wherein the alarm triggering threshold is set based on the safety management requirements and the personnel density tolerance of the monitoring scene, for example, for 5 seconds.
[0110] The embodiment also provides a visual image-based personnel monitoring system, comprising:
[0111] The instruction module is used for capturing the original RGB image of the monitoring scene, analyzing the light conditions and medium types of the original RGB image to obtain the environmental brightness value and medium reflection characteristics; comparing the environmental brightness value with the medium reflection characteristics to generate a light difference, and generating a structured grating trigger instruction when the light difference exceeds the preset difference index;
[0112] The processing module is configured to start a near-infrared coded grating projection action according to the structured grating trigger instruction, synchronously collect a grating deformation image of the target region, call a medium optical parameter set to perform feature enhancement processing on the grating deformation image, and output a clear boundary personnel contour image.
[0113] The mapping module is configured to set a measurement reference point in the monitoring scene, acquire a three-dimensional space absolute position of the measurement reference point, establish a mapping relationship between an image coordinate of the measurement reference point and the three-dimensional space absolute position, extract a human body shape feature from the clear boundary personnel contour image, determine a personnel position coordinate by combining the mapping relationship and the human body shape feature, and output a position coordinate list.
[0114] The alarm module is configured to count a number of position coordinates in the position coordinate list, compare the number of position coordinates with a preset personnel accommodation reference, and start an alarm when the number of position coordinates continuously exceeds the personnel accommodation reference.
[0115] The embodiment further provides a computer device suitable for the case of the personnel counting monitoring method based on a visual image, including a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the personnel counting monitoring method based on a visual image proposed in the above embodiment.
[0116] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to perform wired or wireless communication with an external terminal. The wireless communication can be realized through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device. In addition, the input device can be an external keyboard, touchpad or mouse, etc.
[0117] The embodiment also provides a storage medium on which a computer program is stored, the program being executed by a processor to implement the method for counting people based on a visual image as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.
[0118] To sum up, the application realizes accurate restoration of spatial distribution of people in a complex monitoring scene by establishing a mapping relationship between a two-dimensional image coordinate and an absolute position in a three-dimensional space. Not only does it effectively overcome the positioning deviation problem caused by the traditional visual monitoring which relies on a fixed visual angle or an ideal environment, but also it improves the applicability and stability of the person counting recognition under non-standard installation conditions. Based on the spatial mapping mechanism, it can realize high-precision spatial positioning of a human body without additional depth perception hardware support by using the corresponding relationship between image features and surveying data, thereby providing reliable data support for dynamic people counting and regional behavior analysis, and finally achieving the beneficial effect of improving the overall monitoring judgment ability.
[0119] It should be noted that the above embodiments are only used to illustrate the technical solutions of the application but not limit the application. Although the application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the application, and all of them should be covered in the scope of the claims of the application.
Claims
1. A method for monitoring the number of people based on visual images, characterized by: The application relates to a method for monitoring a target region in a monitoring scene, and belongs to the field of image processing. The method comprises the following steps: capturing a monitoring scene original RGB image, analyzing light conditions and medium types of the original RGB image to obtain an environment brightness value and a medium reflection characteristic, comparing the environment brightness value with the medium reflection characteristic to generate a light difference, and generating a structured light grid triggering instruction when the light difference exceeds a preset difference index; According to the structured light grid triggering instruction, a near-infrared coded light grid projection action is started, a light grid deformation image of a target region is synchronously collected, a medium optical parameter set is called to perform feature enhancement processing on the light grid deformation image, and a clear boundary personnel contour image is outputted; A measurement reference point is set in the monitoring scene, a three-dimensional space absolute position of the measurement reference point is obtained, a mapping relationship between an image coordinate of the measurement reference point and the three-dimensional space absolute position is established, a human body shape feature is extracted from the clear boundary personnel contour image, the personnel position coordinate is determined by combining the mapping relationship and the human body shape feature, and a position coordinate list is outputted; The number of position coordinates in the position coordinate list is counted, and the number of position coordinates is compared with a preset personnel accommodation reference value; when the number of position coordinates continuously exceeds the personnel accommodation reference value, an alarm is started; The environment brightness value is obtained by extracting a pixel-level brightness distribution state from a color-brightness feature set, integrating brightness values of all pixel points in the original RGB image, and generating an index reflecting overall illumination intensity of the current monitoring scene through average processing; The color-brightness feature set is a data set generated by combining color features and brightness features of each pixel point in the original RGB image; The medium reflection characteristic is a medium reflection characteristic describing optical behaviors of glass and rain and fog medium surfaces in the monitoring scene, which is generated by analyzing color, texture and edge features of the original RGB image; The calling of the medium optical parameter set to perform feature enhancement processing on the light grid deformation image to output the clear boundary personnel contour image is specifically as follows: A medium optical parameter set matched with a current medium type is called from a pre-stored medium feature library; Glass refractive index and rain and fog scattering coefficients are extracted from the medium optical parameter set; The brightness values of each pixel point in the light grid deformation image are adjusted according to the glass refractive index and the rain and fog scattering coefficients; Sobel operator edge detection is applied to identify sharp boundaries of the glass reflection area and fuzzy boundaries of the rain and fog scattering area, and a candidate boundary pixel set is generated; For each pixel point in the candidate boundary pixel set, the brightness values of the pixel points in the candidate boundary pixel set in the light grid deformation image are improved to complete edge sharpening processing; Morphological filtering is performed on the light grid deformation image after the edge sharpening processing to remove noise interference, and an enhanced light grid deformation image is generated; Continuous and closed target boundary contour data are obtained from the enhanced light grid deformation image; The boundary contour data meeting a human body size range are filtered from the target boundary contour data as the clear boundary personnel contour image through a human body size reference standard.
2. The method of claim 1, wherein: According to the structured light grid triggering instruction, a near-infrared coded light grid projection action is started, a light grid deformation image of a target region is synchronously collected, a medium optical parameter set is called to perform feature enhancement processing on the light grid deformation image, and a clear boundary personnel contour image is outputted; The structured light grid triggering instruction is used as a starting signal of the near-infrared coded light grid projection; Projection parameters in the structured light grid triggering instruction are read to determine a coding mode. The emission assembly of the near-infrared light source is configured to adjust the projection direction and pattern structure of the near-infrared light source according to the coding mode; After adjusting the projection direction and pattern structure, the near-infrared light source is activated to emit a near-infrared grating pattern to the target area to generate a structured light field covering the target area; At the same time of projecting the near-infrared grating pattern, the grating deformation image of the target area is collected through the near-infrared imaging component.
3. The method of claim 1, wherein: the number of people in the monitored area is counted by counting the number of people in the monitored area based on the visual image. The mapping relationship between the image coordinates of the measurement reference points and the absolute positions in the three-dimensional space is established by correlating the pixel coordinates of the measurement reference points and the absolute positions in the three-dimensional space to generate a mapping relationship table from the two-dimensional image coordinates to the absolute positions in the three-dimensional space.
4. The method of claim 1, wherein: the number of people in the monitored area is counted by counting the number of people in the monitored area based on the visual image. The human body contour features are extracted from the personnel contour image with clear boundaries, and the personnel position coordinates are determined in combination with the mapping relationship and the human body contour features, and a position coordinate list is output, which is specifically as follows, The structural points of the human body contour are detected from the personnel contour image with clear boundaries as the human body contour features; Based on the pixel positions of the structural points of the human body contour in the personnel contour image with clear boundaries, the pixel horizontal coordinate and the pixel vertical coordinate of the human body center point are determined to form an image coordinate; The mapping relationship table between the two-dimensional image coordinates and the absolute positions in the three-dimensional space is called to convert the image coordinate into the longitude, latitude and height of the absolute position in the three-dimensional space; The converted absolute position in the three-dimensional space is recorded as the personnel position coordinate; The absolute positions in the three-dimensional space of all personnel are integrated to generate a position coordinate list containing the position coordinates of all personnel.
5. A people counting system based on visual images, based on the people counting method based on visual images according to any one of claims 1 to 4, characterized in that: It comprises, The instruction module is used for capturing the original RGB image of the monitoring scene, analyzing the light conditions and medium types of the original RGB image to obtain the environmental brightness value and medium reflection characteristics, comparing the environmental brightness value with the medium reflection characteristics to generate a light difference, and generating a structured light grating trigger instruction when the light difference exceeds a preset difference index; The processing module is used for starting the near-infrared coded grating projection action according to the structured light grating trigger instruction, synchronously collecting the grating deformation image of the target area, calling the medium optical parameter set to perform feature enhancement processing on the grating deformation image, and outputting the personnel contour image with clear boundaries; The mapping module is used for setting the measurement reference points in the monitoring scene, obtaining the absolute positions in the three-dimensional space of the measurement reference points, establishing the mapping relationship between the image coordinates of the measurement reference points and the absolute positions in the three-dimensional space, extracting the human body contour features from the personnel contour image with clear boundaries, determining the personnel position coordinates in combination with the mapping relationship and the human body contour features, and outputting a position coordinate list; The alarm module is used for counting the number of position coordinates in the position coordinate list, comparing the number of position coordinates with a preset personnel accommodation reference, and starting the alarm when the number of position coordinates continuously exceeds the personnel accommodation reference.
6. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that: The processor executes the computer program to realize the steps of the staff monitoring method based on visual images according to any one of claims 1-4.
7. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to realize the steps of the staff monitoring method based on visual images according to any one of claims 1-4.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN116883286A
Methods and Systems for Suppressing Non-Document-Boundary Contours in an Image
US20160253571A1