A Visual Detection Method for Abnormal Dress Code Behaviors in a Park

By using cameras to collect videos in the food processing park and combining image processing and machine learning algorithms to detect employee dress behavior, the problem of difficult to detect employee dress abnormalities in the park is solved, efficient and accurate monitoring is achieved, and the risks of cross infection and product contamination are reduced.

CN119580308BActive Publication Date: 2025-06-17QINGDAO HAIZHICHEN IND EQUIP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411746306.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-06-17
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

In food processing parks, it is difficult to detect abnormal dress behaviors of employees in efficient and accurate, resulting in possible cross-infection or product contamination, affecting corporate reputation and economic benefits.

Method used

Videos of employees in the park are collected through cameras, and video feature extraction and pedestrian sub-image recognition are used using image processing and machine learning algorithms. Combined with multi-dimensional clothing feature extraction and similarity calculation, real-time abnormal detection of employee dress behavior is achieved.

Benefits of technology

It improves the efficiency and accuracy of dress monitoring of park staff, can effectively identify and warn of dress abnormalities, and reduces the risks of cross-infection and product contamination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580308B_ABST
    Figure CN119580308B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of dressing detection in a park, and discloses a visual detection method for abnormal dressing behavior in a park. The method includes: collecting a video of pedestrians in the park and extracting features, and using a pedestrian area recognition model in the park to identify a pedestrian sub-image from video frames in the video of pedestrians in the park; extracting multi-dimensional clothing features from the pedestrian sub-image to obtain a pedestrian clothing feature vector; calculating the similarity between the pedestrian clothing feature vector and a standard clothing feature vector as the visual detection result of abnormal dressing. The present invention generates anchor boxes of different pixels and performs pedestrian pixel recognition, corrects the anchor boxes according to the distribution characteristics of pedestrian pixels in the pedestrian pixel recognition result, so that the proportion of pedestrian pixels in the anchor boxes is larger, further filters background pixels, and extracts a pedestrian clothing feature vector representing clothing color change and stripe change according to the color saturation and gray-level co-occurrence matrix of the pedestrian sub-image, realizing visual detection of abnormal dressing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dressing detection in a park, and particularly to a visual detection method for abnormal dressing behavior in a park. Background Art

[0002] In modern society, as an important carrier of economic activities, food processing parks carry a large number of production, R & D and service activities. The management of the park not only involves the reasonable allocation and efficient utilization of resources, but also includes the standardized management of employees' behaviors. Among them, employees' dressing behaviors are an important part of park management, directly affecting corporate image and health and hygiene management. Employees wearing non-compliant clothing may cause cross-infection or product contamination, affecting the reputation and economic benefits of the enterprise. Summary of the Invention

[0003] In view of this, the present invention provides a visual detection method for abnormal dressing behavior in a park, which captures the dressing situations of employees in the park in real time through devices such as cameras, and analyzes them using image processing and machine learning algorithms, effectively improving the efficiency and accuracy of dressing monitoring for park staff.

[0004] To achieve the above object, a visual detection method for abnormal dressing behavior in a park provided by the present invention includes the following steps:

[0005] S1: Collect videos of pedestrians in the park and perform feature extraction to obtain video features of pedestrians in the park;

[0006] S2: Construct a pedestrian area recognition model in the park, and use the pedestrian area recognition model in the park to receive the video features of pedestrians in the park and the videos of pedestrians in the park, and identify pedestrian sub-images from the video frames in the videos of pedestrians in the park;

[0007] S3: Extract multi-dimensional clothing features from the pedestrian sub-images to obtain a clothing feature vector of the pedestrians;

[0008] S4: Calculate the similarity between the clothing feature vector of the pedestrians and the standard clothing feature vector. If the calculated similarity exceeds a preset threshold, it indicates that there is no abnormal dressing for the pedestrians in the pedestrian sub-images, otherwise it indicates that there is abnormal dressing.

[0009] As a further improvement method of the present invention:

[0010] Optionally, in the step S1 of collecting videos of pedestrians in the park, it includes:

[0011] Collect videos of pedestrians in the park, where the videos of pedestrians in the park are composed of consecutive video frames, and the representation form of the collected videos of pedestrians in the park is:

[0012] ;

[0013] ;

[0014] ;

[0015] Wherein:

[0016] represents the pedestrian video in the park, represents the pedestrian video in the park the nth video frame in, N represents the total number of consecutive video frames in the pedestrian video in the park, ;

[0017] represents a pixel matrix of X rows and Y columns, represents the video frame the pixel at the xth row and yth column in, , , X represents the number of pixel rows of the video frame, Y represents the number of pixel columns of the video frame;

[0018] represents the pixel the color values in the R, G, and B color channels respectively;

[0019] Extract the gray value and gradient value of the pixel in the video frame, and based on the gray value and gradient value of the pixel, perform feature extraction on the pedestrian video in the park, wherein the pixel the gray value of is and the gradient value is :

[0020] ;

[0021] ;

[0022] Wherein: represents selecting the maximum value in.

[0023] Optionally, the feature extraction of the pedestrian video in the park includes:

[0024] S11: Calculate the motion information of the pixel in the video frame, wherein the motion information of the pixel is :

[0025] ;

[0026] ;

[0027] ;

[0028] ;

[0029] ;

[0030] ;

[0031] Among them:

[0032] represents the exponential function with the natural constant as the base;

[0033] represents the pixel of the neighborhood pixel gradient sequence, represents the pixel of the j-th neighborhood pixel gradient value, , where the neighborhood pixel is the pixel in the 3×3 pixel area centered on the pixel ;

[0034] represents the pixel of the neighborhood pixel gradient change sequence, represents the pixel of the j-th neighborhood pixel gradient change value;

[0035] represents the pixel of the neighborhood pixel gray weight sequence, represents the pixel of the j-th neighborhood pixel gray weight, represents the pixel of the j-th neighborhood pixel gray value;

[0036] represents the Euclidean distance from the j-th neighborhood pixel of the pixel to the pixel ;

[0037] T represents transpose;

[0038] S12: Combine the motion information of the pixel to perform multi-scale feature extraction on the video frame in multiple dimensions to obtain the multi-scale features of the pixel, where the multi-scale features of the pixel are :

[0039] ;

[0040] ;

[0041] Among them:

[0042] represents the eigenvalue of the pixel at scale c, , C represents the maximum scale;

[0043] Indicated pixel The neighborhood information of the j-th neighborhood pixel;

[0044] Represents the convolution kernel of scale c, , , Successively represent the kernel sizes of the convolution kernel in the width, height, and time dimensions;

[0045] Represents the convolution kernel in the width , height and time dimension values;

[0046] S13: Construct a multi-scale feature matrix of the video frame based on the multi-scale features of the pixels, and use the multi-scale feature matrices of consecutive video frames as the park pedestrian video features :

[0047] ;

[0048] ;

[0049] Wherein:

[0050] Represents the multi-scale feature matrix of the video frame , Represents the feature matrix of the video frame at scale c.

[0051] Optionally, in the S2 step, constructing a park pedestrian area recognition model includes:

[0052] Construct a park pedestrian area recognition model, which takes the park pedestrian video features and the video frames of the park pedestrian video as inputs and outputs the pedestrian sub-images in the video frames, where the park pedestrian area recognition model includes an input layer, an anchor box initialization layer, a pedestrian pixel recognition layer, an anchor box regression layer, and an anchor box correction layer;

[0053] The input layer is used to receive the park pedestrian video features and the video frames of the park pedestrian video;

[0054] The anchor box initialization layer is used to generate anchor boxes of a fixed size for each pixel;

[0055] The pedestrian pixel recognition layer is used to identify the pedestrian pixels that describe pedestrians in the anchor boxes;

[0056] The anchor box regression layer is used to construct a training function and perform regression calculation on the correction parameters of the anchor box by combining the distribution of pedestrian pixels in the anchor box;

[0057] The anchor box correction layer is used to correct the anchor box for each pixel, filter out the anchor boxes with abnormal sizes, and use the image contained in the anchor box in the video frame as the pedestrian sub-image in the video frame;

[0058] The park pedestrian area recognition model is used to receive the park pedestrian video features and the park pedestrian video, and identify the pedestrian sub-images from the video frames in the park pedestrian video.

[0059] Optionally, the step of using the park pedestrian area recognition model to receive the park pedestrian video features and the park pedestrian video, and identifying the pedestrian sub-images from the video frames in the park pedestrian video includes:

[0060] S21: The input layer receives the park pedestrian video features and the video frames of the park pedestrian video;

[0061] S22: The anchor box initialization layer generates anchor boxes with a fixed size for each pixel, where the coordinates of the vertices of the anchor box for the pixel are , , , , where is the pixel width of the anchor box, and

[0062] is the pixel length of the anchor box; S23: The pedestrian pixel recognition layer identifies and marks the pedestrian pixels describing pedestrians in the anchor box according to the multi-scale features of the pixels, where the recognition formula for the pixel

[0063] ;

[0064] ;

[0065] where:

[0066] represents the marking result of the pixel , represents that the pixel is a pedestrian pixel, and represents that the pixel is not a pedestrian pixel;

[0067] represents the multi-scale fusion value of the pixel , and represents a preset threshold;

[0068] Attention weights representing scale c;

[0069] S24: The anchor box regression layer constructs a training function and performs regression calculation on the correction parameters of the anchor box in combination with the distribution of pedestrian pixels in the anchor box;

[0070] where the pixel The correction parameters of the corresponding anchor box are: , represents the offset of the anchor box center coordinate in the horizontal direction, represents the offset of the anchor box center coordinate in the vertical direction, represents the scaling ratio of the anchor box width, represents the scaling ratio of the anchor box length;

[0071] S25: The anchor box correction layer corrects the anchor box for each pixel, where the anchor box vertex coordinate correction result of the pixel is:

[0072] ;

[0073] ;

[0074] ;

[0075] ;

[0076] S26: Count the total number of pixels in the corrected anchor box, mark the anchor box with pixels lower than the preset threshold as an anchor box with abnormal size, and filter it. Use non-maximum suppression to remove overlapping anchor boxes, and use the image contained in the anchor box in the video frame as the pedestrian sub-image in the video frame, where the video frame The set of pedestrian sub-images is:

[0077] ;

[0078] where:

[0079] represents the k-th pedestrian sub-image in the video frame , represents the total number of pedestrian sub-images in the video frame .

[0080] Optionally, in the S3 step, multi-dimensional clothing feature extraction is performed on the pedestrian sub-image, including:

[0081] Perform multi-dimensional clothing feature extraction on the pedestrian sub-image to obtain a pedestrian clothing feature vector, where the multi-dimensional clothing feature extraction process of the pedestrian sub-image is:

[0082] S31: Calculate the color saturation of any pixel in the pedestrian sub-image where the color saturation of the s-th pixel in the pedestrian sub-image is :

[0083] ;

[0084] where:

[0085] , , successively represent the color values of the s-th pixel in the R, G, B color channels of the pedestrian sub-image , , represents the total number of pixels of the pedestrian sub-image ;

[0086] represents selecting the minimum value among;

[0087] S32: Calculate the gray-level co-occurrence matrix of the pedestrian sub-image to obtain the co-occurrence frequencies of different gray levels in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction, where the co-occurrence frequencies of gray levels p and q in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction are successively , ;

[0088] S33: Based on the color saturation of the pixels, calculate the corner feature of the pedestrian sub-image :

[0089] ;

[0090] ;

[0091] where:

[0092] represents a preset color saturation threshold;

[0093] represents the mean color saturation of the pixel region centered on the s-th pixel;

[0094] represents a response threshold, and when exceeds the response threshold, a response is made;

[0095] S34: Based on the gray-level co-occurrence matrix, calculate the pedestrian sub-image ​Gray entropy features in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction , :

[0096] ;

[0097] ;

[0098] S35: Construct a pedestrian sub-image The pedestrian clothing feature vector of:

[0099] 。

[0100] Optionally, in the S4 step, calculating the similarity between the pedestrian clothing feature vector and the standard clothing feature vector includes:

[0101] Collect images of pedestrians wearing standard clothing, and perform multi-dimensional clothing feature extraction. Use the feature extraction result as the standard clothing feature vector , where , , respectively represent the corner point features and gray entropy features of the pedestrian image wearing standard clothing; calculate the similarity between the pedestrian clothing feature vector and the standard clothing feature vector as the similarity between the clothing of the pedestrian in the pedestrian sub-image and the standard clothing, where and The similarity calculation formula between is:

[0102] ;

[0103] Among them:

[0104] represents and The similarity between;

[0105] represents the L2 norm;

[0106] If the calculated similarity exceeds the preset threshold, it means that there is no abnormal clothing for the pedestrian in the pedestrian sub-image, otherwise it means there is abnormal clothing.

[0107] Optionally, in the S24 step, the anchor box regression layer constructs a training function to perform regression calculation on the correction parameters of the anchor box, including:

[0108] S241: Obtain the U - group scene images in different scenarios, perform manual anchor box marking on the pedestrians in the scene images, where the size of the scene images is the same as the anchor box size in step S22. Take the entire scene image as a group of anchor boxes, and the center coordinates of the anchor box are the center coordinates of the scene image, thus constituting the training dataset data:

[0109] ;

[0110] Wherein:

[0111] represents the u - th group of scene images, represents the vertex coordinates of the manually marked anchor boxes of the u - th group of scene images; in the embodiments of the present invention, the number of pedestrians in each group of scene images is 1;

[0112] S242: Perform pedestrian pixel recognition on the pixels in the scene images to obtain a pedestrian pixel distribution matrix, where the value of the pixel in the pedestrian pixel distribution matrix is ;

[0113] S243: Extract the distribution features of the pedestrian pixel distribution matrix, where the distribution features include the pixel coordinates with the most adjacent pedestrian pixels in the pedestrian pixel distribution matrix, the maximum horizontal distance and the maximum vertical distance between any two pedestrian pixels; wherein adjacent pedestrian pixels are pedestrian pixels whose distance from a pedestrian pixel is lower than a preset distance threshold.

[0114] S244: Combine the distribution features and the training dataset to construct a training function for the regression parameters :

[0115] ;

[0116] ;

[0117] Wherein:

[0118] represents the regression parameters, successively represent the regression coefficients of the horizontal offset, vertical offset, scaling ratio of the anchor box width, and scaling ratio of the anchor box length of the anchor box center coordinates;

[0119] is the distribution feature of the u - th group of scene images, is the pixel coordinates with the most adjacent pedestrian pixels, are respectively the maximum horizontal distance and the maximum vertical distance between any two pedestrian pixels;

[0120] represents using the regression parameters Perform regression calculation to obtain the calibration parameters of the initial anchor boxes for the u-th group of scenario images;

[0121] Indicates the coordinate results obtained by correcting the vertex coordinates of the anchor boxes using the calibration parameters;

[0122] Indicates the L1 norm;

[0123] S245: Use the gradient descent algorithm to solve the regression parameters in the training function, and use the regression parameters to perform regression calculation on the distribution characteristics of the pedestrian pixel distribution matrix within the anchor boxes to obtain the calibration parameters of any anchor box.

[0124] To solve the above problems, the present invention provides an electronic device, which includes:

[0125] A memory that stores at least one instruction;

[0126] A communication interface to enable communication of the electronic device; and

[0127] A processor that executes the instructions stored in the memory to implement the above-mentioned visual detection method for abnormal dressing behaviors in the park.

[0128] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned visual detection method for abnormal dressing behaviors in the park.

[0129] Compared with the prior art, the present invention proposes a visual detection method for abnormal dressing behaviors in the park, and this technology has the following advantages:

[0130] First of all, this solution proposes a video feature extraction method. By combining the gray-scale change of neighboring pixels and the gradient change of neighboring pixels of a pixel, the motion information of the pixel is calculated. The higher the motion information, the higher the probability that the pixel is a dynamic pixel. Using the motion information of the pixel as a weight, multi-scale feature extraction is performed to obtain convolutional features representing the temporal and scale changes of the pixel, capturing the spatio-temporal features of the pixel in the video, serving as the video features of park pedestrians, and improving the robustness of pedestrian pixel recognition.

[0131] Meanwhile, this solution proposes a pedestrian image recognition and abnormal clothing detection method. The method uses a park pedestrian area recognition model to generate anchor boxes with different pixels and perform pedestrian pixel recognition. According to the distribution characteristics of pedestrian pixels in the pedestrian pixel recognition results, the anchor boxes are corrected to make the proportion of pedestrian pixels in the anchor boxes larger, further filtering background pixels and highlighting the pedestrian area. Then, a pedestrian sub-image is extracted. Based on the color saturation and gray-level co-occurrence matrix of the pedestrian sub-image, corner features and gray entropy features representing clothing color changes and stripe changes are extracted. The larger the gray entropy, the more obvious the stripes; the larger the corner features, the brighter the clothing color. The similarity between the pedestrian clothing feature vector and the standard clothing feature vector is calculated, and the calculation result is the visual detection result of abnormal clothing. BRIEF DESCRIPTION OF THE DRAWINGS

[0132] Figure 1 FIG. is a schematic flowchart of a method for visually detecting abnormal clothing behavior in a park provided by an embodiment of the present invention.

[0133] The implementation, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0134] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0135] An embodiment of the present application provides a method for visually detecting abnormal clothing behavior in a park. The execution subject of the method for visually detecting abnormal clothing behavior in a park includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the method for visually detecting abnormal clothing behavior in a park can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0136] Embodiment 1:

[0137] A method for visually detecting abnormal clothing behavior in a park includes the following steps:

[0138] S1: Collect a park pedestrian video and perform feature extraction to obtain park pedestrian video features.

[0139] In the S1 step, collecting a park pedestrian video includes:

[0140] Collect a park pedestrian video, where the park pedestrian video consists of continuous video frames, and the representation form of the collected park pedestrian video is:

[0141] ;

[0142] ;

[0143] ;

[0144] Wherein:

[0145] represents the video of pedestrians in the park, represents the video of pedestrians in the park the nth video frame in it, N represents the total number of consecutive video frames in the video of pedestrians in the park, ;

[0146] represents a pixel matrix of X rows and Y columns, represents the video frame the pixel at the xth row and yth column in it, , , X represents the number of pixel rows of the video frame, and Y represents the number of pixel columns of the video frame;

[0147] represents the pixel the color values in the R, G, and B color channels respectively;

[0148] Extract the gray value and gradient value of the pixel in the video frame, and based on the gray value and gradient value of the pixel, perform feature extraction on the video of pedestrians in the park, wherein the pixel the gray value of is and the gradient value is :

[0149] ;

[0150] ;

[0151] Wherein:

[0152] represents selecting the maximum value in.

[0153] The feature extraction of the video of pedestrians in the park includes:

[0154] S11: Calculate the motion information of the pixel in the video frame, wherein the motion information of the pixel is :

[0155] ;

[0156] ;

[0157] ;

[0158] ;

[0159] ;

[0160] ;

[0161] Among them:

[0162] represents the exponential function with the natural constant as the base;

[0163] represents the neighborhood pixel gradient sequence of the pixel , represents the gradient value of the j-th neighborhood pixel of the pixel , , where the neighborhood pixels are the pixels in the 3×3 pixel area centered on the pixel ;

[0164] represents the neighborhood pixel gradient change sequence of the pixel , represents the gradient change value of the j-th neighborhood pixel of the pixel ;

[0165] represents the neighborhood pixel gray weight sequence of the pixel , represents the gray weight of the j-th neighborhood pixel of the pixel , represents the gray value of the j-th neighborhood pixel of the pixel ;

[0166] S12: Combining the motion information of the pixel, perform multi-scale feature extraction on the video frame in multiple dimensions to obtain the multi-scale features of the pixel, where the multi-scale features of the pixel are :

[0167] ;

[0168] ;

[0169] Among them:

[0170] represents the eigenvalue of the pixel at scale c, , C represents the maximum scale;

[0171] represents the pixel The neighborhood information of the j-th neighboring pixel;

[0172] Denote the convolution kernel of scale c, , , Successively denote the kernel sizes of the convolution kernel in the width, height, and time dimensions;

[0173] Denote the convolution kernel in the width , height and time dimension values;

[0174] S13: Construct a multi-scale feature matrix of the video frame based on the multi-scale features of the pixels, and use the multi-scale feature matrices of consecutive video frames as the park pedestrian video features :

[0175] ;

[0176] ;

[0177] Where:

[0178] Denote the multi-scale feature matrix of the video frame , Denote the feature matrix of the video frame at scale c.

[0179] S2: Construct a park pedestrian area recognition model, and use the park pedestrian area recognition model to receive the park pedestrian video features and the park pedestrian video, and identify the pedestrian sub-images from the video frames in the park pedestrian video.

[0180] In the S2 step of constructing the park pedestrian area recognition model, it includes:

[0181] Construct a park pedestrian area recognition model, which takes the park pedestrian video features and the video frames of the park pedestrian video as inputs and the pedestrian sub-images in the video frames as outputs. The park pedestrian area recognition model includes an input layer, an anchor box initialization layer, a pedestrian pixel recognition layer, an anchor box regression layer, and an anchor box correction layer;

[0182] The input layer is used to receive the park pedestrian video features and the video frames of the park pedestrian video;

[0183] The anchor box initialization layer is used to generate anchor boxes of a fixed size for each pixel;

[0184] The pedestrian pixel recognition layer is used to identify the pedestrian pixels that describe pedestrians in the anchor boxes;

[0185] The anchor box regression layer is used to construct a training function and perform regression calculation on the correction parameters of the anchor box by combining the distribution of pedestrian pixels in the anchor box;

[0186] The anchor box correction layer is used to correct the anchor box of each pixel, filter out the anchor boxes with abnormal sizes, and use the image contained in the anchor box in the video frame as the pedestrian sub-image in the video frame;

[0187] The park pedestrian area recognition model is used to receive the park pedestrian video features and the park pedestrian video, and identify the pedestrian sub-images from the video frames in the park pedestrian video.

[0188] The step of using the park pedestrian area recognition model to receive the park pedestrian video features and the park pedestrian video, and identify the pedestrian sub-images from the video frames in the park pedestrian video includes:

[0189] S21: The input layer receives the park pedestrian video features and the video frames of the park pedestrian video;

[0190] S22: The anchor box initialization layer generates anchor boxes with a fixed size for each pixel, where the pixel coordinates of the vertex of the anchor box are , , , , is the pixel width of the anchor box, is the pixel length of the anchor box;

[0191] S23: The pedestrian pixel recognition layer identifies and marks the pedestrian pixels describing pedestrians in the anchor box according to the multi-scale features of the pixels, where the recognition formula for the pixel is:

[0192] ;

[0193] ;

[0194] Where:

[0195] represents the marking result of the pixel , represents that the pixel is a pedestrian pixel, represents that the pixel is not a pedestrian pixel;

[0196] represents the multi-scale fusion value of the pixel , represents a preset threshold;

[0197] Indicates the attention weight of scale c;

[0198] S24: The anchor box regression layer constructs a training function, and calculates the regression of the correction parameters of the anchor box in combination with the distribution of pedestrian pixels in the anchor box;

[0199] where the pixel The correction parameters of the corresponding anchor box are: , Indicates the offset of the anchor box center coordinate in the horizontal direction, Indicates the offset of the anchor box center coordinate in the vertical direction, Indicates the scaling ratio of the anchor box width, Indicates the scaling ratio of the anchor box length;

[0200] S25: The anchor box correction layer corrects the anchor box of each pixel, where the anchor box vertex coordinate correction result of the pixel is:

[0201] ;

[0202] ;

[0203] ;

[0204] ;

[0205] S26: Count the total number of pixels in the corrected anchor box, mark the anchor box with pixels below the preset threshold as an anchor box with abnormal size, and filter it. Use non-maximum suppression to remove overlapping anchor boxes, and use the image contained in the anchor box in the video frame as the pedestrian sub-image in the video frame, where the video frame The set of pedestrian sub-images is:

[0206] ;

[0207] where:

[0208] Indicates the k-th pedestrian sub-image in the video frame , Indicates the video frame The total number of pedestrian sub-images in.

[0209] S3: Extract multi-dimensional clothing features from the pedestrian sub-image to obtain the pedestrian clothing feature vector.

[0210] The extraction of multi-dimensional clothing features from the pedestrian sub-image in the step S3 includes:

[0211] Extract multi-dimensional clothing features from the pedestrian sub-image to obtain the pedestrian clothing feature vector, where the pedestrian sub-image The multi-dimensional clothing feature extraction process is as follows:

[0212] S31: Calculate the color saturation of any pixel in the pedestrian sub-image ; where the color saturation of the s-th pixel in the pedestrian sub-image is :

[0213] ;

[0214] where:

[0215] , , represent the color values of the s-th pixel in the R, G, B color channels of the pedestrian sub-image respectively, , represents the total number of pixels of the pedestrian sub-image ;

[0216] represents selecting the minimum value of;

[0217] S32: Calculate the gray-level co-occurrence matrix of the pedestrian sub-image to obtain the co-occurrence frequencies of different gray levels in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction, where the co-occurrence frequencies of gray levels p and q in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction are , ;

[0218] S33: Based on the color saturation of the pixels, calculate the corner feature of the pedestrian sub-image :

[0219] ;

[0220] ;

[0221] where:

[0222] represents a preset color saturation threshold;

[0223] represents the average color saturation of the pixel region centered on the s-th pixel;

[0224] represents the response threshold. When Respond if it exceeds the response threshold;

[0225] S34: Calculate the pedestrian sub-image based on the gray-level co-occurrence matrix The gray entropy features in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction , :

[0226] ;

[0227] ;

[0228] S35: Construct the pedestrian clothing feature vector of the pedestrian sub-image :

[0229] 。

[0230] S4: Calculate the similarity between the pedestrian clothing feature vector and the standard clothing feature vector. If the calculated similarity exceeds the preset threshold, it means that there is no abnormal clothing for the pedestrian in the pedestrian sub-image; otherwise, it means there is abnormal clothing.

[0231] In the step S4, calculating the similarity between the pedestrian clothing feature vector and the standard clothing feature vector includes:

[0232] Collect the images of pedestrians wearing standard clothing and perform multi-dimensional clothing feature extraction. Use the feature extraction result as the standard clothing feature vector , where , , respectively represent the corner feature and the gray entropy feature of the image of the pedestrian wearing standard clothing; Calculate the similarity between the pedestrian clothing feature vector and the standard clothing feature vector as the similarity between the clothing of the pedestrian in the pedestrian sub-image and the standard clothing, where and The similarity calculation formula between them is:

[0233] ;

[0234] Among them:

[0235] represents and the similarity between them;

[0236] represents the L2 norm;

[0237] If the calculated similarity exceeds the preset threshold, it means that there is no abnormal clothing for the pedestrian in the pedestrian sub-image; otherwise, it means there is abnormal clothing.

[0238] It should be understood that the above embodiments are for illustrative purposes only and the scope of the patent application is not limited by this structure.

[0239] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments. Moreover, the term "comprising", "including" or any other variant thereof in this article is intended to cover a non-exclusive inclusion, so that a process, device, article or method comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or method. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, device, article or method comprising the element.

[0240] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0241] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A method for visually detecting abnormal dressing behavior in a park, characterized in that: The method comprises: S1: Collect the video of pedestrians in the park and extract the features to obtain the features of the video of pedestrians in the park; S2: construct a park pedestrian area recognition model, use the park pedestrian area recognition model to receive park pedestrian video features and park pedestrian videos, and recognize pedestrian sub-images from video frames in the park pedestrian videos; The park pedestrian area recognition model is used to receive park pedestrian video features and park pedestrian videos, and pedestrian sub-images are recognized from video frames in the park pedestrian videos, including: S21: The input layer receives the park pedestrian video feature f and the video frame of the park pedestrian video; S22: Anchor box initialization layer generates a fixed-size anchor box for each pixel, where pixel I n The (x,y) coordinates of the anchor box vertices are wide is the pixel width of the anchor box, and length is the pixel length of the anchor box; S23: The pedestrian pixel recognition layer identifies the pedestrian pixels describing the pedestrians in the anchor frame according to the multi-scale features of the pixels and marks them, where pixel I n The identification formula for (x,y) is: in: α n (x,y) represents pixel I n The labeling result of (x,y), α n (x, y) = 1 indicates pixel I n (x,y) is the pedestrian pixel, α n (x, y) = 0 indicates pixel I n (x,y) is not a pedestrian pixel; β n (x,y) represents pixel I n Multi-scale fusion value of (x,y), Indicates the preset threshold; represents the attention weight of scale c; Represents pixel I n The eigenvalue of (x,y) at scale c; S24: The anchor frame regression layer constructs a training function and performs regression calculation on the correction parameters of the anchor frame based on the distribution of pedestrian pixels in the anchor frame; Where pixel I n The correction parameters of the anchor box corresponding to (x, y) are: Indicates the horizontal offset of the center coordinates of the anchor box. Indicates the vertical offset of the center coordinate of the anchor box. Indicates the scaling ratio of the anchor box width, Indicates the scaling ratio of the anchor box length; S25: The anchor frame correction layer corrects the anchor frame of each pixel, where pixel I n The correction result of the vertex coordinates of the anchor frame corresponding to (x, y) is: S26: Count the total number of pixels in the corrected anchor frame, mark the anchor frame with pixels below the preset threshold as an anchor frame with abnormal size, and filter it, use non-maximum suppression to remove overlapping anchor frames, and use the image contained in the anchor frame in the video frame as the pedestrian sub-image in the video frame, where video frame I n The pedestrian sub-image set is: in: Represents video frame I n The kth pedestrian sub-image in n Represents video frame I n The total number of pedestrian sub-images in the S3: extract multi-dimensional clothing features from pedestrian sub-images to obtain pedestrian clothing feature vectors; S4: Calculate the similarity between the pedestrian clothing feature vector and the standard clothing feature vector. If the calculated similarity exceeds a preset threshold, it means that the pedestrian in the pedestrian sub-image does not have abnormal clothing. Otherwise, it means that there is abnormal clothing.

2. A method for visually detecting abnormal dressing behavior in a park as claimed in claim 1, characterized in that: The step S1 includes collecting the video of pedestrians in the park, including: Collect the video of pedestrians in the park, where the video of pedestrians in the park consists of continuous video frames. The representation of the collected video of pedestrians in the park is: I=(I1,I2,...,I n ,...,I N ) I n =(I n (x,y)) X×Y in: I represents the video of pedestrians in the park, I n represents the nth video frame in the park pedestrian video I, N represents the total number of consecutive video frames in the park pedestrian video, n∈[1,N]; (I n (x,y) X×Y Represents a pixel matrix of X rows and Y columns, I n (x,y) represents the video frame I n The pixel at the xth row and yth column in the video frame, x∈[1,X], y∈[1,Y], X represents the number of pixel rows in the video frame, and Y represents the number of pixel columns in the video frame; Represents pixel I n (x, y) are the color values ​​in the R, G, and B color channels respectively; Extract the grayscale value and gradient value of the pixel in the video frame, and extract the features of the park pedestrian video based on the grayscale value and gradient value of the pixel, where pixel I n The gray value of (x,y) is g n (x,y), gradient value is grad n (x,y).

3. A method for visually detecting abnormal dressing behavior in a park as claimed in claim 2, characterized in that: The feature extraction of the park pedestrian video includes: S11: Calculate the motion information of pixels in the video frame, where pixel I n The motion information of (x,y) is v n (x,y); S12: Combine the motion information of the pixel and extract the multi-scale features of the video frame in multiple dimensions to obtain the multi-scale features of the pixel, where pixel I n The multi-scale feature of (x,y) is f n (x,y); S13: A multi-scale feature matrix of the video frame is obtained based on the multi-scale features of the pixels, and the multi-scale feature matrix of the continuous video frames is used as the park pedestrian video feature f.

4. A method for visually detecting abnormal dressing behavior in a park as claimed in claim 1, characterized in that: The step S2 constructs a park pedestrian area recognition model, including: Constructing a park pedestrian area recognition model, wherein the park pedestrian area recognition model takes park pedestrian video features and video frames of park pedestrian videos as input, and takes pedestrian sub-images in video frames as output, wherein the park pedestrian area recognition model includes an input layer, an anchor frame initialization layer, a pedestrian pixel recognition layer, an anchor frame regression layer, and an anchor frame correction layer; The input layer is used to receive the video features of pedestrians in the park and the video frames of pedestrians in the park; The anchor box initialization layer is used to generate an anchor box of fixed size for each pixel; The pedestrian pixel recognition layer is used to identify the pedestrian pixels describing pedestrians in the anchor frame; The anchor frame regression layer is used to construct a training function and regress the correction parameters of the anchor frame based on the distribution of pedestrian pixels in the anchor frame. The anchor frame correction layer is used to correct the anchor frame of each pixel and filter out anchor frames with abnormal sizes, and use the image contained in the anchor frame in the video frame as the pedestrian sub-image in the video frame; The park pedestrian area recognition model is used to receive park pedestrian video features and park pedestrian videos, and pedestrian sub-images are recognized from video frames in the park pedestrian videos.

5. A method for visually detecting abnormal dressing behavior in a park as claimed in claim 4, characterized in that: The method of using the park pedestrian area recognition model to receive park pedestrian video features and park pedestrian video, and identifying pedestrian sub-images from video frames in the park pedestrian video, includes: S21: The input layer receives the park pedestrian video feature f and the video frame of the park pedestrian video; S22: Anchor box initialization layer generates a fixed-size anchor box for each pixel; S23: The pedestrian pixel recognition layer identifies and marks the pedestrian pixels describing the pedestrians in the anchor frame according to the multi-scale features of the pixels; S24: The anchor frame regression layer constructs a training function and performs regression calculation on the correction parameters of the anchor frame based on the distribution of pedestrian pixels in the anchor frame; S25: the anchor frame correction layer corrects the anchor frame of each pixel according to the correction parameters; S26: Count the total number of pixels in the corrected anchor frame, mark the anchor frame with pixels below the preset threshold as an anchor frame with abnormal size, and filter it, use non-maximum suppression to remove overlapping anchor frames, and use the image contained in the anchor frame in the video frame as the pedestrian sub-image in the video frame, where video frame I n The pedestrian sub-image set is: in: Represents video frame I n The kth pedestrian sub-image in n Represents video frame I n The total number of pedestrian sub-images in the image.

6. A method for visually detecting abnormal dressing behavior in a park as claimed in claim 5, characterized in that: In the step S3, multi-dimensional clothing feature extraction is performed on the pedestrian sub-image, including: The multi-dimensional clothing feature extraction is performed on the pedestrian sub-image to obtain the pedestrian clothing feature vector, where the pedestrian sub-image The multi-dimensional clothing feature extraction process is as follows: S31: Calculate pedestrian sub-image The color saturation of any pixel in the pedestrian sub-image The color saturation of the sth pixel in is S32: Calculate pedestrian sub-image The gray-level co-occurrence matrix is ​​obtained to obtain the co-occurrence frequencies of different gray values ​​in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction, where the co-occurrence frequencies of gray values ​​p and q in the 0-degree co-occurrence direction and the 45-degree co-occurrence direction are respectively S33: Calculate the pedestrian sub-image based on the color saturation of the pixel Corner feature S34: Based on the gray-level co-occurrence matrix, the pedestrian sub-image is calculated Grayscale entropy features at 0 degree co-occurrence direction and 45 degree co-occurrence direction S35: Construct pedestrian sub-image The pedestrian clothing feature vector:

7. A method for visually detecting abnormal dressing behavior in a park as claimed in claim 6, characterized in that: In the step S4, similarity calculation is performed between the pedestrian clothing feature vector and the standard clothing feature vector, including: Collect images of pedestrians wearing standard clothing, and perform multi-dimensional clothing feature extraction. The feature extraction result is used as the standard clothing feature vector F = (F(1), F(2), F(3)), where F(1), F(2), F(3) represent the corner point features and grayscale entropy features of the pedestrian image wearing standard clothing, respectively; calculate the similarity between the pedestrian clothing feature vector and the standard clothing feature vector F as the similarity between the pedestrian clothing and the standard clothing in the pedestrian sub-image, where The similarity calculation formula with F is: in: express The similarity between F; ||·||2 represents the L2 norm; If the calculated similarity exceeds a preset threshold, it means that the pedestrian in the pedestrian sub-image does not have abnormal clothing, otherwise it means that there is abnormal clothing.

8. A method for visually detecting abnormal dressing behavior in a park as claimed in claim 5, characterized in that: In step S24, the anchor frame regression layer constructs a training function to perform regression calculation on the correction parameters of the anchor frame, including: S241: Obtain U groups of scene images under different scenes, and manually mark pedestrians in the scene images with anchor frames, where the size of the scene image is consistent with the size of the anchor frame in step S22. The entire scene image is used as a group of anchor frames, and the center coordinates of the anchor frames are the center coordinates of the scene image, forming a training data set data: data={(I(u),R u )|u∈[1,U]} in: I(u) represents the u-th group of scene images, R u Represents the vertex coordinates of the manually labeled anchor boxes of the u-th group of scene images; S242: performing pedestrian pixel recognition on pixels in the scene image to obtain a pedestrian pixel distribution matrix, wherein the values ​​of the pixels in the pedestrian pixel distribution matrix are {0, 1}; S243: extracting distribution features of pedestrian pixel distribution matrix; S244: Constructing a training function of regression parameters based on the distribution characteristics and the training data set; S245: using a gradient descent algorithm to solve the regression parameters in the training function, and using the regression parameters to perform regression calculation on the distribution characteristics of the pedestrian pixel distribution matrix in the anchor frame to obtain correction parameters of any anchor frame.

Citation Information

Patent Citations

  • Multi-feature fusion overhead pedestrian detection method based on aggregated channel features and a gray level co-occurrence matrix

    CN109190456A

  • Non-standard wearing detection method and device based on deep learning and computer equipment

    CN117197580A