Method for Extracting Texture Features of Multi-Scene Video Images Based on Attention Mechanism
By using the multi-scene video image texture feature extraction method based on attention mechanism in monitoring scenes with uneven lighting, multi-scale Gaussian filtered images are constructed and new texture feature vectors are generated, which solves the problem of low texture classification accuracy and achieves higher texture feature extraction accuracy.
Patent Information
- Application Number
- CN202510045127.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-13
AI Technical Summary
In monitoring scenarios with uneven lighting, the accuracy of texture classification is low, which affects the extraction accuracy of texture features in the image.
A multi-scene video image texture feature extraction method based on attention mechanism is adopted, Gaussian filtered images of multiple scales are constructed through Gaussian filters, the annular neighborhood radius and global contrast of each pixel point are determined, and a positive and negative binary mode values and amplitude binary mode values are combined to generate a new texture feature vector, and a neural network model based on attention mechanism is input for target recognition.
It reduces the impact of light on texture extraction, improves the accuracy of texture classification, and thus improves the extraction accuracy of texture features in the image.
Smart Images

Figure CN119494969B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for extracting texture features of multi-scene video images based on an attention mechanism. Background Art
[0002] Texture is used to describe the arrangement and statistical characteristics of pixel values in a local area of an image. Texture reflects characteristics such as the roughness, smoothness, and regularity of the surface in the image. Texture features usually include gray-scale distribution, spatial arrangement, and frequency components, etc.
[0003] In some scenarios, a texture feature extraction method based on an attention mechanism is often used to extract texture features from images. In this method, the initial operation is to extract texture features using the Center-Symmetric Local Binary Pattern (CLBP). Among them, obtaining a comparison map of the image is a relatively important step. When obtaining the comparison map of the image, it is to calculate the contrast between pixel points and global pixels to obtain the comparison map of the image. In a monitoring scenario with uneven light distribution, it is very easy to lose the texture of local details of the image. For example, in areas with stronger light, the gray scale of pixels is higher, while in areas with lower light, the gray scale of pixels is lower. In this way, the same texture may appear similar in different regions, or different textures in different regions may appear similar. When finally performing texture classification, it is difficult to distinguish the textures in different regions, resulting in a low accuracy rate of texture classification, and further affecting the extraction accuracy of texture features in the image. Summary of the Invention
[0004] In order to solve the technical problem that the accuracy rate of texture classification is low, which in turn affects the extraction accuracy of texture features in the image, the purpose of the present invention is to provide a method for extracting texture features of multi-scene video images based on an attention mechanism. The specific technical solutions adopted are as follows:
[0005] First aspect, an embodiment of the present invention provides a method for extracting texture features of multi-scenario video images based on an attention mechanism, including: obtaining a local image to be recognized from a monitoring video frame of a target area; determining the global contrast of each pixel point according to the gray value of each pixel point in the local image; constructing Gaussian filtered images of multiple scales of the local image through a Gaussian filter, taking each pixel point in the Gaussian filtered image as the central pixel, and determining the circular neighborhood radius of each pixel point; selecting the direction of the point with the largest gray value difference between each pixel point within the range corresponding to the circular neighborhood radius and the central pixel point within the range corresponding to the circular neighborhood radius as the dominant direction, and determining the positive and negative binary pattern value and the amplitude binary pattern value of any pixel point; determining a new texture feature vector of the local image according to the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel point in the Gaussian filtered image of each scale; inputting the new texture feature vector into a neural network model based on the attention mechanism, and identifying the target in the local image through the neural network model.
[0006] Optionally, determining the global contrast of each pixel point according to the gray value of each pixel point in the local image includes: determining the contrast neighborhood radius of each pixel point according to the gray value of each pixel point in the local image; determining the global contrast of each pixel point according to the gray value of each pixel point and the gray values of all pixel points within the range corresponding to the contrast neighborhood radius of each pixel point.
[0007] Optionally, determining the contrast neighborhood radius of each pixel point according to the gray value of each pixel point in the local image includes: determining the light intensity attenuation rate of each column in the local image according to the gray values of adjacent pixel points in the same column of the local image; determining the average value of the light intensity attenuation rates of all columns as the vertical direction attenuation rate of the local image, and determining the horizontal direction attenuation rate of the local image as a preset value; determining the predicted gray value of the pixel point in the current row according to the gray value of the pixel point in the upper row of the current row in the same column and the vertical direction attenuation rate; determining the gray value difference of each pixel point according to the gray value and the predicted gray value of each pixel point; determining the neighborhood range of each pixel point, and determining the first average value of the gray value differences of all pixel points within the neighborhood range; determining the texture change degree of each pixel point according to the gray value difference of each pixel point and the first average value of the neighborhood range of each pixel point; classifying each pixel point according to the texture change degree to obtain multiple texture classifications; determining the core degree of the texture where the pixel point is located by using the distance between any pixel point in the same texture classification and other pixel points in the same texture classification; determining the contrast neighborhood radius of the pixel point according to the preset contrast neighborhood radius and the core degree of the texture where the pixel point is located.
[0008] Optionally, determining the light intensity attenuation rate of each column in the local image according to the gray values of adjacent pixel points in the same column of the local image includes: calculating a first difference between the gray values of adjacent pixel points in the same column of the local image, and superimposing the first differences to obtain a first superimposed value; performing a normalization process on the first superimposed value to obtain the light intensity attenuation rate of each column in the local image.
[0009] Optionally, determining the texture change degree of each pixel point according to the gray difference of each pixel point and the first mean value of the neighborhood range of each pixel point includes: calculating the absolute value of a second difference between the first mean value and the gray difference of the pixel point, and calculating a first ratio of the gray difference of the pixel point to the sum of the absolute value of the second difference and a predetermined value; performing a normalization process on the first ratio to obtain the texture change degree of the pixel point.
[0010] Optionally, determining the core degree of the texture where the pixel point is located by using the distance between any one pixel point in the same texture classification and other pixel points in the same texture classification includes: superimposing the distances between any one pixel point in the same texture classification and other pixel points in the same texture classification to obtain a second superimposed value; performing a normalization process on the second superimposed value to obtain the core degree of the texture where the pixel point is located.
[0011] Optionally, determining the comparison neighborhood radius of the pixel point according to a preset comparison neighborhood radius and the core degree of the texture where the pixel point is located includes: calculating a first product of the preset comparison neighborhood radius and the reciprocal of the core degree of the texture where the pixel point is located; determining the sum value of the preset comparison neighborhood radius and the first product as the comparison neighborhood radius of the pixel point.
[0012] Optionally, determining the global contrast of each pixel point according to the gray value of each pixel point and the gray values of all pixel points within the range corresponding to the comparison neighborhood radius of each pixel point includes: determining a second mean value of the gray values of all pixel points within the range corresponding to the comparison neighborhood radius; determining the global contrast of each pixel point according to the gray value of each pixel point and the second mean value within the range corresponding to the comparison neighborhood radius of each pixel point.
[0013] Optionally, determining the global contrast of each pixel point according to the gray value of each pixel point and the second mean value within the range corresponding to the comparison neighborhood radius of each pixel point includes: when the gray value of the pixel point is greater than or equal to the second mean value, determining the global contrast of the pixel point as a first numerical value; when the gray value of the pixel point is less than the second mean value, determining the global contrast of the pixel point as a second numerical value.
[0014] Optionally, determining a new texture feature vector of the local image according to the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel in the Gaussian filtered image at each scale includes: determining a feature vector of the Gaussian filtered image at each scale according to the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel in the Gaussian filtered image at each scale; and fusing the feature vectors of the Gaussian filtered images at each scale to obtain a new texture feature vector.
[0015] The present invention has the following beneficial effects: First, obtain a local image to be target-recognized from the monitoring video frame of the target area; then determine the global contrast of each pixel according to the gray value of each pixel in the local image; then construct Gaussian filtered images at multiple scales of the local image through a Gaussian filter, and determine the circular neighborhood radius of each pixel with each pixel in the Gaussian filtered image as the central pixel; secondly, select the direction of the point with the largest gray value difference amplitude between each pixel within the range corresponding to the circular neighborhood radius and the central pixel within the range corresponding to the circular neighborhood radius as the dominant direction, and determine the positive and negative binary pattern value and amplitude binary pattern value of any pixel; and determine a new texture feature vector of the local image according to the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel in the Gaussian filtered image at each scale; finally, input the new texture feature vector into a neural network model based on the attention mechanism, and identify the target in the local image through the neural network model.
[0016] In this way, the embodiment of the present invention can determine the fine global contrast of the preliminary texture for the local image in the monitoring screen according to the illumination change, and determine a new texture feature vector of the local image according to the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of the local image, thereby weakening the influence of illumination on texture extraction. Finally, based on the new texture feature vector, a neural network model using the attention mechanism is used to identify the target in the local image. It may cause the same texture to perform similarly in different regions, or different textures in different regions to perform similarly. In this way, the embodiment of the present invention can distinguish the textures in different regions, improve the accuracy of texture classification, and further improve the extraction accuracy of texture features in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0018] Figure 1Flowchart of a method for extracting texture features of multi-scene video images based on an attention mechanism provided by an embodiment of the present invention;
[0019] Figure 2 Installation schematic diagram of a camera provided by an embodiment of the present invention;
[0020] Figure 3 Schematic diagram of the selection method when the circular neighborhood radius is 1 provided by an embodiment of the present invention;
[0021] Figure 4 Schematic diagram of the selection method when the circular neighborhood radius is 3 provided by an embodiment of the present invention;
[0022] Figure 5 Structural schematic diagram of a system for extracting texture features of multi-scene video images based on an attention mechanism provided by an embodiment of the present invention. Detailed implementation manners
[0023] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, structures, features and effects of a method for extracting texture features of multi-scene video images based on an attention mechanism proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0025] The following specifically describes the specific solution of a method for extracting texture features of multi-scene video images based on an attention mechanism provided by the present invention with reference to the accompanying drawings.
[0026] Embodiment 1:
[0027] Please refer to Figure 1 , which shows the flowchart of the method for extracting texture features of multi-scene video images based on an attention mechanism provided by an embodiment of the present invention, including:
[0028] S101, obtain a partial image to be subjected to target recognition from the monitoring video frame of the target area.
[0029] Specifically, the target area can be any area that needs to be monitored in any scenario. Cameras can be installed in the target area to capture video images. The video images include multiple monitored video frames, and each monitored video frame is a separate image, which contains a partial image that needs to be subject to target recognition. For example, a monitored video frame includes a complete image of a worker, and the partial image is the head image of the complete image of the worker. The target to be recognized in the head image is the texture of the safety helmet in the complete image of the worker. Among them, according to different actual scenarios, the monitored video frames and the partial images that need to be subject to target recognition can also be of other types, and the embodiments of the present invention do not make limitations herein.
[0030] Further, when capturing the monitored video frames of the target area, cameras can be installed to capture video images. Exemplarily, as Figure 2 shown, Figure 2 is a schematic diagram of the installation of a camera provided by an embodiment of the present invention. Among them, the width of the entrance and exit of the camera is denoted as X in the embodiments of the present invention, the distance between the camera and the monitoring entrance is D, the height of the camera from the ground is H, and the angle of the camera is not greater than 15 degrees.
[0031] Exemplarily, the relationship among the monitoring width, installation distance, and maximum height at the monitoring entrance in the embodiments of the present invention can be as shown in Table 1 below.
[0032] Table 1 Relationship table of monitoring width, installation distance, and maximum height
[0033]
[0034] For example, at the entrance of a certain community, the width of the monitoring entrance and exit is measured to be 2 meters, and the distance between the monitoring probe allowed to be installed at the community entrance and the opposite side can be greater than about 2 meters to 3 meters. Referring to Table 1 of the above embodiments of the present invention, the process of querying and determining the position is as follows: First, determine that the monitoring width at the entrance and exit is 2 meters; then, check Table 1 and know that it can be installed at a position 1.5 meters to 3 meters away from the entrance and exit; secondly, select to install a texture feature intelligent recognition monitoring probe at a position 3 meters away from the entrance and exit; then, check Table 1 and know that the installation height shall not be higher than 2.2 meters; select to install a texture feature intelligent recognition monitoring probe at a height of 2 meters; finally, adjust the angle of the monitoring probe so that the image of the person standing at the entrance and exit is at the center of the picture, and . Using the monitored video frames obtained by the camera, select the partial images that need to be recognized, and perform grayscale processing on the partial images.
[0035] S102. Determine the global contrast of each pixel point according to the grayscale value of each pixel point in the partial image.
[0036] Specifically, for the local images obtained through the above embodiments of the present invention, due to the working environment of the camera, it is usually easy to cause uneven illumination distribution in the captured images, resulting in a gradual change in the grayscale of the obtained images and affecting texture recognition. On the one hand, since the light source comes from the overhead lighting fixture of the door monitor, the light generated by the fixture is from top to bottom, causing the grayscale value in the local area of the image to become darker from top to bottom. If the selected images are of the same texture, there is an obvious monotonic change trend in grayscale. Therefore, when the monotonic change trend generated by a certain pixel point is abnormal, it may be a pixel point of a certain texture. On the other hand, the distribution of the texture is continuous, and the area where the pixel points with a higher continuous distribution may be of the same texture. Therefore, first, estimate the change degree of the pixel points according to the global image, then determine the distribution of the texture according to the change degree, and finally obtain the comparison range of each pixel point. In addition, due to the operating characteristics of the monitor, it needs to run continuously for 24 hours. When in an environment with poor light illumination, such as night monitoring, artificial light sources in a fixed direction in the monitored area are used to achieve illumination, and then images are obtained. Since the light in such images is irradiated from a single direction, the texture of such images is brighter at the top and darker at the bottom. At this time, if a globally unified comparison area is used to segment the image to obtain a comparison map, the effect of the comparison map is poor, and more texture information is lost in the darker or brighter areas.
[0037] Furthermore, when determining the global contrast of each pixel point, as an optional embodiment of the present invention, first determine the comparison neighborhood radius of each pixel point according to the grayscale value of each pixel point in the local image; then determine the global contrast of each pixel point according to the grayscale value of each pixel point and the grayscale values of all pixel points within the corresponding range of the comparison neighborhood radius corresponding to each pixel point.
[0038] Among them, when determining the comparison neighborhood radius of each pixel, as an optional embodiment of the present invention, the light intensity attenuation rate of each column in the local image is determined according to the gray values of adjacent pixel points in the same column of the local image; the average value of the light intensity attenuation rates of all columns is determined as the vertical direction attenuation rate of the local image, and the horizontal direction attenuation rate of the local image is determined as a preset value; the predicted gray value of the pixel point in the current row is determined according to the gray value of the pixel point in the upper row of the current row in the same column and the vertical direction attenuation rate; the gray difference of each pixel point is determined according to the gray value and the predicted gray value of each pixel point; the neighborhood range of each pixel point is determined, and the first average value of the gray differences of all pixel points within the neighborhood range is determined; the texture change degree of each pixel point is determined according to the gray difference of each pixel point and the first average value of the neighborhood range of each pixel point; each pixel point is classified according to the texture change degree to obtain multiple texture classifications; the core degree of the texture where the pixel point is located is determined by using the distance between any pixel point in the same texture classification and other pixel points in the same texture classification; the comparison neighborhood radius of the pixel point is determined according to the preset comparison neighborhood radius and the core degree of the texture where the pixel point is located.
[0039] Specifically, since the selected local image for monitoring is usually a small-range texture image, it can be considered that the illumination change shows a linear change in the direction from top to bottom. In the horizontal direction of the image, since there are usually multiple illumination sources at the top, the light difference in the horizontal direction is small, and the illumination influence change in the horizontal direction is much smaller than that in the vertical direction. Therefore, it is considered that the illumination attenuation rate in the horizontal direction is a preset value, and the preset value in the embodiment of the present invention is taken as 0. Therefore, assuming that the selected local image only contains one texture, the predicted gray level size should appear, and then the texture change degree of the pixel point is determined according to the predicted gray difference and the actual gray difference. Further, the stronger the gray level of the image, the higher the degree of illumination influence. Since only the illumination change from top to bottom is considered, and the gray level of the image can reflect the illumination change, the illumination change rate is calculated by using the change of the gray level. Taking the upper left corner of the local image as the origin, the right side and the lower side are the positive directions, and the size of the local image is m*n, where m represents the mth row pixel of the local image and n represents the nth column pixel of the local image.
[0040] Further, when determining the light intensity attenuation rate of each column in the local image, as an optional embodiment of the present invention, first, the first difference between the gray values of adjacent pixel points in the same column of the local image is calculated, and the first differences are superimposed to obtain a first superimposed value; then the first superimposed value is normalized to obtain the light intensity attenuation rate of each column in the local image.
[0041] Specifically, the embodiment of the present invention specifically calculates by using the following formula:
[0042]
[0043] In the above formula, represents the light intensity attenuation rate of the k-th column of the local image. m represents the number of rows of pixels in the local image. represents the gray value of the pixel at the i-th row and k-th column in the k-th column. represents the gray value of the pixel at the (i - 1)-th row and k-th column in the k-th column. represents a normalization function, which is used to perform normalization processing.
[0044] Furthermore, by integrating all columns in the local image, the average value of the light intensity attenuation rates of all columns in the local image is obtained, and the attenuation rate in the vertical direction of the selected local image is obtained , and the attenuation rate in the horizontal direction of the local image is 0.
[0045] Furthermore, the above-mentioned light intensity attenuation rate in the vertical direction is determined according to the global gray change. If the selected local image is of the same texture, then in the same column, according to the pixel in the previous row and the change rate, the gray value of the pixel in the next row can be predicted, and the difference between the predicted value and the actual value is small. On the contrary, if the difference between the predicted value and the actual value is large, it means that the two pixel points are not of the same texture and the degree of texture change is large. In the embodiment of the present invention, the following formula is used to calculate the predicted gray value of the pixel in the current row:
[0046]
[0047] In the above formula, represents the predicted gray value of the pixel at the m-th row and k-th column in the local image according to the light intensity attenuation rate in the vertical direction of the local image. represents the gray value of the pixel at the (m - 1)-th row and k-th column. represents the light intensity attenuation rate in the vertical direction affected by light on the local image.
[0048] Furthermore, the gray difference of each pixel point can be the absolute value of the difference between the predicted gray value and the actual gray value of each pixel point. Where there is a texture change in the local image, there is a large difference between the change of the predicted gray value and the change of the actual gray value. And usually the appearance of texture is continuous. Therefore, when the difference between the predicted gray value of a pixel point and the average gray value of the surrounding neighborhood is smaller, it means that the possibility of this pixel point being another texture is higher and the degree of texture change is greater. Therefore, the embodiment of the present invention uses the formula to calculate the gray difference between the predicted gray value and the actual gray value of pixel point i. Wherein, represents the gray - scale difference between the predicted gray - scale value and the actual gray - scale value of pixel point i. represents the actual gray - scale value of pixel point i; represents the predicted gray - scale value of pixel point i.
[0049] Furthermore, in the embodiments of the present invention, the neighborhood range of each pixel point can be selecting 24 pixels around the pixel point as the neighborhood range. Of course, the size of the neighborhood range can also be other values, and the embodiments of the present invention do not limit this here. In the embodiments of the present invention, the mean value of the gray - scale differences within the neighborhood range of pixel point i is obtained respectively, denoted as .
[0050] Furthermore, when determining the texture change degree of each pixel point, as an optional embodiment of the present invention, first calculate the absolute value of the second difference between the first mean value and the gray - scale difference of the pixel point, and calculate the first ratio of the gray - scale difference of the pixel point to the sum of the absolute value of the second difference and a predetermined value; then perform a normalization process on the first ratio to obtain the texture change degree of the pixel point.
[0051] Specifically, the predetermined value can be taken according to the actual situation, and in the embodiments of the present invention, the value is 1.
[0052] In the above formula, represents the texture change degree of pixel point i. represents the gray - scale difference between the predicted gray - scale value and the actual gray - scale value of the j - th pixel point in the neighborhood range of pixel i. represents the mean value of the gray - scale differences between the predicted gray - scale values and the actual gray - scale values of the pixel points within the neighborhood range of the i - th pixel point. represents a normalization function, which is used to perform a normalization process on .
[0053] Specifically, the embodiments of the present invention calculate the texture change degree of the pixel point by the following formula:
[0054] Furthermore, the texture distribution has regional connectivity, so that pixel points with similar texture change degrees may be edge pixels of the same texture, and these edge pixels can be closed to each other. Thus, when the texture change degrees of pixel points are the same, it indicates that the current pixel points may be pixels of the same texture, and textures are usually connected, so that multiple closed regions can be formed, and one closed region is one texture. Among them, the texture is composed of multiple scattered pixel points. When the pixel points are located at the center position of the texture, the linear change affected by light is strong, and the influence degree of light is low, which can improve the contrast segmentation range and accuracy. On the contrary, when the pixel points are located at the edge position of the texture, the linear change affected by light is weak, and the influence degree of light is high, and it is necessary to reduce the contrast segmentation range to reduce the influence of light. Therefore, the embodiments of the present invention cluster similar pixel points according to the texture change degree to obtain the edge pixel features of the texture. Specifically, pixel points with the same texture change degree are classified into one category, and when the pixel points with the same texture change degree indicate that they may be formed by the same texture change.
[0055] Furthermore, calculate the position of each pixel point in the preliminary texture to obtain the size of the comparison region when calculating the comparison graph. If the pixel point i is located at the edge position of the texture, the linear influence change affected by the optical fiber is poor, and the greater the influence degree of light. When selecting a local comparison region and performing segmentation, a smaller comparison range can be selected to reduce the influence of light. On the contrary, if the pixel point is located in the middle position of the texture, the linear influence change affected by light is more obvious, and the influence degree of light is smaller, so that a larger comparison range can be selected to improve the accuracy of contrast segmentation. Therefore, when determining the core degree of the texture where the pixel point is located, as an optional embodiment of the present invention, first add up the distances between any pixel point in the same texture classification and other pixel points in the same texture classification to obtain a second superimposed value; then perform normalization processing on the second superimposed value to obtain the core degree of the texture where the pixel point is located.
[0056] Specifically, the embodiments of the present application calculate the core degree of the texture where the pixel point is located by using the following formula:
[0057]
[0058] In the above formula, represents the core degree of the texture where the pixel point i is located. The smaller the value, the higher the degree that the pixel point i is located at the center of the texture. represents the total number of pixel points in the texture classification where the pixel point i is located. represents the distance between the pixel point i and the pixel point j. represents the normalization function, which is used to perform normalization processing on .
[0059] Furthermore, the closer a pixel is to the center of the texture, the larger the range of contrast segmentation it has, which can improve the accuracy of the contrast segmentation map and retain more texture features. For pixels at the edge, the problem of being highly affected by illumination due to the intersection of texture changes is reduced. Therefore, as an alternative embodiment of the present invention, when determining the contrast neighborhood radius of a pixel, first calculate the first product of the preset contrast neighborhood radius and the reciprocal of the core degree of the texture where the pixel is located; then determine the sum of the preset contrast neighborhood radius and the first product as the contrast neighborhood radius of the pixel.
[0060] Specifically, the contrast neighborhood radius of a pixel can be calculated by the following formula in the embodiment of the present invention:
[0061]
[0062] In the above formula, represents the size of the contrast neighborhood radius of pixel i. represents the size of the preset contrast neighborhood radius set according to experience, usually set as the size of a 3-neighborhood radius window centered on the current pixel, that is is preset to 3. represents the core degree of the texture where pixel i is located. Among them, is at most not greater than half of the selected local image size. It should be noted that when and are not integers, only the integer part is retained, the part after the decimal point is removed, and the integer parts of and are used in the above formula calculation.
[0063] Furthermore, when determining the global contrast of each pixel, as an alternative embodiment of the present invention, first determine the second mean value of the grayscale values of all pixels within the range corresponding to the contrast neighborhood radius; then determine the global contrast of each pixel according to the grayscale value of each pixel and the second mean value within the range corresponding to the contrast neighborhood radius of each pixel.
[0064] Specifically, the grayscale value of each pixel can be compared with the second mean value within the range corresponding to the contrast neighborhood radius of each pixel to obtain the global contrast of each pixel. As an alternative embodiment of the present invention, when the grayscale value of a pixel is greater than or equal to the second mean value, determine the global contrast of the pixel as the first value; when the grayscale value of a pixel is less than the second mean value, determine the global contrast of the pixel as the second value.
[0065] More specifically, in the embodiment of the present invention, the first value can be taken as 1, and the second value can be taken as 0. In the embodiment of the present invention, the following formula is used to calculate the global contrast:
[0066]
[0067] In the above formula, represents the gray value of the i-th pixel point. represents the second mean value of the gray values of all pixel points within the range corresponding to the comparison neighborhood radius. represents the global contrast of the i-th pixel point.
[0068] S103. Construct Gaussian filtered images of multiple scales of the local image through a Gaussian filter, and take each pixel point in the Gaussian filtered image as the central pixel to determine the circular neighborhood radius of each pixel point.
[0069] Specifically, through the pixel point optimized by the above embodiment of the present invention to exclude the influence of light and adapt the contrast range, an optimized contrast map is determined, and other features are extracted by using the Center-Symmetric Local Binary Pattern (CLBP) algorithm. The form of the Gaussian filter can be expressed as , where (x, y) represents the coordinates of the pixel point, represents the standard deviation of the pixel points around the Gaussian filter. The local image is continuously Gaussian filtered n times to obtain Gaussian filtered images of n scales. In the embodiment of the present invention, n can be taken as 10. represents the inverse proportional normalization function, which is used to perform inverse proportional normalization on .
[0070] Furthermore, when determining the circular neighborhood radius, first obtain the central pixel in each texture classification of the local image and set multiple neighborhoods of the central pixel. Since an overly dense neighborhood will introduce a large amount of redundant information additionally and the calculation amount is large. In the embodiment of the present invention, the selected circular neighborhood radii are 1, 3, and 5 respectively, and the number of pixel points in the circular neighborhood is 8, 12, and 16. For example, as Figure 3 and Figure 4 shown, Figure 3 is a schematic diagram of a selection method when the circular neighborhood radius is 1 provided by an embodiment of the present invention, Figure 4 is a schematic diagram of a selection method when the circular neighborhood radius is 3 provided by an embodiment of the present invention. Figure 3 In, taking the pixel point 301 as the center, a circular neighborhood 302 with a radius of 1 is selected. Figure 4 In, taking the pixel point 401 as the center, a circular neighborhood 402 with a radius of 3 is selected.
[0071] S104. Select the direction of the point with the largest gray value difference amplitude between each pixel point within the range corresponding to the annular neighborhood radius and the central pixel point within the range corresponding to the annular neighborhood radius as the dominant direction, and determine the positive and negative binary pattern value and the amplitude binary pattern value of any pixel point.
[0072] Specifically, the direction of the pixel point with the largest gray value difference amplitude between each pixel point within the range corresponding to the annular neighborhood radius and the central pixel point is the dominant direction. Calculate the global contrast of each pixel point counterclockwise. And calculate from the neighborhood radius of the same pixel point from small to large, and then superimpose the global contrast of each pixel point within the range corresponding to each neighborhood radius to obtain the positive and negative binary pattern value and the amplitude binary pattern value of any pixel point. It should be noted that the calculation methods of the positive and negative binary pattern value and the amplitude binary pattern value of the pixel point can refer to the CLBP algorithm in the prior art, and will not be elaborated in the embodiments of the present invention.
[0073] S105. Determine the new texture feature vector of the local image according to the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel point in each Gaussian-filtered image of each scale.
[0074] Specifically, for the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel point in each Gaussian-filtered image of each scale, combine them into a histogram and reduce the dimension to convert them into a row vector. Fuse the feature vectors at multiple scales as a new texture feature vector.
[0075] Further, as an optional embodiment of the present invention, determining the new texture feature vector of the local image includes: determining the feature vector of each Gaussian-filtered image according to the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel point in each Gaussian-filtered image of each scale; fusing the feature vectors of the Gaussian-filtered images of each scale to obtain a new texture feature vector.
[0076] Specifically, use the global contrast, positive and negative binary pattern value, and amplitude binary pattern value of each pixel point as a coordinate system, and then count the number of pixels with the same global contrast, positive and negative binary pattern value, and amplitude binary pattern value, that is, form a dot matrix in three-dimensional space. Finally, represent the dot matrix in the form of a matrix and convert it into a row feature vector, that is, a new texture feature vector.
[0077] S106. Input the new texture feature vector into the neural network model based on the attention mechanism, and identify the target in the local image through the neural network model.
[0078] Specifically, the neural network model based on the attention mechanism can use AlexNet as the backbone network, where the attention mechanism is introduced. The attention mechanism is added to the first layer and the last layer of the AlexNet network. The input new texture feature vector is passed through multiple convolutional layers and pooling layers to obtain the feature map. The feature map is convolved using a 1×1×C convolutional filter to obtain the attention heat map, and then global max pooling is performed on the attention heat map. The maximum response value is selected on the attention heat map to obtain the region with discriminative features, forming an attention-guided deep texture feature learning model.
[0079] Improve the recognition accuracy of the texture in the local image by the AlexNet network.
[0080] In the embodiment of the present invention, the fine global contrast of the preliminary texture can be determined for the local image in the monitoring screen according to the illumination change, and the new texture feature vector of the local image can be determined according to the global contrast, the positive and negative binary pattern values and the amplitude binary pattern values of the local image, so as to weaken the influence of illumination on texture extraction. Finally, based on the new texture feature vector, the neural network model based on the attention mechanism is used for target recognition in the local image, which will cause the same texture to have similar performances in different regions, or different textures in different regions to have similar performances. In this way, the embodiment of the present invention can distinguish the textures in different regions, improve the accuracy of texture classification, and further improve the extraction accuracy of texture features in the image.
[0081] Embodiment 2:
[0082] Corresponding to the multi-scene video image texture feature extraction method based on the attention mechanism provided in the above embodiment, based on the same technical concept, the embodiment of the present invention also provides a multi-scene video image texture feature extraction system based on the attention mechanism. The multi-scene video image texture feature extraction system based on the attention mechanism is used to execute the above multi-scene video image texture feature extraction method based on the attention mechanism. Figure 5 FIG. is a schematic structural diagram of a multi-scene video image texture feature extraction system based on the attention mechanism provided by an embodiment of the present invention. As Figure 5 shown. The multi-scene video image texture feature extraction system based on the attention mechanism may vary greatly due to configuration or performance differences, and may include one or more processors 501 and a memory 502. The memory 502 is used to store computer programs that can run on the processor 501. The processor 501 is used to execute the programs stored in the memory 502 to implement the above Figure 1Each step in the method embodiments. Among them, the memory 502 can be transient storage or persistent storage. The application programs stored in the memory 502 can include one or more modules (not shown in the figure), and each module can include a series of computer-executable instructions in the multi-scenario video image texture feature extraction system based on the attention mechanism.
[0083] Furthermore, the processor 501 can be set to communicate with the memory 502 and execute a series of computer-executable instructions in the memory 502 on the multi-scenario video image texture feature extraction system based on the attention mechanism. The multi-scenario video image texture feature extraction system based on the attention mechanism can also include one or more power supplies 503, one or more wired or wireless network interfaces 504, one or more input / output interfaces 505, and one or more keyboards 506.
[0084] Specifically, in this embodiment, the multi-scenario video image texture feature extraction system based on the attention mechanism includes a processor, a communication interface, a memory, and a communication bus; among them, the processor, the communication interface, and the memory complete communication with each other through the bus; the memory is used to store computer programs; the processor is used to execute the programs stored on the memory to implement the above Figure 1 each step in the method embodiments and has the beneficial effects of the above method embodiments. To avoid repetition, the embodiments of the present invention will not be described in detail here.
[0085] It should be noted that the multi-scenario video image texture feature extraction system provided in the embodiments of the present invention and the multi-scenario video image texture feature extraction method provided in the embodiments of the present invention are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned multi-scenario video image texture feature extraction method and has the same or similar beneficial effects. The repeated parts will not be described in detail.
[0086] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0087] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.
Claims
1. A multi-scene video image texture feature extraction method based on attention mechanism, characterized in that: The multi-scene video image texture feature extraction method based on the attention mechanism includes: Acquire a local image that needs to be identified from the surveillance video of the target area; Determining the global contrast of each pixel in the local image according to the grayscale value of each pixel; Constructing Gaussian filter images of multiple scales of the local image by using a Gaussian filter, taking each pixel point in the Gaussian filter image as a central pixel, and determining the annular neighborhood radius of each pixel point; Select the direction of the point where the grayscale value difference between each pixel point within the range corresponding to the annular neighborhood radius and the central pixel point within the range corresponding to the annular neighborhood radius is the largest as the dominant direction, and determine the positive and negative binary mode values and the amplitude binary mode value of any pixel point; Determine a new texture feature vector of the local image according to the global contrast, the positive and negative binary pattern values, and the amplitude binary pattern values of each pixel in the Gaussian filtered image of each scale; Inputting the new texture feature vector into a neural network model based on an attention mechanism, and identifying the target in the local image through the neural network model; Determining the global contrast of each pixel includes: Determine the comparison neighborhood radius of each pixel point according to the gray value of each pixel point in the local image; Determine the global contrast of each pixel according to the grayscale value of each pixel and the grayscale values of all pixels within a range corresponding to the contrast neighborhood radius corresponding to each pixel; Determining the comparison neighborhood radius of each pixel point according to the grayscale value of each pixel point in the local image comprises: Determining the light intensity attenuation rate of each column in the local image according to the grayscale values of adjacent pixels in the same column of the local image; Determine that the average value of the light intensity attenuation rate of all columns is the vertical attenuation rate of the local image, and determine that the horizontal attenuation rate of the local image is a preset value; Determine a predicted grayscale value of a pixel in the current row according to the grayscale value of a pixel in a row above the current row in the same column and the vertical direction attenuation rate; Determine the grayscale difference of each pixel according to the grayscale value of each pixel and the predicted grayscale value; Determine the neighborhood range of each pixel point, and determine the first mean value of the grayscale difference of all pixels points in the neighborhood range; Determine the texture change degree of each pixel point according to the grayscale difference of each pixel point and the first mean value of the neighborhood range of each pixel point; Classifying each pixel point according to the degree of texture change to obtain multiple texture classifications; Determine the core degree of the texture where any pixel point is located by using the distance between any pixel point in the same texture classification and other pixel points in the same texture classification; The comparison neighborhood radius of the pixel point is determined according to a preset comparison neighborhood radius and the core degree of the texture where the pixel point is located.
2. The multi-scene video image texture feature extraction method based on the attention mechanism according to claim 1 is characterized in that: Determining the light intensity attenuation rate of each column in the local image according to the grayscale values of adjacent pixels in the same column of the local image includes: Calculating a first difference between grayscale values of adjacent pixels in the same column in the local image, and superimposing the first difference values to obtain a first superimposed value; The first superposition value is normalized to obtain the light intensity attenuation rate of each column in the local image.
3. The multi-scene video image texture feature extraction method based on the attention mechanism according to claim 1 is characterized in that: Determining the texture change degree of each pixel point according to the grayscale difference of each pixel point and the first mean value of the neighborhood range of each pixel point includes: Calculating an absolute value of a second difference between the first mean value and the grayscale difference of the pixel point, and calculating a first ratio of the grayscale difference of the pixel point to the sum of the absolute value of the second difference and a predetermined value; The first ratio is normalized to obtain the texture change degree of the pixel point.
4. The multi-scene video image texture feature extraction method based on attention mechanism according to claim 1 is characterized in that: Determining the core degree of the texture where any pixel point is located by using the distance between any pixel point in the same texture classification and other pixel points in the same texture classification includes: Superimposing the distances between any pixel point in the same texture classification and other pixel points in the same texture classification to obtain a second superposition value; The second superposition value is normalized to obtain the core degree of the texture where the pixel point is located.
5. The multi-scene video image texture feature extraction method based on attention mechanism according to claim 1 is characterized in that: Determining the comparison neighborhood radius of the pixel point according to the preset comparison neighborhood radius and the core degree of the texture where the pixel point is located includes: Calculate a first product of the preset comparison neighborhood radius and the inverse of the core degree of the texture where the pixel point is located; The sum of the preset comparison neighborhood radius and the first product is determined as the comparison neighborhood radius of the pixel point.
6. The method for extracting texture features from multi-scene video images based on the attention mechanism according to claim 1, characterized in that: Determining the global contrast of each pixel point according to the grayscale value of each pixel point and the grayscale values of all pixels within a range corresponding to the contrast neighborhood radius corresponding to each pixel point includes: Determine a second mean value of the grayscale values of all pixels within the range corresponding to the comparison neighborhood radius; The global contrast of each pixel is determined according to the grayscale value of each pixel and the second mean value within the range corresponding to the contrast neighborhood radius corresponding to each pixel.
7. The method for extracting texture features from multi-scene video images based on the attention mechanism according to claim 6, characterized in that: Determining the global contrast of each pixel point according to the grayscale value of each pixel point and the second mean value within the range corresponding to the contrast neighborhood radius corresponding to each pixel point comprises: When the grayscale value of the pixel is greater than or equal to the second mean value, determining the global contrast of the pixel to be a first value; When the grayscale value of the pixel point is less than the second mean value, the global contrast of the pixel point is determined to be a second value.
8. The method for extracting texture features from multi-scene video images based on the attention mechanism according to claim 1, characterized in that: Determining a new texture feature vector of the local image according to the global contrast, positive and negative binary pattern values, and amplitude binary pattern values of each pixel in the Gaussian filter image of each scale includes: Determine a feature vector of the Gaussian filtered image of each scale according to the global contrast, positive and negative binary pattern values, and amplitude binary pattern values of each pixel in the Gaussian filtered image of each scale; The feature vectors of the Gaussian filter images of each scale are fused to obtain a new texture feature vector.
Citation Information
Patent Citations
Complete local contrast binary pattern texture image processing method
CN110766023A
Texture image classification method and system based on refined local mode
CN112488123A
Unmanned aerial vehicle crop state visual identification method for air-ground cooperation
CN117409339A