Customer group attribute analysis method and device, electronic equipment and storage medium

CN120112962APending Publication Date: 2025-06-06BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380011006.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, when using human body images collected by the camera for customer attribute analysis, there are problems such as the human body being cut off, the human body being blocked, and the human body image being unclear, resulting in inaccurate analysis results.

Method used

By detecting whether the target image contains a human body, performing human body quality analysis, ensuring that the image quality meets the conditions, human body attribute recognition is carried out, and customer attributes are determined based on the human body attribute recognition results and the position relationship between each individual body.

Benefits of technology

It improves the accuracy of customer attribute analysis and avoids analysis errors caused by incomplete human body, unclear image and dense personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112962A_ABST
    Figure CN120112962A_ABST
Patent Text Reader

Abstract

A customer group attribute analysis method and apparatus, an electronic device and a storage medium. The method comprises: detecting whether a target image collected for a current region contains a human body (S110); performing human body quality analysis on a human body image region in the target image containing the human body (S120); performing human body attribute recognition on the human body image region in which the human body quality satisfies the target quality condition to obtain a human body attribute recognition result (S130); and determining a customer group attribute of the current region according to the human body attribute recognition result and a positional relationship between the human bodies in the target image (S140). Through human body quality analysis, it is ensured that the image for human body recognition has good quality, and inaccurate customer group attribute analysis caused by factors such as incomplete human bodies and unclear images is avoided. The customer group attribute is determined based on the position relation between the human bodies, customer group attribute analysis errors caused by shielding due to dense personnel are avoided, and the accuracy of customer group attribute analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Customer group attribute analysis method, device, electronic device and storage medium Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a customer attribute analysis method, device, electronic device, and storage medium. Background Art

[0002] Public places such as shopping malls and tourist attractions (i.e., parks) have huge daily traffic flows. Intelligent park operations are crucial for park security management, and counting the attributes of park visitors or analyzing the attributes of visitors in designated areas is key to intelligent park operations.

[0003] In related technologies, customer attribute analysis is performed using human body images captured by cameras. However, these images often suffer from issues such as truncation, occlusion, and unclear images of the human body, leading to inaccurate customer attribute analysis results. Therefore, a highly accurate customer attribute analysis method is urgently needed.

[0004] Summary of the Invention

[0005] In view of the above problems, embodiments of the present application provide a customer attribute analysis method, device, electronic device and storage medium to overcome the above problems or at least partially solve the above problems.

[0006] In a first aspect of an embodiment of the present application, a method for analyzing customer attributes is disclosed, the method comprising:

[0007] Detect whether the target image collected for the current area contains a human body;

[0008] performing human body quality analysis on a human body image region in a target image containing a human body;

[0009] Performing human attribute recognition on the human image region whose human body quality meets the target quality condition to obtain a human attribute recognition result;

[0010] The customer group attributes of the current area are determined according to the human attribute recognition result and the positional relationship between the human bodies in the target image.

[0011] A second aspect of the embodiments of the present application discloses a customer group attribute analysis device, the device comprising:

[0012] A detection module is used to detect whether the target image collected for the current area contains a human body;

[0013] An analysis module, configured to perform a human body quality analysis on a human body image region in a target image containing a human body;

[0014] The recognition module is used to perform human attribute recognition on the human image area whose human body quality meets the target quality condition, and obtain the human attribute recognition result;

[0015] The determination module is configured to determine the attributes of the customer group in the current area according to the human attribute recognition result and the positional relationship between the human bodies in the target image.

[0016] The third aspect of the embodiments of the present application discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the customer attribute analysis method described in the first aspect of the embodiments of the present application are implemented.

[0017] A fourth aspect of the embodiments of the present application discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the customer attribute analysis method described in the first aspect of the embodiments of the present application are implemented.

[0018] A fifth aspect of the embodiments of the present application discloses a computer program product, including a computer program, which, when executed by a processor, implements the steps of the customer attribute analysis method described in the first aspect of the embodiments of the present application.

[0019] The embodiments of the present application include the following advantages:

[0020] In an embodiment of the present application, accurate object attribute analysis is achieved through human body quality judgment and human body attribute recognition. For the target image captured in the current area, first, whether it contains a human body is detected, and human body quality analysis is performed on the human body image area in the target image that contains a human body. Then, human body attribute recognition is performed on the human body image area whose body quality meets the target quality conditions to obtain human body attribute recognition results. Finally, based on the human body attribute recognition results and the positional relationship between the various human bodies in the target image, the customer attributes of the current area are determined.

[0021] By performing a human body quality analysis on the target image containing human bodies, the image quality used for human recognition is ensured to be high, avoiding inaccurate customer attribute analysis results due to factors such as incomplete bodies and unclear images. Furthermore, by determining the customer attributes of the current area based on the positional relationships between individual bodies, errors in customer attribute analysis caused by occlusion caused by dense crowds are further avoided, thereby improving the accuracy of customer attribute analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] FIG1 is a flowchart of a method for analyzing customer attributes provided by an embodiment of the present application;

[0024] FIG2 is an example of capturing a target image of a pedestrian from entering the camera's viewing angle to leaving the camera's viewing angle, provided by an embodiment of the present application;

[0025] FIG3 is an example of a human body in multiple dimensions provided in an embodiment of the present application;

[0026] FIG4 is a flow chart of a method for adjusting the size of a human body image region provided by an embodiment of the present application;

[0027] FIG5 is a schematic structural diagram of a human body mass analysis model provided in an embodiment of the present application;

[0028] FIG6 is a schematic diagram of the structure of a spatial attention module provided in an embodiment of the present application;

[0029] FIG7 is a schematic diagram of the structure of a human attribute recognition model provided in an embodiment of the present application;

[0030] FIG8 is a flowchart of another method for analyzing customer attributes provided by an embodiment of the present application;

[0031] FIG9 is a schematic diagram of an application of a customer attribute analysis method provided in an embodiment of the present application;

[0032] FIG10 is a schematic diagram of the structure of a customer attribute analysis device provided in an embodiment of the present application;

[0033] FIG11 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Specific embodiments

[0034] To make the above-mentioned purposes, features, and advantages of this application more clearly understood, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of this application.

[0035] The present application provides a method for analyzing customer attributes, as shown in FIG1 , which is a flowchart of the steps of the method for analyzing customer attributes provided by the present application. As shown in FIG1 , the method for analyzing customer attributes provided by the present application may include steps S110 to S140:

[0036] Step S110: Detect whether the target image captured for the current area contains a human body.

[0037] In this embodiment of the application, the current area refers to the area within the park where customer attribute analysis is required. The current area is determined based on functional requirements. For example, if you need to count the attributes of customers entering the entire park every day, you need to enable the customer attribute analysis function at all entrances to the park, that is, the current area refers to the entrance area of ​​the park; if you need to count the attributes of customers in the square at all times, you need to enable the square customer attribute analysis function, that is, the current area refers to the entire square area.

[0038] For areas where customer attribute analysis is required, a camera is used to capture images to obtain a target image. When pedestrians are neither entering nor leaving the camera's field of view, the captured target image does not contain a human body. Detection of the target image that contains a human body is therefore meaningful. Therefore, after acquiring the target image, the target image must be detected to determine whether it contains a human body. Specifically, a human body detection algorithm is used to detect the target image to obtain a target image that contains a human body. This facilitates subsequent human body quality analysis based on the target image that contains a human body, thereby avoiding further processing of the target image that does not contain a human body and reducing the waste of processing resources.

[0039] Step S120: performing a human body quality analysis on the human body image region in the target image containing the human body.

[0040] In the embodiments of the present application, various scenarios may occur in the target image captured by the camera as a pedestrian enters and leaves the camera's field of view. For example, Figure 2 shows an example of target image capture of a pedestrian entering and leaving the camera's field of view, provided in the embodiments of the present application. When the pedestrian first enters the camera's field of view, the pedestrian is obscured by other obstructions. The pedestrian gradually enters the camera's optimal shooting angle, and finally, when the pedestrian leaves the camera's field of view, only half of the pedestrian appears.

[0041] In a target image sequence containing a human body, if human attribute recognition is performed on each frame of a target image containing a human body, the human attribute recognition of occluded human bodies and target images containing only half of a human body will be inaccurate, resulting in inaccurate human attribute recognition of the entire human body. Therefore, a human quality analysis is performed on the human body image area in the target image containing a human body to determine whether the target image containing a human body can be used for human attribute recognition. By performing a human quality analysis on the human body image area in the target image containing a human body, it is ensured that the human body image area used for human body recognition in the subsequent step S130 has good human body quality, avoiding the problem of inaccurate customer attribute analysis results due to factors such as incomplete human bodies and unclear images.

[0042] In an optional embodiment, a human body quality analysis is performed on a human body image region in a target image containing a human body, including: inputting the human body image region into a pre-trained human body quality analysis model to obtain a human body quality score output by the human body quality analysis model; the training samples of the human body quality analysis model are: sample human body image regions carrying human body quality labels, and the human body quality labels include: a real human body quality score pre-labeled for the sample human body image region, and the quality of the sample human body image region in multiple dimensions.

[0043] In the embodiment of the present application, a body mass score is used to represent the body mass analysis result. A larger body mass score indicates better body quality in the body image area.

[0044] In a specific embodiment, the real human body quality score of the sample human body image area is marked according to the following steps: based on whether the sample human body image area includes a complete human body, whether the human body in the sample human body image area is an upper body or a lower body, whether the sample human body image area includes a real human body, and in combination with the clarity and brightness of the sample human body image area, the human body quality score is marked for the sample human body image area.

[0045] Specifically, in a case where the sample human body image region includes a complete human body, marking the sample human body image region with a human body quality score in a first range;

[0046] In a case where the human body in the sample human body image region is an upper body, marking the sample human body image region with a human body mass score in a second range, where an upper limit value of the second range is less than or equal to a lower limit value of the first range;

[0047] In a case where the human body in the sample human body image region is a lower body body, marking the sample human body image region with a human body mass score in a third range, where an upper limit value of the third range is less than or equal to a lower limit value of the second range;

[0048] In a case where the human body in the sample human body image region is not a real human body, marking the sample human body image region with a human body mass score in a fourth range, where an upper limit value of the fourth range is less than or equal to a lower limit value of the third range;

[0049] The lower the clarity and brightness of the sample human body image region are, the closer the human body quality score marked for the sample human body image region is to the lower limit of the corresponding range.

[0050] For example, the definition of the human body quality score is as follows: the human body quality score of a complete human body is between 0.75 and 1.0 in a first range, the human body quality score of an upper body is between 0.5 and 0.75 in a second range, the human body quality score of a lower body is between 0.2 and 0.5 in a third range, and the human body quality score of a non-human or fake human body is between 0.0 and 0.2 in a fourth range. It is understood that whether the human body in the human body image area is complete, whether there is any occlusion, whether it is the back side, clarity, or brightness, one or more of these factors will affect the human body quality score. If the human body in the human body image area is a complete human body with no occlusion, the human body quality score will be above 0.75. If the human body is from the back side, the human body quality score will be below 0.85. If the human body is facing the camera, the human body quality score will be above 0.85. If the human body is from the back side and the image is clear, the human body quality score will be higher than if the human body is from the back side and the image is blurry. And so on. The human body image area with a human body quality score between 0.0 and 0.2 includes non-human images and fake human bodies (reflections, shadows, posters, etc.). Poor human body quality is caused by a combination of factors such as blurry human body image areas, human body occlusion, and human body stages. In a specific embodiment, the lower the clarity and brightness of the sample human body image area, the closer the human body quality score marked for the sample human body image area is to the lower limit of the corresponding range, that is, the principle of the human body quality score meets the principle of "moving downward". For example, if a complete human body is detected, but the image is blurry and the lighting is dim, the human body quality score will slide towards the 0.5-0.75 range.

[0051] The multiple dimensions refer to factors influencing human body quality, and the multiple dimensions include at least: a complete human body, a truncated human body, an occluded human body, and a non-human body. A truncated human body refers to an incomplete human body captured when a person just enters or leaves the camera range. An occluded human body refers to an object blocking the main human body area. A non-human body refers to a false detection or only the legs. For example, Figure 3 is an example of a multi-dimensional human body provided in an embodiment of the present application.

[0052] Since the sizes of human body image areas in different dimensions may be different, in order to keep the original proportions of the human body image area and to facilitate the processing of the human body quality analysis model, the size of the human body image area needs to be adjusted proportionally to obtain a fixed-size human body image area.

[0053] In a specific embodiment, before inputting the human body image region into a pre-trained human body mass analysis model, the method further includes:

[0054] Adjusting the size of the human body image region to obtain a human body image region whose size meets the input size of the human body mass analysis model;

[0055] The human body mass analysis model is obtained by training a first preset model, and before the sample human body image area is input into the first model, the method further includes:

[0056] The size of the sample human body image region is adjusted to obtain a sample human body image region whose size meets the input size of the first model.

[0057] Whether it is the stage of training the first model to obtain a human body mass model or the stage of using the pre-trained human body mass model to perform human body mass analysis, the human body image area needs to be adjusted. The principle of size adjustment of the human body image area is: set the size of the image input to the first model (or pre-trained human body mass model) to H×W (for example, 64×128), where H is the height of the human body image area, and W is the width of the human body image area, in pixels), and then adjust the size of the human body image area to the size of the image input to the first model (or pre-trained human body mass model) in proportion. Figure 4 is a flow chart of a method for size adjustment of a human body image area provided in an embodiment of the present application.

[0058] Specifically, when the width W of the human body image region is greater than the height H, the width W of the human body image region is adjusted to a fixed size, for example, to 128 pixels. The height H of the human body image region is then proportionally adjusted. The difference between the adjusted height H and the fixed size (for example, 64 pixels) is then calculated. Based on the difference, the upper and lower directions of the human body image region after the height H is adjusted are padded with fixed values ​​(for example, gray pixel values) to obtain a human body image region whose size conforms to the input size of the first model (or a pre-trained human body mass model). For example, the first column, second row, second column, second row, and third column, second row in FIG4 are human body image regions obtained by padded with gray pixel values ​​in the upper and lower directions of the human body image region after the height H is adjusted, and whose size conforms to the input size of the first model (or a pre-trained human body mass model).

[0059] Similarly, if the width W of the human image region is not greater than the height H, the height H of the human image region is adjusted to a fixed size, for example, to 64 pixels. The width W of the human image region is then proportionally adjusted. The difference between the adjusted width W and the fixed size (for example, 128 pixels) is then calculated. Based on the difference, the left and right directions of the human image region after the width H is adjusted are padded with fixed values ​​(for example, gray pixel values) to obtain a human image region whose size conforms to the input size of the first model (or a pre-trained human mass model). For example, the fourth column, second row in FIG4 is a human image region obtained by padded with gray pixel values ​​in both directions of the human image region after the width W is adjusted, and whose size conforms to the input size of the first model (or a pre-trained human mass model).

[0060] The human body quality analysis model performs human body quality analysis on the input fixed-size human body image area. Since the sample human body image area used to train the human body quality analysis model is marked with the true human body quality score and the quality of the sample human body image area in multiple dimensions, the trained human body quality analysis model can calculate the human body quality score in the human body image area based on the quality of each dimension in the human body image area.

[0061] Furthermore, the human body mass analysis model includes at least multiple parallel human body mass analysis branches, and the multiple human body mass analysis branches are used to: learn information affecting the human body mass score in the human body image region from the multiple dimensions at different granularities; input the human body image region into a pre-trained human body mass analysis model to obtain a human body mass score output by the human body mass analysis model, including steps A1 to A3:

[0062] Step A1: inputting the human body image region into the human body quality analysis model to obtain the initial quality prediction information of the multiple dimensions outputted by the multiple human body quality analysis branches.

[0063] Step A2: fusing the initial quality prediction information of the multiple dimensions outputted by the multiple human body quality analysis branches through the human body quality analysis model to obtain final quality prediction information of the human body image region in the multiple dimensions.

[0064] Step A3: obtaining a human body quality score output by the human body quality analysis model for the human body image region according to the respective weights of the multiple dimensions and the final quality prediction information of the human body image region in the multiple dimensions.

[0065] In an embodiment of the present application, in order to extract effective human features from a human image region, in step A1, each human mass analysis branch utilizes a spatial attention mechanism to obtain information that affects the human mass score from the input feature map, thereby obtaining initial quality prediction information in multiple dimensions. Furthermore, the multiple parallel human mass analysis branches include a deep-level human mass analysis branch and a shallow-level human mass analysis branch, with the deep-level human mass analysis branch extracting fine-grained features and the shallow-level human mass analysis branch extracting coarse-grained features. Because each human mass analysis branch learns information that affects the human mass score in a human image region from multiple dimensions at different granularities, each human mass analysis branch outputs initial quality prediction information in multiple dimensions at different scales. In step A2, the initial quality prediction information in multiple dimensions output by each human mass analysis branch is processed into feature maps with the same width and height through downsampling, and then, through splicing, convolution operations, and full-connection processing, the final quality prediction information in multiple dimensions of the human image region, i.e., the quality scores of the multiple dimensions, is obtained. Considering the different impact of each dimension on the human mass score, step A3 determines the human mass score based on the weights of the multiple dimensions. Moreover, the weights of multiple dimensions are not set manually, but are learnable parameters.

[0066] Specifically, taking a complete human body, a truncated human body, an obstructed human body, or a non-human body as an example, the human body mass fraction Q is expressed as: Q = w1 × D1 + w2 × D2 + w3 × D3 + w4 × D4

[0067] Among them, D1, D2, D3, and D4 represent the mass fraction of the complete human body, the mass fraction of the truncated human body, the mass fraction of the occluded human body, and the mass fraction of the non-human body, respectively. w1, w2, w3, and w4 represent the weight of the complete human body, the weight of the truncated human body, the weight of the occluded human body, and the weight of the non-human body, respectively.

[0068] In a specific embodiment, the human body mass analysis model includes: multiple parallel human body mass analysis branches, a first convolution module, multiple basic residual block groups connected in sequence, multiple downsampling modules, a splicing module, a second convolution module, a third convolution module and a fully connected module; wherein, the output end of the first module is connected to the first basic residual block group among the multiple basic residual block groups, the output ends of the multiple basic residual block groups are connected one-to-one with the input ends of the multiple human body mass analysis branches, the output ends of the multiple human body mass analysis branches are connected one-to-one with the input ends of the multiple downsampling modules, and each basic residual block group includes multiple superimposed basic residual blocks.

[0069] For example, Figure 5 is a structural diagram of a human body mass analysis model provided by an embodiment of the present application. Taking the three human body mass analysis branches as an example, the human body mass analysis model processes the human image area as follows: the human body image area is convolved by the first convolution module, and the feature map after convolution is subjected to three residual convolution operations in sequence through three basic residual block groups connected in sequence. Each time, the width and height of the feature map are reduced to half of the original, the number of channels is doubled, and each residual operation performs two basic residual block operations. Afterwards, a spatial attention module is connected to the back of each residual operation. The spatial attention module uses convolution kernels of 7×7, 5×5, and 3×3 to obtain the initial quality prediction information of multiple dimensions output by the three human body mass analysis branches. The downsampling module then performs downsampling processing, processing the initial quality prediction information of multiple dimensions output by each body mass analysis branch into feature maps with the same width and height. This is then concatenated and convolved twice through the splicing module, the second convolution module, and the third convolution module. Finally, the result is expanded into a single dimension and processed using the fully connected module to obtain the final quality prediction information for multiple dimensions (i.e., complete body, truncated body, occluded body, and non-body). Finally, the weight of each dimension is combined to obtain the body quality score.

[0070] In a specific embodiment, the weights of the multiple dimensions are learnable parameters, the human body mass analysis model is obtained by training a first preset model, and the training steps of the human body mass analysis model include steps B1 to B6:

[0071] Step B1: inputting the sample human body image region into the first preset model to obtain final quality prediction information of the sample human body image region in the multiple dimensions.

[0072] Step B2: Obtaining a human body quality score output by the first preset model for the sample human body image region based on the final quality prediction information of the multiple dimensions and the currently learned weights of the multiple dimensions.

[0073] Step B3: Obtain a first human body mass loss value according to the final quality prediction information of the sample human body image region in the multiple dimensions and the quality of the sample human body image region in the multiple dimensions.

[0074] Step B4: Obtain a second body mass loss value according to the body mass score output by the body mass analysis model for the sample body image region and the real body mass score pre-marked for the sample body image region.

[0075] Step B5: Obtaining a total human body mass loss value according to the first human body mass loss value and the second human body mass loss value.

[0076] Step B6: updating the parameters of the first preset model and the learned weights according to the total human body mass loss value.

[0077] In the embodiment of the present application, the first human body mass loss value is calculated based on the cross entropy loss. In order to enable the network to learn the weight coefficients that are adjusted autonomously according to each dimension to obtain the human body mass score, the mean square error loss is calculated based on the real human body mass score and the human body mass score output by the model for the sample human body image area to obtain the second human body mass loss value. Therefore, the total human body mass loss value loss is expressed as: loss = loss1 + loss2

[0078] Here, loss1 is the first body mass loss value, and loss2 is the second body mass loss value. Because the weights of multiple dimensions are determined through learning, multi-dimensional autonomous weight adjustment of body mass analysis is achieved, making the obtained body mass analysis results more accurate.

[0079] For example, Figure 5 is a structural diagram of a human body quality analysis model provided by an embodiment of the present application. Taking three human body quality analysis branches as an example, the human body quality analysis model processes the human body image area as follows: performing a convolution operation on the human body image area, and performing three residual convolution operations on the feature map after convolution. Each time, the width and height of the feature map are reduced to half of the original, the number of channels is doubled, and each residual operation performs two basic residual block operations. Afterwards, a spatial attention module is connected from the back of each residual operation. The spatial attention module uses convolution kernels of 7×7, 5×5, and 3×3 to obtain the initial quality prediction information of multiple dimensions output by the three human body quality analysis branches. Then, a downsampling processing method is used to process the initial quality prediction information of multiple dimensions output by each human body quality analysis branch into a feature map with the same width and height, and splicing and two convolution operations are performed. The obtained result is expanded into one dimension, and full connection is used to obtain the final quality prediction information of multiple dimensions (i.e., complete human body, human body truncation, human body occlusion, and non-human body). Finally, the weight of each dimension is combined to obtain the human body mass score.

[0080] In a specific embodiment, the body mass analysis model includes multiple parallel body mass analysis branches, each of which is a spatial attention module with different convolution kernel sizes. Each spatial attention module includes:

[0081] The first convolution unit, the mean calculation unit, the maximum calculation unit, the splicing unit, the convolution unit of a specific convolution kernel size, the weight calculation unit, and the multiplication unit;

[0082] The input ends of the mean calculation unit and the maximum calculation unit are respectively connected to the output end of the first convolution unit, the output ends of the mean calculation unit and the maximum calculation unit are respectively connected to the splicing unit, and the output ends of the weight calculation unit and the splicing unit are respectively connected to the multiplication unit.

[0083] The structure of the spatial attention module is shown in Figure 6. The spatial attention module processes the input feature map as follows: the first convolution unit convolves the input feature map to obtain a three-dimensional feature map of size W×H×C, where W, H, and C represent the width, height, and number of channels of the feature map, respectively. The mean and maximum values ​​of the feature map are calculated by the mean calculation unit and the maximum value calculation unit according to the channel dimension. The mean means calculating the average value of all channels in the entire feature map, and the maximum value selects the most influential feature channel. The size of the mean and maximum values ​​is W×H×1. The calculated mean and maximum features are then concatenated by the splicing unit to obtain a feature map of size W×H×2. The feature map is then convolved through a convolution unit with a specific convolution kernel size. The convolution kernel can be of different sizes, for example, 7×7, 5×5, 3×3, etc., to obtain a feature map of size W×H×1. The feature map is then sigmoid activated through the weight calculation unit to obtain W×H×1 features, that is, the weight of the input feature map. The W×H×1 features are then multiplied by the original features using the multiplication unit to obtain a feature map of the same size as the original input feature map.

[0084] Step S130: performing human attribute recognition on the human body image region whose human body quality meets the target quality condition to obtain a human body attribute recognition result.

[0085] In an embodiment of the present application, the target quality condition is determined according to the needs of the customer group attribute analysis. For example, if only the age and gender of the customer group need to be analyzed, the analysis can be achieved using the entire human body or part of the human body (i.e., the upper body), and the requirements for human body quality are relatively low at this time; if the customer group's clothing needs to be analyzed, the entire human body needs to be used to achieve the analysis, and the requirements for human body quality are relatively higher at this time. Therefore, different target quality conditions can be specified according to the needs of different customer group attribute analysis. By performing human attribute recognition on the human body image area whose body quality meets the target quality conditions, a human attribute recognition result is obtained, wherein the human attribute recognition result refers to the attribute information of the human body, for example, gender, age, orientation, top type, bottom type, whether to wear a mask, whether to carry a backpack, and other attribute information.

[0086] In an optional embodiment, the human image region for human attribute recognition is determined based on the human body mass score. Specifically, human attribute recognition is performed on the human image region whose human body mass meets the target quality condition to obtain the human attribute recognition result, including the following two methods:

[0087] Method 1: When the customer attribute analysis requirement is to use the complete human body for analysis, human attribute recognition is performed on the human image area whose human body quality meets the first quality threshold to obtain the human attribute recognition result.

[0088] Method 2: When the customer attribute analysis requirement is to use a partial human body for analysis, human attribute recognition is performed on the human image area whose human body quality meets a second quality threshold to obtain a human attribute recognition result; the second quality threshold is lower than the first quality threshold.

[0089] In the embodiment of the present application, human body quality is characterized by a human body mass score. A larger human body mass score indicates better human body quality. Since different customer attribute analysis requirements require different human body mass, different quality thresholds (i.e., a first quality threshold and a second quality threshold) are set according to the customer attribute analysis requirements to capture the human body image area required for customer attribute analysis. Human body quality meeting the first quality threshold means that the quality score is greater than the first threshold, and human body quality meeting the second quality threshold means that the quality score is greater than the second threshold.

[0090] For example, if it is necessary to count the clothing of a human body, the complete human body needs to be used for analysis, and the first quality threshold can be set to 0.75, and then the human clothing recognition is performed on the human body image area with a quality score greater than 0.75; if it is necessary to count the gender of the human body, the upper body (partial human body) can be used for analysis, and the second quality threshold can be set to 0.5, and then the human gender statistics are performed on the human body image area with a quality score greater than 0.5.

[0091] In an optional embodiment, human attribute recognition is performed on a human image region whose human quality meets the target quality conditions to obtain a human attribute recognition result, including: inputting the human image region whose human quality meets the target quality conditions into a pre-trained human attribute recognition model to obtain a human attribute recognition result output by the human attribute recognition model; the training sample of the human attribute recognition model is: a sample human image carrying a human attribute label, and the human attribute label is: a real human attribute pre-labeled for the sample human image.

[0092] In this embodiment of the present application, the human attribute labels refer to gender, age, orientation, top and bottom clothing type, whether a person is wearing a mask, and whether a person is carrying a backpack. Because the sample human images used to train the human attribute recognition model are labeled with real human attribute information, the trained human attribute recognition model is capable of recognizing human attribute information from human image regions.

[0093] Furthermore, the human attribute recognition model includes multiple parallel human attribute recognition branches, and the multiple human attribute recognition branches are used to: learn information affecting human attributes in the human image region from the multiple dimensions at different granularities; input the human image region whose human quality meets the target quality condition into the pre-trained human attribute recognition model, and obtain a human attribute recognition result output by the human attribute recognition model, including steps C1 and C2:

[0094] Step C1: inputting the human body image region whose human body quality meets the target quality condition into the human body attribute recognition model to obtain human body attribute prediction information output by each of the multiple human body attribute recognition branches.

[0095] Step C2: obtaining a human attribute recognition result output by the human attribute recognition model for the human image region according to the respective weights of the plurality of human attribute recognition branches and the human attribute prediction information output by the plurality of human attribute recognition branches.

[0096] In an embodiment of the present application, in order to extract effective human attribute features from the human image region, each human attribute recognition branch in step C1 uses a spatial attention mechanism to obtain effective human attribute features from the input feature map, and the multiple parallel human attribute recognition branches include a deep-level human attribute recognition branch and a shallow-level human attribute recognition branch. The deep-level human attribute recognition branch is used to extract fine-grained human attribute features, while the shallow-level human attribute recognition branch is used to extract coarse-grained human attribute features. Then, convolution kernels of different sizes are used to convert the human attribute features into feature maps of uniform size, and an average pooling operation is used to convert each feature map into a one-dimensional vector. After that, a fully connected layer is used to output human attribute prediction information. Considering that different human attribute recognition branches learn information that affects human attributes in the human image region from multiple dimensions at different granularities, the credibility of the human attribute prediction information output by each human attribute recognition branch is different. Therefore, step C2 determines the human quality score based on the weights of the multiple human attribute recognition branches. In addition, the weights of the multiple human attribute recognition branches are not manually set, but are learnable parameters.

[0097] Specifically, taking the three human attribute recognition branches as an example, the human attribute recognition result A is expressed as: A = (w5×A1+w6×A2+w7×A3) / 3

[0098] Among them, A1, A2, and A3 are the human attribute prediction information of the first human attribute recognition branch, the human attribute prediction information of the second human attribute recognition branch, and the human attribute prediction information of the third human attribute recognition branch, respectively. w5, w6, and w7 are the weights of the first human attribute recognition branch, the weights of the second human attribute recognition branch, and the weights of the third human attribute recognition branch, respectively.

[0099] In a specific embodiment, the weights of the plurality of human attribute recognition branches are learnable parameters, the human attribute recognition model is obtained by training the second preset model, and the training steps of the human attribute recognition model include steps D1 to D6:

[0100] Step D1: inputting the sample human body image into the second preset model to obtain human body attribute prediction information of each of the plurality of human body attribute recognition branches for the sample human body image.

[0101] Step D2: obtaining a human attribute prediction result output by the second preset model for the sample human image based on the human attribute prediction information of the multiple human attribute recognition branches and the currently learned weights of the multiple human attribute recognition branches.

[0102] Step D3: Obtaining loss values ​​corresponding to each of the multiple human attribute recognition branches based on the human attribute prediction information of each of the multiple human attribute recognition branches for the sample human image and the human attribute labels carried by the sample human image.

[0103] Step D4: Obtain a comprehensive human attribute loss value based on the human attribute prediction result output by the second preset model for the sample human image and the human attribute label carried by the sample human image.

[0104] Step D5: Obtaining a total human attribute loss value according to the loss values ​​corresponding to each of the plurality of human attribute recognition branches and the comprehensive human attribute loss value.

[0105] Step D6: Update the parameters of the second preset model and the learned weights according to the total loss value of the human body attributes.

[0106] In this embodiment of the present application, the loss values ​​corresponding to each of the multiple human attribute recognition branches are calculated based on the cross-entropy loss. For each human attribute prediction information output by each human attribute recognition branch, the network adaptively learns the weights of each human attribute recognition branch to obtain human attribute prediction results. The cross-entropy loss is then recalculated based on the human attribute prediction results and the human attribute labels carried by the sample human images to obtain a comprehensive human attribute loss value.

[0107] Taking the three human attribute recognition branches as an example, the total loss loss' of the human attribute recognition model is expressed as:

[0108] loss′=loss3+λ1×loss4+λ2×loss5+λ3×loss6

[0109] Among them, loss3 is the comprehensive human attribute loss value, loss4, loss5, and loss6 are the loss values ​​of the three human attribute recognition branches, and λ1, λ2, and λ3 are the weights of the three human attribute recognition branches. Generally, shallow human attribute recognition branches extract coarse-grained human attribute features, so the human attribute prediction results output by shallow human attribute recognition branches are less reliable, and accordingly, the weights of shallow human attribute recognition branches are smaller.

[0110] For example, Figure 7 is a structural diagram of a human attribute recognition model provided by an embodiment of the present application. With three human attribute recognition branches, the human attribute recognition model processes the human image area whose body quality meets the target quality conditions as follows: performing a convolution operation on the human image area, and performing four residual convolution operations on the feature map after the convolution. Each time, the width and height of the feature map are reduced to half of the original, and the number of channels is doubled. After the basic residual blocks 2, 3, and 4, the spatial attention modules are respectively connected, and then convolution kernels of different sizes are used to convert the feature maps into feature maps of the same size, and the average pooling operation is used to convert each feature map into a 512-dimensional one-dimensional vector. After that, the fully connected layer is used to obtain the human attribute prediction information output by each human attribute recognition branch. Finally, the weights of each human attribute recognition branch are combined to obtain the human attribute recognition result.

[0111] Step S140: determining the customer group attributes of the current area according to the human attribute recognition result and the positional relationship between the human bodies in the target image.

[0112] In the embodiment of the present application, the positional relationship between the various human bodies in the target image can reflect the occlusion between the human bodies. If the positions of two human bodies overlap or are very close, it indicates that there is occlusion between the two human bodies. If the positions of the two human bodies are far apart, it indicates that there is no occlusion between the two human bodies. Therefore, in order to obtain more accurate customer attributes, when the attribute recognition results of each human body are counted, the positional relationship between the various human bodies is combined to determine the customer attributes of the current area. This avoids the customer attribute analysis errors caused by occlusion caused by dense crowds and improves the accuracy of customer attribute analysis.

[0113] In an optional embodiment, another customer group attribute analysis method is provided. As shown in FIG8 , FIG8 is a flowchart of the steps of another customer group attribute analysis method provided in an embodiment of the present application, specifically including steps S810 to S870:

[0114] Step S810: Detect whether the target image captured for the current area contains a human body.

[0115] Step S820: performing a human body quality analysis on the human body image region in the target image containing the human body.

[0116] Step S830: Acquire a human body image sequence with the same human body identification.

[0117] Step S840: Filtering out a plurality of human body image regions whose human body quality meets the target quality condition from the human body image sequence.

[0118] Step S850: performing human attribute recognition on each of the plurality of human image regions.

[0119] Step S860: performing statistics on the human attribute recognition results of the plurality of human image regions to obtain the human attribute recognition result of the same human identifier.

[0120] Step S870: Determine the customer group attributes of the current area according to the human attribute recognition result and the positional relationship between the human bodies in the target image.

[0121] In the embodiment of the present application, the camera captures a target image sequence from the time a pedestrian enters the camera's field of view to the time they leave the camera's field of view. Therefore, each frame of the target image in the target image sequence is detected to determine whether the target image contains a human body. Specifically, a human body detection algorithm is used to detect each frame of the target image sequence in sequence. When a human body is detected, an identifier is assigned to the human body, and the human body is tracked during the next target image detection process (i.e., the same identifier is assigned to the same human body). Then, in step S830, the human body image areas corresponding to all human bodies carrying the same identifier in the target image sequence are used as a human body image sequence with the same human body identifier.

[0122] The quality of the human body in the human body image sequence is characterized by a human body quality score. A quality threshold is set based on the customer attribute analysis requirements. In step S840, human body images in the human body image sequence with a human body quality score greater than the quality threshold are considered as multiple human body image regions that meet the target quality condition. For example, if the human body quality scores of human body images 1, 2, 3, 4, and 5 are 0.2, 0.75, 0.85, 0.80, and 0.5, respectively, and the quality threshold is 0.7, then human body images 2, 3, and 4 are considered as multiple human body image regions that meet the target quality condition.

[0123] Human attribute recognition is performed on multiple human image regions. Each human image region has a human attribute recognition result, and each human attribute recognition result includes all human attribute information. Since different human image regions have different human qualities, the human attribute recognition results corresponding to different human image regions may be different. In order to obtain more accurate human attribute recognition results, in step S860, the human attribute recognition results are counted by voting on the multiple human attribute recognition results. Specifically, for each attribute information in the human attribute recognition result, voting is performed to determine whether the attribute information is accurate. When the voting result for the same attribute information is greater than a voting threshold (e.g., 60%), the attribute information is determined to be accurate attribute information, thereby obtaining an accurate human attribute recognition result.

[0124] In the embodiments of the present application, multiple human image regions corresponding to the same person are used to perform human attribute recognition, resulting in multiple human attribute recognition results. These results are then comprehensively considered to obtain a final human attribute recognition result for the person. Compared to using only a single human image region for human attribute recognition, this method can better eliminate errors in customer attribute analysis caused by factors such as human occlusion, incompleteness, and unclear images, thereby improving the accuracy of customer attribute analysis.

[0125] In an optional embodiment, after obtaining the customer attributes of the current area, the customer attributes of the current area are presented in the following manner: sending the customer attributes of the current area to a server or a display device at target time intervals; and / or when an event is detected from the appearance of the same human body identifier to leaving the current area, sending the customer attributes of the same human body identifier and the customer attributes of the current area before the appearance of the same human body identifier to the server or the display device.

[0126] In the embodiment of the present application, two presentation methods are provided: real-time presentation and delayed presentation. Real-time presentation means that the customer attributes of the current area are sent to the server or display device every target time period; delayed presentation means that when the event of the same human body identifier appearing and leaving the current area is detected, the customer attributes of the same human body identifier and the customer attributes of the current area before the same human body identifier appears are sent to the server or display device. During specific implementation, the presentation method of the customer attributes is determined according to the needs of customer attribute analysis. For example, when counting the customer base of a square at every moment, real-time presentation can be used; when counting the customer attributes of the flow of people in a park, delayed presentation is used, and the customer attribute results are used for later data analysis. In addition, the customer attributes can be presented through charts such as line charts, pie charts and bar charts. For example, a line chart can be used to present the changes in the number of people entering and leaving the current area within a period of time, a pie chart can be used to analyze the ratio of men and women, and a bar chart can be used to present the age distribution of the customer base.

[0127] For example, referring to FIG10 , FIG10 is an application diagram of a customer attribute analysis method provided by an embodiment of the present application. Specifically, an area in the park where customer attribute analysis is required is set. For each area where customer attribute analysis is required, a target image is captured using a camera. It is detected whether the target image captured for the current area contains a human body. A human body quality analysis is performed on the human body image area in the target image containing a human body. Then, human body attribute recognition is performed on the human body image area whose human body quality meets the target quality condition to obtain a human body attribute recognition result. Based on the human body attribute recognition result and the positional relationship between each human body in the target image, the customer attributes of the current area are determined. Finally, the customer attributes of the current area are displayed.

[0128] The present application also provides a customer group attribute analysis device, as shown in FIG10 , which is a schematic structural diagram of a customer group attribute analysis device provided in the present application. The device includes:

[0129] A detection module 1010 is configured to detect whether a target image captured for a current region contains a human body;

[0130] An analysis module 1020 is configured to perform a human body quality analysis on a human body image region in a target image containing a human body;

[0131] The recognition module 1030 is configured to perform human attribute recognition on a human image region whose human body quality meets a target quality condition, and obtain a human attribute recognition result;

[0132] The determination module 1040 is configured to determine the attributes of the customer group in the current area according to the human attribute recognition result and the positional relationship between the human bodies in the target image.

[0133] In an optional embodiment, the target quality condition is determined based on customer attribute analysis requirements; the identification module includes:

[0134] The first recognition submodule is configured to, when the customer attribute analysis requirement is to use a complete human body for analysis, perform human attribute recognition on a human body image region whose human body quality meets a first quality threshold, and obtain a human attribute recognition result;

[0135] The second identification submodule is used to perform human attribute recognition on the human image area whose human body quality meets the second quality threshold when the customer attribute analysis requirement is to use local human body for analysis, and obtain the human attribute recognition result; the second quality threshold is lower than the first quality threshold.

[0136] In an optional embodiment, the analysis module includes:

[0137] The quality model module is used to input the human body image area into a pre-trained human body quality analysis model to obtain a human body quality score output by the human body quality analysis model; the training samples of the human body quality analysis model are: sample human body image areas carrying human body quality labels, and the human body quality labels include: the real human body quality score pre-labeled for the sample human body image area, and the quality of the sample human body image area in multiple dimensions.

[0138] In an optional embodiment, the device further includes:

[0139] A first adjustment module is used to adjust the size of the human body image area to obtain a human body image area whose size meets the input size of the human body mass analysis model;

[0140] The second adjustment module is used to adjust the size of the sample human body image area to obtain a sample human body image area whose size meets the input size of the first model. The human body mass analysis model is obtained by training the first preset model.

[0141] In an optional embodiment, the human body mass analysis model includes at least a plurality of parallel human body mass analysis branches, wherein the plurality of human body mass analysis branches are used to learn information affecting the human body mass score in the human body image region from the plurality of dimensions at different granularities; the quality model module includes:

[0142] A first quality model submodule is configured to input the human body image region into the human body quality analysis model to obtain initial quality prediction information of the multiple dimensions output by each of the multiple human body quality analysis branches;

[0143] A second quality model submodule is configured to fuse the initial quality prediction information of the multiple dimensions outputted by the multiple human mass analysis branches through the human body quality analysis model to obtain final quality prediction information of the human body image region in the multiple dimensions;

[0144] The third quality model submodule is used to obtain the human body quality score output by the human body quality analysis model for the human body image area according to the respective weights of the multiple dimensions and the final quality prediction information of the human body image area in the multiple dimensions.

[0145] In an optional embodiment, the human body mass analysis model also includes: a first convolution module, multiple basic residual block groups connected in sequence, multiple downsampling modules, a splicing module, a second convolution module, a third convolution module and a fully connected module; wherein, the output end of the first module is connected to the first basic residual block group among the multiple basic residual block groups, the output ends of the multiple basic residual block groups are connected one-to-one with the input ends of the multiple human body mass analysis branches, the output ends of the multiple human body mass analysis branches are connected one-to-one with the input ends of the multiple downsampling modules, and each basic residual block group includes multiple superimposed basic residual blocks.

[0146] In an optional embodiment, the multiple body mass analysis branches are respectively multiple spatial attention modules with different convolution kernel sizes, and each spatial attention module includes:

[0147] The first convolution unit, the mean calculation unit, the maximum calculation unit, the splicing unit, the convolution unit of a specific convolution kernel size, the weight calculation unit, and the multiplication unit;

[0148] The input ends of the mean calculation unit and the maximum calculation unit are respectively connected to the output end of the first convolution unit, the output ends of the mean calculation unit and the maximum calculation unit are respectively connected to the splicing unit, and the output ends of the weight calculation unit and the splicing unit are respectively connected to the multiplication unit.

[0149] In an optional embodiment, the weights of each of the multiple dimensions are learnable parameters, the human body mass analysis model is obtained by training a first preset model, and the human body mass analysis model is obtained by training a quality training module, and the quality training module includes:

[0150] A first quality training submodule, configured to input the sample human image region into the first preset model to obtain final quality prediction information of the sample human image region in the multiple dimensions;

[0151] a second quality training submodule, configured to obtain a human body quality score output by the first preset model for the sample human body image region based on the final quality prediction information of the multiple dimensions and the currently learned weights of the multiple dimensions;

[0152] A third quality training submodule is configured to obtain a first human body mass loss value based on the final quality prediction information of the sample human body image region in the multiple dimensions and the quality of the sample human body image region in the multiple dimensions;

[0153] a fourth quality training submodule, configured to obtain a second body mass loss value based on the body mass score output by the body mass analysis model for the sample body image region and a real body mass score pre-marked for the sample body image region;

[0154] a fifth mass training submodule, configured to obtain a total body mass loss value based on the first body mass loss value and the second body mass loss value;

[0155] The sixth mass training submodule is used to update the parameters of the first preset model and the learned weights according to the total human body mass loss value.

[0156] In an optional embodiment, the multiple dimensions include at least: a complete human body, a truncated human body, an occluded human body, and a non-human body; and the true human body quality score of the sample human body image region is marked according to the following steps:

[0157] According to whether the sample human body image area includes a complete human body, whether the human body in the sample human body image area is an upper body or a lower body, whether the sample human body image area includes a real human body, and combined with the clarity and brightness of the sample human body image area, a human body quality score is marked for the sample human body image area.

[0158] In an optional embodiment, when the sample human body image region includes a complete human body, a human body quality score in a first range is marked for the sample human body image region;

[0159] In a case where the human body in the sample human body image region is an upper body, marking the sample human body image region with a human body mass score in a second range, where an upper limit value of the second range is less than or equal to a lower limit value of the first range;

[0160] In a case where the human body in the sample human body image region is a lower body body, marking the sample human body image region with a human body mass score in a third range, where an upper limit value of the third range is less than or equal to a lower limit value of the second range;

[0161] In a case where the human body in the sample human body image region is not a real human body, marking the sample human body image region with a human body mass score in a fourth range, where an upper limit value of the fourth range is less than or equal to a lower limit value of the third range;

[0162] The lower the clarity and brightness of the sample human body image region are, the closer the human body quality score marked for the sample human body image region is to the lower limit of the corresponding range.

[0163] In an optional embodiment, the identification module includes:

[0164] The recognition model module is used to input the human body image area whose human body quality meets the target quality conditions into a pre-trained human body attribute recognition model to obtain the human body attribute recognition result output by the human body attribute recognition model; the training sample of the human body attribute recognition model is: a sample human body image carrying a human body attribute label, and the human body attribute label is: the real human body attribute pre-marked for the sample human body image.

[0165] In an optional embodiment, the human attribute recognition model includes multiple parallel human attribute recognition branches, and the multiple human attribute recognition branches are used to learn information affecting human attributes in the human image region from the multiple dimensions at different granularities; the recognition model module includes:

[0166] A first recognition model submodule is configured to input a human image region whose human body quality meets a target quality condition into the human attribute recognition model to obtain human attribute prediction information output by each of the plurality of human attribute recognition branches;

[0167] The second recognition model submodule is used to obtain the human attribute recognition result output by the human attribute recognition model for the human image area according to the respective weights of the multiple human attribute recognition branches and the human attribute prediction information output by the multiple human attribute recognition branches.

[0168] In an optional embodiment, the weights of the multiple human attribute recognition branches are learnable parameters, the human attribute recognition model is obtained by training the second preset model, and the human attribute recognition model is obtained by training a recognition training module, and the recognition training module includes:

[0169] A first recognition training submodule is configured to input the sample human body image into the second preset model to obtain human attribute prediction information of each of the plurality of human attribute recognition branches for the sample human body image;

[0170] a second recognition training submodule, configured to obtain a human attribute prediction result output by the second preset model for the sample human image based on the human attribute prediction information of the multiple human attribute recognition branches and the currently learned weights of the multiple human attribute recognition branches;

[0171] a third recognition training submodule, configured to obtain a loss value corresponding to each of the plurality of human attribute recognition branches based on the human attribute prediction information of each of the plurality of human attribute recognition branches for the sample human image and the human attribute label carried by the sample human image;

[0172] a fourth recognition training submodule, configured to obtain a comprehensive human attribute loss value based on the human attribute prediction result output by the second preset model for the sample human image and the human attribute label carried by the sample human image;

[0173] a fifth recognition training submodule, configured to obtain a total human attribute loss value based on the loss values ​​corresponding to the plurality of human attribute recognition branches and the comprehensive human attribute loss value;

[0174] The sixth recognition training submodule is used to update the parameters of the second preset model and the learned weights according to the total loss value of the human attributes.

[0175] In an optional embodiment, the identification module includes:

[0176] A sequence acquisition module is used to acquire a sequence of human body images with the same human body identification;

[0177] An image screening module, configured to screen out a plurality of human body image regions whose human body quality meets the target quality condition from the human body image sequence;

[0178] An attribute recognition module, configured to perform human attribute recognition on each of the plurality of human image regions;

[0179] The attribute statistics module is used to collect statistics on the human attribute recognition results of the multiple human image regions to obtain the human attribute recognition results of the same human identifier.

[0180] In an optional embodiment, the device includes:

[0181] A presentation module is configured to send the customer attributes of the current area to a server or a display device at target time intervals; and / or when an event is detected from the appearance of the same human body identifier to the departure from the current area, send the customer attributes of the same human body identifier and the customer attributes of the current area before the appearance of the same human body identifier to the server or the display device.

[0182] The present application also provides an electronic device, as shown in FIG11 , which is a schematic diagram of the structure of an electronic device provided in the present application. As shown in FIG11 , electronic device 1100 includes a memory 1110 and a processor 1120 . Memory 1110 and processor 1120 are connected via a bus. Memory 1110 stores a computer program that can be executed on processor 1120 to implement the steps of the customer attribute analysis method described in the present application.

[0183] The embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the customer attribute analysis method described in the embodiment of the present application are implemented.

[0184] The embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the customer attribute analysis method described in the embodiment of the present application.

[0185] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0186] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices, and equipment according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0187] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0189] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0190] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0191] The above is a detailed introduction to the customer attribute analysis method, device, electronic device and storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A customer group attribute analysis method, characterized in that: include: Detect whether the target image collected for the current area contains a human body; Performing human body quality analysis on a human body image region in a target image containing a human body; Performing human attribute recognition on the human image region whose human quality meets the target quality condition to obtain a human attribute recognition result; The customer group attributes of the current area are determined according to the human attribute recognition result and the positional relationship between the human bodies in the target image.

2. The method according to claim 1, characterized in that The target quality conditions are determined based on the needs of customer group attribute analysis; Perform human attribute recognition on the human image area whose human quality meets the target quality condition, and obtain human attribute recognition results, including: When the customer attribute analysis requirement is to use a complete human body for analysis, human attribute recognition is performed on a human image region whose human quality meets a first quality threshold to obtain a human attribute recognition result; When the customer attribute analysis requirement is to use a partial human body for analysis, human attribute recognition is performed on a human image region whose human quality meets a second quality threshold to obtain a human attribute recognition result; the second quality threshold is lower than the first quality threshold.

3. The method according to claim 1, characterized in that Performing human quality analysis on a human image region in a target image containing a human body, including: Inputting the human body image region into a pre-trained human body mass analysis model to obtain a human body mass score output by the human body mass analysis model; The training samples of the human body quality analysis model are: sample human body image regions carrying human body quality labels, wherein the human body quality labels include: a real human body quality score pre-labeled for the sample human body image region, and the quality of the sample human body image region in multiple dimensions.

4. The method according to claim 3, characterized in that: Before inputting the human body image region into a pre-trained human body quality analysis model, the method further comprises: Adjusting the size of the human body image region to obtain a human body image region whose size meets the input size of the human body mass analysis model; The human body mass analysis model is obtained by training the first preset model. Before the human body image region is input into the first model, the method further includes: The size of the sample human body image region is adjusted to obtain a sample human body image region whose size meets the input size of the first model.

5. The method according to claim 3, characterized in that: The human body mass analysis model comprises at least a plurality of parallel human body mass analysis branches, and the plurality of human body mass analysis branches are used to: learn information affecting the human body mass score in the human body image region from the plurality of dimensions at different granularities; Inputting the human body image region into a pre-trained human body mass analysis model to obtain a human body mass score output by the human body mass analysis model, comprising: Inputting the human body image region into the human body quality analysis model to obtain initial quality prediction information of the multiple dimensions outputted by each of the multiple human body quality analysis branches; The initial quality prediction information of the multiple dimensions outputted by the multiple human body quality analysis branches is merged by the human body quality analysis model to obtain final quality prediction information of the human body image region in the multiple dimensions; According to the respective weights of the multiple dimensions and the final quality prediction information of the human image region in the multiple dimensions, a human body quality score output by the human body quality analysis model for the human body image region is obtained.

6. The method according to claim 5, characterized in that The human body mass analysis model also includes: a first convolution module, multiple basic residual block groups connected in sequence, multiple downsampling modules, a splicing module, a second convolution module, a third convolution module and a fully connected module; wherein the output end of the first module is connected to the first basic residual block group among the multiple basic residual block groups, the output ends of the multiple basic residual block groups are connected one-to-one with the input ends of the multiple human body mass analysis branches, the output ends of the multiple human body mass analysis branches are connected one-to-one with the input ends of the multiple downsampling modules, and each basic residual block group includes multiple superimposed basic residual blocks.

7. The method according to claim 6, characterized in that The multiple human mass analysis branches are respectively multiple spatial attention modules with different convolution kernel sizes, and each spatial attention module includes: A first convolution unit, a mean calculation unit, a maximum calculation unit, a concatenation unit, a convolution unit of a specific convolution kernel size, a weight calculation unit, and a multiplication unit; The input end of each of the mean calculation unit and the maximum calculation unit is connected to the output end of the first convolution unit, and the output end of each of the mean calculation unit and the maximum calculation unit is connected to the output end of the first convolution unit. The weight calculation unit is connected to the splicing unit, and the output ends of the weight calculation unit and the splicing unit are respectively connected to the multiplication unit.

8. The method according to claim 3, characterized in that The weights of each of the multiple dimensions are learnable parameters, and the human body mass analysis model is obtained by training the first preset model. The training steps of the human body mass analysis model include: Inputting the sample human image region into the first preset model to obtain final quality prediction information of the sample human image region in the multiple dimensions; Obtaining a human body quality score output by the first preset model for the sample human body image region according to the final quality prediction information of the multiple dimensions and the currently learned weights of the multiple dimensions; Obtaining a first human body mass loss value according to the final quality prediction information of the sample human body image region in the multiple dimensions and the quality of the sample human body image region in the multiple dimensions; Obtaining a second body mass loss value according to the body mass score output by the body mass analysis model for the sample body image region and the real body mass score pre-marked for the sample body image region; Obtaining a total human body mass loss value according to the first human body mass loss value and the second human body mass loss value; According to the total human body mass loss value, the parameters of the first preset model and the learned weights are updated.

9. The method according to claim 3, characterized in that: The multiple dimensions at least include: a complete human body, a truncated human body, an occluded human body, and a non-human body; the real human body quality score of the sample human body image area is marked according to the following steps: According to whether the sample human body image area includes a complete human body, whether the human body in the sample human body image area is an upper body or a lower body, and whether the sample human body image area includes a real human body, combined with the clarity and brightness of the sample human body image area, a human body quality score is marked for the sample human body image area.

10. The method according to claim 9, characterized in that In a case where the sample human body image region includes a complete human body, marking the sample human body image region with a human body mass score in a first range; In the case where the human body in the sample human body image area is an upper body human body, the sample The human body image area mark is in a second range of human body mass scores, and the upper limit value of the second range is less than or equal to the lower limit value of the first range; In the case where the human body in the sample human body image region is a lower body body, marking the sample human body image region with a human body mass score in a third range, wherein the upper limit value of the third range is less than or equal to the lower limit value of the second range; In the case where the human body in the sample human body image region is not a real human body, marking the sample human body image region with a human body mass score in a fourth range, wherein the upper limit value of the fourth range is less than or equal to the lower limit value of the third range; The lower the clarity and brightness of the sample human body image region, the closer the human body quality score marked for the sample human body image region is to the lower limit value of the corresponding range.

11. The method according to claim 1, characterized in that: Perform human attribute recognition on the human image area whose human quality meets the target quality condition, and obtain human attribute recognition results, including: Inputting the human body image region whose human body mass meets the target quality condition into a pre-trained human body attribute recognition model to obtain a human body attribute recognition result output by the human body attribute recognition model; The training samples of the human attribute recognition model are: sample human images carrying human attribute labels, and the human attribute labels are: real human attributes pre-labeled for the sample human images.

12. The method according to claim 11, characterized in that The human attribute recognition model comprises a plurality of parallel human attribute recognition branches, and the plurality of human attribute recognition branches are used to: learn information affecting human attributes in the human image region from the plurality of dimensions at different granularities; Inputting the human body image region whose human body quality meets the target quality condition into a pre-trained human body attribute recognition model to obtain a human body attribute recognition result output by the human body attribute recognition model, including: Inputting the human body image region whose human body quality meets the target quality condition into the human body attribute recognition model to obtain human body attribute prediction information output by each of the multiple human body attribute recognition branches; According to the respective weights of the plurality of human attribute recognition branches and the human attribute prediction information output by the plurality of human attribute recognition branches, a human attribute recognition result output by the human attribute recognition model for the human image region is obtained.

13. The method according to claim 12, characterized in that The weights of the multiple human attribute recognition branches are learnable parameters. The human attribute recognition model is obtained by training the second preset model. The training steps of the human attribute recognition model include: The sample human body image is input into the second preset model to obtain the multiple human body attribute recognition The branches each predict human attribute information for the sample human image; Obtaining a human attribute prediction result output by the second preset model for the sample human image according to the human attribute prediction information of the multiple human attribute recognition branches and the currently learned weights of the multiple human attribute recognition branches; According to the human attribute prediction information of each of the multiple human attribute recognition branches for the sample human image and the human attribute label carried by the sample human image, obtaining the loss value corresponding to each of the multiple human attribute recognition branches; Obtaining a comprehensive human attribute loss value according to the human attribute prediction result output by the second preset model for the sample human image and the human attribute label carried by the sample human image; Obtaining a total human attribute loss value according to the loss values ​​corresponding to each of the plurality of human attribute recognition branches and the comprehensive human attribute loss value; According to the total loss value of the human body attribute, the parameters of the second preset model and the learned weights are updated.

14. The method according to claim 1, characterized in that Perform human attribute recognition on the human image area whose human quality meets the target quality condition, and obtain human attribute recognition results, including: Obtain a human image sequence of the same human body identification; Screening out a plurality of human body image regions whose human body quality meets the target quality condition from the human body image sequence; Performing human attribute recognition on the multiple human image regions respectively; The human attribute recognition results of the multiple human image regions are counted to obtain the human attribute recognition result of the same human body identifier.

15. The method according to any one of claims 1 to 14, characterized in that: Also includes: Sending the customer group attributes of the current area to a server or a display device at target time intervals; and / or When an event from the same human body identifier appearing to leaving the current area is detected, the customer group attributes of the same human body identifier and the customer group attributes of the current area before the same human body identifier appears are sent to the server or the display device.

16. A customer group attribute analysis device, characterized in that: include: A detection module is used to detect whether the target image collected for the current area contains a human body; The analysis module is used to analyze the human body quality of the human body image area in the target image containing the human body. analyze; The recognition module is used to perform human attribute recognition on the human image area whose human quality meets the target quality condition to obtain the human attribute recognition result; The determination module is used to determine the customer group attributes of the current area according to the human attribute recognition result and the positional relationship between each human body in the target image.

17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the customer group attribute analysis method described in any one of claims 1-15 are implemented.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the customer group attribute analysis method described in any one of claims 1 to 15 are implemented.

19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the customer group attribute analysis method as described in any one of claims 1 to 15 are implemented.