A method and system for mitigating intra-class variation in pedestrian attribute recognition

By introducing the exponential information bottleneck method into the pedestrian attribute recognition system and filtering out redundant feature information, the robustness problem of intra-class changes in pedestrian attribute recognition is solved, and the accuracy and robustness of recognition are improved.

CN117095329BActive Publication Date: 2025-10-24XIAMEN MEIYA PICO INFORMATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310959022.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-01
Publication Date
2025-10-24
Estimated Expiration
2043-08-01

AI Technical Summary

Technical Problem

现有技术中行人属性识别存在类内变化的鲁棒性不强,导致属性相关特征冗余、不相关和噪声问题。

Method used

The exponential information bottleneck method is applied to each convolution module and attention mechanism module of the backbone network. The redundant interference feature information is filtered through the binary cross entropy loss function and the exponential information bottleneck module. The feature information is compressed using conditional probability and the mutual information is maximized to improve the recognition accuracy.

Benefits of technology

有效去除了冗余信息,提高了行人属性识别的准确性和鲁棒性,特别是在不同行人图像中的属性表示一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095329B_ABST
    Figure CN117095329B_ABST
Patent Text Reader

Abstract

A method and system for mitigating intra-class variation in pedestrian attribute recognition are disclosed, including receiving any image x in a dataset i The pedestrian attribute recognition inputting a backbone network for a multi-label classification task adopts binary cross-entropy as a loss function; an exponential information bottleneck is used for each convolution module and an attention mechanism module of the backbone network to filter redundant and interfering feature information existing in the features. The exponential information bottleneck method proposed in the application can be integrated into the attention module to form a novel pedestrian attribute recognition network, which can further process intra-attribute variation based on the attention mechanism. The exponential information bottleneck method is plug-and-play, without any additional computational overhead during inference.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pedestrian attribute recognition, and in particular to a method and system for mitigating intra-class variation of pedestrian attribute recognition. BACKGROUND

[0002] Pedestrian attribute recognition (PAR) has been widely used in video surveillance, which works by recognizing a series of semantic attributes of people, such as age, gender, clothing style, etc. Current research in this field shows that PAR is still challenging. In different pedestrian images, due to the changes of background, lighting, viewing angle and posture, there are significant intra-attribute differences in the captured same attributes. Specifically, even the same pedestrian attribute has different representations in different pedestrian images. Intra-attribute variation will lead to redundancy, irrelevance and noise of attribute-related features.

[0003] To solve the problem of intra-class variation of pedestrian attribute recognition, researchers have made great efforts and achieved a series of significant superior performance. These methods can be divided into body prior-based methods and attention mechanism-based methods in operation. These methods mainly use body prior knowledge or attention mechanism to localize attribute-related regions, aiming to enrich the local representation of attribute-related regions. For body prior knowledge methods, body region proposal methods are used to extract body prior knowledge, which can learn discriminative features. In theory, the extracted features should be robust. However, it largely depends on the performance of the region detector. Since the region proposal method is not specifically designed for PAR, it is difficult to ensure the reliability of the extracted attribute-related regions. On the other hand, by embedding or designing attention modules, the region detector is replaced to drive the network to focus on attribute-related regions and extract effective information. However, these methods with attention modules implicitly assume that the attention module can achieve accurate positioning related to pedestrian attributes, but this is not always accurate. These methods try to locate attribute-related regions and extract robust features to overcome the redundancy, irrelevance and noise caused by intra-attribute variation. SUMMARY

[0004] In order to solve the technical problem that the existing multi-label pedestrian attribute recognition does not have strong robustness to intra-attribute variation, the present application proposes a method and system for mitigating intra-class variation of pedestrian attribute recognition to solve the above technical problems.

[0005] According to a first aspect of the present application, a method for mitigating intra-class variation of pedestrian attribute recognition is provided, comprising:

[0006] S1: receiving any image x in the data set i The pedestrian attribute recognition inputting the backbone network for the multi-label classification task adopts binary cross-entropy as the loss function;

[0007] S2: Each convolutional module and attention mechanism module of the backbone network is affected by an exponential information bottleneck, which filters redundant and interfering feature information in the features.

[0008] In some specific embodiments, the image x i The corresponding pedestrian attribute label y i ∈{0, 1} M where M represents the number of pedestrian attribute categories, and 1 is marked in the attribute y i corresponding to the presence of the attribute, and 0 otherwise.

[0009] In some specific embodiments, the loss function expression of binary cross-entropy is where, represents the image x i is input into the algorithm framework, and the predicted probability value of the jth attribute in the image x i r j is the proportion of positive samples of the jth attribute in the training set.

[0010] In some specific embodiments, a plurality of exponential information bottlenecks L represents the number of exponential information bottleneck modules, and the information fed forward in the framework can be interpreted as a continuous representation of a Markov chain layer by layer, and the expression is where x is the input of the exponential information bottleneck.

[0011] In some specific embodiments, each exponential information bottleneck encodes feature information compression using the conditional probability p(t i |t i-1 ), mainly to minimize the mutual information I(t i ; t i-1 ), filter redundant information between layers t i and layers t i-1 in the network convolutional module or the attention mechanism module, and at the same time, the exponential information bottleneck maximizes the mutual information I(t i ; y) between t i and the output probability y to promote the accuracy of the final attribute recognition.

[0012] In some specific embodiments, the overall constraint of the exponential information bottleneck is where β i i=0,…,L is used to determine the compression degree of the mutual information at each layer.

[0013] ​According to a second aspect of the present application, a computer readable storage medium is provided, having stored thereon one or more computer programs which, when executed by a computer processor, implement the method described above.

[0014] According to a third aspect of the present application, a system for mitigating intra-class variation of pedestrian attribute recognition is provided, comprising:

[0015] a data input module configured to receive any image x in a data set i The pedestrian attribute recognition input into the main network for the multi-label classification task adopts binary cross-entropy as the loss function.

[0016] a redundant interference feature filtering module configured to filter the redundant interference feature information existing in the features by using the exponential information bottleneck acting on each convolution module and the attention mechanism module of the main network.

[0017] In some specific embodiments, the image x i corresponding pedestrian attribute label y i ∈{0,1} M , wherein M represents the number of classes of pedestrian attributes, and 1 is marked in the corresponding y i if the attribute exists, otherwise 0 is marked.

[0018] In some specific embodiments, the loss function expression of binary cross-entropy is , wherein represents the image x i input into the algorithm framework, and the predicted probability value of the jth attribute in the image x i , r j is the proportion of positive samples of the jth attribute in the training set.

[0019] In some specific embodiments, a plurality of exponential information bottlenecks L represents the number of exponential information bottleneck modules, and the information fed forward in the framework can be explained as a continuous representation of a Markov chain layer by layer, and the expression is , wherein x is the input of the exponential information bottleneck; each exponential information bottleneck encodes the feature information compression using the conditional probability p(t i |t i-1 ), mainly minimizes the mutual information I(t i ; t i-1 ), and filters out the redundant information between layer t i and layer t i-1 in the network convolution module or the attention mechanism module, while maximizing t iand the mutual information I(t i ; y) facilitates the accuracy of the final attribute recognition.

[0020] In some specific embodiments, the exponential information bottleneck constrains the overall where β i i = 0, …, L is used to determine the compression degree of mutual information of each layer.

[0021] The present application proposes a method and system for alleviating intra-class variation of pedestrian attribute recognition. The exponential information bottleneck method is proposed to solve intra-class variation. When the exponential information bottleneck method is applied alone in the backbone network, the proposed exponential information bottleneck method can use the compressed mutual information as the main penalty at the beginning of training to remove negative redundant information. When the mutual information penalty is weakened, the binary cross-entropy loss function has a greater contribution to improving the recognition accuracy of PAR. In addition, the proposed exponential information bottleneck method can be integrated into the attention module to form a novel pedestrian familiar recognition network, which can further process intra-class variation based on attention mechanism. In addition, the exponential information bottleneck method of the present application is plug-and-play, without any additional computational overhead during inference. BRIEF DESCRIPTION OF DRAWINGS

[0022] The accompanying drawings are included to provide a further understanding of embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments and serve the purpose of explaining the principles of the application. Other embodiments and many of the intended advantages of the embodiments will be readily appreciated as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings. Other features, objects, and advantages of the application will become apparent from the detailed description of the non-limiting embodiments made by way of illustration:

[0023] Figure 1 is a flowchart of a method for alleviating intra-class variation of pedestrian attribute recognition according to an embodiment of the present application;

[0024] Figure 2 is a block diagram of a method for alleviating intra-class variation of pedestrian attribute recognition according to a specific embodiment of the present application;

[0025] Figure 3 is a system architecture diagram of a method for alleviating intra-class variation of pedestrian attribute recognition according to an embodiment of the present application;

[0026] Figure 4 is a structural schematic diagram of a computer system of an electronic device suitable for implementing embodiments of the present application. DETAILED DESCRIPTION

[0027] The application will be described in further detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related application and not to limit the application. In addition, it should be noted that only parts related to the application are shown in the drawings for ease of description.

[0028] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and embodiments.

[0029] Figure 1 A method flowchart for slowing down intra-class variation of pedestrian attribute recognition is shown according to an embodiment of the present application. As shown in the figure, the method comprises the following steps: Figure 1

[0030] S101: Receiving any image x in the data set i The pedestrian attribute recognition input into the backbone network for the multi-label classification task adopts binary cross-entropy as the loss function.

[0031] In a specific embodiment, the image x i corresponds to the pedestrian attribute label y i ∈{0,1} M , where M represents the number of pedestrian attribute classes, and 1 is marked in the attribute existing corresponding y i , otherwise 0 is marked.

[0032] In a specific embodiment, the loss function expression of binary cross-entropy is wherein, represents that the image x i is input into the algorithm framework, and the predicted probability value of the jth attribute in the image x i , r j is the proportion of positive samples of the jth attribute in the training set.

[0033] S102: Use the exponential information bottleneck to act on each convolutional module and attention mechanism module of the backbone network, and filter the redundant and interfering feature information existing in the features.

[0034] In a specific embodiment, the exponential information bottleneck L represents the number of exponential information bottleneck modules, and the information fed forward in the framework can be explained as a continuous representation of a Markov chain layer by layer, and the expression is wherein, x is the input of the exponential information bottleneck.

[0035] ​In specific embodiments, each exponential information bottleneck adopts conditional probability p(t i |t i-1 ) to encode feature information compression, mainly minimizing mutual information I(t i ; t i-1 ), filtering out redundant information between layer t i and layer t i-1 in the network convolution module or the attention mechanism module, while using the exponential information bottleneck to maximize the mutual information I(t i ; y) between t i and output probability y to promote the accuracy of the final attribute recognition.

[0036] In specific embodiments, the exponential information bottleneck constrains the overall where β i i=0,…,L is used to determine the compression degree of the mutual information of each layer.

[0037] The overall framework of a method for mitigating intra-class variation in pedestrian attribute recognition according to the present application is shown in Figure 2 The entire method includes a backbone network (composed of multiple convolution modules), an attention mechanism module, and an exponential information bottleneck module. To achieve the use of the proposed exponential information bottleneck theory to mitigate intra-class variation in pedestrian attribute recognition algorithms, the following steps are mainly included:

[0038] Step 1: The above Figure 2 receives any image x i in the data set D as input, aiming to identify multiple pedestrian attributes, and the corresponding label of the image x i is y i ∈{0,1} M , M represents the number of pedestrian attribute categories, and 1 is used to mark the presence of the corresponding attribute y i , and 0 is used to mark the absence of the corresponding attribute.

[0039] Step 2: The entire pedestrian attribute recognition problem is regarded as a multi-label classification task, and binary cross-entropy (BCELoss) is used as the loss function, and the expression form of the loss function is as follows: wherein, represents the input of the image x i into the algorithm framework, and ω j is the predicted probability value of the jth attribute in the image x i , and is as follows: wherein, r j is the proportion of positive samples of the jth attribute in the training set.

[0040] Step 3: The exponential information bottleneck proposed in the application respectively acts on each convolution module and attention mechanism module of the backbone network, aiming to compress the mutual information as much as possible, that is, to filter out the redundant and interfering feature information in the features. The entire framework contains multiple exponential information bottlenecks L represents the number of exponential information bottleneck modules, and the information fed forward in the framework can be explained as a continuous representation of a layer-by-layer Markov chain, and the expression is as follows: Among them, x can be regarded as the input of the exponential information bottleneck, and for convenience, it is recorded as t0.

[0041] Step 4: Each exponential information bottleneck adopts conditional probability p(t i |t i-1 ) to encode feature information compression, that is, mainly to minimize the mutual information I(t i ; t i-1 ), filter out the redundant information between layer t i and layer t i-1 in the network convolution module or the attention mechanism module, and at the same time, the exponential information bottleneck maximizes the mutual information I(t i ; y) of t i and output probability y, so as to promote the accuracy of the final attribute recognition. The constraint of the exponential information bottleneck on the entire framework can be expressed as: Among them, β i ,i=0,…,L is used to determine the compression degree of the mutual information of each layer.

[0042] The application proposes a novel and flexible PAR framework, aiming to solve the intra-attribute variation through the proposed exponential information bottleneck method. When the exponential information bottleneck method is applied alone in the backbone network, the proposed exponential information bottleneck method can use the compressed mutual information as the main penalty at the beginning of training to remove negative redundant information. When the mutual information penalty is weakened, the binary cross-entropy loss function has a greater contribution to improving the recognition accuracy of the PAR. In addition, the proposed exponential information bottleneck method can be integrated into the attention module to form a novel pedestrian familiar recognition network, which can further process the intra-attribute variation based on the attention mechanism. In addition, the exponential information bottleneck method of the application is plug-and-play, without any additional computational overhead during inference. Compared with the prior art, the application directly suppresses redundancy, irrelevance and noise, which can significantly help the PAR model to locate the attribute-related region and extract robust features. And the experiments on multiple data sets verify that the PAR model of the method proposed in the application is superior to other advanced methods.

[0043] Figure 3 The system architecture diagram for slowing down the intra-attribute variation of pedestrian attribute recognition of one embodiment of the application is shown in FIG. 1. Figure 3As shown, the system comprises a data input module 301 and a redundant interference feature filtering module 302, wherein the data input module 301 is configured to receive any image x in the data set i The pedestrian attribute recognition using the backbone network for the multi-label classification task adopts binary cross-entropy as a loss function, and the redundant interference feature filtering module 302 is configured to use exponential information bottleneck to act on each convolution module and attention mechanism module of the backbone network, and filter the redundant interference feature information existing in the features.

[0044] Reference will now be made to the following description Figure 4 which shows a structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present application.

[0045] As Figure 4 shown, the computer system comprises a central processing unit (CPU) 401 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 402 or programs loaded from a storage portion 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the system 400 are also stored. The CPU 401, the ROM 402 and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0046] The following components are connected to the I / O interface 405: an input portion 406 including a keyboard, a mouse, etc.; an output portion 407 including a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 408 including a hard disk, etc.; and a communication portion 409 including a network interface card such as a LAN card, a modem, etc. The communication portion 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as necessary. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 410 as necessary, so that a computer program read therefrom is installed in the storage portion 408 as necessary.

[0047] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program in accordance with embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable storage medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from the removable media 41 1. When the computer program is executed by the central processing unit (CPU) 401, the above-described functions defined in the methods of the present application are performed. It should be noted that the computer readable storage medium of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, be - but is not limited to - an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present application, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable storage medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wire line, optical fiber, RF, or any suitable combination of the above.

[0048] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0049] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0050] The modules involved in the embodiments of the present application can be implemented in the form of software, or can be implemented in the form of hardware.

[0051] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately and not be assembled into the electronic device. The above computer readable storage medium carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: receive any image x i The pedestrian attribute recognition inputting the main network for the multi-label classification task adopts the binary cross entropy as the loss function; the exponential information bottleneck is used for each convolution module and the attention mechanism module of the main network to filter the redundant interference feature information existing in the features.

[0052] The above description is only the preferred embodiment of the present application and the explanation of the technical principles. It should be understood by those skilled in the art that the inventive scope of the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present application (but not limited to) with similar functions.

Claims

1. A method of mitigating intra-class variation in pedestrian attribute recognition, the method comprising: Comprising: S1: receive any one image x in the data set i The pedestrian attribute recognition using the input backbone network for the multi-label classification task adopts a binary cross-entropy as a loss function. S2: filtering redundant interference feature information existing in the features by using an exponential information bottleneck on each convolution module and attention mechanism module of the backbone network; Including multiple exponential information bottlenecks L represents the number of exponential information bottleneck modules, the information fed forward in the framework can be interpreted as a continuous representation of a layer-by-layer Markov chain, expressed as Where x is the input of the exponential information bottleneck. Each of the index information bottleneck adopts conditional probability p(t i ) i-1 ) to encode feature information compression, mainly minimizes mutual information I(t i ; t i-1 ), filters out the redundant information between layer t i and layer t i-1 in the network convolution module or the attention mechanism module, and simultaneously utilizes the index information bottleneck to maximize the mutual information I(t i ; y) between t i and output probability y to promote the accuracy of the final attribute recognition. The index information bottleneck constrains the whole where β i where i = 0,..., L is used to determine the compression degree of mutual information of each layer.

2. The method of claim 1, wherein, The image x i Corresponding pedestrian attribute label y i ∈{0,1} M Wherein M represents the number of pedestrian attribute categories, and the corresponding y i Is marked with 1, otherwise 0.

3. The method of claim 1, wherein, The loss function expression of the binary cross-entropy wherein, represents the image x i into the algorithm framework, the image x i The predicted probability value of the jth attribute in the image x, r j is the proportion of positive samples of the jth attribute in the training set.

4. A computer readable storage medium having stored thereon one or more computer programs, The one or more computer programs, when executed by a computer processor, implement the method of any one of claims 1-3.

5. A system for mitigating intra-class variation in pedestrian attribute recognition, the system comprising: Comprising: a data input module configured to receive any image x in a data set i The loss function of the multi-label classification task of the input backbone network for pedestrian attribute recognition is binary cross-entropy. a redundant interference feature filtering module configured to filter redundant interference feature information existing in the features by using an exponential information bottleneck on each convolution module and attention mechanism module of the backbone network; Including multiple exponential information bottlenecks L represents the number of exponential information bottleneck modules, the network feedforward information in the framework can be interpreted as a continuous representation of layer-by-layer Markov chain, expressed as Wherein, x is the input of the exponential information bottleneck; each of the exponential information bottlenecks adopts conditional probability p(t i |t i-1 ) to encode feature information compression, mainly minimizing mutual information I(t i ; t i-1 ), filtering out the redundant information between layer t i and layer t i-1 in the network convolution module or the attention mechanism module, and simultaneously maximizing the mutual information I(t i ; y) of t i and output probability y to promote the accuracy of the final attribute recognition; The index information bottleneck constrains the whole where β i where i = 0,..., L is used to determine the compression degree of mutual information of each layer.

6. The system for mitigating intra-class variation in pedestrian attribute recognition of claim 5, wherein, The image x i Corresponding pedestrian attribute label y i ∈{0,1} M Wherein M represents the number of pedestrian attribute categories, and the corresponding y i Is marked with 1, otherwise 0.

7. The system for mitigating intra-class variation in pedestrian attribute recognition of claim 5, wherein, The loss function expression of the binary cross-entropy wherein, representing the image x i into the algorithm framework, the image x i The predicted probability value of the jth attribute, r j is the proportion of positive samples of the jth attribute in the training set.