Model training method and system for human body weight recognition model and electronic equipment

By regional division and local feature extraction of the global features of the human body, and training the model with the appearance loss function, the problem of the human weight recognition model being affected by posture and occlusion is solved, and the recognition accuracy is improved.

CN120375422APending Publication Date: 2025-07-25CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455809.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing human weight recognition model relies on global characteristics and is susceptible to factors such as human posture, environment and occlusion, resulting in a low recognition accuracy.

Method used

By dividing the global characteristics of the human body, local features of multiple human body areas are extracted, and local appearance features are classified and trained using appearance loss functions to generate a human weight recognition model.

Benefits of technology

It improves the perception ability of the human weight recognition model under fine-grained size and sensitivity to local features, enhances the ability to capture texture and color differences, and improves the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375422A_ABST
    Figure CN120375422A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of model training, and discloses a model training method and system for a human body weight recognition model and electronic equipment. According to the method, the local human body features corresponding to the multiple human body regions are obtained by performing region division on the global human body features, the local appearance features are extracted from the local human body features, feature classification is performed on the local appearance features, and the appearance prediction result is obtained; performing model training on the to-be-trained model through the appearance loss function containing the appearance prediction result, and further generating a human body weight recognition model according to the to-be-trained model after model training. The human body sensing ability of the human body weight recognition model under the fine granularity is enhanced, local appearance features are extracted from human body local features, the capturing ability of fine granularity differences such as textures and colors is enhanced through an appearance loss function, and the recognition accuracy of the human body weight recognition model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model training, and particularly to a model training method, system and electronic device for a human re-identification model. Background Art

[0002] The human re-identification (ReID, Person Re-identification) technology realizes the judgment of identity consistency by matching the target features collected by a camera, and has become the core supporting technology of intelligent security and autonomous driving systems. For example, in the autonomous driving scenario, the human re-identification model extracts pedestrian features through an in-vehicle camera or roadside sensors, and combines multi-source data such as spatio-temporal information, which can provide identity consistency judgment for the backend decision-making module, and then optimize the vehicle path planning and obstacle avoidance strategies, improving the safety and traffic efficiency in complex traffic scenarios.

[0003] Currently, the human re-identification model extracts the global human features from image samples and uses the global human features as the human difference representation for model training. However, not only is the difference between different humans at the individual level very small, but the global human features are easily interfered by factors such as human posture, shooting environment, and line-of-sight occlusion, resulting in the problem of low recognition accuracy of the human re-identification model that relies on a single global feature. Summary of the Invention

[0004] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. The summary is not a comprehensive review, nor is it intended to identify key / important elements or delineate the scope of protection of these embodiments, but rather serves as a preface to the following detailed description.

[0005] In view of the above-mentioned disadvantages of the prior art, the present application provides a model training method, system and electronic device for a human re-identification model to improve the recognition accuracy of the human re-identification model.

[0006] The present application provides a model training method for a human re-identification model, including: obtaining a model to be trained, where the model to be trained includes a feature extraction module and an appearance learning module; extracting features from a preset image sample through the feature extraction module to obtain global human features, and performing region division on the global human features to obtain local human features corresponding to multiple human regions; extracting features from the local human features through the appearance learning module to obtain local appearance features corresponding to the local human features, and performing feature classification on the local appearance features to obtain an appearance prediction result; training the model to be trained through an appearance loss function including the appearance prediction result to generate a human re-identification model according to the model to be trained after model training.

[0007] In one embodiment of the present application, the feature extraction module divides the global human features into regions in the following manner: dividing the global human features according to preset feature division parameters to obtain global feature segments corresponding to each human region, where the feature division parameters include region scale and / or division direction; performing global average pooling on the global feature segments to obtain a first sub-feature, and performing max pooling on the global feature segments to obtain a second sub-feature, so as to perform feature connection on the first sub-feature and the second sub-feature to obtain a local connection feature; adjusting the dimension of the local connection feature through a convolutional layer to obtain local human features.

[0008] In one embodiment of the present application, dividing the global human features according to preset feature division parameters to obtain global feature segments corresponding to each human region includes: extracting human key points from the image sample through a preset pose estimation model, and predicting the human body pose according to the distribution of the human key points to obtain current pose parameters, where the current pose parameters include the current orientation of the human body and / or the occluded region of the human body; establishing a human region distribution map including each human region according to the current pose parameters, and constructing a feature division grid that satisfies the feature division parameters according to the human region distribution map; dividing the global human features according to the feature division grid to obtain global feature segments corresponding to each human region.

[0009] In one embodiment of the present application, the appearance learning module sequentially includes: a feature connection unit, configured to extract features from the local human features to obtain feature shards, and perform feature connection on the feature shards using the preset appearance name of the image sample to obtain local appearance features; an appearance encoding unit, configured to perform feature encoding on the local appearance features to obtain appearance weighted features; a classification unit, configured to perform feature classification on the appearance weighted features to obtain an appearance prediction result corresponding to the appearance weighted features.

[0010] In one embodiment of the present application, the appearance encoding unit performs feature encoding on the local appearance features in the following manner: sequentially performing feature processing on the local appearance features through a convolutional layer and a batch normalization layer to obtain a third sub-feature, and performing feature processing on the local appearance features through a max pooling layer to obtain a fourth sub-feature; performing feature fusion according to the third sub-feature and the fourth sub-feature to obtain appearance weighted features; weighting the texture features and color features in the appearance weighted features through a self-attention mechanism.

[0011] In one embodiment of the present application, an image sample is obtained in the following manner: an image acquisition device is used to acquire an image of a target scene, obtaining a scene image, wherein the target scene includes multiple human bodies; different human bodies in the scene image are respectively marked to obtain human body identification tags, and appearance attributes in the scene image are marked to obtain appearance tags; the scene image with the human body identification tags and the appearance tags is used as the image sample.

[0012] In one embodiment of the present application, the feature extraction module extracts features from a preset image sample in the following manner: if the image sample includes image sub-samples corresponding to multiple sample modalities, then feature extraction is respectively performed on each of the image sub-samples according to different image modalities to obtain modality sub-features respectively corresponding to each of the image samples, wherein the sample modalities include at least a part of a visible light modality, an infrared modality, a multi-spectral modality, and a point cloud modality; the feature distributions of the modality sub-features are respectively constrained by a modality alignment loss, and a self-attention mechanism is used to weight the human body contour features between the modality sub-features; the modality sub-features are fused according to a preset modality weight to obtain a human body global feature.

[0013] In one embodiment of the present application, the method further includes: the to-be-trained model further includes a human body identification learning module; label smoothing processing is performed on the human body identification tags corresponding to the image sample, and the human body local features are classified by the human body identification learning module to obtain a human body identification prediction result; a multi-classification loss function is constructed according to the human body identification prediction result and the human body identification tags after label smoothing processing, so as to jointly train the to-be-trained model through the appearance loss function and the multi-classification loss function.

[0014] The present application provides a model training system for a human re-identification model, including: an acquisition module configured to acquire a to-be-trained model, wherein the to-be-trained model includes a feature extraction module and an appearance learning module; a first execution module configured to extract features from a preset image sample through the feature extraction module to obtain a human body global feature, and perform regional division on the human body global feature to obtain human body local features respectively corresponding to multiple human body regions; a second execution module configured to extract features from the human body local features through the appearance learning module to obtain local appearance features corresponding to the human body local features, and perform feature classification on the local appearance features to obtain an appearance prediction result; a training module configured to perform model training on the to-be-trained model through an appearance loss function including the appearance prediction result, so as to generate a human re-identification model according to the to-be-trained model after model training.

[0015] The present application provides an electronic device, including: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above method.

[0016] Advantages of the present application:

[0017] By dividing the global features of the human body into regions, local human body features corresponding to multiple human body regions are obtained, feature extraction is performed from the local human body features to obtain local appearance features corresponding to the local human body features, and feature classification is performed on the local appearance features to obtain an appearance prediction result. Then, a training model is trained using an appearance loss function including the appearance prediction result. Furthermore, a human re-identification model is generated based on the trained training model. In this way, compared with training the model using global human body features, not only does it enhance the human perception ability of the human re-identification model at a fine-grained level by dividing different local human body features from the global human body features, but also, local appearance features are extracted from the local human body features, and the capture ability of fine-grained differences such as texture and color is strengthened through the appearance loss function, thereby improving the sensitivity of the human re-identification model to local human body features and effectively improving the recognition accuracy of the human re-identification model. Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of a model training method for a human re-identification model in an embodiment of the present application;

[0019] Figure 2 It is a schematic structural diagram of an appearance learning module in an embodiment of the present application;

[0020] Figure 3 It is a schematic structural diagram of an appearance encoding unit in an embodiment of the present application;

[0021] Figure 4 It is a schematic flowchart of another model training method for a human re-identification model in an embodiment of the present application;

[0022] Figure 5 It is a schematic structural diagram of a model training system for a human re-identification model in an embodiment of the present application;

[0023] Figure 6 It is a schematic structural diagram of an electronic device in an embodiment of the present invention.

[0024] Reference Signs:

[0025] 501 - Acquisition Module; 502 - First Execution Module; 503 - Second Execution Module; 504 - Training Module;

[0026] 600 - Computer system; 601 - Central processing unit; 602 - Read-only memory; 603 - Random access memory; 604 - Bus; 605 - Input / output interface; 606 - Input section; 607 - Output section; 608 - Storage section; 609 - Communication section; 610 - Driver; 611 - Removable medium. Detailed implementation

[0027] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and sub-samples in the embodiments can be combined with each other.

[0028] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0029] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0030] The terms "first", "second", etc. in the specification, claims, and above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of this application described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.

[0031] Unless otherwise stated, the term "plurality" means two or more.

[0032] In this application, the character " / " means that the objects before and after are in an "or" relationship. For example, A / B means: A or B.

[0033] The term "and / or" is a description of the association relationship of objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B.

[0034] Combined Figure 1 As shown, the present application provides a model training method for a human body re-identification model, including:

[0035] Step S101, obtain a model to be trained;

[0036] Among them, the model to be trained includes a feature extraction module and an appearance learning module;

[0037] Step S102, extract features from a preset image sample through the feature extraction module to obtain global human body features, and perform regional division on the global human body features to obtain local human body features corresponding to multiple human body regions;

[0038] Step S103, extract features from the local human body features through the appearance learning module to obtain local appearance features corresponding to the local human body features, and perform feature classification on the local appearance features to obtain an appearance prediction result;

[0039] Step S104, train the model to be trained through an appearance loss function including the appearance prediction result, so as to generate a human body re-identification model according to the model to be trained after model training.

[0040] Adopting the model training method for the human body re-identification model provided by the present application, by performing regional division on the global human body features, local human body features corresponding to multiple human body regions are obtained, and features are extracted from the local human body features to obtain local appearance features corresponding to the local human body features, and feature classification is performed on the local appearance features to obtain an appearance prediction result, so as to train the model to be trained through an appearance loss function including the appearance prediction result, and then generate a human body re-identification model according to the model to be trained after model training. In this way, compared with training the model through the global human body features, not only the human body perception ability of the human body re-identification model at a fine-grained level is enhanced by dividing different local human body features from the global human body features, but also the ability to capture fine-grained differences such as texture and color is strengthened by extracting local appearance features from the local human body features through the appearance loss function, thereby improving the sensitivity of the human body re-identification model to local human body features and effectively improving the recognition accuracy of the human body re-identification model.

[0041] Optionally, the image sample is obtained in the following manner: Image acquisition is performed on the target scene through an image acquisition device to obtain a scene image, where the target scene includes multiple human bodies; different human bodies in the scene image are respectively marked to obtain human body identification labels, and the appearance attributes in the scene image are marked to obtain appearance labels; the scene image with human body identification labels and appearance labels is used as the image sample.

[0042] In some embodiments, an image acquisition device is used to acquire an image of a target scene, obtaining a scene image, including: using an in-vehicle camera to acquire an image of the target scene, obtaining a scene image stream, and using any image frame in the scene image stream as the scene image.

[0043] In some embodiments, the appearance attributes include color, material, texture, etc. For example, the color of a backpack, the color of a hat, the material of pants, etc., so as to capture local human features.

[0044] Optionally, the feature extraction module extracts features from a preset image sample in the following manner: if the image sample includes image sub-samples corresponding to multiple sample modalities, then feature extraction is performed on each image sub-sample according to different image modalities, obtaining modality sub-features corresponding to each image sample, where the sample modalities include at least a part of visible light modality, infrared modality, multi-spectral modality, and point cloud modality; using the modality alignment loss to respectively constrain the feature distributions of the modality sub-features, and using the self-attention mechanism to weight the human contour features between the modality sub-features; performing feature fusion on the modality sub-features according to a preset modality weight to obtain the global human feature.

[0045] In some embodiments, the modality alignment loss includes Kullback-Leibler Divergence (KL divergence) or contrastive loss, and by constraining the feature distributions of the modality sub-features, the differences between the sample modalities are reduced.

[0046] In some embodiments, by weighting the human contour features, the weights of the common features between different sample modalities are strengthened.

[0047] In some embodiments, the feature extraction module extracts features from a preset image sample in the following manner: if the image sample only includes an image sub-sample corresponding to the visible light modality, then a deep convolutional neural network is trained based on the image sample dataset to obtain a target detection framework; the human feature map is extracted from the image sub-sample corresponding to the visible light modality through the target detection framework to obtain the global human feature.

[0048] In some embodiments, the image sample dataset uses the COCO training set (Common Objects in Context), and the deep convolutional neural network uses Resnet (Residual Network).

[0049] Optionally, the feature extraction module divides the global human features in the following manner: dividing the global human features according to preset feature division parameters to obtain global feature segments corresponding to each human region, where the feature division parameters include region scale and / or division direction; obtaining a first sub-feature by performing global average pooling on the global feature segments, and obtaining a second sub-feature by performing max pooling on the global feature segments, so as to concatenate the first sub-feature and the second sub-feature to obtain a locally connected feature; adjusting the dimension of the locally connected feature through a convolutional layer to obtain local human features.

[0050] In some embodiments, the region scale is used to represent the number of human regions; if the region scale is too large, the human re-identification model will pay more attention to local human features, resulting in the weakening of global human features; if the region scale is too small, the content contained in local human features is relatively large, and it is difficult for the human re-identification model to learn discriminative local features; after multiple experiments, the value range of the region scale is determined to be between 3 and 5. For example, the region scale is set to 4, so as to enhance the discriminative information of the human re-identification model for local human features within a suitable region scale.

[0051] In some embodiments, the division direction is horizontal division, that is, dividing the human body from top to bottom. For example, the global human features are divided into 4 features by horizontal division, and the local human features corresponding to the head region, upper body region, lower body region, and foot region are obtained in sequence.

[0052] In some embodiments, an image sample I is obtained t ; feature extraction is performed from the image sample I t to obtain global human features F t ; the global human features F t are horizontally divided into 4 human regions according to the feature division parameters to obtain global feature segments F i,j , where i is the index of the global human features and j is the index of the human region corresponding to the global feature segment; global average pooling and max pooling are respectively performed on a single global feature segment, and then local connection features G i,j are obtained through feature fusion; the local connection features G i,j are input into the convolutional layer so that the dimension of the local connection features G i,j is reduced to 256 to obtain H i,j , where the local connection features G i,j , in the case of the same i, represent the descriptions of the same human body in different human regions.

[0053] In some embodiments, the local connection features G i,j are obtained through formula (1):

[0054] G i,j = avgpool(F i,j ) + maxpool(F i,j ) Equation (1)

[0055] In Equation (1), avgpool is global average pooling, and maxpool is max pooling.

[0056] In this way, if only global average pooling is used to strengthen the correspondence between features and categories, in the case where the human body background accounts for a large proportion, the background response value is included in the statistics, and the strong response in a small area will be offset by the weak response of the background, resulting in the loss of the discriminability of local features. By introducing max pooling, only the strongest response within the channel is retained, and the low-value regions of the background and noise are ignored. Even if the discriminative features only occupy a very small area, their maximum values can still be captured, thereby increasing the response of discriminative details.

[0057] Optionally, the global human body features are divided according to preset feature division parameters to obtain global feature segments corresponding to each human body region, including: extracting human body key points from the image sample through a preset pose estimation model, and predicting the human body pose according to the distribution of the human body key points to obtain the current pose parameters, where the current pose parameters include the current orientation of the human body and / or the occluded area of the human body; establishing a human body region distribution map including each human body region according to the current pose parameters, and constructing a feature division grid that meets the feature division parameters according to the human body region distribution map; dividing the global human body features according to the feature division grid to obtain global feature segments corresponding to each human body region respectively.

[0058] In some embodiments, the human body key points with confidence levels lower than a preset threshold are removed, and adjacent frame interpolation or mirror filling is used for the low-confidence points; the main torso vector and the human body direction angle are calculated according to the key point coordinates of the human body key points; the human body orientation is determined according to the human body direction angle. For example, if the human body direction angle is between -45° and 45°, it is determined that the human body is facing forward or backward; the occlusion situation of the key points is counted through the visibility flag bit to determine the occluded area of the human body; a human body region distribution map is generated according to the human body orientation and the occluded area of the human body, and the head region, upper body region, lower body region, and foot region in the global human body features are determined according to the human body region distribution map; a feature division grid is established according to the region scale and the division direction to divide the global feature segments corresponding to the head region, upper body region, lower body region, and foot region respectively from the global human body features.

[0059] In this way, compared with only relying on pose estimation for human body part division, this application does not simply rely on the model accuracy of the pose estimation model. Instead, it combines the regional scale and division direction to unify the division rules between different human global features, and uses pose estimation to improve the adaptability of the division rules to human body occlusion and pose changes, thereby improving the robustness of feature division.

[0060] As shown in Figure 2 , the appearance learning module sequentially includes a feature connection unit 201, an appearance encoding unit 202, and a classification unit 203. Among them, the feature connection unit 201 is used to extract features from human body part features to obtain feature shards, and uses the preset appearance names of the image samples to perform feature connection on the feature shards to obtain local appearance features. The appearance encoding unit 202 is used to perform feature encoding on the local appearance features to obtain appearance weighted features. The classification unit 203 is used to perform feature classification on the appearance weighted features to obtain the appearance prediction results corresponding to the appearance weighted features.

[0061] As shown in Figure 3 , the appearance encoding unit encodes the local appearance features in the following manner: sequentially performs feature processing on the local appearance features through a convolutional layer and a batch normalization layer to obtain a third sub-feature, and performs feature processing on the local appearance features through a max pooling layer to obtain a fourth sub-feature; performs feature fusion according to the third sub-feature and the fourth sub-feature to obtain appearance weighted features; weights the texture features and color features in the appearance weighted features through a self-attention mechanism.

[0062] In some embodiments, appearance labels such as backpack color and hat color are used as supervision information to capture local features that are more detailed than human global features; by using the human body part feature H i,j as the input of the appearance learning module, extracts the corresponding feature shards according to each appearance name, and performs connection to obtain the local appearance feature A i,j ; since the occurrence probabilities of different appearance labels in each human body region are different, it is necessary to weight each local appearance feature through the appearance encoding unit and then input it into the classification unit to obtain the appearance prediction result P i,j .

[0063] In some embodiments, the appearance loss function is as shown in formula (2):

[0064]

[0065] In formula (2), L S is the appearance loss, B is the subset length corresponding to the sample set, CE is the cross-entropy loss, S is the appearance label, where the sample set stores multiple image samples.

[0066] In some embodiments, the image sample includes human sample A and human sample B, the postures of the two humans are similar, appearance labels of human sample A and human sample B are respectively marked for the appearance name "backpack color", wherein, the backpack color of human sample A is red, while the backpack color of human sample B is black; local human features are respectively extracted from human sample A and human sample B to obtain the local human feature H corresponding to human sample A 1,j , the local human feature H corresponding to human sample B 2,j ; according to the appearance labels of the image sample, the local human features corresponding to the upper body area are located from all local human features, that is, the local human features corresponding to j = 2, to obtain the local human feature H corresponding to human sample A 1,2 , the local human feature H corresponding to human sample B 2,2 ; feature slices are extracted from the local human feature H 1,2 , and the extracted feature slices are connected according to the appearance name "backpack color" to obtain the local appearance feature A 1,2 corresponding to the local human feature H 1,2 , similarly, the local appearance feature A 2,2 corresponding to the local human feature H a,2 is obtained; by inputting the local appearance feature A 1,2 and the local appearance feature A 2,2 into the appearance encoding unit, the texture features and image features of the backpack area are enhanced through the self-attention mechanism, and at the same time, interference features such as wrinkle features are suppressed; the output features of the appearance encoding unit are input into the classification unit to obtain the appearance prediction result P 1,2 corresponding to human sample A, the appearance prediction result P 2,2 corresponding to human sample B, and the appearance loss function is expressed as L S = CE(P 1,2 , Red) + CE(P 2,2 , Black).

[0067] Optionally, the method further includes: the model to be trained further includes a human identity learning module; the human identity labels corresponding to the image sample are subjected to label smoothing processing, and the local human features are feature-classified through the human identity learning module to obtain human identity prediction results; a multi-classification loss function is constructed according to the human identity prediction results and the human identity labels after label smoothing processing, so as to jointly train the model to be trained through the appearance loss function and the multi-classification loss function.

[0068] In some embodiments, the local human feature H i,j is further feature-classified through the corresponding classifier FC i,j and the Softmax function to obtain the human identity prediction result P' i,j, where the human identity prediction result is determined by formula (3):

[0069]

[0070] In formula (3), P′ i,j is the human identity prediction result corresponding to the human local feature H i,j , N is the total number of human identities, and W i,j is the classifier weight corresponding to the classifier FC i,j , and I t is the image sample.

[0071] In some embodiments, by smoothing the human identity label G i of the image sample, the strong constraint of the person re-identification model on the true label is alleviated, and model overfitting is avoided, thereby improving the generalization ability of the person re-identification model. Among them, the human identity label G i is smoothed by formula (4):

[0072]

[0073] In formula (4), target is the true target label, and ε is a preset control variable. For example, ε = 0.1.

[0074] In some embodiments, the multi-class loss function is as shown in formula (5):

[0075]

[0076] In formula (5), L G is the multi-class loss.

[0077] Combined with Figure 4 as shown, the present application provides a model training for a person re-identification model, including:

[0078] Step S401, obtaining an image sample, a human identity label corresponding to the image sample, and an appearance label corresponding to the image sample;

[0079] Step S402, extracting features from the image sample to obtain a human global feature;

[0080] Step S403, dividing the human global feature into regions to obtain local connection features corresponding to multiple human regions respectively;

[0081] Step S404, adjusting the dimension of the local connection feature to obtain a human local feature, and jumping to step S405 and step S407;

[0082] Step S405: Perform label smoothing on the human body identification label corresponding to the image sample, and classify the local human body features to obtain the human body identification prediction result;

[0083] Step S406: Combine the human body identification label and the human body identification prediction result after label smoothing to establish a multi-classification loss function to train the model to be trained, and jump to step S410;

[0084] Step S407: Extract features from the local human body features to obtain the local appearance features corresponding to the local human body features;

[0085] Step S408: Classify the local appearance features to obtain the appearance prediction result;

[0086] Step S409: Combine the appearance label and the appearance prediction result to establish an appearance loss function to train the model to be trained, and jump to step S410;

[0087] Step S410: Generate a human body re-identification model according to the model to be trained after model training.

[0088] Using the model training method for the human body re-identification model provided by this application, by dividing the global human body features into regions, the local human body features corresponding to multiple human body regions are obtained, and features are extracted from the local human body features to obtain the local appearance features corresponding to the local human body features, and the local appearance features are classified to obtain the appearance prediction result, so as to train the model to be trained through the appearance loss function including the appearance prediction result, and then generate a human body re-identification model according to the model to be trained after model training, which has the following advantages:

[0089] First, compared with training the model through global human body features, not only does it enhance the human body perception ability of the human body re-identification model at a fine-grained level by dividing different local human body features from the global human body features, but also, local appearance features are extracted from the local human body features, and the ability to capture fine-grained differences such as texture and color is strengthened through the appearance loss function, thereby improving the sensitivity of the human body re-identification model to local human body features and effectively improving the recognition accuracy of the human body re-identification model;

[0090] Second, if only global average pooling is used to strengthen the correspondence between features and categories, in the case where the human body background accounts for a large proportion, the background response value is included in the statistics, and the strong response in a small area will be offset by the weak response of the background, resulting in the loss of the discriminability of local features. By introducing max pooling, only the strongest response within the channel is retained, ignoring the low-value regions of the background and noise. Even if the discriminative features only account for a very small area, their maximum value can still be captured, thereby increasing the response of discriminative details;

[0091] Third, compared with only relying on pose estimation for human body part division, this application does not simply rely on the model accuracy of the pose estimation model. Instead, it combines regional scale and division direction to unify the division rules between different human global features, and uses pose estimation to improve the adaptability of the division rules to human occlusion and pose changes, thereby improving the robustness of feature division.

[0092] Combine Figure 5 As shown in the figure, this application provides a model training system for a human re-identification model, including an acquisition module 501, a first execution module 502, a second execution module 503, and a training module 504.

[0093] The acquisition module 501 is configured to acquire a model to be trained, where the model to be trained includes a feature extraction module and an appearance learning module.

[0094] The first execution module 502 is configured to extract features from a preset image sample through the feature extraction module to obtain human global features, and perform regional division on the human global features to obtain human local features corresponding to multiple human regions.

[0095] The second execution module 503 is configured to extract features from the human local features through the appearance learning module to obtain local appearance features corresponding to the human local features, and perform feature classification on the local appearance features to obtain an appearance prediction result.

[0096] The training module 504 is configured to perform model training on the model to be trained through an appearance loss function including the appearance prediction result, so as to generate a human re-identification model according to the model to be trained after model training.

[0097] Using the model training system for a human re-identification model provided by this application, by performing regional division on human global features, human local features corresponding to multiple human regions are obtained, and features are extracted from the human local features to obtain local appearance features corresponding to the human local features, and feature classification is performed on the local appearance features to obtain an appearance prediction result, so as to perform model training on the model to be trained through an appearance loss function including the appearance prediction result, and then generate a human re-identification model according to the model to be trained after model training. In this way, compared with model training through human global features, not only does it enhance the human perception ability of the human re-identification model at a fine-grained level by dividing different human local features from human global features, but also, by extracting local appearance features from human local features and strengthening the capture ability of fine-grained differences such as texture and color through the appearance loss function, the sensitivity of the human re-identification model to human local features is improved, effectively improving the recognition accuracy of the human re-identification model.

[0098] The present application also provides an electronic device, including: a processor and a memory; the memory is used for storing a computer program, and the processor is used for executing the computer program stored in the memory, so that the electronic device executes the above-mentioned method.

[0099] Figure 6 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that, Figure 6 The shown computer system 600 of the electronic device is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0100] As Figure 6 shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 into the random access memory (RAM) 603, such as executing the method in the above-mentioned embodiments. In the RAM 603, various programs and data required for system operation are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0101] The following components are connected to the I / O interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including, such as, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 608 including a hard disk, etc.; and a communication part 609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required, so that the computer program read from it can be installed into the storage part 608 as required.

[0102] The electronic device disclosed in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication therebetween. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program to enable the electronic device to execute each step of the above method.

[0103] The above description and the drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. Embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and sub-samples of some embodiments may be included in or replace parts and sub-samples of other embodiments. Moreover, the terms used in this application are only for describing embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated sub-samples, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other sub-samples, wholes, steps, operations, elements, components, and / or groupings of these. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device comprising the element. In this document, what each embodiment focuses on may be the differences from other embodiments, and the same or similar parts among the various embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.

[0104] Those skilled in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0105] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some sub-samples can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in this application, the various functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A model training method for a human body weight recognition model, characterized in that, Including: Obtain a model to be trained, where the model to be trained includes a feature extraction module and an appearance learning module; Extract features from a preset image sample through the feature extraction module to obtain global human features, and perform region division on the global human features to obtain local human features corresponding to multiple human regions; Extract features from the local human features through the appearance learning module to obtain local appearance features corresponding to the local human features, and perform feature classification on the local appearance features to obtain an appearance prediction result; Train the model to be trained through an appearance loss function including the appearance prediction result, so as to generate a person re-identification model according to the model to be trained after model training.

2. The method according to claim 1, characterized in that, The feature extraction module divides the global human features in the following manner: Divide the global human features according to preset feature division parameters to obtain global feature segments corresponding to each human region, where the feature division parameters include region scale and / or division direction; Perform global average pooling on the global feature segments to obtain a first sub-feature, and perform max pooling on the global feature segments to obtain a second sub-feature, so as to perform feature connection on the first sub-feature and the second sub-feature to obtain a local connection feature; Adjust the dimension of the local connection feature through a convolutional layer to obtain local human features.

3. The method according to claim 2, wherein Dividing the global human features according to preset feature division parameters includes: Extract human key points from the image sample through a preset pose estimation model, and perform human body pose prediction according to the distribution of the human key points to obtain current pose parameters, where the current pose parameters include the current orientation of the human body and / or the occluded area of the human body; Establish a human region distribution map including each human region according to the current pose parameters, and construct a feature division grid that satisfies the feature division parameters according to the human region distribution map; Divide the global human features according to the feature division grid to obtain global feature segments corresponding to each human region.

4. The method according to claim 1, wherein The appearance learning module sequentially includes: A feature connection unit for extracting features from the local human features to obtain feature shards, and performing feature connection on the feature shards using the appearance name preset in the image sample to obtain local appearance features; An appearance encoding unit for performing feature encoding on the local appearance features to obtain appearance weighted features; A classification unit for performing feature classification on the appearance weighted features to obtain an appearance prediction result corresponding to the appearance weighted features.

5. The method according to claim 4, characterized in that The appearance encoding unit encodes the local appearance features in the following manner: Sequentially perform feature processing on the local appearance features through a convolutional layer and a batch normalization layer to obtain a third sub-feature, and perform feature processing on the local appearance features through a max pooling layer to obtain a fourth sub-feature; Perform feature fusion according to the third sub-feature and the fourth sub-feature to obtain appearance weighted features; Weight the texture feature and color feature in the appearance weighted feature through the self-attention mechanism.

6. The method according to any one of claims 1 to 5, characterized in that, Obtain an image sample in the following manner: Collect an image of a target scene through an image acquisition device to obtain a scene image, where the target scene includes multiple human bodies; Mark different human bodies in the scene image respectively to obtain human body identification labels, and mark the appearance attributes in the scene image to obtain appearance labels; Use the scene image with the human body identification label and the appearance label as the image sample.

7. The method according to any one of claims 1 to 5, characterized in that, The feature extraction module extracts features from a preset image sample in the following manner: If the image sample includes image sub-samples corresponding to multiple sample modalities, extract features from each of the image sub-samples according to different image modalities to obtain modality sub-features corresponding to each of the image samples, where the sample modalities include at least a part of a visible light modality, an infrared modality, a multi-spectral modality, and a point cloud modality; Use the modality alignment loss to respectively constrain the feature distributions of the modality sub-features, and use the self-attention mechanism to weight the human body contour features between the modality sub-features; Fuse the modality sub-features according to a preset modality weight to obtain a human body global feature.

8. The method according to any one of claims 1 to 5, characterized in that The method further includes: The to-be-trained model further includes a human body identification learning module; Perform label smoothing processing on the human body identification label corresponding to the image sample, and perform feature classification on the human body local feature through the human body identification learning module to obtain a human body identification prediction result; Construct a multi-classification loss function according to the human body identification prediction result and the human body identification label after label smoothing processing, so as to jointly train the to-be-trained model through the appearance loss function and the multi-classification loss function.

9. A model training system for a human body weight recognition model, characterized in that, Includes: An acquisition module configured to acquire a to-be-trained model, where the to-be-trained model includes a feature extraction module and an appearance learning module; A first execution module configured to extract features from a preset image sample through the feature extraction module to obtain a human body global feature, and perform regional division on the human body global feature to obtain human body local features corresponding to multiple human body regions; A second execution module configured to extract features from the human body local features through the appearance learning module to obtain local appearance features corresponding to the human body local features, and perform feature classification on the local appearance features to obtain an appearance prediction result; A training module configured to perform model training on the to-be-trained model through an appearance loss function including the appearance prediction result, so as to generate a human body re-identification model according to the to-be-trained model after model training.

10. An electronic device, characterized in that, Includes: A processor and a memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 8.