A Cross-Domain Person Re-Identification Method and Device Based on Bidirectional Semantic Alignment Clustering

By constructing a three-branch model and semantic alignment clustering method, the performance degradation caused by the difference in feature distribution between domains in cross-domain pedestrian re-identification is solved, and a more efficient pedestrian recognition effect is achieved.

CN114913476BActive Publication Date: 2025-07-04PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210440638.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-25
Publication Date
2025-07-04
Estimated Expiration
2042-04-25

AI Technical Summary

Technical Problem

In cross-domain applications, existing pedestrian re-identification technology has degraded model performance due to differences in distribution between domains. The existing methods have problems of insufficient usability and insufficient feature distinction capabilities when narrowing the domain feature distribution distance.

Method used

A cross-domain pedestrian recognition method based on bidirectional semantic alignment clustering is adopted. By constructing a three-branch model, features are learned from the upper body, lower body and full body images, and inter-domain and intra-domain processing modules are introduced. The semantic-based learning of feature space is used for dictionary learning and sparse coding framework, and feature learning is optimized by combining cross-entropy loss and triple loss function.

Benefits of technology

It improves the accuracy and efficiency of pedestrian re-identification, reduces attention to background areas, and enhances the generalization ability of the model in the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913476B_ABST
    Figure CN114913476B_ABST
Patent Text Reader

Abstract

This application relates to the technical fields of computer vision and pattern recognition. More specifically, this application relates to a cross-domain pedestrian re-identification method and device based on bidirectional semantic alignment clustering. The method includes: constructing a cross-domain pedestrian recognition model based on at least three branch models; training the cross-domain pedestrian recognition model based on a source domain dataset and a target domain dataset; pedestrian data; wherein, the three branch models are a first branch model, a second branch model, and a third branch model, and each branch model includes a backbone network and a feature expression module. This application enables the cross-domain pedestrian recognition model to perform three-branch feature learning, introduces three human body segmentation constraints respectively, adopts a dictionary learning and sparse coding framework to learn the semantic basis of the original feature space, and uses more reliable source domain semantic elements to more comprehensively measure the similarity between target domain samples, making the cross-domain pedestrian recognition model more accurate when identifying target pedestrians, thereby improving the efficiency of pedestrian re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of computer vision and pattern recognition. More specifically, the present application relates to a cross-domain pedestrian re-identification method and apparatus based on bidirectional semantic alignment clustering. Background Art

[0002] In recent years, with the development of computer vision and pattern recognition technologies, intelligent monitoring technologies are gradually replacing manual retrieval to automatically analyze the content in surveillance videos. Intelligent monitoring technologies can not only significantly improve the efficiency of data processing, but also mine valuable information from a large amount of video image data to improve the quality of monitoring.

[0003] Existing pedestrian re-identification technologies usually adopt a supervised learning method to learn discriminative multi-level feature expressions from a certain labeled data set, and achieve excellent recognition performance on this data set. However, if the model learned well on the labeled data set (source domain) is directly applied to another unlabeled data set (target domain), and if the distribution differences between these two domains are large, it will lead to a significant decline in the model performance. Summary of the Invention

[0004] Based on the above technical problems, the present invention aims to enable a cross-domain pedestrian recognition model to perform three-branch feature learning, and respectively introduce three human body segmentation constraints to guide the three branch networks to learn human-related features from upper body, lower body, and full body human images, and use the trained cross-domain pedestrian recognition model to recognize target pedestrian data.

[0005] The first aspect of the present invention provides a cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering, and the method includes:

[0006] Constructing a cross-domain pedestrian recognition model based on at least three branch models;

[0007] Training the cross-domain pedestrian recognition model based on a source domain data set and a target domain data set;

[0008] Using the trained cross-domain pedestrian recognition model to recognize target pedestrian data;

[0009] Wherein, the three branch models are a first branch model, a second branch model, and a third branch model, and each branch model includes a backbone network and a feature expression module.

[0010] In some embodiments of the present invention, the source domain dataset has labels, the source domain dataset includes an overall source domain dataset, a first local source domain dataset, and a second local source domain dataset, the target domain dataset includes an overall target domain dataset, a first local target domain dataset, and a second local target domain dataset, the cross-domain pedestrian recognition model further includes an inter-domain processing module and an intra-domain processing module, and training the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset includes:

[0011] Input the overall source domain dataset and the overall target domain dataset into the first branch model simultaneously to obtain a first source domain feature and a first target domain feature;

[0012] Input the first local source domain dataset and the first local target domain dataset into the second branch model simultaneously to obtain a second source domain feature and a second target domain feature;

[0013] Input the second local source domain dataset and the second local target domain dataset into the third branch model simultaneously to obtain a third source domain feature and a third target domain feature;

[0014] Input the first source domain feature, the second source domain feature, the third source domain feature, the first target domain feature, the second target domain feature, and the third target domain feature into the inter-domain processing module to obtain an inter-domain processing result;

[0015] Input the first target domain feature, the second target domain feature, and the third target domain feature into the intra-domain processing module to obtain an intra-domain processing result;

[0016] Obtain a pseudo-label of the target domain dataset based on the inter-domain processing result and the intra-domain processing result;

[0017] Train the cross-domain pedestrian recognition model based on the pseudo-label until the training is completed.

[0018] In some embodiments of the present invention, the backbone network included in each branch model includes five stages of convolutional layers, the five stages of convolutional layers are arranged adjacent to each other, and a deconvolution layer is arranged after the fifth stage convolutional layer; the step of inputting the overall source domain dataset and the overall target domain dataset into the first branch model to obtain a first source domain feature and a first target domain feature includes:

[0019] Input the overall source domain dataset and the overall target domain dataset into the backbone network of the first branch model;

[0020] The five stages of convolutional layers perform convolutional processing on the overall source domain dataset and the overall target domain dataset;

[0021] The deconvolution layer performs segmentation constraints on the result of the convolutional processing to obtain a first source domain original feature and a first target domain original feature;

[0022] Input the original features of the first source domain and the first target domain into the feature expression module for feature extraction to obtain the features of the first source domain and the first target domain.

[0023] In some embodiments of the present invention, the inter-domain processing module includes an inter-domain semantic basis module and an inter-domain sparse expression module; the steps of inputting the features of the first source domain, the second source domain, the third source domain, the first target domain, the second target domain, and the third target domain into the inter-domain processing module to obtain an inter-domain processing result include:

[0024] Input the features of the first source domain, the second source domain, and the third source domain into the semantic basis module across domains;

[0025] Input the features of the first target domain, the second target domain, and the third target domain into the inter-domain sparse expression module;

[0026] The semantic basis module performs dictionary learning on the features of the first source domain, the second source domain, and the third source domain to obtain an inter-domain multi-level semantic basis;

[0027] Input the inter-domain multi-level semantic basis into the inter-domain sparse expression module;

[0028] The inter-domain sparse expression module performs sparse coding framework learning according to the features of the first target domain, the second target domain, the third target domain, and the inter-domain multi-level semantic basis to obtain an inter-domain processing result.

[0029] In some embodiments of the present invention, the intra-domain processing module includes an intra-domain semantic basis module and an intra-domain sparse expression module; the steps of inputting the features of the first target domain, the second target domain, and the third target domain into the intra-domain processing module to obtain an intra-domain processing result include:

[0030] Input the features of the first target domain, the second target domain, and the third target domain into the intra-domain semantic basis module;

[0031] The semantic basis module performs dictionary learning on the features of the first target domain, the second target domain, and the third target domain to obtain an intra-domain semantic basis;

[0032] Input the intra-domain semantic basis into the intra-domain sparse expression module;

[0033] The intra-domain sparse expression module performs sparse coding framework learning on the intra-domain semantic basis to obtain an intra-domain processing result.

[0034] In some embodiments of the present invention, the cross-domain pedestrian recognition model further includes a clustering module; the steps of obtaining the pseudo-labels of the target domain dataset based on the inter-domain processing result and the intra-domain processing result include:

[0035] Input the inter-domain processing result and the intra-domain processing result into a clustering module for clustering to obtain multiple clustering nodes;

[0036] Assign a fake label to each clustering node;

[0037] Use all the fake labels as the fake labels of the target domain dataset.

[0038] In some embodiments of the present invention, the cross-domain pedestrian recognition model is optimized by a cross-entropy loss function and a segmentation constraint.

[0039] In some embodiments of the present invention, the segmentation constraint is implemented based on a pixel-level cross-entropy loss function, and the formula of the segmentation constraint is:

[0040]

[0041] where, represents the segmentation constraint of the overall cross-domain pedestrian recognition model, represents the segmentation constraint of the r-th branch model, represents the segmentation constraint of the first branch model, r is a natural number, and λ represents a hyperparameter.

[0042] The second aspect of the present invention provides a cross-domain pedestrian re-identification device based on bidirectional semantic alignment clustering. The device includes:

[0043] A construction module for constructing a cross-domain pedestrian recognition model based on at least three branch models;

[0044] A training module for training the cross-domain pedestrian recognition model based on a source domain dataset and a target domain dataset;

[0045] An identification module for identifying target pedestrian data by using the trained cross-domain pedestrian recognition model;

[0046] Among them, the three branch models are a first branch model, a second branch model, and a third branch model, and each branch model includes a backbone network and a feature expression module.

[0047] The third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to implement the following steps:

[0048] Construct a cross-domain pedestrian recognition model based on at least three branch models;

[0049] Train the cross-domain pedestrian recognition model based on a source domain dataset and a target domain dataset;

[0050] Identify target pedestrian data by using the trained cross-domain pedestrian recognition model;

[0051] Among them, the three branch models are the first branch model, the second branch model, and the third branch model, and each branch model includes a backbone network and a feature expression module.

[0052] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0053] Construct a cross-domain pedestrian recognition model based on at least three branch models;

[0054] Train the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset;

[0055] Use the trained cross-domain pedestrian recognition model to recognize target pedestrian data;

[0056] Among them, the three branch models are the first branch model, the second branch model, and the third branch model, and each branch model includes a backbone network and a feature expression module.

[0057] A fifth aspect of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0058] Construct a cross-domain pedestrian recognition model based on at least three branch models;

[0059] Train the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset;

[0060] Use the trained cross-domain pedestrian recognition model to recognize target pedestrian data;

[0061] Among them, the three branch models are the first branch model, the second branch model, and the third branch model, and each branch model includes a backbone network and a feature expression module.

[0062] The technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0063] This application proposes to construct a cross-domain pedestrian recognition model based on at least three branch models, train the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset, and use the trained cross-domain pedestrian recognition model to recognize the target pedestrian data. Among them, the three branch models are the first branch model, the second branch model, and the third branch model. Each branch model includes a backbone network and a feature expression module, which realizes the accurate recognition of target pedestrians, thereby improving the efficiency of pedestrian re-identification. Specifically, in the three-branch feature learning network, three human segmentation constraints are respectively introduced to guide the three branch networks to learn human-related features from upper body, lower body, and full body human images, and reduce the attention to the background area. Moreover, the use of segmentation constraints does not increase the computational complexity in the test stage. This application proposes to learn multi-level semantic bases (global and local) from the original feature space of the source domain, and then project the target domain samples into the global subspace and local subspace composed of multi-level semantic bases respectively, so as to comprehensively measure the similarity between target domain samples in each subspace. In particular, the dictionary learning and sparse coding framework is used to learn the semantic bases of the original feature space, and more reliable source domain semantic elements are used to more comprehensively measure the similarity between target domain samples, making the cross-domain pedestrian recognition model more accurate when recognizing target pedestrians, thereby improving the efficiency of pedestrian re-identification.

[0064] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Brief Description of the Drawings

[0065] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0066] Figure 1 Shows a schematic diagram of the steps of a cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering in an exemplary embodiment of this application;

[0067] Figure 2 Shows a schematic diagram of the working process of a cross-domain pedestrian recognition model in an exemplary embodiment of this application;

[0068] Figure 3 Shows a schematic diagram of the structure of a cross-domain pedestrian recognition model in an exemplary embodiment of this application;

[0069] Figure 4 Shows a schematic diagram of the structure of a cross-domain pedestrian re-identification device based on bidirectional semantic alignment clustering in an exemplary embodiment of this application;

[0070] Figure 5 The figure shows a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application;

[0071] Figure 6 The figure shows a schematic diagram of a storage medium provided by an exemplary embodiment of the present application. Detailed implementation manners

[0072] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application. It is obvious to those skilled in the art that the present application can be implemented without one or more of these details. In other examples, some technical features well-known to those skilled in the art are not described to avoid confusion with the present application.

[0073] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or combinations thereof.

[0074] Now, exemplary embodiments according to the present application will be described in more detail with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many different forms and should not be construed as being limited only to the embodiments set forth herein. The drawings are not drawn to scale, where some details may be enlarged for the purpose of clear expression and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary. In practice, there may be deviations due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0075] The following combines the specification appendices Figures 1-6 Several embodiments are given to describe the exemplary embodiments according to the present application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0076] In recent years, with the development of computer vision and pattern recognition technologies, intelligent monitoring technology is gradually replacing manual retrieval to automatically analyze the content in surveillance videos. Intelligent monitoring analyzes the scenes in the areas monitored by each camera, extracts target and event information, such as the identity of the target, the characteristics of the target, the interaction between targets, and the events occurring in the scene. In addition, by integrating information from multiple cameras, intelligent monitoring technology can more comprehensively predict the behavior of targets and the development trend of events, so that when an abnormal event occurs, an alarm signal can be quickly sent to law enforcement agencies. Intelligent monitoring technology can not only significantly improve the efficiency of data processing, but also mine valuable information from a large amount of video image data, improving the quality of monitoring.

[0077] In intelligent monitoring technology, accurate object search is a highly challenging visual task, which aims to accurately search for all instances of a given object from a large amount of video image data. Person re-identification is one of the core research directions of accurate object search, and its goal is to find all instances of a specified person from video images across cameras. Person re-identification helps to achieve cross-camera pedestrian tracking and image retrieval.

[0078] Existing pedestrian re-identification technologies usually adopt the method of supervised learning to learn discriminative multi-level feature representations from a certain labeled dataset and achieve excellent recognition performance on this dataset. However, if the model learned well on the labeled dataset (source domain) is directly applied to another unlabeled dataset (target domain), and the distribution differences between these two domains are large, it will lead to a significant decline in the model performance. The goal of cross-domain pedestrian re-identification technology is to learn discriminative features from the labeled source domain and then adaptively transfer the features of the source domain to the target domain, so as to perform pedestrian re-identification in the target domain. A commonly used solution is to adopt the unsupervised domain adaptation method (Unsupervised Domain Adaptation, UDA), which learns features from the labeled source domain through a convolutional neural network and then transfers the features of the source domain to the unlabeled target domain, thereby improving the generalization ability of the model in the target domain. Its main strategy is to guide the network to learn domain-invariant features by minimizing the distance between the feature distributions of the source domain and the target domain during the feature learning process. If the distance between the source domain and the target domain distributions in the pedestrian re-identification task is directly constrained, it may introduce the problem of negative transfer. To solve this problem, some indirect methods have been proposed to narrow the distance between the distributions. According to the different types of constraints, the existing pedestrian re-identification methods based on unsupervised domain adaptation can be roughly divided into two categories. The first category of methods narrows the distance between the feature distributions of the source domain and the target domain by constructing a shared space between pedestrians in different domains or removing some variables that cause feature distribution differences. The second category of methods is based on an assumption that the differences between different domain distributions mainly come from low-level visual differences. This category of methods uses image style transfer technology to transfer the samples of the source domain to the target domain, so that the content of the samples remains unchanged while the style is the same as that of the target domain, thereby narrowing the visual differences between the samples of the two domains.

[0079] However, there are many problems in the above technical solutions: Indirectly narrowing the distance between the feature distributions of the two domains through middle-level semantic features has low usability in practical applications. On the one hand, it requires greater costs to manually label more attribute tags, and due to the limited number of attribute categories, the discriminative ability of the learned features is insufficient; on the other hand, due to the lack of real annotation information, the reliability of the automatically constructed latent attribute space is insufficient. Moreover, the method based on style transfer can only narrow the low-level visual differences between the two domains, while usually, high-level semantic differences, such as perspective differences, pose differences, and differences in feature types between pedestrians in the two domains, are the main factors affecting the distribution.

[0080] Therefore, in some exemplary embodiments of the present application, a cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering is provided, as Figure 1 shown, the method includes:

[0081] S1. Construct a cross - domain pedestrian recognition model based on at least three branch models;

[0082] S2. Train the cross - domain pedestrian recognition model based on the source - domain dataset and the target - domain dataset;

[0083] S3. Use the trained cross - domain pedestrian recognition model to recognize target pedestrian data;

[0084] Among them, the three branch models are the first branch model, the second branch model, and the third branch model, and each branch model includes a backbone network and a feature expression module.

[0085] In a specific implementation, the source - domain dataset has labels. The source - domain dataset includes the overall source - domain dataset, the first local source - domain dataset, and the second local source - domain dataset. The target - domain dataset includes the overall target - domain dataset, the first local target - domain dataset, and the second local target - domain dataset. As Figure 2 shown, the upper pictures in each of the three groups of inputs represent source - domain data, and the lower pictures in each group represent target - domain data. The cross - domain pedestrian recognition model also includes an inter - domain processing module and an intra - domain processing module. Training the cross - domain pedestrian recognition model based on the source - domain dataset and the target - domain dataset includes: simultaneously inputting the overall source - domain dataset and the overall target - domain dataset into the first branch model to obtain the first source - domain feature and the first target - domain feature; simultaneously inputting the first local source - domain dataset and the first local target - domain dataset into the second branch model to obtain the second source - domain feature and the second target - domain feature; simultaneously inputting the second local source - domain dataset and the second local target - domain dataset into the third branch model to obtain the third source - domain feature and the third target - domain feature; inputting the first source - domain feature, the second source - domain feature, the third source - domain feature, the first target - domain feature, the second target - domain feature, and the third target - domain feature into the inter - domain processing module to obtain an inter - domain processing result; inputting the first target - domain feature, the second target - domain feature, and the third target - domain feature into the intra - domain processing module to obtain an intra - domain processing result; obtaining a false label of the target - domain dataset based on the inter - domain processing result and the intra - domain processing result; training the cross - domain pedestrian recognition model based on the false label until it is trained well.

[0086] The backbone network included in each branch model includes five stages of convolutional layers, as Figure 2As shown, five convolutional layers of different stages are arranged adjacent to each other, and a deconvolution layer is arranged after the convolutional layer of the fifth stage. Inputting the overall source domain dataset and the overall target domain dataset into the first branch model simultaneously to obtain the first source domain features and the first target domain features includes: inputting the overall source domain dataset and the overall target domain dataset into the backbone network of the first branch model simultaneously; performing convolution processing on the overall source domain dataset and the overall target domain dataset by the five convolutional layers of different stages; performing segmentation constraint on the result of the convolution processing by the deconvolution layer to obtain the first source domain raw features and the first target domain raw features; inputting the first source domain raw features and the first target domain raw features into the feature expression module for feature extraction to obtain the first source domain features and the first target domain features.

[0087] Optionally, a pre-trained ResNet50 on the ImageNet dataset can be used as the basic unit, and the above three branch models construct a feature learning network with three branches. For pedestrian images, the overall source domain dataset, the first local source domain dataset, and the second local source domain dataset can respectively represent the full-body image of a pedestrian, the upper-body image of a pedestrian, and the lower-body image of a pedestrian. They all belong to the source domain data. Similarly, the overall target domain dataset, the first local target domain dataset, and the second local target domain dataset also correspond to the full-body image of a pedestrian, the upper-body image of a pedestrian, and the lower-body image of a pedestrian in the target domain. The source domain is data with labels, and the target domain is data collected by a camera in an actual application scenario without labels to facilitate model learning. Here, can be represented as the feature corresponding to the r-th region of the i-th image in the source domain, and can be represented as the feature of the r-th region of the j-th image in the target domain. The above-mentioned deconvolution layer performs segmentation constraint on the result of the convolution processing because the convolutional neural network treats all regions in the image equally due to its inability to effectively focus. If the image contains a complex background, it will affect the distinctiveness of the learned features. In addition, for the cross-domain pedestrian re-identification task, pedestrian images in different domains usually contain different backgrounds. If the model learns too many background-related features, it will lead to a decrease in the transferability of the source domain features. Specifically, in this application, a deconvolution layer with a stride of 2 and a kernel size of 3×3, and a convolutional layer with a stride of 1 and a kernel size of 1×1 are connected behind the convolutional layer of the fifth stage in each branch network to predict the human body part to which each pixel point in the input image belongs. Among them, the role of the deconvolution is to perform upsampling on the feature map, and the role of the 1×1 convolutional layer is to classify each pixel. The present invention uses an existing human parsing algorithm to preprocess the images in the source domain and the target domain, and uses the output segmentation result as the label.

[0088] In some embodiments of the application, the cross-domain pedestrian recognition model is optimized by a cross-entropy loss function and a segmentation constraint, that is, an independent cross-entropy classification loss function is adopted to optimize the feature learning of each branch network. Then, the calculation method of the overall cross-entropy classification loss function is represented by formula (1):

[0089]

[0090] where λ is a hyperparameter used to control the relative importance between the classification loss functions of the global and local branch networks.

[0091] The segmentation constraint is implemented based on a pixel-level cross-entropy loss function, and the segmentation constraint is represented by formula (2):

[0092]

[0093] where represents the overall segmentation constraint of the cross-domain pedestrian recognition model, represents the segmentation constraint of the r-th branch model, represents the segmentation constraint of the first branch model, r is a natural number, and λ represents a hyperparameter.

[0094] In other words, the application combines the classification loss function of the source domain and the segmentation constraints of the two domains to jointly optimize the feature learning of the cross-domain pedestrian re-identification basic network. The overall loss function can be represented by formula (3):

[0095]

[0096] It should be noted that the segmentation constraint introduced in this application is equivalent to a regularization term in a sense. Therefore, it only needs to be calculated during the training stage of the network. In the testing stage, the network layer used to predict the segmentation result is removed in the present invention, and multi-level features are extracted from the pedestrian image through the three trained branch networks for pedestrian re-identification. Therefore, the use of the segmentation constraint does not increase the computational complexity in the testing stage.

[0097] The apparent features of a pedestrian can be sparsely represented by a set of high-level semantic elements, where each semantic element represents a potential attribute. Therefore, to project a pedestrian image from a high-dimensional original feature space to a low-dimensional sparse representation subspace, the key lies in finding a set of semantic bases, and this set of semantic bases needs to have the following two properties: First, the semantic bases can completely describe the original high-dimensional features (i.e., the original feature space can be reconstructed through different linear combinations of the semantic bases); Second, the high-dimensional features can be described by as few semantic base elements as possible. Therefore, in some embodiments of the present application, the inter-domain processing module includes an inter-domain semantic base module and an inter-domain sparse representation module; inputting the first source domain feature, the second source domain feature, the third source domain feature, the first target domain feature, the second target domain feature, and the third target domain feature into the inter-domain processing module, the inter-domain processing result is obtained, including: inputting the first source domain feature, the second source domain feature, and the third source domain feature into the inter-domain semantic base module; as Figure 2 shown, inputting the first target domain feature, the second target domain feature, and the third target domain feature into the inter-domain sparse representation module; the semantic base module performs dictionary learning on the first source domain feature, the second source domain feature, and the third source domain feature to obtain an inter-domain multi-level semantic base; inputting the inter-domain multi-level semantic base into the inter-domain sparse representation module; the inter-domain sparse representation module performs sparse coding framework learning according to the first target domain feature, the second target domain feature, the third target domain feature, and the inter-domain multi-level semantic base to obtain an inter-domain processing result. The intra-domain processing module includes an intra-domain semantic base module and an intra-domain sparse representation module; inputting the first target domain feature, the second target domain feature, and the third target domain feature into the intra-domain processing module, the intra-domain processing result is obtained, including: inputting the first target domain feature, the second target domain feature, and the third target domain feature into the intra-domain semantic base module; the semantic base module performs dictionary learning on the first target domain feature, the second target domain feature, and the third target domain feature to obtain an intra-domain semantic base; inputting the intra-domain semantic base into the intra-domain sparse representation module; the inter-domain sparse representation module performs sparse coding framework learning on the intra-domain semantic base to obtain an intra-domain processing result.

[0098] Specifically, when implementing, the target domain feature matrix is represented as W = [w1,..., w M ∈ R l×M , where w i (i = 1,..., M) represents the feature of the i-th image, and l represents the dimension of w i . The dictionary matrix is represented as A = [a1,..., a g ∈ R l×g , where a j (j = 1,..., g) represents a dictionary item. The coefficient matrix is represented as C = [c1,..., c M ∈ R g×M , where c i (i = 1,..., M) represents wi Sparse representation. To ensure that w i can be obtained as a linear combination of the terms in the dictionary, this application introduces a reconstruction loss function in the dictionary learning process for minimizing the reconstruction error, where ||·||2 represents the L2 norm. Dictionary learning and sparse coding correspond to the objective function of formula (4):

[0099]

[0100] where ||·|| F represents the Frobenius norm of the matrix. ||·||1 represents the L1 norm, which is used to constrain the sparsity of c i so as to ensure that the original feature vector is reconstructed with as few dictionary terms as possible. γ is a hyperparameter used to control the relative importance between the reconstruction loss and the sparse constraint term.

[0101] If both variables A and C are optimized simultaneously, then this optimization problem is non-convex. If either variable A or variable C in the formula is optimized separately, then this optimization problem is convex. Therefore, this application adopts an alternating optimization scheme, first fixing one variable and then optimizing the other variable, and alternating in turn. Thus, the optimization problem defined by formula (4) can be decomposed into two sub-problems, including the least squares problem with L1 regularization and the least squares problem with L2 constraint. Among them, each sub-problem can be solved by existing optimization methods.

[0102] While learning the dictionary A, the l-dimensional original feature vector w of each target domain sample i (i = 1,..., M) is also projected into a lower-dimensional (g-dimensional) sparse representation subspace, and c i is its sparse representation in the subspace. Then, the similarity between the i-th sample and the j-th sample in the target domain can be obtained by calculating the cosine distance between their sparse representations, as expressed by formula (5):

[0103]

[0104] In some embodiments of this application, the cross-domain pedestrian recognition model further includes a clustering module; obtaining the pseudo-labels of the target domain dataset based on the inter-domain processing results and the intra-domain processing results, including: inputting the inter-domain processing results and the intra-domain processing results into the clustering module for clustering to obtain multiple clustering nodes; assigning a pseudo-label to each clustering node; and using all the pseudo-labels as the pseudo-labels of the target domain dataset.

[0105] Intra-domain subspace learning uses the semantic elements learned from the original feature space of the target domain to measure the similarity between samples. However, since only false labels are included, the reliability of the semantic elements learned from the target domain is lower than that of the source domain. On the other hand, there are global or local appearance similarities between pedestrians in the source domain and the target domain. From the perspective of feature expression, pedestrians in different domains share multi-level semantic elements. Based on the above observations, in order to use more reliable source domain semantic elements to more comprehensively measure the similarity between target domain samples, this application learns multi-level semantic bases (global and local) from the original feature space of the source domain, and then projects the target domain samples into the global subspace and local subspace composed of multi-level semantic bases respectively, so as to comprehensively measure the similarity between target domain samples in each subspace.

[0106] Specifically, the present invention represents multiple source domain feature matrices for learning multi-level semantic bases as and where, represents the feature of the i-th sample in the source domain extracted by the r-th (r = 0, 1, 2) branch network in Figure 1 , and o represents the dimension of . At the same time, the dictionary matrix of the source domain is represented as The coefficient matrix of the global dictionary is represented as The coefficient matrix of the local dictionary is represented as It should be noted that since the feature vectors output by the global branch network may also contain features of some local regions, in order to better express the global features of the target domain, this application stacks the global and local feature vectors of the source domain together to form V 0 , for learning more complete semantic bases. Target domain dictionary learning corresponds to the objective function of formula (6):

[0107]

[0108] where γ is set to be the same as γ in formula (4), and each dictionary D r (r = 0, 1, 2) is learned independently.

[0109] After learning the dictionary D r , the present invention calculates the sparse representation of the target domain feature Specifically, it is represented by formula (7):

[0110]

[0111] Therefore, the o-dimensional original feature vector is projected into a sparse representation subspace of a lower dimension (h - dimension), which is the sparse representation in the subspace. Then, by comprehensively calculating the cosine distance between the multi - level sparse representations of the i - th and j - th samples in the target domain, the similarity between them can be obtained, as shown in formula (8):

[0112]

[0113] Furthermore, by comprehensively considering the similarities of the target - domain samples in the intra - domain subspace and the inter - domain subspace, the similarity between the i - th and j - th samples in the target domain can be calculated by equation (9):

[0114]

[0115] where β is a hyper - parameter used to control the relative importance between the similarities of the target - domain samples in the intra - domain subspace and the inter - domain subspace.

[0116] After obtaining the final similarity, the present application uses the DBSCAN algorithm to cluster the target - domain samples, and then automatically assigns a unique fake label to each clustering node. Then, during the training process, the present invention uses these fake labels and adopts a triplet loss function to constrain the relative distance relationship between the target - domain samples, thereby optimizing the feature learning network. Specifically, the algorithm randomly selects P pedestrians from the training set of the target domain according to the fake labels, and then randomly selects K samples from the samples corresponding to each pedestrian, so as to form a sample set of size P×K as the input of the network. Similar to the cross - entropy loss function, the present invention optimizes the feature learning of each branch network through an independent triplet loss function, specifically shown in formula (10):

[0117]

[0118] where, represents the feature of the k - th sample of the i - th pedestrian in the input image extracted by the r - th (r = 0, 1, 2) branch network in Figure 1 , represents the squared Euclidean distance, and α is a hyper - parameter used to control the margin between positive and negative samples, preferably set to 0.5.

[0119] The calculation of the overall triplet loss function is shown in formula (11):

[0120]

[0121] where λ is set to be the same as λ in formula (1).

[0122] It can be seen that the present application uses the cross - entropy loss function Segmentation constraint and the triplet loss function to jointly optimize the feature learning of the three branch networks. The overall loss function can be expressed by Equation (12):

[0123]

[0124] It should be noted that the above branch network and branch model are essentially the same, but only expressed from different perspectives in form. From the perspective of the model, it is called a branch model, and from the perspective of the working process, it is expressed as a branch network. As Figure 3 shown, the cross-domain pedestrian recognition model includes three branch models, an inter-domain processing module, an intra-domain processing module, and a clustering module. Each branch model includes a backbone network and a feature expression module. Among them, the backbone network includes convolutional layers and deconvolutional layers in five stages, and a deconvolutional layer is arranged after the convolutional layer of the fifth stage to be used for segmentation constraint. The inter-domain processing module includes an inter-domain semantic basis module and an inter-domain sparse expression module, and the intra-domain processing module includes an intra-domain semantic basis module and an intra-domain sparse expression module. Of course, the number of branch networks or branch models is not limited to three, and can be multiple greater than three. Other training steps of the cross-domain pedestrian recognition model are based on the existing technologies in this field. During training, the model is optimized according to the above loss function of the present application. After reaching the preset number of training times, the cross-domain pedestrian recognition model is considered to be trained well.

[0125] In the feature learning network with three branches of the present application, three human body segmentation constraints are respectively introduced to guide the three branch networks to learn human-related features from upper body, lower body, and full body human images, and reduce the attention to the background area. Moreover, the use of segmentation constraints will not increase the computational complexity in the test stage. The present application proposes to learn multi-level semantic bases (global and local) from the original feature space of the source domain, and then project the target domain samples into the global subspace and local subspace composed of multi-level semantic bases respectively, so as to comprehensively measure the similarity between target domain samples in each subspace. In particular, a dictionary learning and sparse coding framework is adopted to learn the semantic bases of the original feature space, and more reliable source domain semantic elements are used to more comprehensively measure the similarity between target domain samples, making the cross-domain pedestrian recognition model more accurate when identifying target pedestrians, and thus improving the efficiency of pedestrian re-identification.

[0126] In some exemplary embodiments of the present application, a cross-domain pedestrian re-identification device based on bidirectional semantic alignment clustering is further provided, as Figure 4 shown. The device includes:[[]]

[0127] A construction module 401, configured to construct a cross-domain pedestrian recognition model based on at least three branch models;

[0128] A training module 402, configured to train the cross-domain pedestrian recognition model based on a source domain dataset and a target domain dataset;

[0129] An identification module 403, configured to identify target pedestrian data by using the trained cross-domain pedestrian recognition model;

[0130] Wherein, the three branch models are a first branch model, a second branch model and a third branch model, and each branch model includes a backbone network and a feature expression module.

[0131] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention.

[0132] It should also be emphasized that the system provided in the embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theory, method, technology and application system. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0133] Please refer to the following Figure 5 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. As Figure 5 shown, the electronic device 2 includes: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering provided in any of the foregoing embodiments of the present application.

[0134] Wherein, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 203 (which can be wired or wireless), a communication connection between the system network element and at least one other network element is realized, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0135] The bus 202 can be an ISA bus, a PCI bus, an EISA bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program. Any implementation manner of the cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering disclosed in any implementation manner of the embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.

[0136] The processor 200 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 200 or the instructions in the form of software. The above-mentioned processor 200 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.

[0137] The embodiment of the present application also provides a computer-readable storage medium corresponding to the cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering provided in the foregoing embodiment. Please refer to Figure 6 , Figure 6 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the cross-domain pedestrian re-identification method provided in any of the foregoing embodiments.

[0138] In addition, examples of the computer-readable storage medium may further include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0139] The computer-readable storage medium provided by the above embodiments of the present application and the method for allocating quantum key distribution channels in a space-division multiplexing optical network provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0140] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor implements the steps of the cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering provided by any of the foregoing embodiments, including: constructing a cross-domain pedestrian recognition model based on at least three branch models; training the cross-domain pedestrian recognition model based on a source domain data set and a target domain data set; using the trained cross-domain pedestrian recognition model to recognize target pedestrian data; wherein the three branch models are a first branch model, a second branch model, and a third branch model, and each branch model includes a backbone network and a feature expression module.

[0141] It should be noted that: The algorithms and displays provided here are not inherently related to any specific computer, virtual device, or other equipment. Various general-purpose devices can also be used in conjunction with the teachings provided here. Based on the above description, the structure required to construct such a device is obvious. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described here can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation mode of the present application. In the specification provided here, a large number of specific details are described. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0142] Similarly, it should be understood that, for the purpose of streamlining this application and facilitating the understanding of one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of this application, the various features of this application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim stands on its own as a separate embodiment of this application.

[0143] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0144] The various component embodiments of this application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of this application. This application can also be implemented as a device or device program for executing part or all of the methods described herein. The program implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0145] The above is only the preferred specific embodiment of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering, characterized in that The method includes: Constructing a cross-domain pedestrian recognition model based on at least three branch models; Training the cross-domain pedestrian recognition model based on a source domain dataset and a target domain dataset; Using the trained cross-domain pedestrian recognition model to recognize target pedestrian data; Wherein, the three branch models are a first branch model, a second branch model, and a third branch model, and each branch model includes a backbone network and a feature expression module; The source domain dataset includes an overall source domain dataset, a first local source domain dataset, and a second local source domain dataset, the target domain dataset includes an overall target domain dataset, a first local target domain dataset, and a second local target domain dataset, and the cross-domain pedestrian recognition model further includes an inter-domain processing module; the inter-domain processing module includes an inter-domain semantic basis module and an inter-domain sparse expression module; the training of the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset includes: Simultaneously inputting the overall source domain dataset and the overall target domain dataset into the first branch model to obtain a first source domain feature and a first target domain feature; Simultaneously inputting the first local source domain dataset and the first local target domain dataset into the second branch model to obtain a second source domain feature and a second target domain feature; Simultaneously inputting the second local source domain dataset and the second local target domain dataset into the third branch model to obtain a third source domain feature and a third target domain feature; Inputting the first source domain feature, the second source domain feature, and the third source domain feature into the inter-domain semantic basis module; Inputting the first target domain feature, the second target domain feature, and the third target domain feature into the inter-domain sparse expression module; The inter-domain semantic basis module performs dictionary learning on the first source domain feature, the second source domain feature, and the third source domain feature to obtain an inter-domain multi-level semantic basis; Inputting the inter-domain multi-level semantic basis into the inter-domain sparse expression module; The inter-domain sparse expression module performs sparse coding framework learning according to the first target domain feature, the second target domain feature, the third target domain feature, and the inter-domain multi-level semantic basis to obtain an inter-domain processing result.

2. The cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering according to claim 1, wherein The source domain dataset has labels, and the cross-domain pedestrian recognition model further includes an intra-domain processing module. The training of the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset further includes: Inputting the first target domain feature, the second target domain feature, and the third target domain feature into the intra-domain processing module to obtain an intra-domain processing result; Obtaining a pseudo-label of the target domain dataset based on the inter-domain processing result and the intra-domain processing result; Training the cross-domain pedestrian recognition model based on the pseudo-label until it is trained well.

3. The cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering according to claim 2, wherein The backbone network included in each branch model includes five stages of convolutional layers, the five stages of convolutional layers are arranged adjacent to each other, and a deconvolution layer is arranged after the fifth stage convolutional layer adjacent to each other; the step of simultaneously inputting the overall source domain dataset and the overall target domain dataset into the first branch model to obtain a first source domain feature and a first target domain feature includes: Simultaneously inputting the overall source domain dataset and the overall target domain dataset into the backbone network of the first branch model; The convolutional layers of the five stages perform convolutional processing on the overall source domain dataset and the overall target domain dataset; The deconvolution layer performs segmentation constraints on the results of the convolutional processing to obtain the first source domain original features and the first target domain original features; The first source domain original features and the first target domain original features are input into the feature expression module for feature extraction to obtain the first source domain features and the first target domain features.

4. The cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering according to claim 1, wherein, The intra-domain processing module includes an intra-domain semantic basis module and an intra-domain sparse expression module; The first target domain features, the second target domain features, and the third target domain features are input into the intra-domain processing module to obtain the intra-domain processing results, including: The first target domain features, the second target domain features, and the third target domain features are input into the intra-domain semantic basis module; The intra-domain semantic basis module performs dictionary learning on the first target domain features, the second target domain features, and the third target domain features to obtain the intra-domain semantic basis; The intra-domain semantic basis is input into the intra-domain sparse expression module; The intra-domain sparse expression module performs sparse coding framework learning on the intra-domain semantic basis to obtain the intra-domain processing results.

5. The cross-domain person re-identification method based on bidirectional semantic alignment clustering according to claim 4, wherein, The cross-domain pedestrian recognition model further includes a clustering module; Obtaining the pseudo-labels of the target domain dataset based on the inter-domain processing results and the intra-domain processing results, including: The inter-domain processing results and the intra-domain processing results are input into the clustering module for clustering to obtain a plurality of clustering nodes; Assigning a pseudo-label to each clustering node; All the pseudo-labels are used as the pseudo-labels of the target domain dataset.

6. The cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering according to any one of claims 1-5, characterized in that, The cross-domain pedestrian recognition model is optimized by a cross-entropy loss function and segmentation constraints.

7. The cross-domain pedestrian re-identification method based on bidirectional semantic alignment clustering according to claim 6, wherein The segmentation constraint is implemented based on a pixel-level cross-entropy loss function, and the formula of the segmentation constraint is: Among them, represents the segmentation constraint of the overall cross-domain pedestrian recognition model, represents the segmentation constraint of the r-th branch model, represents the segmentation constraint of the first branch model, r is a natural number, and λ represents a hyperparameter.

8. A cross-domain pedestrian re-identification device based on bidirectional semantic alignment clustering, characterized in that, The device includes: A construction module for constructing a cross-domain pedestrian recognition model based on at least three branch models; A training module for training the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset; An identification module for identifying target pedestrian data by using the trained cross-domain pedestrian recognition model; Wherein, the three branch models are a first branch model, a second branch model, and a third branch model, and each branch model includes a backbone network and a feature expression module; The source domain dataset includes an overall source domain dataset, a first local source domain dataset, and a second local source domain dataset, the target domain dataset includes an overall target domain dataset, a first local target domain dataset, and a second local target domain dataset, and the cross-domain pedestrian recognition model further includes an inter-domain processing module; the inter-domain processing module includes an inter-domain semantic basis module and an inter-domain sparse expression module; training the cross-domain pedestrian recognition model based on the source domain dataset and the target domain dataset includes: The overall source domain dataset and the overall target domain dataset are simultaneously input into the first branch model to obtain the first source domain features and the first target domain features; The first local source domain dataset and the first local target domain dataset are simultaneously input into the second branch model to obtain the second source domain features and the second target domain features; The second local source domain dataset and the second local target domain dataset are simultaneously input into the third branch model to obtain the third source domain features and the third target domain features; Input the first source domain feature, the second source domain feature, and the third source domain feature into the inter-domain semantic basis module; Input the first target domain feature, the second target domain feature, and the third target domain feature into the inter-domain sparse representation module; The inter-domain semantic basis module performs dictionary learning on the first source domain feature, the second source domain feature, and the third source domain feature to obtain an inter-domain multi-level semantic basis; Input the inter-domain multi-level semantic basis into the inter-domain sparse representation module; The inter-domain sparse representation module performs sparse coding framework learning according to the first target domain feature, the second target domain feature, the third target domain feature, and the inter-domain multi-level semantic basis to obtain an inter-domain processing result.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.